当前待处理部门和离线的原因不可纠正的增加,然后减less到零

在一周的时间里,对于3TB希捷硬盘(ST3000DM001-1CH166),smartd报告的离线不可纠正和当前不可读(待定)扇区的数量在缓慢增加,然后是递减的数字,直到最终计数为0,并且错误条件重启。 从日志(只显示更改):

Jul 6 18:04:57 x smartd[462]: Device: /dev/sdb [SAT], 8 Currently unreadable (pending) sectors Jul 6 18:04:58 x smartd[462]: Device: /dev/sdb [SAT], 8 Offline uncorrectable sectors [...] Jul 7 16:34:58 x smartd[462]: Device: /dev/sdb [SAT], 16 Currently unreadable (pending) sectors (changed +8) Jul 7 16:34:58 x smartd[462]: Device: /dev/sdb [SAT], 16 Offline uncorrectable sectors (changed +8) [...] Jul 11 14:04:57 x smartd[462]: Device: /dev/sdb [SAT], 24 Currently unreadable (pending) sectors (changed +8) Jul 11 14:04:57 x smartd[462]: Device: /dev/sdb [SAT], 24 Offline uncorrectable sectors (changed +8) Jul 11 14:34:57 x smartd[462]: Device: /dev/sdb [SAT], 32 Currently unreadable (pending) sectors (changed +8) Jul 11 14:34:58 x smartd[462]: Device: /dev/sdb [SAT], 32 Offline uncorrectable sectors (changed +8) [...] Jul 13 09:04:57 x smartd[462]: Device: /dev/sdb [SAT], 24 Currently unreadable (pending) sectors (changed -8) Jul 13 09:04:57 x smartd[462]: Device: /dev/sdb [SAT], 24 Offline uncorrectable sectors (changed -8) Jul 13 09:34:58 x smartd[462]: Device: /dev/sdb [SAT], 16 Currently unreadable (pending) sectors (changed -8) Jul 13 09:34:58 x smartd[462]: Device: /dev/sdb [SAT], 16 Offline uncorrectable sectors (changed -8) Jul 13 10:04:57 x smartd[462]: Device: /dev/sdb [SAT], 16 Currently unreadable (pending) sectors Jul 13 10:04:57 x smartd[462]: Device: /dev/sdb [SAT], 16 Offline uncorrectable sectors Jul 13 10:34:57 x smartd[462]: Device: /dev/sdb [SAT], No more Currently unreadable (pending) sectors, warning condition reset after 1 email Jul 13 10:34:57 x smartd[462]: Device: /dev/sdb [SAT], No more Offline uncorrectable sectors, warning condition reset after 1 email 

此外,重新分配的扇区数也是0,所以这些扇区似乎没有被重新映射。 以下是驱动器的完整(当前) smartctl -a输出:

 smartctl 6.2 2013-07-26 r3841 [x86_64-linux-3.14.4-100.fc19.x86_64] (local build) Copyright (C) 2002-13, Bruce Allen, Christian Franke, www.smartmontools.org === START OF INFORMATION SECTION === Model Family: Seagate Barracuda 7200.14 (AF) Device Model: ST3000DM001-1CH166 Serial Number: W1F30FK2 LU WWN Device Id: 5 000c50 06129a9a8 Firmware Version: CC27 User Capacity: 3,000,592,982,016 bytes [3.00 TB] Sector Sizes: 512 bytes logical, 4096 bytes physical Rotation Rate: 7200 rpm Device is: In smartctl database [for details use: -P show] ATA Version is: ACS-2, ACS-3 T13/2161-D revision 3b SATA Version is: SATA 3.1, 6.0 Gb/s (current: 3.0 Gb/s) Local Time is: Wed Jul 16 11:23:08 2014 EDT ==> WARNING: A firmware update for this drive may be available, see the following Seagate web pages: http://knowledge.seagate.com/articles/en_US/FAQ/207931en http://knowledge.seagate.com/articles/en_US/FAQ/223651en SMART support is: Available - device has SMART capability. SMART support is: Enabled === START OF READ SMART DATA SECTION === SMART overall-health self-assessment test result: PASSED General SMART Values: Offline data collection status: (0x00) Offline data collection activity was never started. Auto Offline Data Collection: Disabled. Self-test execution status: ( 0) The previous self-test routine completed without error or no self-test has ever been run. Total time to complete Offline data collection: ( 584) seconds. Offline data collection capabilities: (0x73) SMART execute Offline immediate. Auto Offline data collection on/off support. Suspend Offline collection upon new command. No Offline surface scan supported. Self-test supported. Conveyance Self-test supported. Selective Self-test supported. SMART capabilities: (0x0003) Saves SMART data before entering power-saving mode. Supports SMART auto save timer. Error logging capability: (0x01) Error logging supported. General Purpose Logging supported. Short self-test routine recommended polling time: ( 1) minutes. Extended self-test routine recommended polling time: ( 347) minutes. Conveyance self-test routine recommended polling time: ( 2) minutes. SCT capabilities: (0x3085) SCT Status supported. SMART Attributes Data Structure revision number: 10 Vendor Specific SMART Attributes with Thresholds: ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE UPDATED WHEN_FAILED RAW_VALUE 1 Raw_Read_Error_Rate 0x000f 115 099 006 Pre-fail Always - 91131424 3 Spin_Up_Time 0x0003 094 094 000 Pre-fail Always - 0 4 Start_Stop_Count 0x0032 100 100 020 Old_age Always - 12 5 Reallocated_Sector_Ct 0x0033 100 100 010 Pre-fail Always - 0 7 Seek_Error_Rate 0x000f 078 060 030 Pre-fail Always - 59138260 9 Power_On_Hours 0x0032 094 094 000 Old_age Always - 5888 10 Spin_Retry_Count 0x0013 100 100 097 Pre-fail Always - 0 12 Power_Cycle_Count 0x0032 100 100 020 Old_age Always - 12 183 Runtime_Bad_Block 0x0032 100 100 000 Old_age Always - 0 184 End-to-End_Error 0x0032 100 100 099 Old_age Always - 0 187 Reported_Uncorrect 0x0032 100 100 000 Old_age Always - 0 188 Command_Timeout 0x0032 100 100 000 Old_age Always - 0 0 0 189 High_Fly_Writes 0x003a 097 097 000 Old_age Always - 3 190 Airflow_Temperature_Cel 0x0022 053 049 045 Old_age Always - 47 (Min/Max 23/51) 191 G-Sense_Error_Rate 0x0032 100 100 000 Old_age Always - 0 192 Power-Off_Retract_Count 0x0032 100 100 000 Old_age Always - 10 193 Load_Cycle_Count 0x0032 096 096 000 Old_age Always - 8086 194 Temperature_Celsius 0x0022 047 051 000 Old_age Always - 47 (0 22 0 0 0) 197 Current_Pending_Sector 0x0012 100 100 000 Old_age Always - 0 198 Offline_Uncorrectable 0x0010 100 100 000 Old_age Offline - 0 199 UDMA_CRC_Error_Count 0x003e 200 200 000 Old_age Always - 0 240 Head_Flying_Hours 0x0000 100 253 000 Old_age Offline - 5703h+56m+25.808s 241 Total_LBAs_Written 0x0000 100 253 000 Old_age Offline - 11838196191 242 Total_LBAs_Read 0x0000 100 253 000 Old_age Offline - 211237637103 SMART Error Log Version: 1 No Errors Logged SMART Self-test log structure revision number 1 No self-tests have been logged. [To run self-tests, use: smartctl -t] SMART Selective self-test log data structure revision number 1 SPAN MIN_LBA MAX_LBA CURRENT_TEST_STATUS 1 0 0 Not_testing 2 0 0 Not_testing 3 0 0 Not_testing 4 0 0 Not_testing 5 0 0 Not_testing Selective self-test flags (0x0): After scanning selected spans, do NOT read-scan remainder of disk. If Selective self-test is pending on power-up, resume after 0 minute delay. 

我已经下载,但尚未在驱动器上运行Seatools,但驱动器的当前SMART状态基本上看起来不错。 什么可能导致这样的行为?

更新:对于未来的读者来说,这个驱动器实际上还可以工作几个月,此时更多的扇区变得不可纠正/脱机,并且重新分配扇区的数量开始增加到0以上,并且SMART长时间自检开始失败错误。 所以这些信息似乎是一个有用的预警。

我最近几次使用类似的希捷ST3000DM001-1CH166固件CC24。 你不需要Seatools,只需要运行一个长期的智能testing:

 smartctl -t long /dev/sdc 

那么,如果smarctl显示没有错误,你现在可以:

 SMART Self-test log structure revision number 1 Num Test_Description Status Remaining LifeTime(hours) LBA_of_first_error # 1 Extended offline Completed without error 00% 9912 - 

如果失败,则直接发回希捷。 他们在一个星期内取代了我最后一个失败的人。

当前等待扇区就是这样的,磁盘知道的位置数量需要重新分配,但还没有重新分配,因为磁盘没有数据源被重新分配。 一旦写入该位置,磁盘将自动将该区域重新分配到另一个位置,并将新数据写入新位置,并且当前待定扇区的和弦将减less。

这完全正常,磁盘应该如何操作。

您可以使用Linux上的diskscan或Windows上的HD Tune来扫描磁盘以查找错误的位置,并尝试通过将软件写入磁盘来尝试“立即”重新分配位置。