在主机上检测到 SMART 错误 (CurrentPendingSector)

在主机上检测到 SMART 错误 (CurrentPendingSector)

Smart 每天都会在我的一个数据中心服务器 SSD 上报告以下错误:

This message was generated by the smartd daemon running on:

   host name:  x
   DNS domain: x

The following warning/error was logged by the smartd daemon:

Device: /dev/sda [SAT], 1 Currently unreadable (pending) sectors

Device info:
INTEL SSDSC2BB480G7, S/N:PHDV701605XT480BGN, WWN:5-5cd2e4-14d60bdf9, FW:N2010101, 480 GB

For details see host's SYSLOG.

You can also use the smartctl utility for further investigation.
The original message about this issue was sent at Tue Feb  6 18:43:41 2018 EST
Another message will be sent in 24 hours if the problem persists.

Syslog,不显示任何附加信息。

我在其他地方读到过,我需要执行扩展智能测试,然后写入测试结果中指定的扇区以强制驱动器将其标记为坏扇区并重新分配。

我进行了一次长时间的测试:

sudo smartctl -t long /dev/sda1

但是 sudo smartctl -a /dev/sda1 的输出表明测试完成并且没有错误?

smartctl 6.6 2016-05-31 r4324 [x86_64-linux-4.4.98-2-pve] (local build)
Copyright (C) 2002-16, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Device Model:     INTEL SSDSC2BB480G7
Serial Number:    PHDV701605XT480BGN
LU WWN Device Id: 5 5cd2e4 14d60bdf9
Firmware Version: N2010101
User Capacity:    480,103,981,056 bytes [480 GB]
Sector Sizes:     512 bytes logical, 4096 bytes physical
Rotation Rate:    Solid State Device
Form Factor:      2.5 inches
Device is:        Not in smartctl database [for details use: -P showall]
ATA Version is:   ACS-3 T13/2161-D revision 5
SATA Version is:  SATA 3.1, 6.0 Gb/s (current: 6.0 Gb/s)
Local Time is:    Wed Feb 21 04:58:20 2018 EST
SMART support is: Available - device has SMART capability.
SMART support is: Enabled

=== START OF READ SMART DATA SECTION ===
SMART overall-health self-assessment test result: PASSED

General SMART Values:
Offline data collection status:  (0x02) Offline data collection activity
                                        was completed without error.
                                        Auto Offline Data Collection: Disabled.
Self-test execution status:      (   0) The previous self-test routine completed
                                        without error or no self-test has ever
                                        been run.
Total time to complete Offline
data collection:                (   72) seconds.
Offline data collection
capabilities:                    (0x79) SMART execute Offline immediate.
                                        No Auto Offline data collection support.
                                        Suspend Offline collection upon new
                                        command.
                                        Offline surface scan supported.
                                        Self-test supported.
                                        Conveyance Self-test supported.
                                        Selective Self-test supported.
SMART capabilities:            (0x0003) Saves SMART data before entering
                                        power-saving mode.
                                        Supports SMART auto save timer.
Error logging capability:        (0x01) Error logging supported.
                                        General Purpose Logging supported.
Short self-test routine
recommended polling time:        (   1) minutes.
Extended self-test routine
recommended polling time:        (   2) minutes.
Conveyance self-test routine
recommended polling time:        (   2) minutes.
SCT capabilities:              (0x003d) SCT Status supported.
                                        SCT Error Recovery Control supported.
                                        SCT Feature Control supported.
                                        SCT Data Table supported.

SMART Attributes Data Structure revision number: 1
Vendor Specific SMART Attributes with Thresholds:
ID# ATTRIBUTE_NAME          FLAG     VALUE WORST THRESH TYPE      UPDATED  WHEN_FAILED RAW_VALUE
  5 Reallocated_Sector_Ct   0x0032   100   100   000    Old_age   Always       -       0
  9 Power_On_Hours          0x0032   100   100   000    Old_age   Always       -       8753
 12 Power_Cycle_Count       0x0032   100   100   000    Old_age   Always       -       11
170 Unknown_Attribute       0x0033   100   100   010    Pre-fail  Always       -       0
171 Unknown_Attribute       0x0032   100   100   000    Old_age   Always       -       0
172 Unknown_Attribute       0x0032   100   100   000    Old_age   Always       -       0
174 Unknown_Attribute       0x0032   100   100   000    Old_age   Always       -       9
175 Program_Fail_Count_Chip 0x0033   100   100   010    Pre-fail  Always       -       270807480494
183 Runtime_Bad_Block       0x0032   100   100   000    Old_age   Always       -       0
184 End-to-End_Error        0x0033   100   100   090    Pre-fail  Always       -       0
187 Reported_Uncorrect      0x0032   100   100   000    Old_age   Always       -       0
190 Airflow_Temperature_Cel 0x0022   070   064   000    Old_age   Always       -       30 (Min/Max 16/37)
192 Power-Off_Retract_Count 0x0032   100   100   000    Old_age   Always       -       9
194 Temperature_Celsius     0x0022   100   100   000    Old_age   Always       -       30
197 Current_Pending_Sector  0x0012   100   100   000    Old_age   Always       -       1
199 UDMA_CRC_Error_Count    0x003e   100   100   000    Old_age   Always       -       0
225 Unknown_SSD_Attribute   0x0032   100   100   000    Old_age   Always       -       415161
226 Unknown_SSD_Attribute   0x0032   100   100   000    Old_age   Always       -       1126
227 Unknown_SSD_Attribute   0x0032   100   100   000    Old_age   Always       -       51
228 Power-off_Retract_Count 0x0032   100   100   000    Old_age   Always       -       525179
232 Available_Reservd_Space 0x0033   100   100   010    Pre-fail  Always       -       0
233 Media_Wearout_Indicator 0x0032   099   099   000    Old_age   Always       -       0
234 Unknown_Attribute       0x0032   100   100   000    Old_age   Always       -       0
241 Total_LBAs_Written      0x0032   100   100   000    Old_age   Always       -       415161
242 Total_LBAs_Read         0x0032   100   100   000    Old_age   Always       -       433149
243 Unknown_Attribute       0x0032   100   100   000    Old_age   Always       -       979238

SMART Error Log Version: 1
No Errors Logged

SMART Self-test log structure revision number 1
Num  Test_Description    Status                  Remaining  LifeTime(hours)  LBA_of_first_error
# 1  Extended offline    Completed without error       00%      8753         -
# 2  Extended offline    Completed without error       00%      8752         -
# 3  Extended offline    Completed without error       00%      8464         -
# 4  Short offline       Completed without error       00%       888         -
# 5  Short offline       Completed without error       00%       886         -
# 6  Short offline       Completed without error       00%       885         -
# 7  Short offline       Completed without error       00%         0         -
# 8  Short offline       Completed without error       00%         0         -
# 9  Short offline       Completed without error       00%         0         -

SMART Selective self-test log data structure revision number 1
 SPAN  MIN_LBA  MAX_LBA  CURRENT_TEST_STATUS
    1        0        0  Not_testing
    2        0        0  Not_testing
    3        0        0  Not_testing
    4        0        0  Not_testing
    5        0        0  Not_testing
Selective self-test flags (0x0):
  After scanning selected spans, do NOT read-scan remainder of disk.
If Selective self-test is pending on power-up, resume after 0 minute delay.

真的有问题吗?如果是,该如何正确诊断?

答案1

对于驱动器来说,只有 1 个待处理扇区不是问题。只需忽略它,直到它显著增加。

虽然您可以尝试使用以下说明来修复它(首先备份数据!):https://github.com/smartmontools/smartmontools/blob/master/www/BadBlockHowTo.txt

您还可以将您的驾驶 SMART 报告与其他人进行比较这个 SMART 报告库

答案2

最大的问题是该扇区是否有您关心的数据,请尝试读取您的数据文件并查看是否缺少某些内容。

无论如何,为了避免此错误消息重复出现,您需要覆盖数据,以便让驱动器忽略此消息。如果块没有其他用途,修剪块也应该有效。您可以尝试在文件系统上执行 fstrim 来清除此消息。

SMART 长测试对媒体进行随机测试,而不是完整测试,因此即使出现错误也可能会成功,特别是如果这是一种可恢复的错误,因为制造商不希望单个待处理的错误导致您对驱动器进行 RMA。

相关内容