如何解释MCE消息?


10

我注意到最近在/var/log/messages我们的一台服务器上出现了很多错误(如下)。但是,与syslog中的解码条目相比,mce客户端似乎不太确定错误源。是否有某种键可以用来解释MCE输出?

Nov 12 04:19:19 areion kernel: [14698753.176035] Machine check events logged
Nov 12 04:19:19 areion mcelog: HARDWARE ERROR. This is *NOT* a software problem!
Nov 12 04:19:19 areion mcelog: Please contact your hardware vendor
Nov 12 04:19:19 areion mcelog: MCE 0
Nov 12 04:19:19 areion mcelog: CPU 0 BANK 8
Nov 12 04:19:19 areion mcelog: MISC 640738dd0009159c ADDR 96236c6c0
Nov 12 04:19:19 areion mcelog: TIME 1352711959 Mon Nov 12 04:19:19 2012
Nov 12 04:19:19 areion mcelog: MCG status:
Nov 12 04:19:19 areion mcelog: MCi status:
Nov 12 04:19:19 areion mcelog: MCi_MISC register valid
Nov 12 04:19:19 areion mcelog: MCi_ADDR register valid
Nov 12 04:19:19 areion mcelog: MCA: MEMORY CONTROLLER RD_CHANNELunspecified_ERR
Nov 12 04:19:19 areion mcelog: Transaction: Memory read error
Nov 12 04:19:19 areion mcelog: STATUS 8c0000400001009f MCGSTATUS 0
Nov 12 04:19:19 areion mcelog: MCGCAP 1c09 APICID 20 SOCKETID 1
Nov 12 04:19:19 areion mcelog: CPUID Vendor Intel Family 6 Model 44

所有错误似乎都与相同的存储库有关:

areion:~# awk -F'mcelog:' '/mcelog:.*BANK/{ print $2; }' < /var/log/messages |uniq
 CPU 0 BANK 8 

我正在运行mcelog守护程序,当我检查错误信息时,它似乎不知道错误从何而来。只有它们与之关联CPU0(此框中只有一个CPU):

Memory errors
SOCKET 1 CHANNEL any DIMM any
corrected memory errors:
        77 total
        77 in 24h
uncorrected memory errors:
        0 total
        0 in 24h
Per page corrected memory statistics:
359ffc000: total 2 2 in 24h online

3b93cc000: total 2 2 in 24h online

3ce45c000: total 2 2 in 24h online

96236c000: total 20 20 in 24h online triggered

96545c000: total 9 9 in 24h online

96a82c000: total 9 9 in 24h online

96a8ec000: total 1 1 in 24h online

96fb6c000: total 15 15 in 24h online triggered

9c2edc000: total 15 15 in 24h online triggered

9c5eac000: total 1 1 in 24h online

9c6a1c000: total 1 1 in 24h online

我如何解释这些信息还不清楚。一方面,mce客户端不指示通道或DIMM,但是解码后的消息表明DIMM 8上发生了错误。dmesg似乎表明只记录了42条消息:

[14698753.176035] Machine check events logged
[14698753.629174] Machine check events logged
[14698815.338595] __ratelimit: 38 callbacks suppressed
[14698815.338628] Machine check events logged
[14698816.020797] Machine check events logged

我似乎收到了各种信息,这使我想知道根据来自各种来源的信息所做出的假设。

其他信息:

areion:~# grep 'model name' /proc/cpuinfo |uniq
model name      : Intel(R) Xeon(R) CPU           X5670  @ 2.93GHz

areion:~# apt-cache policy mcelog |grep Installed
  Installed: 1.0~pre3-3

areion:~# lsb_release -a
No LSB modules are available.
Distributor ID: Debian
Description:    Debian GNU/Linux 6.0.6 (squeeze)
Release:        6.0.6
Codename:       squeeze

Answers:


2

您可能想尝试更换有问题的DIMM(CPU 0,SOCKET 8),并查看是否继续生成MCE消息。

mcelog软件包配置了一些随时间推移发生的各种MCE事件的默认阈值。查看/etc/mcelog/mcelog.conf详细信息。对于内存页面错误,阈值是24小时内的10个事件。(我不确定该数字来自何处,但这可能是一个合理的参考点)。您的帖子提到24页整整24小时内有77个可纠正的事件,因此DIMM很可能出现了问题,可能会或可能不会变得更严重。

对于从不同来源收到不一致的信息,我不会感到沮丧。总的来说,我发现固件级别的任何东西都是特定于平台的(即特定于该特定硬件型号)。对于固件相关问题,我的经验法则是供应商工具通常最准确,但使用最少。更通用的开源工具更易于使用,但可能无法提供足够的信息来准确显示正在发生的事情。

By using our site, you acknowledge that you have read and understand our Cookie Policy and Privacy Policy.
Licensed under cc by-sa 3.0 with attribution required.