Communication failure determination method and apparatus
By parsing register error information in the BMC and using OOBMSM to determine the MCTP communication link configuration, the problem of MCTP communication fault location in servers is solved, enabling fast and accurate fault type identification and reducing operation and maintenance costs.
Patent Information
- Application Number
- CN202511212221.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-08-28
AI Technical Summary
Existing technologies struggle to quickly and accurately determine the type of server fault, especially VGA display anomalies and system freezes caused by MCTP communication failures, resulting in high maintenance costs and low efficiency.
By receiving register error information from the central processing unit in the Baseboard Management Controller (BMC), parsing routing identification information, determining the type of communication fault, and using the Out-of-Band Management Service Module (OOBMSM) and MCTP communication link configuration information, the MCTP communication error can be accurately determined.
Quickly identify MCTP communication errors, reduce misjudgments, save manpower and resources, improve fault location efficiency, and reduce operation and maintenance costs.
Smart Images

Figure CN120723526B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of communication technology and fault management, and more particularly, to a communication fault determination method and device. BACKGROUND
[0002] As the core equipment for data storage, processing and transmission, servers need to have high concurrent processing capability, continuous stable operation, powerful I / O throughput performance and elastic expansion capability, and are the key infrastructure supporting emerging technologies such as cloud computing, big data and artificial intelligence.
[0003] With the deep expansion of server application scenarios, its reliability has become one of the core evaluation indicators. Server failure can cause key business interruption, and even trigger data loss and other problems, so it is crucial to quickly determine the fault and efficiently repair it to ensure business continuity. However, since a single abnormal phenomenon may correspond to multiple potential fault sources (such as hardware compatibility, network link interruption, software configuration error, etc.), the related fault determination method relies more on post-log grabbing and manual backtracking analysis, which not only has poor timeliness, but also consumes a large amount of computing resources and operation and maintenance manpower. SUMMARY
[0004] In view of the above problems, the present application provides a communication fault determination method and device, and further provides a storage medium and a program product.
[0005] According to one aspect of the present application, a communication fault determination method is provided, applied to a baseboard management controller, comprising: in response to receiving register error information from a central processor, analyzing the register error information to obtain routing identification information, the routing identification information being used to locate an error triggering device; wherein the register error information is obtained from a register related to the baseboard management controller based on a target error event by the central processor; in the case that the routing identification information indicates an out-of-band management service module, determining that the register error information is target communication type information; wherein the baseboard management controller, the central processor and the out-of-band management service module are connected via a communication link of the target communication type; determining a communication fault type according to configuration information of the communication link, and sending the communication fault type to the central processor.
[0006] According to another aspect of the present application, a communication fault determination method is provided, applied to a central processor, comprising: in response to a target interrupt event being triggered, obtaining error event information; analyzing the error event information to determine an error event type; in the case that the error event type is a target error event, obtaining register error information in a target register based on a basic input / output system, and sending the register error information to a baseboard management controller.
[0007] Another aspect of the present application provides an electronic device, comprising: a baseboard management controller, one or more processors; a memory for storing one or more computer programs, wherein the baseboard management controller and the one or more processors execute the one or more computer programs to implement the steps of the method.
[0008] Another aspect of the present application also provides a computer readable storage medium having stored thereon a computer program or instructions, which when executed by a processor implement the steps of the method.
[0009] Another aspect of the present application also provides a computer program product comprising a computer program or instructions, which when executed by a processor implement the steps of the method. BRIEF DESCRIPTION OF DRAWINGS
[0010] The above and other objects, features and advantages of the present application will become more apparent from the following description when taken in conjunction with the accompanying drawings, in which:
[0011] Figure 1 An application scenario diagram of the communication fault determination method, device, medium and program product according to an embodiment of the present application is shown.
[0012] Figure 2 A flowchart of the communication fault determination method according to an embodiment of the present application is shown.
[0013] Figure 3 A flowchart of the communication fault determination method according to another embodiment of the present application is shown.
[0014] Figure 4 A communication fault determination flowchart according to an embodiment of the present application is shown.
[0015] Figure 5 A block diagram of an electronic device suitable for implementing the communication fault determination method according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0016] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. It is to be understood, however, that the description is merely exemplary and is not intended to limit the scope of the present application. In the following detailed description of the embodiments of the present application, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that one or more embodiments of the present application can be practiced without these specific details. In other instances, well-known structures and functions have not been described in detail in order to avoid obscuring aspects of the present application.
[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used herein, the terms "comprises", "comprising", "includes", "including" and the like are, where used, meant to be inclusive, but not limiting in any way.
[0018] All terms used herein, including technical and scientific terms, have the meanings commonly understood by one of ordinary skill in the art unless otherwise defined herein. It should be noted that the use of terms herein, such as should be construed in accordance with the meaning that is consistent with the context of the present description, and should not be interpreted in an idealized or overly formal manner.
[0019] In the case of using expressions similar to "at least one of A, B, and C, etc.", in general, it should be interpreted as having the meaning of including at least one of A, B, and C, etc. (for example, "a system having at least one of A, B, and C" should include but not be limited to a system having A alone, a system having B alone, a system having C alone, a system having both A and B, a system having both A and C, a system having both B and C, and / or a system having A, B, and C, etc.).
[0020] In the technical solutions of the present application, the user information (including but not limited to user personal information, user image information, user equipment information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved are information and data authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data comply with relevant laws, regulations and standards, necessary security measures are taken, do not violate public order and good customs, and provide corresponding operation portal for user to choose authorization or refusal.
[0021] The embodiments of the present application relate to the fields of computer technology and fault management technology, for example, can be applied to the field of server fault management technology. For the convenience of understanding, part of the technical terms mentioned in this paper are introduced and explained as follows:
[0022] CPU (Central Processing Unit, Central Processing Unit), as the operation and control core of computer system, is the final execution unit of information processing and program running.
[0023] BMC (Baseboard Management Controller) and IPMI (Intelligent Platform Management Interface) are the basic core function subsystems of the server, responsible for the core functions of server hardware state management, operating system management, health state management, power consumption management, etc. IPMI is an independent component on the motherboard that can run independently of the host system CPU, BIOS (Basic Input / Output System) and OS (Operating System), and its core component is BMC. BMC interacts with other components such as BIOS, CPU, etc., all through IPMI.
[0024] BIOS is a boot firmware stored in the motherboard flash memory, responsible for Power On Self Test (POST), hardware initialization and booting the operating system.
[0025] OOB (Out-of-Band Management) is a way to remotely manage servers, even if the server's operating system crashes, the network is unavailable or the power is off, the administrator can still access and control the server. Out-of-band management is achieved through a management interface independent of the operating system (such as BMC, IPMI, etc.) to remotely manage the server. It is usually used for maintenance, diagnosis and repair of servers without physical contact.
[0026] OOBMSM (Out of Band Management Service Module) is used to implement out-of-band management functions independent of the operating system. In one embodiment, OOBMSM can exist as a PCIe device, directly interacting with BMC, providing CPU information query, configuration and fault diagnosis capabilities.
[0027] PCIe (Peripheral Component Interconnect express) is a high-speed serial bus interface used to connect various external devices of a computer, such as graphics cards, network cards and storage controllers, etc. The computer can use this interface to expand one or more peripheral devices. Exemplarily, a PCIe device can be connected to the motherboard through a PCIe slot to expand the computer's functions, such as increasing graphics processing capability, network connection capability, storage expansion capability, etc.
[0028] MCTP (Management Component Transport Protocol) is a management and control transport protocol used for communication and control between devices. It defines a set of communication specifications that enable devices to perform effective management and control operations, including device discovery, configuration, monitoring, and fault handling. MCTP protocol can run on different transport layers such as PCIe, Ethernet, etc. Devices send MCTP messages to target devices through the transport layer and receive MCTP messages from other devices. MCTP protocol uses MCTP messages for communication, each of which contains control and management information between devices. MCTP messages are composed of multiple fields, including message type, message length, source device identification, and target device identification. The basic transmission unit of MCTP messages is MCTP packet, and one MCTP message transmission may contain one or more MCTP packets. Endpoint on MCTP link is used to receive MCTP message packets and handle MCTP control instructions. MCTP uses Endpoint ID (EID) as a logical address for MCTP packet transmission and routing between endpoints.
[0029] In one embodiment, MCTP protocol can be used for communication between management controller (BMC) and management controller, as well as communication between management controller and management device. For example, BMC can use MCTP protocol to access management devices through sending and receiving MCTP format messages on multiple different bus types (such as PCIe).
[0030] MCTP errors usually occur in hardware devices such as servers, workstations, or high-end motherboards, and are related to communication with system hardware management components such as baseboard management controller (BMC), out-of-band management functions of processors, etc. Such errors may be caused by hardware failure, firmware / driver problem, or system configuration anomaly.
[0031] In server hardware architecture, RP register (Root Port Register) refers to the configuration register of PCIe root port (Root Port), which is used to control and manage the communication link between CPU and PCIe devices (such as BMC).
[0032] In PCIe architecture, BDF (Bus:Device:Function) is a three-tuple used to uniquely identify each PCIe device or its function in the system. BDF (Bus number + Device number + Function number) constitutes the identity card number of each PCIe device node, which determines the unique target device.
[0033] In the server startup process, the display output is usually realized by the VGA (Video Graphic Array) interface of the BMC. The CPU needs to establish a communication link with the BMC through the MCTP protocol, dynamically transmit display data to the frame buffer of the BMC, and finally output to the display device through the VGA interface. However, if the MCTP communication link between the CPU and the BMC fails, the BMC will not be able to obtain the correct display data frame, resulting in abnormal pictures (such as black screen, static image or garbled code) output by the VGA interface. This exception is manifested as a "server stuck" state at the user end, but since there can be many reasons for the machine to be stuck, it is often difficult to quickly and accurately determine that the MCTP communication fault has occurred, and it is easy to be misjudged as a hardware fault or system startup failure.
[0034] In addition, the related art mainly relies on log analysis to handle the problem of "server stuck", and if effective information cannot be collected at the first time of the problem, the operation and maintenance personnel need to spend a lot of time and effort to reproduce the problem at that time after the fault occurs, collect BMC logs, MCTP protocol packets and hardware register information, and then locate the communication link abnormal point by cross-comparing massive data, which takes several hours for a single troubleshooting, significantly increasing the operation and maintenance cost.
[0035] Therefore, the embodiments of the present application provide a communication fault determination method, comprising: in response to receiving register error information from a central processor, analyzing the register error information to obtain routing identification information, the routing identification information being used to locate an error triggering device; wherein the register error information is obtained by the central processor from a register related to a baseboard management controller based on a target error event; in the case that the routing identification information indicates an out-of-band management service module, determining that the register error information is target communication type information; wherein the baseboard management controller, the central processor and the out-of-band management service module are connected via a communication link of a target communication type; determining a communication fault type according to configuration information of the communication link, and sending the communication fault type to the central processor.
[0036] According to an embodiment of the present application, when the BMC reports an ACS error to the CPU, the BIOS can automatically obtain the register error information of the RP end of the BMC and send the register error information to the BMC. The BMC can parse the received TLP message, determine the Target ID based on the TLP Header. In the case that the Target ID indicates an OOB MSM, it can be determined that the TLP message is MCTP information, so it can be determined that an MCTP communication error has occurred. Further, the configuration information related to the MCTP communication link where the BMC is located can be obtained to determine the specific fault type / reason.
[0037] It can be understood that the communication fault determination method (hereinafter referred to as the present method) provided by the embodiments of the present application can automatically collect the register error information related to the BMC and report it to the BMC through the BIOS at the first time when the BMC reports an ACS error. The BMC can parse the Target ID based on the TLP Header. In the case that the Target ID indicates an OOB MSM (MCTP bus owner), the communication link between the OOB MSM, the CPU and the BMC can be used to exclude non-MCTP interference factors, determine that an MCTP communication error has occurred, and then accurately determine the fault type according to the MCTP link configuration information. For the fault symptoms such as VGA output abnormality and system "stuck", the present method can quickly distinguish whether it is caused by MCTP communication error, so as to reduce misjudgment, save a lot of manpower and material resources, and help to solve the problem of fault type positioning behind abnormal symptoms.
[0038] The embodiments of the present application provide a method for determining MCTP communication abnormality between the CPU and the BMC, which can be used to solve the problem that the fault type / reason cannot be efficiently positioned when the MCTP error causes the VGA to fail to display during the server startup process.
[0039] Figure 1 An application scenario diagram of the communication fault determination method, device, medium and program product according to the embodiments of the present application is shown.
[0040] As shown in Figure 1 , the application scenario 100 according to the embodiment can include a baseboard management controller 101 (BMC), a video graphics array interface 102 (VGA interface), a host 103, an out-of-band management service module 104 (OOB MSM) and a peripheral component interconnect express (PCIe) link. The host 103 includes the out-of-band management service module 104, and the baseboard management controller 101 includes the video graphics array interface 102.
[0041] As shown in Figure 1As shown, the baseboard management controller 101 and the host 103 are connected through a PCIe link, and the out-of-band management service module 104 can exist as a PCIe device to directly interact with the baseboard management controller 101 and provide CPU information query, configuration, and fault diagnosis capabilities. On the hardware, the baseboard management controller 101 has a PCIe lane connected to the CPU, and the out-of-band management service module 104 communicates with the baseboard management controller 101 through this physical link. In an embodiment, the baseboard management controller 101, the video graphics array interface 102, the CPU, and the out-of-band management service module 104 can be connected via the communication link of MCTP.
[0042] According to an embodiment of the present application, the out-of-band management service module 104 can communicate with the baseboard management controller 101 as a bus owner (MCTP bus owner). When the server completes the startup process and enters the OS, the out-of-band management service module 104 can send an MCTP message to the baseboard management controller 101 through the PCIe link every 5s. After receiving the MCTP message, the baseboard management controller 101 can obtain the EID information, can return an MCTP response according to the MCTP protocol specification, and can return the MCTP response to the out-of-band management service module 104 through the PCIe link. At the same time, the baseboard management controller 101 can also act as a control node to initiate an MCTP request through the PCIe link in the reverse direction, access other PCIe devices (such as GPU acceleration cards, NVMe controllers) in the host that support the MCTP protocol, and implement out-of-band device state monitoring.
[0043] Figure 2 A flowchart of a communication fault determination method according to an embodiment of the present application is shown.
[0044] As shown in Figure 2 The method 200 includes operations S210-S230. The method 200 can be applied to a BMC.
[0045] At operation S210, in response to receiving register error information from a central processor, the register error information is parsed to obtain routing identification information, and the routing identification information is used to locate an error triggering device; wherein the register error information is obtained from a register related to a baseboard management controller based on a target error event by the central processor.
[0046] At operation S220, in a case where the routing identification information indicates an out-of-band management service module, the register error information is determined to be target communication type information; wherein the baseboard management controller, the central processor, and the out-of-band management service module are connected via a communication link of a target communication type.
[0047] At operation S230, a communication fault type is determined according to the configuration information of the communication link, and the communication fault type is sent to the central processor.
[0048] According to one embodiment of the present application, the method 200 can be applied to determine the MCTP communication error occurred when the server system startup process CPU communicates with the BMC, and the VGA display problem can occur in the case of the aforementioned MCTP error.
[0049] In one embodiment, the target error event is that the BMC reports to the CPU that an ACS (Access Control Services) error occurs. The register error information is the register error information of the RP end (Root Port Register) of the BMC collected by the BIOS when the target error event occurs, and the register error information includes the TLP (Transaction Layer Packet) that is in error. The routing identification information is the Target ID (target ID) in the Header field of the TLP.
[0050] Specifically, when the CPU / OOB MSM sends an MCTP message to the BMC, if the BMC returns an error or does not respond, an ACS error can be generated. When the ACS error is upgraded to a UCE error (uncorrect error), an SMI (System Management Interrupt) is triggered, at which time the BIOS acquires the register information of the RP end of the BMC and sends the register information to the BMC. The RP end, also known as the Root Port end, is the root port on the PCIe Root Complex in the CPU / chipset. The RP end of the BMC refers to the Root Port and its related register area on the PCIe link connected to the BMC on the CPU / chipset side. When the BMC communicates with the CPU, all TLPs, error logs, and AER (Advanced Error Reporting) information are collected by the Root Port and reported. When an error occurs in the MCTP communication between the BMC and the CPU, the Root Port (RP) records the TLP in error in the TLP Header Log area of the AER extension register.
[0051] In the PCIe architecture, TLP is the core data transmission unit for transferring data and control information between devices. TLP is transmitted through the serial connection of PCIe, involving multiple levels of processing, including transaction layer, data link layer and physical layer. The structure of TLP includes one or more Header fields and an optional Data field. The Header field of TLP contains key information such as request type, transmission length, requester ID, target ID and data validity flag.
[0052] According to an embodiment of the present application, the BMC receives the RP end register information (i.e. register error information) reported by the BIOS, which includes the TLP message. After receiving the TLP message, the BMC can parse the Header field of the TLP to obtain the Target ID. In the case where the Target ID indicates an OOB MSM, since the BMC, CPU and OOB MSM are connected via the MCTP communication link, it can be determined that the TLP message is MCTP information, and further that an MCTP communication error has occurred. The logic is as follows: the OOB MSM as a bus owner means that it has control over the PCIe bus, and the ACS error is essentially a bus access conflict, which indicates that the OOB MSM attempts to access an unauthorized or unreachable device. Since the OOB MSM, CPU, BMC and VGA are all on the MCTP communication link, if it is determined that the error is triggered by the OOB MSM, it can be further traced back to an MCTP fault.
[0053] In an embodiment, in the case where it is determined that it is MCTP information, the configuration information related to the MCTP communication link where the BMC is located can be obtained, and the communication fault type can be determined according to the configuration information, and then the communication fault type can be sent to the CPU.
[0054] According to an embodiment of the present application, when the BMC reports to the CPU that an ACS error has occurred, the BIOS can automatically obtain the register error information of the RP end of the BMC and send the register error information to the BMC. The BMC can parse the received TLP message and determine the Target ID based on the TLP Header. In the case where the Target ID indicates an OOB MSM, it can be determined that the TLP message is MCTP information, and thus it can be determined that an MCTP communication error has occurred. Further, the configuration information related to the MCTP communication link where the BMC is located can be obtained to determine the specific fault type / reason.
[0055] It can be understood that the communication fault determination method (hereinafter referred to as the method) provided by the embodiments of the present application can report the register error information related to the BMC to the BMC through the BIOS automatically at the first time when the BMC reports the ACS error. The BMC can analyze the Target ID based on the TLP Header, and in the case that the Target ID indicates the OOB MSM (MCTP bus owner), the communication link among the OOB MSM, the CPU and the BMC can be used to exclude the non-MCTP interference factors, determine that the MCTP communication error occurs, and then the fault type can be accurately determined according to the MCTP link configuration information. For the fault symptoms such as VGA output abnormality and system "stuck", the method can quickly distinguish whether it is caused by the MCTP communication error, so as to reduce the misjudgment, save a lot of manpower and material resources, and help to solve the problem of fault type positioning behind abnormal symptoms.
[0056] According to the embodiments of the present application, the method further comprises: before receiving the register error information from the central processor, in response to receiving the communication request from the out-of-band management service module, sending a communication response to the out-of-band management service module; in the case that the communication response sending fails, triggering a target error event.
[0057] In one embodiment, the BIOS and the BMC can be connected through PCIe in hardware. The BIOS can initialize the VGA display of the BMC in the boot process, allocate a bus number to the BMC, and then initialize the OOB MSM, and the OOB MSM is the bus owner by default. The BMC needs to update the MCTP BDF information (BUS is the BUS allocated to the BMC by the BIOS) of itself after the BIOS completes the POST (power-on self-test).
[0058] The OOB MSM can communicate with the BMC as the bus owner (MCTP bus owner). When the server completes the startup process and enters the OS, the OOB MSM can send the MCTP message to the BMC through the PCIe link every 5s. The BMC can obtain the EID information after receiving the MCTP message, can return the MCTP response according to the MCTP protocol specification, and return the MCTP response to the OOB MSM through the PCIe link. When the CPU / OOB MSM sends the MCTP message to the BMC, if the BMC returns an error or does not respond, an ACS error can be generated.
[0059] According to the embodiments of the present application, the register error information is analyzed to obtain the routing identification information, comprising: based on a preset communication protocol, analyzing the header of the register error information to obtain the routing identification information.
[0060] In one embodiment, the Requester ID and Completer ID (Target ID) fields in the TLP Header are each 16 bits long and have the format Bus(8) + Device(5) + Function(3). The BMC can parse the TLP Header according to the PCIe protocol to obtain the Target ID after receiving the TLP message.
[0061] According to an embodiment of the present application, in the case where the routing identification information indicates the out-of-band management service module, determining the register error information as the target communication type information comprises: determining an error triggering device according to a preset routing mapping relationship and the routing identification information; and in the case where the error triggering device is the out-of-band management service module, determining the register error information as the target communication type information.
[0062] In one embodiment, the Target ID in the PCIe architecture is the Completer ID, which is a fixed field in each TLP Header and is composed of Bus:Device:Function (BDF), and is used to tell the switching network "where to send this TLP". The destination of the TLP can be determined according to a preset static routing table and the Target ID.
[0063] For example, when the value of the Target ID is fixedly mapped to 0x18, it indicates that the final destination of the TLP is the OOBMSM, and thus the TLP is MCTP information.
[0064] According to an embodiment of the present application, in the case where the substrate management controller is connected to the central processor as a high-speed bus expansion device, the method further comprises: in the case where the register error information is determined as the target communication type information, obtaining communication protocol stack information of the substrate management controller and a bridge bus number and a device bus number allocated to the substrate management controller; and wherein the communication protocol stack information comprises a bus device function number of the substrate management controller.
[0065] In one embodiment, the BMC and the host are connected through a PCIe link. In the case where the TLP (register error information) is determined as MCTP information, the BMC can obtain the BDF (i.e. the bus device function number recorded by the BMC) in the internal MCTP structure data and the Bridge bus number and the Device bus number of the internal register.
[0066] According to an embodiment of the present application, the determining the communication fault type according to the configuration information of the communication link comprises: matching the bus device function number respectively with the bridge bus number and the device bus number; in the case that the bus device function number does not match at least one of the bridge bus number and the device bus number, determining that the communication fault type is a bus device function configuration error of the baseboard management controller; in the case that the bus device function number matches both the bridge bus number and the device bus number, determining that the communication fault type is a communication error between the baseboard management controller and the central processing unit.
[0067] In one embodiment, the BMC can check whether the BDF in the internal MCTP structure data is consistent with the internal register Bridge bus number and Device bus number. Wherein, if the BDF is not consistent, it means that the actual BDF of the device does not match the BDF claimed in the firmware / configuration / packet, which will lead to routing failure or permission check failure, thus triggering an error interrupt or a log alarm.
[0068] As an example, if the BDF in the MCTP is inconsistent with the register, the BMC can record a log, explicitly recording that the communication fault type is that the BDF in the MCTP is inconsistent with the BUS information of the register. Further, maintenance suggestions can also be given, for example, asking the BIOS to check whether the bus assigned to the BMC by the BIOS is correct and final, and asking the BMC to check whether the BDF in the MCTP structure data of the BMC is the BUS finally written by the BIOS.
[0069] As another example, if the BDF in the MCTP is consistent with the register, then the communication fault type can be recorded as MCTP communication error. Further, maintenance suggestions can also be given, for example, the MCTP can be closed for testing, such as closing the MCTP at the BIOS end or closing the MCTP at the BMC end.
[0070] In one embodiment, the BMC can send the communication fault type and the corresponding maintenance suggestion to the CPU.
[0071] Figure 3 A flowchart of a communication fault determination method according to another embodiment of the present application is shown.
[0072] As Figure 3 shown, the method 300 comprises operations S310-S330. The method 300 can be applied to the CPU.
[0073] At operation S310, in response to a target interrupt event being triggered, error event information is acquired.
[0074] In operation S320, the error event information is analyzed to determine an error event type.
[0075] In operation S330, in a case where the error event type is a target error event, register error information in a target register is acquired based on a basic input output system, and the register error information is sent to a baseboard management controller.
[0076] According to one embodiment of the present disclosure, the target interrupt event is an SMI (System Management Interrupt). In response to the SMI, relevant error event information can be acquired, and the error event information can be analyzed to determine an error event type.
[0077] In one embodiment, in a case where the error event type is that an ACS error is reported on an RP end where the BMC is located, i.e., in a case where the error event type is a target error event, register error information of the RP end where the BMC is located can be acquired based on BIOS, and the register error information is sent to the BMC.
[0078] According to an embodiment of the present disclosure, the method further includes, in response to the target interrupt event being triggered, sending, based on the out-of-band management service module, a communication request to the baseboard management controller, so as to trigger the target error event in a case where a communication failure occurs between the baseboard management controller and the central processing unit.
[0079] In one embodiment, when the server completes a startup process and enters an OS, the OOBSM can send an MCTP message to the BMC through a PCIe link every 5 s. After the BMC receives the MCTP message, the BMC can acquire EID information, can return an MCTP response according to an MCTP protocol specification, and can return the MCTP response to the OOBSM through the PCIe link. When the CPU / OOBMS sends an MCTP message to the BMC, if the BMC returns an error or does not respond, an ACS error can be generated.
[0080] According to an embodiment of the present disclosure, the method further includes, in response to receiving the communication failure type from the baseboard management controller, sending alarm information according to the communication failure type.
[0081] In one embodiment, the CPU receives the communication failure type and maintenance suggestions sent by the BMC, and can send alarm information to maintenance personnel, so that the maintenance personnel can quickly locate a fault and solve the problem.
[0082] In an optional embodiment, register error information and determined communication fault type and the like can also be stored to establish a fault data set. The fault determination model can be trained based on the fault data set, and the register error information can be analyzed based on the trained fault determination model to achieve the effect of accurate and rapid fault determination.
[0083] Figure 4 A communication fault determination flowchart according to an embodiment of the application is shown.
[0084] As Figure 4 shown, the basic input and output system (BIOS) can perform initialization during the boot process. The initialization operation may, for example, include initializing the VGA display of the BMC, assigning a bus number to the BMC, and then initializing the OOB MSM, with the default OOB MSM as the bus owner. The baseboard management controller (BMC) needs to wait until the BIOS completes the POST (power-on self-test) before updating its own MCTP BDF information. Then, as Figure 4 shown, the error reporting mechanism can be started. The error reporting mechanism is as follows: during the process of entering the OS, the CPU starts to send MCTP messages to the outside through the OOB MSM, and if the BMC returns an error message, an ACS error may be generated. This UCE error triggers an SMI, at which time the BIOS reports the register information of the RP side, including the TLP HEADER information.
[0085] As Figure 4 shown, in the case that the target error event is reported on the root port register where the baseboard management controller is located, that is, in the case that the ACS error information is reported on the RP where the BMC is located, the basic input and output system (BIOS) can send the register error information to the baseboard management controller (BMC). As Figure 4 shown, after receiving the register error information, the baseboard management controller (BMC) can parse the header (TLP Header) and determine the Target ID (routing identification information) based on the TLP Header. If the Target ID indicates the OOB MSM, it can be determined that the TLP message is an MCTP message (target communication type information), so that it can be determined that an MCTP communication error has occurred.
[0086] As Figure 4As shown, in a case where it is determined that the register error information is target communication type information, it can be further determined whether the bus device function number matches the bridge bus number and the device bus number. That is, in a case where it is determined that the TLP message is an MCTP message, the BMC can further determine whether the MCTP EID of the BMC is consistent with the BDF reported by the BIOS. If it is determined that the BDF in the MCTP of the BMC is inconsistent with the BDF reported by the BIOS, it can be determined that the cause of the failure is that the BDF is inconsistent (that is, it is determined that the communication failure type is that a bus device function configuration error of the baseboard management controller occurs), and an alarm and a solution suggestion are given. If the BDF in the MCTP of the BMC is consistent with the BDF reported by the BIOS, it can be determined that an MCTP communication error occurs (that is, it is determined that the communication failure type is that a communication error occurs between the baseboard management controller and the central processor).
[0087] The embodiment of the present application provides a method for identifying an MCTP error. When a PCIe error is triggered, the BIOS reports the error to the BMC through SMI, so that the BMC automatically collects other related information, analyzes TLP header information, and determines what type of message causes the error. In a case where it is determined that the message is an MCTP message, it is further analyzed whether the EID in the MCTP is the bus allocated to the BMC by the BIOS. If the EID in the MCTP does not match the bus allocated by the BIOS, error information is recorded in time, and alarm information is given.
[0088] Figure 5 A block diagram of an electronic device suitable for implementing the communication failure determination method according to the embodiment of the present application is shown.
[0089] As shown, Figure 5 The electronic device 500 according to the embodiment of the present application can include a baseboard management controller 510, one or more processors 520, and a memory 530.
[0090] The memory 530 is configured to store one or more computer programs.
[0091] The baseboard management controller 510 and the one or more processors 520 are configured to execute the one or more computer programs to implement the steps according to the above method.
[0092] The storage 530 can include, for example, but not limited to, read only memory (ROM), random access memory (RAM), and the like. For example, in the RAM, various programs and data required for the operation of the electronic device 500 are stored. The processor 520, the ROM, and the RAM are connected to each other through a bus. The processor 520 performs various operations of the method procedures according to embodiments of the present application by executing the programs in the ROM and / or the RAM. Note that the programs can also be stored in one or more memories other than the ROM and the RAM. The processor 520 can also perform various operations of the method procedures according to embodiments of the present application by executing the programs stored in the one or more memories.
[0093] The processor 520 can perform various appropriate actions and processes according to programs stored in a read only memory (ROM) or loaded from a storage section into a random access memory (RAM). The processor 520 can include, for example, a general purpose microprocessor (e.g., a CPU), an instruction set processor, and / or a related chip set, and / or a dedicated microprocessor (e.g., an application specific integrated circuit (ASIC)), and the like. The processor 520 can also include an on-board memory for cache use. The processor 520 can include a single processing unit or a plurality of processing units for performing different actions of the method procedures according to embodiments of the present application.
[0094] According to embodiments of the present application, the electronic device 500 can further include an input / output (I / O) interface, which is also connected to the bus. The electronic device 500 can further include one or more of the following components connected to the input / output (I / O) interface: an input section including a keyboard, a mouse, and the like; an output section including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), and the like, and a speaker, and the like; a storage section including a hard disk, and the like; and a communication section including a network interface card such as a LAN card, a modem, and the like. The communication section performs communication processing via a network such as the Internet. A drive is also connected to the input / output (I / O) interface as necessary. A removable medium such as a magnetic disk, a magneto-optical disk, a semiconductor memory, and the like is attached to the drive as necessary, so that a computer program read out from the removable medium is installed into the storage section as necessary.
[0095] The present application also provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments; or can exist separately without being assembled into the device / apparatus / system. The above computer readable storage medium carries one or more programs, which, when executed, implement the method according to embodiments of the present application.
[0096] According to an embodiment of the present application, the computer readable storage medium can be a non-transitory computer readable storage medium, for example, can include but not limited to: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In this application, a computer readable storage medium can be any tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, a computer readable storage medium can include one or more memory components described above and / or one or more memories other than the ROM and / or RAM.
[0097] Embodiments of the present application also include a computer program product, which includes a computer program containing program codes for executing the methods shown in the flowcharts. When the computer program product is run in a computer system, the program codes are used to make the computer system implement the communication fault determination method provided by the embodiments of the present application.
[0098] The above functions defined in the system / device / apparatus of the embodiments of the present application are performed when the computer program is executed by the processor 520. According to an embodiment of the present application, the system, device, module, unit, etc. described above can be implemented by computer program modules.
[0099] In one embodiment, the computer program can rely on tangible storage media such as optical storage media, magnetic storage media, etc. In another embodiment, the computer program can also be transmitted, distributed, and downloaded in the form of signals on a network medium, and be downloaded and installed through the communication part, and / or installed from a detachable medium. The program codes contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the foregoing.
[0100] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part, and / or installed from the detachable medium. When the computer program is executed by the processor 520, the above functions defined in the system of the embodiments of the present application are performed. According to an embodiment of the present application, the system, device, apparatus, module, unit, etc. described above can be implemented by computer program modules.
[0101] According to embodiments of the present application, program code for implementing the computer programs provided by embodiments of the present application can be written in any combination of one or more programming languages, and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming language includes, but is not limited to, such languages as Java, C++, python, "C" language, or the like. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the remote computing device, or entirely on the remote computing device or server. In the latter scenario, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing device, such as through the Internet using an Internet Service Provider.
[0102] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0103] Those skilled in the art will understand that features recited in the various embodiments of the present application can be combined and / or integrated in a variety of ways, even if such combinations or integrations are not expressly noted in the present application. In particular, features recited in the various embodiments of the present application can be combined and / or integrated in ways that are not expressly noted in the present application, without departing from the spirit and teachings of the present application. All such combinations and / or integrations are within the scope of the present application.
[0104] The embodiments of the present application are described above. However, these embodiments are merely for illustration purposes, and are not intended to limit the scope of the present application. Although the embodiments are described above separately, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of the present application, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications shall fall within the scope of the present application.
Claims
1. A communication fault determination method, applied to a baseboard management controller, characterized in that, The method includes: In response to receiving register error information from the central processing unit, the register error information is parsed to obtain routing identification information, which is used to locate the error triggering device; wherein, the register error information is obtained by the central processing unit from registers related to the baseboard management controller based on the triggering of a target error event; the target error event is an access control service error event, which is triggered when a communication failure occurs between the baseboard management controller and the central processing unit; If the routing identification information indicates an out-of-band management service module, the register error information is determined to be management component transmission protocol information; wherein, the baseboard management controller, the central processing unit, and the out-of-band management service module are connected via a communication link of the management component transmission protocol; the out-of-band management service module is used as the owner of the management component transmission protocol bus; Based on the configuration information of the communication link of the management component transmission protocol, the communication fault type is determined and the communication fault type is sent to the central processing unit; the communication fault type includes: the baseboard management controller has a bus device function configuration error; or the management component transmission protocol communication error occurs between the baseboard management controller and the central processing unit.
2. The method according to claim 1, characterized in that, The step of parsing the register error information to obtain the routing identifier information includes: Based on a preset communication protocol, the header of the register error message is parsed to obtain the routing identification information.
3. The method according to claim 1, characterized in that, When the routing identification information indicates an out-of-band management service module, determining that the register error information is management component transmission protocol information includes: The error triggering device is determined based on the preset routing mapping relationship and the routing identification information; When the error triggering device is an out-of-band management service module, the register error information is determined to be management component transmission protocol information.
4. The method according to claim 1, characterized in that, When the baseboard management controller is connected to the central processing unit as a high-speed bus expansion device, the method further includes: If the register error information is determined to be management component transmission protocol information, the communication protocol stack information for the baseboard management controller, as well as the bridge bus number and device bus number assigned to the baseboard management controller are obtained. The communication protocol stack information includes the bus device function number of the baseboard management controller.
5. The method according to claim 4, characterized in that, The step of determining the communication fault type based on the configuration information of the communication link includes: The bus device function number is matched with the bridge bus number and the device bus number, respectively; If the bus device function number does not match at least one of the bridge bus number and the device bus number, the communication fault type is determined to be a bus device function configuration error of the baseboard management controller. If the bus device function number matches both the bridge bus number and the device bus number, the communication fault type is determined to be a management component transmission protocol communication error between the baseboard management controller and the central processing unit.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: before receiving register error information from the central processing unit, In response to receiving a communication request from the out-of-band management service module, a communication response is sent to the out-of-band management service module; In the event that the communication response fails to be sent, the target error event is triggered.
7. A method for determining communication faults, applied to a central processing unit, characterized in that, The method includes: In response to a target interrupt event being triggered, obtain error event information; The error event information is analyzed to determine the error event type; In the case where the error event type is a target error event, register error information in the target register is obtained based on the Basic Input / Output System (PIS), and the register error information is sent to the Baseboard Management Controller (BMC). This allows the BMC to determine the communication fault type based on the configuration information of the communication link of the Management Component Transmission Protocol (MTP). The target error event is an Access Control Service (ACCS) error event. The communication fault type includes: a bus device function configuration error occurring in the BMC; or a MTP communication error occurring between the BMC and the CPU. The method further includes: in response to the target interrupt event being triggered, sending a communication request to the baseboard management controller based on the out-of-band management service module, so as to trigger the target error event in the event of a communication failure between the baseboard management controller and the central processing unit; the baseboard management controller is connected to the central processing unit as a high-speed bus extension device; the baseboard management controller, the central processing unit, and the out-of-band management service module are connected via a communication link of the management component transmission protocol; the out-of-band management service module is used as the owner of the management component transmission protocol bus.
8. The method according to claim 7, characterized in that, The method further includes: In response to receiving a communication fault type from the baseboard management controller, an alarm message is sent according to the communication fault type.
9. An electronic device, comprising: Baseboard management controller; One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the substrate management controller executes the one or more computer programs to implement the steps of the method according to any one of claims 1 to 6; the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 7 to 8.
Citation Information
Patent Citations
Fault processing method and device and server
CN111414268A
Fault information acquisition method and device, baseboard management controller, system and medium
CN116974809A