Abnormality monitoring method and device for internal module of data processor, chip, network interface card, computer equipment, readable storage medium and program product

By storing and reporting abnormal information of the internal module of the data processor when an exception is detected by the RAS node, the problem of failure of the internal module of the DPU in the prior art is solved, and higher reliability and maintainability are achieved.

CN120276934APending Publication Date: 2025-07-08NANJING JAGUAR MICROSYSTEMS CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510358034.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing RAS detection system cannot effectively support the fault reporting and failure recovery of IP Cores such as VPE, DPE, RPE, etc. in DPU internal modules, resulting in insufficient reliability and maintainability of the data processor.

Method used

When an RAS node detects an internal module exception, it stores exception information and obtains the exception event reporting method, and uses mechanisms such as interrupt controller, kernel state driver or user state driver to report and process abnormal information to achieve fault detection and recovery.

Benefits of technology

It improves the fault detection capability and system stability of the internal modules of the data processor, enhances maintainability, and ensures reliable operation and rapid failure recovery of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276934A_ABST
    Figure CN120276934A_ABST
Patent Text Reader

Abstract

The invention relates to an exception monitoring method and device for an internal module of a data processor, a chip, a network interface card, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: under the condition that an internal module is detected to be abnormal through an RAS node, storing abnormal information of the abnormal internal module; obtaining an abnormal event reporting mode corresponding to the RAS node; and reporting the abnormal information through an event based on the abnormal event reporting mode. By adopting the method, the faults of the internal modules of the data processor can be reported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and particularly to a method, apparatus, chip, network interface card, computer device, computer-readable storage medium, and computer program product for abnormal monitoring of internal modules of a data processor. Background Art

[0002] A DPU (Data Processing Unit) is a dedicated processor constructed with data as the center, which completes acceleration processing tasks for networks, storage, and security. The DPU utilizes internally integrated IP cores such as CSUB (CPU subsystem), DPE (Dataplane Process Engine), VPE (Virtio Process Engine), RPE (Rdma Process Engine), HAC (Hardware Accelerate), and PSUB (PCIE subsystem) to provide a complete solution for cloud service providers for offloading and accelerating networks, computing, storage, security, etc. The long-term stable operation of the internal IP cores of the DPU is crucial for the DPU.

[0003] Reliability, Availability, and Serviceability (RAS) is a concept used in servers to measure their robustness. The RAS goal is to enable the system to run as reliably as possible for a long time without downtime, reduce system downtime; provide a hardware detection and reporting mechanism so that the administrator can be notified in time to replace the hardware before data loss or downtime caused by hardware errors; provide a hardware error recovery mechanism and correct the errors as much as possible to enable the system to run sustainably and reliably.

[0004] In traditional technologies, when a peripheral module of the CPU RAS processing solution fails, the OS (Operation System) is notified through methods such as interrupt, polling, and event reporting, and the driver on the OS completes the collection of fault information and fault recovery actions.

[0005] However, current RAS detection systems can only count RAS reports of CPUs, memories, and PCIEs and do not support fault reporting and fault recovery of IP cores such as the internal modules VPE (Virtio Process Engine), DPE (Dataplane Process Engine), and RPE (Rdma Process Engine) of the DPU. Summary of the Invention

[0006] Based on this, in view of the above technical problems, it is necessary to provide an abnormal monitoring method, device, chip, network interface card, computer device, computer-readable storage medium, and computer program product for internal modules of a data processor that can perform fault reporting on internal modules of the data processor.

[0007] In a first aspect, the present application provides an abnormal monitoring method for internal modules of a data processor, which is applied to the data processor. The method includes:

[0008] When an internal module exception is detected through the RAS node, store the exception information of the abnormal internal module;

[0009] Obtain the abnormal event reporting method corresponding to the RAS node;

[0010] Based on the abnormal event reporting method, report the abnormal information through an event.

[0011] In one embodiment, before obtaining the abnormal reporting method corresponding to the RAS node, it further includes:

[0012] Configure the abnormal event reporting method through the extensible firmware interface or basic input / output system of the central processing unit subsystem module;

[0013] Store the abnormal event reporting method in the RAS control node of the data processor;

[0014] The obtaining of the abnormal event reporting method corresponding to the RAS node includes:

[0015] Obtain the abnormal event reporting method corresponding to the RAS node from the RAS control node.

[0016] In one embodiment, the reporting of the abnormal information through an event based on the abnormal event reporting method includes any one of the following:

[0017] The interrupt of the exception event corresponding to the exception information is notified to the basic input / output system of the manageable control processor through the interrupt controller of the data processor, and the exception information corresponding to the exception event is reported to the management node on the basic event controller through the basic input / output system of the manageable control processor, and the exception information in the management node is provided for the user to query;

[0018] The interrupt of the exception event corresponding to the exception information is notified to the extensible firmware interface or the basic input / output system of the central processing unit subsystem module through the interrupt controller of the data processor, and after the exception information is collected and processed through the extensible firmware interface or the basic input / output system of the central processing unit subsystem module, the exception information corresponding to the exception event is reported through the software delegate exception interface method, and the exception information is persistently stored through the kernel tracing tool;

[0019] The exception information stored in the RAS node corresponding to the internal module is read through the kernel-mode driver or the user-mode driver of the central processing unit kernel system in a timed polling manner, and the exception information is reported, and the exception information is persistently stored through the kernel tracing tool.

[0020] In one embodiment, the method further includes:

[0021] When an internal module exception is detected in the RAS node, determine the type of the exception;

[0022] When the type of the exception is a correctable exception, trigger a fault handling interrupt, and the fault handling interrupt is used to report the exception information corresponding to the exception;

[0023] When the type of the exception is an uncorrectable exception, perform at least one of the following operations: trigger an error recovery interrupt, and the error recovery interrupt is used to report the exception information corresponding to the exception; return an in-band error response to the processing unit, and the in-band error response is used to instruct the processing unit to report a synchronous external data abort signal or an asynchronous error interrupt signal.

[0024] In one embodiment, storing the exception information of the internal module with exceptions includes:

[0025] Storing the exception information of the internal module with exceptions based on a standard error record, where the standard error record includes a status register for storing a common status field and an optional address register for storing the address where the error occurs; or the standard error record includes a status register for storing a common status field, an optional address register for storing the address where the error occurs, and a custom register for storing auxiliary information.

[0026] In a second aspect, the present application further provides an abnormal monitoring method for internal modules of a data processor, and the method includes:

[0027] Receiving and storing abnormal information reported by a RAS node, where the abnormal information is reported by the RAS node based on the abnormal monitoring method for internal modules of the above data processor;

[0028] Returning corresponding abnormal information based on a user request.

[0029] In one embodiment, the receiving and storing abnormal information reported by a RAS node includes at least one of the following:

[0030] Receiving and storing the abnormal information through a management node on a basic event controller, and the abnormal information in the management node is for user query;

[0031] Receiving an abnormal event through a scalable firmware interface or a basic input / output system of a central processing unit subsystem module, and storing the abnormal information into a cross-platform error report based on the abnormal event; calling an event processing function hooked by a hardware error driver through a software delegation abnormal interface method to read and process the abnormal processing information, and persistently storing the abnormal information through a kernel tracing tool;

[0032] Reading the abnormal information stored in the RAS node corresponding to the internal module in a timed polling manner through a kernel-mode driver or a user-mode driver, reporting the abnormal information, and persistently storing the abnormal information through a kernel tracing tool.

[0033] In one embodiment, the returning corresponding abnormal information based on a user request includes:

[0034] Returning corresponding abnormal information through a subscription method; and / or

[0035] Returning corresponding abnormal information through a RAS interrupt method.

[0036] In a third aspect, the present application further provides an abnormal monitoring device for internal modules of a data processor, which is applied to a data processor, and the device includes:

[0037] A storage module, configured to store abnormal information of an abnormal internal module when an internal module abnormality is detected through a RAS node;

[0038] A reporting method acquisition module, configured to acquire an abnormal event reporting method corresponding to a RAS node;

[0039] A reporting module, configured to report the abnormal information through an event based on the abnormal event reporting method.

[0040] Fourth aspect, the present application further provides an abnormal monitoring device for an internal module of a data processor, and the device includes:

[0041] A receiving module, configured to receive and store abnormal information reported by an RAS node, where the abnormal information is reported by the RAS node based on the abnormal monitoring device for the internal module of the above-mentioned data processor;

[0042] A returning module, configured to return corresponding abnormal information based on a user request.

[0043] Fifth aspect, the present application further provides a chip, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the method in any one of the above-mentioned embodiments are implemented.

[0044] Sixth aspect, the present application further provides a network interface card, including the chip in any one of the above-mentioned embodiments and a plurality of interfaces, where the chip processes data or communicates externally through the interfaces.

[0045] Seventh aspect, the present application further provides a computer device, including the network interface card in any one of the above-mentioned embodiments, where the network interface card is used to process data or communicate externally.

[0046] Eighth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method in any one of the above-mentioned embodiments are implemented.

[0047] Ninth aspect, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method in any one of the above-mentioned embodiments are implemented.

[0048] For the above-mentioned abnormal monitoring method, device, chip, network interface card, computer device, computer-readable storage medium and computer program product of the internal module of the data processor, the internal module of the data processor all has an RAS node. In this way, when the RAS node detects an abnormality in the internal module, the abnormal information of the abnormal internal module is stored, and the reporting method of the corresponding abnormal event of the RAS node is obtained; based on the reporting method of the abnormal event, the abnormal information is reported through an event. In this way, the data processor side can implement the fault detection and reporting of the internal module of the data processor, which helps to improve the maintainability and stability of the data processor. Description of the Drawings

[0049] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the accompanying drawings required for the description of the embodiments of the present application or related technologies. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related accompanying drawings can also be obtained based on these drawings.

[0050] Figure 1 Schematic diagram of the application environment of the abnormal monitoring method for the internal modules of the data processor in one embodiment;

[0051] Figure 2 Schematic diagram of the process of the abnormal monitoring method for the internal modules of the data processor in one embodiment;

[0052] Figure 3 Schematic diagram of the process of reporting the RAS interrupt in the DPU scenario of the data processor to the Baseboard Management Controller (BMC) in one embodiment;

[0053] Figure 4 Schematic diagram of the process of reporting the RAS interrupt in the DPU scenario of the data processor to the Unified Extensible Firmware Interface (UEFI) in one embodiment;

[0054] Figure 5 Schematic diagram of the process of the kernel mode or user mode driver in the RAS of the DPU scenario of the data processor polling the RAS exception information regularly in one embodiment;

[0055] Figure 6 Schematic diagram of the process of the configuration steps of the reporting method in one embodiment;

[0056] Figure 7 Schematic diagram of the reporting process in one embodiment;

[0057] Figure 8 Schematic diagram of the structure of the RAS node in one embodiment;

[0058] Figure 9 Schematic diagram of the process of the abnormal monitoring method for the internal modules of the data processor in another embodiment;

[0059] Figure 10 Schematic diagram of the process of the abnormal monitoring method for the internal modules of the data processor in still another embodiment;

[0060] Figure 11 Schematic diagram of the process of the abnormal monitoring method for the internal modules of the data processor corresponding to the second reporting method in one embodiment;

[0061] Figure 12 Schematic diagram of the interaction process of the initialization process SDEI in one embodiment;

[0062] Figure 13 It is a schematic block diagram of an abnormal monitoring device for an internal module of a data processor in an embodiment;

[0063] Figure 14 It is a schematic block diagram of an abnormal monitoring device for an internal module of a data processor in another embodiment;

[0064] Figure 15 It is a schematic diagram of the internal structure of a computer device in an embodiment. Specific embodiments

[0065] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0066] The method for monitoring abnormal conditions of internal modules of a data processor provided by an embodiment of the present application can be applied to an application environment as Figure 1 shown. Among them, the data processor DPU includes at least one internal module, such as VPE (Virtio Process Engine, virtual device queue engine), DPE (Dataplane Process Engine, data processing engine), RPE (Rdma Process Engine, RDMA service processing engine), HAC (Hardware Accelerate, offloading of network data security services), etc. Figure 1 Only some internal modules are shown in, and in other embodiments, the data processor DPU may further include other internal modules, which are not specifically limited herein.

[0067] The central processing unit CPU includes a virtualization system, and the virtualization system includes a VM (Virtual Machine, virtual machine).

[0068] Among them, the uplink network data is connected from the Switch (network switch) to the MAC (Media Access Control, media access control) module of the data processor DPU through the Fiber (optical fiber). The DPU uses the internal integrated CSUB (CPU subsystem, CPU subsystem), DPE, VPE, RPE, HAC, PSUB (PCIE subsystem, PCIE subsystem, not shown in Figure 1 ) and other internal module IP cores to provide network, computing, storage, etc. offloading and acceleration for cloud service providers, and send the network traffic to the Host CPU and VM through the PCIE interface.

[0069] Downlink network data is sent from the Host CPU and VMs to the Switch through the Data Processor Unit (DPU). The internal module IP core integrated inside the DPU provides a complete solution for network, computing, storage, etc. offloading and acceleration for cloud service providers.

[0070] Among them, each internal module generally includes a dedicated memory (RAM) component, a cache component, a clock bus component, etc. for the internal module IP Core; these components report errors to the RAS Node (RAS node) after detecting errors; for the DPU scenario, the RAS Node includes: CSUB Node, Memory Node, PSUB Node, DPE Node, VPENode, RPE Node, HAC Node, etc.

[0071] In an exemplary embodiment, as Figure 2 shown, an exception monitoring method for the internal module of a data processor is provided. Taking the data processor DPU in Figure 1 as an example, the method includes the following steps 202 to 206. Among them:

[0072] S202: When an internal module exception is detected through the RAS node, store the exception information of the exception internal module.

[0073] The data processor includes multiple internal modules. In this application, a corresponding RAS node is configured for each internal module. The RAS node can detect exceptions of each component of the internal module and store the exception information of the exception internal module when an exception is detected.

[0074] Among them, the RAS nodes corresponding to each internal module do not interfere with each other and operate independently.

[0075] Each RAS node includes an exception record register, which is used to store exception information. Optionally, the exception information includes: the severity level of the error, the address where the error occurs, the value of the correctable error (CE) counter that may be included, information for identifying the Field Replaceable Units (FRUs) and locating errors within the FRUs, and other implementation-defined information.

[0076] In other alternative embodiments, the exception record register further includes a control register, which stores control information for error detection, correction, and reporting at the node. For example, the control information may include the target value of the CE counter, etc., which is not specifically limited here.

[0077] S204: Obtain the abnormal event reporting method corresponding to the RAS node.

[0078] S206: Report abnormal information through events based on the abnormal event reporting method.

[0079] The abnormal event reporting method is pre-configured, and this abnormal event reporting method is also the way for the RAS node to report abnormal information.

[0080] In some alternative embodiments, the abnormal event reporting method may include at least one of three types. The specific definitions of these three abnormal event reporting methods can be found below.

[0081] Among them, the abnormal event reporting method can be stored in the RAS control node. In this way, after the RAS detects an abnormality, it can obtain the abnormal event reporting method from the RAS control node and report the abnormal information based on the obtained abnormal event reporting method.

[0082] In one alternative embodiment, reporting abnormal information through events based on the abnormal event reporting method includes any of the following:

[0083] Notify the interrupt of the abnormal event corresponding to the abnormal information to the basic input / output system of the manageable control processor through the interrupt controller of the data processor, and report the abnormal information corresponding to the abnormal event to the management node on the basic event controller through the basic input / output system of the manageable control processor. The abnormal information in the management node is available for users to query.

[0084] Notify the interrupt of the abnormal event corresponding to the abnormal information to the extensible firmware interface or the basic input / output system of the central processing unit subsystem module through the interrupt controller of the data processor, and after collecting and processing the abnormal information through the extensible firmware interface or the basic input / output system of the central processing unit subsystem module, report the abnormal information corresponding to the abnormal event through the software delegation abnormal interface method, and persistently store the abnormal information through the kernel tracing tool.

[0085] Read the abnormal information stored in the RAS node corresponding to the internal module in a timed polling manner through the kernel-mode driver or the user-mode driver of the central processing unit kernel system, report the abnormal information, and persistently store the abnormal information through the kernel tracing tool.

[0086] Among them, in combination with Figure 3 as shown Figure 3It is a flowchart of reporting RAS interrupts of the data processor DPU scenario to the Baseboard Management Controller (BMC) in an embodiment. In this embodiment, the RAS Node (RAS node) notifies the interruption of the RAS Event corresponding to the exception information to the BIOS (Basic Input-Output System) of the MCP (Manageability Control Processor). The MCP BIOS reports it to the management node on the BMC through IPMI (Intelligent Platform Management Interface). Users and operation and maintenance personnel can view the RAS exception records of the data processor DPU chip through the BMC.

[0087] Among them, the MCP can be regarded as a small core in the data processor DPU, and this small core can report the exception information to the management node on the BMC. The management node on the BMC can provide query interfaces for the exception information uploaded by the RAS node, etc., to facilitate users to query the exception information.

[0088] Among them, combined with Figure 4 shown, Figure 4 It is a flowchart of reporting RAS interrupts of the data processor DPU scenario to the Unified Extensible Firmware Interface (UEFI) in an embodiment. In this embodiment, the RAS Node notifies the interruption of the RAS Event corresponding to the exception information to the Unified Extensible Firmware Interface or the Basic Input-Output System UEFI BIOS in the Central Processing Unit Subsystem Module (CSUB). After the UEFI BIOS completes information collection and processing, it reports the exception event to the kernel module, such as the Linux OS, through the SDEI (Software Delegated Exception Interface) method. Specifically, it can report the exception event to the kernel module through the GHES Driver (Generic Hardware Error Source Driver) through the SDEI method. Subsequently, the exception information in the kernel module is persistently stored in the kernel user state through the kernel tracing tool for users to query. For example, the exception information in the kernel module can be persistently stored in the kernel user state through the kernel tracing tool in the daemon thread rasdaemon.

[0089] Among them, combined with Figure 5 shown, Figure 5It is a flowchart of the kernel-mode or user-mode driver in the RAS of the data processor DPU scenario in an embodiment for periodically polling RAS exception information. In this embodiment, the operating system of the central processing unit, such as the kernel-mode driver or user-mode driver in LinuxOS, polls the exception record register of the internal module IPCore of the data processor DPU in a timer polling manner to complete the collection and reporting process of exception information. Subsequently, the exception information in the kernel module is persistently stored in the kernel user mode through a kernel tracing tool for the user to query. For example, the exception information in the kernel module can be persistently stored in the kernel user mode by the daemon thread rasdaemon through the kernel tracing tool.

[0090] For the above-mentioned method for monitoring exceptions of the internal modules of the data processor, each internal module of the data processor is equipped with a RAS node. In this way, when an internal module exception is detected through the RAS node, the exception information of the internal module storing the exception is stored, and the exception event reporting method corresponding to the RAS node is obtained; based on the exception event reporting method, the exception information is reported through the event. In this way, the data processor side can implement the fault detection and reporting of the internal modules of the data processor, which helps to improve the maintainability and stability of the data processor.

[0091] In one optional embodiment, before obtaining the exception reporting method corresponding to the RAS node, it further includes: configuring the exception event reporting method through the extensible firmware interface or basic input / output system of the central processing unit subsystem module; storing the exception event reporting method in the RAS control node of the data processor; obtaining the exception event reporting method corresponding to the RAS node includes: obtaining the exception event reporting method corresponding to the RAS node from the RAS control node.

[0092] Combined with Figure 6 shown, Figure 6 It is a flowchart of the reporting method configuration step in an embodiment. In this embodiment, the user configures the reporting method of the exception event RAS Event in the UEFI (Unified Extensible Firmware Interface) / BIOS (Basic Input-Output System) of the CSUB (CPU subsystem) module, and stores the reporting method of the exception event RAS Event in the RAS Controller (RAS control node) of the data processor DPU. Thus, when the interrupt controller reports the corresponding exception event, the reporting method can be obtained from the RAS Controller (RAS control node) of the data processor DPU.

[0093] Specifically, combined with Figure 7 ,Figure 7 It is a schematic diagram of the reporting process in an embodiment. In this embodiment, the RAS node of the internal module of the data processor DPU detects abnormal information, formats and stores it in the abnormal record register, and then notifies the RAS Controller (RAS control node) in the form of an event to obtain the reporting method from the RAS Controller (RAS control node), and further controls the interrupt controller GIC of the data processor DPU to report the corresponding abnormal information according to the reporting method.

[0094] In the above embodiment, the configuration of the reporting method can be realized, and the corresponding abnormal information is reported based on the configured reporting method, which is more flexible and convenient.

[0095] In one optional embodiment, the method further includes: when the RAS node detects an internal module abnormality, determining the type of the abnormality; when the type of the abnormality is a correctable abnormality, triggering a fault handling interrupt, where the fault handling interrupt is used to report the abnormal information corresponding to the abnormality; when the type of the abnormality is an uncorrectable abnormality, performing at least one of the following operations: triggering an error recovery interrupt, where the error recovery interrupt is used to report the abnormal information corresponding to the abnormality; returning an in-band error response to the processing unit, where the in-band error response is used to instruct the processing unit to report a synchronous external data abort signal or an asynchronous error interrupt signal.

[0096] For easy understanding, in combination with Figure 8 as shown in Figure 8 It is a schematic structural diagram of the RAS node in an embodiment. In this embodiment, the RAS node includes three parts. The first part is the in-band error, the second part is the error interrupt, and the third part is the abnormal record register.

[0097] When the RAS node detects an abnormality, it records the abnormality and generates an error handling interrupt sent to the interrupt controller GICController. In the DPU scenario, the RAS node records all errors in the abnormal record register to implement error recovery and error handling. The specific definition of the abnormal information can be referred to above.

[0098] Among them, for correctable errors, the RAS node will correct the errors. For example, for single-bit ECC of RAM, if the error cannot be corrected, the RAS Node will postpone the error handling by setting the Posions flag.

[0099] In the case where the type of the exception is a correctable exception, a Fault Handling Interrupt (FHI) is triggered, and the Fault Handling Interrupt is used to report the exception information corresponding to the exception. Among them, if the component implements a CE counter, a Fault Handling Interrupt is generated only when the CE counter overflows. For example, each time an exception occurs, the value of the CE counter is incremented by 1, and a Fault Handling Interrupt is generated only when the CE counter overflows.

[0100] In the case where the type of the exception is an uncorrectable exception, at least one of the following operations is performed: triggering an Error Recovery Interrupt, where the Error Recovery Interrupt is used to report the exception information corresponding to the exception; returning an in-band error response to the Processing Element (PE, also referred to as the accessed Processing Element), where the in-band error response is used to instruct the Processing Element to report a synchronous external data abort signal or an asynchronous error interrupt signal, and the generation of the signal by the Processing Element after receiving the in-band error response depends on the implementation of the Processing Element and is not specifically limited herein.

[0101] Continue to combine Figure 8 , where the exception vector includes a synchronous data abort signal vector, an asynchronous error interrupt signal, an Error Recovery Interrupt, and a Fault Handling Interrupt, and the exception vector is generated and stored in the arm core cluster after the software processes the exception information.

[0102] In one optional embodiment, the exception information of the internal module storing the exception includes: storing the exception information of the internal module of the exception based on a standard error record, where the standard error record includes a status register for storing a common status field and an optional address register for storing the address where the error occurs; or the standard error record includes a status register for storing a common status field, an optional address register for storing the address where the error occurs, and a custom register for storing auxiliary information.

[0103] Among them, in the DPU scenario of the present application, a standard error record and an access error record are defined as mechanisms for system registers or memory-mapped components.

[0104] The above-mentioned exception record register can be based on a standard error record, and the standard error at least includes a status register for storing a common status field and an optional address register for storing the address where the error occurs; optionally, the standard error record can also include a custom register for storing auxiliary information.

[0105] Among them, the status register ERR <n>STATUS, for common status fields such as type of error and gross characteristics.

[0106] Optional address register ERR <n>ADDR is used to store the address where an error occurs.

[0107] The status register customized for the data processor DPU scenario is called ERR <n>MISC <m>These custom status registers are used to store auxiliary information, which is used to identify the FRU, locate errors within the FRU, and optionally correct error counters for software to poll the correction error rate.

[0108] In some alternative embodiments, each exception message records 4 ERRs <n>MISC <m>: ERR <n>MISC0, ERR <n>MISC1, ERR <n>MISC2 and ERR <n>MISC3. In other embodiments, ERR in each abnormal information record can be set based on differences in DPU chips and the like <n>MISC <m>The quantity.

[0109] In one alternative embodiment, the present application further provides an abnormal monitoring method for an internal module of a data processor. Specifically, in combination with Figure 9 As shown, when applied to a central processing unit, the method includes:

[0110] S902: Receive and store the abnormal information reported by the RAS node. The abnormal information is reported by the RAS node based on the abnormal monitoring method for the internal module of the data processor in any of the above embodiments.

[0111] In combination with Figure 10 As shown, the RAS node reports abnormal information through an interrupt, so that the subsequent RAS detection module can read the register information of the RAS node and notify the RAS processing module. The RAS processing module completes abnormal classification processing according to the severity level of the abnormal information and notifies the abnormal information to the RAS storage module. Subsequently, the RAS storage module completes the persistent storage of RAS information and provides it for user query.

[0112] Among them, the reporting method of the abnormal information includes at least one of three types. Thus, receiving and storing the abnormal information reported by the RAS node includes at least one of the following: receiving and storing the abnormal information through the management node on the basic event controller, and the abnormal information in the management node is provided for user query. Receiving an abnormal event through the extensible firmware interface or the basic input / output system of the central processing unit subsystem module, and storing the abnormal information into the cross-platform error report based on the abnormal event; invoking the event processing function hooked by the hardware error driver (such as the general hardware error driver in the above text) through the software delegated abnormal interface method to read and process the abnormal processing information, and persistently storing the abnormal information through the kernel tracing tool. Reading the abnormal information stored in the RAS node corresponding to the internal module in a timed polling manner through the kernel-mode driver or the user-mode driver, reporting the abnormal information, and persistently storing the abnormal information through the kernel tracing tool.

[0113] Specific limitations on the three methods can be referred to above and will not be elaborated here.

[0114] S904: Return the corresponding abnormal information based on the user request.

[0115] In one alternative embodiment, returning the corresponding abnormal information based on the user request includes: returning the corresponding abnormal information through a subscription method; and / or returning the corresponding abnormal information through the RAS interrupt method.

[0116] Specifically, the virtual machine VM can subscribe to the exception information of the data processor DPU in a subscription manner, such as through a socket, or it can also be notified by injecting a RAS interrupt into the virtual machine VM by the data processor DPU to implement the query of exception information.

[0117] For ease of understanding, in combination with Figure 11 as shown in Figure 11 a schematic diagram of the exception monitoring method of the internal module of the data processor corresponding to the second reporting method in an embodiment. In this embodiment, the exception monitoring method of the internal module of the data processor is applied to the DPU processor, and the central processing unit subsystem is taken as an example of the ARM system for illustration, which specifically includes:

[0118] An exception signal is detected by the RAS node in the data processor DPU and connected to the interrupt controller (GIC controller) of the data processor DPU through an interrupt line. The interrupt controller is connected to the exception controller of the CPU core inside the data processor DPU. The RAS node is an abstraction of an internal module IP core of a DPU. The internal module IP core includes CSUB, DPE, VPE, RPE, HAC, PSUB, etc. The data processor DPU can perform exception monitoring after startup.

[0119] When the RAS Node detects an exception, it records the exception and generates an exception handling interrupt sent to the interrupt controller. In the data processor DPU scenario, the RAS Node records all exceptions into the exception record register to implement exception recovery and exception handling.

[0120] The RAS Node includes an optional in-band error. When the processing unit PE receives an in-band error response, the RAS Node will generate a synchronous external data abort (SEA) or an asynchronous SError interrupt (SEI).

[0121] The exception signal is transmitted to the interrupt controller of the data processor DPU through the RAS node in the data processor DPU, and the exception is reported by the interrupt controller.

[0122] Among them, the reporting method of the exception information can be configured by the user using the UEFI or BIOS configuration menu; the configuration method includes enabling or disabling the RAS node, and configuring the reporting method of the RAS Node. The reporting methods include at least one of reporting to the MCP through an interrupt, reporting to the ATF through an interrupt, and configuring the polling method.

[0123] The notification mechanism of ATF (ARM Trusted Firmware) and Linux OS adopts SDEI (Software Delegated Exception Interface). The interaction process of initialization process SDEI can be combined with Figure 12 As shown, Figure 12 EL3 in can be equivalent to Figure 11 ATF in it can interact with Linux OS through SMC (Secure Monitor Call) or HVC (Hypervisor Call), where SMC and HVC are two instructions in ARM architecture, used to implement security extension and virtualization extension respectively. Their main function is to switch and communicate between different execution environments (such as secure world and non-secure world, virtual machine and hypervisor).

[0124] Software Delegated Exception (SDE) is a mechanism that can preempt other exceptions and exclusive mechanisms to pass special system events to the operating system (or hypervisor). After receiving system events, the firmware (such as UEFI) uses SDEI to notify the Normal world and execute the registered function handler.

[0125] Combination Figure 12 ,In the initialization phase, the SDEI client binds a non-secure interrupt, and the SDEI Dispatch program returns a platform dynamic event number.,Then the client registers a handler for the event and enables the event,unmasking all events on the current processing unit PE.

[0126] SDEI can support OS and / or hypervisor to perform the following operations: subscribe to and process system events, shield system events, migrate system event processing to different processing units PE, add or delete processing units PE involved in event processing, convert existing interrupts into SDEI event sources, generate software events and notify SDEI client.

[0127] After initialization, ATF receives an interrupt from the interrupt controller in the data processor DPU and jumps to the MM Driver in S-EL0 for processing.

[0128] The MM Driver collects specific exception information of the corresponding module according to the interrupt type. The MM Driver creates a CPER (Cross-Platform Error Reporting) based on the event identifier Event Id corresponding to the interrupt and writes it into the corresponding GHES (Generic Hardware Error Source) table. After the MM Driver finishes execution, it returns to the ATF.

[0129] The SDEI Dispatcher in the ATF notifies the SDEI Driver (equivalent to the SDEI client) in the kernel OS through SDEI Notify. The SDEI Driver calls back the event handling function Handler hooked by the GHES Driver for processing. Among them, the exception information in the kernel GHES is read into the GHES error list, and the execution of the error handling function IRQ_WORK is triggered. Then, it notifies the ATF that the current interrupt handling is completed. The GHES Driver of the OS calls the target function one by one according to the GHES error list, such as ghes_do_proc for processing. If it is a Memory type error, it reports it to the EDAC Driver for processing, and other errors are default recorded in ftrace.

[0130] The RASdaemon in the user space records the RAS event into the database and reports it to the RAS collector. The virtual machine VM can subscribe to the exception information of the data processor DPU in a subscription manner, such as through a socket, or it can also be notified by injecting a RAS interrupt into the virtual machine VM by the data processor DPU to implement the query of exception information.

[0131] In the above embodiments, in the RAS fault detection and fault repair of the data processor DPU scenario, when an exception occurs in the internal module of the data processor DPU, the exception information is sent to the kernel OS Kernel of the data processor DPU through interrupts, polling, event reporting, etc. The driver program in the kernel OS Kernel records the fault information in ftrace. The user space program rasdaemon on the data processor DPU side obtains it by monitoring ftrace, or it can also inject a ras interrupt into the VM to achieve the RAS perception of the virtual machine VM.

[0132] It should be understood that although the steps in the flowcharts involved in the above embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0133] Based on the same inventive concept, an embodiment of the present application further provides an abnormal monitoring device for an internal module of a data processor for implementing the abnormal monitoring method of the internal module of the data processor involved above. The implementation solution provided by this device to solve the problem is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the abnormal monitoring device for the internal module of the data processor provided below can refer to the limitations on the abnormal monitoring method of the internal module of the data processor in the above text, and will not be repeated here.

[0134] In an exemplary embodiment, as Figure 13 shown, an abnormal monitoring device for an internal module of a data processor is provided, including: a storage module 1301, a reporting method acquisition module 1302, and a reporting module 1303, where:

[0135] The storage module 1301 is configured to store the abnormal information of the abnormal internal module when an internal module abnormality is detected through the RAS node;

[0136] The reporting method acquisition module 1302 is configured to acquire the abnormal event reporting method corresponding to the RAS node;

[0137] The reporting module 1303 is configured to report the abnormal information through an event based on the abnormal event reporting method.

[0138] In one of the optional embodiments, the above device further includes: a configuration module, configured to configure the abnormal event reporting method through the extensible firmware interface or the basic input / output system of the central processor subsystem module; and store the abnormal event reporting method in the RAS control node in the data processor.

[0139] The above reporting method acquisition module 1302 is specifically configured to acquire the abnormal event reporting method corresponding to the RAS node from the RAS control node.

[0140] In one alternative embodiment, the above reporting module 1303 is configured to report exception information through an event by any of the following exception event reporting methods: notify the interrupt of the exception event corresponding to the exception information to the basic input / output system of the manageable control processor through the interrupt controller of the data processor, and report the exception information corresponding to the exception event to the management node on the basic event controller through the basic input / output system of the manageable control processor, where the exception information in the management node is for user query; notify the interrupt of the exception event corresponding to the exception information to the extensible firmware interface or the basic input / output system of the central processing unit subsystem module through the interrupt controller of the data processor, and after collecting and processing the exception information through the extensible firmware interface or the basic input / output system of the central processing unit subsystem module, report the exception information corresponding to the exception event through the software delegation exception interface method, and persistently store the exception information through the kernel tracing tool; read the exception information stored in the RAS node corresponding to the internal module at regular intervals through the kernel-mode driver or the user-mode driver of the central processing unit kernel system, report the exception information, and persistently store the exception information through the kernel tracing tool.

[0141] In one alternative embodiment, the above device further includes an interrupt generation module, configured to determine the type of exception when an internal module exception is detected in the RAS node; when the type of exception is a correctable exception, trigger a fault handling interrupt, where the fault handling interrupt is used to report the exception information corresponding to the exception; when the type of exception is an uncorrectable exception, perform at least one of the following operations: trigger an error recovery interrupt, where the error recovery interrupt is used to report the exception information corresponding to the exception; return an in-band error response to the processing unit, where the in-band error response is used to instruct the processing unit to report a synchronous external data abort signal or an asynchronous error interrupt signal.

[0142] In one alternative embodiment, the above storage module 1301 is further configured to store the exception information of the abnormal internal module based on a standard error record, where the standard error record includes a status register for storing a common status field and an optional address register for storing the address where the error occurs; or the standard error record includes a status register for storing a common status field, an optional address register for storing the address where the error occurs, and a custom register for storing auxiliary information.

[0143] In an exemplary embodiment, as Figure 14 shown, there is provided an exception monitoring device for an internal module of a data processor, including:

[0144] a receiving module 1401, configured to receive and store the exception information reported by the RAS node, where the exception information is reported by the RAS node based on the exception monitoring method for the internal module of the data processor in any of the above embodiments;

[0145] A return module 1402, configured to return corresponding exception information based on a user request.

[0146] In one optional embodiment, the above-mentioned receiving module 1401 is further configured to receive and store the exception information reported by the RAS node in at least one of the following manners: receiving and storing the exception information through a management node on a basic event controller, and the exception information in the management node is provided for user query; receiving an exception event through a scalable firmware interface or a basic input / output system of a central processing unit subsystem module, and storing the exception information into a cross-platform error report based on the exception event; calling an event processing function hooked by a hardware error driver through a software delegated exception interface to read and process the exception handling information, and persistently storing the exception information through a kernel tracing tool; reading the exception information stored in the corresponding RAS node of an internal module in a timed polling manner through a kernel-mode driver or a user-mode driver, reporting the exception information, and persistently storing the exception information through a kernel tracing tool.

[0147] In one optional embodiment, the above-mentioned return module 1402 is further configured to return the corresponding exception information by way of subscription; and / or return the corresponding exception information by way of a RAS interrupt.

[0148] Each module in the exception monitoring device of the internal module of the above-mentioned data processor can be implemented in whole or in part by software, hardware, and their combination. Each of the above-mentioned modules can be embedded in a processor in a computer device in a hardware form or be independent of the processor, or can be stored in a memory in the computer device in a software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above-mentioned modules.

[0149] In an exemplary embodiment, a chip is provided, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method described in any one of the above-mentioned embodiments are implemented.

[0150] In an exemplary embodiment, a network interface card is provided, which includes the chip described in any one of the above-mentioned embodiments and a plurality of interfaces. Among them, the interfaces include wired / wireless communication interfaces such as PCI / PCIE interfaces, UART interfaces, SPI interfaces, USB interfaces, and network interfaces, and the chip processes data or communicates externally through the interfaces.

[0151] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 15 As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a network interface card, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the network interface card, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. When the computer program is executed by the processor, it implements an abnormal monitoring method for internal modules of a data processor. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0152] Those skilled in the art can understand that Figure 15 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0153] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0154] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0155] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0156] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0157] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0158] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.

[0159] The above-described embodiments merely represent several implementation manners of this application, and their descriptions are relatively specific and detailed. However, it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.< / m> < / n> < / n> < / n> < / n> < / n> < / m> < / n> < / m> < / n> < / n> < / n>

Claims

1. An abnormal monitoring method for an internal module of a data processor, characterized in that Applied to a data processor, the method includes: When an internal module exception is detected by the RAS node, storing exception information of the exception internal module; Obtaining an exception event reporting method corresponding to the RAS node; Based on the exception event reporting method, reporting the exception information through an event.

2. The method according to claim 1, characterized in that Before obtaining the exception reporting method corresponding to the RAS node, it further includes: Configuring an exception event reporting method through the extensible firmware interface or basic input / output system of the central processing unit subsystem module; Storing the exception event reporting method in the RAS control node in the data processor; The obtaining of the exception event reporting method corresponding to the RAS node includes: Obtaining the exception event reporting method corresponding to the RAS node from the RAS control node.

3. The method according to claim 2, wherein The reporting of the exception information through an event based on the exception event reporting method includes any one of the following: Notifying, through the interrupt controller of the data processor, an interrupt of the exception event corresponding to the exception information to the basic input / output system of the manageable control processor, and reporting, through the basic input / output system of the manageable control processor, the exception information corresponding to the exception event to the management node on the basic event controller, where the exception information in the management node is for user query; Notifying, through the interrupt controller of the data processor, an interrupt of the exception event corresponding to the exception information to the extensible firmware interface or basic input / output system of the central processing unit subsystem module, and after collecting and processing the exception information through the extensible firmware interface or basic input / output system of the central processing unit subsystem module, reporting, through a software delegation exception interface method, the exception information corresponding to the exception event, and persistently storing the exception information through a kernel tracing tool; Reading, in a timed polling manner, the exception information stored in the RAS node corresponding to the internal module through the kernel-mode driver or user-mode driver of the central processing unit kernel system, reporting the exception information, and persistently storing the exception information through a kernel tracing tool.

4. The method according to claim 1, wherein The method further includes: When an internal module exception is detected by the RAS node, determining the type of the exception; When the type of the exception is a correctable exception, triggering a fault handling interrupt, where the fault handling interrupt is used to report the exception information corresponding to the exception; When the type of the exception is an uncorrectable exception, performing at least one of the following operations: triggering an error recovery interrupt, where the error recovery interrupt is used to report the exception information corresponding to the exception; returning an in-band error response to the processing unit, where the in-band error response is used to instruct the processing unit to report a synchronous external data abort signal or an asynchronous error interrupt signal.

5. The method according to any one of claims 1 to 4, characterized in that The storing of the exception information of the exception internal module includes: Exception information of an internal module storing exceptions based on a standard error record, where the standard error record includes a status register for storing a common status field and an optional address register for storing the address where an error occurred; or the standard error record includes a status register for storing a common status field, an optional address register for storing the address where an error occurred, and a custom register for storing auxiliary information.

6. A method for monitoring anomalies of internal modules of a data processor, characterized in that, The method includes: Receiving and storing exception information reported by a RAS node, where the exception information is reported by the RAS node based on the exception monitoring method of the internal module of the data processor according to any one of claims 1 to 5; Returning corresponding exception information based on a user request.

7. The method according to claim 6, characterized in that The receiving and storing exception information reported by a RAS node includes at least one of the following: Receiving and storing the exception information through a management node on a basic event controller, where the exception information in the management node is available for user query; Receiving an exception event through an extensible firmware interface or a basic input / output system of a central processing unit subsystem module, and storing the exception information in a cross-platform error report based on the exception event; Invoking an event handling function hooked by a hardware error driver through a software delegated exception interface method to read and process the exception handling information, and persistently storing the exception information through a kernel tracing tool; Reading the exception information stored in the corresponding RAS node of the internal module in a timed polling manner through a kernel-mode driver or a user-mode driver, reporting the exception information, and persistently storing the exception information through a kernel tracing tool.

8. The method according to claim 6, characterized in that, The returning corresponding exception information based on a user request includes: Returning corresponding exception information through a subscription method; and / or Returning corresponding exception information through a RAS interrupt method.

9. An abnormal monitoring device for an internal module of a data processor, characterized in that, Applied to a data processor, the apparatus includes: A storage module for storing exception information of an abnormal internal module when an internal module exception is detected through a RAS node; A reporting method acquisition module for acquiring an exception event reporting method corresponding to a RAS node; A reporting module for reporting the exception information through an event based on the exception event reporting method.

10. An abnormal monitoring device for an internal module of a data processor, characterized in that, The apparatus includes: A receiving module for receiving and storing exception information reported by a RAS node, where the exception information is reported by the RAS node based on the exception monitoring apparatus of the internal module of the data processor according to claim 9; A returning module for returning corresponding exception information based on a user request.

11. A chip, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

12. A network interface card, characterized in that, Including a chip as claimed in claim 11 and a plurality of interfaces, where the chip processes data or communicates externally through the interfaces.

13. A computer device, characterized in that, Including a network interface card as claimed in claim 12, where the network interface card is used to process data or communicate externally.

14. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.