Hygon platform-based fault handling system, method and apparatus, device, and medium
Patent Information
- Application Number
- PCT/CN2025/129163
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-27
- Filing Date
- 2025-10-22
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025129163_01102026_PF_FP_ABST
Abstract
Description
Fault handling system, method, device, equipment and media based on the Hygon platform
[0001] This application claims priority to Chinese Patent Application No. 202510371806X, filed on March 27, 2025, entitled "Fault Handling System, Method, Apparatus, Device and Medium Based on Hygon Platform", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates to the field of computer technology, and in particular to a fault handling system, method, apparatus, device and medium based on the Hygon platform. Background Technology
[0003] The Hygon platform is a comprehensive information technology platform centered on the Hygon series x86 architecture CPU. The RAS (Reliability, Availability, and Serviceability) function of the Hygon platform generally ensures stable instruction execution by the CPU. However, for unrepairable or uncorrectable hardware failures occurring in the system, current error handling solutions struggle to quickly and accurately locate the cause of the error and resolve the fault, thus affecting the normal operation of the equipment. Summary of the Invention
[0004] This disclosure provides a fault handling system, method, apparatus, equipment, and medium based on the Hygon platform.
[0005] According to a first aspect of this disclosure, a fault handling system based on the Hygon platform is provided, comprising:
[0006] The central processing unit is used to send target fault detection signals to the component management unit;
[0007] The component management unit is used to acquire processor error data based on the target fault detection signal, determine at least one target central processing unit associated with the processor error data, acquire component attribute data of each target central processing unit, and determine fault location information based on the component attribute data and the processor error data.
[0008] In one possible implementation, the system further includes a baseboard management controller;
[0009] The component management unit is further configured to determine and generate a fault analysis result based on the fault location information, and send the fault analysis result to the substrate management controller;
[0010] The baseboard management controller is used to determine a fault handling strategy based on the fault analysis results.
[0011] In one possible implementation, the generated fault analysis results include the results of faults in the target memory region; the fault handling strategies include a memory replacement strategy or a power restart strategy.
[0012] In one possible implementation, the baseboard management controller is also used to display the fault analysis results.
[0013] In one possible implementation, the component attribute data of the target central processing unit includes memory address information corresponding to each processing module of the target central processing unit;
[0014] The component management unit is also used to match the memory address information with the processor error data to obtain a matching target memory address as fault location information.
[0015] In one possible implementation, the component management unit is further configured to extract memory address data included in the processor error data; calculate the matching degree between the memory address information corresponding to each processing module and the memory address data; and determine the memory address information with a matching degree exceeding a preset matching degree threshold as fault location information.
[0016] In one embodiment, the component management unit is further configured to identify various processing signals sent by the central processing unit; if each of the processing signals includes a target fault detection signal, it receives message data sent by the central processing unit that generated the target fault detection signal as processor error data.
[0017] According to a second aspect of this disclosure, a fault handling method based on the Hygon platform is provided, comprising:
[0018] In response to the target fault detection signal, acquire processor error data;
[0019] Identify at least one target central processing unit associated with the processor error data;
[0020] Obtain component attribute data for each of the target central processing units;
[0021] Based on the component attribute data and the processor error data, the fault location information is determined.
[0022] In one possible implementation, the method further includes:
[0023] Based on the fault location information, a fault analysis result is generated.
[0024] The fault analysis results are sent to the baseboard management controller so that the baseboard management controller can determine a fault handling strategy based on the fault analysis results.
[0025] In one possible implementation, the generated fault analysis results include the results of faults in the target memory region; the fault handling strategies include a memory replacement strategy or a power restart strategy.
[0026] In one embodiment, the method further includes: the baseboard management controller displaying the fault analysis results.
[0027] In one possible implementation, the component attribute data of the target central processing unit includes memory address information corresponding to each processing module of the target central processing unit;
[0028] The step of determining the fault location information based on the component attribute data and the processor error data includes:
[0029] The memory address information is matched with the processor error data to obtain the matching target memory address as the fault location information.
[0030] In one possible implementation, the step of matching the memory address information with the processor error data to obtain a matching target memory address as fault location information includes:
[0031] Extract the memory address data included in the processor error data;
[0032] Calculate the matching degree between the memory address information corresponding to each processing module and the memory address data;
[0033] Memory address information with a matching degree exceeding a preset matching degree threshold is identified as fault location information.
[0034] In one possible implementation, the method further includes:
[0035] Identify the various processing signals sent by the central processing unit;
[0036] If each of the processing signals includes a target fault detection signal, the message data sent by the central processing unit that generates the target fault detection signal is received as processor error data.
[0037] According to a third aspect of this disclosure, a fault handling device based on the Hygon platform is provided, comprising:
[0038] The data acquisition module is used to acquire processor error data in response to a target fault detection signal; determine at least one target central processing unit associated with the processor error data; and acquire component attribute data of each target central processing unit.
[0039] The fault location module is used to determine the fault location information based on the component attribute data and the processor error data.
[0040] In one embodiment, the fault location module is further configured to determine and generate a fault analysis result based on the fault location information; and send the fault analysis result to the baseboard management controller so that the baseboard management controller can determine a fault handling strategy based on the fault analysis result.
[0041] In one possible implementation, the generated fault analysis results include the results of faults in the target memory region; the fault handling strategies include a memory replacement strategy or a power restart strategy.
[0042] In one possible implementation, the baseboard management controller displays the fault analysis results.
[0043] According to a fourth aspect of this disclosure, an electronic device is provided, comprising:
[0044] A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the methods described herein.
[0045] According to a fifth aspect of this disclosure, a storage medium comprising computer-executable instructions is provided, which, when executed by a computer processor, are used to perform the methods described herein.
[0046] The fault handling system based on the Hygon platform provided in this embodiment allows the central processing unit (CPU) to send a target fault detection signal to the component management unit (BMU). The BMU can then acquire processor error data based on the target fault detection signal, determine at least one target CPU associated with the processor error data, acquire component attribute data for each target CPU, and determine the fault location information based on the component attribute data and the processor error data. In other words, in this embodiment, the BMU can perform in-depth analysis of the error data using the component attribute data of the target CPU, refine the causes of the fault, and thus obtain a more accurate fault location.
[0047] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0048] The above and other objects, features, and advantages of this disclosure will become readily apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. Several embodiments of this disclosure are illustrated in the drawings by way of example and not limitation, in which:
[0049] In the accompanying drawings, the same or corresponding reference numerals indicate the same or corresponding parts.
[0050] Figure 1 shows a schematic diagram of a fault handling system based on the Hygon platform provided in an embodiment of this application;
[0051] Figure 2 shows a schematic diagram of an implementation flow of the fault handling method based on the Hygon platform provided in an embodiment of this application;
[0052] Figure 3 shows a schematic diagram of a fault handling device based on the Hygon platform provided in an embodiment of this application;
[0053] Figure 4 shows a schematic diagram of the composition structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation
[0054] To make the objectives, features, and advantages of this disclosure more apparent and understandable, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0055] Current error handling solutions for unrepairable or uncorrectable hardware faults in the Hygon platform system struggle to quickly and accurately pinpoint the cause of the error and resolve the problem, thus affecting the normal operation of the equipment. Therefore, to accurately locate the cause of the error and resolve the problem, this application provides a fault handling system, method, apparatus, electronic device, and storage medium based on the Hygon platform. The electronic device provided in this application can be a mobile phone, computer, tablet computer, or server, etc.
[0056] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings.
[0057] Figure 1 shows a schematic diagram of a fault handling system based on the Hygon platform provided in an embodiment of this application. As shown in Figure 1, it includes:
[0058] Central processing unit 101 is used to send target fault detection signals to component management unit 102;
[0059] The component management unit 102 is used to acquire processor error data based on the target fault detection signal, determine at least one target central processing unit associated with the processor error data, acquire component attribute data of each target central processing unit, and determine fault location information based on the component attribute data and the processor error data.
[0060] Using the fault handling system based on the Hygon platform provided in this embodiment, the central processing unit 101 can send a target fault detection signal to the component management unit 102. The component management unit 102 can obtain processor error data based on the target fault detection signal, determine at least one target central processing unit associated with the processor error data, obtain component attribute data for each target central processing unit, and determine fault location information based on the component attribute data and the processor error data. That is, in this embodiment, the component management unit can perform in-depth analysis of the error data through the component attribute data of the target central processing unit, refine the cause of the fault, and thus obtain a more accurate fault location.
[0061] In this embodiment, the central processing unit can be a CPU from the Hygon platform, such as the Hygon 7000 series CPU, Hygon 5000 series CPU, or Hygon 3000 series CPU. The component management unit can be a BIOS (Basic Input / Output System) from the Hygon platform.
[0062] In this embodiment, the component management unit can detect the status of various hardware components in the system, such as the power supply, hard drive, optical drive, keyboard, network adapter, and memory, to determine whether each hardware component is working properly. For example, when the device is powered on, the component management unit can detect the power-on status of the power supply to determine that the power supply is providing normal power. If the system experiences abnormal conditions such as memory corruption, power failure, hard drive track damage, or system communication interruption, each hardware component can send a fault alarm interrupt signal to the CPU. After recognizing the fault alarm interrupt signal, the CPU can jump to the location corresponding to the fault alarm interrupt signal, collect fault data, generate a fault message, and send a target fault detection signal to the component management unit 102. The target fault detection signal may include a CPU identifier and fault indication information. After receiving the target fault detection signal, the component management unit 102 can locate the corresponding CPU based on the CPU identifier and hardware fault indication information, and collect a fault message from the CPU as processor error data.
[0063] In this embodiment of the disclosure, the component management unit can detect the status of each hardware component when the device is powered on, or it can detect the status of each hardware component every first preset period. The first preset period can be set to 10 minutes or 20 minutes, etc.
[0064] In this embodiment of the disclosure, the component management unit is further configured to identify various processing signals sent by the central processing unit; if each of the processing signals includes a target fault detection signal, the unit receives message data sent by the central processing unit that generates the target fault detection signal as processor error data.
[0065] In this embodiment, each central processing unit (CPU) in the system can send a processing signal to the component management unit every second preset period. The processing signal may include information such as the CPU's operating status, power consumption, and hardware fault indication information. In this embodiment, the component management unit may also receive the processing signals sent by each CPU every third preset period. The second preset period can be set to 1 minute or 2 minutes, etc. The third preset period can be set to 2 minutes or 3 minutes, etc. If the processing signal received by the component management unit includes hardware fault indication information, the received processing signal can be used as a target fault detection signal. The component management unit can then use the fault message of the target CPU corresponding to the target fault detection signal as processor error data. The fault message may include information such as the faulty hardware type and the CPU corresponding to the faulty hardware.
[0066] In this embodiment of the disclosure, the component attribute data of the target central processing unit may include attribute data characterizing features such as the performance and component architecture of the target central processing unit. For example, the component attribute data may include data such as the type, model, clock frequency, number of cores, cache size, and memory address of the target central processing unit.
[0067] In one possible implementation, the component attribute data of the target central processing unit includes memory address information corresponding to each processing module of the target central processing unit; the component management unit is further configured to match the memory address information with the processor error data to obtain a matching target memory address as fault location information.
[0068] The processing modules of a target CPU may include its cores, cache levels, memory controller, input / output interfaces, and arithmetic logic units (ALUs). Each target CPU can access a different range of memory addresses. The memory address range corresponding to each processing module of the target CPU can be used as the memory address information for that processing module. For example, target CPU A includes core 1, core 2, L1 cache, L2 cache, L3 cache, memory controller, input / output interfaces, and ALUs. The memory address range that target CPU A can access is 0x00000000000000000 to 0x0000007FFFFFFFFF. The memory address range corresponding to core 1 of target CPU A is 0x0000000000000000 to 0x00000000000000F. The memory address range corresponding to Core 2 of target CPU A is 0x000000000000000F to 0x00000000000000FF. The memory address range corresponding to L1 cache of target CPU A is 0x00000000000000FF to 0x0000000000000FFF. The memory address range corresponding to L2 cache of target CPU A is 0x0000000000000FFF to 0x000000000000FFFF. The memory address range corresponding to L3 cache of target CPU A is 0x000000000000FFFF to 0x00000000000FFFFF. The memory address range corresponding to the memory controller of target CPU A is 0x00000000000FFFFF to 0x0000000000FFFFFF. The memory address range corresponding to the input / output interfaces of the target CPU A is 0x0000000000FFFFFF to 0x000000000FFFFFFF. The memory address range corresponding to the arithmetic logic unit of the target CPU A is 0x000000000FFFFFFF to 0x00000000FFFFFFFF. In this embodiment, the component management unit can obtain the memory address range corresponding to each processing module of the CPU through the registers of the target CPU. The memory address range corresponding to each processing module of the CPU refers to the memory address range determined by the lower limit and upper limit of the memory address of each processing module. For example, the component management unit can read the DRAM Base Address register of CPU0 and the DRAM Limit Address register of CPU1.The lower memory address limit A1 and upper memory address limit A2 of CPU0's processing module 01, and the lower memory address limit A3 and upper memory address limit A4 of processing module 02 were read from the DRAM Base Address. The lower memory address limit B1 and upper memory address limit B2 of CPU1's processing module 11, and the lower memory address limit B3 and upper memory address limit B4 of processing module 12 were read from the DRAM Limit Address.
[0069] In this embodiment of the disclosure, the component management unit can determine the fault location information by comparing and matching component attribute data and processor error data. Optionally, memory address information can be matched with processor error data to obtain a matching memory address as the target memory address, and the target memory address can be used as the fault location information. Alternatively, the memory address range corresponding to the processing module corresponding to the target memory address can be used as the fault location information. For example, if the processor error data includes memory address D1, and memory address D1 belongs to the memory address range corresponding to the L1 cache of the target CPU A, then memory address D1 can be used as the fault location information. Alternatively, the memory address range corresponding to the L1 cache of the target CPU A can be used as the fault location information.
[0070] Processor error data can be error messages in a preset format. For example, the format of processor error data can include error types. In this embodiment of the disclosure, the error message can predefine the error types of the hardware. For example, the error type "poison comsumption" represents an error or malfunction caused by memory usage, and the error type "T1" represents an error or malfunction caused by high hardware temperature. The component management unit can query whether the processor error data includes preset error types and locate the fault for the specific error type. For example, if the component management unit finds that the processor error data includes the error type "poison comsumption", it can determine that the device has caused a CPU fault due to memory usage. The component management unit can then match the memory address information of each processing module of the target central manager with the memory address data included in the processor error data, and determine the memory address data in the processor error data that matches the memory address information of each processing module of the target central manager as the target fault data. Alternatively, after finding that the processor error data includes the error type "poison comsumption", the component management unit can determine the memory address data related to the error type "poison comsumption" as the target fault data and record it.
[0071] In this embodiment of the disclosure, since the fault handling system based on the Hygon platform may include multiple central processing units, in order to accurately locate the fault location, when the component management unit finds that the error type "poison comsumption" is included in the processor error data, it can also associate and record the memory address data related to the error type "poison comsumption" with the corresponding central processing unit.
[0072] In this embodiment of the disclosure, the component management unit is further configured to extract memory address data included in the processor error data; calculate the matching degree between the memory address information corresponding to each processing module and the memory address data; and determine the memory address information with a matching degree exceeding a preset matching degree threshold as fault location information.
[0073] In one possible implementation, the memory address data included in the processor error data can be used as the target memory address. The component management unit can calculate, for each processing module, the proportion of target memory addresses belonging to the memory address range corresponding to that processing module, and use this proportion as the matching degree between the memory address information corresponding to that processing module and the memory address data. For example, if the memory address range corresponding to processing module 01 of the target CPU is 0x00000000 to 0x000000FF, and the memory address range corresponding to processing module 02 of the target CPU is 0x00000100 to 0x000001FF, the target memory address includes [0x00000010, 0x00000050, 0x00000180, 0x00000250]. For processing module 01, the memory addresses belonging to the memory address range corresponding to processing module 01 are 0x00000010 and 0x00000050, meaning that 50% of the target memory addresses belong to the memory address range corresponding to processing module 01. For processing module 02, the memory address belonging to the memory address range corresponding to processing module 02 is 0x00000180, meaning that 25% of the target memory addresses belong to the memory address range corresponding to processing module 01. Therefore, the matching degree between the memory address information and memory address data corresponding to processing module 01 is 50%, and the matching degree between the memory address information and memory address data corresponding to processing module 02 is 25%.
[0074] In another possible implementation, the memory address data included in the processor error data can be a memory address range, which can be used as the target memory address range. The component management unit can, for each processing module, calculate the proportion of memory addresses that overlap between the target memory address range and the memory address information corresponding to that processing module, relative to the memory address range corresponding to that processing module, and use this proportion as the matching degree between the memory address information corresponding to that processing module and the memory address data.
[0075] In this embodiment of the disclosure, the preset matching degree threshold can be set according to actual application requirements, for example, set to 70% or 75%.
[0076] In this embodiment, the component management unit can analyze fault information, obtain the memory address range of each processing module of the central processing unit (CPU) by combining the CPU registers, and match the memory address range of each processing module with the processor error data to refine the fault scope to the processing module level. This enables precise fault location, helps managers analyze the cause of the fault more clearly, and resolve the problem in a timely manner, avoiding the cost and resource waste caused by blindly replacing the CPU. Furthermore, the fault handling system provided in this embodiment can automatically locate the fault at the processing module level, avoiding other problems caused by operator error and improving system reliability.
[0077] In this embodiment of the disclosure, the system further includes a baseboard management controller; the component management unit is further configured to determine and generate a fault analysis result based on the fault location information, and send the fault analysis result to the baseboard management controller; the baseboard management controller is configured to determine a fault handling strategy based on the fault analysis result.
[0078] In this embodiment, the component management unit can generate fault analysis results according to a target format based on fault location information. For example, it can generate formatted fault analysis results based on fault type, fault hardware type, and fault memory address. In this embodiment, the component management unit can send the fault analysis results to the baseboard management controller (BMC) via a serial communication protocol (I2C, Inter-Integrated Circuit) bus or a standard interface (IPMI, Intelligent Platform Management Interface) bus.
[0079] The baseboard management controller can determine a fault handling strategy based on the fault analysis results. In one possible implementation, generating the fault analysis results may include the result of a fault in the target memory region; the fault handling strategy may include a memory replacement strategy or a power-on strategy. For example, if the fault location information includes an abnormal memory address range corresponding to the processing module 01 of the CPU 0, the memory address range corresponding to the processing module 01 can be used as the target memory region, and the result of a fault in the target memory region can be generated as the fault analysis result. Then, the baseboard management controller can generate a strategy to replace the memory corresponding to the processing module 01 of the CPU 0 based on the fault analysis result, thereby resolving the problem caused by the memory fault. Alternatively, since the abnormal memory address range corresponding to the processing module 01 may also be caused by other data occupying memory addresses within that memory address range, the baseboard management controller can also generate a power-on strategy based on the fault analysis result, resolving this type of fault by restarting the power. Alternatively, the baseboard management controller can also generate a strategy to replace the CPU 0, resolving the problem caused by the hardware fault by replacing the CPU 0.
[0080] In this embodiment of the disclosure, the baseboard management controller is further configured to display the fault analysis results.
[0081] In this embodiment of the disclosure, the baseboard management controller can display fault analysis results in the form of text or charts on a display, making it easier for staff to understand the cause of the fault. For example, if it is a memory fault of the CPU processor, the baseboard management controller can generate a fault result table, which can display information such as the faulty CPU, the processing module of the faulty CPU, and the memory address range of the processing module, making it easier for staff to intuitively obtain fault information.
[0082] For scenarios where it is inconvenient for staff to obtain fault information through a display screen, the baseboard management controller can also display the fault analysis results via a monitor and in the form of voice broadcast. For example, the baseboard management controller can generate a text version of the fault analysis results, convert the text version into a voice version, and then play the voice version of the fault analysis results through a voice playback device, making it easier for staff to understand the fault information in a timely manner.
[0083] In this embodiment of the disclosure, if the fault problem is not resolved within a preset time period after the baseboard management controller displays the fault analysis results, it may be because the staff did not receive the fault analysis results or ignored the fault analysis results. In order to resolve the fault as soon as possible, the baseboard management controller can generate fault handling alarm prompts and display the fault handling alarm prompts in the form of voice or text.
[0084] The fault handling system based on the Hygon platform provided in this disclosure allows the component management unit to obtain the memory address range of each processing module of the central processing unit by combining the CPU registers. By matching the memory address range of each processing module of the central processing unit with the processor error data, the fault range is refined to the processing module level, enabling precise fault location. This helps managers to analyze the cause of the fault more clearly and resolve the fault problem in a timely manner, avoiding the cost and resource waste caused by blindly replacing the central processing unit. At the same time, it also avoids other problems caused by staff misoperation, thus improving the reliability of the system.
[0085] Figure 2 illustrates a schematic flowchart of an implementation of a fault handling method based on the Hygon platform provided in this application embodiment. This method is applied to a component management unit in a fault handling system based on the Hygon platform. The fault handling system may further include one or more central processing units. As shown in Figure 2, the method includes:
[0086] S201, in response to the target fault detection signal, acquires processor error data.
[0087] In this embodiment, the component management unit can monitor the status of each piece of hardware in the system. If an abnormality occurs in the system, each piece of hardware can send a fault alarm interrupt signal to the CPU. After recognizing the fault alarm interrupt signal, the CPU can jump to the location corresponding to the fault alarm interrupt signal, collect fault data, generate a fault message, and send a target fault detection signal to the component management unit. The target fault detection signal may include a CPU identifier and fault indication information. After receiving the target fault detection signal, the component management unit can locate the corresponding CPU based on the CPU identifier and hardware fault indication information, and collect a fault message from the CPU as processor error data.
[0088] S202, determine at least one target central processing unit associated with the processor error data.
[0089] In this embodiment of the disclosure, the component management unit can identify the CPU identifier to obtain the CPU that caused the fault in the processor error data, and use it as the target CPU associated with the processor error data.
[0090] S203, Obtain component attribute data for each of the target central processing units.
[0091] In this embodiment of the disclosure, the component attribute data may include attribute data characterizing features such as the performance and component architecture of the target central processing unit.
[0092] S204, Based on the component attribute data and the processor error data, determine the fault location information.
[0093] In this embodiment of the disclosure, the memory address that matches the component attribute data and the processor error data can be used as the fault location information.
[0094] Using the fault handling method based on the Hygon platform provided in this disclosure, the component management unit can acquire processor error data based on the target fault detection signal, determine at least one target CPU associated with the processor error data, acquire component attribute data for each target CPU, and determine fault location information based on the component attribute data and the processor error data. That is, in this disclosure, the component management unit can perform in-depth analysis of the error data using the component attribute data of the target CPU, refine the causes of the fault, and thus obtain a more accurate fault location.
[0095] In this embodiment of the present disclosure, the fault handling method further includes: determining and generating a fault analysis result based on the fault location information; and sending the fault analysis result to a baseboard management controller so that the baseboard management controller determines a fault handling strategy based on the fault analysis result.
[0096] In this embodiment of the disclosure, the generation of fault analysis results includes the results of faults in the target memory region; the fault handling strategy includes a memory replacement strategy or a power restart strategy.
[0097] In this embodiment of the disclosure, the fault handling method further includes: the baseboard management controller displaying the fault analysis results.
[0098] In this embodiment of the disclosure, the component attribute data of the target central processing unit includes memory address information corresponding to each processing module of the target central processing unit; the step of determining the fault location information based on the component attribute data and the processor error data includes: matching the memory address information with the processor error data to obtain a matching target memory address as the fault location information.
[0099] In this embodiment of the disclosure, the step of matching the memory address information with the processor error data to obtain a matching target memory address as fault location information includes: extracting memory address data included in the processor error data; calculating the matching degree between the memory address information corresponding to each processing module and the memory address data; and determining the memory address information with a matching degree exceeding a preset matching degree threshold as fault location information.
[0100] In this embodiment of the disclosure, the fault handling method further includes: identifying various processing signals sent by the central processing unit; if each of the processing signals includes a target fault detection signal, receiving message data sent by the central processing unit that generated the target fault detection signal as processor error data.
[0101] Based on the same inventive concept, and according to the fault handling method based on the hypogonal platform provided in the above embodiments of this disclosure, correspondingly, another embodiment of this disclosure also provides a fault handling device based on the hypogonal platform, the structural schematic diagram of which is shown in Figure 3, specifically including:
[0102] The data acquisition module 301 is used to acquire processor error data in response to a target fault detection signal; determine at least one target central processing unit associated with the processor error data; and acquire component attribute data of each target central processing unit.
[0103] The fault location module 302 is used to determine the fault location information based on the component attribute data and the processor error data.
[0104] Using the fault handling device based on the Hygon platform provided in this embodiment, the component management unit can acquire processor error data based on the target fault detection signal, determine at least one target CPU associated with the processor error data, acquire component attribute data for each target CPU, and determine fault location information based on the component attribute data and the processor error data. That is, in this embodiment, the component management unit can perform in-depth analysis of the error data using the component attribute data of the target CPU, refine the causes of the fault, and thus obtain a more accurate fault location.
[0105] In one embodiment, the fault location module 302 is further configured to determine and generate a fault analysis result based on the fault location information; and send the fault analysis result to the baseboard management controller so that the baseboard management controller can determine a fault handling strategy based on the fault analysis result.
[0106] In one possible implementation, the generated fault analysis results include the results of faults in the target memory region; the fault handling strategies include a memory replacement strategy or a power restart strategy.
[0107] In one possible implementation, the baseboard management controller displays the fault analysis results.
[0108] In one embodiment, the component attribute data of the target central processing unit includes memory address information corresponding to each processing module of the target central processing unit; the fault location module 302 is specifically used to match the memory address information with the processor error data to obtain a matching target memory address as fault location information.
[0109] In one possible implementation, the fault location module 302 is specifically used to extract memory address data included in the processor error data; calculate the matching degree between the memory address information corresponding to each processing module and the memory address data; and determine the memory address information with a matching degree exceeding a preset matching degree threshold as fault location information.
[0110] In one embodiment, the fault location module 302 is further configured to identify various processing signals sent by the central processing unit; if each of the processing signals includes a target fault detection signal, the module receives message data sent by the central processing unit that generated the target fault detection signal as processor error data.
[0111] According to embodiments of this disclosure, this disclosure also provides an electronic device and a readable storage medium.
[0112] Figure 4 illustrates a schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0113] As shown in Figure 4, the electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. The RAM 403 may also store various programs and data required for the operation of the electronic device 400. The computing unit 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0114] Multiple components in electronic device 400 are connected to I / O interface 405, including: input unit 406, such as keyboard, mouse, etc.; output unit 407, such as various types of displays, speakers, etc.; storage unit 408, such as disk, optical disk, etc.; and communication unit 409, such as network card, modem, wireless transceiver, etc. Communication unit 409 allows electronic device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0115] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above, such as the fault handling method based on the Hygon platform. For example, in some embodiments, the fault handling method based on the Hygon platform can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 400 via ROM 402 and / or communication unit 409. When the computer program is loaded into RAM 403 and executed by the computing unit 401, one or more steps of the fault handling method based on the Hygon platform described above can be performed. Alternatively, in other embodiments, computing unit 401 may be configured to perform a fault handling method based on the Hygon platform by any other suitable means (e.g., by means of firmware).
[0116] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.
[0117] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0118] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0119] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0120] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0121] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0122] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0123] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means two or more, unless otherwise explicitly specified.
[0124] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.
Claims
1. A fault handling system based on the Hygon platform, characterized in that, include: The central processing unit is used to send target fault detection signals to the component management unit; The component management unit is used to acquire processor error data based on the target fault detection signal, determine at least one target central processing unit associated with the processor error data, acquire component attribute data of each target central processing unit, and determine fault location information based on the component attribute data and the processor error data.
2. The system according to claim 1, characterized in that, The system also includes a baseboard management controller; The component management unit is further configured to determine and generate a fault analysis result based on the fault location information, and send the fault analysis result to the substrate management controller; The baseboard management controller is used to determine a fault handling strategy based on the fault analysis results.
3. The system according to claim 2, characterized in that, The generated fault analysis results include the results of faults in the target memory region; the fault handling strategies include memory replacement strategies or power restart strategies.
4. The system according to claim 2, characterized in that, The baseboard management controller is also used to display the fault analysis results.
5. The system according to claim 1, characterized in that, The component attribute data of the target central processing unit includes the memory address information corresponding to each processing module of the target central processing unit; The component management unit is also used to match the memory address information with the processor error data to obtain a matching target memory address as fault location information.
6. The system according to claim 5, characterized in that, The component management unit is also used to extract memory address data included in the processor error data; and to calculate the matching degree between the memory address information corresponding to each processing module and the memory address data. Memory address information with a matching degree exceeding a preset matching degree threshold is identified as fault location information.
7. The system according to claim 1, characterized in that, The component management unit is also used to identify various processing signals sent by the central processing unit; if each of the processing signals includes a target fault detection signal, it receives the message data sent by the central processing unit that generated the target fault detection signal as processor error data.
8. A fault handling method based on the Hygon platform, characterized in that, A component management unit applied to a fault handling system based on the Hygon platform, the method comprising: In response to the target fault detection signal, acquire processor error data; Identify at least one target central processing unit associated with the processor error data; Obtain component attribute data for each of the target central processing units; Based on the component attribute data and the processor error data, the fault location information is determined.
9. The method according to claim 8, characterized in that, The method further includes: Based on the fault location information, a fault analysis result is generated. The fault analysis results are sent to the baseboard management controller so that the baseboard management controller can determine a fault handling strategy based on the fault analysis results.
10. The method according to claim 9, characterized in that, The generated fault analysis results include the results of faults in the target memory region; the fault handling strategies include memory replacement strategies or power restart strategies.
11. The method according to claim 9, characterized in that, The method further includes: the baseboard management controller displaying the fault analysis results.
12. The method according to claim 8, characterized in that, The component attribute data of the target central processing unit includes the memory address information corresponding to each processing module of the target central processing unit; The step of determining the fault location information based on the component attribute data and the processor error data includes: The memory address information is matched with the processor error data to obtain the matching target memory address as the fault location information.
13. The method according to claim 12, characterized in that, The step of matching the memory address information with the processor error data to obtain a matching target memory address as fault location information includes: Extract the memory address data included in the processor error data; Calculate the matching degree between the memory address information corresponding to each processing module and the memory address data; Memory address information with a matching degree exceeding a preset matching degree threshold is identified as fault location information.
14. The method according to claim 8, characterized in that, The method further includes: Identify the various processing signals sent by the central processing unit; If each of the processing signals includes a target fault detection signal, the message data sent by the central processing unit that generates the target fault detection signal is received as processor error data.
15. A fault handling device based on a marine optical platform, characterized in that, include: The data acquisition module is used to acquire processor error data in response to the target fault detection signal; Identify at least one target central processing unit associated with the processor error data; Obtain component attribute data for each of the target central processing units; The fault location module is used to determine the fault location information based on the component attribute data and the processor error data.
16. The apparatus according to claim 15, characterized in that, The fault location module is further configured to determine and generate fault analysis results based on the fault location information; and send the fault analysis results to the baseboard management controller so that the baseboard management controller can determine a fault handling strategy based on the fault analysis results.
17. The apparatus according to claim 16, characterized in that, The generated fault analysis results include the results of faults in the target memory region; the fault handling strategies include memory replacement strategies or power restart strategies.
18. The apparatus according to claim 16, characterized in that, The baseboard management controller displays the fault analysis results.
19. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as described in any one of claims 1-7.
20. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the method as described in any one of claims 1-7.