Server processing system, method, electronic device and storage medium

By building a communication structure between the interrupt processor, management processor and data processor in the server, state synchronization and data sharing between server processors are achieved, solving the problems of data synchronization delay and hardware-software disconnection, improving the efficiency and accuracy of fault handling, and meeting the response requirements of PCIe 5.0 devices.

CN120407265BActive Publication Date: 2025-09-12INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510898120.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-12
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

The data synchronization delay between processors in the server is high, and the hardware signals and software status are disconnected, resulting in untimely and inefficient fault handling, making it difficult to meet the sub-microsecond fault response requirements of PCIe 5.0 devices.

Method used

A special communication structure is constructed to communicate interrupt signals and status data through preset physical links and PCIe links between the interrupt processor, management processor and data processor. The interrupt processor determines the current interrupt level based on the interrupt signal, dynamically selects the target processor for fault processing, and uses shared storage areas to achieve rapid sharing and synchronization of fault information and status data.

Benefits of technology

It achieves state synchronization and data sharing between processors, quickly responds to and handles faults, reduces the burden on data processors, improves the efficiency and accuracy of fault handling, and meets the sub-microsecond fault response requirements of PCIe 5.0 devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407265B_ABST
    Figure CN120407265B_ABST
Patent Text Reader

Abstract

The present invention discloses a processing system, method, electronic device and storage medium for a server, relating to the technical field of electronic digital data processing, including an interrupt processor, a management processor and a data processor. The interrupt processor, the management processor and the data processor can communicate with each other through corresponding links to realize signal and status data sharing. After the interrupt signal is triggered, the interrupt processor can determine the current interrupt level according to the interrupt signal to determine the target processor, so as to control the target processor to perform fault processing based on the shared information. The technical problem that the data synchronization delay between the processors of the server is high and the hardware signal and software status are disconnected in the related technology, making it difficult to effectively handle the fault when the server fails, is solved. The technical effect of timely status synchronization and data sharing between the processors through a special architecture is achieved, breaking the fault information island and quickly responding to and handling the fault is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic digital data processing, and in particular to a processing system, method, electronic equipment and storage medium of a server. Background Art

[0002] The data processed by servers and the scale of server hardware in data centers are experiencing explosive growth. The field of server management technology has placed increasingly higher demands on the reliability and availability of server components. In related technologies, the server processing architecture has certain defects. Among them, the data synchronization delay between processors in the server is high, and the hardware signal and software status are out of sync, making it difficult to effectively handle faults when a server fails. Improvements are urgently needed. Summary of the Invention

[0003] The present invention provides a server processing system, method, electronic device and storage medium to at least solve the problems in the related art such as high data synchronization delay between processors of the server, disconnection between hardware signals and software status, and difficulty in effectively handling faults when a server fails.

[0004] The present invention provides a processing system for a server, including an interrupt processor, a management processor and a data processor. After the interrupt processor detects an interrupt signal, the interrupt processor communicates the interrupt signal and status data with the data processor based on a preset physical link. The interrupt processor communicates the interrupt signal and status data with the management processor based on a preset access link. The data processor communicates the interrupt signal and status data based on a PCIe link. The interrupt processor determines the current interrupt level of the server based on the interrupt signal, and determines the target processor based on the current interrupt level, so as to control the target processor to perform corresponding processing actions on at least one device involved in the interrupt signal based on the interrupt signal or status data.

[0005] The present invention also provides a server processing method, including: after the interrupt processor detects an interrupt signal, determining the current interrupt level of the server according to the interrupt signal; determining the target processor based on the current interrupt level to control the target processor to perform corresponding processing actions on the device based on the interrupt signal or status data.

[0006] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned server processing methods when executing the computer program.

[0007] The present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned server processing methods are implemented.

[0008] The present invention also provides a computer program product, comprising a computer program, which implements the steps of any of the above-mentioned server processing methods when executed by a processor.

[0009] Through the present invention, since the interrupt processor, management processor and data processor can communicate through corresponding links, thereby realizing signal and status data sharing, after the interrupt signal is triggered, the interrupt processor can determine the current interrupt level according to the interrupt signal, and determine the target processor according to the current interrupt level, so as to control the target processor to perform fault processing based on the shared information, thereby solving the technical problems in the related technology that the data synchronization delay between the processors of the server is high, the hardware signal and the software status are disconnected, and it is difficult to effectively handle the fault when the server fails. The present invention achieves the technical effect of timely status synchronization and data sharing between the processors through a special architecture, breaking the fault information island, and quickly responding to and handling the fault. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0011] Figure 1 A schematic structural diagram of a processing system of a server provided according to an embodiment of the present invention;

[0012] Figure 2 A schematic structural diagram of a processing system of a server provided according to an embodiment of the present invention;

[0013] Figure 3 The present invention provides a flowchart of a processing method of a server according to an embodiment of the present invention.

[0014] Among them, 10 is the server processing system, 100 is the interrupt processor, 200 is the management processor, 201 is the shared storage area, 300 is the data processor, 401 is the interface card 1, 402 is the interface card 2, 403 is the interface card 3, 404 is the interface card 4, and 500 is the software. DETAILED DESCRIPTION

[0015] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0016] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.

[0017] It is understandable that the related art has the following defects:

[0018] Communication protocol fragmentation: The CPU (Central Processing Unit) and BMC (Baseboard Management Controller) use the IPMI (Intelligent Platform Management Interface) protocol (software protocol stack), while the BMC and CPLD (Complex Programmable Logic Device) use I2C / GPIO (low-speed hardware interface). The lack of a direct connection between the CPU and CPLD results in high fault status synchronization delays, and untimely fault handling often leads to fault propagation. Furthermore, hardware signals and software status are disconnected. While the CPU can access the PCIe (Peripheral Component Interconnect Express) configuration space, out-of-band hardware signals are detected in the CPLD. This lack of data synchronization affects the efficiency and accuracy of fault handling.

[0019] Information synchronization is inefficient, and server management and fault repair rely on out-of-band low-speed bus signal polling to access device data, with a delay of more than 120ms, which cannot meet the sub-microsecond fault response requirements of PCIe5.0 devices.

[0020] Fault isolation often uses a single BMC management domain decision-making model, which may lead to "false isolation" (for example, a temporary hardware link error is misidentified as a hardware failure), resulting in increased service interruption rate.

[0021] The fault isolation mechanism lacks refined management. The server interface card is the core channel for service external data interaction. Traditional methods only propose relevant isolation measures for power failures, lacking refined classification and repair management of faults. At the same time, server operating status, fault history information, etc. are not included in the fault repair isolation scope.

[0022] In order to solve the above technical problems, the embodiments of the present invention can realize the rapid sharing of fault information and status data by constructing a special communication structure, and after triggering the interrupt signal, determine different target processors according to different priorities, so as to combine data such as fault history information to improve the fault processing efficiency and fault processing accuracy.

[0023] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0024] like Figure 1 As shown, an embodiment of the present invention provides a server processing system 10 , including an interrupt processor 100 , a management processor 200 and a data processor 300 .

[0025] After the interrupt processor 100 detects an interrupt signal, the interrupt processor 100 communicates the interrupt signal and status data with the data processor 300 based on a preset physical link. The interrupt processor 100 communicates the interrupt signal and status data with the management processor 200 based on a preset access link. The data processor 300 communicates the interrupt signal and status data based on a PCIe link.

[0026] The interrupt handler 100 determines the current interrupt level of the server according to the interrupt signal, and determines the target processor based on the current interrupt level to control the target processor to perform corresponding processing actions on at least one device involved in the interrupt signal based on the interrupt signal or status data.

[0027] In actual implementation, the embodiment of the present invention can be applied to a typical server architecture, including an interrupt processor 100 , a management processor 200 , and a data processor 300 .

[0028] Interrupt handler 100 may be a CPLD, a type of programmable logic device (PLD). CPLDs are digital integrated circuits with user-definable logic functions. Their core structure consists of programmable logic macrocells and a programmable interconnect matrix. Users can input design logic using hardware description language or schematics. After generating a target file, the code is burned into the CPLD chip via a download cable. Programming is based on E2PROM or FLASH memory, supporting non-volatile storage. Programs are not lost after power failures, and programming cycles can reach tens of thousands.

[0029] The management processor 200 may be a BMC, which is a dedicated microcontroller independent of the host system and integrated on the motherboard of hardware such as servers and network devices. It is used to monitor and manage various parameters of the physical environment (such as temperature, voltage, and fan speed) and hardware status (such as power supply, CPU, and memory). Even if the main CPU or operating system crashes, the BMC can still work.

[0030] The data processor 300 may be a CPU, which is a main computing unit that executes an operating system and application program codes and processes core business data.

[0031] An interrupt signal is an electronic signal sent by a hardware device (such as a network card, hard drive controller, or fan fault sensor) or software when it requires the processor's immediate attention. It interrupts the processor's current task and requires it to handle an urgent event (such as data arrival, an error, or excessive temperature).

[0032] The interrupt handler 100 may detect an interrupt signal through a plurality of interface cards. After detecting the interrupt signal, the interrupt handler 100 may implement rapid data sharing through a pre-built communication architecture.

[0033] The interrupt processor 100 and the management processor 200 can communicate state data and share interrupt signals with each other through direct access methods such as I3C and DMA.

[0034] Between the interrupt handler 100 and the data processor 300, state data can be communicated and interrupt signals can be shared through a physical link such as LPC.

[0035] The management processor 200 and the data processor 300 can communicate status data and share interrupt signals with each other through a PCIe link.

[0036] Through the above links, the three processors can complete real-time synchronization of status, thereby facilitating subsequent fault response.

[0037] Furthermore, in addition to interrupt delivery, the interrupt handler 100 can also determine the level of the current interrupt according to the received interrupt signal, and dynamically select a target processor based on the determined interrupt level:

[0038] For example, for interrupt signals that require urgent processing, are high-level, or require computational processing, data processor 300 may be selected as the target processor. For interrupt signals that do not require urgent processing, are management-related, or occur when data processor 300 is unavailable (or under heavy load), management processor 200 may be selected as the target processor.

[0039] After receiving the interrupt signal and related status data, the selected target processor (management processor 200 or data processor 300) executes a predefined processing action.

[0040] Based on the above architecture, the embodiment of the present invention can reduce the burden on the data processor 300, i.e., the CPU: the interrupt processor 100, i.e., the CPLD, serves as a dedicated interrupt controller, responsible for preliminary aggregation, filtering, and classification of interrupts, thereby avoiding a large number of original interrupt signals from directly impacting the CPU, significantly reducing the CPU's interrupt processing overhead (context switching), and allowing the CPU to focus more on core computing tasks.

[0041] Based on the interrupt level, the CPLD intelligently decides whether to hand the interrupt over to the CPU for fast processing or to the BMC for out-of-band management, thereby improving the overall efficiency and pertinence of interrupt processing.

[0042] Low-latency core communication: The physical link between the CPLD and the CPU is designed for fast interrupt signal transmission, ensuring that high-priority interrupts can receive timely response from the CPU.

[0043] The preset access link between the CPLD and BMC uses a stable bus designed specifically for management, ensuring reliable transmission of out-of-band management information.

[0044] PCIe devices communicate directly with the CPU through PCIe links, fully utilizing the high bandwidth and low latency characteristics of PCIe to handle device interrupts and data transmission.

[0045] Optionally, in one embodiment of the present invention, the management processor 200 has a shared memory area, wherein the interrupt signal detected by the interrupt processor 100 and the status data of the interrupt processor 100 are written to the shared memory area through a preset access link.

[0046] The abnormal status signal detected by the data processor 300, the status data of the data processor 300, and the server operation data collected by the data processor 300 are written to the shared storage area through the PCIe link.

[0047] The monitoring abnormality signal detected by the management processor 200, the status data of the management processor 200, and the hardware status data collected by the management processor 200 are written into the shared storage area.

[0048] In addition to interrupt handler 100 detecting interrupt signals, management processor 200 and data processor 300 can also perform corresponding fault detection. For example, management processor 200 can detect interface card monitoring information (temperature, power consumption), while data processor 300 can detect interface card PCIe link status information (bit errors, bandwidth, and rate). When any processor detects an abnormal signal, such as an interrupt signal, abnormal status signal, or abnormal monitoring signal, it can directly write the corresponding signal and status data to the shared memory area, allowing for subsequent access to relevant data during fault handling. Alternatively, the fault corresponding to the signal can be recorded so that subsequent processing strategies can be adjusted based on historical information.

[0049] As a possible implementation method, a shared PCIe device configuration space memory solution is adopted between the data processor 300 and the management processor 200 using a PCIe link interconnection. The management processor 200, as an EP (Endpoint Device) device, can be directly read and written by the BIOS (Basic Input / Output System) to access the shared memory, that is, the shared storage area.

[0050] The interrupt handler 100 may write data into the shared memory by accessing the link, such as by direct access.

[0051] That is, in this embodiment of the present invention, all key events and status data from three independent sources (hardware interrupts from the interrupt processor 100, system operation information from the data processor 300, and hardware monitoring data from the management processor 200) can be aggregated into the same shared storage area of ​​the management processor 200, greatly facilitating problem diagnosis, root cause analysis, and historical tracing.

[0052] Optionally, in one embodiment of the present invention, the interrupt processor 100, the management processor 200 and the data processor 300 write the interrupt signal and the corresponding status data to the shared storage area based on a preset atomic operation to prohibit the interrupt processor 100, the management processor 200 and the data processor 300 from writing data at the same time.

[0053] Atomic operations are defined as indivisible read and write memory operations that are predefined at the hardware or underlying firmware level. An atomic operation cannot be interrupted by operations from other processors or threads during execution. It either succeeds completely (all intended modifications take effect) or fails completely (the memory state remains unchanged), without any intermediate states or partial writes.

[0054] In an embodiment of the present invention, atomic operations serve as the underlying mechanism for mutual exclusion locks. When the interrupt processor 100 (CPLD), management processor 200 (BMC), or data processor 300 (CPU) needs to write to a shared memory area, they must "declare" write permission by executing this preset atomic operation. The core purpose of this atomic operation is to ensure that only one processor can successfully obtain write permission and perform the actual write operation at a time. This strictly prevents the CPLD, BMC, and CPU (or any two of them) from writing to the same block or associated area of ​​the shared memory area at the same time. This prevents concurrent write conflicts and ensures data integrity and consistency.

[0055] For example, if no other processor currently holds the lock (that is, a specific flag bit in the shared memory area is in the "free" state), the processor performing the atomic operation will successfully set the flag bit to "locked" and immediately obtain write permission and start executing its data writing operation.

[0056] If another processor currently holds the lock (the flag is already in the "locked" state), the processor attempting to perform the atomic operation will fail and will not be able to obtain write access. It must wait (perhaps through polling or interrupt notification) until it detects that the lock has been released (the flag changes back to "free").

[0057] Writing and "unlocking": A processor that has been granted write permission can safely write its data to the target location in the shared memory area. After completing the write, it must perform another predefined atomic operation (or set a specific flag bit) to release the lock (set the flag bit to "free"), indicating that other waiting processors can now attempt to acquire the lock and write.

[0058] Optionally, in one embodiment of the present invention, the interrupt handler 100 includes: an encoding unit, a matching unit, and a determining unit.

[0059] The encoding unit is used to perform fault encoding on each fault standard of the interface card based on a preset encoding rule to generate corresponding encoding information.

[0060] The matching unit is used to match the corresponding fault status code in combination with the coding information and the interrupt signal.

[0061] The determination unit is used to determine the interrupt level according to the fault status code.

[0062] Fault standardization and coding (coding unit)

[0063] In some embodiments, a set of fault coding rules (hardware logic implementation) may be pre-set inside the interrupt handler 100 .

[0064] The encoding unit encodes the fault criteria of each interface card (such as power supply error, link disconnection, temperature limit exceedance, verification error, etc.) at the hardware level.

[0065] When encoding output, encoding information corresponding to each specific fault is generated (for example, a 5-bit binary code represents a fault type), thereby digitizing and standardizing complex physical fault signals, providing a basis for subsequent matching and priority determination.

[0066] When the hardware detects an actual interrupt signal (such as an interface card error) input into the interrupt handler 100, the matching unit is triggered. The matching unit performs a hardware logic comparison of the received interrupt signal with the encoding information library generated by the encoding unit. This determines the fault status code corresponding to the interrupt signal. This status code not only contains the fault type (derived from the encoding information) but may also imply or be associated with basic priority information (for example, certain severe fault types inherently correspond to high-priority codes). This information then determines the interrupt level.

[0067] For example, the interrupt handler 100 may encode and transmit the finally determined interrupt level via the INTx interrupt signal lines (eg, 4 groups*8 signals) or a more economical solution (eg, 3 IO pins).

[0068] Here is an example of using 3 IOs (Input / Output) for 8-level encoding:

[0069] The three binary IO pins can be combined to produce 2^3=8 states (000 to 111).

[0070] Each state represents an interrupt level (e.g. 000 = level 0 / no interrupt, 001 = level 1, 010 = level 2, ..., 111 = level 7).

[0071] Based on different interrupt levels, this embodiment of the present invention can identify the target processor for fault handling, optimizing resources. The data processor 300 only handles truly urgent interrupts, avoiding being overwhelmed by a flood of low-priority interrupts. The management processor 200 handles deferrable tasks (such as logging and fan control), freeing up the computing power of the data processor 300 for core services.

[0072] Optionally, in one embodiment of the present invention, the interrupt handler 100 includes: a first sending unit and a second sending unit.

[0073] The first sending unit is configured to generate a corresponding emergency fault handling signal based on the interrupt signal when the interrupt level is greater than a preset level, and send the emergency fault handling signal to the data processor 300 through a preset physical link.

[0074] The second sending unit is configured to generate a corresponding non-emergency fault processing signal based on the interrupt signal when the interrupt level is less than or equal to a preset level, and send the non-emergency fault processing signal to the management processor 200 through a preset access link.

[0075] According to different interrupt levels, the embodiment of the present invention can select different target processors.

[0076] In the embodiment of the present invention, a critical threshold (eg, level 4) may be preset to distinguish between urgent interrupts (needing to be immediately processed by the data processor 300) and non-urgent interrupts (which can be processed by the management processor 200).

[0077] For example, if the preset level is 7, an interruption level greater than 7 is considered an emergency interruption, and an interruption level ≤ 7 is considered a non-emergency interruption.

[0078] The first sending unit can send emergency fault processing signals such as CPU power supply anomaly, memory uncorrectable error, system watchdog timeout, etc. to the data processor 300, allowing the data processor 300 to capture and process emergency events (such as triggering a kernel crash dump) with minimal latency.

[0079] The second sending unit can send non-emergency fault processing signals such as correctable memory errors, interface card hot plug events, fan speed adjustment requests, etc. to the management processor 200. The management processor 200 performs out-of-band recording, alarming or delayed processing to avoid occupying computing resources of the data processor 300.

[0080] In addition, there are some faults that require the management processor 200 and the data processor 300 to handle in coordination. According to different interrupt levels, the interrupt signals that need to be handled in coordination can also be screened out, and then the first sending unit and the second sending unit are used to send signals to the management processor 200 and the data processor 300 at the same time to perform fault handling.

[0081] For example, as shown in Table 1, the embodiment of the present invention can determine the priority based on the interrupt signal and the trigger source, and then determine the target processor. Table 1 is a fault collaborative processing table.

[0082] Table 1

[0083]

[0084] Optionally, in one embodiment of the present invention, the management processor 200 includes: a second response unit and a second processing unit.

[0085] Among them, the second response unit is used to respond to non-emergency fault processing signals, read fault area data from the shared storage area of ​​the management processor 200, and calculate the fault score of each fault based on the controller status data, the data processor 300 status data and the fault area data, and prioritize multiple faults based on the fault score to obtain a processing queue.

[0086] The second processing unit is configured to execute corresponding processing actions based on the processing queue.

[0087] During actual execution, the second response unit of the embodiment of the present invention may receive a non-emergency fault processing signal from the interrupt handler 100 (via a preset access link).

[0088] The second response unit can read the original information of the faulty device (such as sensor address, error register value), controller status data (such as hardware status written by the interrupt processor 100, such as power / clock stability), and status data of the data processor 300 (operating status reported by the main processor, such as load, temperature, memory ECC (Error-Correcting Code) count) from the shared storage area.

[0089] Based on the read data, the second response unit can score the fault. For example, it can score based on the preset weight of the fault type (such as fan failure = 0.8, correctable memory error = 0.3), whether the fault may affect other components (such as power supply failure risk value > high temperature), the degree of system resource utilization by the fault (such as the weight of fan failure increases under high load), etc., and then prioritize the fault handling according to the score.

[0090] The second processing unit can perform corresponding fault processing according to the sorting result of the second response unit.

[0091] Optionally, in one embodiment of the present invention, the second response unit includes: an acquisition subunit, an assignment subunit and a calculation subunit.

[0092] The acquisition subunit is used to obtain the fault severity level and historical fault frequency of each fault from the fault area data, and obtain the current delay sensitivity of the system based on the status data of the data processor 300.

[0093] The assignment subunit is used to assign corresponding weights to the fault severity level, historical fault frequency and current delay sensitivity of the system.

[0094] The calculation subunit is used to calculate the fault score of each fault using the weight, fault severity level, historical fault frequency and current delay sensitivity of the system.

[0095] The calculation expression of the fault score is:

[0096] P dynamic = α • S severity + β • L latency + c • H history ,

[0097] in, P dynamic Score the fault. S severity is the fault severity level, L latency is the current delay sensitivity of the system, H history is the historical failure frequency, α 、 β 、 c They are the weights of fault severity level, system current delay sensitivity, and historical fault frequency.

[0098] The embodiment of the present invention can use a dynamic factor model to implement a dynamic priority adjustment algorithm to achieve server interface card fault isolation. The core principle of the algorithm is as follows:

[0099] The calculation expression of the fault score is:

[0100] P dynamic = α • S severity + β • L latency + c • H history ,

[0101] in, P dynamic Score the fault. S severity is the fault severity level (1-10, e.g. PCIe link training failure = level 10, temperature overrun = level 5), L latency is the current delay sensitivity of the system (calculated based on the load rate of the data processor 300, with a linear mapping from 0.1 to 1.0), H history is the historical failure frequency (exponential decay model H =1- eT - λt , l =0.05), α 、 β 、 c The fault severity levels (e.g. α =0.6, β =0.3, c =0.1), the weights of the system's current delay sensitivity and historical failure frequency.

[0102] PCIe link training is a series of operations performed by PCIe devices during initialization or recovery, aiming to establish and optimize the communication link between devices, ensuring link stability, maximizing data transmission rates, and reducing bit error rates.

[0103] The fault scoring adjustment algorithm can solve the core problem of complex fault causes and recurring faults. According to the fault scoring adjustment algorithm, the root cause of the fault can be dynamically identified to improve the accuracy and effectiveness of fault isolation and repair.

[0104] Optionally, in one embodiment of the present invention, the second processing unit includes: a matching subunit, a first processing unit, a second processing unit, a third processing unit, a fourth processing unit and a fifth processing unit.

[0105] The matching subunit is used to match a processing action for any fault based on the fault score;

[0106] The first processing unit is configured to, when the fault score is within a first preset score range, determine a processing action of confirming a target faulty device of any fault, powering off and isolating the target faulty device, and generating a corresponding alarm signal based on the target faulty device.

[0107] The second processing unit is configured to determine, when the fault score is within a second preset score interval, that a processing action is to shut down a high-speed peripheral component interconnect interface of a target faulty device and record any fault in a black box log, wherein a lower limit value of the first preset score interval is greater than an upper limit value of the second preset score interval.

[0108] The third processing unit is used to determine, when the fault score is in a third preset score interval, that a processing action is to reduce the data transmission speed of the high-speed peripheral component interconnection interface of the target faulty device to a preset speed threshold and to initiate a link repair action, wherein the lower limit value of the second preset score interval is greater than the upper limit value of the third preset score interval.

[0109] The fourth processing unit is configured to determine, when the fault score is within a fourth preset score interval, a processing action of limiting the bandwidth of the target faulty device to a preset value and reporting any fault to the operating system, wherein the lower limit value of the third preset score interval is greater than the upper limit value of the fourth preset score interval.

[0110] The fifth processing unit is configured to determine, when the fault score is within a fifth preset score interval, that the processing action is to record any fault in a black box log, wherein a lower limit value of the fourth preset score interval is greater than an upper limit value of the fifth preset score interval.

[0111] Among them, the matching subunit can realize the matching of fault scores and processing actions.

[0112] The embodiment of the present invention can pre-construct mapping rules, determine the corresponding dynamic scoring interval according to the fault score, determine the corresponding fault handling level within different dynamic scoring intervals, and use the processing actions corresponding to the fault handling level to handle the fault, forming a fault hierarchical processing system to avoid excessive processing and optimize resources.

[0113] For example, it can be shown in Table 2, which is a dynamic scoring interval-level-processing action comparison table.

[0114] Table 2

[0115]

[0116] Optionally, in one embodiment of the present invention, the data processor 300 includes: a first responding unit and a first processing unit.

[0117] The first response unit is configured to respond to the emergency fault handling signal and determine at least one fault and a corresponding fault isolation level based on the status data of the interrupt processor 100 and the interrupt signal, so as to match a corresponding fault isolation strategy for each fault based on the fault isolation level;

[0118] The first processing unit is configured to execute corresponding processing actions based on the fault isolation strategy.

[0119] Furthermore, in an embodiment of the present invention, the first response unit can receive an emergency fault signal sent by the interrupt processor 100 through a physical direct link, determine the fault isolation level based on the status data of the interrupt processor 100 (obtained directly from the interrupt processor 100 through the physical link) and the interrupt signal, and perform corresponding processing actions through the first processing unit.

[0120] The fault isolation level can be determined based on the interrupt level of the interrupt signal or by analysis. For example, the interrupt signal feature can be analyzed by edge type (rising edge = instantaneous fault, continuous low level = permanent fault).

[0121] Directly acquiring status data through physical links can effectively improve response speed.

[0122] Optionally, in one embodiment of the present invention, the data processor 300 is further configured to control the interrupt processor 100 to perform processing actions.

[0123] It is understandable that the interrupt handler 100 contains a programmable logic device, and the data processor 300 can send instructions to it through a preset register interface or through a pre-built physical link.

[0124] In terms of processing actions, the interrupt handler 100 can perform a wide range of tasks. For example, it can dynamically adjust sensor sampling rates, which is very useful when troubleshooting intermittent faults; it can also switch redundant power modules to achieve seamless failover; and it can even forcibly isolate a faulty PCIe device, which is more thorough than software isolation. These operations embody the principle of hardware autonomy: when the data processor 300 detects an anomaly, it can directly direct the interrupt handler 100 to resolve the problem at the underlying level without waiting for the OS (operating system) to intervene.

[0125] Optionally, in one embodiment of the present invention, the processing system 10 of the server further includes: a determination module and a repair module.

[0126] The determination module is used to obtain historical fault data and business information of the fault corresponding to the interrupt signal, and determine whether the fault meets the preset repair conditions based on the historical fault data and business information;

[0127] The repair module is used to match the repair strategy corresponding to the fault that meets the preset repair conditions and perform the corresponding repair action.

[0128] As a possible implementation method, the determination module of an embodiment of the present invention can extract historical fault data from a shared storage area, and determine whether the fault corresponding to the interrupt signal can be repaired based on the historical fault data and business information (such as real-time business value weight, business tolerance, resource dependency map, etc.).

[0129] When making a judgment, one can determine whether the repair conditions are met by considering multiple aspects, such as technical feasibility (such as whether the fault is known to be repairable), economic rationality (such as the comparison between the time consumption of automatic repair and manual processing), and risk controllability (such as the completeness of the rollback plan when the repair fails).

[0130] After the repair conditions are met, the embodiment of the present invention can perform corresponding repair actions.

[0131] Optionally, in one embodiment of the present invention, the repair module includes: a determination unit, a first repair unit, a second repair unit and a third repair unit.

[0132] The determining unit is used to determine the fault type of the fault that meets the preset repair condition.

[0133] The first repair unit is configured to execute a preset hot reset attempt action to repair the fault when the fault type is a preset zombie fault type.

[0134] The second repairing unit is configured to execute a preset firmware reloading action to repair the fault when the fault type is a preset firmware fault type.

[0135] The third repairing unit is configured to, when the fault type is a preset connection fault type, execute a preset link or interface reset action to repair the fault.

[0136] Among them, the repair module can perform corresponding fault repair according to the fault type.

[0137] For the zombie fault type, the first repair unit can perform hot reset attempts such as hot restart and configuration recovery to repair the fault;

[0138] For firmware fault types, the second repair unit may perform a firmware reload action to repair the fault;

[0139] For the connection fault type, the third repair unit may perform a link or interface reset action to repair the fault.

[0140] In addition, for voltage drift faults, dynamic compensation power supply can be performed, data pollution faults can be cache flushed and memory remapped, and transient link errors can be link retrained.

[0141] The embodiment of the present invention can record each fault, repair judgment result, repair action and repair result, and then store the corresponding data when the repair result is successful for subsequent call. For repair actions that require manual intervention, corresponding records can also be made, and it can be determined whether they can be handled by themselves, so that automatic response can be achieved when the corresponding fault occurs again next time.

[0142] Optionally, in one embodiment of the present invention, the processing system 10 of the server further includes: a judgment module and an alarm module.

[0143] The judgment module is used to judge whether the processing action meets the preset server operation impact condition before the control data processor 300 or the management processor 200 performs the corresponding processing action.

[0144] The alarm module is used to terminate the processing action and generate an alarm signal when the preset server operation impact conditions are met.

[0145] As a possible implementation method, the judgment module can perform a preliminary assessment of the impact of the processing action.

[0146] Evaluation dimensions may include: performance impact index (predicting the extent of CPU, memory, and IO performance degradation), business continuity risk (analyzing the SLA tolerance of related businesses), and hardware safety boundaries (detecting the safety margin of parameters such as voltage and temperature). Among them, the thresholds related to each evaluation can be dynamically adjusted according to the specific business scenario.

[0147] When it is determined that the fault handling will affect the operation of the server, the embodiment of the present invention can stop the processing action and report a warning so as to perform corresponding fault handling according to the user's decision to avoid affecting the normal business processing of the server.

[0148] Optionally, in one embodiment of the present invention, after the management processor 200 detects a monitoring exception signal, the management processor 200 determines whether the interrupt processor 100 detects an interrupt signal, and determines whether the data processor 300 detects a status exception signal, and performs corresponding exception handling actions based on the data in the shared storage area if no interrupt signal and status exception signal are detected.

[0149] During the actual execution process, exception detection is not performed only by the interrupt processor 100. In fact, the interrupt processor 100, the management processor 200 and the data processor 300 can all detect exception signals. Generally speaking, when a fault occurs, two or more of the interrupt processor 100, the management processor 200 and the data processor 300 will detect the exception signal. At this time, if the interrupt processor 100 detects an interrupt signal, the interrupt level of the interrupt processor 100 is used as the priority processing method. If the interrupt processor 100 does not detect an interrupt signal, it is possible that the management processor 200 detects a monitoring exception signal while the interrupt processor 100 does not detect an interrupt signal. At this time, the management processor 200 handles the related faults of the monitoring exception signal.

[0150] For example, the management processor 200 may monitor the hardware status in real time through an integrated sensor array (temperature / voltage / fan speed sensor), and generate a monitoring abnormality signal when it detects that a preset threshold is exceeded.

[0151] At this time, through the special link structure of the embodiment of the present invention, the management processor 200 can directly obtain whether the interrupt processor 100 and the data processor 300 have detected an interrupt signal and a status abnormality signal.

[0152] When no interrupt signal or status abnormality signal is detected, the data processor 300 can perform in-depth data analysis based on the data in the shared storage area, for example, parsing the sensor data matrix collected by the data processor 300 itself, constructing a time series waveform, cross-analyzing the performance counter data of the data processor 300, analyzing the frequency domain characteristics of the sensor data, and using a pre-trained decision tree model to identify abnormal patterns.

[0153] Based on the analysis results, the management processor 200 can perform hierarchical processing. For example, if the anomaly type is data drift, sensor calibration and historical data correction are performed; if the anomaly type is latent failure, preventive frequency reduction and spare part preheating are performed; if the anomaly type is environmental interference, filtering enhancement is performed; and for unknown anomalies, black box recording can also be performed.

[0154] Optionally, in one embodiment of the present invention, after the data processor 300 detects a status exception signal, the data processor 300 determines whether the interrupt processor 100 detects an interrupt signal, and determines whether the management processor 200 detects a monitoring exception signal, and performs corresponding exception handling actions based on the data in the shared storage area if no interrupt signal and monitoring exception signal are detected.

[0155] Similarly, there may be a situation where only the data processor 300 detects a status abnormality signal. In this case, the data processor 300 can also confirm the status of the interrupt processor 100 and the management processor 200 through the special communication link of the embodiment of the present invention, and then perform corresponding exception handling actions based on the data in the shared storage area when no interrupt signal and monitoring abnormality signal are detected.

[0156] For example, when the data processor 300 detects an abnormality in the PCIe link status information of the interface card, it can determine the actual abnormality based on the data in the shared storage area and perform corresponding exception handling actions. For example, if the abnormality is that the PCIe link communicating with the interface card is disconnected, it can try to reconnect or directly disconnect the link communicating with the interface card, rely on the data obtained by the interrupt processor 100 and the management processor 200 to maintain normal operation, and report the fault.

[0157] Combine Figure 2 As shown, the working principle of the server processing system 10 of an embodiment of the present invention is described in detail using an embodiment.

[0158] Currently, interface card component fault detection typically relies on BMC firmware polling the interface card chip to collect information such as component sensor temperature and chip link error counts. The BMC management link typically uses low-speed I2C access. The access method and cycle times cannot guarantee fault detection response speed, hindering subsequent fault isolation and hardware solution implementation. The interface card itself is a PCIe device, and the CPU can access the add-in PCIe device via the PCIe link. During server startup or operation, the BIOS (Basic Input / Output System) firmware can collect fault information from the add-in device by detecting PCIe link failures. Although the BIOS can transmit fault information to the BMC through the IPMI channel in existing solutions, independent out-of-band diagnostic troubleshooting can potentially impact services and lacks collaborative diagnostic capabilities and accurate fault resolution. The embodiment of the present invention can set up a special communication structure, the state synchronization bus (StateSync Bus). This hardware bus adopts a three-terminal heterogeneous interconnection architecture. Through the hardware-level interconnection bus and collaboration, it can achieve state synchronization and sharing among the three devices of CPU, BMC and CPLD. At the same time, it can break the fault information island in the system, making the diagnostic information under the system quickly respond and interoperable. The hardware design of the state synchronization bus can be as follows Figure 2 As shown, taking the CPLD as the interrupt processor 100 , the BMC as the management processor 200 , and the CPU as the data processor 300 as an example, the processing system 10 of the server may include the interrupt processor 100 , the management processor 200 , and the data processor 300 .

[0159] Among them, the interrupt processor 100 can communicate with multiple interface cards (interface card 1 401, interface card 2 402, interface card 3 403 and interface card 4 404) to obtain interrupt signals; the management processor 200 can also monitor the hardware status data of multiple interface cards through the BMC access channel via SW500 to obtain monitoring abnormality signals; the data processor 300 can communicate with multiple interface cards through the PCIe link to obtain status abnormality signals.

[0160] The data processor 300 and management processor 200 are interconnected using a shared PCIe device configuration space memory solution. As an EP (Endpoint Device), the management processor 200 can directly read and write to the shared memory area 201 using the BIOS. The data processor 300 and interrupt handler 100, as well as the interrupt handler 100 and management processor 200, use interrupts for rapid diagnosis and status monitoring. The data processor 300 and interrupt handler 100 utilize an LPC (Low Pin Count Physical Link, a serial communication interface used to connect low-bandwidth, low-speed peripherals on a computer motherboard) physical link. The data processor 300 can access the interrupt handler 100 through memory-mapped I / O (memory-mapped I / O, which maps the registers or storage space of I / O devices to the main memory address space, allowing the data processor 300 to access these devices using standard memory read and write instructions) to operate and control the interrupt controller. The interrupt processor 100 and the management processor 200 communicate directly with each other through an I3C (Improved Inter Integrated Circuit) link. The management processor 200 driver layer uses DMA (Direct Memory Access) technology to achieve rapid data synchronization between the interrupt processor 100 and the management processor 200. The interrupt signal line of the interrupt processor 100 is connected to the NMI interrupt pin of the data processor 300 and the interrupt control pin of the management processor 200, respectively.

[0161] In actual implementation, taking the process of implementing fault processing after the interrupt processor 100 detects an interrupt signal as an example, the embodiment of the present invention may include the following steps:

[0162] Step S1, interrupt detection. During the server startup or operation, the interrupt processor 100 monitors the PCIe key signals of the interface card component. This process is implemented by the interrupt processor 100 logic code. The specific implementation logic includes:

[0163] The interrupt processor 100 detects status signal input (interface card PCIe PRSNT in-position signal and PWR_GOOD power status, HeartBeat firmware status, AER_ALERT fault alarm);

[0164] The internal priority arbiter of the interrupt processor 100 performs interrupt arbitration according to the defined priority based on the current detection status code. The interrupt processor 100 uses logic code to implement hardware-level interrupt aggregation and supports 32 levels of priority. The interrupt controller is divided into an interrupt mask controller (32 bits), a priority arbiter (Fixed Priority), and an INTx interrupt signal line (4*8 groups). In practice, several IOs can be used for priority encoding. For example, 3 IOs are used as a group for priority encoding, and 8 groups of different interrupt levels can be implemented according to the priority.

[0165] The interrupt handler 100 triggers an interrupt according to the interrupt mask register mask, and the triggering rules refer to the hardware fault interrupt vector allocation table.

[0166] Step S2, interrupt synchronization. The interrupt processor 100 implements a memory management module internally, and updates the fault status to the shared memory through the status synchronization bus. That is, when the hardware fault signal is triggered, the interrupt processor 100 can update the status data in the shared storage area 201 according to the interrupt type. The writing process adopts atomic operations to avoid data synchronization anomalies caused by access by the interrupt processor 100 or the management processor 200. The I3C link between the interrupt processor 100 and the management processor 200, the interrupt processor 100 shared memory access driver is implemented internally by the management processor 200, and the shared memory area is accessed internally using DMA. This solution can avoid the efficiency of the management processor 200 business affecting the alarm. The specific synchronization process is as follows:

[0167] The interrupt handler 100 acquires the atomic lock atomic_lock, which can be implemented by using a fixed location in the shared memory or an external physical pin to achieve atomic access to the shared memory;

[0168] Calculate the target address (address translation), encode each PCIe device in the system, and implement data unique mapping address access based on the address translation algorithm;

[0169] The interrupt processor 100 logic implements fault data framing, and the frame data includes: fault type, timestamp, reserved bits, and checksum;

[0170] The interrupt handler 100 initiates a DMA write operation to write data into the shared memory;

[0171] The interrupt handler 100 releases the atomic lock.

[0172] Step S3: software collaborative processing.

[0173] Interrupt handler 100 detects hardware signals such as power, presence, and abnormal status from add-in card components. When interrupt handler 100 detects a hardware anomaly, it triggers an interrupt to the BMC and BIOS for processing. Management processor 200 makes decisions based on diagnostic results from specific areas within shared memory area 201, enabling both in-band and out-of-band collaborative diagnosis without the operating system's awareness of the service.

[0174] Based on the fault interrupt vector allocation table implemented in the interrupt handler 100 logic, the interrupt handler 100 prioritizes high-priority faults by triggering interrupts to the management processor 200 or data processor 300. The interrupt handler 100 triggers an NMI (non-maskable interrupt) to wake up the BIOS interrupt service routine. The BIOS reads the faulty area in shared memory, determines the isolation level, and responds with isolation measures. The BMC's running status polling monitoring service detects data updates in shared memory, invokes a dynamically adjusted isolation algorithm, evaluates the fault, and applies different repair strategies based on the evaluation score.

[0175] The embodiments of the present invention can detect interface card component faults through the above-mentioned structure and steps. A hierarchical fault isolation mechanism is employed to isolate interface card component faults. Fault levels are categorized as: startup isolation, operation isolation, pending correction isolation, and minor fault. The startup isolation level indicates that the fault severely impacts system startup and operation. During startup, the device is hardware-disconnected from the system or its power is cut off to ensure that it does not affect normal device startup. At the operation isolation level, if a device experiences a sudden fault or persistent malfunction during server operation, the device is isolated and a backup channel or port is switched, or the communication link is disconnected, to prevent the fault from spreading. At the pending correction isolation level, if a device component fails and collaborative diagnosis determines that the device is repairable, the component is restored through repair measures. This hierarchical fault isolation mechanism ensures that the fault does not spread and minimizes the impact on service continuity. Minor faults do not impact service communication and are only recorded in the alarm log. External interface card interrupt status input, interrupt priority type, and fault information state synchronization via a state synchronization bus output a server interface card hardware fault interrupt vector allocation table, enabling collaborative fault handling, as shown in Table 1.

[0176] Based on the above hardware design, the embodiment of the present invention can implement a dynamic priority adjustment algorithm based on a dynamic factor model to achieve server interface card fault isolation. The core principle of the algorithm is as follows:

[0177] The calculation expression of the fault score is:

[0178] P dynamic = α • Sseverity + β • L latency + c • H history ,

[0179] in, P dynamic Score the fault. S severity is the fault severity level (1-10, e.g. PCIe link training failure = level 10, temperature overrun = level 5), L latency is the current delay sensitivity of the system (calculated based on the load rate of the data processor 300, with a linear mapping from 0.1 to 1.0), H history is the historical failure frequency (exponential decay model H =1- eT - λt , l =0.05), α 、 β 、 c The fault severity levels (e.g. α =0.6, β =0.3, c =0.1), the weights of the system's current delay sensitivity and historical failure frequency.

[0180] In summary, during actual fault handling, embodiments of the present invention can collaboratively detect faults during server interface card startup / operation through the interrupt processor 100, management processor 200, and data processor 300 (CPLD / BMC / BIOS). The management processor 200 reads fault information from the shared memory area 201 and invokes a dynamic priority assessment algorithm. This dynamic assessment process fully considers parameters such as the fault type, historical fault information, and current system sensitivity to calculate a dynamic fault score and distinguish whether the fault is a high-priority or low-priority fault. Priority levels can be determined using a preset method (e.g., high priority P <= 7, low priority P >= 8). High-priority faults are immediately isolated by an interrupt, resulting in a power / link disconnection or a fault alarm. Low-priority faults are placed in a delayed processing queue by default, awaiting processing by the management processor 200. Once an alarm is triggered, a recovery policy decision is made. Based on historical fault data and service importance, a decision is made on whether to repair the component. Repair policies include warm reset attempts, firmware reloads, and PCIe link port resets. Different repair policies can be pre-set and bound to different device types.

[0181] Step S4, interruption masking strategy.

[0182] The embodiment of the present invention may reserve an interrupt mask register, and the data processor 300 or the management processor 200 may perform a masking design for related faults according to actual scenarios to ensure that the isolated repair does not affect the normal function startup.

[0183] In summary, the three-level interrupt fusion architecture of the embodiment of the present invention can integrate hardware signal triggering, in-band interrupts, and out-of-band alarms into a unified state bus, breaking the fault information islands in the system and improving the timeliness and effectiveness of fault alarms; through the fault priority grading strategy in the interrupt processor 100, different response entities are allocated to reduce state synchronization delay and concurrent fault processing capabilities; through the unified state bus data synchronization mechanism, each processor under the server system 10 can share fault information, and the fault processing process can be synchronized and shared in real time to avoid the risk of fault propagation under the system; and a fault scoring adjustment algorithm is proposed to solve the core problem of complex fault causes and repeated faults. According to the fault scoring adjustment algorithm, the root cause of the fault can be dynamically determined, and the accuracy and effectiveness of fault isolation and repair can be improved.

[0184] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0185] like Figure 3 As shown, an embodiment of the present invention further provides a processing method for a server, comprising the following steps:

[0186] In step S301, after the interruption processor detects an interruption signal, the current interruption level of the server is determined according to the interruption signal.

[0187] In step S302, a target processor is determined based on the current interrupt level to control the target processor to perform corresponding processing actions on the device based on the interrupt signal or status data.

[0188] For the description of the features in the embodiment corresponding to the processing method of the server, please refer to the relevant description of the embodiment corresponding to the processing system of the server, which will not be repeated here.

[0189] An embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned server processing method embodiments.

[0190] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned server processing method embodiments when running.

[0191] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0192] An embodiment of the present invention further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned server processing method embodiments are implemented.

[0193] An embodiment of the present invention further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned server processing method embodiments are implemented.

[0194] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0195] The above is a detailed introduction to the processing system, method, electronic device and storage medium of a server provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, the present invention can also be improved and modified in several ways, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.

Claims

1. A server processing system, characterized in that: The system comprises an interrupt processor, a management processor and a data processor. After the interrupt processor detects an interrupt signal, the interrupt processor communicates the interrupt signal and status data with the data processor based on a preset physical link. The interrupt processor communicates the interrupt signal and status data with the management processor based on a preset access link. The data processor communicates the interrupt signal and status data based on a PCIe link. The interrupt processor determines a current interrupt level of the server according to the interrupt signal, and determines a target processor based on the current interrupt level, so as to control the target processor to perform a corresponding processing action on at least one device involved in the interrupt signal based on the interrupt signal or the status data; In which, the management processor has a shared storage area, wherein the interrupt signal detected by the interrupt processor and the status data of the interrupt processor are written to the shared storage area through the preset access link, the status abnormality signal detected by the data processor, the status data of the data processor and the server operation data collected by the data processor are written to the shared storage area through the PCIe link, and the monitoring abnormality signal detected by the management processor, the status data of the management processor and the hardware status data collected by the management processor are written to the shared storage area.

2. The server processing system according to claim 1, characterized in that: The interrupt processor, the management processor, and the data processor write the interrupt signal and corresponding status data into the shared memory area based on a preset atomic operation, so as to prohibit the interrupt processor, the management processor, and the data processor from writing data at the same time.

3. The server processing system according to claim 1, wherein: After the management processor detects the monitoring exception signal, the management processor determines whether the interrupt processor detects the interrupt signal and determines whether the data processor detects the status exception signal, and performs corresponding exception handling actions based on the data in the shared storage area if the interrupt signal and the status exception signal are not detected.

4. The server processing system according to claim 1, wherein: After the data processor detects the status exception signal, the data processor determines whether the interrupt processor detects the interrupt signal and determines whether the management processor detects the monitoring exception signal, and performs corresponding exception handling actions based on the data in the shared storage area if the interrupt signal and the monitoring exception signal are not detected.

5. The server processing system according to claim 1, wherein: The interrupt handler comprises: an encoding unit, configured to perform fault encoding on each fault standard of the interface card based on a preset encoding rule to generate corresponding encoding information; A matching unit, configured to match a corresponding fault status code in combination with the coding information and the interrupt signal; A determining unit is configured to determine the interruption level according to the fault status code.

6. The server processing system according to claim 1, wherein the interrupt handler comprises: a first sending unit, configured to generate a corresponding emergency fault handling signal based on the interrupt signal when the interrupt level is greater than a preset level, and send the emergency fault handling signal to the data processor through the preset physical link; The second sending unit is configured to generate a corresponding non-emergency fault processing signal based on the interrupt signal when the interrupt level is less than or equal to the preset level, and send the non-emergency fault processing signal to the management processor through the preset access link.

7. The server processing system according to claim 6, characterized in that: The data processor comprises: a first response unit, configured to determine, in response to the emergency fault handling signal, at least one fault and a corresponding fault isolation level based on the state data of the interrupt handler and the interrupt signal, so as to match a corresponding fault isolation strategy for each fault based on the fault isolation level; The first processing unit is configured to execute corresponding processing actions based on the fault isolation strategy.

8. The server processing system according to claim 6, characterized in that: The management processor includes: a second response unit, configured to respond to the non-urgent fault processing signal, read fault region data from a shared memory area of ​​the management processor, calculate a fault score for each fault based on controller status data, data processor status data, and the fault region data, and prioritize multiple faults based on the fault scores to obtain a processing queue; The second processing unit is configured to execute corresponding processing actions based on the processing queue.

9. The server processing system according to claim 8, characterized in that: The second response unit includes: an acquisition subunit, configured to acquire the fault severity level and historical fault frequency of each fault from the fault area data, and obtain the current delay sensitivity of the system based on the status data of the data processor; an assignment subunit, configured to assign corresponding weights to the fault severity level, the historical fault frequency, and the current delay sensitivity of the system; A calculation subunit is configured to calculate a fault score for each fault by using the weight, the fault severity level, the historical fault frequency, and the current delay sensitivity of the system.

10. The server processing system according to claim 9, characterized in that: The calculation expression of the fault score is: P dynamic = α • S severity + β • L latency + γ • H history , in, P dynamic Score the fault, S severity is the fault severity level, L latency is the current delay sensitivity of the system, H history is the historical fault frequency, α 、 β 、 γ are the weights of the fault severity level, the current delay sensitivity of the system, and the historical fault frequency, respectively.

11. The server processing system according to claim 8, wherein: The second processing unit includes: a matching subunit, configured to match the processing action for any fault based on the fault score; a first processing unit, configured to, when the fault score is within a first preset score range, determine that the processing action is to identify a target faulty device of any fault, power off and isolate the target faulty device, and generate a corresponding alarm signal based on the target faulty device; a second processing unit, configured to, if the fault score is within a second preset score interval, determine that the processing action is to shut down a high-speed peripheral component interconnect interface of the target faulty device and record any fault in a black box log, wherein a lower limit of the first preset score interval is greater than an upper limit of the second preset score interval; a third processing unit, configured to, when the fault score is within a third preset score interval, determine that the processing action is to reduce the data transmission speed of the high-speed peripheral component interconnect interface of the target faulty device to a preset speed threshold and initiate a link repair action, wherein the lower limit of the second preset score interval is greater than the upper limit of the third preset score interval; a fourth processing unit, configured to, if the fault score is within a fourth preset score interval, determine that the processing action is to limit the bandwidth of the target faulty device to a preset value and report the fault to an operating system, wherein a lower limit of the third preset score interval is greater than an upper limit of the fourth preset score interval; a fifth processing unit, configured to determine, when the fault score is in a fifth preset score interval, that the processing action is to record any one of the faults in the black box log, wherein a lower limit value of the fourth preset score interval is greater than an upper limit value of the fifth preset score interval.

12. The server processing system according to claim 1, wherein: Also includes: A determination module, configured to obtain historical fault data and service information of the fault corresponding to the interrupt signal, and determine whether the fault meets a preset repair condition based on the historical fault data and service information; The repair module is used to match the repair strategy corresponding to the fault that meets the preset repair conditions and perform the corresponding repair action.

13. The server processing system according to claim 12, wherein: The repair module includes: A determination unit, configured to determine a fault type of the fault that meets the preset repair condition; A first repair unit is configured to, when the fault type is a preset deadlock fault type, execute a preset hot reset attempt action to repair the fault; a second repairing unit, configured to, when the fault type is a preset firmware fault type, execute a preset firmware reloading action to repair the fault; The third repairing unit is configured to, when the fault type is a preset connection fault type, execute a preset link or interface reset action to repair the fault.

14. The server processing system according to claim 1, wherein: Also includes: A judgment module, configured to judge whether a processing action satisfies a preset server operation impact condition before controlling the data processor or the management processor to execute a corresponding processing action; The alarm module is used to terminate the processing action and generate an alarm signal when the preset server operation influencing condition is met.

15. The server processing system according to claim 1, wherein: The data processor is further configured to control the interrupt processor to execute the processing action.

16. A processing method of a server, characterized in that: A processing system applied to a server according to any one of claims 1 to 15, wherein the method comprises the following steps: After the interruption processor detects an interruption signal, determining a current interruption level of the server according to the interruption signal; A target processor is determined based on the current interrupt level, so as to control the target processor to perform a corresponding processing action on the device based on the interrupt signal or the status data.

17. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the server processing method as claimed in claim 16 when executing the computer program.

18. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the processing method of the server according to claim 16.

19. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the processing method of the server as claimed in claim 16 are implemented.

Citation Information

Patent Citations

  • Apparatus and method for shortening interrupt latency

    JP2010140239A