Server processing system and method, electronic equipment and storage medium

By building a link communication structure between the interrupt processor, management processor and data processor in the server, and dynamically selecting the target processor for troubleshooting, the problem of data synchronization delay and hardware and software status in the server is solved, and fast and accurate fault response and processing is achieved.

CN120407265AActive Publication Date: 2025-08-01INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510898120.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-08-01
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

The data synchronization delay between the processors in the server is high, and the hardware signal and software status are disconnected, resulting in untimely and inefficient fault handling, making it difficult to effectively manage.

Method used

By constructing a special communication structure, link communication between the interrupt processor, management processor and data processor is realized, fault information and status data are shared, and target processor is dynamically selected for fault processing according to the interrupt level.

Benefits of technology

It realizes state synchronization and data sharing between processors, responds quickly and handles faults, improves the efficiency and accuracy of fault handling, reduces CPU burden, and reduces the risk of fault spread.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407265A_ABST
    Figure CN120407265A_ABST
Patent Text Reader

Abstract

The invention discloses a server processing system and method, electronic equipment and a storage medium, and relates to the technical field of electric digital data process.The server processing system comprises an interrupt processor, a management processor and a data processor, and the interrupt processor, the management processor and the data processor can communicate through corresponding links; and after the interrupt signal is triggered, the interrupt processor can determine the current interrupt level according to the interrupt signal so as to determine the target processor and control the target processor to carry out fault processing based on the shared information, so that the problems that the data synchronization delay among the processors of the server is relatively high and the data synchronization time is short in the prior art are solved. The technical problem that fault processing is difficult to effectively carry out when a server breaks down due to disconnection of hardware signals and software states is solved, and the technical effects that through a special framework, state synchronization and data sharing can be carried out among processors in time, fault information islands can be broken, and faults can be quickly responded and processed are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic digital data processing, and particularly to a processing system, method, electronic device and storage medium of a server. Background Art

[0002] The data processed by the server and the server hardware scale in the data center show an explosive growth phenomenon. The server management technology field has put forward higher and higher requirements for the reliability and availability of server components. In related technologies, there are certain defects in the processing architecture of the server. Among them, the data synchronization delay between each processor in the server is relatively high, and the hardware signal is out of touch with the software state, making it difficult to effectively handle faults when the server fails, and urgent improvement is needed. Summary of the Invention

[0003] The present invention provides a processing system, method, electronic device and storage medium of a server, so as to at least solve the problem that the data synchronization delay between each processor of the server in related technologies is relatively high, the hardware signal is out of touch with the software state, and it is difficult to effectively handle faults when the server fails.

[0004] The present invention provides a processing system of a server, including an interrupt processor, a management processor and a data processor. After the interrupt processor detects an interrupt signal, the interrupt processor communicates the interrupt signal and status data with the data processor based on a preset physical link, and the interrupt processor communicates the interrupt signal and status data with the management processor based on a preset access link. The data processor communicates the interrupt signal and status data based on a PCIe link. Among them, the interrupt processor determines the current interrupt level of the server according to the interrupt signal, and determines a target processor based on the current interrupt level, so as to control the target processor to perform corresponding processing actions on at least one device involved in the interrupt signal based on the interrupt signal or status data.

[0005] The present invention also provides a processing method of a server, including: after the interrupt processor detects an interrupt signal, determining the current interrupt level of the server according to the interrupt signal; determining a target processor based on the current interrupt level, so as to control the target processor to perform corresponding processing actions on the device based on the interrupt signal or status data.

[0006] The present invention also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any one of the above-mentioned processing methods of the server when executing the computer program.

[0007] The present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned processing methods of the server are implemented.

[0008] The present invention also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of any one of the above-mentioned server processing methods.

[0009] Through the present invention, since the interrupt processor, the management processor and the data processor can communicate through corresponding links, thereby realizing signal and status data sharing. Thus, after the interrupt signal is triggered, the interrupt processor can determine the current interrupt level according to the interrupt signal, so as to determine the target processor according to the current interrupt level, and control the target processor to perform fault handling based on the shared information, solving the technical problem in the related art that the data synchronization delay between the processors of the server is relatively high, and the hardware signals and software status are out of touch, making it difficult to effectively perform fault handling when the server fails, achieving the technical effect that through a special architecture, the processors can perform status synchronization and data sharing in a timely manner, breaking the fault information island, and quickly responding to and handling faults. Description of the Drawings

[0010] In order to more clearly illustrate the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0011] Figure 1 FIG. is a schematic structural diagram of a server processing system according to an embodiment of the present invention; Figure 2 FIG. is a schematic structural diagram of a server processing system according to an embodiment of the present invention; Figure 3 FIG. is a flowchart of a server processing method according to an embodiment of the present invention.

[0012] Among them, 10 - server processing system, 100 - interrupt processor, 200 - management processor, 201 - shared storage area, 300 - data processor, 401 - interface card 1, 402 - interface card 2, 403 - interface card 3, 404 - interface card 4, 500 - software. Detailed Embodiments

[0013] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0014] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects and not to describe a specific order or sequence.

[0015] It can be understood that the following defects exist in the related art: Communication protocol fragmentation: The CPU (Central Processing Unit) and BMC (Baseboard Management Controller) use the IPMI (Intelligent Platform Management Interface) protocol (software protocol stack), the BMC and CPLD (Complex Programmable Logic Device) use I2C / GPIO (low-speed hardware interface), and there is no direct connection channel between the CPU and CPLD, resulting in a high synchronization delay of fault states and frequent occurrence of fault diffusion caused by untimely fault handling. In addition, the hardware signals and software states are disjointed. The CPU can access the content of the PCIe (Peripheral Component Interconnect Express) configuration space, while the out-of-band hardware signals are detected in the CPLD, and the data lacks interaction and synchronization, affecting the efficiency and accuracy of fault handling.

[0016] Low information synchronization efficiency. Server management and fault repair rely on polling access to device data through the out-of-band low-speed bus signal, with a delay of more than 120 ms, which cannot meet the requirements of PCIe5.0 devices for sub-microsecond fault response.

[0017] Fault isolation often adopts the single management domain decision-making mode of the BMC, and there is a phenomenon of "mis-isolation" (for example, a temporary hardware link error is misjudged as a hardware fault), resulting in an increase in the service interruption rate.

[0018] The fault isolation mechanism lacks refined management. The server interface card is the core channel for external data interaction of the service. The traditional method only proposes relevant isolation measures for power failures, lacks refined classification and repair management of faults, and at the same time, the server operating state, fault history information, etc. are not included in the scope of fault repair and isolation.

[0019] To solve the above technical problems, embodiments of the present invention can build a special communication structure to achieve rapid sharing of fault information and status data. After an interrupt signal is triggered, different target processors are determined according to different priorities, so as to improve the processing efficiency and accuracy of fault handling by combining data such as fault history information.

[0020] To enable those skilled in the art of this technology to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0021] As Figure 1 shown, an embodiment of the present invention provides a processing system 10 for a server, including: an interrupt processor 100, a management processor 200, and a data processor 300.

[0022] Among them, after the interrupt processor 100 detects an interrupt signal, the interrupt processor 100 communicates the interrupt signal and status data with the data processor 300 based on a preset physical link, and the interrupt processor 100 communicates the interrupt signal and status data with the management processor 200 based on a preset access link. The data processor 300 communicates the interrupt signal and status data based on a PCIe link, where the interrupt processor 100 determines the current interrupt level of the server according to the interrupt signal, and determines a target processor based on the current interrupt level, so as to control the target processor to perform corresponding processing actions on at least one device involved in the interrupt signal based on the interrupt signal or status data.

[0023] In the actual execution process, embodiments of the present invention can be applied to a typical server architecture, including an interrupt processor 100, a management processor 200, and a data processor 300.

[0024] Among them, the interrupt processor 100 can be a CPLD. A CPLD is a type of programmable logic device and belongs to a digital integrated circuit with user-customizable logic functions. Its core structure consists of programmable logic macro cells and a programmable interconnect matrix. Users can design logic through a hardware description language or schematic input, generate a target file, and then burn the code into the CPLD chip via a download cable. Its programming method is based on E2PROM or FLASH memory, supports non-volatile storage, the program does not lose power after power-off, and the programming times can reach tens of thousands of times.

[0025] The management processor 200 can be a BMC, which is a dedicated microcontroller independent of the host system and is integrated on the hardware mainboards of servers, network devices, etc. It is used to monitor and manage various parameters of the physical environment (such as temperature, voltage, fan speed) and hardware status (such as power supply, CPU, memory). Even if the main CPU or operating system crashes, the BMC can still work.

[0026] The data processor 300 can be a CPU, which is the main computing unit, executes the operating system and application program code, and processes core business data.

[0027] An interrupt signal can refer to an electronic signal sent by a hardware device (such as a network card, a hard disk controller, a fan failure sensor, etc.) or software when it needs the processor's immediate attention. It interrupts the task that the processor is currently executing and requires it to handle an emergency event (such as data arrival, error occurrence, overheating).

[0028] The interrupt processor 100 can detect the interrupt signal through multiple interface cards. After detecting the interrupt signal, the interrupt processor 100 can achieve fast data sharing through a pre-built communication architecture.

[0029] Between the interrupt processor 100 and the management processor 200, communication of each other's status data and sharing of interrupt signals can be carried out through direct access methods such as I3C and DMA.

[0030] Between the interrupt processor 100 and the data processor 300, communication of each other's status data and sharing of interrupt signals can be carried out through a physical link such as LPC.

[0031] Between the management processor 200 and the data processor 300, communication of each other's status data and sharing of interrupt signals can be carried out through a PCIe link.

[0032] Through the above links, real-time synchronization of status can be completed among the three processors, which is convenient for subsequent fault response.

[0033] Furthermore, in addition to interrupt transfer, the interrupt processor 100 can also determine the current interrupt level based on the received interrupt signal, and dynamically select a target processor based on the determined interrupt level: For example, for interrupt signals that require urgent processing, are of high level, or require computational processing, the data processor 300 can be selected as the target processor. For interrupt signals that do not require urgent processing, are related to management, or when the data processor 300 is unavailable (or has a high load), the management processor 200 may be selected as the target processor.

[0034] After receiving the interrupt signal and related status data, the selected target processor (the management processor 200 or the data processor 300) executes a predefined processing action.

[0035] Based on the above architecture, embodiments of the present invention can reduce the burden on the data processor 300, i.e., the CPU: The interrupt processor 100, i.e., the CPLD, as a dedicated interrupt controller, is responsible for initially summarizing, filtering, and classifying interrupts, avoiding a large number of raw interrupt signals directly impacting the CPU, significantly reducing the overhead (context switching) of the CPU in processing interrupts, and enabling the CPU to focus more on core computing tasks.

[0036] The CPLD intelligently decides whether to hand over the interrupt to the CPU for quick processing or to the BMC for out-of-band management according to the interrupt level, improving the overall efficiency and pertinence of interrupt processing.

[0037] Low-latency core communication: The physical link between the CPLD and the CPU is designed specifically for quickly transmitting interrupt signals, ensuring that high-priority interrupts can be promptly responded to by the CPU.

[0038] The preset access link between the CPLD and the BMC uses a stable bus designed specifically for management, ensuring the reliable transmission of out-of-band management information.

[0039] The PCIe device communicates directly with the CPU through the PCIe link, making full use of the high bandwidth and low latency characteristics of PCIe to handle device interrupts and data transmission.

[0040] Optionally, in an embodiment of the present invention, the management processor 200 has a shared storage area, wherein the interrupt signals detected by the interrupt processor 100 and the status data of the interrupt processor 100 are written to the shared storage area through a preset access link. The status exception signals detected by the data processor 300, the status data of the data processor 300, and the server operation data collected by the data processor 300 are written to the shared storage area through the PCIe link. The monitoring exception signals detected by the management processor 200, the status data of the management processor 200, and the hardware status data collected by the management processor 200 are written to the shared storage area.

[0041] In addition to the interrupt processor 100 being able to detect interrupt signals, the management processor 200 and the data processor 300 can also perform corresponding fault detections. For example, the management processor 200 can detect the monitoring information (temperature, power consumption) of the interface card; the data processor 300 can detect the PCIe link status information (error code, bandwidth, rate) of the interface card. When any processor detects an abnormal signal, such as an interrupt signal, a status exception signal, or a monitoring exception signal, the corresponding signal and the corresponding status data can be directly written to the shared storage area for subsequent call of relevant data during fault handling, or the faults corresponding to the signals can be recorded for subsequent policy adjustment according to historical information during processing.

[0042] As a possible implementation, in the shared PCIe device configuration space memory solution where the data processor 300 and the management processor 200 are interconnected by a PCIe link, the management processor 200, as an EP (Endpoint Device) device, can be directly read and written to by the BIOS (Basic Input / Output System), that is, the shared memory area.

[0043] The interrupt processor 100 can write data to the shared memory through an access link, such as a direct access method.

[0044] That is to say, in the embodiments of the present invention, all key events and status data from three independent sources (hardware interrupts of the interrupt processor 100, system operation information of the data processor 300, and hardware monitoring data of the management processor 200) can be aggregated in the same shared memory area of the management processor 200, which greatly facilitates problem diagnosis, root cause analysis, and historical tracing.

[0045] Optionally, in an embodiment of the present invention, the interrupt processor 100, the management processor 200, and the data processor 300 write the interrupt signal and the corresponding status data to the shared memory area based on a preset atomic operation to prevent the interrupt processor 100, the management processor 200, and the data processor 300 from writing data simultaneously.

[0046] It can be understood that an atomic operation refers to a read and write memory operation that is predefined at the hardware or low-level firmware level and is indivisible. An atomic operation will not be interrupted by the operations of other processors or threads during execution. It will either succeed completely (all expected modifications take effect) or fail completely (the memory state remains unchanged), without an intermediate state or partial write.

[0047] In the embodiments of the present invention, the atomic operation serves as the underlying mechanism of the mutex. When the interrupt processor 100 (CPLD), the management processor 200 (BMC), or the data processor 300 (CPU) needs to write to the shared memory area, they must "declare" the write permission by executing this preset atomic operation. The core purpose of this atomic operation is to ensure that only one processor can successfully obtain the write permission and perform the actual write action at the same time. It strictly prevents the CPLD, BMC, and CPU (or any two of them) from writing to the same block or associated area of the shared memory area at the same moment. Thus, it prevents concurrent write conflicts and ensures data integrity and consistency.

[0048] For example, if no other processor currently holds the lock (i.e., a specific flag in the shared storage area is in the "idle" state), the processor performing the atomic operation will successfully set the flag to "locked" and immediately obtain the write permission to start its data writing operation.

[0049] If another processor currently holds the lock (the flag is already in the "locked" state), then the processor attempting to perform the atomic operation will fail and cannot obtain the write permission. It must wait (possibly by polling or interrupt notification) until it detects that the lock is released (the flag changes back to "idle").

[0050] Writing and "unlocking": The processor that has obtained the write permission can safely write its data to the target location in the shared storage area. After completing the writing, it must perform another preset atomic operation (or set a specific flag) to release the lock (set the flag to "idle"), informing other waiting processors that they can now attempt to acquire the lock and perform writing.

[0051] Optionally, in an embodiment of the present invention, the interrupt processor 100 includes: an encoding unit, a matching unit, and a determining unit.

[0052] Among them, the encoding unit is used to perform fault encoding on each fault criterion of the interface card based on a preset encoding rule to generate corresponding encoding information.

[0053] The matching unit is used to match the corresponding fault status code by combining the encoding information and the interrupt signal.

[0054] The determining unit is used to determine the interrupt level according to the fault status code.

[0055] Fault standardization and encoding (encoding unit) In some embodiments, a set of fault encoding rules (implemented by hardware logic) can be pre-set inside the interrupt processor 100.

[0056] The encoding unit performs hardware-level encoding on the fault criteria of each interface card (such as power error, link disconnection, temperature overrun, parity error, etc.).

[0057] When performing encoding output, encoding information corresponding to each specific fault is generated (for example, a 5-bit binary code represents a fault type). To digitize and standardize complex physical fault signals, providing a basis for subsequent matching and priority determination.

[0058] When the hardware detects an actual interrupt signal (such as an error reported by an interface card) input to the interrupt processor 100, the matching unit is triggered. The matching unit performs a hardware logic comparison between the received interrupt signal and the encoded information library generated by the encoding unit to determine the fault status code corresponding to the interrupt signal. This status code not only contains the fault type (from the encoded information), but may also implicitly or be associated with the basic priority information (for example, some severe fault types inherently correspond to high-priority encodings), thereby determining the interrupt level.

[0059] For example, the interrupt processor 100 can encode and transmit the finally determined interrupt level through the INTx interrupt signal line (such as 4 groups * 8 signals) or a more economical solution (such as 3 IO pins).

[0060] Among them, an example of 8-level encoding using 3 IO (Input / Output): 3 binary IO pins can be combined into 2^3 = 8 states (from 000 to 111).

[0061] Each state represents an interrupt level (for example, 000 = level 0 / no interrupt, 001 = level 1, 010 = level 2,..., 111 = level 7).

[0062] According to different interrupt levels, embodiments of the present invention can determine the target processor to complete the fault handling, achieve resource optimization. The data processor 300 only processes truly urgent interrupts and avoids being overwhelmed by a large number of low-priority interrupts. The management processor 200 undertakes delayable tasks (such as logging, fan regulation), releasing the computing power of the data processor 300 for core services.

[0063] Optionally, in an embodiment of the present invention, the interrupt processor 100 includes: a first sending unit and a second sending unit.

[0064] Among them, the first sending unit is used to generate a corresponding emergency fault handling signal based on the interrupt signal and send the emergency fault handling signal to the data processor 300 through a preset physical link when the interrupt level is greater than the preset level.

[0065] The second sending unit is used to generate a corresponding non-emergency fault handling signal based on the interrupt signal and send the non-emergency fault handling signal to the management processor 200 through a preset access link when the interrupt level is less than or equal to the preset level.

[0066] According to different interrupt levels, embodiments of the present invention can select different target processors.

[0067] An embodiment of the present invention can preset a critical threshold (e.g., level 4) to distinguish between emergency interrupts (which require immediate processing by the data processor 300) and non-emergency interrupts (which can be processed by the management processor 200).

[0068] For example: if the preset level is 7, an interrupt level > 7 is an emergency interrupt, and an interrupt level ≤ 7 is a non-emergency interrupt.

[0069] The first sending unit can send emergency fault handling signals such as CPU power supply abnormality, uncorrectable memory error, system watchdog timeout, etc. to the data processor 300, enabling the data processor 300 to capture and process emergency events with the lowest latency (e.g., triggering a kernel crash dump).

[0070] The second sending unit can send non-emergency fault handling signals such as correctable memory errors, interface card hot plug events, fan speed regulation requests, etc. to the management processor 200. The management processor 200 performs out-of-band recording, alarming, or delayed processing to avoid occupying the computing resources of the data processor 300.

[0071] In addition, there are also some faults that require the collaborative processing of the management processor 200 and the data processor 300. According to different interrupt levels, interrupt signals that require collaborative processing can also be filtered out, and then the first sending unit and the second sending unit are used to send signals to the management processor 200 and the data processor 300 simultaneously for fault handling.

[0072] For example, as shown in Table 1, an embodiment of the present invention can determine the priority based on the interrupt signal and the trigger source, and then determine the target processor. Table 1 is a fault collaborative processing table.

[0073] Table 1

[0074] Optionally, in an embodiment of the present invention, the management processor 200 includes: a second response unit and a second processing unit.

[0075] Among them, the second response unit is used to respond to non-emergency fault handling signals, read the fault area data from the shared storage area of the management processor 200, calculate the fault score of each fault based on the status data of the controller, the status data of the data processor 300, and the fault area data, and arrange the priorities of multiple faults based on the fault scores to obtain a processing queue.

[0076] The second processing unit is used to perform corresponding processing actions based on the processing queue.

[0077] In the actual execution process, the second response unit of the embodiment of the present invention can receive non-emergency fault handling signals from the interrupt processor 100 (through a preset access link).

[0078] The second response unit can read the original information of the faulty device (such as sensor address, error register value), controller status data (such as the hardware status written by the interrupt processor 100, such as power / clock stability), and status data of the data processor 300 (the operating status reported by the main processor, such as load, temperature, memory ECC (Error-Correcting Code) count) from the shared storage area.

[0079] Based on the read data, the second response unit can perform a fault score. For example, preset weights based on the fault type (such as fan fault = 0.8, correctable memory error = 0.3), whether the fault is likely to affect other components (such as the risk value of power failure > overheating), the degree of occupation of system resources by the fault (such as the weight of the fan fault increases under high load), etc., and then prioritize the fault handling according to the score.

[0080] The second processing unit can then perform corresponding fault handling according to the sorting result of the second response unit.

[0081] Optionally, in an embodiment of the present invention, the second response unit includes: an acquisition subunit, an assignment subunit, and a calculation subunit.

[0082] Among them, the acquisition subunit is used to obtain the fault severity level and historical fault frequency of each fault from the fault area data, and obtain the current system latency sensitivity based on the status data of the data processor 300.

[0083] The assignment subunit is used to assign corresponding weights to the fault severity level, historical fault frequency, and current system latency sensitivity respectively.

[0084] The calculation subunit is used to calculate the fault score of each fault by using the weights, fault severity level, historical fault frequency, and current system latency sensitivity.

[0085] Among them, the calculation expression of the fault score is: P dynamic = α • S severity + β • L latency + γ • H history , Among them, P dynamic is the fault score, S severity is the fault severity level, Llatency is the current latency sensitivity of the system, H history is the historical fault frequency, α 、 β 、 γ are the weights of the fault severity level, the current latency sensitivity of the system, and the historical fault frequency, respectively.

[0086] An embodiment of the present invention can use a dynamic factor model to implement a dynamic priority adjustment algorithm to achieve the function of isolating server interface card faults. The core principle of the algorithm is as follows: The calculation expression of the fault score is: P dynamic = α • S severity + β • L latency + γ • H history , where, P dynamic is the fault score, S severity is the fault severity level (1 - 10 levels, e.g., PCIe link training failure = 10 levels, temperature overrun = 5 levels), L latency is the current latency sensitivity of the system (calculated based on the load rate of the data processor 300, linearly mapped from 0.1 - 1.0), H history is the historical fault frequency (exponential decay model H =1 - eT - λt , λ =0.05), α 、 β 、 γ are the weights of the fault severity level (e.g., α =0.6, β =0.3, γ =0.1), the current latency sensitivity of the system, and the historical fault frequency, respectively.

[0087] Among them, PCIe link training is a series of operations performed by PCIe devices during initialization or recovery, aiming to establish and optimize the communication link between devices. Ensure link stability, maximize data transfer rate, reduce bit error rate, etc.

[0088] Through the fault scoring adjustment algorithm, the core problems of complex fault causes and repeated faults are solved. According to the fault scoring adjustment algorithm, the root cause of the fault can be dynamically identified, and the accuracy and effectiveness of fault isolation and repair can be improved.

[0089] Optionally, in an embodiment of the present invention, the second processing unit includes: a matching subunit, a first processing unit, a second processing unit, a third processing unit, a fourth processing unit, and a fifth processing unit.

[0090] Among them, the matching subunit is used to match a processing action for any fault based on the fault score. The first processing unit is used to, when the fault score is in the first preset score range, determine that the processing action is to confirm the target fault device of any fault, cut off the power supply of the target fault device for isolation, and generate a corresponding alarm signal based on the target fault device.

[0091] The second processing unit is used to, when the fault score is in the second preset score range, determine that the processing action is to close the high-speed peripheral component interconnect interface of the target fault device and record any fault in the black box log, where the lower limit value of the first preset score range is greater than the upper limit value of the second preset score range.

[0092] The third processing unit is used to, when the fault score is in the third preset score range, determine that the processing action is to reduce the data transmission speed of the high-speed peripheral component interconnect interface of the target fault device to a preset speed threshold and start a link repair action, where the lower limit value of the second preset score range is greater than the upper limit value of the third preset score range.

[0093] The fourth processing unit is used to, when the fault score is in the fourth preset score range, determine that the processing action is to limit the bandwidth of the target fault device to a preset value and report any fault to the operating system, where the lower limit value of the third preset score range is greater than the upper limit value of the fourth preset score range.

[0094] The fifth processing unit is used to, when the fault score is in the fifth preset score range, determine that the processing action is to record any fault in the black box log, where the lower limit value of the fourth preset score range is greater than the upper limit value of the fifth preset score range.

[0095] Among them, the matching subunit can implement the matching of the fault score and the processing action.

[0096] Embodiments of the present invention can pre-construct mapping rules, determine corresponding dynamic score ranges according to the fault score, determine corresponding fault processing levels in different dynamic score ranges, and use the processing actions corresponding to the fault processing levels to perform fault processing, forming a fault classification processing system, avoiding over-processing, and optimizing resources.

[0097] For example, as shown in Table 2, Table 2 is a dynamic scoring range - level - processing action comparison table.

[0098] Table 2

[0099] Optionally, in an embodiment of the present invention, the data processor 300 includes: a first response unit and a first processing unit.

[0100] Among them, the first response unit is configured to respond to an emergency fault handling signal, determine at least one fault and the corresponding fault isolation level based on the status data and the interrupt signal of the interrupt processor 100, so as to match a corresponding fault isolation policy for each fault based on the fault isolation level; The first processing unit is configured to perform corresponding processing actions based on the fault isolation policy.

[0101] Furthermore, in an embodiment of the present invention, the first response unit may receive an emergency fault signal sent by the interrupt processor 100 through a physical direct connection link, determine the fault isolation level according to the status data of the interrupt processor 100 (obtained directly from the interrupt processor 100 through the physical link) and the interrupt signal, and perform corresponding processing actions through the first processing unit.

[0102] Among them, the fault isolation level may be determined according to the interrupt level of the interrupt signal, or may be determined by parsing. For example, the interrupt signal feature may be a parsed edge type (rising edge = instantaneous fault, continuous low level = permanent fault).

[0103] Directly obtaining the status data through the physical link can effectively improve the response speed.

[0104] Optionally, in an embodiment of the present invention, the data processor 300 is further configured to control the interrupt processor 100 to perform processing actions.

[0105] It can be understood that the interrupt processor 100 is a programmable logic device inside, and the data processor 300 can send instructions to it through a preset register interface, or can also be implemented through a pre - constructed physical link.

[0106] In terms of processing actions, the interrupt processor 100 can do many things. For example, dynamically adjusting the sensor sampling rate, which is very useful when troubleshooting intermittent faults; or switching redundant power modules to achieve seamless fault transfer; and can also forcibly isolate faulty PCIe devices, which is more thorough than software isolation. These operations all embody the idea of hardware autonomy - after the data processor 300 discovers an anomaly, it does not wait for the OS (operating system) to intervene, but directly commands the interrupt processor 100 to solve the problem at the underlying layer.

[0107] Optionally, in an embodiment of the present invention, the processing system 10 of the server further includes: a determination module and a repair module.

[0108] Among them, the determination module is used to obtain the historical fault data and service information of the fault corresponding to the interruption signal, and determine whether the fault meets the preset repair conditions in combination with the historical fault data and service information; The repair module is used to match the repair strategy corresponding to the fault that meets the preset repair conditions and perform the corresponding repair actions.

[0109] As a possible implementation manner, the determination module of the embodiment of the present invention can extract the historical fault data from the shared storage area, and determine whether the fault corresponding to the interruption signal can be repaired according to the historical fault data and service information (such as real-time service value weight, service tolerance, resource dependency graph, etc.).

[0110] When making a judgment, it can be determined from multiple aspects such as technical feasibility (such as whether the fault is known to be repairable), economic rationality (such as the comparison of the time consumed by automatic repair and manual processing), and risk controllability (such as the completeness of the rollback plan when the repair fails) to determine whether the repair conditions are met.

[0111] After the repair conditions are met, the embodiment of the present invention can perform the corresponding repair actions.

[0112] Optionally, in an embodiment of the present invention, the repair module includes: a determination unit, a first repair unit, a second repair unit, and a third repair unit.

[0113] Among them, the determination unit is used to determine the fault type of the fault that meets the preset repair conditions.

[0114] The first repair unit is used to perform a preset hot reset attempt action to repair the fault when the fault type is a preset zombie fault type.

[0115] The second repair unit is used to perform a preset firmware reloading action to repair the fault when the fault type is a preset firmware fault type.

[0116] The third repair unit is used to perform a preset link or interface reset action to repair the fault when the fault type is a preset connection fault type.

[0117] Among them, the repair module can perform corresponding fault repairs according to the fault type.

[0118] For the zombie fault type, the first repair unit can perform hot reset attempt actions such as hot restart and configuration recovery to repair the fault; For the firmware fault type, the second repair unit can perform a firmware reloading action to repair the fault; For the connection fault type, the third repair unit can perform a link or interface reset operation to repair the fault.

[0119] In addition, for the voltage drift fault, dynamic compensation power supply can be performed, for the data contamination fault, cache flushing and memory remapping can be performed, and for the transient link error, link retraining can be performed, etc.

[0120] The embodiment of the present invention can record each fault, repair determination result, repair operation and repair result, and then, when the repair result is successful, store the corresponding data for subsequent invocation. For the repair operations that require manual intervention, corresponding records can also be made, and it can be determined whether they can be handled automatically, so as to achieve automatic response when the corresponding fault occurs again next time.

[0121] Optionally, in an embodiment of the present invention, the processing system 10 of the server further includes: a judgment module and an alarm module.

[0122] Wherein, the judgment module is used to judge whether the processing operation meets the preset server operation impact conditions before the control data processor 300 or the management processor 200 executes the corresponding processing operation.

[0123] The alarm module is used to terminate the processing operation and generate an alarm signal when the preset server operation impact conditions are met.

[0124] As a possible implementation manner, the judgment module can perform a pre-evaluation of the impact of the processing operation.

[0125] The evaluation dimensions can include: performance impact index (predicting the decline range of CPU, memory, and IO performance), business continuity risk (analyzing the tolerance of associated service SLAs), and hardware safety boundary (detecting the safety margin of parameters such as voltage and temperature), etc. Among them, the thresholds related to each evaluation can be dynamically adjusted according to the specific business scenario.

[0126] When it is determined that the fault handling will affect the server operation, the embodiment of the present invention can stop the processing operation and report a warning, so as to perform corresponding fault handling according to the user's decision and avoid affecting the normal business processing of the server.

[0127] Optionally, in an embodiment of the present invention, after the management processor 200 detects a monitoring abnormal signal, the management processor 300 judges whether the interrupt processor 100 detects an interrupt signal, and judges whether the data processor 300 detects a status abnormal signal, and based on the data in the shared storage area, performs corresponding abnormal processing operations when the interrupt signal and the status abnormal signal are not detected.

[0128] In the actual execution process, exception detection is not only carried out by the interrupt processor 100. In fact, the interrupt processor 100, the management processor 200, and the data processor 300 can all detect exception signals. Generally speaking, when a failure occurs, two or more of the interrupt processor 100, the management processor 200, and the data processor 300 will detect the exception signal. At this time, if the interrupt processor 100 detects an interrupt signal, the interrupt level of the interrupt processor 100 is used as the priority processing method. If the interrupt processor 100 does not detect an interrupt signal, there may be a situation where the management processor 200 detects a monitoring exception signal while the interrupt processor 100 does not detect an interrupt signal. At this time, the management processor 200 processes the relevant faults of the monitoring exception signal.

[0129] For example, the management processor 200 can monitor the hardware status in real time through an integrated sensor array (temperature / voltage / fan speed sensor), and generate a monitoring exception signal when it detects that the preset threshold is exceeded.

[0130] At this time, through the special link structure of the embodiment of the present invention, the management processor 200 can directly obtain whether the interrupt processor 100 and the data processor 300 have detected an interrupt signal and a status exception signal.

[0131] In the case where no interrupt signal and status exception signal are detected, the data processor 200 can perform in-depth data analysis based on the data in the shared storage area. For example, parsing the sensor data matrix collected by the data processor 200 itself, constructing a time series waveform, cross-analyzing the performance counter data of the data processor 300, analyzing the frequency domain characteristics of the sensor data, using a pre-trained decision tree model to identify abnormal patterns, etc. [[ID=eleven]]

[0132] According to the analysis results, the management processor 200 can perform hierarchical processing. For example, in the case where the exception type is data drift type, sensor calibration and historical data correction are performed; in the case where the exception type is latent failure type, preventive frequency reduction and spare part preheating are performed; in the case where the exception type is environmental interference type, filtering enhancement is performed; for unknown exceptions, black box recording can also be performed.

[0133] Optionally, in an embodiment of the present invention, after the data processor 300 detects a status exception signal, the data processor 300 determines whether the interrupt processor 100 has detected an interrupt signal, determines whether the management processor 200 has detected a monitoring exception signal, and performs corresponding exception handling actions based on the data in the shared storage area in the case where no interrupt signal and monitoring exception signal are detected.

[0134] Similarly, there may also be a situation where only the data processor 300 detects an abnormal status signal. In this case, the data processor 300 can also confirm the status of the interrupt processor 100 and the management processor 200 through the special communication link of the embodiments of the present invention, and then, in the case of not detecting an interrupt signal and a monitoring abnormal signal, perform corresponding abnormal handling actions based on the data in the shared storage area.

[0135] For example, when the data processor 300 detects that the PCIe link status information of the interface card is abnormal, it can determine the actual abnormality according to the data in the shared storage area and perform corresponding abnormal handling actions. For example, if the abnormality is that the PCIe link communicating with the interface card is disconnected, it can attempt to reconnect or directly disconnect the link communicating with the interface card, rely on the data obtained by the interrupt processor 100 and the management processor 200 to maintain normal operation, and report the fault.

[0136] Combined with Figure 2 shown, the working principle of the processing system 10 of the server according to the embodiments of the present invention will be described in detail with an embodiment.

[0137] Currently, for the fault detection of the interface card component, it is generally collected by the BMC firmware polling and accessing the interface card chip to collect information such as the temperature of the component sensor and the chip link error count to determine. The BMC management link usually uses the low-speed I2C for access, and neither the access method nor the access period can guarantee the response speed of the fault detection, which is not conducive to the subsequent fault isolation or the implementation of the hardware solution. The interface card itself is a PCIe device, and the CPU can access the externally plugged PCIe device through the PCIe link. During the server startup phase or the running phase, the BIOS (Basic Input / Output System) firmware can collect the fault information of the externally plugged card device by detecting the PCIe link fault. Although the BIOS can transfer the fault information through the IPMI channel between the BIOS and the BMC in the existing solution, the separate out-of-band diagnostic fault handling may affect the service and lacks the ability of collaborative diagnosis and accurate fault handling. The embodiments of the present invention can set up a special communication structure, the StateSync Bus. This hardware bus adopts a three-terminal heterogeneous interconnection architecture, and through the cooperation of the hardware-level interconnection bus, it realizes the status synchronization and sharing among the three devices of the CPU, BMC, and CPLD, and at the same time can break the fault information islands in the system, enabling the rapid response and mutual access of the diagnostic information in the system. The hardware design of the StateSync Bus can be as Figure 2 shown. Taking the CPLD as the interrupt processor 100, the BMC as the management processor 200, and the CPU as the data processor 300 as an example, the processing system 10 of the server can include an interrupt processor 100, a management processor 200, and a data processor 300.

[0138] Among them, the interrupt processor 100 can communicate with multiple interface cards (interface card 1 401, interface card 2 402, interface card 3 403, and interface card 4 404) to obtain interrupt signals; the management processor 200 can also monitor the hardware status data of multiple interface cards through the BMC access channel via SW500 to obtain monitoring exception signals; the data processor 300 can communicate with multiple interface cards through the PCIe link to obtain status exception signals.

[0139] A shared PCIe device configuration space memory scheme with PCIe link interconnection is adopted between the data processor 300 and the management processor 200. As an EP (Endpoint Device) device, the management processor 200 can be directly read and written by the BIOS to access the shared storage area 201. Interrupts are used for rapid diagnosis and status monitoring between the data processor 300 and the interrupt processor 100, and between the interrupt processor 100 and the management processor 200. The data processor 300 and the interrupt processor 100 adopt an LPC (Low Pin Count Physical Link, a serial communication interface for connecting low-bandwidth and low-speed peripherals on a computer motherboard) physical link. The data processor 300 can access the interrupt processor 100 through Memory-Mapped I / O (Memory-Mapped I / O, by mapping the registers or storage space of I / O devices to the main memory address space, enabling the data processor 300 to access these devices through ordinary memory read and write instructions) to operate and control the interrupt controller. The interrupt processor 100 and the management processor 200 are directly interconnected for data communication through an I3C (Improved Inter Integrated Circuit) link. At the same time, the management processor 200 drives the layer application of DMA (Direct Memory Access) technology, which can realize the rapid synchronization of data between the interrupt processor 100 and the management processor 200. The interrupt signal lines of the interrupt processor 100 are respectively connected to the NMI interrupt pins of the data processor 300 and the interrupt control pins of the management processor 200.

[0140] In the actual execution process, taking the process of implementing fault handling after the interrupt processor 100 detects an interrupt signal as an example, the embodiments of the present invention may include the following steps: Step S1, interrupt detection. During the startup or operation of the server, the interrupt processor 100 monitors the PCIe key signals of the interface card components. This process is implemented by the logic code of the interrupt processor 100. The specific implementation logic includes: The interrupt processor 100 detects the status signal input detection (PCIe PRSNT in-position signal and PWR_GOOD power status, HeartBeat firmware status, AER_ALERT fault alarm) of the interface card; The internal priority arbiter of the interrupt processor 100 performs interrupt arbitration according to the current detection status code and the defined priority. The interrupt processor 100 uses logical codes to implement hardware-level interrupt aggregation, supporting 32 levels of priority. The interrupt controller is divided into an interrupt mask controller (32 bits), a priority arbiter (Fixed Priority), and INTx interrupt signal lines (4 * 8 groups). Actually, several IOs can be used for priority encoding. For example, 3 IOs are used as a group for priority encoding, and 8 different interrupt levels can be implemented corresponding to the priority; The interrupt processor 100 triggers an interrupt according to the interrupt mask register mask, and the triggering rule refers to the hardware fault interrupt vector allocation table.

[0141] Step S2, interrupt synchronization. The interrupt processor 100 internally implements a memory management module, and updates the fault status to the shared memory through the status synchronization bus. That is, when the hardware fault signal is triggered, the interrupt processor 100 can update the status data in the shared storage area 201 according to the interrupt type. The writing process uses atomic operations to avoid data synchronization exceptions caused by access by the interrupt processor 100 or the management processor 200. For the I3C link between the interrupt processor 100 and the management processor 200, the management processor 200 internally implements a shared memory access driver for the interrupt processor 100, and internally uses DMA to access the shared memory area. This solution can avoid the impact of the management processor 200's business on the alarm efficiency. The specific synchronization process is as follows: The interrupt processor 100 acquires the atomic lock atomic_lock. The implementation of this lock can utilize a fixed position in the shared memory or an external physical pin to achieve atomic access to the shared memory; Calculate the target address (address conversion), encode each PCIe device in the system, and implement data-unique mapped address access according to the address conversion algorithm; The interrupt processor 100 logically implements fault data framing. The frame data includes: fault type, timestamp, reserved bit, checksum; The interrupt processor 100 initiates a DMA write operation to write the data into the shared memory; The interrupt processor 100 releases the atomic lock.

[0142] Step S3, software collaborative processing.

[0143] Hardware signals such as the power signal, presence signal, and abnormal status of the external card component are detected by the interrupt processor 100. When the interrupt processor 100 detects an abnormal hardware signal, it triggers an interrupt to the BMC and BIOS for processing. The management processor 200 makes decisions based on the diagnostic results in a specific area of the shared storage 201, realizing the operation of the operating system without perception and the coordinated diagnosis between in-band and out-of-band.

[0144] According to the fault interrupt vector allocation table implemented in the logic of the interrupt processor 100, the interrupt processor 100 can preferentially trigger interrupts of the management processor 200 or the data processor 300 for high-priority faults. The interrupt processor 100 triggers an NMI (non-maskable interrupt) to wake up the BIOS interrupt service program. The BIOS reads the fault area data in the shared memory, judges the isolation level, and responds to the isolation measures. The BMC side runs a service for polling and monitoring the operating status. When it detects data updates in the shared memory, it calls a dynamic adjustment isolation algorithm to evaluate the fault and adopts different repair strategies according to the evaluation score.

[0145] The embodiments of the present invention can detect the faults of the interface card component through the above structure and steps, and adopt a mechanism for grading and isolating the faults of the interface card component. The fault levels are divided into: power-on isolation level, running isolation level, to-be-corrected isolation level, and minor fault level. The power-on isolation level indicates that the fault seriously affects system startup and operation. During the startup process, the device is hardware-disconnected from the system or powered off to ensure that it does not affect the normal startup of the device. The running isolation level means that during the operation of the server, when the device suddenly fails or has persistent functional abnormalities, the device is isolated, and the standby channel or port / communication link is switched to ensure that the fault does not spread. The to-be-corrected isolation level means that when a device component fails, after collaborative diagnosis, it is determined that the device has the possibility of repair, and the component will be restored through repair means. By adopting the fault grading and isolation mechanism, the purpose of ensuring that the fault does not spread and the minimum impact on the continuous operation of the service is achieved. Minor faults do not affect the service communication status and only record warning logs. The input of the external interface card interrupt status, the type of interrupt priority, and the fault information status synchronization implemented by the status synchronization bus can output the server interface card hardware fault interrupt vector allocation table, and the fault collaborative processing can be realized as shown in Table 1.

[0146] According to the above hardware design, the embodiments of the present invention can adopt a dynamic factor model to implement a dynamic priority adjustment algorithm to realize the fault isolation function of the server interface card. The core principle of the algorithm is as follows: The calculation expression of the fault score is: P dynamic = α • S severity + β• L latency + γ • H history , wherein P dynamic is the fault score, S severity is the fault severity level (1 - 10 levels, e.g., PCIe link training failure = level 10, temperature over - limit = level 5), L latency is the current system latency sensitivity (calculated based on the load rate of the data processor 300, linearly mapped from 0.1 - 1.0), H history is the historical fault frequency (exponential decay model H = 1 - eT - λt , λ = 0.05), α 、 β 、 γ are the weights of the fault severity level (e.g., α = 0.6, β = 0.3, γ = 0.1), the current system latency sensitivity, and the historical fault frequency respectively.

[0147] In summary, during actual fault handling, the embodiments of the present invention can cooperatively detect faults during the startup / operation process of the server interface card through the interrupt processor 100, the management processor 200, and the data processor 300 (CPLD / BMC / BIOS). The management processor 200 reads the fault information in the shared storage area 201, calls the dynamic priority evaluation algorithm, and the dynamic evaluation process needs to fully consider parameters such as the fault type, fault historical information, and the current system sensitivity, calculates the dynamic fault score, and distinguishes whether the fault is a high - priority fault or a low - priority fault. The high and low priorities can be set in a preset manner (e.g., high - priority P <= 7, low - priority P >= 8). High - priority faults trigger an immediate isolation process by interrupt, and adopt strategies such as power - off / link - off or triggering a fault alarm, while low - priority faults default to the delayed processing queue and wait for the management processor 200 to queue for processing. After the alarm is triggered, a recovery strategy decision is triggered, and it is decided whether to repair the component based on information such as historical fault data and business importance. The repair strategy includes operations such as attempting a warm reset, re - loading firmware, and resetting the PCIe link port. Different repair strategies can be preset and bound according to different device types.

[0148] Step S4, interrupt masking strategy.

[0149] In the embodiments of the present invention, an interrupt mask register can be reserved. The data processor 300 or the management processor 200 can perform mask design for relevant faults according to the actual scenario to ensure that the isolation and repair will not affect the normal function startup.

[0150] In summary, the three-level interrupt fusion architecture of the embodiments of the present invention can integrate hardware signal triggering, in-band interrupts, and out-of-band alarms into a unified status bus, break the fault information islands in the system, and improve the timeliness and effectiveness of fault alarms; through the fault priority classification strategy in the interrupt processor 100, different response entities are allocated to reduce the status synchronization delay and the concurrent fault handling ability; through the unified status bus data synchronization mechanism, each processor under the system 10 of the server can share fault information, and the fault handling process can be synchronized and shared in real time to avoid the risk of fault spread under the system; and a fault score adjustment algorithm is proposed to solve the core problems of complex fault causes and repeated occurrence of faults. According to the fault score adjustment algorithm, the root cause of the fault can be dynamically identified, and the accuracy and effectiveness of fault isolation and repair can be improved.

[0151] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0152] As Figure 3 shown, the embodiments of the present invention also provide a processing method for a server, including the following steps: In step S301, after the interrupt processor detects an interrupt signal, the current interrupt level of the server is determined according to the interrupt signal.

[0153] In step S302, a target processor is determined based on the current interrupt level to control the target processor to perform corresponding processing actions on the device based on the interrupt signal or status data.

[0154] For the description of the features in the corresponding embodiments of the processing method of the server, reference can be made to the relevant descriptions in the corresponding embodiments of the processing system of the server, which will not be elaborated here one by one.

[0155] The embodiments of the present invention also provide an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above embodiments of the processing method of the server.

[0156] The embodiments of the present invention also provide a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above embodiments of the processing method of the server when running.

[0157] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media that can store computer programs, such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disks, magnetic disks, or optical discs.

[0158] An embodiment of the present invention also provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described server processing method embodiments.

[0159] An embodiment of the present invention also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described server processing method embodiments.

[0160] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0161] The above has introduced in detail a server processing system, method, electronic device, and storage medium provided by the present invention. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A processing system for a server, characterized in that, It includes an interrupt processor, a management processor, and a data processor. After the interrupt processor detects an interrupt signal, the interrupt processor communicates the interrupt signal and status data with the data processor based on a preset physical link, the interrupt processor communicates the interrupt signal and status data with the management processor based on a preset access link, and the data processor communicates the interrupt signal and status data based on a PCIe link. Among them, the interrupt processor determines the current interrupt level of the server according to the interrupt signal, and determines a target processor based on the current interrupt level, so as to control the target processor to perform corresponding processing actions on at least one device involved in the interrupt signal based on the interrupt signal or the status data.

2. The processing system of the server according to claim 1, characterized in that The management processor has a shared storage area. Among them, the interrupt signal detected by the interrupt processor and the status data of the interrupt processor are written into the shared storage area through the preset access link. The status exception signal detected by the data processor, the status data of the data processor, and the server operation data collected by the data processor are written into the shared storage area through the PCIe link. The monitoring exception signal detected by the management processor, the status data of the management processor, and the hardware status data collected by the management processor are written into the shared storage area.

3. The processing system of the server according to claim 2, characterized in that, The interrupt processor, the management processor, and the data processor write the interrupt signal and the corresponding status data into the shared storage area based on a preset atomic operation to prevent the interrupt processor, the management processor, and the data processor from writing data simultaneously.

4. The processing system of the server according to claim 2, characterized in that After the management processor detects the monitoring exception signal, the management processor determines whether the interrupt processor detects the interrupt signal, and determines whether the data processor detects a status exception signal, and in the case of not detecting the interrupt signal and the status exception signal, performs corresponding exception handling actions based on the data in the shared storage area.

5. The processing system of the server according to claim 2, wherein After the data processor detects the status exception signal, the data processor determines whether the interrupt processor detects the interrupt signal, and determines whether the management processor detects a monitoring exception signal, and in the case of not detecting the interrupt signal and the monitoring exception signal, performs corresponding exception handling actions based on the data in the shared storage area.

6. The processing system of the server according to claim 1, characterized in that, The interrupt processor includes: an encoding unit for performing fault encoding on each fault criterion of the interface card based on a preset encoding rule to generate corresponding encoding information; a matching unit for matching a corresponding fault status code by combining the encoding information and the interrupt signal; a determining unit for determining the interrupt level according to the fault status code.

7. The processing system of the server according to claim 1, characterized in that the interrupt processor includes: A first sending unit, configured to, when the interruption level is greater than the preset level, generate a corresponding emergency fault handling signal based on the interruption signal, and send the emergency fault handling signal to the data processor through the preset physical link; A second sending unit, configured to, when the interruption level is less than or equal to the preset level, generate a corresponding non-emergency fault handling signal based on the interruption signal, and send the non-emergency fault handling signal to the management processor through the preset access link.

8. The processing system of the server according to claim 7, characterized in that The data processor includes: A first response unit, configured to, in response to the emergency fault handling signal, determine at least one fault and a corresponding fault isolation level based on the status data of the interruption processor and the interruption signal, so as to match a corresponding fault isolation policy for each fault based on the fault isolation level; A first processing unit, configured to perform corresponding processing actions based on the fault isolation policy.

9. The processing system of the server according to claim 7, wherein The management processor includes: A second response unit, configured to, in response to the non-emergency fault handling signal, read fault area data from the shared storage area of the management processor, calculate a fault score for each fault based on the status data of the controller, the status data of the data processor, and the fault area data, and arrange multiple faults in priority order based on the fault scores to obtain a processing queue; A second processing unit, configured to perform corresponding processing actions based on the processing queue.

10. The processing system of the server according to claim 9, wherein The second response unit includes: An obtaining subunit, configured to obtain the fault severity level and the historical fault frequency of each fault from the fault area data, and obtain the current system delay sensitivity based on the status data of the data processor; An assignment subunit, configured to assign corresponding weights to the fault severity level, the historical fault frequency, and the current system delay sensitivity respectively; A calculation subunit, configured to calculate the fault score of each fault by using the weights, the fault severity level, the historical fault frequency, and the current system delay sensitivity.

11. The processing system of the server according to claim 10, wherein, The calculation expression of the fault score is: P dynamic = α • S severity + β • L latency + γ • H history , wherein, P dynamic is the fault score, S severity is the fault severity level, L latency is the current delay sensitivity of the system, H history is the historical fault frequency, α and β and γ are the weights of the fault severity level, the current delay sensitivity of the system, and the historical fault frequency, respectively.

12. The processing system of the server according to claim 9, characterized in that, The second processing unit includes: A matching subunit, configured to match a processing action for any one of the faults based on the fault score; A first processing unit, configured to, when the fault score is within a first preset score range, determine the processing action as identifying a target fault device of any one of the faults, performing power-off isolation on the target fault device, and generating a corresponding alarm signal based on the target fault device; A second processing unit, configured to, when the fault score is within a second preset score range, determine the processing action as closing the high-speed peripheral component interconnect interface of the target fault device, and recording any one of the faults in a black box log, where the lower limit value of the first preset score range is greater than the upper limit value of the second preset score range; A third processing unit, configured to determine, when the fault score is in a third preset score range, that the processing action is to reduce the data transmission speed of the high-speed peripheral component interconnect interface of the target faulty device to a preset speed threshold, and initiate a link repair action, where a lower limit value of the second preset score range is greater than an upper limit value of the third preset score range; A fourth processing unit, configured to determine, when the fault score is in a fourth preset score range, that the processing action is to limit the bandwidth of the target faulty device to a preset value, and report the any fault to an operating system, where a lower limit value of the third preset score range is greater than an upper limit value of the fourth preset score range; A fifth processing unit, configured to determine, when the fault score is in a fifth preset score range, that the processing action is to record the any fault in a black box log, where a lower limit value of the fourth preset score range is greater than an upper limit value of the fifth preset score range.

13. The processing system of the server according to claim 1, wherein Further included: A determination module, configured to obtain historical fault data and service information of a fault corresponding to the interrupt signal, and determine, in combination with the historical fault data and the service information, whether the fault meets a preset repair condition; A repair module, configured to match a repair strategy corresponding to a fault that meets the preset repair condition, and execute a corresponding repair action.

14. The processing system of the server according to claim 13, wherein The repair module includes: A determination unit, configured to determine a fault type of a fault that meets the preset repair condition; A first repair unit, configured to execute a preset hot reset attempt action for fault repair when the fault type is a preset frozen fault type; A second repair unit, configured to execute a preset firmware reloading action for fault repair when the fault type is a preset firmware fault type; A third repair unit, configured to execute a preset link or interface reset action for fault repair when the fault type is a preset connection fault type.

15. The processing system of the server according to claim 1, characterized in that, Further included: A judgment module, configured to judge whether the processing action meets a preset server operation influence condition before controlling the data processor or the management processor to execute a corresponding processing action; An alarm module, configured to terminate the processing action and generate an alarm signal when the preset server operation influence condition is met.

16. The processing system of the server according to claim 1, characterized in that, The data processor is further configured to control the interrupt processor to execute the processing action.

17. A processing method for a server, characterized in that Applied to a processing system of a server as described in any one of claims 1-16, where the method includes the following steps: After the interrupt processor detects an interrupt signal, determining a current interrupt level of the server according to the interrupt signal; Determining a target processor based on the current interrupt level, so as to control the target processor to execute a corresponding processing action on the device based on the interrupt signal or the status data.

18. An electronic device, characterized in that, Including: A memory, configured to store a computer program; A processor, configured to implement the steps of the processing method of the server as described in claim 17 when executing the computer program.

19. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein when the computer program is executed by a processor, the steps of the processing method of the server as described in claim 17 are implemented.

20. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the processing method of the server as described in claim 17 are implemented.

Citation Information

Patent Citations

  • Interrupt processing method, device and system

    CN103473191A

  • Apparatus and method for shortening interrupt latency

    JP2010140239A

  • Interrupt thresholding for SMT and multi processor systems

    US20060112208A1

Cited By

  • Signal processor access method and computer equipment

    CN120892095A

  • A signal processor access method and computer device

    CN120892095B

  • Hardware fault real-time detection method and system based on cooperation of CPU and BMC

    CN121008967A

  • PCIe equipment loss retry method, device and system

    CN122019241A