Computing device and interrupt storm management method of computing device

By dynamically adjusting the interrupt storm detection cycle and threshold, combining AI model and user configuration, the problem of inadequate interrupt storm judgment standards for computing equipment is solved, and system reliability and performance are improved.

CN120276892APending Publication Date: 2025-07-08XFUSION DIGITAL TECH CO LTD

Patent Information

Application Number
CN202510112802.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, the interrupt storm detection cycle and interrupt storm threshold of computing equipment are fixed and cannot adapt to the difference in memory CE errors in different scenarios, resulting in system performance being affected or problems not being discovered in time, affecting system reliability and performance.

Method used

By dynamically adjusting the interrupt storm detection cycle and interrupt storm threshold in the kernel running state, combining AI model and user configuration, it can be flexibly adjusted according to memory CE error feature information to meet the needs of different scenarios.

Benefits of technology

It improves the system reliability and performance of computing devices, avoids the impact of system performance, and realizes intelligent management and resource optimization of interrupt storms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276892A_ABST
    Figure CN120276892A_ABST
Patent Text Reader

Abstract

The invention discloses computing equipment and an interrupt storm management method of the computing equipment, the computing equipment comprises a kernel and a memory, the memory is configured with memory variables related to correctable machine check interruptions, and the memory variables are used for storing interrupt storm detection periods and interrupt storm thresholds corresponding to the correctable machine check interruptions. Under the condition that the kernel is in a running state, in response to a received configuration instruction for an interrupt storm detection period and an interrupt storm threshold value, modifying a variable value, corresponding to the interrupt storm detection period, of the memory variable into a first target value, and modifying a variable value, corresponding to the interrupt storm threshold value, of the memory variable into a second target value, the configuration instruction carries a first target value and a second target value. Thus, under the condition that the kernel is in the running state, the interrupt storm detection period and the interrupt storm threshold can be dynamically adjusted so as to be suitable for detection and determination of the interrupt storm in different scenes, and therefore the system reliability of the computing device can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a computing device and a method for managing interrupt storms of a computing device. Background Art

[0002] With the development of computing devices such as servers, the memory capacity supported by the computing device is getting larger and the operating frequency is getting higher, resulting in more and more memory errors, especially that the memory frequently generates correctable errors (CE). Currently, if the computing device detects that the occurrence times of the memory CE error reach the set error threshold, a corrected machine check interrupt (CMCI) will be triggered. Moreover, if the computing device detects that the number of CMCI interrupts within a certain interrupt storm detection period reaches the set interrupt storm threshold, it is considered that an interrupt storm has occurred. Currently, both the interrupt storm detection period and the interrupt storm threshold are fixed values set in the computing device during the initialization process of the computing device. Therefore, the judgment criterion for the interrupt storm is fixed and cannot meet the different requirements for the interrupt storm judgment criterion in different scenarios, resulting in problems such as affecting the system performance of the computing device. Summary of the Invention

[0003] This application provides a computing device and a method for managing interrupt storms of a computing device, which can realize the dynamic adjustment of the interrupt storm detection period and the interrupt storm threshold to meet the different requirements for the interrupt storm judgment criterion in different scenarios, thereby improving the system reliability of the computing device, ensuring the system performance or avoiding the impact on the system performance.

[0004] To solve the above technical problems, in a first aspect, an embodiment of this application provides a method for managing interrupt storms of a computing device. The computing device includes a kernel and a memory, and a memory variable related to the corrected machine check interrupt is configured in the memory. The memory variable is used to store the interrupt storm detection period and the interrupt storm threshold corresponding to the corrected machine check interrupt, and the interrupt storm detection period and the interrupt storm threshold are used for determining the interrupt storm. The method includes: when the kernel is in a running state, in response to receiving a configuration instruction for the interrupt storm detection period and the interrupt storm threshold, modifying the variable value of the memory variable corresponding to the interrupt storm detection period to a first target value, and modifying the variable value of the memory variable corresponding to the interrupt storm threshold to a second target value. The configuration instruction carries the first target value and the second target value.

[0005] Subsequently, determine whether an interrupt storm has occurred according to the occurrence time and occurrence times of the corrected machine check interrupt, and the interrupt storm detection period and the interrupt storm threshold stored in the memory variable.

[0006] The computing device is, for example, a server. There are differences in the process, physical quality, model, etc. of the memory included in different types of servers. Therefore, different servers with different memories have different adaptabilities to memory CE errors and corresponding interrupt storms. The generation of an interrupt storm causes the processor to continuously respond to these interrupt requests, thereby occupying a large amount of processor resources and memory resources and affecting the normal operation of the system. Therefore, if the interrupt storm judgment criterion is too strict, the system performance will be affected. However, if the interrupt storm judgment criterion is too loose, there will sometimes be problems such as the timely discovery of system problems being affected. Therefore, for different memories, different types of memory CE errors, and other different scenarios, only setting a fixed interrupt storm detection period and interrupt storm threshold cannot accurately reflect the differences in memory process, physical quality, etc., and cannot well meet the different requirements for the interrupt storm judgment criterion in different scenarios, resulting in problems such as affecting the system reliability of the computing device and thus affecting the system performance.

[0007] Based on the interrupt storm management method provided in this embodiment, when the kernel of the computing device is in a running state, the interrupt storm detection period and the interrupt storm threshold in the computing device can be dynamically adjusted, and it is determined whether an interrupt storm occurs according to the dynamically adjusted interrupt storm detection period and interrupt storm threshold, which can better meet the different requirements for the interrupt storm judgment criterion in different scenarios, thereby improving the system reliability of the computing device, ensuring the system performance or avoiding the system performance from being affected.

[0008] In a possible implementation of the first aspect above, the first target value and the second target value can be obtained based on the user's setting operations for the interrupt storm detection period and the interrupt storm threshold, and the configuration instruction can be generated in response to the obtaining of the first target value and the second target value for dynamically adjusting the interrupt storm detection period and the interrupt storm threshold.

[0009] In this way, the user can set the required interrupt storm detection period and interrupt storm threshold according to the specific requirements and usage scenarios of the computing device, which can better meet the different requirements of the user for the interrupt storm judgment criterion in different scenarios, thereby improving the system reliability of the computing device, ensuring the system performance or avoiding the system performance from being affected.

[0010] In a possible implementation of the above first aspect, the first target value and the second target value can be obtained based on the prediction results of the target model for the interruption storm detection period and the interruption storm threshold, and the configuration instruction can be generated in response to the obtaining of the first target value and the second target value for dynamically adjusting the interruption storm detection period and the interruption storm threshold. Additionally, the target model is obtained by training based on the historical information corresponding to the correctable machine check interruption, and the historical information includes the feature information of the historical memory correctable error and the corresponding historical interruption storm detection period and historical interruption storm threshold. Further, the feature information can be, for example, type information, and of course, it can also be other feature information of the memory correctable error.

[0011] With the continuous development of Artificial Intelligence (AI) technology, more and more computing devices have AI capabilities. Therefore, the prediction of the interruption storm detection period and the interruption storm threshold can be realized relying on the AI capabilities of the computing devices. That is, a target model for predicting the interruption storm detection period and the interruption storm threshold can be trained based on the AI capabilities, and the target model is used to predict the interruption storm detection period and the interruption storm threshold, so as to more flexibly and dynamically adapt to the different requirements of users for the interruption storm judgment criteria in different scenarios, thereby improving the system reliability of the computing device, ensuring system performance or avoiding the impact on system performance. Moreover, the adjustment efficiency of the interruption storm detection period and the interruption storm threshold can also be improved.

[0012] In a possible implementation of the above first aspect, obtaining the first target value and the second target value based on the prediction results of the target model for the interruption storm detection period and the interruption storm threshold includes: obtaining target information, where the target information includes the feature information of the current memory correctable error; inputting the target information into the target model for model analysis and processing, and obtaining the prediction results of the interruption storm detection period and the interruption storm threshold as the first target value and the second target value respectively.

[0013] Based on the feature information such as the type information of the memory correctable error generated in the current scenario in the computing device, the interruption storm detection period and the interruption storm threshold of the memory correctable error that are more adaptable to this feature can be analyzed through the target model, so as to better manage the memory correctable error, thereby improving the system reliability of the computing device, ensuring system performance or avoiding the impact on system performance.

[0014] In a possible implementation of the above first aspect, the first target value and the second target value are obtained based on the business scenario type of the computing device, and the configuration instruction can be generated in response to the obtaining of the first target value and the second target value for dynamically adjusting the interruption storm detection period and the interruption storm threshold.

[0015] In a possible implementation of the first aspect above, obtaining the first target value and the second target value based on the business scenario type of the computing device includes: determining the business scenario type of the computing device; and obtaining the first target value and the second target value according to the business scenario type and the configuration policy corresponding to the business scenario type for the interrupt storm detection period and the interrupt storm threshold.

[0016] In a possible implementation of the first aspect above, the configuration policies corresponding to different business scenario types may be different. Of course, the configuration policies corresponding to some business scenario types may also be the same.

[0017] In a possible implementation of the first aspect above, determining the business scenario type of the computing device includes determining through at least one of the following pieces of information: the system type information of the computing device; the system performance information of the computing device; the business characteristic information of the computing device.

[0018] In a possible implementation of the first aspect above, according to the difference in the system type of the computing device or the difference in the involved business, the business scenario type may include, for example, a high-sensitivity scenario, a balanced sensitivity and performance scenario, a low-sensitivity scenario, a dynamic adjustment scenario, a grading scenario, etc. Of course, it may also include other scenarios, which can be set as needed.

[0019] In a possible implementation of the first aspect above, the configuration policy includes the pre-configured first target value and second target value.

[0020] In a possible implementation of the first aspect above, the configuration policy includes determining the first target value and the second target value according to at least one of the following pieces of information: the characteristic information of correctable machine check interrupts; the characteristic information of correctable errors; the system performance information of the computing device; the business characteristic information of the computing device.

[0021] In this way, based on the business scenario type corresponding to the computing device and the corresponding configuration policy, the interrupt storm detection period and the interrupt storm threshold of memory correctable errors that are more suitable for this business scenario can be analyzed and obtained, so that the memory correctable errors can be better managed, thereby improving the system reliability of the computing device, ensuring system performance or preventing system performance from being affected.

[0022] Of course, in a possible implementation of the first aspect above, the first target value and the second target value may also be obtained based on other methods.

[0023] In a possible implementation of the first aspect above, the method further includes: during the kernel startup process, initializing the variable values of the memory variables corresponding to the interrupt storm detection period and the variable values corresponding to the interrupt storm threshold according to the configuration file, and creating an interface file for modifying the memory variables.

[0024] In a possible implementation of the above first aspect, the configuration file includes a default interrupt storm detection period and a default interrupt storm threshold for interrupt storm detection period interruption and interrupt storm threshold initialization.

[0025] Based on this, during the kernel startup process of the computing device, according to the default interrupt storm detection period and the default interrupt storm threshold in the configuration file, the computing device initializes the variable values corresponding to the interrupt storm detection period and the variable values corresponding to the interrupt storm threshold in the memory variables, so as to store the default interrupt storm detection period and the default interrupt storm threshold into the memory variables as the default interrupt storm judgment criteria. After the kernel of the computing device is in the running state, it determines whether an interrupt storm has occurred according to the default interrupt storm detection period and the default interrupt storm threshold stored in the memory variables. If the default interrupt storm detection period and the default interrupt storm threshold stored in the memory variables are dynamically modified subsequently, the modified new interrupt storm detection period and interrupt storm threshold are used as the new interrupt storm judgment criteria to determine whether an interrupt storm has occurred.

[0026] In addition, during the kernel startup process, an interface file for modifying the memory variables can also be created, and when it is necessary to modify the interrupt storm detection period and the interrupt storm threshold, the modification can be based on the interface file. In this way, the modification of the interrupt storm detection period and the interrupt storm threshold can be conveniently implemented.

[0027] In a possible implementation of the above first aspect, the kernel is also configured with a flag information for indicating whether it is allowed to perform a modification operation on the memory variables. The method further includes: when it is determined according to the flag information that it is allowed to perform a modification operation on the memory variables, modifying the variable value corresponding to the interrupt storm detection period in the memory variables to a first target value, and modifying the variable value corresponding to the interrupt storm threshold in the memory variables to a second target value.

[0028] In a possible implementation of the above first aspect, before modifying the variable value corresponding to the interrupt storm detection period in the memory variables to a first target value and modifying the variable value corresponding to the interrupt storm threshold in the memory variables to a second target value, the method further includes: modifying the flag information so that the flag information indicates that the memory variables are currently being modified and parallel modification operations on the memory variables are not allowed; after modifying the variable value corresponding to the interrupt storm detection period in the memory variables to a first target value and modifying the variable value corresponding to the interrupt storm threshold in the memory variables to a second target value, the method further includes: modifying the flag information so that the flag information indicates that it is allowed to perform a modification operation on the memory variables.

[0029] In a possible implementation of the above first aspect, the method further includes: when it is determined according to the marking information that a modification operation on the memory variable is not allowed, determining whether to allow a modification operation on the memory variable based on a preset period according to the marking information.

[0030] Based on the setting of the marking information and determining whether to perform a modification operation on the memory variable according to the marking information, parallel modification of the memory variable can be effectively prevented, ensuring the accuracy and security of the modification of the memory variable.

[0031] In a possible implementation of the above first aspect, modifying the variable value of the memory variable corresponding to the interrupt storm detection period to a first target value and modifying the variable value of the memory variable corresponding to the interrupt storm threshold to a second target value includes: parsing the configuration instruction to obtain the first target value and the second target value; verifying the first target value and the second target value; when the verification passes, determining whether to allow a modification operation on the memory variable according to the marking information, and when allowing a modification operation on the memory variable, modifying the variable value of the memory variable corresponding to the interrupt storm detection period to the first target value and modifying the variable value of the memory variable corresponding to the interrupt storm threshold to the second target value.

[0032] In a possible implementation of the above first aspect, when the verification fails, an error message is generated.

[0033] The verification can be, for example, verifying the compliance of the field lengths of the first target value and the second target value, etc., to improve the accuracy of writing the interrupt storm detection period and the interrupt storm threshold into the memory variable and ensure the normal progress of subsequent processing. When the verification fails, generating an error message can be used for the user to perform corresponding processing or subsequent processing, etc.

[0034] In a possible implementation of the above first aspect, determining whether an interrupt storm has occurred according to the occurrence time and occurrence frequency of the correctable machine check interrupt, and the interrupt storm detection period and the interrupt storm threshold stored in the memory variable includes: if it is determined according to the occurrence time and occurrence frequency of the correctable machine check interrupt that the occurrence frequency of the correctable machine check interrupt within the interrupt storm detection period is greater than or equal to the interrupt storm threshold, it is determined that an interrupt storm has occurred; if it is determined according to the occurrence time and occurrence frequency of the correctable machine check interrupt that the occurrence frequency of the correctable machine check interrupt within the interrupt storm detection period is less than the interrupt storm threshold, it is determined that no interrupt storm has occurred.

[0035] Second aspect, an embodiment of the present application provides a computing device, which includes a kernel and a memory. A memory variable related to a correctable machine check interrupt is configured in the memory. The memory variable is used to store an interrupt storm detection period and an interrupt storm threshold corresponding to the correctable machine check interrupt. The interrupt storm detection period and the interrupt storm threshold are used for determining an interrupt storm. The kernel is configured with a first function, and the first function includes: when the kernel is in a running state, in response to receiving a configuration instruction for the interrupt storm detection period and the interrupt storm threshold, modifying the variable value corresponding to the interrupt storm detection period of the memory variable to a first target value, and modifying the variable value corresponding to the interrupt storm threshold of the memory variable to a second target value for determining an interrupt storm. The configuration instruction carries the first target value and the second target value.

[0036] In a possible implementation of the above second aspect, the kernel is configured with a second function, and the second function includes: during the kernel startup process, according to a configuration file, initializing the variable value corresponding to the interrupt storm detection period of the memory variable and the variable value corresponding to the interrupt storm threshold for determining an interrupt storm, and creating an interface file for modifying the memory variable for modifying the variable value corresponding to the interrupt storm detection period of the memory variable and the variable value corresponding to the interrupt storm threshold.

[0037] Third aspect, an embodiment of the present application provides a computing device, including: a memory for storing a computer program, where the computer program includes program instructions; a processor for executing the program instructions so that the computing device executes the interrupt storm management method provided in the first aspect and / or any possible implementation manner of the first aspect as described above.

[0038] Fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions, and the program instructions are run by a computing device so that the computing device executes the interrupt storm management method provided in the first aspect and / or any possible implementation manner of the first aspect as described above.

[0039] Fifth aspect, the present application provides a computer program product. When the computer program product runs on a computing device, it causes the computing device to execute the interrupt storm management method provided in the first aspect and / or any possible implementation manner of the first aspect as described above.

[0040] It can be understood that the beneficial effects of the above second aspect to the fifth aspect can refer to the relevant descriptions in the first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] To more clearly illustrate the technical solutions provided by the embodiments of the present application, the accompanying drawings used in the description of the embodiments will be briefly introduced below.

[0042] Figure 1 Schematic diagram of a system architecture of a server provided by an embodiment of the present application;

[0043] Figure 2 Schematic diagram of a function description of a server provided by an embodiment of the present application;

[0044] Figure 3 Schematic diagram of a process of an interrupt storm management method for a server provided by an embodiment of the present application;

[0045] Figure 4 Schematic diagram of an implementation process of an interrupt storm management method for a server provided by an embodiment of the present application;

[0046] Figure 5 Another schematic diagram of an implementation process of an interrupt storm management method for a server provided by an embodiment of the present application;

[0047] Figure 6 Schematic diagram of a structure of a computing device provided by an embodiment of the present application. Detailed implementation manners

[0048] The computing device provided by the embodiment of the present application is, for example, a server. Exemplarily, it may be, for example, an x86 (Intel) server.

[0049] The main structure of the server provided by the embodiment of the present application and the main professional terms involved in the embodiment of the present application will be explained below.

[0050] As Figure 1 shown, a server generally includes a hardware layer, an operating system layer, and an application layer. Of course, a server may also include a network layer, a database layer, etc., which will not be elaborated here.

[0051] The hardware layer is the basis of the server architecture and generally includes hardware such as a central processing unit (CPU), a memory controller, a memory, and hardware registers.

[0052] The CPU, as the operation and control core of a computer system, is the final execution unit for information processing and program operation.

[0053] The memory is usually used to store operation data in the CPU and data such as data exchanged with external memories such as hard disks.

[0054] The memory controller is an important part of the computer system that controls the memory internally and is responsible for the data exchange between the memory and the CPU.

[0055] Hardware registers are some small storage devices inside the CPU used to store data.

[0056] The operating system layer is the interface layer between server software applications and hardware. It is responsible for managing and controlling server hardware resources, providing various services and functions, and can implement functions related to the server operation and maintenance system.

[0057] The operating system layer usually includes the operating system kernel, simply referred to as the kernel. The kernel is an intermediate layer between the hardware and the application layer, responsible for passing requests from the application layer to the hardware and acting as a low-level driver to address various devices and components in the operating system.

[0058] The kernel includes a reliability, availability, and serviceability (RAS) subsystem. RAS is a system that combines hardware and software, used to detect, report, and repair hardware errors to ensure the long-term normal operation of the system.

[0059] The application layer is a collection of software applications running on the server, including multiple software applications such as Application 1, Application 2, etc. Application 1 and Application 2 can be, for example, applications for managing hardware errors.

[0060] Memory corrected error (CE) refers to an error that occurs in the memory and can be detected and corrected by the hardware itself. The hardware can detect and repair errors through error correction codes, avoiding impacts on system performance.

[0061] Corrected machine check interrupt (CMCI) is a mechanism used to control that when the number of memory CE errors detected by the CPU or the memory controller reaches a specific value (such as the set interrupt threshold), a CMCI interrupt is triggered to notify the corresponding application for processing. By setting the interrupt storm threshold corresponding to the CMCI interrupt, the interrupt storm caused by frequent memory CE errors can be avoided, which affects system performance.

[0062] Interrupt storm refers to the system receiving a large number of interrupt requests corresponding to memory CE errors within a short period. The generation of an interrupt storm will cause the processor to continuously respond to these interrupt requests, thus occupying a large amount of processor resources and memory resources and affecting the normal operation of the system.

[0063] Memory variables are variables that exist independently in memory outside the table structure. Memory variables can be used to store data. When defining a memory variable, a name (i.e., variable name) and an initial value (i.e., variable value) need to be assigned to it. After a memory variable is created, it is stored in memory.

[0064] In the server field, the memory reliability in the kernel RAS subsystem significantly affects system performance. Although these characteristics usually do not directly cause serious impacts on the reliability of system data, they will significantly affect system performance and increase operation and maintenance costs. In various business scenarios, the application must be able to perceive the occurrence of memory CE errors in a timely or delayed manner so as to perform corresponding processing.

[0065] Currently, the control strategy of the RAS subsystem for reporting memory CE errors is as follows: set the interrupt threshold corresponding to the memory CE error. The interrupt threshold is, for example, 1, and set the interrupt storm detection period and interrupt storm threshold corresponding to the CMCI interrupt of the memory CE error. The interrupt storm detection period is, for example, 1 second, and the interrupt storm threshold is, for example, 15. The interrupt threshold can be stored through the aforementioned hardware register, and the interrupt storm detection period and interrupt storm threshold can be stored through the aforementioned memory variable. The memory variable includes a variable name and a variable value. Then, the variable name corresponding to the interrupt storm detection period is, for example, "interrupt storm detection period", and the corresponding variable value is "1", and the variable name corresponding to the interrupt storm threshold is, for example, "interrupt storm threshold", and the corresponding variable value is "15". Of course, different variable names and variable values can also be represented by binary data, characters and other information.

[0066] During the initialization process of the kernel startup, the RAS subsystem initializes the interrupt threshold stored in the hardware register according to the preset configuration file, and initializes the interrupt storm detection period and interrupt storm threshold stored in the memory variable according to the preset configuration file. During the kernel operation phase, the CPU or the memory processor determines whether to trigger the CMCI interrupt and report the CMCI interrupt to the RAS subsystem according to the number of detected memory CE errors and the interrupt threshold in the hardware register. The RAS subsystem determines whether an interrupt storm has occurred according to the occurrence time and number of the received CMCI interrupts and the interrupt storm detection period and interrupt storm threshold stored in the memory variable.

[0067] Specifically, during the memory operation, if the CPU or the memory processor detects that the number of memory CE errors reaches the interrupt threshold 1, a CMCI interrupt is triggered and reported to the RAS subsystem, and the RAS subsystem reports it to the corresponding application related to CMCI interrupt management, such as application 1, for processing. Further, if the RAS subsystem detects that the number of CMCI interrupts generated within 1 second reaches 15 times, that is, greater than or equal to 15 times, it means that an interrupt storm has occurred and processing related to interrupt storm management needs to be carried out.

[0068] If the interruption storm detection period and the interruption storm threshold are fixed and unchanged, and are only initialized and configured once during the initialization of the RAS subsystem when the system starts. During the entire kernel operation, the fixed interruption storm detection period and the interruption storm threshold are used as the measurement criteria for determining the interruption storm. As mentioned above, since the processes, qualities, and models of the memory included in different types of servers will be different, different servers including different memories have different adaptabilities to memory CE errors and corresponding interruption storms. The occurrence of an interruption storm will cause the processor to continuously respond to these interruption requests, thereby occupying a large amount of processor resources and memory resources and affecting the normal operation of the system. Therefore, if the interruption storm judgment criteria are too strict, the system performance will be affected. However, if the interruption storm judgment criteria are too loose, there will sometimes be problems such as the timely discovery of system problems being affected. Therefore, for different memories, different types of memory CE errors, and other different scenarios, only setting fixed interruption storm detection periods and interruption storm thresholds cannot accurately reflect the differences in memory processes, qualities, etc., and cannot well meet the different requirements for interruption storm judgment criteria in different scenarios, resulting in problems such as affecting the system reliability of the computing device system and thus affecting the system performance of the computing device system.

[0069] Specifically, on the one hand, the fixed interruption storm detection period and the interruption storm threshold cannot effectively adapt to different service loads and operating environments, and cannot effectively handle the actual workload of the server, easily leading to frequent occurrence of interruption storms, or causing the system to be unable to respond to redundant interruptions in high-load situations in a timely manner, thereby affecting the system performance and response speed of the computing device and increasing the system burden. On the other hand, the fixed interruption storm detection period and the interruption storm threshold cannot be reasonably adjusted according to specific scenarios and requirements, which may lead to overly frequent interruption storms or interruption storms not being discovered in time, thereby increasing the operation and maintenance costs. On the third hand, the single interruption storm detection period and the interruption storm threshold cannot be integrated with the AI-based operation and maintenance management system, and cannot perform operation and maintenance management based on AI decisions, resulting in the system being unable to effectively use AI decisions for real-time monitoring and automatic adjustment, lacking intelligent response capabilities, and restricting the automation and intelligent level of operation and maintenance. On the fourth hand, in a dynamic and changeable application environment, the fixed interruption storm detection period and the interruption storm threshold make the system unable to quickly adapt to environmental changes and unable to achieve better resource utilization, thereby affecting the overall efficiency of the system.

[0070] Based on this, the server provided by the embodiment of the present application has the ability to dynamically adjust the interruption storm detection period and the interruption storm threshold during the kernel operation.

[0071] Exemplarily, such as Figure 1As shown, the application 2 included in the server may be an application related to the dynamic adjustment of the interrupt storm detection period and the interrupt storm threshold.

[0072] In one implementation manner of this application, the application 2 is, for example, an application that provides user services to users. The server can, for example, provide an operation interface for users to set and modify the interrupt storm detection period and the interrupt storm threshold through the application 2. Through this operation interface, users can set a new interrupt storm detection period and an interrupt storm threshold according to specific application scenarios and requirements.

[0073] Therefore, in this embodiment, users are allowed to flexibly configure the threshold policy for memory CE interrupt storm detection. This flexibility enables users to flexibly configure and dynamically adjust the interrupt storm detection period and the interrupt storm threshold of the memory CE interrupt according to actual usage conditions such as the actual operating environment and actual reliability and other business requirements, which can better adapt to different system requirements and operating environments, find the best balance between system performance and reliability, and ensure system performance.

[0074] Moreover, based on the user's setting operations for the interrupt storm detection period and the interrupt storm threshold, the application 2 generates a configuration instruction including the interrupt storm detection period and the interrupt storm threshold set by the user, and sends it to the RAS subsystem. The RAS subsystem modifies the variable values corresponding to the interrupt storm detection period and the interrupt storm threshold stored in the memory variable according to the configuration instruction to store the new interrupt storm detection period and the interrupt storm threshold.

[0075] In another implementation manner of this application, the application 2 may be a program application with AI decision-making functions. The application 2 includes a trained target model, which is obtained by training the model based on historical information such as the type information of historical memory CE errors corresponding to CMCI interrupts and the corresponding historical interrupt storm detection period and historical interrupt storm threshold. When the kernel is in the running state, the application 2 can obtain target information such as the type information of memory CE errors generated in the server, input the target information into the target model for model analysis and processing, and the target model automatically analyzes and obtains a new interrupt storm detection period and an interrupt storm threshold based on the AI decision-making function.

[0076] In this way, different error types can correspond to different interrupt storm detection periods and interrupt storm thresholds. For example, the type information of memory CE errors may include Memory Controller Errors Memory ScrubbingError, MemoryController Errors Memory read error, Microcode ROM Parity Error, Generic Cache Hierarchy, TLB Errors, etc.

[0077] For example, when 200 errors of the type Memory Controller Errors Memory read error are detected within 1 minute, this means that there must be an application frequently reading memory data with CE errors. At this time, the memory access performance of the service will decrease because the memory controller corrects the memory CE errors and occupies a certain memory bandwidth. If the interrupt storm detection period and the interrupt storm threshold remain small at this time, the application will also be frequently interrupted by CMCI interrupts and will also be affected by the interrupts. If the interrupt storm detection period and the interrupt storm threshold are increased, the frequency of triggering the interrupt storm will decrease significantly, and the impairment of the application performance will be alleviated.

[0078] Furthermore, the above type information may also be other characteristic information of memory CE errors. Of course, the target information and historical information used for predicting the interrupt storm detection period and the interrupt storm threshold may also be other information affecting system performance, such as the system CPU occupancy rate, etc., which can be set as needed.

[0079] Therefore, in this embodiment, the AI decision-making function is allowed to flexibly configure the threshold strategy for detecting memory CE interrupt storms, and the application program 2 can perform dynamic adjustment of the interrupt storm detection period and the interrupt storm threshold based on the AI-based strategy. This flexibility enables the system to adaptively configure and dynamically adjust the interrupt storm detection period and the interrupt storm threshold of the memory CE interrupt storm according to the actual operating environment and business requirements, in cooperation with the AI intelligent operation and maintenance management system, based on real-time data analysis and integrating AI technology. This intelligent decision-making ability enables the system to dynamically adapt to the changing operating environment, helps to more intelligently manage and handle interrupt storms, reduce system failures and downtime, and optimize performance and resource utilization.

[0080] The AI decision-making function can be implemented based on algorithms such as the reinforcement learning (RL) algorithm.

[0081] Moreover, the application 2 generates a configuration instruction including the interruption storm detection period and the interruption storm threshold, i.e., the prediction result, based on the target model, and sends it to the RAS subsystem. The RAS subsystem modifies the variable values corresponding to the interruption storm detection period and the interruption storm threshold stored in the memory variable according to the configuration instruction to store the new interruption storm detection period and the interruption storm threshold.

[0082] In another implementation manner of this application, the application 2 can be a program application with information analysis and processing capabilities. For example, the application 2 can determine the business scenario type of the server, and obtain a new interruption storm detection period and an interruption storm threshold according to the business scenario type and the configuration strategy for the interruption storm detection period and the interruption storm threshold corresponding to the business scenario type.

[0083] Exemplarily, multiple different business scenario types and the corresponding configuration strategies can be preset in advance. The configuration strategies corresponding to different business scenario types can be different. Of course, the configuration strategies corresponding to some business scenario types can also be the same.

[0084] For servers corresponding to different business scenario types, a new interruption storm detection period and an interruption storm threshold can be determined based on this to implement dynamic setting of the interruption storm detection period and the interruption storm threshold. Determine whether an interruption storm occurs based on the new interruption storm detection period and the interruption storm threshold. In the case of an interruption storm, trigger the subsequent protection mechanism, so as to better ensure the reliability and stability of the server system.

[0085] Furthermore, determining the business scenario type of the server includes determining the business scenario type through at least one of the system type information of the server, the system performance information of the server, and the business characteristic information of the server.

[0086] Among them, the system type information can refer to, for example, a key business system with high requirements for system performance configured on the server, a relatively general business system with ordinary requirements for system performance, a business system with low requirements for system performance, etc.

[0087] The system performance information of the server, for example, can be information such as the load and memory usage rate of the server, which reflects the current or recent performance of the system.

[0088] The business characteristic information of the server, for example, can be, corresponding to the above system type information, such as a business with high requirements for system performance, a business with ordinary requirements for system performance, a business with low requirements for system performance, etc., or can also refer to information such as the workload size of the business processed by the server currently or recently.

[0089] Of course, the type of business scenario can also be determined based on information other than the above information.

[0090] Furthermore, the configuration policy can include a pre-configured fixed interrupt storm detection period and interrupt storm threshold. The configuration policy can also include dynamically determining a new interrupt storm detection period and interrupt storm threshold based on at least one of the characteristic information of CMCI, the characteristic information of memory CE errors, the system performance information of the server, and the business characteristic information of the server.

[0091] Among them, the characteristic information of CMCI can include, for example, the number of CMCI interrupts per unit time, the interrupt growth rate of CMCI interrupts, and other information reflecting the characteristics of CMCI interrupts.

[0092] The characteristic information of memory CE errors can include, for example, the historical trend of memory CE errors and other information reflecting the characteristics of memory CE errors.

[0093] The system performance information of the server can include, for example, some key performance indicators of the system such as CPU utilization, memory utilization, and I / O throughput. Of course, the system performance information of the server can also include other information reflecting the system performance of the server.

[0094] The business characteristic information of the server can include, for example, business critical parameters such as business type. The business type includes, for example, key businesses such as financial transactions and real-time data processing.

[0095] Of course, the configuration policy can also include dynamically determining a new interrupt storm detection period and interrupt storm threshold based on information other than the above information.

[0096] Furthermore, according to the different system types of the server or the different businesses involved, the type of business scenario can include, for example, high-sensitivity scenarios, balanced sensitivity and performance scenarios, low-sensitivity scenarios, dynamic adjustment scenarios, grading scenarios, etc.

[0097] The interrupt storm management processes corresponding to high-sensitivity scenarios, balanced sensitivity and performance scenarios, low-sensitivity scenarios, dynamic adjustment scenarios, and grading scenarios are further described below.

[0098] Regarding high-sensitivity scenarios, the description is as follows:

[0099] Exemplarily, in some key business systems involving business processes such as financial transactions and real-time data processing, even a short-term system performance degradation may cause serious impacts. Therefore, such systems need to have a high sensitivity to interrupt storms and be able to detect and take measures quickly. For a server configured with such a system, the type of its business scenario can be considered a high-sensitivity scenario.

[0100] For the policy configuration of the interruption storm detection period and interruption storm threshold corresponding to the high-sensitivity scenario, for example, it can be as shown in Table 1 below. The interruption storm detection period (n seconds) is 1 second, and the interruption storm threshold (m times) is 5 times.

[0101] Table 1

[0102] Interrupt storm detection period (n seconds) Interrupt storm threshold (m times) 1 5

[0103] Based on this, in this scenario, the system is set to determine that an interruption storm has occurred when 5 CMCI interruptions are detected within 1 second. Such a setting can quickly respond to the outbreak of memory errors and promptly trigger protection mechanisms, such as load migration, restart migration, or isolating the faulty memory area, to minimize the impact on the business.

[0104] Regarding the balanced sensitivity and performance scenario, the description is as follows:

[0105] Exemplarily, in most general server system environments, it is necessary to find a balance between sensitivity and system performance. The system should not trigger memory isolation due to a small number of memory CE errors, otherwise it will more easily increase the overhead of system memory resources. Therefore, for a server configured with such a general server system, its business scenario type can be considered as a balanced sensitivity and performance scenario.

[0106] For the policy configuration of the interruption storm detection period and interruption storm threshold corresponding to the balanced sensitivity and performance scenario, for example, it can be as shown in Table 2. The interruption storm detection period (n seconds) is 1 second, and the interruption storm threshold (m times) is 15 times.

[0107] Table 2

[0108] Interrupt storm detection period (n seconds) Interrupt storm threshold (m times) 1 15

[0109] Based on this, in this configuration, the system is set to consider an interruption storm when 15 CMCI interruptions are detected within 1 second. Such a setting not only ensures timely response to memory errors but also does not overreact to occasional errors, thus maintaining the availability of system resources.

[0110] Regarding the low-sensitivity scenario, the description is as follows:

[0111] Exemplarily, in some systems with high fault tolerance, such as big data processing platforms, they may be more inclined to accept a certain degree of memory errors to ensure that more memory resources can be applied for. Therefore, for a server configured with such a system, its business scenario type can be considered as a low-sensitivity scenario.

[0112] The policy configuration for the interruption storm detection period and interruption storm threshold corresponding to the low-sensitivity scenario can be, for example, as shown in Table 3. The interruption storm detection period (n seconds) is 1 second, and the interruption storm threshold (m times) is 45 times.

[0113] Table 3

[0114] Interrupt storm detection period (n seconds) Interrupt storm threshold (m times) 1 45

[0115] Based on this, in this scenario, the system is set to determine an interruption storm only when 45 CMCI interruptions are detected within 1 second. Such a high threshold setting allows the system to continue running while tolerating more memory errors and triggers protective measures only when the errors reach a severe level, thus maximizing the utilization of system resources.

[0116] Regarding the dynamic adjustment scenario, the description is as follows:

[0117] Exemplarily, for some system environments with dynamic changes such as cloud computing platforms, the system load and memory usage may change frequently. Therefore, a method that can dynamically adjust the interruption storm detection policy according to the current system state is required. For a server configured with such a system, its business scenario type can be considered as a dynamic adjustment scenario.

[0118] The policy configuration for the interruption storm detection period and interruption storm threshold corresponding to the dynamic adjustment scenario can be, for example, as shown in Table 4. The interruption storm detection period (n seconds) is a dynamic value, and the interruption storm threshold (m times) is a dynamic value. And, exemplarily, the dynamic value of the interruption storm detection period can be adjusted according to the system load. For example, it is set to 5 seconds under low load and 10 seconds under high load, etc. The dynamic value of the interruption storm threshold can be adjusted according to the memory usage rate. For example, it is set to 30 times when the memory usage rate is high and 20 times when the usage rate is low. Of course, the dynamic values of the interruption storm detection period and the interruption storm threshold can also be determined according to other information and can be set as needed.

[0119] Table 4

[0120] Interrupt storm detection period (n seconds) Interrupt storm threshold (m times) Dynamic Dynamic

[0121] Based on this, by dynamically adjusting the interruption storm detection period and the interruption storm threshold, the system can better adapt to different operating states. When the load is high or the memory usage rate is high, the threshold is appropriately increased to increase the tolerance; when the load is low or the memory usage rate is low, the threshold is decreased to improve the detection sensitivity.

[0122] Regarding the grading scenario, the description is as follows:

[0123] Exemplarily, in some complex systems, different countermeasures may need to be taken according to the severity of the interrupt storm. Therefore, multiple levels of interrupt storm levels and corresponding interrupt storm detection strategies can be set. For a server configured with such a system, its business scenario type can be considered a level division scenario.

[0124] The policy configuration for the interrupt storm detection period and interrupt storm threshold corresponding to the level division scenario, exemplarily, for example, can be as shown in Table 5 below, where the interrupt storm includes three levels, such as Level 1, Level 2, and Level 3. Among them, the interrupt storm detection period (n seconds) corresponding to Level 1 is 1 second, and the interrupt storm threshold (m times) is 10 times. The interrupt storm detection period (n seconds) corresponding to Level 2 is 5 seconds, and the interrupt storm threshold (m times) is 30 times. The interrupt storm detection period (n seconds) corresponding to Level 3 is 10 seconds, and the interrupt storm threshold (m times) is 50 times.

[0125] Table 5

[0126] Level Interrupt storm detection period (n seconds) Interrupt storm threshold (m times) Level 1 1 10 Level 2 5 30 Level 3 10 50

[0127] Based on this, according to the severity of the interrupt storm, the system can be divided into three levels, for example. Level 1 corresponds to the most severe interrupt storm and triggers the most urgent response measures; Levels 2 and 3 correspond to moderate and less severe interrupt storms, respectively, and corresponding handling measures are taken. This grading strategy enables the system to manage memory errors more precisely and effectively balance system stability and resource utilization.

[0128] Through the above configuration examples of different scenarios, it can be seen that the flexibility and configurability of the interrupt storm detection strategy enable the system to adjust the detection sensitivity and response method for the interrupt storm according to specific application requirements and environmental conditions, thereby better ensuring the reliability and stability of the system.

[0129] Furthermore, regarding the division of interrupt storm levels, the following is an explanation:

[0130] In the level division scenario, determining the interrupt storm level can depend on a combination of multiple parameters, which can, for example, reflect the occurrence frequency of memory errors, the degree of impact on system performance, and the criticality of the business. Exemplarily, these parameters can include, for example, the number of CMCI interrupts per unit time, the growth rate of CMCI interrupts, system performance metric parameters, business criticality parameters, the historical trend of memory CE errors, etc. Of course, they can also be other parameters, which can be set as needed.

[0131] The following further explains some strategies and parameters for determining the interrupt storm level.

[0132] First, the selection of parameters and their meanings are described. The following are the key parameters for determining the interruption storm level and their meanings.

[0133] Regarding the number of CMCI interruptions (INT_COUNT) per unit time, the number of CMCI interruptions per unit time refers to the number of CMCI interruptions detected by the system within the set interruption storm detection period (such as 1 second, 5 seconds). It can directly reflect the occurrence frequency of memory errors and is the core indicator for determining the interruption storm level. The interruption storm level is determined based on the number of CMCI interruptions per unit time. For example, if there are 10 interruptions in 1 second, the interruption storm level is considered level 1; if there are 30 interruptions in 5 seconds, the interruption storm level is considered level 2; if there are 50 interruptions in 10 seconds, the interruption storm level is considered level 3.

[0134] Regarding the interruption growth rate (INT_RATE), the interruption growth rate refers to the growth speed of the number of CMCI interruptions per unit time. It can reflect the severity of the storm. The higher the growth rate, the greater the threat of the interruption storm to the system. The interruption growth rate is calculated according to the formula: INT_RATE = (current interruption count - previous interruption count) / time interval. The interruption storm level is determined based on the interruption growth rate. For example, the greater the interruption growth rate, the higher the interruption storm level. Moreover, when the interruption count increases from 10 times per second to 20 times per second (growth rate 100%), a higher level may be triggered.

[0135] Regarding the system performance indicators (SYS_PERF), the system performance indicators refer to some key performance indicators of the system, such as CPU utilization, memory utilization, I / O throughput, etc. It can reflect the impact degree of the interruption storm on the overall system performance and can indirectly judge the severity of the interruption storm. The interruption storm level is determined based on the system performance indicators. For example, if the CPU utilization exceeds 80% and the number of interruptions is high, it can be determined as a high-level storm. When the memory utilization is close to saturation and the number of interruptions increases, it can also be judged as a high-level storm. Conversely, it is judged as a low-level storm.

[0136] Regarding the business criticality (BUSINESS_PRIORITY), the business criticality refers to the classification of levels according to the importance of the business, such as key businesses like financial transactions and real-time data processing. Different businesses have different tolerances for interruption storms, and key businesses require more stringent detection and response. The interruption storm level is determined based on the business criticality. For example, for a financial transaction system, the interruption storm threshold is relatively lower (such as 5 interruptions in 1 second trigger level 1), and for a big data processing system, the interruption storm threshold is higher (such as 50 interruptions in 10 seconds trigger level 3).

[0137] For the HISTORY_TREND of memory CE errors, the historical trend refers to judging the trend of memory CE errors based on historical data, such as whether there is a periodic interruption storm or a suddenly increasing storm. Predict the severity of the interruption storm through historical data to assist in level determination. Determine the interruption storm level according to the historical trend of memory CE errors. For example, if the historical data shows that the number of interruptions is gradually increasing, the current storm may be judged as a higher level.

[0138] Based on the above parameters, the determination logic for the interruption storm level can be as follows.

[0139] Exemplarily, according to the above parameters, the interruption storm level can be divided into multiple levels, Figures 3 - 5 several levels. First, define the levels, including the following 4 levels, and the level judgment conditions corresponding to each level can be as shown in Table 6.

[0140] Level 1 (Emergency): The number of interruptions is extremely high, the system performance drops severely, and the business criticality is high.

[0141] Level 2 (Severe): The number of interruptions is relatively high, the system performance drops to a certain extent, and the business criticality is medium.

[0142] Level 3 (Medium): The number of interruptions is relatively large, the system performance is slightly affected, and the business criticality is low.

[0143] Level 4 (Minor): The number of interruptions is small, and the system performance is not significantly affected.

[0144] Table 6

[0145]

[0146]

[0147] Based on the above parameters and levels, the process of dynamically determining the interruption storm level by combining real-time data with a rule engine to determine a new interruption storm detection period and interruption storm threshold can be described as follows.

[0148] To achieve flexible determination of the interruption storm level, the above rules can be integrated into a real-time rule engine to dynamically adjust the level in combination with system real-time data. The specific steps are as follows:

[0149] Step 1, data collection: For example, collect the number of interruptions, interruption growth rate, system performance indicators, business criticality, and historical trend data of the aforementioned CMCI interruptions in real time.

[0150] Step 2, rule matching: According to the preset rule table, match the real-time data with the level determination conditions.

[0151] Step 3, Dynamic Adjustment: Dynamically adjust the interruption storm level according to the matching result. Further, each interruption storm level corresponds to a corresponding interruption storm detection period and an interruption storm threshold, so that corresponding new interruption storm detection periods and interruption storm thresholds can be obtained. Further, corresponding protection measures (such as reducing the load, isolating the faulty memory, restarting the service, etc.) can also be triggered according to the determined interruption storm level.

[0152] Step 4, Logging: Record the determination results and processing measures of each interruption storm level for subsequent analysis and optimization. This analysis includes the analysis of system performance, etc., and this optimization includes the optimization processing of configuration policies, etc.

[0153] The following describes the methods for determining the interruption storm level in some example scenarios.

[0154] As shown in Table 7, the interruption storm level can be determined according to the real-time data obtained in different scenarios.

[0155] Table 7

[0156]

[0157] Further, new interruption storm detection periods and interruption storm thresholds can be determined according to the interruption storm level.

[0158] In addition, regarding triggering corresponding protection measures according to the determined interruption storm level, it includes: for the triggered level 1 (urgent), the system needs to immediately take emergency measures. For the triggered level 3 (medium), the system takes medium protection measures. For the triggered level 4 (minor), the system does not need to intervene immediately.

[0159] Further, the configuration policies corresponding to the above business scenario types can also be other policies other than the above configuration policies, and the business scenario types can also include other scenarios other than the above scenarios, all of which can be set as needed.

[0160] After Application 2 obtains the new interruption storm detection period and interruption storm threshold based on the above business scenario type, it generates a configuration instruction including the interruption storm detection period and interruption storm threshold, and sends it to the RAS subsystem. The RAS subsystem modifies the variable values corresponding to the interruption storm detection period and interruption storm threshold stored in the memory variable according to the configuration instruction to store the new interruption storm detection period and interruption storm threshold.

[0161] In this way, based on the business scenario type of the server, a new interruption storm detection period and an interruption storm threshold can be obtained to achieve dynamic setting of the interruption storm detection period and the interruption storm threshold. Determine whether an interruption storm occurs based on the new interruption storm detection period and the interruption storm threshold. In the case of an interruption storm, trigger subsequent protection mechanisms, thereby better ensuring the reliability and stability of the server system.

[0162] In the implementation manner of this application, after the application 2 obtains the new interruption storm detection period and the interruption storm threshold, it is considered to have obtained the first target value corresponding to the interruption storm detection period and the second target value corresponding to the interruption storm threshold. By cooperating with the kernel RAS subsystem, the variable values corresponding to the foregoing memory variables can be adjusted to obtain the new interruption storm detection period and the interruption storm threshold for the detection and determination of the interruption storm.

[0163] Of course, the application 2 can also be an application that implements the setting of the new interruption storm detection period and the interruption storm threshold through other means.

[0164] Furthermore, the server can also set the new interruption storm detection period and the interruption storm threshold in other ways other than the above three ways of setting the new interruption storm detection period and the interruption storm threshold by the user, automatically analyzing the new interruption storm detection period and the interruption storm threshold by the AI decision-making function, and obtaining the new interruption storm detection period and the interruption storm threshold according to the business scenario type of the server. It can be set as needed during the kernel startup process.

[0165] In summary, the server has the function of initializing the interruption storm detection period and the interruption storm threshold during the kernel startup process, and the function of dynamically adjusting the interruption storm detection period and the interruption storm threshold during the kernel operation process.

[0166] Furthermore, the function of the server for initializing and dynamically adjusting the interruption storm detection period and the interruption storm threshold can be implemented by the kernel RAS subsystem. Therefore, the kernel RAS subsystem in the server has the first function and the second function.

[0167] Such as Figure 2As shown, the second function includes that during the kernel startup process, the kernel RAS subsystem initializes the memory variables according to the configuration file, obtaining the variable values corresponding to the interrupt storm detection period and the variable values corresponding to the interrupt storm threshold, so as to obtain the initialized interrupt storm detection period and interrupt storm threshold for determining the interrupt storm. The initialized interrupt storm detection period is, for example, 1 second, and the initialized interrupt storm threshold is, for example, 15. Then, if 15 CMCI interrupts are generated within 1 second, it is considered that an interrupt storm has occurred. Of course, the initialized interrupt storm detection period can also be any value greater than 0 other than 1 second, and the initialized interrupt storm threshold can also be any positive integer greater than 0 other than 15, and the two can be set according to needs. Moreover, the kernel RAS subsystem creates an interface file for modifying the memory variables to modify the variable values corresponding to the interrupt storm detection period and the variable values corresponding to the interrupt storm threshold of the memory variables, realizing the configuration ability for the interrupt storm detection period and the interrupt storm threshold.

[0168] The first function includes: when the kernel is in the running state, the kernel RAS subsystem, in response to receiving the configuration instructions for the interrupt storm detection period and the interrupt storm threshold, modifies the variable value corresponding to the interrupt storm detection period of the memory variable to the first target value, and modifies the variable value corresponding to the interrupt storm threshold of the memory variable to the second target value, obtaining the new interrupt storm detection period and interrupt storm threshold for determining the interrupt storm, where the configuration instructions carry the first target value and the second target value. Among them, the new interrupt storm detection period is, for example, N seconds, and the new interrupt storm threshold is, for example, M. Then, if M CMCI interrupts are generated within N seconds, it is considered that an interrupt storm has occurred. N can be any number greater than 0, such as N being 1, 2, 4.5, 10.6, etc., and M can be any positive integer greater than 0, such as M being 30, 46, etc., and the values of N and M are determined according to the specific scenario. Moreover, the kernel RAS subsystem can also implement read-write synchronization protection for the memory variables according to the marking information of the memory variables to avoid multiple programs parallelly performing modification operations on the memory variables.

[0169] In this way, the server can dynamically adjust the interrupt storm detection period and the interrupt storm threshold based on different application scenarios or system running states, etc., when the kernel is in the running state, preventing the problem of frequent occurrence or missed reporting of interrupt storms, improving the device stability and reliability, while reducing the occupation of related resources such as the processor resources and memory space in the server by the interrupt, ensuring the system performance of the server, reducing the operation and maintenance cost, and making the server have stronger usability and practicality.

[0170] Next, the process of setting the interrupt storm detection period and the interrupt storm threshold of the server provided by the embodiments of the present application, as well as the process of performing interrupt storm detection, will be described in conjunction with the accompanying drawings.

[0171] Exemplarily, such as Figure 3 As shown, in an implementation manner of the present application, the process of setting the interruption storm detection period and the interruption storm threshold of the server, as well as the process of performing interruption storm detection, includes the following steps:

[0172] S110, during the kernel startup process, according to the configuration file, initialize the variable values of the memory variables corresponding to the interruption storm detection period and the variable values corresponding to the interruption storm threshold, and create an interface file for modifying the memory variables.

[0173] The server stores a preset configuration file, and the configuration file may include default values for initializing the interruption storm detection period and the interruption storm threshold. Exemplarily, the default value of the interruption storm detection period is, for example, 1 second, and the default value of the interruption storm threshold is, for example, 15.

[0174] Furthermore, the configuration file may also include a configuration program related to the creation of the interface. The RAS subsystem creates an interface for modifying the memory variables by reading the configuration program in the configuration file. This interface may exist in the form of a file. Therefore, the creation of the interface is the creation of the interface file. Of course, the interface may also exist in other forms, which can be set according to needs. The RAS subsystem can communicate with, for example, the aforementioned application 2 based on the interface file, and implement read and write operations on the memory variables, so as to modify the interruption storm detection period and the interruption storm threshold stored in the memory variables.

[0175] Exemplarily, this configuration file may be a static configuration file, which is a file written in advance and stored on the disk of this server.

[0176] Exemplarily, such as Figure 4 As shown, during the kernel startup initialization process, the RAS subsystem first parses the configuration parameters in the configuration file. The configuration parameters correspond to the default interruption storm detection period and the interruption storm threshold pre-configured by the user, and obtain the default interruption storm detection period and the interruption storm threshold.

[0177] Parsing the configuration parameters means parsing the parameter names and parameter values in the configuration parameters to obtain the parameter values. For example, the configuration parameters include <period> =<1>, <threshold>If it is equal to <15>, the parameter names period (i.e., the interrupt storm detection period) and threshold (i.e., the interrupt storm threshold) are parsed, and the corresponding parameter values are 1 and 15 respectively.

[0178] Then, verify the parsed interrupt storm detection period and interrupt storm threshold. By judging whether the parsed interrupt storm detection period and interrupt storm threshold are within the range of valid bits, etc., to determine whether the parsed interrupt storm detection period and interrupt storm threshold are legal, and thus determine whether the verification passes.

[0179] For example, if the number of valid bits corresponding to the interrupt storm detection period and the interrupt storm threshold in binary representation is 15 bits, then if it is verified that the valid bits of the interrupt storm detection period and the interrupt storm threshold in binary representation are within these 15 valid bits, the verification passes. Otherwise, the verification fails.

[0180] If the verification passes, modify the corresponding memory variables according to the initialized interrupt storm detection period and interrupt storm threshold, that is, the variable values corresponding to the interrupt storm detection period and the variable values corresponding to the interrupt storm threshold, so as to save the initialized interrupt storm detection period and interrupt storm threshold to the corresponding memory variables. In this way, the parameter values can be saved for subsequent use.

[0181] If the verification fails, an error message can be generated for recording, or reported to the corresponding application program for relevant processing.

[0182] Furthermore, the RAS subsystem creates, for example, an interface file for modifying memory variables by reading the configuration program in the configuration file.

[0183] Based on the interface file created from the configuration file, on the one hand, it defines the ways and rules for the RAS subsystem to access memory variables, ensuring that the RAS subsystem can correctly read and write the variable values of memory variables. For example, the interface file defines the addresses, names, and functions of the memory variables to be accessed, the access permissions for memory variables (such as read operations or write operations, etc.), to ensure the security and stability of the system; and defines the specific ways for the RAS subsystem to access memory variables, such as accessing through memory mapping, I / O ports, etc., to ensure the efficiency and accuracy of access. Of course, other content related to the interaction between the RAS subsystem and memory variables can also be defined.

[0184] On the other hand, the interface file created based on the configuration file also defines the ways and rules for interacting with, for example, the aforementioned application 2, etc., to ensure that the RAS subsystem can correctly interact with application 2. For example, the interface file defines the address, name, and functions of application 2 that need to be accessed, the access permissions to application 2, etc., to ensure the security and stability of the system; and defines the content related to the communication channel between the RAS subsystem and application 2, as well as the specific ways for the RAS subsystem to access application 2. Of course, other content related to the interaction between the RAS subsystem and application 2 can also be defined.

[0185] In this way, in the case where application 2 is the aforementioned application that provides user services, a user interface and channels can be provided, allowing users to adjust the two detection thresholds, namely, the interrupt storm detection period and the interrupt storm threshold, of the memory CE interrupt storm according to specific application scenarios, so as to realize the adjustment of the interrupt storm detection period and the interrupt storm threshold by users to cope with different workloads. Moreover, users can read and set the interrupt storm detection period and the interrupt storm threshold through the interface file, which provides greater flexibility and control to users based on the interface file, enhances the user's control over the interrupt storm detection period and the interrupt storm threshold, and can meet the personalized needs of different users.

[0186] S120. When the kernel is in the running state, determine whether an interrupt storm has occurred according to the occurrence time and occurrence count of CMCI interrupts, as well as the initialized interrupt storm detection period and interrupt storm threshold stored in the memory variable.

[0187] Among them, if it is determined according to the occurrence time and occurrence count of CMCI interrupts that the occurrence count of CMCI interrupts within the interrupt storm detection period is greater than or equal to the interrupt storm threshold, it is determined that an interrupt storm has occurred; if it is determined according to the occurrence time and occurrence count of CMCI interrupts that the occurrence count of CMCI interrupts within the interrupt storm detection period is less than the interrupt storm threshold, it is determined that no interrupt storm has occurred.

[0188] Exemplarily, as Figure 5 shown, if the RAS subsystem detects that the number of CMCI interrupts received within 1 second is greater than or equal to 15 times, it indicates that an interrupt storm has occurred, and it is reported to the corresponding application such as the aforementioned application 2 for processing.

[0189] S130. When the kernel is in the running state, in response to receiving a configuration instruction for the interrupt storm detection period and the interrupt storm threshold, modify the variable value of the memory variable corresponding to the interrupt storm detection period to the first target value, and modify the variable value of the memory variable corresponding to the interrupt storm threshold to the second target value. The configuration instruction carries the first target value and the second target value.

[0190] The configuration instruction can be sent by the aforementioned application 2 to the RAS subsystem. It can be a new policy setting instruction corresponding to the aforementioned user setting function, or a new policy setting instruction corresponding to the aforementioned AI intelligent operation and maintenance function, or a new policy setting instruction corresponding to the aforementioned function of setting according to the business scenario type of the server, etc. As Figure 4 shown, when the kernel is in the running state, application 2 sends a configuration instruction through the communication interface and communication channel with the RAS subsystem. The configuration instruction includes the parameter values corresponding to the interrupt storm detection period and the interrupt storm threshold.

[0191] After the RAS subsystem obtains the configuration instruction sent by application 2 through the application layer, it first parses the configuration instruction, reads the new parameter information in the configuration instruction, and parses the parameter name and parameter value in the new parameter information to obtain the target value. For example, the new parameter information is <period> =<2>, <threshold>If it is equal to <30>, the parameter names period (i.e., the interruption storm detection period) and threshold (i.e., the interruption storm threshold) are parsed out, and the corresponding parameter values are 2 (i.e., the first target value) and 30 (i.e., the second target value) respectively.

[0192] The server includes multiple memories. Multiple memories can share the interruption storm detection period and the interruption storm threshold as a memory group, and can include multiple memory variables corresponding to different memory groups. Therefore, the configuration instruction can include a parameter information for performing read and write operations on a memory variable; or, based on a parameter information, for performing read and write operations on multiple memory variables. Or, the configuration instruction can also include multiple parameter information for performing read and write operations on different memory variables respectively. Thus, the flexible and timely adjustment of the interruption storm detection period and the interruption storm threshold in different memory variables under different service scenarios can be realized.

[0193] Then, the RAS subsystem verifies the target values of the obtained new parameters. By determining whether the parsed target values are within the range of valid bits and other methods, it is determined whether the target values are legal to determine whether the verification passes. For example, the valid bits corresponding to the interruption storm detection period in the binary representation are 15 bits, and it is verified whether the first target value is within the 15 bits of the valid bits in the binary representation. If so, the verification passes. Further, the valid bits corresponding to the interruption storm threshold in the binary representation are 10 bits, and it is verified whether the second target value is within the 10 bits of the valid bits in the binary representation. If so, the verification passes. If both the first target value and the second target value pass the verification, it is necessary to set the memory variable based on the first target value, adjust the interruption storm detection period in the memory variable to the first target value, and it is necessary to set the memory variable based on the second target value, adjust the interruption storm threshold in the memory variable to the second target value. If one of the verifications fails, no modification operation is performed on the memory variable, and the default interruption storm detection period and the interruption storm threshold remain unchanged. And, the foregoing error message can be generated.

[0194] Furthermore, in the case where both the first target value and the second target value pass the verification, first determine the marking information of the memory variable. The marking information is used to indicate whether the modification operation is allowed to be performed on the memory variable. The marking information can be, for example, a Flag mark. The marking information can be stored through the corresponding memory or hardware register and is associated with the memory variable. Before performing a write operation on the memory variable, first obtain the marking information to determine whether the modification operation is allowed to be performed on the memory variable to avoid concurrent modification of the memory variable, resulting in system instability.

[0195] Exemplarily, if the Flag is marked as ON, it indicates that there is currently another program modifying the memory variable, and parallel modification operations on the memory variable are not allowed. After waiting for a period of time according to the preset period, check again whether the Flag is OFF. The value of the preset period can be set as needed, for example, it can be any value such as 1 second, 2 seconds, 6.5 seconds, etc. If the Flag is marked as OFF, it means that there is no other program modifying the memory variable currently, and modification operations on the memory variable are allowed. The program for modifying the memory variable can be, for example, a program that counts CMCI interrupts when a CMCI interrupt occurs.

[0196] If the Flag is marked as OFF, first modify the Flag to ON to indicate that the memory variable is currently being modified and parallel modification operations on the memory variable are not allowed, so as to prevent other programs from modifying the memory variable in parallel. Then, set the memory variable based on the first target value and the second target value, modify the variable value corresponding to the interrupt storm detection period in the memory variable to the first target value (i.e., time N), and modify the variable value corresponding to the interrupt storm threshold in the memory variable to the second target value (i.e., count M). In this way, new parameter values can be set and saved for subsequent use.

[0197] After the modification is completed, modify the Flag back to OFF to indicate that there is no program currently modifying the memory variable, and allow other programs to modify the memory variable.

[0198] By setting and modifying the Flag, a synchronization protection mechanism for the interrupt storm detection period, interrupt storm threshold, and other configuration information stored in the memory variable can be provided, preventing concurrent modification of the configuration information, ensuring the successful and correct modification of the interrupt storm detection period and interrupt storm threshold, and avoiding problems such as information chaos that may be caused by concurrent modification of the configuration information, ensuring the security and consistency of system operation.

[0199] S140, determine whether an interrupt storm has occurred based on the occurrence time and occurrence count of the CMCI interrupt, as well as the new interrupt storm detection period and interrupt storm threshold stored in the memory variable.

[0200] Exemplarily, as Figure 4 shown, when the number of memory CE errors detected by the CPU is greater than the aforementioned interrupt threshold, a CMCI interrupt is triggered, and the CPU reports the CMCI interrupt to the RAS subsystem through the CMCI interrupt processing entry.

[0201] The RAS subsystem determines the Flag of the memory variable in response to receiving a CMCI interrupt. If the Flag is ON, it indicates that another program is currently modifying the memory variable, and parallel modification operations on the memory variable are not allowed. After waiting for a period of time according to a preset cycle, the Flag is rechecked to see if it is OFF. If the Flag is OFF, it means that no other program is modifying the memory variable, and modification operations on the memory variable are allowed. At this time, the program that modifies the memory variable can be, for example, the program that modifies the interrupt storm detection period and the interrupt storm threshold in the memory variable as described above.

[0202] If the Flag is OFF, the Flag is first modified to ON to indicate that a program is currently modifying the memory variable, so as to prevent other programs from modifying the memory variable in parallel. Then, it is judged whether the currently received CMCI interrupt is the Mth CMCI interrupt within N seconds.

[0203] Exemplarily, as Figure 5 shown, for example, within N seconds, the occurrence times of CMCI interrupts are T1, T2, T3... TM in sequence, and the corresponding CMCI interrupts are CMCI interrupt 1, CMCI interrupt 2, CMCI interrupt 3... CMCI interrupt M in sequence.

[0204] Then, as Figure 4 shown, if the current CMCI interrupt is not the Mth CMCI interrupt within N seconds, the count of CMCI interrupts is incremented by 1, or if this CMCI interrupt is the first CMCI interrupt in a new interrupt storm detection period, the count of CMCI interrupts is set to 1. Then, the Flag is modified to OFF to indicate that no program is modifying the memory variable, allowing other programs to modify the memory variable, and the next CMCI interrupt is continuously detected.

[0205] If the current CMCI interrupt is the Mth CMCI interrupt within the Nth second, it means that an interrupt storm has occurred.

[0206] After the modification is completed, the Flag is modified to OFF to indicate that no program is modifying the memory variable, allowing other programs to modify the memory variable.

[0207] Through the setting and modification of the Flag, a synchronization protection mechanism for the interrupt storm detection period, interrupt storm threshold, and other configuration information stored in memory variables can be provided to prevent concurrent modification of the configuration information, ensuring the successful and correct modification of the interrupt storm detection period and interrupt storm threshold, ensuring the security and consistency of the configuration, and avoiding potential risks such as information disorder that may be caused by concurrent modification of the configuration information, ensuring the security and consistency of the system operation.

[0208] Furthermore, as Figure 5 shown, in another implementation of the present application, the RAS subsystem can also have the function of reporting the interrupt storm event to the aforementioned application 2 to notify the storm event to the application 2. When the application 2 is the aforementioned user service application, the application 2 can present the interrupt storm event to the user, and the user can further adaptively configure the interrupt storm detection period and interrupt storm threshold. Moreover, the user can dynamically adjust the interrupt storm detection period and interrupt storm threshold based on real-time monitoring data and continuously optimize the system configuration through the feedback mechanism to achieve more refined system management.

[0209] When the application 2 is the aforementioned AI decision-making function application, the application 2 can further train the model based on the received interrupt storm event. For example, the interrupt storm event can include the type information of the corresponding memory CE error and the current interrupt storm detection period and interrupt storm threshold. Thus, through continuous attempts and adjustments, the optimal interrupt storm detection period and interrupt storm threshold setting strategy can be learned, and a new target model can be obtained for the predictive analysis of the interrupt storm detection period and interrupt storm threshold to adaptively configure the interrupt storm detection period and interrupt storm threshold.

[0210] For example, if the system's performance metrics (such as latency, throughput) are significantly improved and the growth rate of memory errors slows down significantly after the AI target model adjusts the interrupt storm detection period and interrupt storm threshold, the AI target model will record and strengthen this adjustment strategy. On the contrary, if the system performance does not improve after adjusting the interrupt storm detection period and interrupt storm threshold, or the memory error situation remains serious, the AI decision-making module will further refine the error classification by analyzing the error cause and adopt a more precise interrupt storm detection period and interrupt storm threshold setting; or, perform deeper hardware detection and repair. This adaptive feedback mechanism ensures that the management strategy for memory correctable errors can respond in a timely manner and has long-term optimization space, thus achieving the goals of intelligence and control.

[0211] Accordingly, the type and error frequency of the current memory CE error are input into the AI target model. Based on the AI target model, the interrupt threshold is adjusted according to the current memory CE type and error frequency, the change in system performance is observed, and the performance change is used as feedback to determine the target value of the interrupt threshold. The AI target model generates a configuration instruction based on the determined target value and sends it to the kernel through an interface. The RAS subsystem of the kernel adjusts the interrupt storm detection period and the interrupt storm threshold to this target value by accessing the interface file of the register.

[0212] Based on the interrupt storm management method provided by the embodiment of the present application, during the kernel startup process of the server, the RAS subsystem in the server can initialize the variable values of the memory variables corresponding to the interrupt storm detection period and the variable values corresponding to the interrupt storm threshold according to the default interrupt storm detection period and the default interrupt storm threshold in the configuration file, so as to store the default interrupt storm detection period and the default interrupt storm threshold in the memory variables as the default interrupt storm judgment criteria. After the kernel of the server is in the running state, it is determined whether an interrupt storm has occurred according to the default interrupt storm detection period and the default interrupt storm threshold stored in the memory variables.

[0213] In addition, during the kernel startup process, the RAS subsystem can also create an interface file for modifying the memory variables. When it is necessary to modify the interrupt storm detection period and the interrupt storm threshold, corresponding modifications can be made based on the configuration information in the interface file. In this way, the modification of the interrupt storm detection period and the interrupt storm threshold can be conveniently realized.

[0214] When the kernel is in the running state, the RAS subsystem can support flexible configuration and adaptive adjustment of the interrupt storm detection period and the interrupt storm threshold. For example, the interrupt storm detection period and the interrupt storm threshold in the server can be dynamically adjusted according to user operations, so that the user can flexibly adjust the interrupt storm detection period and the interrupt storm threshold corresponding to the memory CE interrupt according to actual needs, or the interrupt storm detection period and the interrupt storm threshold in the server can be dynamically adjusted according to the analysis result of the AI model to achieve adaptive interrupt storm detection period and interrupt storm threshold management for the interrupt storm detection corresponding to the memory CE error, that is, to realize the adaptive strategy adjustment of the interrupt storm detection period and the interrupt storm threshold, and can also better adapt to the current AI-based intelligent operation and maintenance management. Moreover, the RAS subsystem determines whether an interrupt storm is generated according to the dynamically adjusted interrupt storm detection period and the interrupt storm threshold, and uses the modified new interrupt storm detection period and the interrupt storm threshold as the new interrupt storm judgment criteria to determine whether an interrupt storm has occurred.

[0215] In this way, it is possible to better meet the different requirements for the interruption storm judgment criteria in different scenarios, which helps to manage and handle interruption storms more intelligently, reduce system failures and downtime, thereby improving the system reliability of computing devices, ensuring system performance or preventing system performance from being affected.

[0216] Furthermore, the interruption storm detection period and the interruption storm threshold can be set by the user to achieve dynamic adjustment, which can meet the different requirements of the user for the interruption storm detection period and the interruption storm threshold in different usage scenarios and needs, so as to better meet the usage scenarios, improve system reliability, ensure system performance or prevent system performance from being affected.

[0217] Furthermore, the interruption storm detection period and the interruption storm threshold can be analyzed and obtained by an AI model based on information such as errors of memory CE in the server to achieve dynamic adjustment, which can meet the different requirements of the server for the interruption storm detection period and the interruption storm threshold in different usage scenarios and needs, so as to better meet the usage scenarios, improve system reliability, ensure system performance or prevent system performance from being affected.

[0218] Moreover, the server provided by the implementation mode of this application, based on the compatibility of the AI-based intelligent operation and maintenance management system, improves the system fault response ability, optimizes performance by analyzing and adjusting the reporting strategy in real time through AI. In addition, by providing the flexible configuration ability for the interruption storm detection period and the interruption storm threshold, it supports the intelligent operation and maintenance management strategy, enabling the system to automatically adapt to different environmental changes and business requirements and maximize the utilization of resources.

[0219] Furthermore, the interruption storm detection period and the interruption storm threshold can be obtained from the foregoing service scenario information and configuration strategy, which can meet the different requirements of the server for the interruption storm detection period and the interruption storm threshold in different usage scenarios and needs, so as to better meet the usage scenarios, improve system reliability, ensure system performance or prevent system performance from being affected.

[0220] Furthermore, by adopting a flexible interface file, the user can conveniently read and set the interruption storm detection period and the interruption storm threshold of CMCI interruptions. The interface functions include operations such as obtaining the current interruption storm detection period and the interruption storm threshold, and configuring new interruption storm detection periods and interruption storm thresholds, which improves the friendliness of the interaction between the user and the system.

[0221] In some other implementations of the present application, the above-mentioned server may also be a mobile phone, a tablet computer, a laptop computer, a palmtop computer, a mobile internet device (MID), a wearable device (such as including: smart watch, smart bracelet, pedometer, etc.), a vehicle-mounted device, a personal digital assistant, a portable media player, a navigation device, a video game device, a set-top box, a virtual reality and / or augmented reality device, an Internet of Things device, an industrial control device, a streaming media client device, an e-book, a reading device, a POS machine, and other electronic devices.

[0222] The embodiments of the present application also provide a computing device, which can be implemented in various forms. For example, the computing device described in the embodiments of the present application may be the aforementioned electronic device.

[0223] Please refer to Figure 6 As shown, the computing device includes a processor 110, a memory 120, and a communication bus 130. The processor 110 is communicatively connected to the memory 120 through the communication bus 130.

[0224] Among them, the memory 120 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing an operating system, instructions for at least one function, and instructions for implementing the interrupt storm management method provided in the above embodiments, etc.; the data storage area can store data involved in the interrupt storm management method provided in the above embodiments, etc.

[0225] The processor 110 may include one or more processing cores. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 120, the processor 110 invokes the data stored in the memory 120, executes various functions of the embodiments of the present application, and processes data. The processor 110 may be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that for different devices, the electronic devices for implementing the functions of the foregoing processor 110 may also be others, and the embodiments of the present application do not make specific limitations.

[0226] Undoubtedly, in addition to the processor 110, the memory 120, and the communication bus 130, the computing device may further include other more components, such as an input unit (such as at least one of a keyboard, a touch panel, and an audio input unit), and an output unit (such as at least one of a display screen and a speaker).

[0227] This embodiment also provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the foregoing interrupt storm management method when executed by an electronic device.

[0228] This embodiment also provides a computer program product including a computer program, which implements the foregoing interrupt storm management method when executed by an electronic device.

[0229] It should be noted that terms such as "first" and "second" are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0230] Note that in the accompanying drawings, some structural or method features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or ordering may not be required. Instead, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Additionally, the inclusion of a structural or method feature in a particular figure does not imply that such a feature is required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.

[0231] Although the present application has been illustrated and described by referring to certain preferred embodiments thereof, those of ordinary skill in the art should understand that the above content is a further detailed description of the present application in combination with specific embodiments, and it cannot be determined that the specific implementation of the present application is limited only to these descriptions. Those skilled in the art can make various changes in form and detail, including making several simple deductions or substitutions, without departing from the spirit and scope of the present application.< / threshold> < / period> < / threshold> < / period>

Claims

1. A method for interrupt storm management of a computing device, characterized in that, The computing device includes a kernel and a memory. Memory variables related to correctable machine check interrupts are configured in the memory. The memory variables are used to store the interrupt storm detection period and the interrupt storm threshold corresponding to the correctable machine check interrupts. The interrupt storm detection period and the interrupt storm threshold are used for determining an interrupt storm. The method includes: When the kernel is in a running state, in response to receiving a configuration instruction for the interrupt storm detection period and the interrupt storm threshold, modify the variable value of the memory variable corresponding to the interrupt storm detection period to a first target value, and modify the variable value of the memory variable corresponding to the interrupt storm threshold to a second target value. The configuration instruction carries the first target value and the second target value.

2. The method according to claim 1, characterized in that, The first target value and the second target value are obtained based on any one of the following methods: Obtained based on a user's setting operation for the interrupt storm detection period and the interrupt storm threshold; Obtained based on the prediction results of a target model for the interrupt storm detection period and the interrupt storm threshold. The target model is trained based on historical information corresponding to the correctable machine check interrupts. The historical information includes feature information of historical memory correctable errors, as well as corresponding historical interrupt storm detection periods and historical interrupt storm thresholds; or, Obtained based on the business scenario type of the computing device.

3. The method according to claim 2, wherein Obtaining the first target value and the second target value based on the prediction results of the target model for the interrupt storm detection period and the interrupt storm threshold includes: Obtain target information, where the target information includes feature information of current memory correctable errors; Input the target information into the target model for model analysis and processing, and obtain the prediction results of the interrupt storm detection period and the interrupt storm threshold as the first target value and the second target value respectively.

4. The method according to claim 2, characterized in that Obtaining the first target value and the second target value based on the business scenario type of the computing device includes: Determine the business scenario type of the computing device; According to the business scenario type and the configuration strategy for the interrupt storm detection period and the interrupt storm threshold corresponding to the business scenario type, obtain the first target value and the second target value.

5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: During the kernel startup process, initialize the variable value of the memory variable corresponding to the interrupt storm detection period and the variable value corresponding to the interrupt storm threshold according to a configuration file, and create an interface file for modifying the memory variable.

6. The method according to any one of claims 1-5, characterized in that, The kernel is also configured with a flag information, which is used to indicate whether to allow modifying the memory variable. The method further includes: When it is determined that the memory variable is allowed to be modified according to the flag information, modify the variable value of the memory variable corresponding to the interrupt storm detection period to a first target value, and modify the variable value of the memory variable corresponding to the interrupt storm threshold to a second target value.

7. According to the method of claim 6, wherein Before modifying the variable value of the memory variable corresponding to the interrupt storm detection period to a first target value and modifying the variable value of the memory variable corresponding to the interrupt storm threshold to a second target value, the method further includes: Modify the flag information so that the flag information indicates that the memory variable is currently being modified and parallel modification operations on the memory variable are not allowed; After modifying the variable value of the memory variable corresponding to the interrupt storm detection period to a first target value and modifying the variable value of the memory variable corresponding to the interrupt storm threshold to a second target value, the method further includes: Modify the flag information so that the flag information indicates that modification operations on the memory variable are allowed.

8. The method according to claim 6 or 7, characterized in that, Modifying the variable value of the memory variable corresponding to the interrupt storm detection period to a first target value and modifying the variable value of the memory variable corresponding to the interrupt storm threshold to a second target value includes: Parse the configuration instruction to obtain the first target value and the second target value; Verify the first target value and the second target value; When the verification passes, determine whether modification operations on the memory variable are allowed according to the flag information. When modification operations on the memory variable are allowed, modify the variable value of the memory variable corresponding to the interrupt storm detection period to the first target value and modify the variable value of the memory variable corresponding to the interrupt storm threshold to the second target value.

9. A computing device, characterized in that, The computing device includes a kernel and a memory. The memory is configured with memory variables related to correctable machine check interrupts. The memory variables are used to store the interrupt storm detection period and the interrupt storm threshold corresponding to the correctable machine check interrupts. The interrupt storm detection period and the interrupt storm threshold are used for determining an interrupt storm. The kernel is configured with a first function. The first function includes: when the kernel is in a running state, in response to receiving a configuration instruction for the interrupt storm detection period and the interrupt storm threshold, modify the variable value of the memory variable corresponding to the interrupt storm detection period to a first target value and modify the variable value of the memory variable corresponding to the interrupt storm threshold to a second target value. The configuration instruction carries the first target value and the second target value.

10. The computing device according to claim 9, wherein The kernel is configured with a second function. The second function includes: during the startup process of the kernel, initialize the variable value of the memory variable corresponding to the interrupt storm detection period and the variable value corresponding to the interrupt storm threshold according to a configuration file, and create an interface file for performing modification operations on the memory variable.

Citation Information

Patent Citations

  • Alarm threshold determination method and device, equipment, storage medium and alarm system

    CN115202802A

  • Memory fault-tolerant method and device capable of correcting error storm and medium

    CN116954986A

  • Interrupt request processing method and device, electronic equipment and readable storage medium

    CN117271095A

  • Method and apparatus for controlling interrupt storms

    US20050182879A1

Cited By

  • Interrupt processing method and device, equipment, storage medium and program

    CN121210062A

  • Interrupt processing method, apparatus, device, storage medium, and program

    CN121210062B