A server-based dynamic real-time multi-memory fault collection method
The dynamic real-time memory fault collection method addresses the limitations of existing server memory fault collection by enabling immediate and comprehensive fault detection through server communication, ECC register monitoring, and system log analysis, enhancing fault diagnosis and repair efficiency.
Patent Information
- Application Number
- CN202411959681.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-12-30
AI Technical Summary
Traditional server memory failure collection methods only focus on memory data in limited dimensions, miss key exception information, and rely on fixed-time interval sampling strategies to inadequate memory failure data.
The dynamic real-time multivariate memory fault collection method is adopted, and by monitoring ECC registers and system logs, combining text processing and unique identifier (ID) systems, the memory status is monitored in real time, and error data is reported in the shortest communication time, setting the second time threshold to end the monitoring.
It realizes instant response to memory failures, improves fault diagnosis and repair efficiency, provides a comprehensive perspective on fault analysis, reduces the number of error reports, and reduces the burden of system maintenance.
Smart Images

Figure CN119782026B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer system monitoring and fault management, and specifically to a method for dynamically and real-time collecting multi-source memory faults based on a server. Background Art
[0002] The technical field of computer system monitoring and fault management mainly focuses on how to monitor the running state of a computer system in real time, discover and manage potential faults to ensure the stability, efficiency, and reliability of the system;
[0003] In today's digital age, as the core hub for various data processing, storage, and network services, servers bear the heavy responsibility of ensuring the stable operation of a large amount of business. With the continuous improvement of server performance, the complexity of the memory system has been increasing day by day, and the requirements for its reliability have become more and more stringent;
[0004] There are many limitations in traditional server memory fault collection methods, which constitute the background of the problems to be solved urgently. On the one hand, past collection methods focused on obtaining memory data with limited dimensions, only paying attention to the basic information at the memory hardware level, and could only capture extremely limited fault manifestations, missing a large amount of key abnormal information hidden in the complex memory operation mechanism. On the other hand, timeliness is also a major shortcoming of traditional methods. Most existing memory fault collections follow a fixed time interval sampling strategy, with a relatively long interval period, ranging from several minutes to several hours. There is an urgent need for a method to solve the problems of insufficient memory fault data and timeliness caused by the single collection of server memory data. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for dynamically and real-time collecting multi-source memory faults based on a server to solve the problems proposed in the prior art.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A method for dynamically and real-time collecting multi-source memory faults based on a server, and the method for dynamically and real-time collecting multi-source memory faults specifically includes the following steps:
[0007] Step S1: After the server is powered on and started, communicate between each pair of servers in the server cluster, and respectively collect the time required for communication;
[0008] Step S2: Among the collected time required for communication, obtain the shortest time required for communication, and use the shortest time required for communication as the first time threshold;
[0009] Step S3: Start monitoring the memory error register in the ECC register;
[0010] Step S4: When a memory error is detected, obtain the error data in the memory error register, and based on text processing, obtain the data related to the memory error in the system log;
[0011] Step S5: Determine the query timestamp through the data related to the memory error in the system log. The query timestamp refers to, during the process of obtaining the data related to the memory error in the system log based on text processing, marking the matching position of the memory error in the last system log obtained each time. This mark is the query timestamp;
[0012] Step S6: Start querying from the query timestamp, gradually obtain the data related to the memory error in the system log, and report it to the system. During the gradual acquisition process, if the data related to the memory error in the system log is not fully obtained within the first time threshold, first report the newly obtained memory error to the system, and then continue to obtain the data related to the memory error in the system log until the data related to the memory error in the system log is fully obtained, and mark the query timestamp at this position;
[0013] Step S7: The system sets a second time threshold. When the preset second time threshold is reached, stop listening; if the preset second time threshold is not reached, repeat Steps S3 to S6.
[0014] In Step S1, after the server is powered on and starts up, complete the kernel initialization and load each system service, including hardware inspection, kernel loading, and system service startup, to make the server in a normal usage state.
[0015] The text processing includes regular expressions, template filling and searching, and keyword matching.
[0016] The error data in the memory error register includes error type, error quantity, and error occurrence address information.
[0017] The data related to the memory error in the system log includes error type, quantity, occurrence address, and occurrence time information.
[0018] When the time from listening for and obtaining the memory error to reporting it to the system is less than the first time threshold, feedback an alarm to the administrator port.
[0019] The second time threshold is the total listening duration set in advance, specifically the time of the longest operable cycle of a server.
[0020] The dynamic real-time multi-memory fault collection method includes the following architecture: communication module, register monitoring module, system log monitoring module, text processing module, and fault reporting module;
[0021] The communication module is responsible for managing and controlling the information exchange between servers, including data transmission and time synchronization, and takes the shortest communication time required by collecting the communication time as the first time threshold;
[0022] The register monitoring module is used to monitor the memory error register in the ECC register and obtain error information if an error is found;
[0023] The system log monitoring module is used to obtain data related to memory errors in the system log and obtain error information if an error is found;
[0024] The text processing module is responsible for processing and integrating the error data obtained from the register monitoring module and the system log monitoring module, including deduplication, sorting, and formatting operations;
[0025] The fault reporting module is responsible for reporting the processed error data to the system to ensure that the system discovers and processes problems, and decides whether to end the monitoring according to the preset second time threshold.
[0026] The core of step S6 lies in innovatively adopting a unique identifier (ID) system based on the server identification code, memory channel number, memory error occurrence time, and type, enabling accurate identification and merging of similar or duplicate error messages even from different sources. This process not only improves the efficiency of data processing but also greatly reduces the number of error reports, thus reducing the burden on system maintenance personnel;
[0027] The specific implementation steps for querying and gradually obtaining data related to memory errors in the system log and reporting them to the system starting from the query timestamp include:
[0028] Step S6-1: Establish an ID for each memory error. The ID consists of 36 digits, 3 delimiters "-", and 4 decimal points ".", and is separated by the delimiters into a 12-bit IP address, a 3-bit memory channel number, a 16-bit time, a 6-bit microsecond, and a 1-bit memory error type code from left to right in sequence; among them, for the IP address, the 4 groups of numbers separated by decimal points are padded with zeros to 3 digits from the left and then concatenated into a 12-bit number after removing the decimal points; for the time format, zeros are added to the left of the hour, minute, and second that are less than 2 digits to make them 2 digits;
[0029] Step S6-2: Integrate and deduplicate the data related to memory errors in the obtained system log, and merge the memory errors that occur on the same memory channel of the same server within the preset second time threshold into 1 memory error;
[0030] Step S6-3: Report the deduplicated data related to memory errors in the system log to the local system and the management node system.
[0031] ECC (Error-Correcting Code): That is, error checking and correction; the ECC register is a register that stores ECC-related information and is used in a computer system to detect and correct memory data errors.
[0032] A server cluster is a group of interconnected independent servers that work together and are connected through a network;
[0033] The system is a local system and a management node system;
[0034] The "immediately" means to respond within the shortest time interval required for communication within the server cluster.
[0035] The second time threshold is the total duration of the pre-set monitoring, specifically the time of the longest operable cycle of a server. Among them, the second time threshold can be an empirical value, can also be obtained through a certain method, and can be adjusted according to the actual application situation of the server to make it more reasonable. It should be noted that the method of adjusting the second time threshold includes but is not limited to one or more of the following methods: grid search, random search, Bayesian optimization, gradient descent method, genetic algorithm, particle swarm optimization, simulated annealing, differential evolution, model-based optimization, reinforcement learning, rule-based method, expert system, meta-heuristic algorithm, automated machine learning.
[0036] Compared with the prior art, the beneficial effects of the present invention are: The present invention proposes a dynamic real-time multi-source memory fault collection method based on a server, which can dynamically monitor the memory status, allows the system to respond immediately when a problem occurs, rather than relying on regular inspections, which improves the efficiency of fault diagnosis and repair; real-time collects relevant information on memory faults, and the system can quickly respond to faults to improve system stability; diversified data collection provides a more comprehensive perspective for fault analysis, helps to accurately identify the root cause of faults, and thus accelerates the fault troubleshooting process. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a flowchart of the working process of a dynamic real-time multi-source memory fault collection method based on a server of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0039] Embodiment: As Figure 1As shown in the figure, the present invention provides a technical solution, a dynamic real-time multi-memory fault collection method based on a server. The dynamic real-time multi-memory fault collection method specifically includes the following steps:
[0040] Step S1: After the server is powered on and started, communicate between each pair of servers in the server cluster, and respectively collect the time required for communication.
[0041] Step S2: Among the collected communication required times, obtain the shortest communication required time, and use the shortest communication required time as the first time threshold; and find the longest communication time between each pair within the server cluster to control the time for monitoring and reporting memory errors.
[0042] Step S3: Start monitoring the memory error register in the ECC register.
[0043] When the memory DIMM detects a memory error, it will send it to the cache and MCU memory controllers at all levels. After the CPU obtains this information, it stores it in the Aarch64 system registers inside it; the register monitoring module will monitor the Aarch64 system registers related to RAS in the CPU, and obtain relevant detailed information after the memory error occurs.
[0044] Step S4: When it is monitored that a memory error occurs, obtain the error data in the memory error register, and based on text processing, obtain the data related to the memory error in the system log.
[0045] Step S5: Determine the query timestamp through the data related to the memory error in the system log. The query timestamp refers to, during the process of obtaining the data related to the memory error in the system log based on text processing, marking the matching position of the memory error in the last system log obtained each time, and this marking is the query timestamp.
[0046] There are multiple moments at intervals of the first time threshold, , , , …, , …, where is the initial moment. During the time interval from to , lines of information to be processed are generated in the system log of a certain server. If the lines of information to be processed have not been processed yet at the moment of , then first report the processed information to the system and continue to process the remaining information to be processed. If it is still not processed later, continue to report at intervals of the first time threshold at each moment.
[0047] Step S6: Start querying from the query timestamp, gradually obtain the data related to memory errors in the system log, and report it to the system. During the gradual acquisition process, if the data related to memory errors in the system log is not fully obtained within the first time threshold, first report the newly obtained memory errors to the system, and then continue to obtain the data related to memory errors in the system log until all the data related to memory errors in the system log is obtained, and mark the query timestamp at this position;
[0048] Step S7: The system sets a second time threshold. When the preset second time threshold is reached, stop listening; if the preset second time threshold is not reached, repeat steps S3 to S6.
[0049] In step S1, after the server is powered on and started, complete the kernel initialization and load each system service, including hardware inspection, kernel loading, and system service startup, so that the server is in a normal use state.
[0050] The text processing includes regular expressions, template filling and searching, and keyword matching.
[0051] The error data in the memory error register includes error type, error quantity, and error occurrence address information.
[0052] The data related to memory errors in the system log includes error type, quantity, occurrence address, and occurrence time information.
[0053] When the time from listening for memory errors to reporting to the system is less than the first time threshold, feedback an alarm to the administrator port.
[0054] The second time threshold is the total duration of the preset listening, specifically the time of the longest operable cycle of a server.
[0055] The dynamic real-time multi-memory fault collection method includes the following architecture: a communication module, a register listening module, a system log listening module, a text processing module, and a fault reporting module;
[0056] The communication module is responsible for managing and controlling the information exchange between servers, including data transmission and time synchronization, and takes the shortest communication required time as the first time threshold by collecting the time required for communication;
[0057] The register listening module is used to listen to the memory error register in the ECC register and obtain error information when an error is found;
[0058] The system log listening module is used to obtain the data related to memory errors in the system log and obtain error information when an error is found;
[0059] The text processing module is responsible for processing and integrating the error data obtained from the register monitoring module and the system log monitoring module, including deduplication, sorting, and formatting operations;
[0060] The fault reporting module is responsible for reporting the processed error data to the system to ensure that the system discovers and processes problems, and decides whether to end the monitoring according to the preset second time threshold.
[0061] The core of step S6 lies in innovatively adopting a unique identifier (ID) system based on the server identification code, memory channel number, memory error occurrence time, and type, enabling precise identification and merging of similar or duplicate error messages even from different sources. This process not only improves the efficiency of data processing but also greatly reduces the number of error reports, thus alleviating the burden on system maintenance personnel;
[0062] The specific implementation steps of querying the memory error-related data in the system log starting from the query timestamp and reporting it to the system include:
[0063] Step S6-1: Establish an ID for each memory error. The ID consists of 36 digits, 3 delimiters "-", and 4 decimal points ".", and is separated by the delimiters into a 12-bit IP address, a 3-bit memory channel number, a 16-bit time, a 6-bit microsecond, and a 1-bit memory error type code from left to right in sequence; among them, for the IP address, the 4 groups of numbers separated by decimal points are padded to 3 digits from the left and then the decimal points are removed and concatenated into a 12-bit number; for the time format, 0 is added to the left of the hour, minute, and second with less than 2 digits to make it 2 digits;
[0064] Step S6-2: Integrate and deduplicate the memory error-related data obtained from the system log, and merge the memory errors that occur on the same memory channel of the same server within the preset second time threshold into 1 memory error;
[0065] Among them, the memory error type code comparison table is specifically:
[0066]
[0067] At the moment of 1730427497.296820, a server with an IP address of 192.168.1.15 had a correctable memory error on memory channel 8, and the ID of this memory error is 192.168.001.015-008-20241101101817.296820-1.
[0068] Step S6-3: Report the deduplicated memory error-related data in the system log to the local system and the management node system.
[0069] The second time threshold is the total duration of the pre-set monitoring in advance, specifically the time of the longest operable cycle of a server. Among them, the second time threshold can be an empirical value or can be obtained through certain methods, and can be adjusted according to the actual application situation of the server to make it more reasonable. It should be noted that the method of adjusting the second time threshold is carried out by including but not limited to one or more of the following methods: grid search, random search, Bayesian optimization, gradient descent method, genetic algorithm, particle swarm optimization, simulated annealing, differential evolution, model-based optimization, reinforcement learning, rule-based method, expert system, meta-heuristic algorithm, automated machine learning.
[0070] A server cluster is a group of interconnected independent servers that work together, and these servers are connected through a network;
[0071] ECC (Error-Correcting Code): that is, error checking and correction; an ECC register is a register that stores ECC-related information and is used in a computer system to detect and correct memory data errors.
[0072] The system is a local system and a management node system;
[0073] The "immediately" means to respond within the shortest time interval required for communication within the server cluster.
[0074] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.
Claims
1. A server-based dynamic real-time multi-memory fault collection method, characterized in that: The specific steps of the dynamic real-time multi-source memory fault collection method are as follows: Step S1: After the server is powered on and starts up, communicate between servers in the server cluster pairwise, and collect the time required for communication respectively. Step S2: Among the collected time required for communication, obtain the shortest time required for communication, and use the shortest time required for communication as the first time threshold. Step S3: Start monitoring the memory error register in the ECC register. Step S4: When a memory error is monitored, obtain the error data in the memory error register, and based on text processing, obtain the data related to the memory error in the system log. Step S5: Determine the query timestamp through the data related to the memory error in the system log. The query timestamp refers to, during the process of obtaining the data related to the memory error in the system log based on text processing, marking the matching position of the memory error in the last system log obtained each time. The marked position is the query timestamp. Step S6: Start querying from the query timestamp, gradually obtain the data related to the memory error in the system log, and report it to the system. Specifically: Step S6-1: Establish an ID for each memory error. The ID consists of 36 digits, 3 separators, and 4 decimal points, and is separated by the separators from left to right into a 12-bit IP address, a 3-bit memory channel number, a 16-bit time, a 6-bit microsecond, and a 1-bit memory error type code. Among them, for the IP address, first make up the 4 groups of numbers separated by decimal points to 3 digits from the left, and then splice them into 12 digits after removing the decimal points; for the time format, add 0 to 2 digits on the left side of the hour, minute, and second that are less than 2 digits. Step S6-2: Integrate and deduplicate the data related to the memory error in the obtained system log, and merge the memory errors that occur on the same memory channel of the same server within the preset second time threshold into 1 memory error. Step S6-3: Report the deduplicated data related to the memory error in the system log to the local system and the management node system. During the gradual acquisition process, if the data related to the memory error in the system log is not completely acquired within the first time threshold, first report the newly acquired memory error to the system, and then continue to acquire the data related to the memory error in the system log until the data related to the memory error in the system log is completely acquired, and mark the query timestamp at the position of the last newly acquired memory error reported. Step S7: The system sets the second time threshold. When the preset second time threshold is reached, end the monitoring; if the preset second time threshold is not reached, repeat steps S3 to S6.
2. The dynamic real-time multi-memory fault collection method based on a server according to claim 1, characterized in that: In step S1, after the server is powered on and starts up, complete the kernel initialization and load each system service, including hardware inspection, kernel loading, and system service startup, so that the server is in a normal use state.
3. A dynamic real-time multi-memory fault collection method based on a server according to claim 1, characterized in that: The text processing includes regular expressions, template filling and searching, and keyword matching.
4. A method for dynamically and real - time collecting multiple memory faults based on a server according to claim 1, characterized in that: The error data in the memory error register includes error type, error quantity, and error occurrence address information.
5. A method for dynamically and real-time collecting multiple memory faults based on a server according to claim 1, characterized in that: The data related to the memory error in the system log includes error type, quantity, occurrence address, and occurrence time information.
6. A method for dynamically and real-time collecting multiple memory faults based on a server according to claim 1, characterized in that: When the time from monitoring and obtaining a memory error to reporting it to the system is less than the first time threshold, an alarm is fed back to the administrator port.
7. A method for dynamically and real-time collecting multiple memory faults based on a server according to claim 1, characterized in that: The second time threshold is the total duration of the preset monitoring, specifically the time of the longest operable cycle of a server.
8. A method for dynamically and real-time collecting multiple memory faults based on a server according to claim 1, characterized in that: The dynamic real-time multi-memory fault collection method includes the following architecture: a communication module, a register monitoring module, a system log monitoring module, a text processing module, and a fault reporting module; The communication module is responsible for managing and controlling the information exchange between servers, including data transmission and time synchronization, and takes the shortest communication time required as the first time threshold by collecting the time required for communication. The register monitoring module is used to monitor the memory error register in the ECC register and obtain error information when an error is found. The system log monitoring module is used to obtain data related to memory errors in the system log and obtain error information when an error is found. The text processing module is responsible for processing and integrating the error data obtained from the register monitoring module and the system log monitoring module, including deduplication, sorting, and formatting operations. The fault reporting module is responsible for reporting the processed error data to the system to ensure that the system discovers and processes problems, and decides whether to end the monitoring according to the preset second time threshold.
Citation Information
Patent Citations
Operating system memory fault processing method and system
CN118689690A
Development of internet and service, and method for enhancing security
JP2023021877A