Simple Network Management Protocol Extended Memory ECC Error Monitoring Method and System
By extending SNMP, real-time monitoring and analysis of server memory ECC errors is achieved, and the problem of difficult real-time monitoring of memory ECC errors in the existing technology is solved, which improves the inspection efficiency and reduces the number of system failures.
Patent Information
- Application Number
- CN202510179903.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-19
AI Technical Summary
The existing technology is difficult to monitor memory ECC errors in real time during server delivery, resulting in operating system downtime, low troubleshooting efficiency, and difficult to achieve early warning.
Through the extension of Simple Network Management Protocol (SNMP), real-time monitoring of memory ECC errors is realized, including server initialization and startup of SNMP, polling and collecting the master node, conducting online statistics and analysis, and rating each server health through a health grading algorithm, and classification and trend prediction based on the number of single-bit errors and double-bit errors.
Real-time monitoring of memory ECC errors in the computer room, alarm and prevention are made in advance, effectively improving the inspection efficiency and reducing the number of system crashes or crashes.
Smart Images

Figure CN119645718B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a monitoring method and system for simple network management protocol extended memory ECC error reporting, belonging to the field of network communication. Background Art
[0002] At present, when the server is delivered to the customer's site to run business, the memory ECC error will cause the operating system to crash. Specifically, the CE error (Correctable ECC Error) is usually caused by an error in a single bit in the memory module. Although this error can be corrected by the system, if the error occurs in a critical system or application area, or the error occurs frequently, it may cause the system to crash or freeze, and then cause the customer's important data to be lost or the service business to be stopped, posing a huge security risk.
[0003] Although the BMC (Baseboard Management Controller) has the function of recording ECC errors, the error record viewing process needs to be operated separately, and the troubleshooting efficiency is too low. In addition, usually, the administrator only starts the troubleshooting program after the system failure occurs, making it difficult to achieve early warning. Summary of the invention
[0004] The present invention provides a monitoring method and system for simple network management protocol extended memory ECC error reporting, aiming to at least solve one of the technical problems existing in the prior art.
[0005] The technical solution of the present invention relates to a monitoring method for a simple network management protocol extended memory ECC error. The method according to the present invention comprises the following steps:
[0006] S100, server initialization starts SNMP;
[0007] S200, polling and collecting data from the master node;
[0008] S300, conducts online (real-time) statistics and analysis, including:
[0009] S310, receiving and processing data;
[0010] S320, using a health grading algorithm, scoring the health of each server and classifying them based on the number of single-bit errors and double-bit errors;
[0011] S330, perform data analysis and trend forecasting to predict future error trends and provide early warning of potential problems based on historical data and using time series models;
[0012] S400. Visualize the data.
[0013] Further, the step S310 includes:
[0014] S311, defining ECC error data using a dictionary structure to define original data;
[0015] S312. Use pd.DataFrame to convert the original data dictionary into a Pandas data frame to create a data frame. In timestamp formatting, convert the Timestamp field into the Pandas date and time format.
[0016] S313, fillna(0) is used to fill the missing values in the data frame to perform missing value filling;
[0017] S314. Output standardized data.
[0018] Further, in step S311:
[0019] Define Server as the IP address of the server; define SingleBitErrors as the number of single-bit errors; define DoubleBitErrors as the number of double-bit errors; and define Timestamp as the time when the data was collected.
[0020] Further, the step S320 includes:
[0021] S321, setting a classification threshold;
[0022] S322, setting a classification function; wherein, defining a function classify_health(row) to evaluate each row of data and return the corresponding health status according to the threshold condition: Healthy, Mildly Abnormal or Severely Abnormal;
[0023] S323. Use the DataFrame.apply() method to apply the classify_health function to each row of data and store the results in a new column HealthStatus.
[0024] S324. Output the classification result, wherein the output server health status table content items include Server, SingleBitErrors, DoubleBitErrors and the classification result HealthStatus.
[0025] Further, in step S321,
[0026] The threshold settings used to define the health status include:
[0027] When single-bit errors < 20 and double-bit errors < 2, it is considered healthy;
[0028] When the single-bit error is between 20 and 50, and the double-bit error is between 2 and 5, it is judged as a mild abnormal state;
[0029] When single-bit errors > 50 or double-bit errors > 5, it is judged as a serious abnormality.
[0030] Further, the step S100 includes:
[0031] S110. Use MODULE-IDENTITY to define the module identifier (OID), update time, development organization and description information to define the basic information of the module and assign a unique enterprise OID to the module;
[0032] S120, introduce the standard SNMPv2-SMI module, and use its MODULE-IDENTITY and OBJECT-TYPE basic types to import the basic module;
[0033] S130, using OBJECT-TYPE to define specific monitoring objects; wherein, the monitoring object singleBitErrors is defined as being used to record the number of single-bit ECC errors, and the monitoring object oubleBitErrors is defined as being used to record the number of double-bit ECC errors, and the access rights of the objects are set to read-only, and the range of the limit value is set;
[0034] S140. Use the OID hierarchy to define a unique identifier for each object to organize the object hierarchy.
[0035] Further, the step S200 includes:
[0036] S210. Use the high-level interface provided by pysnmp.hlapi to simplify SNMP operations, and introduce the getCmd method to initiate SNMP GET requests and collect data of specific objects to import necessary modules;
[0037] S220, determining a data collection function, including: a function fetch_ecc_data(server_ip) for receiving the IP address of the server as input, and returning the collected ECC data result;
[0038] S230, creating the necessary parameters of the SNMP request to set the SNMP request; including: SnmpEngine() for initializing the SNMP engine; CommunityData() for specifying the SNMP community string and protocol version; UdpTransportTarget() for setting the IP address of the target server and the SNMP default port; ContextData() for providing context information; and ObjectType() for specifying the MIB object to be collected;
[0039] S240, using the iterator returned by getCmd to parse the SNMP response one by one, to iteratively obtain data, and to judge the collection result; wherein, if there is an error, the error information is recorded, otherwise, the returned OID and the corresponding value are extracted;
[0040] S250: Store the collected results in a dictionary format and return them to the caller.
[0041] Further, the step S400 includes:
[0042] S410, using a list containing the number of historical ECC errors as input data and for training a time series model;
[0043] S420, specifying parameters of the ARIMA model, including: parameter p is specified as the order of the autoregressive (AR) term, which is used to indicate how many past values the model uses to predict the current value; parameter d is specified as the number of differences, which is used to convert a non-stationary time series into a stationary series; parameter q is specified as the order of the moving average (MA) term, which is used to indicate how many past prediction errors the model uses;
[0044] S430, using the model.fit() method to perform model training on historical data to generate a fitted time series model;
[0045] S440, calling the forecast(steps=n) method to predict the number of errors in the next n time steps, and output the time series model prediction result to obtain the future ECC error trend.
[0046] The technical solution of the present invention also relates to a computer-readable storage medium on which program instructions are stored, and the above-mentioned method is implemented when the program instructions are executed by a processor.
[0047] The technical solution of the present invention also relates to a monitoring system for simple network management protocol extended memory ECC errors, the system comprising a computer device, and the computer device comprises the above-mentioned computer-readable storage medium.
[0048] The beneficial effects of the present invention are as follows:
[0049] The present invention is extended according to the functions of the current SNMP, and provides an SNMP extended memory ECC error monitoring method and system, which can realize real-time monitoring of memory ECC in a computer room for early warning and prevention, effectively improve the troubleshooting efficiency, and reduce the number of occurrences that lead to system crashes or freezes. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a basic flow chart of the method according to the present invention.
[0051] Figure 2 is a distribution diagram of visualized health status according to the method of the present invention. DETAILED DESCRIPTION
[0052] The concept, specific structure and technical effects of the present invention will be clearly and completely described below in combination with the embodiments and drawings to fully understand the purpose, scheme and effect of the present invention.
[0053] It should be noted that, unless otherwise specified, when a feature is referred to as being "fixed" or "connected" to another feature, it may be directly fixed or connected to another feature, or it may be indirectly fixed or connected to another feature. The singular forms "a", "said" and "the" used herein are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art. The terms used in this specification are intended only to describe specific embodiments and are not intended to limit the invention. The term "and / or" used herein includes any combination of one or more of the related listed items.
[0054] It should be understood that, although the term first, second, third etc. may be adopted to describe various elements in the present disclosure, these elements should not be limited to these terms. These terms are only used to distinguish the same type of elements from each other. For example, without departing from the scope of the present disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element. The use of any and all examples or exemplary language ("for example", "such as" etc.) provided herein is only intended to better illustrate embodiments of the present invention, and unless otherwise required, will not impose limitations on the scope of the present invention.
[0055] Reference Figure 1 to Figure 2 In some embodiments, the method for monitoring the SNMP extended memory ECC error according to the present invention comprises at least the following steps:
[0056] S100, server initialization starts SNMP;
[0057] S200, polling and collecting data from the master node;
[0058] S300, conducts online (real-time) statistics and analysis, including:
[0059] S310, receiving and processing data;
[0060] S320, using a health grading algorithm, scoring the health of each server and classifying them based on the number of single-bit errors and double-bit errors;
[0061] S330, perform data analysis and trend forecasting to predict future error trends and provide early warning of potential problems based on historical data and using time series models;
[0062] S400. Visualize the data.
[0063] The present invention is extended according to the current SNMP function, and provides an SNMP extended memory ECC error monitoring method and system, which can realize real-time (real-time online) monitoring of memory ECC in the computer room for early warning and prevention. It can be understood that SNMP (Simple Network Management Protocol) can realize effective management and monitoring of various types of equipment, such as routers, switches, servers and other types of equipment. Specifically, network administrators can collect information about network devices through the SNMP protocol. For example, in the present invention, the memory error occurrence recorded by the ECC module (Error Checking and Correcting) can be collected and analyzed by the network management system (NMS) through the SNMP protocol.
[0064] In some embodiments, when the server of the present invention initializes and starts SNMP, SNMPAgent is configured on the server side, and the MIB file is expanded to support ECC data collection. It can be understood that the existing BMC has the service function of SNMP and can be directly started. At the same time, a monitoring statistics service is deployed on the monitoring master node to realize the monitoring of each server.
[0065] Specifically, first define the basic information of the module, use MODULE-IDENTITY to define the module identifier (OID), update time, development organization and description information, and ensure that a unique enterprise OID is assigned to the module. Then import the basic module, introduce the standard SNMPv2-SMI module, and use its MODULE-IDENTITY and OBJECT-TYPE basic types. Then define the object type, use OBJECT-TYPE to define specific monitoring objects, for example, the monitoring object singleBitErrors is defined to record the number of single-bit ECC errors, and the monitoring object doubleBitErrors is defined to record the number of double-bit ECC errors, and set the object's access permission to read-only, and set the range of limit values, for example, the range of limit values is set to 0..4294967295. Then organize the object hierarchy and use the OID hierarchy to define the unique identifier of each object, for example, 1.3.6.1.4.1.63895.1 corresponds to singleBitErrors; 1.3.6.1.4.1.63895.2 corresponds to doubleBitErrors. After completing the definition and saving the file, end the definition with END at the end of the file. Finally, save the MIB file as ECC-MIB.txt to provide SNMP Agent to load.
[0066] In some embodiments, when the present invention polls and collects data from the master node, the master node is set to poll all servers every 5 minutes to obtain batch ECC error statistics, summarize error data within a specified time period, and perform real-time analysis and statistics.
[0067] Specifically, first import the necessary modules, use the high-level interface provided by pysnmp.hlapi to simplify SNMP operations, and introduce the getCmd method to initiate SNMP GET requests to collect data of specific objects. Then define the data collection function, such as the function name definition fetch_ecc_data(server_ip), receive the server's IP address as input, and return the collected ECC data results. Then set the SNMP request and create the necessary parameters of the SNMP request, including: SnmpEngine() for initializing the SNMP engine; CommunityData('public', mpModel=0) for specifying the SNMP community string (such as public) and protocol version (such as SNMPv1); UdpTransportTarget((server_ip, 161)) for setting the target server's IP address and SNMP default port (such as 161); ContextData() for providing context information (usually empty); and ObjectType(ObjectIdentity('OID')) for specifying the MIB object to be collected (OID of ECC-MIB). Then iterate to obtain data, use the iterator returned by getCmd to parse the SNMP response one by one, and judge the collection results. If there is an error (error_indication or error_status), record the error information, otherwise, extract the returned OID and corresponding value. Finally, store and return data, store the collection results in dictionary form (OID: value), and return it to the caller.
[0068] In some embodiments, when the present invention performs real-time statistics and analysis, it first receives and processes data, that is, imports the collected ECC error data into the analysis module, and performs basic preprocessing, which includes operations such as reading, cleaning and standardizing data. Then, a health grading calculation is performed, and a health score is given to each server, and classification is performed based on the number of single-bit errors and double-bit errors. Finally, data analysis and trend prediction are performed, so that in addition to basic classification, trend prediction based on historical data is also used, and the time series method ARIMA (AutoRegressive Integrated Moving Average) is used to predict future error trends, so as to achieve early warning of potential problems.
[0069] In an application embodiment, when the present invention receives and processes data, the original data is first defined, and the ECC error data is defined using a dictionary structure, including: defining Server as the IP address of the server; defining SingleBitErrors as the number of single-bit errors; defining DoubleBitErrors as the number of double-bit errors; and defining Timestamp as the time of data collection. Then a data frame is created, and the original data dictionary is converted into a Pandas data frame using pd.DataFrame to facilitate subsequent processing and analysis. In timestamp formatting, the Timestamp field is converted into the date and time format of Pandas to facilitate time series analysis and sorting operations. Then missing value filling is performed, and fillna(0) is used to fill the missing values in the data frame to avoid the problem of outlier processing in subsequent analysis. Finally, the standardized data is output, so that after the data frame is preprocessed, the data type of each column is consistent and there are no missing values, which can provide a good foundation for subsequent data analysis (such as health grading or trend prediction).
[0070] In an application embodiment, when the present invention performs health grading calculation, the classification threshold is first defined, and then the classification function is defined. Specifically, the function classify_health(row) is defined to evaluate each row of data (each server) and return the corresponding health status according to the threshold condition: Healthy, Mildly Abnormal or Severely Abnormal. Then the classification function is applied, and the classify_health function is applied to each row of data using the DataFrame.apply() method, and the result is stored in a new column HealthStatus for viewing with other information. Finally, the classification result is output, such as printing out the health status table of the server, including Server, SingleBitErrors, DoubleBitErrors and the classification result HealthStatus. Further, a specific embodiment is used here to illustrate that a health conversion data module is provided in the health grading algorithm of the embodiment method of the present invention, which uses Table 1 as input data, and the obtained health conversion data is shown in Table 2.
[0071] Table 1: Input data
[0072]
[0073] Table 2: Health conversion data
[0074]
[0075] Furthermore, in the health grading algorithm of the present invention, the health status is defined by setting a threshold. For example, when the single-bit error is < 20 and the double-bit error is < 2, it is judged to be in a healthy state (Healthy); when the single-bit error is between 20 and 50, and the double-bit error is between 2 and 5, it is judged to be a slightly abnormal state (Mildly Abnormal); when the single-bit error is > 50 or the double-bit error is > 5, it is judged to be a severely abnormal state (Severely Abnormal).
[0076] In an application embodiment, when the present invention performs data analysis and trend forecasting, data is first collected, and a list containing the number of historical ECC errors (such as historical_errors) is used as input data for training a time series model. Then, an ARIMA model is defined, and the parameters of the ARIMA model are specified as order=(p, d, q), where the parameter p is specified as the order of the autoregressive (AR) term, which is used to indicate how many past values the model uses to predict the current value; the parameter d is specified as the number of differences, which is used to convert a non-stationary time series into a stationary series; the parameter q is specified as the order of the moving average (MA) term, which is used to indicate how many past prediction errors the model uses, for example, order=(1, 1, 1) represents first-order autoregression, first difference, and first-order moving average. Then, the model is fitted, and the model.fit() method is used to train the model on the historical data to generate a fitted time series model. Then, a forecast is made, and the forecast(steps=n) method is called to predict the number of errors in the next n time steps (such as predicting the number of errors in the next two days). Finally, the prediction results are output and the predicted values are printed to help administrators understand future ECC error trends.
[0077] In some embodiments, the present invention visualizes the data using Matplotlib to visualize the distribution of health status (see Figure 2 ), and monitor the error trend, and finally generate a monitoring report for later reporting and recording. Further, a specific embodiment is used here to illustrate, and the monitoring report generated by the method of the embodiment of the present invention is shown as follows:
[0078] Health Report - 2024-11-26
[0079] Server: 192.168.1.100
[0080] Single Bit Errors: 15
[0081] Double Bit Errors: 2
[0082] Health Status: Healthy
[0083] Server: 192.168.1.101
[0084] Single Bit Errors: 23
[0085] Double Bit Errors: 1
[0086] Health Status: Mildly Abnormal
[0087] Server: 192.168.1.102
[0088] Single Bit Errors: 5
[0089] Double Bit Errors: 0
[0090] Health Status: Healthy
[0091] It should be appreciated that the method steps in the embodiments of the present invention can be implemented or implemented by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer readable memory. The method can use standard programming techniques. Each program can be implemented in a high-level process or object-oriented programming language to communicate with a computer system. However, if necessary, the program can be implemented in an assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, the program can be run on a programmed ASIC for this purpose.
[0092] In addition, the operations of the processes described herein may be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The processes described herein (or variations and / or combinations thereof) may be performed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors, by hardware, or a combination thereof. The computer program includes a plurality of instructions that may be executed by one or more processors.
[0093] Further, the method can be implemented in any type of computing platform that is operably connected to a suitable computer, including but not limited to a personal computer, a minicomputer, a mainframe, a workstation, a network or distributed computing environment, a separate or integrated computer platform, or in communication with a charged particle tool or other imaging device, etc. Various aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, an optical read and / or write storage medium, an RSM, a ROM, etc., so that it can be read by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the process described herein. In addition, the machine-readable code, or portions thereof, can be transmitted via a wired or wireless network. When such media includes instructions or programs that implement the steps described above in conjunction with a microprocessor or other data processor, the invention described herein includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention can also include the computer itself.
[0094] The computer program can be applied to input data to perform the functions described herein, thereby converting the input data to generate output data stored in a non-volatile memory. The output information can also be applied to one or more output devices, such as a display. In a preferred embodiment of the present invention, the converted data represents physical and tangible objects, including specific visual depictions of physical and tangible objects produced on a display.
[0095] The above is only a preferred embodiment of the present invention. The present invention is not limited to the above implementation. As long as the technical effect of the present invention is achieved by the same means, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the scope of protection of the present invention. Within the scope of protection of the present invention, its technical scheme and / or implementation method may have various modifications and changes.
Claims
1. A method for monitoring simple network management protocol extended memory ECC error, characterized in that: The method comprises the following steps: S100, server initialization starts SNMP; S200, polling and collecting data from the master node; S300, conducts real-time online statistics and analysis, including: S310, receiving and processing data; S320, using a health grading algorithm, scoring the health of each server and classifying them based on the number of single-bit errors and double-bit errors; S330, perform data analysis and trend forecasting to predict future error trends and provide early warning of potential problems based on historical data and using time series models; S400, visualize the data; Wherein, the step S320 includes: S321, setting a classification threshold; S322, setting a classification function; wherein, defining a function classify_health(row) to evaluate each row of data and return the corresponding health status according to the threshold condition: Healthy, Mildly Abnormal or Severely Abnormal; S323. Use the DataFrame.apply() method to apply the classify_health function to each row of data and store the results in a new column HealthStatus. S324, outputting the classification result, wherein the outputted server health status table content items include Server, SingleBitErrors, DoubleBitErrors and the classification result HealthStatus; Wherein, in step S321, the threshold setting for defining the health status includes: When single-bit errors < 20 and double-bit errors < 2, it is considered healthy; When the single-bit error is between 20 and 50, and the double-bit error is between 2 and 5, it is judged as a mild abnormal state; When single-bit errors > 50 or double-bit errors > 5, it is considered a serious abnormality; Wherein, the step S330 further includes: S331, using a list containing the number of historical ECC errors as input data and for training a time series model; S332, specifying parameters of the ARIMA model, including: parameter p is specified as the order of the autoregressive (AR) term, which is used to indicate how many past values the model uses to predict the current value; parameter d is specified as the number of differences, which is used to convert a non-stationary time series into a stationary series; parameter q is specified as the order of the moving average (MA) term, which is used to indicate how many past prediction errors the model uses; S333. Use the model.fit() method to train the model on the historical data and generate a fitted time series model; S334. Call the forecast(steps=n) method to predict the number of errors in the next n time steps, and output the time series model prediction result to obtain the future ECC error trend.
2. The monitoring method according to claim 1, characterized in that: The step S310 includes: S311, defining ECC error data using a dictionary structure to define original data; S312. Use pd.DataFrame to convert the original data dictionary into a Pandas data frame to create a data frame. In timestamp formatting, convert the Timestamp field into the Pandas date and time format. S313, fillna(0) is used to fill the missing values in the data frame to perform missing value filling; S314. Output standardized data.
3. The monitoring method according to claim 2, characterized in that: In the step S311: Define Server as the IP address of the server; define SingleBitErrors as the number of single-bit errors; define DoubleBitErrors as the number of double-bit errors; and define Timestamp as the time when the data was collected.
4. The monitoring method according to claim 1, characterized in that: The step S100 includes: S110. Use MODULE-IDENTITY to define the module identifier (OID), update time, development organization and description information to define the basic information of the module and assign a unique enterprise OID to the module; S120, introduce the standard SNMPv2-SMI module, and use its MODULE-IDENTITY and OBJECT-TYPE basic types to import the basic module; S130, using OBJECT-TYPE to define specific monitoring objects; wherein the monitoring object singleBitErrors is defined as being used to record the number of single-bit ECC errors, and the monitoring object doubleBitErrors is defined as being used to record the number of double-bit ECC errors, and the access rights of the objects are set to read-only, and the range of the limit value is set; S140. Use the OID hierarchy to define a unique identifier for each object to organize the object hierarchy.
5. The monitoring method according to claim 1, characterized in that: The step S200 includes: S210. Use the high-level interface provided by pysnmp.hlapi to simplify SNMP operations, and introduce the getCmd method to initiate SNMP GET requests and collect data of specific objects to import modules; S220, determining a data collection function, including: a function fetch_ecc_data(server_ip) for receiving the IP address of the server as input, and returning the collected ECC data result; S230, create SNMP request parameters to set the SNMP request; including: SnmpEngine() for initializing the SNMP engine; CommunityData() for specifying the SNMP community string and protocol version; UdpTransportTarget() for setting the IP address of the target server and the SNMP default port; ContextData() for providing context information; and ObjectType() for specifying the MIB object to be collected; S240, using the iterator returned by getCmd to parse the SNMP response one by one, to iteratively obtain data, and to judge the collection result; wherein, if there is an error, the error information is recorded, otherwise, the returned OID and the corresponding value are extracted; S250: Store the collected results in a dictionary format and return them to the caller.
6. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the monitoring method according to any one of claims 1 to 5 is implemented.
7. A monitoring system for simple network management protocol extended memory ECC error reporting, characterized in that: include: A computer device comprising a computer readable storage medium according to claim 6.
Citation Information
Patent Citations
Node fault model training method, detection method, equipment, medium and product
CN114726713A
Memory detection method and computing device
CN116775351A