Memory fault diagnosis method, system and device, storage medium and program product
By extracting the characteristic sequence of memory operation information within the preset time period and performing initial and target fault diagnosis, the problem of low memory fault diagnosis accuracy in the prior art is solved, and higher diagnostic accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510922846.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-04
AI Technical Summary
The prior art is difficult to provide dynamic feedback based on the actual operation of memory, resulting in a decrease in the accuracy of memory fault diagnosis.
By extracting the feature sequence of memory operation information within the preset time period, using the operation feature sequence for initial fault diagnosis, determining the feature sequence of diagnostic strategy, and finally troubleshooting, improving diagnostic accuracy.
It realizes dynamic feedback based on memory operation information evolution over time, improving the accuracy and reliability of memory fault diagnosis.
Smart Images

Figure CN120407271A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fault diagnosis, and more particularly to a memory fault diagnosis method, system, device, storage medium and program product. Background Art
[0002] With the rapid development of information technology, servers undertake important business computing and data storage tasks. As one of the core components of a server, the stability and reliability of the memory module are very important for the overall system performance and business continuity.
[0003] However, in the actual operating environment, due to various factors such as memory manufacturing processes, usage environmental conditions, load fluctuations, temperature changes, and device aging, the memory is prone to various types of faults. Related technologies are difficult to provide dynamic feedback based on the actual operating conditions of the memory, resulting in a reduction in the accuracy of memory fault diagnosis. Summary of the Invention
[0004] In view of the above problems, the present invention provides a memory fault diagnosis method, system, device, storage medium and program product.
[0005] According to a first aspect of the present invention, there is provided a memory fault diagnosis method, including: extracting features from memory operation information within a preset time period according to time information within the preset time period to obtain an operation feature sequence, the operation feature sequence including at least one operation feature sorted according to the time information; performing an initial fault diagnosis on the memory using the operation feature sequence to obtain an initial diagnosis result; determining a diagnostic strategy feature sequence of the memory according to the initial diagnosis result and the operation feature sequence; and performing a fault diagnosis on the memory using the diagnostic strategy feature sequence to obtain a target diagnosis result.
[0006] According to a second aspect of the present invention, there is provided a memory fault diagnosis system deployed on a server, including: a collection device for collecting memory operation information within a preset time period; and a memory fault diagnosis device for implementing the steps of the above-mentioned memory fault diagnosis method when executed.
[0007] A third aspect of the present invention provides a memory fault diagnosis device, including: a feature extraction module for extracting features from memory operation information within a preset time period according to time information within the preset time period to obtain an operation feature sequence, the operation feature sequence including at least one operation feature sorted according to the time information; a first fault diagnosis module for performing an initial fault diagnosis on the memory using the operation feature sequence to obtain an initial diagnosis result; a first determination module for determining a diagnostic strategy feature sequence of the memory according to the initial diagnosis result and the operation feature sequence; and a second fault diagnosis module for performing a fault diagnosis on the memory using the diagnostic strategy feature sequence to obtain a target diagnosis result.
[0008] A fourth aspect of the present invention provides an electronic device, including: one or more device processors; a memory for storing one or more computer programs, wherein the one or more device processors execute the one or more computer programs to implement the steps of the memory fault diagnosis method described above.
[0009] A fifth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a device processor, the steps of the memory fault diagnosis method described above are implemented.
[0010] A sixth aspect of the present invention further provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a device processor, the steps of the memory fault diagnosis method described above are implemented.
[0011] According to an embodiment of the present invention, according to the time information within a preset time period, feature extraction is performed on the memory operation information within the preset time period to obtain an operation feature sequence. Therefore, the operation feature sequence has time dimension information. The memory is initially diagnosed for faults using the operation feature sequence to obtain an initial diagnosis result; then, according to the initial fault result and the operation feature sequence, a diagnostic strategy feature sequence of the memory is determined. Therefore, the diagnostic strategy feature sequence includes the initial fault result with dynamic feedback based on the actual memory operation information and the memory operation information with time dimension information. The memory is diagnosed for faults using the diagnostic strategy feature to obtain a target fault result. Therefore, fault diagnosis is performed using the trend of the memory operation information evolving over time and the dynamic feedback information, improving the accuracy of memory fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features, and advantages of the present invention will become clearer. In the drawings:
[0013] Figure 1 An application scenario diagram of the memory fault diagnosis method according to an embodiment of the present invention is shown;
[0014] Figure 2 A flowchart of the memory fault diagnosis method according to an embodiment of the present invention is shown;
[0015] Figure 3A A schematic diagram of a processor card slot according to an embodiment of the present invention is shown;
[0016] Figure 3B A schematic diagram of a first dual in-line memory module slot according to an embodiment of the present invention is shown;
[0017] Figure 4Shows a schematic diagram of a memory fault diagnosis method according to another embodiment of the present invention;
[0018] Figure 5A Shows a structural block diagram of a memory fault diagnosis system according to an embodiment of the present invention;
[0019] Figure 5B Shows a structural block diagram of a server according to an embodiment of the present invention;
[0020] Figure 5C Shows a structural block diagram of a memory fault diagnosis device according to an embodiment of the present invention;
[0021] Figure 6 Shows a block diagram of an electronic device suitable for implementing the memory fault diagnosis method according to an embodiment of the present invention. Detailed implementation manners
[0022] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a thorough understanding of the embodiments of the present invention. However, it is obvious that one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.
[0023] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0024] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0025] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but not be limited to a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0026] With the rapid development of information technology, servers undertake important business computing and data storage tasks. As one of the core components of a server, the stability and reliability of memory are very important for the overall system performance and business continuity.
[0027] However, in the actual operating environment, due to various factors such as memory manufacturing processes, usage environmental conditions, load fluctuations, temperature changes, and device aging, memory is prone to various types of failures. Related technologies are difficult to provide dynamic feedback based on the actual operating conditions of the memory, resulting in a reduction in the accuracy of memory fault diagnosis.
[0028] Embodiments of the present invention provide a memory fault diagnosis method, including: extracting features from memory operation information within a preset time period according to the time information within the preset time period to obtain an operation feature sequence, where the operation feature sequence includes at least one operation feature sorted according to the time information; performing an initial fault diagnosis on the memory using the operation feature sequence to obtain an initial diagnosis result; determining a diagnostic strategy feature sequence of the memory according to the initial diagnosis result and the operation feature sequence; and performing a fault diagnosis on the memory using the diagnostic strategy feature sequence to obtain a target diagnosis result.
[0029] Figure 1 The application scenario diagram of the memory fault diagnosis method according to an embodiment of the present invention is shown.
[0030] As Figure 1 shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0031] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).
[0032] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0033] The server 105 can be a server that provides various services. For example, it can be a background management server (for example only) that supports memory fault diagnosis requests sent by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server can process the received memory fault diagnosis requests from users to obtain memory-related information to be diagnosed. The memory-related information includes a preset time period, memory operation information, and the like. According to the time information within the preset time period, feature extraction is performed on the memory operation information within the preset time period to obtain a running feature sequence; the running feature sequence is used to perform an initial fault diagnosis on the memory to obtain an initial diagnosis result; according to the initial diagnosis result and the running feature sequence, a diagnostic policy feature sequence of the memory is determined; the diagnostic policy feature sequence is used to perform a fault diagnosis on the memory to obtain a target diagnosis result, and the target diagnosis result (such as a web page, information, or data obtained or generated according to a user request) is fed back to the terminal device.
[0034] It should be noted that the memory fault diagnosis method provided in the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the memory fault diagnosis device provided in the embodiments of the present invention can generally be set in the server 105. The memory fault diagnosis method provided in the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the memory fault diagnosis device provided in the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0035] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0036] Figure 2 shows a flowchart of the memory fault diagnosis method according to an embodiment of the present invention.
[0037] As Figure 2 shown, the memory fault diagnosis method of this embodiment includes operation S210 to operation S240, and this memory fault diagnosis method can be executed by a server.
[0038] In operation S210, according to the time information within the preset time period, feature extraction is performed on the memory operation information within the preset time period to obtain a running feature sequence.
[0039] According to an embodiment of the present invention, the preset time period may be a time window for extracting memory operation information. For example, the time window may be 5 minutes, 15 minutes, 30 minutes, 1 hour, 6 hours, 24 hours, 48 hours, 72 hours, 7 days, etc.
[0040] According to an embodiment of the present invention, the time information may include the start time to the end time of the time window. For example, the preset time period may be 5 minutes, the start time is 08:00, the end time is 08:05, and the time information may include any moment between 08:00 and 08:05.
[0041] According to an embodiment of the present invention, the memory operation information may characterize the operation state of the memory. For example, the memory operation information may include log information about the memory, read information of the memory, memory usage rate, the environment where the memory is located, etc. After obtaining the memory operation information, data cleaning and preprocessing may be performed first to improve the data quality.
[0042] According to an embodiment of the present invention, the operation feature sequence includes at least one operation feature sorted according to the time information. For example, the operation feature sequence Ht may be {H1, H2, H3... Hn}. H1, H2, H3... Hn may be sorted from early to late according to the time information. H1, H2, H3... Hn may respectively have error behaviors, memory usage conditions, memory environments, historical fault information that has occurred in the memory, etc.
[0043] Exemplarily, according to the time information within the preset time period, calculating the average value of the memory operation information within the preset time period can capture the long-term trend of the memory operation. For example, the average value of the memory usage rate within the preset time period. The operation feature may be determined based on the average value of the memory usage rate within the preset time period.
[0044] Exemplarily, the maximum value of the memory operation and the minimum value of the memory operation can be obtained from the memory operation information within the preset time period. Based on the maximum value of the memory operation and the minimum value of the memory operation, the change range of the memory operation within the preset time period can be obtained, and the stability feature of the memory operation can be extracted.
[0045] Exemplarily, the preset time period may be 24 hours. The time series decomposition method may be used to decompose the memory operation information within the preset time period into a trend part, a seasonal part, and a residual part. The trend part can reflect whether there is a long-term growth or decline trend in the memory operation information. The seasonal part can reflect whether there are regular fluctuations in the memory operation information. The residual part can characterize that the memory operation information has no obvious rules or patterns. The operation feature may be determined based on the trend, seasonal, and residual parts.
[0046] In operation S220, an initial fault diagnosis of the memory is performed using the operation feature sequence to obtain an initial diagnosis result.
[0047] Exemplarily, the running feature can be obtained based on the average value of the memory running information within a preset time period. The memory running information at each time information can be compared with the average value to obtain the target number of running features greater than the average threshold. In the case where the target number is greater than the number threshold, it can be determined that the initial fault result is that the memory has a fault within the preset time period. In the case where the target number is less than the number threshold, it can be determined that the initial fault result is that the memory has no fault within the preset time period.
[0048] Exemplarily, the running feature can be a stability feature. The stability feature is classified using a machine learning algorithm (such as a random forest, a deep learning model, etc.) to obtain the memory fault type corresponding to the stability feature. The initial diagnosis result is determined according to the memory fault type.
[0049] Exemplarily, the trend part, the seasonal part, and the residual part can be used as running features and input into a classification algorithm (such as a decision tree, a random forest, a support vector machine, etc.) to obtain the initial diagnosis result of the memory. In the case where the residual part shows obvious abnormal fluctuations, an anomaly detection algorithm can be used to further detect faults. Isolation Forest is a tree-based algorithm that isolates outliers by randomly selecting features and randomly partitioning data. The anomaly detection algorithm can be Isolation Forest, One-Class SVM (One-Class Support Vector Machine), etc. One-Class SVM is an algorithm based on support vector machines, and the residual part that exceeds the normal sample boundary is regarded as an anomaly.
[0050] In operation S230, according to the initial diagnosis result and the running feature sequence, the diagnostic strategy feature sequence of the memory is determined.
[0051] According to the embodiments of the present invention, each fault type corresponds to a different quantization value. For example, the fault type can be that the memory has an Error Correcting Code (ECC) check error, and the quantization value can be 2; the fault type can be that the memory has an Uncorrectable Error (UE), and the quantization value can be 1; the fault type can be that the memory has a Correctable Error (CE), and the quantization value can be 0.
[0052] Exemplarily, the initial diagnosis result and the running feature sequence can be concatenated to obtain the diagnostic strategy feature sequence.
[0053] Exemplarily, the memory is repaired according to the initial diagnosis result to obtain a repair result. The initial diagnosis result, the repair result, and the running feature sequence are concatenated to obtain the diagnostic strategy feature sequence.
[0054] Exemplarily, model parameters in a machine learning algorithm related to an initial diagnosis result are determined. The initial diagnosis result, the model parameters, and a sequence of operation characteristics are concatenated to obtain a diagnostic strategy feature sequence.
[0055] In operation S240, the memory is diagnosed for faults using the diagnostic strategy feature sequence to obtain a target diagnosis result.
[0056] According to an embodiment of the present invention, the diagnostic strategy feature sequence can be used to determine a fault diagnosis strategy.
[0057] Exemplarily, the diagnostic strategy feature sequence is matched with reference diagnostic strategy feature sequences of multiple fault diagnosis strategies in a database to obtain a target fault diagnosis strategy that matches the diagnostic strategy feature sequence. The memory is diagnosed for faults using the target fault diagnosis strategy to obtain a target diagnosis result.
[0058] Exemplarily, the diagnostic strategy feature sequence can be input into a strategy classification algorithm (such as a decision tree, random forest, support vector machine, etc.) to obtain a target fault diagnosis strategy. The target fault diagnosis strategy can be to adjust the model parameters of the initial diagnosis result or to determine a target feature extraction method from multiple optional feature extraction methods (such as average value, stability feature, etc.). The specific form of the target fault diagnosis strategy is not limited, but it should be noted that the target fault diagnosis strategy is a feedback for operations S210 to S230 to adjust the fault diagnosis process and improve the accuracy of fault diagnosis.
[0059] According to an embodiment of the present invention, based on the time information within a preset time period, feature extraction is performed on the memory operation information within the preset time period to obtain a sequence of operation characteristics. Therefore, the sequence of operation characteristics has time dimension information. The memory is initially diagnosed for faults using the sequence of operation characteristics to obtain an initial diagnosis result; then, based on the initial fault result and the sequence of operation characteristics, a diagnostic strategy feature sequence of the memory is determined. Therefore, the diagnostic strategy feature sequence includes the initial fault result with dynamic feedback based on the actual memory operation information and the memory operation information with time dimension information. Fault diagnosis is performed on the memory using the diagnostic strategy features to obtain a target fault result, thereby performing fault diagnosis using the trend of the memory operation information evolving over time and the dynamic feedback information, improving the accuracy of memory fault diagnosis.
[0060] According to an embodiment of the present invention, the memory operation information includes at least one of the following: historical fault diagnosis information of the memory, electrical environment information, operating system access information, operating system load information, and memory fault events recorded in the operating system log.
[0061] According to an embodiment of the present invention, operating system access information, operating load information of the operating system, memory fault events recorded in the operating system log, and historical fault diagnosis information can be obtained from the operating system through periodic scanning, real-time monitoring of data interfaces, etc.
[0062] Extract memory fault events that occurred within a preset time period from the operating system log. Memory fault events may include fault types (such as ECC, UE, CE), memory addresses where the memory fault events occurred, and occurrence times, etc. Multiple memory fault events can be sorted according to the occurrence time to record the fault evolution time series. For example, memory fault events may include the number of UEs or CEs that occurred, growth rate, hot row number, error entropy value, propagation speed, error density, etc.
[0063] Operating load information may include the utilization rate of the server's service processor, memory usage rate, time when the service processor is waiting for input / output operations to complete, I / O (input / output) activities related to swapping, task queue length, etc.
[0064] Operating system access information may include the current hot row, page hit rate, read / write bandwidth, cache miss rate, access mode jump situation, etc. A row heat map and behavior map of the memory can also be constructed to assist in spatial anomaly detection.
[0065] Since faults often have an evolution path, the memory operation information includes historical fault diagnosis information and memory fault events recorded in the log. Feature extraction can be performed on the memory operation information within a preset time period according to the time information within the preset time period to obtain an operation feature sequence, so that the operation feature sequence contains the trend of fault evolution over time.
[0066] Historical fault diagnosis information may include process information of historical fault diagnosis, historical diagnosis results, historical repair results, etc. Historical diagnosis results may include whether they are consistent with the actual fault (such as the location where the fault occurred, fault type, etc.); historical repair results include whether actual repair operations are triggered (such as page migration), repair time, etc.
[0067] According to an embodiment of the present invention, power environment information can be obtained through a baseboard management controller. Power environment information may include the temperature, voltage, current, power consumption slope, etc. of the DIMM (Dual Inline Memory Module) at the address where the memory fault event occurred.
[0068] According to an embodiment of the present invention, the memory operation information may have multi-dimensional information, such as the historical fault diagnosis information of the memory, the electrical environment information, the operating system access information, the operating load information of the operating system, and the memory fault events recorded in the operating system log. Therefore, the operation feature sequence may have trends such as faults, electrical environment, operating system access, and operating load evolving over time. Using the operation feature sequence for fault diagnosis improves the reliability of memory fault diagnosis.
[0069] According to an embodiment of the present invention, the above method further includes: obtaining memory fault events within a preset time period from the log; locating the memory address where the memory fault event occurs according to the preset access path of the operating system, and the memory operation information further includes the memory address; obtaining the electrical environment information at the memory address.
[0070] According to an embodiment of the present invention, the memory address may be a preset access path set according to the processor socket (CPU Socket), memory channel (Channel), dual in-line memory module slot, memory bank hierarchy, chip, and row / column.
[0071] According to an embodiment of the present invention, the memory address where the memory fault event occurs can be located step by step according to the preset access path to the smallest unit.
[0072] Figure 3A The schematic diagram of the processor card slot according to an embodiment of the present invention is shown.
[0073] As Figure 3A shown, the processor card slot is an interface on the motherboard for inserting a service processor. The server supports multiple processor card slots, and each processor card slot is installed with a multi-core service processor. One or more integrated memory controllers (IMCs) are integrated inside each processor card slot. It interacts with the system memory through the first memory channel and the second memory channel. In a Non-Uniform Memory Access (NUMA) architecture, the memory and input / output resources directly controlled by each processor card slot form a NUMA node. In actual deployment, a NUMA node is composed of a processor card slot, the first dual in-line memory module slot and the second dual in-line memory module slot on the first memory channel, and the third dual in-line memory module slot and the fourth dual in-line memory module slot on the second memory channel.
[0074] Figure 3B The schematic diagram of the first dual in-line memory module slot according to an embodiment of the present invention is shown.
[0075] AsFigure 3B As shown, each memory array level of the first dual in-line memory module slot has multiple particles, such as particle 1, particle 2, particle 3, particle 4, particle 5, etc. The memory array levels are usually organized in the form of a two-dimensional array. For example, particle 5 is divided into rows and columns. Row errors, column errors, byte errors, etc. may occur on particle 5.
[0076] According to an embodiment of the present invention, historical fault diagnosis information at a memory address, operating system access information related to the memory address, and operating load information of the operating system can be obtained. Furthermore, the obtained memory operation data is related to memory fault events, improving data quality and reducing the amount of memory operation information, thereby enhancing the computing performance of the server.
[0077] According to an embodiment of the present invention, memory fault events within a preset time period are obtained from the log; according to a preset access path of the operating system, the memory address where the memory fault event occurs is located; then the electrical environment information at the memory address is obtained. Thus, the obtained memory operation information is information related to memory fault events, which can reduce the amount of memory operation information, enhance the computing performance of the server, and improve the response timeliness of fault diagnosis.
[0078] According to an embodiment of the present invention, feature extraction is performed on the memory operation information within a preset time period according to the time information within the preset time period to obtain a running feature sequence, including: determining attribute statistical features according to the time information and multiple attribute data in the memory operation information; performing vectorization processing on the memory operation information according to the time information to obtain a running state feature sequence; and splicing the attribute statistical features and the running state feature sequence to obtain the running feature sequence.
[0079] According to an embodiment of the present invention, the attribute statistical features can represent the change of the attribute data within a preset time. For example, the attribute statistical features may include the maximum value, minimum value, extreme value, variance, average value, cumulative value, etc. of the attribute data within a preset time period.
[0080] According to an embodiment of the present invention, the vectorized memory operation information is sorted according to the time information to obtain a running state feature sequence. The running state feature sequence includes at least one running state feature, and the running state feature can be obtained by splicing multiple attribute dimensions. The running state feature can represent the instantaneous running state of the memory.
[0081] To enhance the model's ability to perceive the fault evolution trend, the memory operation information is vectorized according to the time information to obtain the operation state feature sequence. For example, there are a total of w moments within a preset time period, which can be from t - w + 1, t - w + 2 to the moment t. The memory operation information can include n attribute data. The memory operation information of w moments is concatenated to obtain the memory operation information sequence within the preset time period 。
[0082] (1);
[0083] w is the number of moments within the preset time period, and n is the number of attribute data. The memory operation information is vectorized according to the time information to obtain the vector feature sequence. The vector feature sequence is unfolded into a one-dimensional feature vector to obtain the operation state feature sequence ,R wn represents a vector space of w×n dimensions. The operation state feature sequence includes at least one operation state feature. For example, the operation state feature sequence includes the operation state features vec(S t-w+1 ), vec(S t-w+2 )... vec(S t ).
[0084] (2);
[0085] Exemplarily, there are operation state features of 2 moments within the preset time period, and the memory operation information can include 2 attribute data. The vector feature sequence is shown in formula (3):
[0086] (3);
[0087] The vector feature sequence is unfolded into a one-dimensional feature vector as the operation state feature sequence vec( ), vec( ) = [10 20 12 22].
[0088] According to an embodiment of the present invention, the attribute statistical feature is concatenated with the operation state feature sequence unfolded into a one-dimensional feature vector to obtain the operation feature sequence.
[0089] According to an embodiment of the present invention, by determining the attribute statistical feature based on the time information and multiple attribute data in the memory operation information; concatenating the attribute statistical feature and the operation state feature sequence to obtain the operation feature sequence, so that the operation feature sequence has the change trend of the attribute data over time. Therefore, the operation feature sequence is used for initial fault diagnosis to improve the accuracy of the initial diagnosis result.
[0090] According to an embodiment of the present invention, the attribute statistical feature includes at least one of the following: a gradient feature representing the operation change trend of the memory within a preset time period, a fluctuation feature representing the fluctuation condition of the attribute data within a preset time period, and a neighboring state feature representing the change trend of the attribute data at adjacent moments.
[0091] According to an embodiment of the present invention, the gradient feature k t has the following formula:
[0092] (4);
[0093] where i represents the i-th moment within the preset time period, , .
[0094] According to an embodiment of the present invention, the fluctuation feature has the following formula:
[0095] (5);
[0096] According to an embodiment of the present invention, for the neighboring state feature the formula is as follows:
[0097] (6);
[0098] ε can be 0.0001 to avoid the denominator being 0.
[0099] For example, the attribute statistical feature may include a gradient feature, a fluctuation feature, and a neighboring state feature, and the attribute statistical feature is a one-dimensional feature vector. The attribute statistical feature and the operation state feature sequence are concatenated to obtain the operation feature sequence X t , that is, .
[0100] According to an embodiment of the present invention, the three attribute statistical features, namely, the gradient feature representing the operation change trend of the memory within a preset time period, the fluctuation feature representing the fluctuation condition of the attribute data within a preset time period, and the neighboring state feature representing the change trend of the attribute data at adjacent moments, can reflect the change of the memory operation within the preset time period and improve the reliability of the fault diagnosis result.
[0101] According to an embodiment of the present invention, an initial fault diagnosis of the memory is performed using a sequence of operating characteristics to obtain an initial diagnosis result, including: performing time encoding on the operating characteristics according to the order of the operating characteristics in the sequence of operating characteristics to obtain time characteristics; embedding the time characteristics into a preset position of the operating characteristics to obtain time-aware characteristics; processing the time-aware characteristics using a multi-head self-attention mechanism to obtain target attention characteristics; and performing an initial fault diagnosis of the memory using the sequence of target attention characteristics to obtain an initial diagnosis result, where the sequence of target attention characteristics includes at least one target attention characteristic.
[0102] According to an embodiment of the present invention, when the sequence number of the operating characteristic in the sequence of operating characteristics is even, that is, i is even, the even time characteristics are as follows:
[0103] (7);
[0104] When the sequence number of the operating characteristic in the sequence of operating characteristics is odd, that is, i is odd, the odd time characteristics are as follows:
[0105] (8);
[0106] In formulas (7) and (8), j is any positive integer, and d h = Z / h, where Z is the dimension of the sequence of operating characteristics and h is the number of attention heads of the multi-head attention mechanism. The time characteristics PE include even time characteristics and odd time characteristics.
[0107] According to an embodiment of the present invention, the preset position may be the head or tail of the operating state characteristic in the operating characteristics. Embedding the time characteristics PE into the preset position of the operating characteristics to obtain a sequence of time-aware characteristics, the sequence of time-aware characteristics is as follows:
[0108] (9);
[0109] According to an embodiment of the present invention, performing time encoding on the operating characteristics according to the order of the operating characteristics in the sequence of operating characteristics to obtain time characteristics; embedding the time characteristics into a preset position of the operating characteristics to obtain time-aware characteristics. By processing the time-aware characteristics using a multi-head self-attention mechanism to obtain target attention characteristics, the multi-head attention mechanism extracts the global dependence characteristics between each moment, that is, the target attention characteristics have the global characteristics of the memory operation within a preset time period. Performing an initial fault diagnosis of the memory using the sequence of target attention characteristics to obtain an initial diagnosis result can reduce the influence of local information on the diagnosis and improve the accuracy of the initial diagnosis result.
[0110] According to an embodiment of the present invention, the time perception features are processed using a multi-head self-attention mechanism to obtain target attention features, including: processing the time perception features using a multi-head self-attention mechanism to obtain initial self-attention features; normalizing the time perception features and the initial self-attention features to obtain intermediate attention features; and inputting the intermediate attention features into a position feed-forward network to obtain target attention features.
[0111] According to an embodiment of the present invention, the time perception features are mapped to queries (Q), keys (K), and values (V) using a multi-head attention mechanism. The formula is as follows:
[0112] (10);
[0113] W Q 、W K 、W V are the weight matrices for queries, keys, and values respectively, and the weight matrices are trainable parameters.
[0114] The output of the i-th head attention mechanism is calculated based on Q, K, and V. The formula is as follows:
[0115] (11);
[0116] Q K T is the dot product between the query and the key, is a scaling operation to avoid the dot product result from being too large. The output head i of the single-head attention mechanism is obtained through an activation function (such as the softmax function).
[0117] The initial self-attention feature MultiHead(X) is as follows:
[0118] (12);
[0119] W O is the output weight matrix to maintain the stability of the multi-head attention mechanism during training. The outputs head1 of the first head attention mechanism to headh of the h-th head attention mechanism are concatenated to obtain the initial self-attention feature.
[0120] LayerNorm is a normalization function. After concatenating the time perception feature sequence and the initial self-attention feature MultiHead(X), the feature sequence Z1 to be normalized is obtained.
[0121] (13);
[0122] Input the feature sequence to be normalized into the position feed-forward network (FFN) to obtain the feedback feature sequence FFN(Z1).
[0123] (14);
[0124] The inputs W1 and W2 are the network weight matrices of the position feed-forward network, and b1 and b2 are the two biases of the position feed-forward network respectively. ReLU is the activation function.
[0125] After concatenating the feedback feature sequence FFN(Z1) and the feature sequence Z1 to be normalized, then perform normalization to obtain the target attention feature sequence, and the target attention feature sequence includes at least one target attention feature. The target attention feature sequence H t The formula is as follows:
[0126] (15);
[0127] The above steps constitute a block (Block). For memory fault diagnosis, it can include 3 Blocks.
[0128] According to an embodiment of the present invention, input the target attention feature sequence into the fault prediction model to obtain the initial diagnosis result. For example, the fault prediction model can be a decision tree, and the output of the decision tree is:
[0129] (16);
[0130] is the m-th decision tree, represents the probability of a memory fault occurring in the memory within the future Δt time, f GBDT represents the decision tree function.
[0131] The training objective of the decision tree can be that the total loss reaches the minimum value.
[0132] (17);
[0133] represents the binary cross-entropy loss, is the first regularization coefficient, is the second regularization coefficient, represents the sample pair (u, v) in the set β, which is used to calculate the difference loss between sample pairs. are respectively the squares of the Euclidean norms of the network weight matrices W1 and W2. is the output obtained by inputting the sample u into the m-th decision tree, is the output obtained by inputting sample v into the m-th decision tree. The memory addresses corresponding to sample u and sample v are the same. For example, both sample u and sample v correspond to the same row or column.
[0134] According to an embodiment of the present invention, by using the multi-head self-attention mechanism to process the time-aware features, the initial self-attention features are obtained, which can capture diverse information of the time-aware features and enhance the expressive ability of the fault prediction model for fault diagnosis. The multi-attention mechanism performs parallel computing, improving the computing efficiency. The time-aware features and the initial self-attention features are normalized to obtain the intermediate attention features; the intermediate attention features are input into the position feed-forward network to obtain the target attention features. Therefore, the intermediate attention features can perform non-linear transformation independently at each position, improving the stability of the fault prediction model.
[0135] According to an embodiment of the present invention, the memory operation information includes the historical fault diagnosis information of the memory, and the historical fault diagnosis information includes at least one of the following: the historical repair result corresponding to the historical target prediction strategy, the historical risk threshold.
[0136] Exemplarily, the start time of the preset time period is t, and the historical risk threshold can be the risk threshold set for the fault prediction model in the fault diagnosis at time t - 1. The historical repair result corresponding to the historical target prediction strategy can be the repair result of the fault at time t - 1, and the repair result can be 1 for successful repair and 0 for failed repair.
[0137] Exemplarily, the memory operation information can also include the change rate of the number of memory fault events within the preset time period, the change rate of the electrical environment information within the preset time period, and the change rate of the memory access frequency within the preset time period.
[0138] According to an embodiment of the present invention, according to the initial diagnosis result and the operation feature sequence, the diagnostic strategy feature sequence of the memory is determined, including: determining the diagnostic strategy feature sequence according to the time-aware feature sequence and the initial diagnosis result.
[0139] According to an embodiment of the present invention, the time-aware features and the initial diagnosis result are concatenated to obtain the diagnostic strategy feature sequence.
[0140] Exemplarily, the diagnostic strategy feature sequence G t has the following formula:
[0141] (18);
[0142] According to an embodiment of the present invention, the diagnostic strategy feature sequence can also include the risk threshold of the initial fault diagnosis, the repair method of the initial fault diagnosis, the number of wrong behaviors earlier than the preset time period, the transformation rate of the electrical environment, etc.
[0143] The risk threshold of the initial fault diagnosis and the repair method of the initial fault diagnosis can truly reflect the process of the initial fault diagnosis, so as to determine a suitable target prediction strategy, enabling the fault detection model to quickly adapt to the changes in the scenario.
[0144] The number of error behaviors earlier than the preset time period, the transformation rate of the electrical environment, etc. can analyze the operation state information more comprehensively, and can judge the reliability of the memory operation data at the start of the preset time period, thereby improving the data quality of the diagnostic strategy features and improving the computing performance.
[0145] According to an embodiment of the present invention, according to the time-aware feature sequence and the initial diagnosis result, a diagnostic strategy feature sequence of the memory is determined. That is, the diagnostic strategy feature sequence is dynamically feedback according to the initial fault result obtained from the actual memory operation information, and has the trend information and feedback information of the memory fault evolving over time.
[0146] According to an embodiment of the present invention, the memory is fault-diagnosed using the diagnostic strategy feature sequence to obtain a target diagnosis result, including: inputting the diagnostic strategy feature sequence into a strategy prediction model to obtain a target prediction strategy, and the strategy prediction model is trained based on the reward function values of multiple prediction strategies of the fault prediction model corresponding to the historical diagnostic strategy feature sequence. The fault prediction model is used to perform an initial fault diagnosis on the memory using the operation feature sequence to obtain an initial diagnosis result; using the fault prediction model to perform a fault diagnosis on the memory according to the target prediction strategy to obtain a target diagnosis result.
[0147] According to an embodiment of the present invention, the strategy prediction model can be a multi-layer perceptron network.
[0148] Exemplarily, the strategy prediction model can adopt a three-layer multi-layer perceptron.
[0149] (19);
[0150] Diagnostic strategy feature sequence G t Function for inputting into a 3-layer multi-layer perceptron (MLP) network , the dimension of the input layer is consistent with the state space, the hidden layer can be 128-dimensional, and the output layer is the corresponding prediction strategy A t The spatial size of, and the dimension of the output vector can be determined according to the number of prediction strategies. The probabilities of multiple prediction strategies being selected are output after passing through the softmax function .
[0151] According to an embodiment of the present invention, the fault prediction model and the strategy prediction model are used in combination. By inputting the operation feature sequence into the fault prediction model, an initial diagnosis result can be obtained. The initial diagnosis result is concatenated with the operation feature sequence to obtain the diagnosis strategy feature. The diagnosis strategy feature is input into the strategy prediction model to obtain the target prediction strategy, enabling a prediction feedback mechanism in the fault diagnosis process. Therefore, in dynamic scenarios (such as unstable load, sudden increase in memory temperature), the parameters of the fault prediction model can be adaptively adjusted to suit diverse scenario changes and improve the accuracy of fault diagnosis in different scenarios.
[0152] Figure 4 The figure shows a schematic diagram of a memory fault diagnosis method according to an embodiment of the present invention.
[0153] As Figure 4 shown, by inputting the target attention feature sequence into the fault prediction model 410, an initial diagnosis result can be obtained. The initial diagnosis result is concatenated with the time-aware feature sequence to obtain the diagnosis strategy feature sequence 420. The diagnosis strategy feature sequence 420 is input into the strategy prediction model 430 to obtain the target prediction strategy. The memory is fault-diagnosed using the fault prediction model 410 according to the target prediction strategy to obtain the target diagnosis result 440.
[0154] According to an embodiment of the present invention, the prediction strategy includes one of the following: processing the target attention feature sequence using the fault prediction model to obtain the target diagnosis result; determining a target operation feature subsequence corresponding to the preset memory operation information from the operation feature sequence; updating the risk threshold of the fault prediction model based on the target operation feature subsequence and using the updated fault prediction model to perform fault diagnosis on the memory. The preset memory operation information includes the electrical environment information of the operating system and the operating system access information. The risk threshold is used to compare with the output value of the fault prediction model to determine the diagnosis result of the memory according to the comparison result; determining the target memory operation information for fault diagnosis from the memory operation information within the preset time period according to the preset duration; adopting a repair method for scheduling the storage space of the memory; temporarily not performing any operation.
[0155] According to an embodiment of the present invention, the prediction strategy may involve the data dimension of the input to the fault prediction model, whether to link with the fault modification module, the memory operation information for feature extraction, etc.
[0156] The prediction strategy is shown in Table 1.
[0157] Table 1
[0158]
[0159] According to an embodiment of the present invention, the preset time length may be 10 minutes, and the time length within the preset time period may be 15 minutes. For example, target memory operation information within 10 minutes for fault diagnosis is determined from the memory operation information within the preset time period.
[0160] A repair method of the storage space of the scheduling memory is adopted. For example, the repair method of the storage space of the scheduling memory may be to execute memory page migration or adjust the storage capacity of the memory in a linked manner.
[0161] Temporarily not performing any operation means not triggering any behavior, that is, observing state.
[0162] According to embodiments of the present invention, prediction strategies can include factors such as the data dimensions input to the fault prediction model, whether a fault correction module needs to be linked, and memory operation information used for feature extraction. These strategies can be adjustments to the initial fault diagnosis process. Utilizing a fault prediction model and following a targeted prediction strategy for memory fault diagnosis results in more accurate results, improving the robustness of the fault prediction model.
[0163] According to an embodiment of the present invention, the strategy prediction model is trained based on the following method: using the fault prediction model to process the historical diagnosis strategy features according to multiple prediction strategies to obtain the reward function values of multiple prediction strategies; determining the target reward value based on the multiple reward function values; updating the parameters of the pre-trained strategy prediction model based on the target reward value until the target reward value reaches the preset reward value to obtain the strategy prediction model.
[0164] According to an embodiment of the present invention, prediction strategy A t The reward function value R t+1 The formula is as follows:
[0165] (20);
[0166] Indicates whether the diagnosis result obtained according to the prediction strategy is the same as the actual fault result, 1 if they are the same, and 0 if they are not the same. Indicates whether the diagnosis result obtained according to the prediction strategy has an error (memory address, fault type, etc.). If so, it is 1, otherwise it is 0. Indicates the resource cost consumed in executing the prediction strategy. Indicates the repair result according to the diagnosis result obtained by the prediction strategy. If the repair is successful, it is 1, otherwise it is 0. α1, α2, α3 and α4 are 、 、 、 The preset reward weights.
[0167] The preset reward value can be within the time step, the prediction strategy At The target reward value reaches the maximum value. The formula for the target reward value J(θ) is as follows:
[0168] (21);
[0169] γ is the discount factor (γ can take the value of 0.9), T is the time step of the policy prediction model, is the expected function, π is the pi, and the parameter of the policy prediction model is θ. By continuously updating the parameter θ, J(θ) reaches the maximum value, and the policy prediction model is obtained.
[0170] According to an embodiment of the present invention, the fault prediction model is used to process the historical diagnosis policy features according to multiple prediction strategies to obtain the reward function values of multiple prediction strategies; according to the multiple reward function values, the target reward value is determined; based on the target reward value, the parameters of the pre-trained policy prediction model are updated until the target reward value reaches the preset reward value, and the policy prediction model is obtained, realizing training the policy prediction model using the output result of the fault prediction model, making the policy prediction model more suitable for the scenario of the fault prediction model, and improving the accuracy of the memory diagnosis result.
[0171] Figure 5A Shows a structural block diagram of a memory fault diagnosis system according to an embodiment of the present invention.
[0172] The memory fault diagnosis system 500 can be deployed on the server 105. The memory fault diagnosis system 500 can include a collection device 510 and a memory fault diagnosis device 520. The collection device 510 is used to collect the memory operation information within a preset time period. The memory fault diagnosis device 520 is used to implement the steps of the memory fault diagnosis method when executed.
[0173] Figure 5B Shows a structural block diagram of a server according to an embodiment of the present invention.
[0174] According to an embodiment of the present invention, the server 105 can include a baseboard management controller 1051, a memory 1052, an operating system 1053, and a service processor 1054. The baseboard management controller 1051 is connected to the memory controller to monitor the electrical environment information of the memory 1052; the memory 1052 is connected to the operating system 1053, and the operating system 1053 is used to schedule the storage space of the memory 1052.
[0175] According to an embodiment of the present invention, the memory 1052 is connected to the service processor 1054 of the server 105 via a memory channel. The server 105 can include multiple service processors 1054.
[0176] According to an embodiment of the present invention, the acquisition device 510 is configured to: acquire the electrical environment information of the memory 1052 from the baseboard management controller 1051; acquire the operating system access information, the operating load information of the operating system 1053, and the memory failure events recorded in the log of the operating system 1053 from the operating system 1053; acquire the historical failure diagnosis information of the memory 1052 from the memory failure diagnosis device 520, and the memory operation information includes at least one of the following: historical failure diagnosis information, electrical environment information, operating system access information, operating load information, and memory failure events.
[0177] For example, the operating system access information can be obtained through a performance analysis tool. The operating system access information may include current hot rows, page hit rate, read / write bandwidth utilization, etc. The performance analysis tool can be perf (Performance Events) event sampling.
[0178] For example, the acquisition device 510 can use a tool for monitoring and recording hardware errors to acquire memory failure events. The tool for monitoring and recording hardware errors can be mcelog (Machine Check Exception Logging), EDAC (Error Detection And Correction) framework, RasDaemon (Reliability, Availability, and Serviceability Daemon). The memory failure events can be scanned periodically or listened to in real time through the kernel interface.
[0179] For example, the acquisition device 510 can access the temperature, voltage, current, power consumption data, etc. of the dual in-line memory module under the operating load through IPMI (Intelligent Platform Management Interface). Access the SPD (Serial Presence Detec) element of each dual in-line memory module through the system management bus to obtain the module temperature and calibrated electrical parameters.
[0180] For example, the acquisition device 510 can obtain the model parameters corresponding to the historical target prediction strategy in the historical failure diagnosis information through the internal log of the memory failure diagnosis device 520, obtain the historical risk threshold through the early warning linkage system, and obtain the historical repair results corresponding to the historical target prediction strategy through the basic input / output system.
[0181] According to an embodiment of the present invention, the acquisition device 510 may include a plurality of acquisition probes, which are respectively deployed on the baseboard management controller 1051, the operating system 1053, and the memory fault diagnosis device 520.
[0182] Figure 5C The structural block diagram of a memory fault diagnosis device according to an embodiment of the present invention is shown.
[0183] As Figure 5C shown, the memory fault diagnosis device 520 of this embodiment includes a feature extraction module 521, a first fault diagnosis module 522, a first determination module 523, and a second fault diagnosis module 524.
[0184] The feature extraction module 521 is used to extract features from the memory operation information within a preset time period according to the time information within the preset time period, so as to obtain an operation feature sequence, and the operation feature sequence includes at least one operation feature sorted according to the time information. In one embodiment, the feature extraction module 521 may be used to execute the operation S210 described above, which will not be elaborated here.
[0185] The first fault diagnosis module 522 is used to perform an initial fault diagnosis on the memory by using the operation feature sequence to obtain an initial diagnosis result. In one embodiment, the first fault diagnosis module 522 may be used to execute the operation S220 described above, which will not be elaborated here.
[0186] The first determination module 523 is used to determine the diagnostic strategy feature sequence of the memory according to the initial diagnosis result and the operation feature sequence. In one embodiment, the first determination module 523 may be used to execute the operation S230 described above, which will not be elaborated here.
[0187] The second fault diagnosis module 524 is used to perform a fault diagnosis on the memory by using the diagnostic strategy feature sequence to obtain a target diagnosis result. In one embodiment, the second fault diagnosis module 524 may be used to execute the operation S240 described above, which will not be elaborated here.
[0188] According to an embodiment of the present invention, the first fault diagnosis module 522 includes a time encoding sub-module, an embedding sub-module, a processing sub-module, and a first fault diagnosis sub-module. The time encoding sub-module is used to perform time encoding on the operation features according to the order of the operation features in the operation feature sequence to obtain time features; the embedding sub-module is used to embed the time features into preset positions of the operation features to obtain time-aware features; the processing sub-module is used to process the time-aware features by using a multi-head self-attention mechanism to obtain target attention features; the first fault diagnosis sub-module is used to perform an initial fault diagnosis on the memory by using the target attention feature sequence to obtain an initial diagnosis result, and the target attention feature sequence includes at least one target attention feature.
[0189] According to an embodiment of the present invention, the processing sub-module includes a processing unit, a normalization unit, and an input unit. The processing unit is configured to process the time-aware features by using a multi-head self-attention mechanism to obtain initial self-attention features; the normalization unit is configured to normalize the time-aware features and the initial self-attention features to obtain intermediate attention features; the input unit is configured to input the intermediate attention features into a position feed-forward network to obtain target attention features.
[0190] According to an embodiment of the present invention, the first determination module 523 includes a first determination sub-module. The first determination sub-module is configured to determine a diagnostic policy feature sequence according to the time-aware feature sequence and the initial diagnostic result.
[0191] According to an embodiment of the present invention, the above device further includes a first acquisition module, a positioning module, and a second acquisition module. The first acquisition module is configured to acquire memory failure events within a preset time period from the log; the positioning module is configured to locate the memory address where the memory failure event occurs according to a preset access path of the operating system, and the memory operation information further includes the memory address; the second acquisition module is configured to acquire the electrical environment information at the memory address.
[0192] According to an embodiment of the present invention, the feature extraction module 521 includes a second determination sub-module, a vector processing sub-module, and a splicing sub-module. The second determination sub-module is configured to determine attribute statistical features according to multiple attribute data in the time information and the memory operation information; the vector processing sub-module is configured to perform vectorization processing on the memory operation information according to the time information to obtain a running state feature sequence; the splicing sub-module is configured to splice the attribute statistical features and the running state feature sequence to obtain a running feature sequence.
[0193] According to an embodiment of the present invention, the second fault diagnosis module 524 includes an input sub-module and a second fault diagnosis sub-module. The input sub-module is configured to input the diagnostic policy feature sequence into a policy prediction model to obtain a target prediction policy, and the policy prediction model is trained based on the reward function values of multiple prediction policies of a fault prediction model corresponding to a historical diagnostic policy feature sequence, and the fault prediction model is configured to perform an initial fault diagnosis on the memory by using the running feature sequence to obtain an initial diagnostic result; the second fault diagnosis sub-module is configured to perform a fault diagnosis on the memory according to the target prediction policy by using the fault prediction model to obtain a target diagnostic result.
[0194] According to an embodiment of the present invention, any plurality of modules among the feature extraction module 521, the first fault diagnosis module 522, the first determination module 523, and the second fault diagnosis module 524 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the feature extraction module 521, the first fault diagnosis module 522, the first determination module 523, and the second fault diagnosis module 524 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or any other reasonable manner that can integrate or package circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the feature extraction module 521, the first fault diagnosis module 522, the first determination module 523, and the second fault diagnosis module 524 may be at least partially implemented as a computer program module, and when the computer program module is run, it can execute corresponding functions.
[0195] Figure 6 The block diagram of an electronic device suitable for implementing the memory fault diagnosis method according to an embodiment of the present invention is shown.
[0196] As Figure 6 shown, the electronic device 600 according to an embodiment of the present invention includes a device processor 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage part 608 into the random access memory (RAM) 603. The device processor 601 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The device processor 601 may also include on-board memory for caching purposes. The device processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0197] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The device processor 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The device processor 601 performs various operations of the method flow according to the embodiments of the present invention by executing the programs in the ROM 602 and / or the RAM 603. It should be noted that the programs can also be stored in one or more memories other than the ROM 602 and the RAM 603. The device processor 601 can also perform various operations of the method flow according to the embodiments of the present invention by executing the programs stored in the one or more memories.
[0198] According to an embodiment of the present invention, the electronic device 600 may further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the input / output (I / O) interface 605: an input part 606 including a keyboard, a mouse, etc.; an output part 607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage part 608 including a hard disk, etc.; and a communication part 609 including a network interface card such as a LAN card, a modem, etc. The communication part 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed into the storage part 608 as needed.
[0199] The present invention also provides a computer-readable storage medium, which may be included in the device / device / system described in the above embodiments; or may exist separately without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.
[0200] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-described ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603.
[0201] An embodiment of the present invention also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the memory fault diagnosis method provided by the embodiment of the present invention.
[0202] When the computer program is executed by the device processor 601, it executes the above functions defined in the system / apparatus of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0203] In one embodiment, the computer program can rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program can also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 609, and / or be installed from the removable medium 611. The program code contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0204] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or be installed from the removable medium 611. When the computer program is executed by the device processor 601, it executes the above functions defined in the system of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0205] According to an embodiment of the present invention, program code for executing the computer program provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0206] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0207] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0208] The above describes the embodiments of the present invention. However, these embodiments are only for illustrative purposes and are not intended to limit the scope of the present invention. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. A method for diagnosing memory faults, characterized in that, The memory fault diagnosis method includes: According to the time information within a preset time period, extract features from the memory operation information within the preset time period to obtain an operation feature sequence, where the operation feature sequence includes at least one operation feature sorted according to the time information; Use the operation feature sequence to perform an initial fault diagnosis on the memory to obtain an initial diagnosis result; According to the initial diagnosis result and the operation feature sequence, determine a diagnostic strategy feature sequence of the memory; Use the diagnostic strategy feature sequence to perform a fault diagnosis on the memory to obtain a target diagnosis result.
2. The memory fault diagnosis method according to claim 1, characterized in that The step of using the operation feature sequence to perform an initial fault diagnosis on the memory to obtain an initial diagnosis result includes: According to the order of the operation features in the operation feature sequence, perform time encoding on the operation features to obtain time features; Embed the time features into preset positions of the operation features to obtain time-aware features; Use a multi-head self-attention mechanism to process the time-aware features to obtain target attention features; Use the target attention feature sequence to perform an initial fault diagnosis on the memory to obtain the initial diagnosis result, where the target attention feature sequence includes at least one of the target attention features.
3. The memory fault diagnosis method according to claim 2, wherein The step of using a multi-head self-attention mechanism to process the time-aware features to obtain target attention features includes: Use the multi-head self-attention mechanism to process the time-aware features to obtain initial self-attention features; Normalize the time-aware features and the initial self-attention features to obtain intermediate attention features; Input the intermediate attention features into a position feed-forward network to obtain the target attention features.
4. The memory fault diagnosis method according to claim 2, wherein The step of determining a diagnostic strategy feature sequence of the memory according to the initial diagnosis result and the operation feature sequence includes: Determine the diagnostic strategy feature sequence according to the time-aware feature sequence and the initial diagnosis result.
5. The memory fault diagnosis method according to any one of claims 1 to 4, characterized in that, The memory operation information includes at least one of the following: historical fault diagnosis information of the memory, electrical environment information, operating system access information, operating system operation load information, memory fault events recorded in the operating system log.
6. The memory fault diagnosis method according to claim 5, characterized in that, The memory fault diagnosis method further includes: Obtain memory fault events within the preset time period from the operating system log; According to a preset access path of the operating system, locate the memory address where the memory fault event occurs, and the memory operation information further includes the memory address; Obtain the electrical environment information at the memory address.
7. The memory fault diagnosis method according to claim 1, characterized in that, The step of extracting features from the memory operation information within a preset time period according to the time information within the preset time period to obtain an operation feature sequence includes: According to the time information and multiple attribute data in the memory operation information, determine attribute statistical features; Perform vectorization processing on the memory operation information according to the time information to obtain an operation state feature sequence; Concatenate the attribute statistical features and the operation state feature sequence to obtain the operation feature sequence.
8. The memory fault diagnosis method according to claim 7, characterized in that The attribute statistical features include at least one of the following: a gradient feature representing the running change trend of the memory within the preset time period, a fluctuation feature representing the fluctuation of the attribute data within the preset time period, and a proximity state feature representing the change trend of the attribute data at adjacent moments.
9. The memory fault diagnosis method according to claim 2, wherein Using the diagnostic policy feature sequence to perform a fault diagnosis on the memory to obtain a target diagnosis result includes: Inputting the diagnostic policy feature sequence into a policy prediction model to obtain a target prediction policy. The policy prediction model is trained based on the reward function values of multiple prediction policies of a fault prediction model corresponding to a historical diagnostic policy feature sequence. The fault prediction model is used to perform an initial fault diagnosis on the memory using the operation feature sequence to obtain an initial diagnosis result; Using the fault prediction model to perform a fault diagnosis on the memory according to the target prediction policy to obtain the target diagnosis result.
10. The memory fault diagnosis method according to claim 9, wherein, The prediction policy includes one of the following: Using the fault prediction model to process the target attention feature sequence to obtain the target diagnosis result; Determining a target operation feature subsequence corresponding to preset memory operation information from the operation feature sequence; Updating the risk threshold of the fault prediction model based on the target operation feature subsequence, and using the updated fault prediction model to perform a fault diagnosis on the memory. The preset memory operation information includes the electrical environment information of the operating system and the operating system access information. The risk threshold is used to compare with the output value of the fault prediction model to determine the diagnosis result of the memory according to the comparison result; Determining target memory operation information for fault diagnosis from the memory operation information within the preset time period according to a preset duration; Adopting a repair method of scheduling the storage space of the memory; Temporarily not performing any operation.
11. The memory fault diagnosis method according to claim 9, characterized in that The memory operation information includes the historical fault diagnosis information of the memory. The historical fault diagnosis information includes at least one of the following: the historical repair result corresponding to the historical target prediction policy, the historical risk threshold.
12. The memory fault diagnosis method according to claim 9, wherein, The policy prediction model is trained based on the following method: Using the fault prediction model to process the historical diagnostic policy feature according to multiple prediction policies to obtain the reward function values of the multiple prediction policies; Determining a target reward value according to the multiple reward function values; Updating the parameters of the pre-trained policy prediction model based on the target reward value until the target reward value reaches a preset reward value to obtain the policy prediction model.
13. A memory fault diagnosis system is deployed on a server, characterized in that, Including: An acquisition device for acquiring memory operation information within a preset time period; A memory fault diagnosis device, which when executed, implements the steps of the memory fault diagnosis method according to any one of claims 1 to 12.
14. The memory fault diagnosis system according to claim 13, characterized in that, The server includes a baseboard management controller, a memory, and an operating system. The baseboard management controller is connected to a memory controller to monitor the electrical environment information of the memory; The memory is connected to the operating system, and the operating system is used to schedule the storage space of the memory.
15. The memory fault diagnosis system according to claim 14, wherein The acquisition device is used for: Acquiring the electrical environment information of the memory from the baseboard management controller; Collect operating system access information, operating load information of the operating system, and memory fault events recorded in the log of the operating system; Collect historical fault diagnosis information of the memory from the memory fault diagnosis device, where the memory operation information includes at least one of the following: the historical fault diagnosis information, the electrical environment information, the operating system access information, the operating load information, and the memory fault events.
16. The memory fault diagnosis system according to claim 14, wherein The memory is connected to the service processor of the server via a memory channel.
17. The memory fault diagnosis system according to claim 14, characterized in that, The collection device includes a plurality of collection probes, and the plurality of collection probes are respectively deployed on the baseboard management controller, the operating system, and the memory fault diagnosis device.
18. An electronic device, comprising: One or more device processors; A memory for storing one or more computer programs, wherein the one or more device processors execute the one or more computer programs to implement the steps of the memory fault diagnosis method according to any one of claims 1 to 12.
19. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instruction is executed by the device processor, it implements the steps of the memory fault diagnosis method according to any one of claims 1 to 12.
20. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instruction is executed by the device processor, it implements the steps of the memory fault diagnosis method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Early warning method and device for memory fault
CN113297046A
Power grid fault diagnosis method and device, electronic equipment and storage medium
CN117554747A
Battery fault diagnosis method and device, computer equipment and readable storage medium
CN118795349A
Fault prediction method and device, storage medium and electronic equipment
CN119201617A
Fault analysis method for automatic testing equipment of vehicle machine
CN119537079A