Method and system for power hardware fault prediction based on multi-modal data fusion

CN122817002APending Publication Date: 2026-09-25XINJIANG KERONG CLOUD DIGITAL TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611309691.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-27
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本发明提供了一种基于多模态数据融合的算力硬件故障预测方法及系统,以解决预测差异化配置内存条的槽位时,难以对低配槽中的故障位置进行预测隔离,致使存储运营成本高的问题

Benefits of technology

[0016]进一步,内存条交换后,第二服务器根据预设的第一更换故障范围和对照范围之间的差异,获取再次交换内存条的时间作为预测更换时间,获取当前时间和预测更换时间之间第一服务器增加的读写数据量,再结合业务相关关系获取当前时间和预测更换时间之间本地增加的读写数据量,作为第二预测增加量;第二服务器将第二预测增加量与第二业务故障关系结合,获取当前时间和预测更换时间之间本地连接的内存条增加的故障范围,并作为第二增加故障范围,再整合参照范围,得到预测更换时间内存条的故障范围作为第二更新故障范围;根据第二更新故障范围生成为高配槽位更换新内存条的提示。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817002A_ABST
    Figure CN122817002A_ABST
Patent Text Reader

Abstract

The present application relates to the field of storage hardware failure, and particularly relates to a computing power hardware failure prediction method and system based on multi-modal data fusion. The computing power hardware failure prediction method based on multi-modal data fusion comprises the following steps: deploying a first server connected with a high-configuration slot and a second server connected with a low-configuration slot; when the first server is running, the read-write address and fault information of the local memory are collected in real time. The first server collects fault data and forms a first log, the second server analyzes and constructs multi-dimensional business relationships, and combines memory swap to indirectly infer the fault range of the low-configuration slot. When the second server cannot locate the fault range, accurate prediction of the fault range and active maintenance are realized; when the slots of the differentially configured memory are isolated, the fault location in the memory is isolated by predicting the fault range, thereby solving the problem of high storage operation cost, and optimizing resource allocation, improving utilization and reducing overall cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of storage hardware failure, and in particular to a method and system for predicting computing hardware failures based on multimodal data fusion. Background Technology

[0002] In large-scale cloud data centers and high-performance computing scenarios, server memory failure is one of the main causes of system downtime or data loss. Server memory is prone to failure under continuous high-load computing scenarios, and memory failure often leads to business interruption, data corruption, and other problems that can cause significant economic losses. Therefore, predicting memory failure is crucial for server operation.

[0003] Currently, some technologies attempt to predict memory failures using machine learning models. For example, Chinese patent application CN121210203A discloses a training method for a memory failure prediction model. This method uses historical memory operation logs from both faulty and non-faulty servers, extracts spatial error data from memory structural units (such as banks, rows, and cols), and trains the prediction model through data augmentation and feature extraction to predict server memory failures.

[0004] However, the effective operation of this type of prediction method relies on servers having the same memory storage structure to ensure that the dimensions, channel meanings, and feature distributions of spatial error data are aligned. For small and medium-sized data centers with limited budgets, the services carried by servers have significant differences in priority and performance. To balance business needs and cost expenditures, a differentiated configuration strategy is usually adopted, resulting in different storage structures between servers. When using the above method for prediction, because the memory structure corresponding to the low-configuration slots cannot be aligned with the high-configuration slots or standard structures in the training data, it is difficult to accurately predict the fault location within the memory modules of the low-configuration slots. Even if most of the storage area of ​​the memory modules in the low-configuration slots is normal and usable, they must be forcibly shut down, which wastes resources and increases hardware replenishment costs. Meanwhile, the high-configuration slots need to bear high-pressure loads such as high-frequency read and write operations and large-size continuous data processing. As more storage locations are isolated based on the prediction results, the speed of fault propagation accelerates, and the high IO load also accelerates memory module wear and shortens their lifespan. The combination of these two factors ultimately leads to a sharp increase in the overall storage operation costs of small and medium-sized data centers. Summary of the Invention

[0005] This invention provides a computing hardware fault prediction method and system based on multimodal data fusion to solve the problem of high storage operation costs caused by the difficulty in predicting and isolating fault locations in low-configuration slots when predicting the slots of memory modules with differentiated configurations.

[0006] To solve the above-mentioned technical problems, this application provides the following technical solution: The method for predicting computing hardware failures based on multimodal data fusion includes the following steps: Deploy a first server connected to a high-configuration slot and a second server connected to a low-configuration slot; When the first server is running, it collects the read and write addresses and fault information of the local memory in real time and stores the first log containing the amount of read and write data, the scope of the fault, and the timestamp. The second server retrieves and analyzes the first log to obtain the first business fault relationship between the amount of read and write data and the scope of the fault, and the first business time relationship between the amount of read and write data and the change of the amount of read and write data over time. At the same time, it collects the local read and write data in real time, and analyzes the business relationship between the local and first server read and write data after aligning the data collection time. The second server combines the preset relationship between the first time and the first business time to accumulate the amount of data read and written by the first server within the first time period. Then, it combines the relationship between the first business failures to obtain the failure range of the first server within the first time period. Finally, it integrates with the first log to obtain the first predicted failure range. Based on the first predicted failure range, it generates a prompt for swapping memory modules in low-configuration slots and high-configuration slots. Before the memory modules are swapped, the second server uses the latest fault range in the first log as a reference range. After the memory modules are swapped, the first server checks the memory modules and overwrites the fault range and check time in its local first log. After the memory modules are swapped again, the second server obtains the earliest fault range from the overwritten first log as a comparison range. Combining this with the cumulative read and write data volume during the two memory module swaps, the second server analyzes and obtains the fault relationships of the second service. The second server combines the cumulative read / write data volume of the first server within a preset second time period with the business-related relationship to obtain the local cumulative read / write data volume. Then, it combines the second business fault relationship to infer the local newly added fault range, and then integrates the reference range to obtain the second predicted fault range. The second predicted fault range is output as the final fault prediction result.

[0007] The basic principle and beneficial effects of this invention are as follows: This invention utilizes the address mapping function of the high-end slots of the first server to collect and store its memory read / write data volume, fault range, and timestamps to form a first log. The second server retrieves this log and analyzes it to obtain the correlation between the read / write data volume and the fault range and time, as well as the correlation between its own read / write data volume and that of the first server. Then, through two memory module swaps, combined with the local cumulative read / write data volume during the swap period, it constructs its own business fault relationship. Finally, based on this relationship and the cumulative read / write data volume predicted by the first server within the time period, it indirectly infers its own fault range, solving the problem of the difficulty in predicting the fault range of low-end slots. At the same time, by swapping memory modules, the second server provides memory modules with fewer faults to the first server, providing a better storage foundation for the first server to complete core operations. Furthermore, the second server performs low-frequency data read / write on memory modules with a clearly defined fault range, allowing the remaining usable storage units in the memory modules to continue to play their value, reducing the impact of faults and lowering the memory module vacancy rate, thus reducing storage costs.

[0008] This invention achieves core business assurance and efficient resource utilization through functional division of labor and computing power optimization. The second server independently undertakes non-core tasks such as data analysis, fault relationship construction, and fault range prediction, without requiring additional computing power from the first server. This effectively reduces the operational burden on the first server, avoids the occupation of core computing resources, and ensures that the first server can focus on the efficient and stable operation of core businesses. At the same time, it reduces the coupling between functions and improves the overall reliability of the invention. The first server obtains a better storage foundation by leveraging the memory modules with fewer failures provided by the second server, while the second server carries low-priority businesses on memory modules with a clearly defined fault range, reducing the impact of failures and lowering the memory module vacancy rate. The two servers complement each other's functions, balancing core business performance and low-cost requirements without increasing investment in high-end hardware, and significantly reducing storage costs.

[0009] The first and second servers construct a precise and efficient positive cycle through data sharing and bidirectional empowerment. The first log generated by the first server provides the second server with rich and accurate data analysis samples, enabling the second server to overcome the limitation of difficulty in predicting fault locations. By constructing multi-dimensional business relationships, it indirectly infers the scope of faults. The prediction results are continuously optimized with data accumulation, which not only improves the pain point of difficulty in predicting the scope of faults in low-configuration slots, but also provides accurate basis for memory module replacement, avoiding the accumulation of faults that lead to server malfunctions. The analysis results of the second server support memory module swapping, making resource allocation more targeted, and ensuring that the first server continuously obtains high-quality memory suitable for high-load demands. This forms a virtuous cycle of data exchange, model optimization, resource optimization, and more accurate data, maximizing the overall memory resource utilization and operating efficiency of this invention.

[0010] Dynamic resource allocation and proactive predictive maintenance further enhance the practical value of the solution. The two servers alternate the use of memory modules, allowing memory modules in different fault states to precisely adapt to the differentiated needs of core and low-priority services. This extends the overall lifespan of the memory modules, avoids resource waste, and better suits the limited budgets of small and medium-sized data centers. The proactive fault prediction of the second server shifts fault handling from reactive response to proactive prevention, ensuring the continuity of low-priority services while providing a stable guarantee for the memory supply of the first server, reducing the risk of service interruption. Furthermore, this collaborative model significantly increases the utilization rate of the second server, fully leveraging its data analysis capabilities. It not only separates computing power from memory management but also strengthens the adaptability of this invention to business expansion, further reducing operating costs.

[0011] In summary, this invention collects fault data and generates a first log through a first server, and analyzes and constructs multi-dimensional business relationships through a second server. It indirectly infers the fault range of low-configuration slots by combining memory module swapping. When the second server cannot locate the fault range, it achieves accurate prediction and proactive maintenance of the fault range. This invention addresses the problem of high storage operation costs caused by isolating fault locations within memory modules by predicting the fault range when using memory modules with differentiated configurations. At the same time, it optimizes resource allocation, improves utilization, and reduces overall costs.

[0012] Furthermore, when the second server retrieves the first log from the first server, it first obtains the growth rate of the first server's read and write data volume from the first business time relationship, and adjusts the frequency of retrieving the first log according to the growth rate of the read and write data volume.

[0013] Low-configuration slots cannot pinpoint the scope of the fault, forcing the second server to passively wait for the first server to upload large amounts of logs, wasting bandwidth and increasing the burden on the first server. This invention allows the second server to increase its retrieval frequency when the first server's read / write volume surges and decrease its frequency when read / write volume is stable. This enables the second server to obtain critical fault data in a timely manner during peak business periods, while reducing unnecessary data transmission during off-peak periods, thus reducing the network and storage pressure on the first server.

[0014] Furthermore, after the memory modules are swapped, the second server obtains the number of correction CEs from the fault information in the first log, calculates the relationship between the number of correction CEs and the amount of read / write data between two adjacent memory module swap times as a data correction relationship, and generates a prompt to replace the high-spec memory module with a new one based on the data correction relationship corresponding to the memory module.

[0015] High-performance memory slots, due to their high load and rapid increase in failures, often require frequent memory module replacements, leading to increased costs. This invention allows a second server to assess the rate of memory module health deterioration based on the number of corrective errors (CE) corrections, proactively identifying the risk of a large number of uncorrectable errors occurring in the short term. For example, when the first server is undertaking large-scale training tasks, memory errors often erupt in a concentrated manner. This method can prompt replacement before the number of errors reaches a threshold, preventing the first server from interrupting core tasks due to sudden failures. Simultaneously, data correction relationships also help the second server more accurately assess the remaining lifespan of the memory modules, shifting the replacement strategy for high-performance memory slots from passive repair to proactive prevention, significantly reducing the risk of core business interruptions and improving the overall reliability of this invention.

[0016] Furthermore, after the memory modules are swapped, the second server, based on the difference between the preset first replacement fault range and the reference range, obtains the time of the next memory module swap as the predicted replacement time. It then obtains the amount of read / write data added by the first server between the current time and the predicted replacement time, and combines this with the business-related relationship to obtain the amount of read / write data added locally between the current time and the predicted replacement time, as the second predicted increase. The second server combines the second predicted increase with the second business fault relationship to obtain the fault range of the locally connected memory modules between the current time and the predicted replacement time, and uses this as the second increased fault range. It then integrates the reference range to obtain the fault range of the memory modules at the predicted replacement time as the second updated fault range. Based on the second updated fault range, a prompt to replace the memory module in the high-configuration slot with a new one is generated.

[0017] Low-spec memory slots cannot directly determine the scope of a fault, and can only isolate the entire memory module, resulting in wasted resources. This invention allows a second server to indirectly predict the fault propagation trend of its own memory modules based on changes in the read / write data volume of the first server, enabling the second server to make replacement decisions before the fault impacts business operations. For example, when the second server is handling low-priority tasks such as log storage or backup, even if a small number of memory modules fail, replacement time can be scheduled in advance through prediction, avoiding data write failures due to sudden failures during peak business periods. This ability to predict in advance allows the second server to plan memory resources more rationally, reducing business risks caused by accumulated faults, while also allowing the first server to obtain healthier storage media when swapping memory modules, improving the overall system resource utilization.

[0018] Furthermore, the second server obtains the fault range of CE correction from the fault information in the first log as the corrected range, obtains the position distribution, proportion, and overlap relationship between the corrected range and all fault ranges in the first log, and generates a prompt to replace the memory stick in the high-configuration slot based on the obtained relationship.

[0019] Low-spec memory slots cannot pinpoint fault locations, relying solely on the first server's detection results, leading to inaccurate assessments of memory module status. This invention allows a second server to analyze fault distribution patterns from multiple dimensions, such as whether faults are concentrated in specific areas, recurring, or exhibiting a spreading trend. For example, when the second server handles cold data storage, repeated error corrections in certain areas may indicate that the memory module is about to enter a rapid degradation phase; this method can identify this risk in advance. Through these distribution characteristics, the second server can more accurately determine whether the memory module is suitable for continued use in low-spec slots or needs to be replaced prematurely, thereby further improving the system's sensitivity to faults and prediction accuracy, and reducing resource waste caused by misjudgments.

[0020] Furthermore, when the first replacement fault range is not greater than the second update fault range with the predicted replacement time, the second server generates a prompt to move the high-configuration memory module to the low-configuration slot and configure a new memory module for the high-configuration slot based on the predicted replacement time.

[0021] High-performance memory slots experience rapid failures due to high load, while low-performance memory slots, lacking address mapping capabilities, cannot effectively utilize faulty memory modules. This invention compares these two failure ranges and can replace high-performance memory modules before they fully deteriorate, ensuring the primary server always uses the least faulty memory resources. For example, when the primary server handles real-time inference tasks, timely memory module replacement can prevent inference delays or service degradation due to sudden errors. Simultaneously, transferring used memory modules to a secondary server to handle lower-priority tasks maximizes the remaining lifespan of the memory modules, reduces hardware procurement costs, and achieves tiered resource utilization and dynamic resource transfer.

[0022] Furthermore, when the second server is running, it obtains the number of faults that increase between two adjacent memory module replacement times as the verification quantity; after the previous memory module replacement, the first server obtains the reference range, the second predicted fault range, and the verification quantity; the difference between the reference range and the second predicted fault range is taken as the suspected range, and the fault information outside the reference range is investigated from the suspected range until the number of fault information is not less than the verification quantity.

[0023] Low-configuration slots cannot pinpoint the fault range, potentially leading to inaccuracies in the second server's assessment of its own fault status. This invention allows the first server to perform targeted troubleshooting based on the suspected fault range after a swap, using the number of verified faults as a termination condition to ensure that the number of identified faults is sufficient to cover actual new faults. For example, when the first server handles high-consistency tasks such as database operations, accurate fault range troubleshooting can prevent data write errors due to missed fault addresses. This verification mechanism makes the fault range obtained by the second server more reliable, providing a more accurate data foundation for subsequent fault prediction and replacement prompts, thus improving the overall reliability of the fault detection in this invention.

[0024] Furthermore, after the first server investigates fault information outside the reference range, it takes the fault range of the investigated fault information as the actual second newly added range and sends the actual second newly added range to the second server. The second server combines the increase in read and write data of the first server between the two most recent memory module replacement times with the business-related relationship to obtain the local increase in read and write data between the two most recent memory module replacement times. Then, it combines this with the second business fault relationship to obtain the predicted newly added fault range of the locally connected memory module between the two most recent memory module replacement times. Finally, it compares this with the received actual second newly added range and adjusts the second business fault relationship based on the comparison difference.

[0025] The second server lacks the ability to directly obtain the fault range, which may lead to biases in its prediction model. This invention corrects the fault relationships of the second service by introducing actual fault data, continuously optimizing these relationships over time. For example, when the second server undertakes continuously running background tasks such as video transcoding, the continuous calibration of the fault relationships allows for increasingly accurate fault prediction, preventing service interruptions due to prediction errors. This closed-loop optimization mechanism continuously improves the predictive capabilities of the second server, making fault management more intelligent and reliable.

[0026] Furthermore, when the first replacement fault range is not greater than the second update fault range of the predicted replacement time, the second server combines the difference between the first replacement fault range and the second update fault range with the first business fault relationship to obtain the amount of read and write data between the predicted replacement time and the time of the next memory module swap. Then, it combines the data correction relationship to obtain the number of correction CEs between the predicted replacement time and the time of the next memory module swap as a reference number, obtains the data correction relationship between the predicted replacement time and the time of the next memory module swap, and generates a prompt to replace the high-configuration slot with a new memory module based on the obtained relationship.

[0027] High-performance memory slots experience rapid failures due to high load, often requiring frequent memory module replacements and increasing costs. This invention allows a second server to predict future error trends based on differences in failure scope and time intervals, proactively identifying the risk of a large number of errors occurring in the short term. For example, when the first server handles large-scale matrix operations, memory errors tend to increase rapidly; this method can prompt replacement before errors erupt, preventing task failure. This cross-time-period predictive capability allows for more accurate planning of high-performance memory slot maintenance cycles, further reducing the probability of core business interruptions and improving stability and cost-effectiveness. Attached Figure Description

[0028] Figure 1 This is a flowchart of the computing hardware fault prediction method based on multimodal data fusion in Example 1. Detailed Implementation

[0029] The following will describe the concept and technical effects of the present invention clearly and completely with reference to embodiments, so as to fully understand the purpose, features and effects of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are all within the scope of protection of the present invention. Example 1 In this embodiment, memory and external storage differ fundamentally in their data characteristics and fault mechanisms. External storage is used for long-term data storage, and its access patterns are relatively stable; faults are mostly related to the aging of the storage medium. Memory, on the other hand, is used for temporary storage of running program data; its access frequency is high, read / write volume varies greatly, and faults are mostly caused by electrical signal interference, chip fatigue, and other factors, directly related to read / write pressure. Therefore, the occurrence and spread of memory faults are significantly correlated with the amount of data read / written. Analyzing the amount of data read / written can effectively reflect the degree of wear and tear on memory cells, while the stored content itself does not directly contribute to fault prediction. Therefore, this embodiment selects the amount of data read / written as the core analysis indicator.

[0030] Error Correction (CE) is a type of data error that server memory can automatically correct through ECC (Error Checking and Correction) mechanisms. It is mostly caused by single-bit flips and is a basic error correction function commonly supported by server memory. Correction of this type of error relies on a clearly defined fault address. Specifically, when reading or writing data, the memory controller detects the data error using ECC checksums and locates the specific memory address where the error occurred (i.e., the fault address). It then recalculates and restores the correct data at that address based on the checksum. The entire process requires no manual intervention, does not interrupt system operation, cause program crashes, or result in data loss, and has no direct impact on the server's current business operations. However, CE is an important signal of memory chip aging and potential failure. Its frequency increases with increased memory read / write pressure and usage time. A continuously increasing number of CEs significantly increases the risk of serious memory errors. Therefore, monitoring the frequency of CEs, the distribution range of fault addresses, and their changing trends are key indicators for assessing the health status of memory modules.

[0031] UCE is a critical error that cannot be repaired at the memory hardware level by the ECC mechanism. It is often caused by multi-bit errors, memory module damage, hardware connection failures, or controller failures.

[0032] like Figure 1 As shown, the computing hardware fault prediction method based on multimodal data fusion includes the following steps: Deploy a first server connected to the high-performance slot and a second server connected to the low-performance slot. In this embodiment, it is assumed that during deployment, the memory module in the high-performance slot is memory module A, and the memory module in the low-performance slot is memory module B.

[0033] When the first server is running, it collects the read / write addresses and fault information of local memory in real time. That is, when the first server is running, it collects the read / write addresses and fault information of data in memory stick A. It stores a first log containing the amount of read / write data, the scope of the fault, and a timestamp (at this time, the first log stores the amount of read / write data and the scope of the fault in memory stick A).

[0034] The second server retrieves and analyzes the first log to obtain the first business fault relationship between the amount of read and write data and the scope of the fault (the relationship between the amount of read and write data of the first server and the fault information caused to memory module A), and the first business time relationship between the amount of read and write data and the time change (the relationship between the amount of read and write data of the first server and the time change). At the same time, it collects the local read and write data in real time, and after aligning with the collection time, it analyzes the business correlation relationship between the local (corresponding to memory module B) and the first server's read and write data (the relationship between the amount of read and write data of the first server and the amount of read and write data of the second server at the same time).

[0035] Specifically, after retrieving the first log, the second server performs multi-dimensional analysis to obtain three types of relationships: first business time relationships, first business failure relationships, and business-related relationships. Simultaneously, the obtained first log is stored in external storage. The multi-dimensional analysis process is as follows: The first service failure relationship reflects the correspondence between the amount of data read and written by the first server and the failure range. In this embodiment, a multilayer perceptron network is used to construct the first service failure relationship. The multilayer perceptron network predicts the failure probability of the available physical address range by fusing explicit input features and implicit features, specifically including the following steps: Using the vector concatenation operator, the available physical address range indicator vector (with dimension ) is concatenated. This is concatenated with the physical read / write data scalar (obtained after maximum value normalization) of the previous sampling period, representing the availability status of each physical interval (1 for available, 0 for disabled), to construct the generated dimension. joint input feature vector The formula for calculating the maximum value normalization method is as follows: , In the formula, This is the dimensionless physical loading factor obtained after normalization. This is a scalar representing the amount of physical read / write data obtained in the previous sampling period; This is a preset boundary value for the maximum physical read / write volume per cycle (set based on the maximum physical bandwidth of the memory or the historical maximum read / write data volume). This method maps the data volume to the [0,1] interval by truncating the upper limit of the calculated ratio to 1.

[0036] At the same time, the vector of historical read / write activity, reflecting the historical degradation of the media (with dimensions of 1), will be used to determine the proportion of such activity. ) and the frequency vector of historical cumulative corrected errors (CE) (dimension: Specifically, this refers to the cumulative frequency of corrected errors occurring in each physical address range of the memory module up to the end of the current sampling period or the start of the prediction period. This is then concatenated to construct the generation dimension. implicit eigenvectors .

[0037] Before splicing, the cumulative corrected error frequency for each address range is preprocessed using physical normalization. A dynamic frequency upper limit is set. (If an impedance failure threshold is set, it is recommended to replace the device directly if the value is exceeded.) Calculate the dimensionless CE reflection coefficient. and using dimensionless Concatenate with the proportion vector.

[0038] The joint input feature vector is projected onto the input layer, while the implicit feature vector is input into the hidden layer. The hidden layer state is modulated by the feature projection matrix of the implicit features to obtain the output state vector of the hidden layer. , In the formula, This is the output state vector of the hidden layer; This is the weight matrix for the joint input feature vector; This is the feature projection matrix for the implicit feature vector; The bias vector of the hidden layer; This is the activation function.

[0039] After the output state vector is mapped by the output layer, it is physically filtered through a masking mechanism, and finally outputs a fault probability vector for the available physical address range: , In the formula, The final output is a failure probability vector of the available physical address range (dimension: ...). ); This is the output layer weight matrix; This is the output layer bias vector; A vector indicating the available physical address range; This represents the Hadamard integrator. Thus, disabled physical address ranges are restricted from participating in the output calculation; only the fault prediction probability of the available physical address range at the end of the sampling period is output. If the fault prediction probability is greater than or equal to a fault prediction probability threshold preset by the maintenance personnel, then the physical address range is defined as the predicted fault range, thereby generating the fault prediction result.

[0040] When training the first business fault relationship model and the second business fault relationship model, the known physical availability of memory and the actual fault occurrence in the first log are used as training labels. The deviation between the predicted value and the true label is calculated using the binary cross-entropy loss function commonly used in this field and preset by maintenance personnel. The backpropagation algorithm is executed through the Adam optimizer to adjust the model weights until the model converges, thereby completing the network training.

[0041] The first business time relationship reflects the trend of the change of read and write data volume over time. In this embodiment, a time series model is used by default. By fitting the read and write data volume at consecutive time points, the rate and fluctuation pattern of the increase of read and write data volume over time are obtained.

[0042] Business correlation is used to describe the degree of correlation between the read and write data volumes of the second server and the first server. In this embodiment, the read and write data volumes of the two servers (the first server and the second server) within the same time interval are aligned, and the similarity of their changing trends is calculated to obtain the business correlation. Specifically, the same historical monitoring time periods of the two servers are divided into... Two equal sampling time windows ( (where the integer is greater than 1) Calculate the cumulative read and write data volume of the two servers within each sampling time window, thereby constructing the time-series vector of the read and write traffic of the first server. Second server read / write traffic time sequence vector Subsequently, a pre-set Pearson correlation analysis algorithm (e.g., directly calling the pearsonr interface built into the SciPy scientific computing library in Python) is used to calculate the similarity coefficient of the trends in read and write data volume changes between the two servers. When the similarity coefficient is greater than a preset confidence threshold, it is determined that the business read and write characteristics of the two servers are highly correlated. Under this confidence premise, a univariate linear regression is performed on the aligned historical read and write data volume of the two servers to establish a proportional mapping formula between the read and write data volume of the two servers: , In the above formula, This indicates the predicted read / write data volume on the second server. This indicates the actual amount of data read and written by the first server. The proportionality coefficients obtained by fitting using the least squares method, For the regression line at The fitting intercept constant on the above. The above proportional mapping formula serves as the business relevance relationship.

[0043] Specifically, the proportional slope coefficient With the fitting intercept constant The calculation method is as follows: obtain the contents Sample set of historical time-series aligned data ,in, This represents the actual amount of local read / write data collected by the second server under normal historical conditions; it is calculated using the following formula based on the least squares principle. and : , , The above proportional mapping formula serves as the business-related relationship.

[0044] Before the memory modules are swapped (before memory module A and memory module B are swapped, the first log still contains data related to memory module A, such as fault data), the second server uses the latest fault range in the first log (i.e. the fault range in memory module A just before the swap) as the reference range, which is the fault range when memory module A first connects with the second server.

[0045] After each memory module is swapped, the local first log is cleared first, and then the first log corresponding to the current memory module is written. That is, the first log corresponding to the current memory module is used to overwrite the first log corresponding to the memory module before the swap.

[0046] After the memory modules are swapped (memory module A and memory module B are swapped), the first server checks the memory modules (at this time, the first server checks the relevant data of memory module B), and overwrites the local first log with the fault scope and detection time (at this time, the first log stores the relevant data of memory module B, but does not contain the relevant data of memory module A. At this time, the first log of memory module A has been stored on the external storage by the second server).

[0047] After the memory modules are swapped again (assuming this swap involves swapping memory module A to another first server and setting memory module B back to the low-configuration slot), the second server retrieves the earliest fault range (the fault range of memory module A immediately after the swap) from the overwritten first log (which now stores fault information of memory module A recorded during the first server's runtime) as a reference range. The difference between the reference range and the comparison range represents the increased fault range of memory module A on the second server during the two memory module swaps. Combining this with the accumulated read / write data volume locally during the two memory module swaps, the second service fault relationship is analyzed.

[0048] When analyzing the second service fault relationship, the reference range (the fault address range of memory module A) is represented as the first fault feature vector. (where the elements corresponding to the fault physical intervals are marked as) The remaining normal intervals are marked as The control range is represented as the second fault feature vector. By performing element-by-element difference elimination calculations on the control range and the reference range, the difference in fault distribution actually increased by memory module A during the two physical memory module swaps on the second server was determined: , In the formula, This refers to the fault range vector added during operation on the second server. The second server synchronously reads the total physical read / write data accumulated by the local controller during the two exchanges and maps it to a dimensionless cumulative physical load factor using the aforementioned maximum normalization method. Specifically, the cumulative physical load factor. The normalization calculation method is the same as the aforementioned formula: , In the formula, This is a scalar representing the cumulative amount of read and write data locally during the two memory module swaps. A preset cumulative maximum read / write limit value to match the total time of two exchanges (e.g., the maximum limit value per cycle). The total number of sampling periods included between the two exchanges The product of, i.e. ).

[0049] Based on the increased fault range vector With the corresponding cumulative physical load factor The second service fault relationship model is updated and trained using binary classification supervision signals.

[0050] The second service failure relationship model uses physical address range indicator vectors. With cumulative physical load factor The input is direct, and the fault prediction probability of the available physical address range at the end of the sampling period is the direct output. The forward propagation calculation formula is as follows: , In the formula, The predicted output is the fault prediction probability vector for each available physical address range at the end of the current sampling period; The input is a vector indicating the available physical address range (where the corresponding element of the available range is...). The element corresponding to the isolated interval that has been disabled by the system is ); The input is a normalized cumulative physical read / write data scalar (dimensionless, with a value range of 1). ); The pre-defined physical interval spatial association weight matrix is ​​used to characterize the fault collaborative damage probability relationship between adjacent physical intervals; This is the weight vector for mapping the read / write load to the fault-induced intensity of each physical interval; It is the bias vector; For activation functions; This represents the Hadamard product operator, used to filter and mask predictions based on available states.

[0051] The second server uses training data obtained from the first log, including the difference in the fault range of the memory module before and after it was replaced with the second server, and the second cumulative read / write data volume collected after it was replaced with the second server.

[0052] In the inference and prediction phase, the current moment is used as the starting moment for prediction. At this point, memory module A is installed on the first server (high-spec slot), and memory module B is installed on the second server (low-spec slot). Maintenance personnel first preset the duration of the first time period. Second time period duration The first and second servers are based on... Obtain the end time of the first time period as the first prediction end time. (Right now Then, the first prediction end time will be... Postponed The second prediction end time is obtained. (Right now ).

[0053] The specific prediction and deduction process is as follows: Step 1: Predict the changes in faults and available space for each memory module during the first time period.

[0054] Specifically, for memory module A deployed on the first server, the second server obtains the predicted start time. Initial available physical space for memory module A. Duration of the first time period. Given the first business time relationship, predict the cumulative read / write data volume of the first server within that time period and normalize it into a physical load factor; simultaneously obtain... The historical cumulative corrected error (CE) frequency is used to construct an implicit feature vector. The physical load factor and the initial available physical space indicator vector are input into the pre-trained first service fault relationship model, which outputs the fault probability of each physical interval on memory module A. A preset probability threshold is used for filtering to obtain the first set of newly predicted fault base addresses for memory module A within the first time period. This first set of newly predicted fault base addresses is merged with the initial fault address range of memory module A to obtain the first predicted fault range. The first predicted fault range is excluded from the total physical space of memory module A, thus obtaining the fault range of memory module A. The available physical space at the start of the moment. If the predicted first fault range of memory module A reaches the preset first replacement fault range, the second server will automatically generate a swap prompt, alerting maintenance personnel to... Always swap the memory module A in the high-spec slot with the memory module B in the low-spec slot.

[0055] For memory module B deployed on the second server, the second server uses the predicted cumulative read / write data volume from the first server, combined with the business correlation between the two servers (regression mapping formula), to calculate the predicted cumulative read / write data volume of the second server (local) within the first time period. This predicted cumulative read / write data volume is normalized and, along with the initial available physical space indicator vector of memory module B, is input into the pre-trained second business fault relationship to obtain the fault probability of memory module B in each available physical address range. A preset probability threshold is used for filtering to obtain the second set of newly predicted fault base addresses generated by memory module B within the first time period. This second set of newly predicted fault base addresses is excluded from the current available physical space of memory module B, thereby determining the fault probability of memory module B in each available physical address range. The initial available physical space when switching to the first server.

[0056] Step 2: Conduct a prediction and simulation of the second phase after the proposed exchange.

[0057] Specifically, assuming the maintenance personnel are... Physical swaps were performed according to the exchange prompts, that is, memory stick A was moved to the second server (low-end slot), and memory stick B was moved to the first server (high-end slot).

[0058] Second time period Within this period, the following fault evolution prediction is made for memory module A that has been moved to the second server: The second server will [perform the second time period]. Input the first business time relationship model to obtain the predicted read and write data volume of the first server in the time period. Then, combine the business relationship between the two servers to calculate and predict the predicted cumulative read and write data volume of the second server in the second time period.

[0059] The second server normalizes the cumulative read / write data volume predicted locally during the second time period, and inputs it, along with the physical space state corresponding to memory module A during the continuous degradation phase, into the second business fault relationship model to calculate the value at the end of the second prediction. The failure probability of each available physical address range on memory module A. The failure probabilities are filtered using a preset probability threshold to identify newly predicted failure addresses for memory module A within the second time period. These newly predicted failure addresses are then compared with the end time of phase one (i.e.,...). At the time of prediction, the existing fault ranges (first predicted fault ranges) of memory module A are combined and superimposed to obtain the corresponding value at the end of the second prediction. The cumulative fault range of memory module A is used as the second predicted fault range; the second server outputs the second predicted fault range as the final fault prediction result to the maintenance terminal.

[0060] When the second server retrieves the first log from the first server, it first obtains the growth rate of the first server's read / write data volume from the first business time relationship, and adjusts the frequency of retrieving the first log proportionally based on the growth rate of read / write data volume. In this embodiment, maintenance personnel set a proportional coefficient as a baseline value based on the relationship between the increase rate of read / write data volume and the expansion rate of the fault scope; at the same time, they set an initial value for the frequency of retrieving the first log based on the expansion rate of the first server's fault scope; when the growth rate accelerates, the initial value of the frequency or the current frequency of retrieving the first log is increased proportionally to the baseline value to ensure timely acquisition of the latest fault data. When the growth rate slows down, the initial value of the frequency or the current frequency of retrieving the first log is decreased proportionally to the baseline value to reduce resource consumption.

[0061] After the memory modules are swapped, the second server retrieves the number of corrective CE (Corrective Error) attempts from the fault information in the first log. Assuming this refers to the fault information of memory module A in the first log during the two swaps, the relationship between the number of corrective CE attempts and the amount of read / write data between two adjacent memory module swaps is used as the data correction relationship. Specifically, in this embodiment, a classification algorithm is used by default to divide the read / write data amount into multiple intervals, and the number of corrective CE attempts within each interval is counted to obtain the correspondence between the read / write data amount intervals and the number of corrective CE attempts. When monitoring the data read volume of the first server, based on this correspondence, the current accumulated data read volume is first matched with each read / write data amount interval to obtain the corresponding number of corrective CE attempts. This number of corrective CE attempts is used as a threshold. If the current accumulated number of corrective CE attempts in the first log is not less than this threshold, a prompt to replace the memory module in the higher-spec slot (i.e., the memory module with a smaller fault range) is generated. If the current accumulated number of corrective CE attempts in the first log is less than this threshold, the current accumulated number of corrective CE attempts and the amount of read / write data in the first log are continuously monitored.

[0062] Furthermore, the year-over-year coefficient can be dynamically adjusted based on the number of CE corrections performed on the memory module. In this embodiment, a default baseline value for the number of CE corrections is set for each read / write data volume range. When the actual cumulative number of CE corrections exceeds this baseline value, it indicates that the health of the memory module has declined, and the year-over-year coefficient increases accordingly, thereby increasing the log retrieval frequency. When the number of CE corrections is lower than the baseline value, the year-over-year coefficient decreases or remains unchanged, thereby reducing or maintaining the retrieval frequency.

[0063] The second server obtains the fault range of CE correction from the fault information in the first log as the corrected range, obtains the position distribution, proportion, and overlap relationship between the corrected range and all fault ranges in the first log, and generates a prompt to replace the memory stick in the high-configuration slot based on the obtained relationship.

[0064] Specifically, the second server extracts all fault addresses corresponding to the corrected CE from the fault information in the first log, and organizes these addresses into continuous or discrete intervals according to their physical location to form the corrected range.

[0065] Regarding location distribution, the second server maps the corrected ranges to the physical address space of the memory modules, observing whether these ranges are concentrated in a certain area, or whether they exhibit a continuous or scattered distribution. For example, if a large number of corrected ranges are concentrated in a certain address range of the memory module, it indicates that the memory cells in that area may have a continuous deterioration trend. If the corrected ranges are scattered across multiple discontinuous ranges, it indicates that the memory module failure exhibits random characteristics, possibly caused by external interference or accidental factors.

[0066] Specifically, after obtaining the corrected range, the second server sorts all ranges by their starting address and checks the address intervals of adjacent ranges sequentially. If the interval is extremely small (less than the corresponding threshold set by the maintenance personnel) or there are no gaps, they are merged into continuous intervals; if the interval is significant (not less than the corresponding threshold set by the maintenance personnel), they are retained as independent intervals. The number of merged continuous intervals and the coverage area of ​​each interval are then counted. If the number of intervals is small and the coverage area is large (compared to the number threshold and coverage area threshold set by the maintenance personnel), it is judged as a continuous distribution; if the number of intervals is large and the distribution is scattered, it is judged as a discrete distribution. Simultaneously, address span is used for verification: concentrated spans support continuous distribution, while scattered spans support discrete distribution. Through this discrimination of distribution characteristics, the second server can determine whether the fault has a localized concentration, thereby assessing the risk of future UCE (Unified Error Correction) of the memory module.

[0067] Regarding the percentage, the second server compares the total length of corrected fault ranges with the total length of all fault ranges in the first log to calculate the proportion of corrected ranges to all fault ranges. If the percentage of corrected ranges is high (compared to the corresponding threshold set by maintenance personnel), it indicates that most faults can be corrected through ECC, and the memory modules are currently in a relatively controllable state. If the percentage of corrected ranges is low, it indicates that there are a large number of uncorrected or uncorrectable faults, and the memory modules may have entered a rapid deterioration phase. In addition, the second server also compares the percentage of corrected ranges with historical data to observe its changing trend. If the percentage continues to decline, it indicates that the health of the memory modules is deteriorating, and replacement needs to be prepared in advance.

[0068] Regarding overlap counts, the second server tracks the overlap between corrected ranges and historical fault ranges. Specifically, the second server checks whether each corrected range overlaps with previously recorded fault ranges, as well as the number of overlaps and their locations. If a corrected error (CE) occurs multiple times within a certain address range, it indicates a decline in the stability of the memory cells in that region, potentially leading to an uncorrectable error. If the overlap count increases over time, it indicates that the fault is spreading or worsening. By analyzing the changes in overlap counts over time, the second server can determine the persistence and spread of the fault, thus more accurately assessing the remaining lifespan of the memory module.

[0069] Through the multimodal and multidimensional comprehensive analysis described in this embodiment, the second server can fully understand the fault characteristics of the memory module, determine its health status and future fault trends, and generate replacement prompts accordingly, ensuring that the high-configuration slot always uses a healthy memory module, while also providing a more accurate data foundation for the second server's own fault prediction.

[0070] In this embodiment, the second server performs a comprehensive quantitative analysis of the corrected error (CE) physical address range recorded in the first log across the following three dimensions: Spatial distribution characteristic analysis includes: first, sorting the corrected physical addresses in ascending order; when the physical distance between adjacent addresses is less than the gap merging threshold preset by maintenance personnel... If the intervals are consecutive, they are merged into the same continuous interval; otherwise, they are retained as independent intervals. This gives the number of merged intervals. Then calculate the spatial distribution continuity index. ( (where the interval length is...). If the number of intervals... The number is less than the preset threshold by maintenance personnel. and Not less than the continuity threshold for maintenance personnel If the distribution is continuous, it indicates that the fault is highly localized; otherwise, it is determined to be discrete.

[0071] The analysis of the percentage and trend of corrected faults includes: the ratio of the physical length of corrected faults to the total fault length within the calculation period. .like Below the safety ratio threshold set by maintenance personnel This indicates an increased probability of uncorrectable errors occurring; regarding historical... A first-order linear fit is performed on the proportion data of each period; specifically, for each sampling period... Calculate the ratio of the corrected fault length to the total fault length. The ratio time series sequence is obtained. Subsequently, a first-order linear regression was performed on the ratio sequence with time as the independent variable and the ratio as the dependent variable. The slope of the fitted line was then calculated as the slope of the deterioration trend. .when (Within multiple consecutive sampling periods, the specific number of periods is set by the maintenance personnel) and if the absolute value increases, it is determined that the health status of the memory module is in an accelerated deterioration trend.

[0072] Spatiotemporal overlap frequency and diffusion analysis includes: counting the cumulative number of corrected errors occurring on each physical page within a preset sliding time window. If a specific physical page Exceeding the preset frequency threshold If the unit is determined to have degenerated into a permanent physical fault, then the fault is determined to have topological diffusion if the range of the physical area where the overlapping error occurs expands in a stepwise manner within adjacent periods (i.e., the fault suddenly expands from being distributed in a single physical row or physical page space to other adjacent physical rows or banks within two adjacent monitoring periods).

[0073] Decision output: When any two or more of the above conditions of strong spatial concentration (continuous distribution), accelerated deterioration trend (decreasing proportion) and permanent fault propagation (overlapping limits) are triggered, the second server will automatically output a slot replacement prompt.

[0074] In practice, the first server is connected to the high-performance memory slot, and the second server is connected to the low-performance memory slot. Initially, the high-performance memory slot uses memory module A, and the low-performance memory slot uses memory module B. During operation, the first server collects the read / write addresses and fault information of memory module A in real time, forming a first log containing the amount of read / write data, the scope of the fault, and a timestamp.

[0075] After retrieving the log, the second server constructs the first business fault relationship between the amount of read and write data, available physical space, and fault scope. It also constructs the first business time relationship of the change of the amount of read and write data over time through a time series model, and obtains the business correlation relationship between itself and the amount of read and write data of the first server through correlation analysis.

[0076] The second server combines the relationship between the first time and the first business time to accumulate the read and write data volume of the first server, then predicts the fault range through the fault relationship of the first business, and integrates it with the first log to obtain the first predicted fault range. When the first predicted fault range reaches the first replacement fault range, a swap prompt is generated, and maintenance personnel swap memory module A with memory module B.

[0077] Before the swap, the second server uses the latest fault range in the first log as a reference range. After the swap, the first server checks memory module B and updates the first log.

[0078] After the second exchange, the second server obtains the earliest fault range from the updated logs as a reference range. Combining its accumulated read and write data volume during the two exchanges with the available physical space at the time of the first exchange, it constructs a second business fault relationship through a regression model to describe the corresponding pattern between its own read and write data volume and fault growth.

[0079] The second server combines the second time, the cumulative read and write data volume of the first server, and business-related relationships to infer its own read and write data volume. Then, it predicts the scope of new local faults through the second business fault relationships. After integrating the reference range, it obtains the second predicted fault range, which is used to generate a replacement prompt for the low-configuration slot.

[0080] The second server also dynamically adjusts the log retrieval frequency based on the time relationship of the first business, and can further optimize the frequency based on the number of CE corrections.

[0081] After the memory modules are swapped, the second server calculates the relationship between the number of CE corrections and the amount of data read and written, forming a data correction relationship to determine whether the high-spec slot needs to be replaced with a new memory module.

[0082] In addition, the second server extracts the fault range of the CE correction as the corrected range, and analyzes its distribution pattern by sorting, merging and interval judgment (see the above content for specific analysis methods), judges the controllability of the fault by the proportion, and judges whether the fault continues or spreads by the number of overlaps.

[0083] Example 2 The only difference between this embodiment and Embodiment 1 is that after the memory modules are swapped (assuming memory module B is installed in the high-end slot and memory module A is installed in the low-end slot after the swap), the second server obtains the time for swapping the memory modules again (the time it takes for the fault range of memory module B to increase to the first replacement fault range while the first server is running) as the predicted replacement time based on the difference between the preset first replacement fault range (the maximum fault range of the memory module connected to the first server) and the comparison range (the fault range when memory module B is swapped to the high-end slot). This process first obtains the increased fault range. The second server solves the predicted traffic volume using the following forward step-iterative algorithm: The initial read / write data volume increment is set to... , with the set step size Gradually accumulate After each accumulation, the current cumulative read / write data volume (i.e., the current read / write volume and the sum of the read / write volume) is calculated. The sum of the available physical addresses obtained each time is used as input and substituted into the first service fault relationship. Following the method shown in Example 1, the predicted fault range is obtained. The union of this predicted fault range and the control range is taken, and it is determined whether it completely covers the first preset replacement fault range. When it is determined that it completely covers the control range, the iteration is terminated, and the current value is... The predicted business volume is determined. Then, the predicted business volume is obtained by combining the first business time relationship. The corresponding predicted replacement time period is combined with the current time to obtain the predicted replacement time.

[0084] The system calculates the increase in read / write data on the first server between the current time and the predicted replacement time. Then, it combines this with business-related data to calculate the increase in read / write data locally between the current time and the predicted replacement time, using this as the second predicted increase. Where possible, the second predicted increase can be obtained through the relationship between the predicted replacement time period and the second business timeframe.

[0085] The second server combines the second predicted increase with the second service fault relationship (the combination method can be referred to in Example 1) to obtain the fault range of the locally connected memory module between the current time and the predicted replacement time (i.e., the fault range of memory module A), and uses it as the second increased fault range. It then integrates this with a reference range (the fault range when memory module A is moved to a lower-spec slot) to obtain the fault range of the memory module at the predicted replacement time, which is used as the second updated fault range. Based on the second updated fault range, a prompt to replace the memory module in the higher-spec slot is generated.

[0086] Specifically, when the first replacement fault range is no greater than the second update fault range predicted for replacement time, the second server generates a prompt based on the predicted replacement time to move the high-spec memory module to the low-spec slot or to configure a new memory module for the high-spec slot. That is, when the second update fault range is no less than the first replacement fault range, a prompt is generated to replace the memory module with one of smaller fault ranges (assuming memory module C in this case) in the high-spec slot at the predicted replacement time, or to move the high-spec memory module to the low-spec slot (i.e., move memory module B to the low-spec slot).

[0087] In practice, after the swap, the high-end slot is occupied by memory module B, and the low-end slot by memory module A. The second server first calculates the difference between the first replacement fault range and the fault range when memory module B is swapped to the high-end slot. Using a forward iterative search method, based on the actual cumulative read and write data volume of the current server, the simulated read and write volume is gradually increased by a preset step size set by the maintenance personnel. Combined with the first business fault relationship in Example 1, the predicted fault range is obtained until the predicted fault range obtained by the output fault probability includes the first replacement fault range. At this point, the iterative training of the first business fault relationship is stopped, and the cumulative incremental simulated read and write volume is used as the predicted business volume. Combined with the first business time relationship, the predicted replacement time period is obtained, and the current time is added to obtain the predicted replacement time.

[0088] Next, take the increase in read / write data of the first server from the current time to the predicted replacement time. Through the business correlation (assuming a mapping coefficient of 0.75), obtain the second predicted increase (if conditions permit, this can be obtained through the relationship between the predicted replacement time period and the second business time). Input this into the second business fault relationship to obtain the second increased fault range of memory stick A. Take the union of this with the reference range when memory stick A is replaced with a low-configuration slot to obtain the second updated fault range.

[0089] When the second update failure range is not less than the first replacement failure range, a prompt is generated indicating that the predicted replacement time is to replace memory stick C in the high-configuration slot and to replace memory stick B in the low-configuration slot.

[0090] Example 3 The only difference between this embodiment and embodiments 1-2 is that, assuming before the first swap, the high-configuration slot contains memory module A and the low-configuration slot contains memory module B; after the first swap, the high-configuration slot contains memory module B and the low-configuration slot contains memory module A; after the second swap, the other high-configuration slot contains memory module A and the low-configuration slot contains memory module B. When the second server is running, it obtains the number of faults added between two adjacent memory module swap times as the verification quantity, specifically the number of faults added during the two swaps of memory module A (i.e., the number of faults added by memory module A in the low-configuration slot as obtained by the second server) as the verification quantity.

[0091] After the previous (first) memory module swap, the first server obtains a reference range (the fault range of memory module A, and the fault range when memory module A is moved to the lower-spec slot), a second predicted fault range (the fault range of memory module A predicted by the second server for the next memory module swap), and a verification count. The difference between the reference range and the second predicted fault range is taken as the suspected range (the fault range predicted by the second server for memory module A in the lower-spec slot). Starting from the suspected range, fault information outside the reference range is checked until the number of fault information is not less than the verification count. The difference range refers to the set of physical address intervals formed by the remaining physical addresses after excluding the physical addresses contained in the physical address set corresponding to the reference range from the physical address set corresponding to the second predicted fault range.

[0092] After the first server checks the fault information outside the reference range, it takes the fault range of the fault information it has checked as the actual second newly added range (the fault range that memory module A has actually added in the low-configuration slot, as checked by the first server), and sends the actual second newly added range to the second server.

[0093] The second server combines the increase in read / write data on the first server between the two most recent memory module replacement times with business-related relationships to obtain the local increase in read / write data between the two most recent memory module replacement times. In other words, after the second server obtains the actual second increase range, it obtains the local cumulative read / write data between the two replacements. Some servers can directly obtain the local cumulative read / write data.

[0094] The second predicted fault range obtained above is compared with the actual second newly added range, and the second business fault relationship is adjusted according to the comparison difference, that is, the second business fault relationship model is trained and optimized.

[0095] In practice, before the initial swap, the high-configuration slot is occupied by memory module A and the low-configuration slot by memory module B. After the initial swap, the high-configuration slot is occupied by memory module B and the low-configuration slot by memory module A. After the second swap, memory module A is moved to another high-configuration slot. The number of new faults added between two adjacent swaps of memory module A in the low-configuration slot on the second server is the checksum.

[0096] The first server uses the fault range when memory module A is replaced with a lower-spec version as a reference range, and investigates based on the difference (suspected range) between the second predicted fault range and the reference range until the number of faults is not less than the number of verifications. After investigation, the actual newly added fault range of memory module A is sent to the second server.

[0097] The second server combines the local read / write data volume between the two exchanges, compares the predicted range of new faults when memory module A is in the low-configuration slot with the actual range of new faults, and adjusts and optimizes the fault relationship model of the second service.

[0098] Example 4 The only difference between this embodiment and embodiments 1-3 is that when the first replacement fault range is not greater than the second update fault range of the predicted replacement time (i.e., the second server anticipates that the memory module in the lower-spec slot can be replaced with a memory module in the higher-spec slot during the next memory module replacement), the second server combines the difference between the first replacement fault range and the second update fault range with the first service fault relationship to obtain the amount of read and write data between the predicted replacement time and the time of the next memory module swap. That is, the amount of read and write data accumulated by the first server during the period after the next replacement when the fault range of the memory module in the higher-spec slot expands to the first replacement fault range. Since the cumulative read / write data volume is positively correlated with the physical fault range and monotonically increases, the second server constructs a test read / write data volume input to a known first business fault relationship model within a feasible read / write value range to obtain the corresponding simulated fault range. The simulated fault range is numerically compared with the physical space difference between the first replacement fault range and the second update fault range to determine the deviation between the two. The deviation is iteratively adjusted to adjust the test read / write data volume and re-input to the model until the simulated fault range matches the physical space difference. At this point, the iteration stops, and the current test read / write data volume is taken as the determined read / write data volume.

[0099] Then, by combining the data correction relationship to obtain the number of correction CE counts between the predicted replacement time and the next memory module replacement time (i.e., the cumulative number of correction CE counts between the next replacement time and the time after that), the data correction relationship between the predicted replacement time and the time after that is obtained. Based on the obtained relationship, a prompt to replace the high-spec memory module is generated. The specific acquisition of the data correction relationship and the generation of the prompt to replace the new memory module are shown in Example 1.

[0100] The above are merely embodiments of the present invention. The invention is not limited to the fields covered by these embodiments. Commonly known structures and characteristics in the solutions are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are able to access all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, under the guidance of this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.

Claims

1. A method for predicting computing hardware faults based on multimodal data fusion, characterized in that, Includes the following steps: Deploy a first server connected to a high-configuration slot and a second server connected to a low-configuration slot; When the first server is running, it collects the read and write addresses and fault information of the local memory in real time and stores the first log containing the amount of read and write data, the scope of the fault and the timestamp. The second server retrieves and analyzes the first log to obtain the first business fault relationship between the amount of read and written data and the scope of the fault, and the first business time relationship between the amount of read and written data and the time change of the amount of read and written data. At the same time, it collects the local read and written data in real time, and analyzes the business relationship between the local and first server read and written data after aligning the collection time. The second server combines the preset relationship between the first time and the first business time to accumulate the amount of data read and written by the first server within the first time period. Then, it combines the relationship between the first business failures to obtain the failure range of the first server within the first time period. Finally, it integrates with the first log to obtain the first predicted failure range. Based on the first predicted failure range, it generates a prompt for swapping memory modules in low-configuration slots and high-configuration slots. Before the memory modules are swapped, the second server uses the latest fault range in the first log as a reference range; after the memory modules are swapped, the first server checks the memory modules and overwrites its local first log with the fault range and the check time. After the memory modules were swapped again, the second server obtained the earliest fault range from the overwritten first log as a reference range. Combined with the local cumulative read and write data volume during the two memory module swaps, the fault relationships of the second service were analyzed. The second server combines the cumulative read / write data volume of the first server within a preset second time period with the business-related relationship to obtain the local cumulative read / write data volume. Then, it combines the second business fault relationship to infer the local newly added fault range, and then integrates the reference range to obtain the second predicted fault range. The second predicted fault range is output as the final fault prediction result.

2. The computing hardware fault prediction method based on multimodal data fusion according to claim 1, characterized in that: When the second server retrieves the first log from the first server, it first obtains the growth rate of the first server's read and write data volume from the first business time relationship, and adjusts the frequency of retrieving the first log according to the growth rate of the read and write data volume.

3. The computing hardware fault prediction method based on multimodal data fusion according to claim 1, characterized in that: After the memory modules are swapped, the second server obtains the number of correction CEs from the fault information in the first log, calculates the relationship between the number of correction CEs and the amount of read / write data between two adjacent memory module swap times as a data correction relationship, and generates a prompt to replace the high-spec memory module with a new one based on the data correction relationship corresponding to the memory module.

4. The computing hardware fault prediction method based on multimodal data fusion according to claim 1, characterized in that: After the memory modules are swapped, the second server obtains the time for swapping the memory modules again as the predicted replacement time based on the difference between the preset first replacement fault range and the comparison range. It also obtains the amount of read / write data added by the first server between the current time and the predicted replacement time, and then obtains the amount of read / write data added locally between the current time and the predicted replacement time by combining the business-related relationships, as the second predicted increase. The second server combines the second predicted increase with the second business fault relationship to obtain the fault range of the locally connected memory module between the current time and the predicted replacement time, and uses it as the second increased fault range. Then, it integrates the reference range to obtain the fault range of the memory module at the predicted replacement time as the second updated fault range. Based on the second updated fault range, it generates a prompt to replace the high-configuration slot with a new memory module.

5. The computing hardware fault prediction method based on multimodal data fusion according to claim 3, characterized in that: The second server obtains the fault range of CE correction from the fault information in the first log as the corrected range, obtains the position distribution, proportion, and overlap relationship between the corrected range and all fault ranges in the first log, and generates a prompt to replace the memory stick in the high-configuration slot based on the obtained relationship.

6. The computing hardware fault prediction method based on multimodal data fusion according to claim 4, characterized in that: When the first replacement fault range is no greater than the second update fault range with the predicted replacement time, the second server generates a prompt to move the high-configuration memory module to the low-configuration slot and configure a new memory module for the high-configuration slot based on the predicted replacement time.

7. The computing hardware fault prediction method based on multimodal data fusion according to claim 1, characterized in that: When the second server is running, the number of faults that increase between two consecutive memory module replacement times is used as the verification count. After the previous memory module swap, the first server obtains the reference range, the second predicted fault range, and the number of checks; the difference between the reference range and the second predicted fault range is taken as the suspected range, and the fault information outside the reference range is checked from the suspected range until the number of fault information is not less than the number of checks.

8. The computing hardware fault prediction method based on multimodal data fusion according to claim 7, characterized in that: After the first server checks the fault information outside the reference range, it takes the fault range of the fault information found as the actual second newly added range and sends the actual second newly added range to the second server. The second server combines the increase in read / write data on the first server between the two most recent memory module replacement times with the business-related relationship to obtain the local increase in read / write data between the two most recent memory module replacement times. Then, it combines this with the second business fault relationship to obtain the predicted range of new faults in the locally connected memory modules between the two most recent memory module replacement times. Finally, it compares this with the received actual second new range and adjusts the second business fault relationship based on the comparison difference.

9. The computing hardware fault prediction method based on multimodal data fusion according to claim 3 or 6, characterized in that: When the first replacement fault range is not greater than the second update fault range of the predicted replacement time, the second server combines the difference between the first replacement fault range and the second update fault range with the first business fault relationship to obtain the read and write data volume between the predicted replacement time and the next memory module swap time. Then, it combines the data correction relationship to obtain the number of correction CEs between the predicted replacement time and the next memory module swap time as a reference number, obtains the data correction relationship between the predicted replacement time and the next memory module swap time, and generates a prompt to replace the high-configuration slot with a new memory module based on the obtained relationship.

10. A computing hardware fault prediction system based on multimodal data fusion, characterized in that, The computing hardware fault prediction method based on multimodal data fusion as described in any one of claims 1-9 was used.

Citation Information

Patent Citations

  • Memory fault prediction model training method, memory fault processing method and electronic equipment

    CN121210203A