A memory data recovery method in a high temperature environment

By using a cloud-based fault location and capture network and dynamic resource scheduling, accurate location and efficient recovery of memory data under high-temperature conditions are achieved. This solves the problems of inaccurate fault location, unreasonable resource allocation, and transmission interference in existing technologies, and improves the success rate and efficiency of data recovery.

CN121116711BActive Publication Date: 2026-05-05CHENGDU YUWEIGUI TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHENGDU YUWEIGUI TECHNOLOGY CO LTD
Filing Date
2025-09-04
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

In high-temperature environments, existing technologies struggle to accurately locate faulty memory data segments, resulting in a lack of targeted recovery efforts, unreasonable resource allocation, and a lack of real-time error correction mechanisms that make data transmission susceptible to interference and data loss.

Method used

By using a cloud-generated fault location and capture network, faulty memory data segments can be accurately identified, computing resources can be dynamically allocated, a dedicated data recovery transmission channel can be established, and interference can be detected and corrected in real time during transmission to optimize resource scheduling and improve recovery efficiency.

Benefits of technology

It achieves accurate location and efficient recovery of memory data under high temperature conditions, ensuring data integrity, improving recovery success rate and efficiency, and solving the problems of resource waste and interference in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121116711B_ABST
    Figure CN121116711B_ABST
Patent Text Reader

Abstract

The application discloses a memory data recovery method in a high-temperature environment and relates to the technical field of data recovery.The method comprises the following steps: capturing a fault positioning capture network, detecting and identifying a fault data segment in the memory in a high-temperature environment, and generating a data recovery transmission channel and a recovery computer mechanism between the memory address space; at each node of the data recovery transmission channel, an interference correction mechanism is used to detect and correct transmission interference in real time, and the recovery computer mechanism is used to recover the fault data segment being transmitted in real time; when the task state of the data recovery transmission channel, the node or the recovery computer mechanism is idle, the computing resources occupied thereby are released and returned to the cloud, and the computing resources are intelligently scheduled. The method realizes accurate positioning of faults, dynamic allocation and scheduling of computing resources through the cloud, real-time detection and correction of transmission interference, and synchronous data recovery, thereby optimizing the overall recovery efficiency and ensuring the complete and accurate recovery of data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data recovery technology, specifically a method for recovering memory data in a high-temperature environment. Background Technology

[0002] In high-temperature environments, memory data is prone to failure, leading to data loss or corruption. With the continuous increase in data volume and the increasing requirements for data reliability, how to quickly and accurately recover memory data in harsh environments such as high temperatures has become an important issue.

[0003] Existing technologies struggle to accurately locate faulty data segments in memory under high-temperature environments, failing to precisely determine the location, range, and total number of faults in the memory address space. This results in a lack of focus and inefficiency in recovery efforts. Furthermore, the data recovery process suffers from inefficient allocation of computing resources, failing to dynamically allocate and schedule resources based on the actual needs of the faulty data segments. This leads to slow recovery of some faulty data segments due to insufficient resources, while other resources remain idle and wasted. Moreover, high-temperature environments are often accompanied by strong magnetic fields, making data transmission susceptible to interference. Existing technologies lack effective real-time detection and correction mechanisms, failing to address interference issues during transmission in a timely manner, leading to errors or data loss during recovery.

[0004] Therefore, a method for recovering memory data in high-temperature environments is urgently needed to address the above problems. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a method for recovering memory data in high-temperature environments. This method solves the problems of inaccurate memory data fault location, low recovery efficiency due to unreasonable resource allocation, and data transmission errors or loss due to the lack of a real-time error correction mechanism in high-temperature environments.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for recovering memory data in a high-temperature environment, comprising the following steps: detecting and identifying faulty data segments in memory under high-temperature conditions through a fault location and capture network generated by the cloud, and determining the fault information of each faulty data segment; based on the fault information, generating a data recovery transmission channel between the cloud and the memory address space by allocating computing resources, and generating a recovery computing mechanism uniquely bound to each faulty data segment; transmitting each faulty data segment in the data recovery transmission channel, and at each node of the data recovery transmission channel, using an interference correction mechanism to detect and correct transmission interference caused by the high-temperature and strong magnetic environment in real time, while simultaneously performing real-time data recovery of the faulty data segment being transmitted through the recovery computing mechanism; when the task status of the data recovery transmission channel, its nodes, or the recovery computing mechanism is detected to be idle, releasing the computing resources occupied by them and transmitting them back to the cloud, prioritizing the scheduling of the computing resources transmitted back to the cloud to faulty data segments whose recovery is paused due to insufficient computing resources, or scheduling them to the recovery computing mechanism bound to the faulty data segment with the slowest recovery progress, thereby optimizing the overall recovery efficiency.

[0007] Furthermore, the fault location and capture network is constructed by dividing the memory address space in a high-temperature environment into regular grids according to a preset grid size, and arranging detection nodes at the intersection of each grid. The detection nodes integrate an access latency monitoring unit, an error frequency measurement unit, and a temperature and magnetic field sensing unit, which are used to collect access latency anomaly rate, unit soft error occurrence rate, and local temperature and magnetic field signals in real time to form a comprehensive response value. The network determines whether a grid has a fault based on the comparison between the comprehensive response value and a threshold. Faulty grids that are continuous and whose comprehensive response value difference is lower than a preset difference threshold are merged into the same fault data segment, and the fault data segment number and fault information are output. The fault information includes location, range, and total number.

[0008] Furthermore, the detection node is also used to classify and identify fault modes unique to high-temperature environments, extract fault type identifiers from the identified fault modes, set differentiated error correction and recovery strategies for fault data segments based on the fault type identifiers, and transmit the set error correction and recovery strategies and fault grid coordinates back to the cloud.

[0009] Furthermore, the detection node feeds back the location of the faulty data segment in the memory address space to the cloud. The cloud then generates a data recovery channel for each faulty data segment in its memory address space. Each data recovery transmission channel consists of multiple nodes connected in series. Each node includes an interference correction mechanism and a cache unit. Each node uses the interference correction mechanism to detect and correct transmission interference caused by high temperature and strong magnetism in real time. When a data recovery transmission channel is detected to have completed transmission or a node has no subsequent faulty data segment transmission, the data recovery transmission channel or node releases the computing resources it occupies and reclaims them to the cloud.

[0010] Furthermore, the interference correction mechanism is deployed on each node of the data recovery transmission channel, including an interference detection module, an interference correction module, and a state management module. Specifically, the analysis includes: the interference detection module monitors the checksum, bit flip rate, node temperature, and magnetic field signals of the transmitted faulty data segment based on a pre-configured high-temperature strong magnetic noise model to detect the specific interference type; the interference correction module calls a preset error correction algorithm to process the detected interference type, including rereading, reconstruction, or correction based on forward error correction codes; and the state management module releases the computing resources occupied by the node and sends them back to the cloud for reallocation when the node completes the transmission of the current data segment and has no subsequent tasks.

[0011] Furthermore, the number of recovery computing mechanisms is consistent with the number of faulty data segments. The cloud identifies the resource requirement ratio of each faulty data segment based on its size and transmission distance, and divides the remaining computing resources in the cloud into multiple groups according to their resource requirement ratio, forming multiple independent recovery computing mechanisms. Each recovery computing mechanism has a unique matching tag and a transmission carrier. The unique matching tag includes the faulty data segment number, location, and resource requirement ratio, used for binding and identification with the recovery computing mechanism during transmission and recovery. The transmission carrier carries computing resources to form the recovery computing mechanism. It carries computing resources to the location of the faulty data segment responsible for matching and binding in the data recovery transmission channel, uses the unique matching tag to match the faulty data segment, and binds the successfully matched recovery computing mechanism to that faulty data segment. Each recovery computing mechanism includes a recovery process scheduling unit and a data reconstruction unit. The recovery process scheduling unit is responsible for receiving and managing the transmission tasks of the bound faulty data segments, and the data reconstruction unit executes recovery operations according to the recovery strategy set by the fault type identifier.

[0012] Furthermore, the continuous monitoring and recovery process specifically monitors the task status of each data recovery transmission channel, its nodes, and each recovery computing mechanism. When it is detected that a data recovery transmission channel has no faulty data segment transmission recovery requirement within a predetermined time window, the computing resources it occupies are immediately released and transmitted back to the cloud. When the faulty data segment bound to the recovery computing mechanism reaches the cloud through the data recovery transmission channel and completes recovery, the remaining computing resources carried by the transmission carrier of the recovery computing mechanism are immediately released, and the transmission carrier is destroyed in the cloud. The cloud prioritizes scheduling the transmitted computing resources to the cached faulty data. According to the segment, when a cached faulty data segment is identified in the data recovery transmission channel, the returned computing resources are transmitted to the location of the cached faulty data segment in the data recovery transmission channel, and combined with the transmission carrier bound to the cached faulty data segment to update the recovery computing mechanism of the cached faulty data segment; when no cached faulty data segment is found in any data recovery transmission channel, the returned computing resources are scheduled to the location of the data recovery transmission channel of the faulty data segment with the slowest recovery progress, and the computing resources are combined with the computing resources existing in the recovery computing mechanism bound to the faulty data segment.

[0013] Furthermore, the cached faulty data segment specifically refers to the situation where, during transmission, the faulty data segment's recovery is paused due to the exhaustion of computing resources in its bound recovery computing mechanism. The faulty data segment is then guided to the nearest node in the current data recovery transmission channel for caching. Simultaneously, the transmission carrier of the recovery computing mechanism is temporarily stored in that node along with the faulty data segment, awaiting cloud-allocated computing resources for combination. The node records the caching time, and the caching status and caching time are transmitted back to the cloud using the faulty data segment's unique matching tag. The computing resources received by the cloud are allocated to the node containing the faulty data segment with the longest caching time, and the faulty data segment is matched and bound using its unique matching tag.

[0014] The present invention has the following beneficial effects:

[0015] This high-temperature environment memory data recovery method utilizes a cloud-generated fault location and capture network to actively detect and accurately identify faulty data segments in memory under high-temperature conditions, clearly defining their location, range, and total number. This avoids the problems of blindly scanning memory and failing to define fault boundaries in traditional recovery methods. Based on fault segment information, the cloud allocates a dedicated data recovery transmission channel and a uniquely bound recovery calculation mechanism to each faulty data segment, rather than using a single channel and general recovery logic. This avoids recovery failures or data distortion caused by the inability of general recovery mechanisms to adapt to some faulty segments, improving the success rate of data recovery and the integrity of the recovered data under high-temperature conditions. To address the transmission interference easily caused by medium-high temperature and strong magnetic fields in high-temperature environments, and the ease with which data is lost or damaged during transmission, a dual protection mechanism is designed to solve key obstacles posed by the environment. In terms of anti-interference, each data recovery transmission channel... The node employs an interference correction mechanism to detect and correct interference during transmission in real time, preventing data corruption caused by high temperatures and strong magnetic fields. This ensures reliable transmission of faulty data segments from memory to the recovery node. In terms of efficiency, the faulty data segments are simultaneously restored in real time during transmission through a recovery computing mechanism, partially overlapping transmission and recovery times to shorten the overall recovery cycle. This is particularly suitable for scenarios where memory data may deteriorate further under high temperatures. By continuously monitoring task status, the node achieves intelligent scheduling of idle computing resources, solving the problems of rigid resource allocation and idle waste in traditional recovery. Resources are prioritized for transmission to faulty segments that are paused due to insufficient resources or have the slowest recovery progress. This ensures that limited cloud computing resources are tilted towards bottleneck links, avoiding the imbalance of some faulty segments having excess resources while others are stalled due to lack of resources, thus improving the overall efficiency of parallel recovery of multiple faulty segments.

[0016] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0017] Figure 1 This is a flowchart of a high-temperature environment memory data recovery method according to the present invention.

[0018] Figure 2 This is a schematic diagram of the logical framework of a high-temperature environment memory data recovery method according to the present invention. Detailed Implementation

[0019] This application embodiment provides a high-temperature environment memory data recovery method, which enables precise fault location in the cloud, dynamic allocation and scheduling of computing resources, and real-time detection and correction of transmission interference to synchronously recover data, thereby optimizing the overall recovery efficiency and ensuring complete and accurate data recovery.

[0020] The overall concept of this application's embodiments is as follows:

[0021] By utilizing a cloud-generated fault location and capture network, faulty data segments in memory are accurately located under high-temperature conditions. Based on the faulty data segment information, computing resources are allocated by the cloud to generate corresponding data recovery transmission channels and recovery computing mechanisms. During data transmission, an interference correction mechanism is used to detect and correct transmission interference in real time, while the recovery computing mechanism is used to recover the faulty data segments in real time. The recovery process is continuously monitored, idle computing resources are dynamically released and returned, and rationally scheduled to where needed to optimize overall recovery efficiency and achieve efficient and accurate recovery of memory data under high-temperature conditions.

[0022] Please see Figure 1 , Figure 2 This invention provides a technical solution: a method for recovering memory data in a high-temperature environment, comprising the following steps: detecting and identifying faulty data segments in memory under high-temperature conditions through a fault location and capture network generated by the cloud, and determining the fault information of each faulty data segment; based on the fault information, generating a data recovery transmission channel between the cloud and the memory address space by allocating computing resources, and generating a recovery computing mechanism uniquely bound to each faulty data segment; transmitting each faulty data segment in the data recovery transmission channel, and at each node of the data recovery transmission channel, using an interference correction mechanism to detect and correct transmission interference caused by the high-temperature and strong magnetic environment in real time, while simultaneously performing real-time data recovery of the faulty data segment being transmitted through the recovery computing mechanism; when the task status of the data recovery transmission channel, its nodes, or the recovery computing mechanism is detected to be idle, releasing the computing resources occupied by them and transmitting them back to the cloud, prioritizing the scheduling of the computing resources transmitted back to the cloud to faulty data segments whose recovery is paused due to insufficient computing resources, or scheduling them to the recovery computing mechanism bound to the faulty data segment with the slowest recovery progress, thereby optimizing the overall recovery efficiency.

[0023] Specifically, the fault location and capture network detects fault locations by dividing the memory address space into a regular grid. The specific analysis of the fault location and capture network includes: pre-setting a grid size based on memory capacity and fault distribution characteristics; dividing the entire memory address range into several equally sized grid cells; the grid size can be determined based on empirical parameters or dynamically adjusted according to detection requirements; and through grid division, partitioned monitoring of large-capacity memory is achieved, controlling the detection overhead of individual grids while ensuring detection coverage.

[0024] At the intersection of each grid, detection nodes are deployed to collect operational status indicators of the memory in that area. This is equivalent to injecting a monitoring network into the memory space under high temperature conditions, enabling comprehensive perception of the health status of each area. The detection nodes are implemented through a combination of hardware and software: at the hardware level, miniature sensors or test circuits are pre-embedded in the memory modules as detection nodes, and these sensors are deployed at the corresponding grid locations after the memory is divided into grids; at the software level, instructions are issued from the cloud to control the field equipment to perform test operations within the address range of each grid, simulating the function of the detection nodes with software logic.

[0025] Each probe node integrates three main functional units: an access latency monitoring unit, an error frequency measurement unit, and a temperature and magnetic field sensing unit. The access latency monitoring unit times the read and write operations of the grid's memory addresses, measures the actual access latency, and compares it with a normal benchmark value to detect abnormal access latency. In high-temperature environments, if a potential fault occurs in a certain area, its access latency may be significantly higher than normal, and this unit will generate a latency anomaly signal. The error frequency measurement unit counts the number of storage error events occurring within the grid. Specifically, it detects the memory controller's error correction log or obtains the bit flip or read / write error count generated in that area within a given time window by actively writing known test data to the storage unit and then reading it back for comparison. The temperature and magnetic field sensing unit uses sensors deployed at the node to measure the local temperature and magnetic field strength at that location in real time.

[0026] The detection node comprehensively analyzes the three monitoring results mentioned above to generate a comprehensive response value to quantify the health status of the grid. The comprehensive response value is determined according to a pre-defined weighted model, such as the weighted sum of factors like the proportion of access latency exceeding normal values, the number of errors per unit time, and the magnitude of local temperature or magnetic field exceeding standard values. For the comprehensive response value of each grid, a fault judgment threshold is set: when the response value exceeds this threshold, the grid is judged to be faulty. This threshold is set through experimental calibration or according to memory device specifications and environmental tolerances. For example, the average comprehensive response value under normal conditions plus a preset safety margin can be selected as the threshold to minimize false alarms and false negatives. Once the comprehensive response value of a grid exceeds the threshold, the detection node marks the grid as a faulty grid and uses the detection node to feed back the location of the faulty data segment in the memory address space to the cloud. The cloud then generates a data recovery channel for each faulty data segment's memory address space.

[0027] To determine the extent of faulty data segments, the fault location and capture network not only identifies individual faulty grids but also merges adjacent faulty grids. Specifically, when adjacent grids are both marked as faulty and their combined response values ​​differ by less than a preset difference threshold, these grids are considered to be affected by the same fault source and are therefore merged into a single continuous faulty data segment. The difference threshold measures the difference in fault severity between adjacent faulty grids and can be set, for example, based on the relative difference rate of their combined response values. If the combined response values ​​of two adjacent faulty grids differ by less than the threshold, it indicates that the two fault conditions are similar and likely belong to the same fault region; conversely, if the difference is significant, they may be independent faults and should be identified separately. This continuous faulty grid merging strategy accurately characterizes the starting position, size range, and total number of faulty data segments in memory. After detection, the fault location and capture network assigns a unique number to each faulty data segment, records its address range, and transmits it back to the cloud via the network, providing a basis for subsequent data recovery.

[0028] In this implementation scheme, the detection node also undertakes the task of classifying and identifying fault modes specific to high-temperature environments. After collecting multi-dimensional data such as latency, error rate, and environmental parameters, the fault mode classification algorithm built into the node analyzes the data patterns to determine the specific type of fault. This classification can be achieved through preset rules or through machine learning models. For example, a decision tree algorithm can be used for classification. If a high frequency of errors is detected, but most are single-bit random flips and the access latency is not significantly abnormal, it can be identified as a soft error type fault and marked as a bit flip fault, such as a bit flip induced by high temperature. If a large area of ​​read and write operations within a certain grid fails or errors are clustered, and the access latency is significantly increased, it may belong to the type of storage data block corruption or row and column line failure, and is marked as data block corruption. When the temperature is normal but the error pattern shows periodic flips, it can be suspected that the data misalignment is caused by an address line fault and is marked as an address line error. Each identified fault mode corresponds to a fault type identifier. After the fault mode identification is completed, the fault type identifier is sent back to the cloud along with the fault grid coordinate information. The cloud then formulates differentiated error correction and recovery strategies for different types of faults, laying the foundation for selecting appropriate algorithms for the recovery calculation mechanism in subsequent steps.

[0029] By designing a fault location and capture network, it is possible to locate faulty areas in memory under high-temperature environments in a timely and precise manner, and to identify the fault type.

[0030] Specifically, after locating the faulty data segment, a data recovery transmission channel is established in the cloud for each faulty data segment. The faulty data in the high-temperature environment is gradually transmitted to the cloud for recovery. This transmission channel is considered a logical data link, which in actual implementation consists of multiple cascaded physical nodes. Several serially connected nodes are deployed from the memory location where the fault occurred to the remote cloud server. Every two adjacent nodes are connected by a communication link, forming an end-to-end transmission path. Each node can be an independent physical hardware device, such as an edge computing node or signal relay station distributed between the field and the cloud, or a virtual software unit allocated by the cloud, such as a container or process running in different network locations, undertaking intermediate processing tasks during data transmission.

[0031] In this implementation scheme, a data recovery channel is generated from the cloud to the memory address space of each faulty data segment. The specific steps for generating the data recovery channel are as follows: After the detection node reports the location of the faulty data segment to the cloud, the cloud dynamically deploys and connects several relay nodes through the communication interface with the memory device, constructing a dedicated data recovery transmission channel between the cloud and the target memory address space. This transmission channel consists of multiple sequentially connected nodes, starting from the memory address of the faulty data segment and extending to the cloud server. Each node is equipped with an interference correction mechanism and a cache unit to detect and correct signal interference caused by high-temperature and strong magnetic environments in real time during data segment transmission, and to temporarily store data fragments to ensure continuous and stable transmission. Through this communication channel, the faulty data segment can be reliably transmitted from the memory environment to the cloud for recovery. When a data recovery channel completes its transmission task or when a node within it has no further faulty data segments to transmit, the corresponding channel or node immediately releases its occupied computing and network resources and returns these resources to the cloud for reallocation to the recovery of other faulty segments.

[0032] Each transmission node is equipped with an interference correction mechanism and a buffer unit. The buffer unit is used to temporarily store received faulty data segments, balance the processing rates between different nodes, and act as temporary data storage when necessary, such as when subsequent nodes are temporarily unable to receive data. The interference correction mechanism includes an interference detection module, an interference correction module, and a state management module, used to combat signal interference caused by high temperature and strong magnetic environments at each step of data transmission. When a faulty data segment arrives at an intermediate node, the node first checks the integrity and accuracy of the received faulty data segment through the interference detection module. The interference detection module, based on a pre-configured high temperature and strong magnetic noise model, compares the check information, bit flip rate, and other indicators in the faulty data segment, and combines the node's local temperature or magnetic field sensing data to determine whether interference has occurred during transmission and the specific type of interference. The bit flip rate is a quantitative indicator used to measure the frequency or probability of such bit flip errors occurring per unit time or per unit of data volume. Bit flip represents a typical error type in memory or data transmission, referring to a storage unit or data bit unexpectedly changing from its original 0 state to 1, or from 1 to 0 due to external interference. The noise model is a pre-established reference model based on the statistical laws of data errors under high temperature and strong magnetic field environments, such as the probability of random bit flips or burst error distribution that may occur per unit amount of data under given temperature and magnetic field strength. The detection module uses this model as a benchmark, comparing the measured data error characteristics with the model's predictions. If the received faulty data segment detects errors through its own checksum and the error bit distribution matches the random noise type described by the model, random noise interference is identified. If the errors exhibit a concentrated distribution of multiple bits, and the node detects instantaneous magnetic field fluctuations exceeding a threshold, it is identified as strong magnetic pulse interference. If the overall checksum of the faulty data segment repeatedly mismatches, it indicates that the interference may damage the entire data segment. By combining monitoring indicators with the noise model, the interference detection module can distinguish different modes of interference, such as random single-bit flips, continuous multi-bit errors, and signal interruptions.

[0033] When an interference type is identified, the interference correction module of the interference error correction mechanism will perform corresponding error correction processing. For detected transient random noise interference, the interference correction module can request the previous node to retransmit the faulty data segment to utilize the randomness of noise at different times to obtain correct data. If the interference causes some bits to remain erroneous and cannot be corrected after rereading, the interference correction module will initiate a reconstruction process, combining existing redundant information or context data to reconstruct the damaged content, such as using the parity bits or parity information attached to the data to infer the original data. In addition, for data with forward error correction codes attached, the interference correction module will call a preset algorithm to directly correct the detected errors, using redundant parity bits to restore the erroneous bits. After error correction is completed, the faulty data segment in the cache unit will be updated to the corrected, error-free version and continue to be transmitted to the next node. If there are still unrecoverable residual errors after correction, the node will record this situation and mark it for transmission along with the faulty data segment, so that the cloud can take further measures in the final stage.

[0034] During data transmission, each node's state management module monitors the node's task status and releases and reclaims resources at appropriate times. When a node has successfully transmitted all faulty data segments of the current faulty data segment and there is no new faulty data to process within the predetermined time window, the node determines that its task is complete. It then releases its occupied computing resources through the state management module and reports to the cloud for reallocation. Once all data on the entire transmission channel has been transmitted, each node in the channel sequentially enters an idle state and releases its resources, thus preventing nodes from occupying resources for extended periods. The release of node resources does not affect the faulty data segments that have already been transmitted.

[0035] The design of the data recovery transmission channel and nodes ensures the reliability of faulty data segment transmission in harsh environments. Real-time interference detection and correction at each node ensures the integrity of transmitted data. In addition, when nodes are idle, resources are immediately released and reclaimed, improving overall resource utilization efficiency and accelerating the data recovery process.

[0036] Specifically, while the faulty data segments are being transmitted through the transmission channel, the cloud initiates a recovery computing mechanism corresponding to each faulty data segment to perform real-time recovery processing on the received faulty data. Since on-site computing resources are limited in high-temperature environments, the recovery calculation work relies on cloud computing resources. Based on information such as the size and transmission distance of each faulty data segment returned by the previous fault location and capture network, the cloud assesses the resource requirement ratio of each faulty data segment. Specifically, the larger the data volume of a faulty data segment and the farther it is from the cloud, the higher the transmission latency or the more interference along the way, thus requiring more computing resources to complete recovery. A weight value is calculated for each faulty data segment to represent its relative resource requirement ratio. For example, considering both the number of bytes in the faulty segment and its transmission path length, these two factors are weighted proportionally. The cloud then allocates currently available computing resources according to the weight ratio of each faulty segment, forming multiple independent recovery computing mechanisms. For instance, if two faulty data segments are detected, one approximately twice the size of the other, and both have similar transmission conditions to the cloud, the cloud will allocate computing resources in a ratio of approximately 2:1, establishing a larger recovery computing mechanism for the former and a smaller one for the latter. Each recovery computing mechanism is a set of isolated computing resource units allocated by the cloud to ensure that the recovery processes of each faulty data segment do not interfere with each other and can be carried out in parallel.

[0037] A unique matching tag and transmission carrier are generated for each recovery computing mechanism. This unique matching tag contains the faulty data segment's number, location, and allocated resource requirements. This tag is attached to the corresponding faulty data segment and recovery computing mechanism to match and bind the data segment with the correct recovery mechanism during transmission and recovery. The transmission carrier carries computing resources to form the recovery computing mechanism. It travels to the location of the faulty data segment responsible for matching and binding within the data recovery transmission channel, and uses the unique matching tag to match the faulty data segment. The successfully matched recovery computing mechanism is then bound to that faulty data segment. When a faulty data segment arrives at the data recovery transmission channel, its unique matching tag is read, and it is automatically assigned to a recovery computing mechanism with the same unique matching tag for processing, ensuring that each faulty data segment is exclusively handled by the corresponding recovery computing mechanism.

[0038] Within each recovery computing mechanism, there are two core components: a recovery process scheduling unit and a data reconstruction unit. The recovery process scheduling unit is responsible for managing the recovery task flow of the faulty data segments bound to the mechanism, including receiving faulty data segments arriving at the data recovery transmission channel, scheduling the data reconstruction unit to perform recovery operations, and monitoring the recovery progress and resource usage. During the recovery process, the recovery process scheduling unit also monitors the mechanism's resource load and recovery speed in real time. If data backlog or processing speed significantly lags behind the receiving speed is detected, it indicates that current computing resources may be insufficient. This status is then notified to the cloud to trigger subsequent resource scheduling or data caching measures. If the data of the faulty data segment is processed ahead of schedule, the recovery process scheduling unit will mark the recovery computing mechanism as idle, preparing to release computing resources.

[0039] The data reconstruction unit is a functional module that executes specific recovery algorithms. This unit executes a predetermined recovery strategy for the bound faulty data segment, reconstructing the damaged data into the most accurate possible original content. The recovery strategy is determined during the fault location phase for different fault modes. For example, for bit-flip faulty data segments, the data reconstruction unit uses error checking and correction to correct randomly occurring bit-flip errors, or determines the correct bit value through multiple reads and a majority voting mechanism. For data block corruption faulty data segments, if some storage units are completely damaged and unreadable, the data reconstruction unit will use redundant data or verification information to reconstruct the missing content, such as using pre-stored backup blocks or parity data to deduce the original value. For address line error faulty data segments, the data reconstruction unit reassembles the data according to the specific fault mode, such as detecting and correcting data misalignment caused by constant address bit errors, rearranging the misplaced bytes back to the correct order. During the data reconstruction process, the data reconstruction unit continuously writes the recovered correct data to the target storage area in the cloud, gradually piecing together complete recovered data. Once all fragments of the faulty data segment have been reconstructed and verified to be error-free, the recovery calculation mechanism completes its task.

[0040] In this implementation scheme, the recovery computing mechanism is essentially a process module that performs error correction and data reconstruction on faulty data segments under the influence of computing resources allocated in the cloud. The data recovery process can be described as follows: When a faulty data segment is extracted through the transmission channel and bound to its corresponding recovery computing mechanism, the data reconstruction unit within the mechanism performs repair operations on the data segment according to a pre-set recovery strategy. For example, for soft bit flip errors occurring under high-temperature environments, the recovery computing mechanism can repeatedly read the memory unit to obtain stable data bits, or apply forward error correction codes to detect and correct the erroneous bits. Similarly, for multi-bit errors caused by strong magnetic interference, the data reconstruction unit can use redundancy check information or parity bits to reconstruct the data segment and recover lost or tampered bits. Throughout the process, the recovery process scheduling unit is responsible for coordinating data acquisition and error correction operations, ensuring that an appropriate recovery algorithm is selected according to the fault type and that the repaired data segment is uploaded back to the cloud in a timely manner. Through the above methods, the recovery computing mechanism can gradually correct and reconstruct faulty data segments while transmitting data, until the original data content is fully restored.

[0041] By designing a recovery computing mechanism, a one-to-one computing unit can be provided for each faulty data segment. Appropriate computing resources are allocated according to the degree and type of fault for parallel recovery, ensuring that the recovery work of each faulty data segment does not interfere with each other, resulting in higher overall resource utilization. This avoids situations where some tasks are blocked due to insufficient resources while other resources are idle, and significantly improves the efficiency of memory data recovery in high-temperature environments.

[0042] Specifically, throughout the recovery process, the cloud continuously monitors the operational status of each transmission channel, node, and recovery computing mechanism to dynamically schedule and release computing resources. A predetermined time window is set as the idle determination condition: when it is detected that a data recovery transmission channel has no faulty data segment transmission recovery needs within the predetermined time window, the computing resources it occupies are immediately released and transmitted back to the cloud. When the faulty data segment bound to the recovery computing mechanism arrives at the cloud through the data recovery transmission channel and completes recovery, the remaining computing resources carried by the transmission carrier of the recovery computing mechanism are immediately released, and the transmission carrier is destroyed in the cloud.

[0043] The cloud performs unified scheduling and allocation of reclaimed resources. To utilize these resources most effectively, the scheduling strategy adopts a priority order of pausing first and then slowing down: resources are prioritized for scheduling back-transmitted computing resources to cached faulty data segments. When a cached faulty data segment is identified in the data recovery transmission channel, the back-transmitted computing resources are transmitted to the data recovery transmission channel location of the cached faulty data segment, combined with the transmission carrier bound to the cached faulty data segment, and the recovery computing mechanism of the cached faulty data segment is updated. When no cached faulty data segments are found in any data recovery transmission channel, the back-transmitted computing resources are scheduled to the data recovery transmission channel location of the faulty data segment with the slowest recovery progress, and the computing resources are combined with the computing resources existing in the recovery computing mechanism bound to the faulty data segment.

[0044] In this implementation scheme, the recovery computing mechanism is essentially an independent computing unit dynamically generated by the cloud for each faulty data segment. It carries a certain amount of computing resources and is bound to a specific faulty data segment to perform the recovery task for that data segment. The process of generating and binding a new recovery computing mechanism is as follows: The cloud packages the corresponding computing resources into a transmission carrier based on the location of the faulty data segment and the required resource ratio, and assigns it a unique matching tag. This transmission carrier is sent to the node where the faulty data segment is located via the data recovery channel, and automatically matches and binds to the faulty data segment through the unique tag, thereby establishing a recovery computing mechanism for that data segment on-site. Once the mechanism is established, it starts running its recovery process scheduling unit and data reconstruction unit to perform error correction and reconstruction tasks on the bound data segment. When the recovery computing mechanism resources are exhausted during the recovery process, resulting in the faulty data segment being temporarily cached, the cloud will continuously monitor the status of each mechanism, and immediately prioritize the allocation of the recovered resources to the node where the temporarily cached faulty data segment is located after other channels or nodes release idle computing resources. Specifically, the cloud locates the corresponding node based on the unique matching tag of the faulty data segment. Newly allocated computing resources are then delivered via a transmission carrier and combined with the waiting recovery computing mechanism on that node to form an updated recovery computing mechanism. This updated mechanism has more computing resources than the previous one: its binding relationships remain unchanged, but its internal resource configuration is expanded, and its processing power is correspondingly improved. This means that the recovery computing mechanism can utilize more parallel computing or more complex error correction algorithms to accelerate the recovery of the remaining data. Through this dynamic update, the system ensures that even if an initial resource shortage causes a pause, the faulty data segment can be recovered as quickly as possible after resources become available, thus optimizing overall data recovery efficiency.

[0045] When the cloud receives a resource reclamation notification, it first checks if any faulty data segments are in a recovery cache state due to exhaustion of computing resources. The cache state occurs because the computing resources allocated to a certain recovery mechanism are insufficient to complete the recovery of the faulty data segment, preventing further processing and causing it to be temporarily redirected to a node in the transmission channel for caching. For each paused faulty data segment, the cloud records its unique matching tag and the time of the pause. If available computing resources are available in the cloud resource pool at this time, the segment with the longest cache time is selected from all paused faulty segments, and the newly reclaimed computing resources are immediately scheduled to the transmission carrier of the recovery mechanism corresponding to that faulty segment. The cloud locates the transmission channel node where the faulty data segment is located using its unique matching tag and notifies that node to release the cache state and continue transmitting the cached faulty data fragment to the recovery mechanism for processing. This priority strategy can promptly wake up recovery tasks that have been suspended due to resource shortages, preventing faulty data segments from remaining on cache nodes for extended periods.

[0046] When there are no faulty data segments cached in the cloud, the resource scheduling strategy will shift to accelerating the slowest recovery task. The real-time progress of all data segments currently being recovered is compared to identify the slowest recovering faulty data segment. Slowest recovery speed is specifically defined as processing the least amount of data per unit time or having the highest proportion of remaining unrecovered data. The slow recovery may be due to a large data volume in the segment, complex error correction calculations, or insufficient previously allocated computing resources. In this case, the cloud will allocate newly reclaimed computing resources to the recovery mechanism bound to the faulty data segment, adding extra computing power. An example of allocating returned computing resources to the recovery mechanism bound to the slowest recovering faulty data segment is that the cloud directly incorporates released idle CPU cores, memory blocks, and other computing resources into this recovery mechanism, enabling it to process more data segments in parallel or execute more complex error correction algorithms. Through this resource fusion, the originally slowest recovery task will be accelerated, thereby narrowing the progress gap with other tasks.

[0047] By employing dynamic scheduling strategies, every idle computing resource can be fully utilized, prioritizing the timely resumption of tasks paused due to insufficient computing resources, balancing the progress of parallel tasks, and avoiding resource idleness or long waiting times for a single task, thereby improving the overall efficiency and integrity of data recovery.

[0048] Specifically, a specific example of this application embodiment is as follows: Taking a 1GB memory module placed in a 150℃, 0.5T strong magnetic environment as an example, multiple data errors occurred during operation. This application embodiment restores the memory module, and the specific process is as follows:

[0049] A cloud-based fault location and capture network performs gridded detection of the memory address space. Assuming a grid size of 64KB, the detection nodes detect three faulty data segments in the memory. Faulty data segment 1 has a large number of random bit flip errors, faulty data segment 2 has a complete memory cell corruption, and faulty data segment 3 has data misalignment caused by address line abnormalities. The detection nodes in the fault location and capture network identify the location range and fault type information of each faulty data segment and feed it back to the cloud. The cloud assesses resource requirements based on the data volume and transmission distance of each faulty data segment and allocates available computing resources proportionally to the three recovery computing mechanisms. Faulty data segment 1 is approximately 16MB, faulty data segment 2 is approximately 8MB, and faulty data segment 3 is approximately 4MB. Therefore, the cloud allocates approximately 50% of resources to segment 1, 30% to segment 2, and 20% to segment 3, and generates a recovery mechanism with a unique matching tag for each segment. Three transmission channels are established to transmit the data of segments 1, 2, and 3 from the field memory to the cloud. During transmission, the deployed nodes perform real-time correction for interference: if random noise appears in the data segment of segment 1... Interference caused a verification error, and the node requested a retransmission for correction. When segment 2 data had multi-bit errors, the node used forward error correction codes to repair it. Although segment 3 data did not have obvious transmission errors, the node marked its address order as disordered. After error correction, the three faulty data segments arrived at the cloud one after another. At the same time, three recovery computing mechanisms in the cloud ran in parallel. Segment 3, due to its smallest data volume and only needing address order rearrangement, completed recovery the fastest, and its occupied computing resources were immediately released back to the cloud resource pool. At this time, segment 2 was paused because the computational overhead was large due to data block corruption, and the original computing resources were exhausted. The unprocessed data was temporarily stored in the nodes of the transmission channel. The computing resources released by segment 3 were immediately allocated to the transmission carrier of the recovery computing mechanism corresponding to segment 2, allowing the recovery work of segment 2 to continue. Segment 2 then completed data reconstruction, and its resource units were released. The cloud allocated the newly released computing resources to the still-running segment 1 recovery computing mechanism. With the support of additional computing power, the remaining data errors of segment 1 were quickly corrected. Finally, the memory data of segments 1, 2, and 3 were all successfully recovered and verified in the cloud.

[0050] In summary, this application has at least the following effects:

[0051] By using a cloud-generated fault location and capture network, faulty data segments in memory under high-temperature environments can be accurately detected and identified, and their location, range, and total number can be determined. This provides an accurate basis for subsequent targeted recovery, improving the accuracy and effectiveness of the recovery process. Based on the faulty data segment information, computing resources are allocated to generate corresponding data recovery transmission channels and recovery computing mechanisms, enabling on-demand resource allocation. During the recovery process, resource status is continuously monitored, and idle resources are dynamically released and returned, prioritizing allocation to faulty data segments with insufficient computing resources or slow recovery progress, thus optimizing overall recovery efficiency and avoiding resource waste. During data transmission, an interference correction mechanism is used to detect and correct transmission interference caused by high-temperature and strong magnetic environments in real time. Simultaneously, the recovery computing mechanism performs real-time data recovery of faulty data segments being transmitted, ensuring the integrity and accuracy of data during transmission and recovery, and reducing data loss and errors.

[0052] Those skilled in the art will understand that embodiments of the present invention can be provided as methods. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0053] This invention is described with reference to a flowchart of a method according to embodiments of the invention. It should be understood that the combination of each step in the flowchart can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, generate instructions for implementing the process. Figure 1 A device for a function specified in one or more processes.

[0054] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 The function specified in one or more processes.

[0055] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 Steps of a specified function in one or more processes.

[0056] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0057] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for recovering memory data in a high-temperature environment, characterized in that, Includes the following steps: A fault location and capture network generated by the cloud is used to detect and identify faulty data segments in memory under high-temperature conditions, and to determine the fault information of each faulty data segment. The fault location and capture network is constructed by dividing the memory address space in the high-temperature environment into regular grids according to a preset grid size, and arranging detection nodes at each grid intersection. The detection nodes integrate access latency monitoring units, error frequency measurement units, and temperature and magnetic field sensing units to collect access latency anomaly rate, unit soft error occurrence rate, and local temperature and magnetic field signals in real time to form a comprehensive response value. The network determines whether the grid is faulty based on the comparison between the comprehensive response value and a threshold. Faulty grids that are continuous and whose comprehensive response value difference is lower than a preset difference threshold are merged into the same faulty data segment, and the faulty data segment number and fault information are output. The fault information includes location, range, and total number. Based on the fault information, the cloud allocates computing resources to generate a data recovery transmission channel between the cloud and the memory address space, and generates a recovery computing mechanism uniquely bound to each faulty data segment. The number of recovery computing mechanisms is consistent with the number of faulty data segments. The cloud identifies the resource requirement ratio of each faulty data segment based on the size of each faulty data segment and the transmission distance from the cloud. The remaining computing resources in the cloud are divided into multiple groups according to the resource requirement ratio to form multiple independent recovery computing mechanisms. Each recovery computing mechanism has a unique matching tag and transmission carrier. Each faulty data segment is transmitted in the data recovery transmission channel. At each node of the data recovery transmission channel, an interference correction mechanism is used to detect and correct transmission interference caused by high temperature and strong magnetic environment in real time. The interference correction mechanism is deployed on each node of the data recovery transmission channel and includes an interference detection module, an interference correction module and a status management module. At the same time, the faulty data segment being transmitted is recovered in real time through the recovery calculation mechanism. When the task status of the data recovery transmission channel, its node, or the recovery computing mechanism is detected to be idle, the computing resources occupied by it are released and transmitted back to the cloud. The computing resources transmitted back to the cloud are prioritized for the faulty data segment whose recovery is suspended due to insufficient computing resources, or are scheduled to the recovery computing mechanism bound to the faulty data segment with the slowest recovery progress, thereby optimizing the overall recovery efficiency.

2. The method for recovering memory data in a high-temperature environment according to claim 1, characterized in that, The detection node is also used to classify and identify fault modes unique to high-temperature environments, extract fault type identifiers from the identified fault modes, set differentiated error correction and recovery strategies for fault data segments based on the fault type identifiers, and transmit the set error correction and recovery strategies and fault grid coordinates back to the cloud.

3. The method for recovering memory data in a high-temperature environment according to claim 1, characterized in that, The detection node feeds back the location of the faulty data segment in the memory address space to the cloud. The cloud then generates a data recovery channel for each faulty data segment in its memory address space. Each data recovery transmission channel consists of multiple nodes connected in series. Each node includes an interference correction mechanism and a cache unit. Each node uses the interference correction mechanism to detect and correct transmission interference caused by high temperature and strong magnetism in real time. When a data recovery transmission channel is detected to have completed transmission or a node has no subsequent faulty data segment transmission, the data recovery transmission channel or node releases the computing resources it occupies and reclaims them to the cloud.

4. The method for recovering memory data in a high-temperature environment according to claim 3, characterized in that, The specific analysis of the interference correction mechanism includes: The interference detection module monitors the checksum, bit flip rate, node temperature, and magnetic field signal of the transmitted faulty data segment according to the pre-configured high temperature and strong magnetic noise model, and detects the specific type of interference. The interference correction module calls a preset error correction algorithm to process the detected interference type, including rereading, reconstruction, or correction based on forward error correction codes; The state management module releases the computing resources occupied by a node and sends them back to the cloud for reallocation when the node completes the transmission of the current data segment and has no subsequent tasks.

5. The method for recovering memory data in a high-temperature environment according to claim 1, characterized in that, The unique matching tag includes the fault data segment number, location, and resource requirement ratio, which is used to bind and identify the fault data segment with the recovery calculation mechanism during transmission and recovery. The transmission carrier is used to carry computing resources to form a recovery computing mechanism. The transmission carrier carries computing resources to the location of the faulty data segment responsible for matching and binding in the data recovery transmission channel. A unique matching tag is used to match the faulty data segment, and the successfully matched recovery computing mechanism is bound to the faulty data segment. Each recovery computing mechanism includes a recovery process scheduling unit and a data reconstruction unit, wherein the recovery process scheduling unit is responsible for receiving and managing the transmission tasks of the bound fault data segments, and the data reconstruction unit performs recovery operations according to the recovery strategy set by the fault type identifier.

6. The method for recovering memory data in a high-temperature environment according to claim 5, characterized in that, The recovery process specifically monitors the task status of each data recovery transmission channel, its nodes, and each recovery computing mechanism. When it is detected that a data recovery transmission channel has no faulty data segment transmission recovery requirement within a predetermined time window, the computing resources it occupies are immediately released and transmitted back to the cloud. When the faulty data segment bound to the recovery computing mechanism reaches the cloud through the data recovery transmission channel and completes the recovery, the remaining computing resources carried by the transmission carrier of the recovery computing mechanism are immediately released, and the transmission carrier is destroyed in the cloud. The cloud prioritizes scheduling the returned computing resources to the cached faulty data segments. When a cached faulty data segment is identified in the data recovery transmission channel, the returned computing resources are transmitted to the location of the cached faulty data segment in the data recovery transmission channel, combined with the transmission carrier bound to the cached faulty data segment, and the recovery computing mechanism of the cached faulty data segment is updated. If there are no cached faulty data segments in any of the data recovery transmission channels, the computing resources to be returned will be scheduled to the data recovery transmission channel where the faulty data segment with the slowest recovery progress is located, and the computing resources will be combined with the computing resources in the recovery computing mechanism bound to that faulty data segment.

7. The method for recovering memory data in a high-temperature environment according to claim 6, characterized in that, The cached faulty data segment specifically refers to the situation where, during transmission, a faulty data segment's recovery is paused due to the exhaustion of computing resources in its bound recovery computing mechanism. The faulty data segment is then guided to the nearest node in the current data recovery transmission channel for caching. Simultaneously, the transmission carrier of the recovery computing mechanism is temporarily stored in that node along with the faulty data segment, awaiting cloud-allocated computing resources for combination. The node records the caching time, and the caching status and caching time are transmitted back to the cloud using the faulty data segment's unique matching tag. The computing resources received by the cloud are allocated to the node containing the faulty data segment with the longest caching time, and the faulty data segment is matched and bound using its unique matching tag.

Citation Information

Patent Citations

  • LPWANs edge cloud cooperative anti-interference method based on fuzzy detection recovery

    CN113938485A

  • Memory fault repairing method and device, equipment, medium and computer program product

    CN120353632A