Data recovery method, apparatus and computing device

WO2026194625A1PCT designated stage Publication Date: 2026-09-24HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/080689
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-21
Filing Date
2026-02-28
Publication Date
2026-09-24

Smart Images

  • Figure CN2026080689_24092026_PF_FP_ABST
    Figure CN2026080689_24092026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computers. Provided are a data recovery method, an apparatus and a computing device. The computing device has a local memory and a remote memory. The method comprises: determining an associated memory of a first memory area; and, when a memory error occurs in the first memory area, performing data recovery on the first memory area on the basis of first data stored in the associated memory, the first memory area being a local memory for which a first process applies, the associated memory comprising a second memory area, the second memory area and the first memory area both storing the first data, and the second memory area comprising a local memory for which a second process applies or a remote memory for which the first process applies. The technical solution provided in the present application can reduce the effect of data recovery mechanisms on service performance, and does not require additional storage space.
Need to check novelty before this filing date? Find Prior Art

Description

Data recovery methods, devices and computing equipment

[0001] This application claims priority to Chinese Patent Application No. 202510344380.9, filed on March 21, 2025, entitled “Data Recovery Method, Apparatus and Computing Device”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of computer technology, and in particular to a data recovery method, apparatus and computing device. Background Technology

[0003] Memory is a critical component of computing devices such as servers. Memory failures can lead to data loss, system crashes, and other problems. To improve reliability, current computing devices widely employ error checking and correcting (ECC) mechanisms, which detect and correct memory errors by attaching ECC codes during data storage.

[0004] The ECC mechanism can detect and automatically correct single-bit errors, as well as multi-bit errors. Single-bit errors are also known as correctable errors (CE), while multi-bit errors are known as uncorrectable errors (UCE). That is, the ECC mechanism cannot correct UCE errors. Checkpointing is another widely used fault-tolerance technique. This mechanism can compensate for the shortcomings of the ECC mechanism and improve the system's fault tolerance. Checkpointing involves periodically storing the operational state of services to a remote storage system during business operations, forming checkpoints. When a system failure occurs, services can be restored from the nearest checkpoint.

[0005] When the checkpoint mechanism generates a checkpoint, the business process is in a blocked state, and writing the checkpoint takes a long time, which can have a significant impact on business performance; in addition, this mechanism requires additional storage space. Summary of the Invention

[0006] This application provides a data recovery method, apparatus, and computing device that can reduce the impact on business performance and requires no additional storage space.

[0007] Firstly, a data recovery method is provided, which can be applied to a computing device having near-end memory and far-end memory; the method can be executed by the computing device, or by a module applied in the computing device (such as a processor, chip, or chip system, etc.), or by a logical node, logical module, or software that can implement all or part of the functions of the computing device.

[0008] The method includes: determining the associated memory of a first memory region; and, in the event of a memory error in the first memory region, recovering data from the first memory region based on the first data stored in the associated memory. The first memory region is the near-end memory allocated by a first process, the associated memory includes a second memory region, both the second and first memory regions store the first data, and the second memory region includes either the near-end memory allocated by the second process or the far-end memory allocated by the first process.

[0009] In the aforementioned data recovery scheme, for the first memory region requested by the first process, a memory region containing the same data is determined from the near-end memory region requested by other processes or the far-end memory region requested by this process, and this region is designated as the associated memory region of the first memory region. When a memory error occurs in the first memory region, data recovery can be performed on the first memory region based on the data stored in its associated memory region. In this way, while achieving data fault tolerance, on the one hand, there is no need to set up additional storage space for data backup; on the other hand, there is no need to block process write checkpoints, and data can be recovered from the point of business interruption, thus reducing the impact on business performance.

[0010] In one possible implementation of the first aspect, the second memory region is the near-end memory requested by the second process, and the first process and the second process share the first memory region and the second memory region.

[0011] In the above implementation, the first process and the second process share the first memory area and the second memory area. In this way, the first process and the second process can easily obtain relevant data in the first memory area and the second memory area, thereby saving the performance overhead caused by inter-process communication and making it easy to implement.

[0012] In one possible implementation of the first aspect, the associated memory further includes a third memory region storing the first data, the third memory region being near-end memory requested by the third process.

[0013] The above implementation method allows the first memory area to have more associated memory options when a memory error occurs, thereby improving the reliability of data recovery.

[0014] In one possible implementation of the first aspect, the near-end memory of the computing device includes a plurality of memory regions, including a first memory region and a second memory region. Each of the plurality of memory regions has a corresponding data identifier, which is used to identify the data stored in the corresponding memory region. The data identifiers of the first memory region and the second memory region have the same value.

[0015] Through the above implementation method, the first process can conveniently determine the associated memory through data identifiers, thereby improving processing efficiency.

[0016] In one possible implementation of the first aspect, the data identifier is a checksum of the data stored in the corresponding memory area.

[0017] In one possible implementation of the first aspect, the near-end memory of the computing device includes a plurality of memory regions, including a first memory region and a second memory region, each of the plurality of memory regions having a corresponding status flag, the status flag including any one of the following:

[0018] The first value indicates whether the memory area is readable and writable;

[0019] The second value is used to indicate that the memory area is not readable;

[0020] The third value is used to indicate that the memory area is not writable.

[0021] The above implementation method enables processes to identify and recognize the read / write status of memory areas through status flags, thereby improving the reliability of data reading and writing.

[0022] In one possible implementation of the first aspect, after a memory error occurs in the first memory area, the status flag of the first memory area is updated from a first value to a second value, and the status flag of the second memory area is updated from a first value to a third value; after the first memory area completes data recovery, the status flag of the first memory area is updated from a second value to a first value, and the status flag of the second memory area is updated from a third value to a first value.

[0023] By implementing the above methods, the second process will not write new data to the second memory area during the process of the first process recovering data from the second memory area, thereby improving the reliability of data recovery from the first memory area.

[0024] In one possible implementation of the first aspect, the second memory region is the address space in remote memory requested by the first process to store the first data. This allows the first process to more easily retrieve the first data in the event of a subsequent memory error, thereby improving data recovery efficiency.

[0025] In one possible implementation of the first aspect, the first process and the second process are used to perform the matrix factorization task.

[0026] The above implementation methods can improve the computational efficiency of matrix factorization tasks.

[0027] In one possible implementation of the first aspect, in the matrix to be decomposed in the matrix decomposition task, each data block corresponds to a business process; for any business process, the near-end memory requested by the business process includes: a computing memory area, a first communication memory area, and a second communication memory area, wherein the first memory area is any one of the near-end memory areas requested by the first process.

[0028] The computing memory area is used to cache data related to computing tasks; the first communication memory area is used to cache communication data between the business process and at least one corresponding first target process; and the second communication memory area is used to cache communication data between the business process and at least one corresponding second target process. The first target process and the second target process are other business processes besides the business process. The data block corresponding to the first target process is located in the same row as the data block corresponding to the business process, and the data block corresponding to the second target process is located in the same column as the data block corresponding to the business process.

[0029] The associated memory of the computing memory area of ​​the business process includes: the remote memory requested by the business process; the associated memory of the first communication memory area of ​​the business process includes: the first communication memory area of ​​at least one first target process corresponding to the business process; the associated memory of the second communication memory area of ​​the business process includes: the second communication memory area of ​​at least one second target process corresponding to the business process.

[0030] Secondly, a data recovery device is provided, which can be applied to a computing device having near-end memory and far-end memory. The device may include:

[0031] The processing module is used to determine the associated memory of the first memory area; the first memory area is the near memory requested by the first process, and the associated memory includes the second memory area. Both the second memory area and the first memory area store the first data. The second memory area includes the near memory requested by the second process or the far memory requested by the first process.

[0032] The data recovery module is used to recover data from the first memory area based on the first data stored in the associated memory area in the event of a memory error in the first memory area.

[0033] In one possible implementation of the second aspect, the second memory region is the near-end memory requested by the second process, and the first process and the second process share the first memory region and the second memory region.

[0034] In one possible implementation of the second aspect, the associated memory further includes a third memory region that stores the first data, and the third memory region is near-end memory requested by the third process.

[0035] In one possible implementation of the second aspect, the near-end memory of the computing device includes multiple memory regions, including a first memory region and a second memory region. Each of the multiple memory regions has a corresponding data identifier, which is used to identify the data stored in the corresponding memory region. The data identifier values ​​of the first memory region and the second memory region are the same.

[0036] In one possible implementation of the second aspect, the data identifier is a checksum of the data stored in the corresponding memory area.

[0037] In one possible implementation of the second aspect, the near-end memory of the computing device includes a plurality of memory regions, including a first memory region and a second memory region, and each of the plurality of memory regions has a corresponding status identifier, the status identifier including any one of the following:

[0038] The first value indicates whether the memory area is readable and writable;

[0039] The second value is used to indicate that the memory area is not readable;

[0040] The third value is used to indicate that the memory area is not writable.

[0041] In one possible implementation of the second aspect, after a memory error occurs in the first memory area, the status flag of the first memory area is updated from a first value to a second value, and the status flag of the second memory area is updated from a first value to a third value; after the first memory area completes data recovery, the status flag of the first memory area is updated from a second value to a first value, and the status flag of the second memory area is updated from a third value to a first value.

[0042] In one possible implementation of the second aspect, the second memory region is the address space in the remote memory requested by the first process that stores the first data.

[0043] In one possible implementation of the second aspect, the first process and the second process are used to perform the matrix decomposition task.

[0044] Thirdly, a computing device is provided, comprising: a memory and a processor, wherein the memory is used to store a program; and the processor is used to execute the method described in the first aspect or any embodiment thereof when the program is invoked.

[0045] Fourthly, a readable storage medium is provided having a program stored thereon that, when executed by a processor, implements the method described in the first aspect or any embodiment of the first aspect.

[0046] Fifthly, a program product is provided that, when run on a device, causes the device to perform the method described in the first aspect or any embodiment of the first aspect.

[0047] In a sixth aspect, a chip system is provided, including a processor coupled to a memory, the processor executing a program stored in the memory to implement the method described in the first aspect or any embodiment thereof. The chip system may be a single chip or a chip module composed of multiple chips.

[0048] It is understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0049] Figure 1 is a schematic diagram of the service recovery principle of the checkpoint mechanism provided in the embodiment of this application;

[0050] Figure 2 is a schematic diagram of the structure of the computing device provided in an embodiment of this application;

[0051] Figure 3 is a flowchart illustrating the data recovery method provided in an embodiment of this application;

[0052] Figure 4 is a schematic diagram illustrating the principle of determining associated memory according to an embodiment of this application;

[0053] Figure 5 is a schematic diagram illustrating another principle for determining associated memory provided in an embodiment of this application;

[0054] Figure 6 is a schematic diagram of the data recovery process provided in an embodiment of this application;

[0055] Figure 7 is a schematic diagram of the data recovery principle provided in the embodiment of this application;

[0056] Figure 8 is a schematic diagram of the matrix decomposition principle provided in an embodiment of this application;

[0057] Figure 9 is a schematic diagram of the data recovery device provided in an embodiment of this application. Detailed Implementation

[0058] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is only for explaining specific embodiments and is not intended to limit the application. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0059] To facilitate understanding of the embodiments of this application, the two fault-tolerance technologies, ECC mechanism and checkpoint mechanism, will be explained below.

[0060] ECC (Error Correction Control) is a hardware-level fault-tolerance mechanism primarily used to detect and correct data errors (i.e., memory errors) occurring in the memory of a computing device. In this mechanism, when the processor writes data to memory, the memory controller encodes the data to be written using ECC, generating additional checksum information (i.e., an ECC code, hereinafter referred to as the first ECC code), and then writes the data and the first ECC code together into memory. When the processor reads data from memory, the memory controller reads the data and its corresponding first ECC code simultaneously, recalculates the ECC code based on the read data (hereinafter referred to as the second ECC code), and then compares the second ECC code with the first ECC code.

[0061] If the second ECC code matches the first ECC code, it indicates that the data is error-free and the verification is successful. The memory controller can then pass the data to the processor. If the second ECC code does not match the first ECC code, it indicates that a data error has occurred and the verification has failed. The memory controller can determine whether the error is a single-bit error or a multi-bit error based on the first and second ECC codes.

[0062] Among them, a single-bit error is called a CE error. For a single-bit error, the memory controller can locate the bit where the error occurred, automatically correct the error of that bit, and then return the corrected data to the processor.

[0063] Multi-bit errors are called UCE errors. For multi-bit errors, the memory controller cannot locate the bit position where the error occurred, and therefore cannot correct such errors. In this case, the memory controller can report the memory error information to the processor through the interrupt mechanism. The processor then reports the memory error information to the operating system (OS). The operating system can terminate the corresponding business process (i.e., the process executing the business) based on the memory error information.

[0064] If the business process is terminated, the previous data calculations will become invalid, resulting in a waste of computing power.

[0065] Checkpointing is a software-level fault-tolerance mechanism that allows data recovery based on stored checkpoints. It can be combined with ECC (Error Correction Control) to enhance system fault tolerance. For example, in the event of a UCE (Uninterruptible Code Correction) error, the operating system can perform data recovery based on checkpoints.

[0066] The checkpoint mechanism saves the operational state of a business process periodically (i.e., checkpoints) during business operations, allowing recovery from the most recent checkpoint in the event of a system failure. The core of the checkpoint mechanism is the checkpoint itself. To ensure system state consistency when generating a checkpoint, the system first suspends data input for each business process; that is, each business process pauses processing new task data. After each business process has finished processing its current task data, its state is copied to a remote storage system, resulting in a checkpoint. Then, data input for each business process is resumed. When a system failure occurs, such as a UCE error, all or affected business processes can be restarted. The most recently saved checkpoint is then read from the remote storage system, and the state of each business process is reset to the state corresponding to that checkpoint. Other business tasks can then continue execution without having to run the entire business process from scratch.

[0067] When the local storage space of the computing device is sufficient, checkpoints can also be stored in local memory or hard disk storage devices. This example illustrates the use of a remote storage system as an example. The remote storage system can store multiple checkpoints, or only the most recently saved checkpoint. The state saved in the checkpoint can include memory data, register values, thread states, etc., for each business process. The following example uses memory data as an illustration.

[0068] For example, Figure 1 illustrates a schematic diagram of the service recovery principle of a checkpoint mechanism. As shown in Figure 1, R1, R2, and R3 represent three processes of a certain service, which process task data based on their respective allocated memory. The system sequentially generates checkpoint CK1 at time point T1, checkpoint CK2 at time point T4, and checkpoint CK3 at time point T7. Taking checkpoint CK1 as an example, checkpoint CK1 includes the memory data of processes R1, R2, and R3 at time point T1. Assuming that process R1 experiences a UCE error at time point T9, the operating system can read checkpoint CK3 and restore the memory data of processes R1, R2, and R3 to the state at time point T7 based on checkpoint CK3.

[0069] As described above, the checkpointing mechanism requires additional storage space. Furthermore, when generating a checkpoint, the business process is in a blocked state; new data input is cached, and each business process can only resume processing new task data after the checkpoint is written. Writing a checkpoint can take several seconds to several minutes, which is unacceptable for businesses requiring low latency. In addition, business recovery can only be based on the most recently saved checkpoint, meaning that business recovery has a certain lag and cannot resume from the point of interruption. All these factors will have a certain impact on business performance.

[0070] Based on this, this application provides a data recovery method that does not require additional storage space and can reduce the impact on business performance when recovering data.

[0071] The data recovery method provided in this application embodiment can be applied to computing devices. Figure 2 is a schematic diagram of the structure of the computing device provided in this application embodiment. The computing device can be a server or a terminal device. As shown in Figure 2, the computing device may include: a processor 110, a memory 120, and a communication interface 130. The processor 110, the memory 120, and the communication interface 130 can each be one or more.

[0072] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the computing device. In other embodiments of this application, the computing device may include more or fewer components than illustrated, or combine some components, or split some components, or arrange the components differently. The functional characteristics of the illustrated components may be implemented in hardware, in software, or in a combination of software and hardware.

[0073] In this embodiment, the processor 110, memory 120, and communication interface 130 can communicate with each other via bus 140. This bus 140 can be an industry standard architecture (ISA) bus, a peripheral component interconnect express (PCIe) bus, or an extended industry standard architecture (EISA) bus, etc. The bus 140 may include an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used to represent it in the figure, but this does not indicate that there is only one bus or one type of bus.

[0074] Processor 110 is the control center of the computing device. Processor 110 can be a computing unit with computing capabilities, such as a central processing unit (CPU), graphics processing unit (GPU), or neural-network processing unit (NPU). Processor 110 can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. For ease of description, the following embodiments use a CPU as an example to illustrate processor 110.

[0075] In some embodiments, processor 110 may include one or more CPUs.

[0076] In some embodiments, the computing device may include a plurality of processors 110, each of which may be a single-core processor or a multi-core processor. Here, processor 110 may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0077] In some embodiments, the processor 110 may include a memory controller 111 for managing the memory 120 and handling data exchange between the memory 120 and the processor 110. For example, after receiving a write request from the processor 110, the memory controller 111 can store the data in the write request in the memory 120. As another example, after receiving a read request from the processor 110, the memory controller 111 can read data from the memory 120 according to the memory address carried in the read request and return the read data to the processor 110.

[0078] The memory controller 111 can use the above-mentioned ECC mechanism to read and write data, thereby realizing the detection and correction of memory errors.

[0079] As one possible implementation, the memory controller 111 can also be an external device to the processor 110, connected to the processor 110 via a bus.

[0080] Memory 120, also known as main memory, refers to the internal memory that directly exchanges data with the processor. Memory 120 is characterized by its ability to read and write data at any time and its high speed, and can be used as temporary data storage for the operating system or other running programs. Memory 120 may include volatile memory, which can be random access memory (RAM). Its use as an external cache is an example and not a limitation. RAM includes various forms such as: static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced dynamic random access memory (EDRAM), high bandwidth memory (HBM), and direct rambus RAM (DRRAM).

[0081] In this embodiment of the application, the memory 120 connected to the processor 110 may include a near-end memory 121 and a far-end memory 122. The near-end memory 121 is typically located closer to the processor 110, while the far-end memory 122 is typically located farther from the processor 110. The near-end memory 121 and the far-end memory 122 have one or more of the following differences: the access speed of the near-end memory 121 is higher than that of the far-end memory 122; the access latency of the near-end memory 121 is lower than that of the far-end memory 122; the storage capacity of the near-end memory 121 is smaller than that of the far-end memory 122; and the cost of the near-end memory 121 is higher than that of the far-end memory 122.

[0082] For example, the near-end memory 121 can be high bandwidth memory (HBM), and the far-end memory 122 can be dynamic random access memory (DRAM) or double data rate synchronous dynamic random access memory (DDR SDRAM), etc.

[0083] The near-end memory 121 can be used as a computation buffer or a communication buffer. The computation buffer refers to the buffer between the processor 110 and the far-end memory 122. This buffer can serve as a temporary data storage location between the processor 110 and the far-end memory 122, storing data currently being processed by the processor 110 or frequently accessed data, thereby improving data access efficiency. The communication buffer can be used for data transfer between processes within the processor 110 and other processes (referred to here as communication data). That is, the communication buffer can be used to cache inter-process communication data, reducing inter-process communication latency and improving the overall system communication efficiency.

[0084] In some embodiments, the computing device may further include non-volatile memory, which may include at least one of the following: read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, solid-state disk (SSD), etc.

[0085] Communication interface 130 is used for communicating with other devices or communication networks. Communication interface 130 may include a receiving unit to implement receiving functions and a sending unit to implement sending functions. For example, communication interface 130 may be a network interface card (NIC). In some examples, communication interface 130 may also be referred to as a communication module or communication unit, etc.

[0086] In some embodiments, the computing device described above may be a computing node in a computer cluster. For example, the computing device may be a computing node in a supercomputer. A supercomputer is also known as a supercomputing center. A supercomputing center may include multiple computing nodes to efficiently process data using the computing resources on multiple computing nodes. A supercomputing center may also be called a supercomputing system, etc., and the name of the supercomputing center is not limited in the embodiments of this application.

[0087] The aforementioned computing devices can independently or collaboratively with other computing devices (such as other computing nodes in a supercomputing center) to implement one or more services, such as high-performance LINPACK (HPL) services, artificial intelligence (AI) services, and other high-performance computing (HPC) services. HPL is a benchmark test program used to test the floating-point performance of supercomputers. LINPACK is short for linear system package, and HPL is an advanced version of LINPACK. HPL tests and evaluates the floating-point performance of supercomputers by solving dense systems of linear equations.

[0088] A computing device can implement related services through several processes running on the processor. That is, when implementing related services, the computing device can distribute the various tasks of the service to multiple processes (i.e., service processes), and these processes will work together to implement the service. Each process can execute its related tasks using the allocated memory.

[0089] As mentioned above, the memory connected to the processor 110 includes near memory 121 and far memory 122. Near memory 121 can be used as a calculation buffer or a communication buffer. That is, when each process performs related tasks, it can request far memory 122 or near memory 121 as a calculation buffer or a communication buffer.

[0090] Specifically, a process can request remote memory 122 to store data to be processed by the process, as well as intermediate data generated during data processing. A process can request near memory 121 as a computation buffer to cache data that the process is currently processing or frequently accessing in remote memory 122; a process can also request near memory 121 as a communication buffer to cache communication data between the process and other processes.

[0091] For near memory 121 and far memory 122, each process can request one or more memory areas. For example, process A requests one memory area in far memory 122 (referred to as the far memory area) and two memory areas in near memory 121 (referred to as the near memory area). One of the near memory areas is used as a calculation buffer, and the other is used as a communication buffer for transmitting data with process B. For example, data transmitted by process A to process B or data obtained by process A from process B can be stored in the communication buffer. Process B can request a corresponding near memory area as a communication buffer for exchanging data with process B's communication buffer.

[0092] For a near-end memory area requested by a process, if it is used as a computation buffer, the data stored in the near-end memory area can be found in the same data in the far-end memory area requested by the same process; if it is used as a communication buffer, the data stored in the near-end memory area can be found in the same data in the near-end memory areas of other processes.

[0093] Based on this, in this embodiment, for any near-end memory region requested by a process, a memory region containing the same data as the near-end memory region can be determined from near-end memory regions requested by other processes or far-end memory regions requested by the current process, and this region can be used as the associated memory of the near-end memory region. When a memory error occurs in the near-end memory region, data recovery can be performed based on the data stored in its associated memory. This eliminates the need for additional storage space for data backup and avoids blocking process write checkpoints, allowing data recovery from the point of business interruption, thereby reducing the impact of the data recovery mechanism on business performance. The data recovery method is described in detail below.

[0094] Figure 3 is a flowchart illustrating the data recovery method provided in this embodiment. This method can be executed by the aforementioned computing device, or by a module applied in the computing device (e.g., a processor, chip, or chip system), or by a logical node, logical module, or software capable of implementing all or part of the functions of the computing device. As shown in Figure 3, the method may include the following steps:

[0095] S100. Determine the associated memory of the first memory area; the first memory area is the near memory requested by the first process, and the associated memory includes the second memory area. Both the second memory area and the first memory area store the first data. The second memory area includes the near memory requested by the second process or the far memory requested by the first process.

[0096] The first process can be any business process of a certain business executed by the computing device. The first process can request one or more near-end memory regions, and the first memory region can be any near-end memory region requested by the first process.

[0097] The associated memory acts similarly to backup memory, storing initial data (i.e., the data stored in the first memory area) to recover data from the first memory area in the event of a memory error. As mentioned earlier, the near-end memory area requested by the business process can be used as a communication buffer or a computation buffer; that is, the first memory area can be either a communication buffer or a computation buffer. The associated memory corresponding to the first memory area in these two scenarios is explained below.

[0098] The first memory area is a communication buffer.

[0099] As mentioned earlier, when a process allocates a near-end memory area as a communication buffer, the data stored in this near-end memory area can be found in the near-end memory areas of other processes. That is, the data stored in the first memory area can be found in the near-end memory areas allocated by other processes. The first memory area allocated by the first process can be used to transmit communication data with one or more other processes. Correspondingly, the associated memory of the first memory area can include one or more memory areas allocated by other processes (referred to here as associated memory areas).

[0100] For example, the associated memory of the first memory region includes a second memory region requested by the second process, which stores the first data; in some examples, the associated memory of the first memory region also includes a third memory region requested by the third process, which also stores the first data; in some examples, the associated memory of the first memory region may also include more associated memory regions.

[0101] For example, as shown in Figure 4, the business processes running in the computing device include processes 1 to 4. Memory area NM1 is the near-end memory requested by process 1, memory area NM2 is the near-end memory requested by process 2, memory area NM3 is the near-end memory requested by process 3, and memory area NM4 is the near-end memory requested by process 4. Memory areas NM1 to NM4 are all communication buffers. The first process can be any process among processes 1 to 4. Here, process 1 is used as an example for illustration. Correspondingly, the first memory area is memory area NM1.

[0102] It is understood that this example illustrates the business process as including four processes. The business process running on the computing device may include more or fewer processes, and this embodiment does not make any special limitation on this.

[0103] Assuming that memory region NM1 of process 1 is used to transfer data with processes 3 and 4, that is, the data stored in memory region NM1 can be found in memory regions NM3 and NM4, then the associated memory of memory region NM1 can include memory regions NM3 and / or NM4. For example, the associated memory determined by memory region NM1 includes memory regions NM3 and NM4.

[0104] In practical implementation, after the first process obtains the first memory area, it can search for associated memory that stores the same data as the first memory area during business operation, and then record the found associated memory so that the data can be recovered based on the associated memory in the event of a memory error.

[0105] To facilitate the search for associated memory, in some embodiments, as shown in Figure 4, the near-end memory area requested by a process can be assigned a data identifier Tag. This data identifier Tag can be used to identify the data stored in the corresponding memory area. After the first process requests the first memory area, it can set the value of the data identifier Tag of the first memory area (referred to here as the target value) during business operation. Then, it can search for near-end memory areas requested by other processes with the data identifier Tag of the target value, and determine the associated memory based on the found near-end memory area.

[0106] Continuing with Figure 4 as an example, the data identifier Tag values ​​for memory areas NM1, NM3, and NM4 are all 0001, while the data identifier Tag value for memory area NM2 is 0002. Process 1 can then determine, based on this data identifier Tag, that the data stored in memory areas NM3 and NM4 is the same as the data stored in memory area NM1. Therefore, memory areas NM3 and / or NM4 can be considered as associated memory areas NM1. It is understood that the data identifier Tag values ​​here are merely an example to indicate whether the data identifier values ​​of memory areas are the same, and are not intended to limit this application. For information on how to set the data identifier Tag, please refer to the following description.

[0107] In some embodiments, the data identifier Tag can be a checksum of the data stored in the memory area. For example, the checksum can be a checksum. After storing data in the first memory area, the first process can calculate the checksum of the data stored in the first memory area and set the value of the data identifier Tag in the first memory area to the checksum.

[0108] It is understood that the checksum can also be other forms of checksum, such as a hash value, etc., and this application embodiment does not particularly limit it. In addition, during the execution of business by the first process, the data stored in the first memory area will be updated. The first process can update the data identifier Tag of the first memory area each time the data in the first memory area is updated, and then update the associated memory of the first memory area based on the updated data identifier Tag.

[0109] In some embodiments, data tags in memory areas can be set based on business logic. Specifically, based on business logic, it can be determined which processes will transfer data to each other during each business operation phase. This allows for the design of the data tag setting method for the near-end memory areas of each process, ensuring that the data tag values ​​of the near-end memory areas storing the same data are identical. For example, when executing subtask A, process 1 and process 2 will transfer data to each other, and their communication buffers for data transfer can be set to the same data tag based on preset rules. When executing subtask B, process 1 transfers data to process 3 and process 4 respectively, and the communication buffers for data transfer among these three processes can be set to the same data tag based on preset rules.

[0110] In some implementations, when determining associated memory, the first process can search for all memory regions that meet the target requirements (store the same data as the first memory region) and determine all the memory regions found as associated memory regions of the first memory region.

[0111] In some implementations, the number of memory regions included in the associated memory of the first memory region can be set to an upper limit (target number). The first process can stop searching when the number of memory regions that meet the target requirements reaches the target number, and the target number of memory regions found are designated as the associated memory of the first memory region. It is understandable that the number of memory regions that meet the target requirements may be less than the target number. In this case, all memory regions that meet the target requirements can be designated as the associated memory of the first memory region. The target number can be set as needed, for example, it could be 10.

[0112] For each associated memory region included in the first memory region, the first process can record the memory identification information of these memory regions during recording. If a memory error occurs subsequently, the corresponding memory region can be accessed based on this memory identification information. This memory identification information can specifically be a pointer or starting address of the associated memory region. For example, in Figure 4, process 1 can specifically record the pointers or starting addresses of memory regions NM3 and NM4. The memory identification information can also be other identification information that can indicate a memory region; this embodiment does not impose any particular limitation on this.

[0113] In some embodiments, the first process may communicate with other processes using message queues or other methods to obtain relevant data from the memory areas of other processes, such as data identifiers (Tags), memory identifier information, and data stored in the memory areas.

[0114] In some embodiments, in order to facilitate the first process to obtain data from other processes, the first process and other processes can use a shared memory mechanism to request memory, and the communication buffers of each process can be shared with each other. This can effectively save the performance overhead caused by inter-process communication and is also easy to implement.

[0115] For example, in Figure 4 above, processes 1 to 4 can use a shared memory mechanism to request near-end memory areas. Each process shares memory areas NM1 to NM4. In this way, each process can easily obtain memory data from each memory area, which facilitates the determination of associated memory and subsequent data recovery.

[0116] In this embodiment of the application, when a process requests a near-end memory area as a communication buffer, it can use a target interface to request memory. The target interface can call a memory request function (such as the malloc function) to request memory from the kernel, and after the memory is requested, it can return a structure, which can include the metadata of the requested memory area.

[0117] For example, when the first process requests the first memory area, it can do so through a target interface, which can return a structure corresponding to the first memory area, including the metadata of the first memory area.

[0118] The target interface can be a predefined target function, etc., and the specific implementation method and interface name are not specifically limited here. The structures of each memory area of ​​the process can be stored in the process's heap space or stack space.

[0119] As shown in Figure 4, the metadata of a memory region may include the address (specifically, the starting address) and length of the memory region, and may also include the data identifier Tag of the memory region mentioned above. In some examples, the metadata of the memory region may also include the status identifier Sta of the memory region to indicate the read / write status of the memory region (i.e., whether it is readable and writable), as described in subsequent embodiments.

[0120] Initially, the value of the data identifier Tag in the structure of the memory area can be empty. After the process writes data to the memory area, it can set the value of the data identifier Tag.

[0121] When a process determines the associated memory of a memory region, it can record the memory identification information of the associated memory region in the structure of that memory region. That is, the structure of the memory region can also include related information of the associated memory (referred to as associated memory information). Initially, the associated memory information in the structure of the memory region can be empty. After the process determines the associated memory of the memory region, it can record the memory identification information of each associated memory region included in the associated memory in the structure of that memory region.

[0122] The first memory area is a calculation buffer.

[0123] As mentioned earlier, when a process uses a near-end memory region as a computation buffer, the data stored in that near-end memory region can be found in the same far-end memory region allocated by the same process. That is, the data stored in the first memory region can be found in the same far-end memory region allocated by the first process. In this case, the associated memory of the first memory region can include the far-end memory region allocated by the first process; that is, the second memory region is the far-end memory allocated by the first process.

[0124] Continuing with the example of business processes running in the computing device, including processes 1 to 4, as shown in Figure 5, the memory areas requested by process 1 include the near memory area NM5 and the far memory area FM1, the memory areas requested by process 2 include the near memory area NM6 and the far memory area FM2, the memory areas requested by process 3 include the near memory area NM7 and the far memory area FM3, and the memory areas requested by process 4 include the near memory area NM8 and the far memory area FM4; the near memory areas NM5 to NM8 are all computing buffers.

[0125] The near memory region NM5 of process 1 can find the same data in the far memory region FM1 requested by process 1. Therefore, the associated memory of the near memory region NM5 can include the far memory region FM1. Similarly, the associated memory of the near memory regions of processes 2 to 4 includes their respective far memory regions.

[0126] Similar to Figure 4, the first process can be any process from process 1 to process 4. Here, we will still take process 1 as the first process as an example for illustration. Correspondingly, the first memory area is the near memory area NM5, and the second memory area (the associated memory of the first memory area) can be the far memory area FM1 of process 1.

[0127] Considering that the data stored in the near memory region NM5 (i.e., the first data) is a portion of the data in the far memory region FM1, in some implementations, the second memory region can specifically be the address space in the far memory allocated by the first process that stores the first data. That is, the second memory region is a segment of memory in the far memory region FM1 that stores the first data. In this way, the first data can be retrieved more conveniently in the event of a memory error.

[0128] Similar to the communication buffer, after the first process allocates a computation buffer (i.e., the first memory area), it can record the memory identification information of the associated memory area (i.e., the second memory area) of the computation buffer. Specifically, this memory identification information can be the address information of the second memory area. For example, if the second memory area is the address space in the first process's remote memory area where the first data is stored, the memory identification information of the second memory area can include the start and end addresses of the second memory area.

[0129] Similar to requesting a communication buffer, the first process can use the aforementioned target interface to request a first memory area as a computation buffer. The associated memory information of the computation buffer can be recorded in the corresponding structure; this structure can include metadata such as the address and length of the computation buffer.

[0130] In some examples, the computation buffer can be configured with a data identifier (Tag) and a status identifier (Sta). That is, the metadata of the computation buffer can include a data identifier (Tag) and a status identifier (Sta). The way the data identifier (Tag) and the status identifier (Sta) are configured is similar to that of the communication buffer, and will not be described again here.

[0131] In some examples, the computation buffer may not have a data identifier (Tag) and a status identifier (Sta). That is, the metadata of the computation buffer may not include the data identifier (Tag) and the status identifier (Sta), or in other words, the data identifier (Tag) and the status identifier (Sta) in the metadata of the computation buffer can be empty. The following explanation will mainly focus on the example of the computation buffer not having a data identifier (Tag) and a status identifier (Sta).

[0132] In this embodiment of the application, when a process (such as the first process) requests a computation buffer, it can use a non-shared memory mechanism to request it. That is, each process's computation buffer can be its own private and not shared with other processes.

[0133] S200: In the event of a memory error in the first memory area, data recovery is performed on the first memory area based on the first data stored in the associated memory.

[0134] As described above, the first memory area stores the first data. The first process uses the memory area of ​​another process that also stores the first data or the remote memory of its own process as the associated memory of the first memory area. If a memory error occurs in the first memory area, the first data can be retrieved from the associated memory to recover the data in the first memory area.

[0135] Specifically, the memory controller can detect memory errors based on the ECC mechanism. As mentioned earlier, memory errors can include CE errors and UCE errors. Considering that the memory controller can automatically correct CE errors, in some examples, the memory error occurring in the first memory area mentioned above can specifically be a UCE error. That is, when a CE error occurs in the first memory area, the ECC mechanism can be used to correct the CE error; when a UCE error occurs in the first memory area, data recovery can be performed on the first memory area based on the first data stored in the associated memory of the first memory area.

[0136] As mentioned earlier, when the memory controller detects a UCE error, it can report the memory error information to the processor via an interrupt mechanism. The processor then reports the memory error information to the operating system. After receiving the memory error information, the operating system can send a signal to the application process to inform it that a UCE error has occurred in the first memory area. The application process can then perform corresponding memory error handling, such as terminating the process. In this embodiment, the signal handling process can be customized using the sigaction mechanism, i.e., a customized memory error handling process, which does not directly terminate the process but triggers a data recovery process.

[0137] During data recovery, the first process can read the first data from the associated memory using CPU operations or direct memory access (DMA) operations, copy the first data to the first memory area, and complete the data recovery. For example, in Figure 4 above, if a UCE error occurs in memory area NM1, process 1 can select a memory area from memory areas NM2 and NM3, and copy the data stored in that memory area to memory area NM1. As another example, in Figure 5 above, if a UCE error occurs in the near-end memory area NM5, process 1 can read the first data from the far-end memory area FM1 and copy the first data to the near-end memory area NM5.

[0138] In the case where the first memory area is a communication buffer, the associated memory of the first memory area is a memory area requested by other processes. Since different processes process data at different speeds, the data update speed of the near memory area of ​​each process is different. Therefore, after the first process finishes recording the associated memory of the first memory area, the data of one or more memory areas in the associated memory may have been updated during the time between when the first memory area encounters a memory error.

[0139] In view of this situation, in some embodiments, if a memory error occurs in the first memory area, the first process can determine the memory area currently storing the first data (referred to here as the target memory area) from the associated memory, and retrieve the first data from the target memory area for data recovery.

[0140] In practice, the first process can access and obtain the current data identifier Tag of each associated memory area based on the recorded memory identifier information of each associated memory area, and determine the associated memory area whose data identifier Tag is the target value (the value of the data identifier Tag of the first memory area) as the target memory area.

[0141] In some examples, the first process can obtain the data identifier (Tag) for each associated memory region and identify any associated memory region whose data identifier (Tag) matches the target value as the target memory region. In other examples, the first process can traverse the associated memory regions sequentially, and after identifying the associated memory region whose data identifier (Tag) matches the target value, it can stop the traversal process and identify that associated memory region as the target memory region.

[0142] When the first process reads the first data from the target memory area, the process to which the target memory area belongs may write data to the target memory area. Considering this situation, in some embodiments, each memory area can be set with a status flag Sta, so that the process can identify and recognize the read and write status of the memory area through the status flag Sta, thereby improving the reliability of data reading and writing.

[0143] The status flag Sta can include three values: a first value D, a second value E, and a third value M. The first value D indicates that the memory area is readable and writable; the second value E indicates that the memory area is not readable; and the third value M indicates that the memory area is not writable.

[0144] Under normal circumstances, the status flag Sta of the memory area can be the first value D; when a memory error occurs in the memory area, its status flag Sta can be set to the second value E; when the data in the memory area is used for data recovery, its status flag Sta can be set to the third value M.

[0145] For example, if a memory error occurs in the first memory area of ​​the first process, its status flag Sta can be set to the second value E to indicate that a memory error has occurred in the first memory area and its data is unreadable. In addition, the status flag Sta of the target memory area used for data recovery can be set to the third value M to indicate that the data in the memory area is being transferred (i.e. copied) and data cannot be written to the memory area at present. When the first memory area completes data recovery, the status flag Sta of the first memory area and the target memory area can be updated to the first value D to indicate that the two memory areas can be read and written normally.

[0146] The data recovery process will be illustrated below using process 1 shown in Figure 4 as an example.

[0147] As mentioned earlier, when process 1 determines the associated memory of memory area NM1, the data identifier Tag of memory areas NM3 and NM4 is the same as the data identifier Tag of memory area NM1, as shown in Figure 6(a). Process 1 determines memory areas NM3 and NM4 as the associated memory of memory area NM1. Each memory area can be read and written normally, and the status identifier Sta is the first value D.

[0148] As shown in Figure 6(b), a UCE error occurs in memory area NM1. At this time, process 1 can update the status flag Sta of memory area NM1 to the second value E. Assuming that the data flag Tag of memory area NM3 has been updated to 0003, and the data flag Tag of memory area NM4 has not been updated and is still 0001, then as shown in Figure 6(c), process 1 can identify memory area NM4 as the target memory area, and then set the status flag Sta of memory area NM4 to the third value M. At this time, process 4 cannot write data to memory area NM4. Afterwards, process 1 can read the first data from memory area NM4, copy the first data to memory area NM1, and restore the data in memory area NM1.

[0149] As shown in Figure 6(d), after process 1 copies the first data from memory area NM4 to memory area NM1, it can set the status flag Sta of both memory area NM4 and memory area NM1 to the first value D; or, process 1 can set the status flag Sta of memory area NM4 to the first value D after reading the first data from memory area NM4, and set the status flag Sta of memory area NM1 to the first value D after writing the first data to memory area NM1.

[0150] In some examples, after a UCE error occurs in memory region NM1, when process 1 is determining the target memory region, it detects that the data identifier tag of memory region NM3 has been updated. At this time, process 1 can also delete memory region NM3 from the associated memory.

[0151] The aforementioned status flag Sta can also be used by the first process to determine the target memory region. For example, when a memory error occurs in the first memory region of the first process, there may also be memory regions with memory errors among the various associated memory regions of the first memory region. Considering this situation, when determining the target memory region, in addition to obtaining the current data flag Ta of the associated memory region, the first process can obtain the current status flag Sta of the associated memory region, and determine the associated memory region where the data flag Ta is the target value and the status flag Sta is not the second value E as the target memory region. That is, the data flag Ta of the target memory region is the target value, and the status flag Sta is the first value D or the third value M.

[0152] In some examples, there may not be a target memory region that meets the requirements. In such cases, the first process can use the original memory error handling methods, such as terminating the process.

[0153] When the first memory area serves as a computation buffer, the associated memory (second memory area) of the first memory area is the remote memory of the first process. Correspondingly, when a UCE error occurs, the first process can read the first data from the second memory area and write the first data back to the first memory area to complete data recovery. For example, the first process can copy the data stored in the address space corresponding to the start and end addresses in the remote memory to the first memory area based on the start and end addresses of the second memory area.

[0154] In some examples, the first process may also update the status flag Sta of the first memory area from the first value D to the second value E after a memory error occurs in the first memory area; and update the status flag Sta of the first memory area from the second value E to the first value D after the data recovery of the first memory area is completed.

[0155] The differences between the above data recovery scheme and the checkpoint mechanism will be explained below with reference to Figure 7.

[0156] Continuing with the example of the business processes shown in Figure 1, including processes R1 to R3, as shown in Figure 7, the memory areas requested by process R1 include computation buffer A1, communication buffer B1, communication buffer C1, and remote memory area D1; the memory areas requested by process R2 include computation buffer A2, communication buffer B2, communication buffer C2, and remote memory area D2; and the memory areas requested by process R3 include computation buffer A3, communication buffer B3, communication buffer C3, and remote memory area D3. Based on the above data recovery method, process R1 can determine that the associated memory of computation buffer A1 includes remote memory area D1; the associated memory of communication buffer B1 includes communication buffers B2 and B3; and the associated memory of communication buffer C1 includes communication buffers C2 and C3. Other processes are similar; here, process R1 is used as an example for illustrative purposes.

[0157] Suppose that a UCE error occurs in the computation buffer A1 of process R1 at time T9. Process R1 can then read data from the associated memory (remote memory area D1) of computation buffer A1, restore the data stored in computation buffer A1, and then continue executing the relevant tasks. Suppose that a UCE error occurs in the communication buffer B1 of process R1 at time T11. Process R1 can then read data from the associated memory (communication buffer B2 or communication buffer B3) of communication buffer B1. For example, here, data is read from communication buffer B3 to restore the data stored in communication buffer B1, and then the relevant tasks can continue executing.

[0158] Compared with the checkpoint mechanism shown in Figure 1, the above data recovery scheme does not require additional storage space for each business process to store checkpoints. Moreover, it does not require blocking processes from writing checkpoints and can recover data from the point of business interruption after a UCE error occurs.

[0159] The data recovery method provided in this application, for a first memory area requested by a first process, determines a memory area containing the same data as the first memory area from a near-end memory area requested by another process or a far-end memory area requested by the current process, and uses this as the associated memory of the first memory area. When a memory error occurs in the first memory area, data recovery can be performed based on the data stored in its associated memory. This solution eliminates the need for additional storage space for data backup and avoids blocking process write checkpoints, allowing data recovery from the point of business interruption, thereby reducing the impact on business performance.

[0160] The data recovery method provided in this application can be applied to businesses such as HPL or AI. The following description uses HPL as an example.

[0161] As mentioned earlier, HPL is a benchmark test program used to test the floating-point performance of supercomputers. HPL tests and evaluates the floating-point performance of supercomputers by solving dense linear equations.

[0162] The HPL testing process mainly includes matrix generation, matrix decomposition, back-substitution solution, result verification, and performance evaluation. The core task of HPL is to solve the linear equation system A*x=b, where A is an N×N dense matrix, x represents an unknown vector, b represents the right-hand side vector, and N is a positive integer.

[0163] In the matrix generation stage, a pseudo-random dense matrix A can be generated using the linear congruential algorithm (LCA), and a random vector b can also be generated.

[0164] In the matrix decomposition stage, matrix A can be decomposed into a lower triangular matrix L and an upper triangular matrix U using the LU decomposition algorithm (Gaussian elimination), i.e., A = LU.

[0165] In the back-substitution stage, based on L and U obtained from LU decomposition, L*y = b is solved by forward substitution, and then U*x = y is solved by back substitution to obtain the final solution x.

[0166] During the result verification phase, the correctness of the calculation results is verified by calculating the residuals. Specifically, the calculation result is considered correct when the norm of the residuals is less than the preset machine precision.

[0167] During the performance evaluation phase, floating point operations per second (FLPOS) can be used to measure the performance of computing devices.

[0168] Matrix decomposition is the most time-consuming stage in the entire HPL process and a crucial factor in determining computer performance. The matrix decomposition process is briefly described below using matrix A shown in Figure 8 as an example.

[0169] When performing matrix factorization, matrix A can be divided into multiple data blocks, each containing M×M data items, where M is a positive integer. These data blocks can be assigned to multiple processes (i.e., business processes) for processing, with each data block corresponding to one process, and each process handling one or more data blocks. For example, here, each data block is assigned to processes R00 to R22. It is understood that the size of matrix A, the number of business processes executing the matrix factorization task, and the correspondence between data blocks and processes are merely examples, and this application embodiment does not impose any particular limitations on these aspects.

[0170] LU decomposition employs an iterative approach, processing one row and one column of data in each iteration and updating the tail matrix. After multiple iterations, the entire matrix A is decomposed. Specifically, referring to Figure 8, A11 represents the data block in the first row and first column, A12 includes the other data blocks in the row containing A11, and A21 includes the other data blocks in the column containing A11. In the first iteration, process R00 first decomposes data block A11, obtaining decomposition results L11 and U11. Then, L11 is broadcast to the process corresponding to A12 (i.e., the other data blocks in the row containing A11) (referred to as the row process) for data processing, obtaining U12 for each data block; U11 is broadcast to the process corresponding to A21 (i.e., the other data blocks in the column containing A11) (referred to as the column process) for data processing, obtaining L21 for each data block; then, the tail matrix A22 is updated based on each U12 and L21. In the second iteration, a similar decomposition and update is performed on the updated A22 until the entire matrix A is decomposed. It is understood that this is only a brief description of the matrix decomposition process, and the specific matrix decomposition process may include more processing steps.

[0171] In this embodiment, each process executing the matrix factorization task can apply for remote memory to store the data blocks to be processed, and can apply for multiple near-end memory areas as communication buffers and computation buffers. The near-end memory areas applied for by each process can employ the data recovery method described in the preceding embodiments. That is, the first process mentioned in the preceding embodiments can be any process among the processes executing the matrix factorization task. The specific implementation process is described below.

[0172] For any process performing the above matrix factorization task, the near-end memory requested by the process may include: computation memory area C, first communication memory area ML, and second communication memory area MU; the computation memory area C, first communication memory area ML, and second communication memory area MU requested by the process may each include one or more.

[0173] The computation memory area C is a computation buffer used to cache data related to the computation task, such as the currently processed data block and data generated during the computation process. The first communication memory area ML and the second communication memory area MU are both communication buffers. The first communication memory area ML is used to cache communication data between the process and at least one corresponding row process (also referred to as the first target process), and the second communication memory area MU is used to cache communication data between the process and at least one corresponding column process (also referred to as the second target process). Row processes and column processes are processes other than the current process; the data block corresponding to a row process is located in the same row as the data block corresponding to the current process, and the data block corresponding to a column process is located in the same column as the data block corresponding to the current process.

[0174] For example, in Figure 8, process R00 corresponding to data block A11 can request a computation memory area C00, a first communication memory area ML00, and a second communication memory area MU00 for data block A11. The computation memory area C11 can be used to cache data block A11 and data generated during computation; the first communication memory area ML00 can be used to cache L11 obtained from the decomposition of data block A11, and other row processes in the same row as data block A11 can obtain L11 from this first communication memory area ML00; the second communication memory area MU00 can be used to cache U11 obtained from the decomposition of data block A11, and other column processes in the same column as data block A11 can obtain U11 from this second communication memory area MU00. Process R00 can also request computation memory areas for other data blocks to be processed in the current iteration to process the corresponding data blocks. Similarly, other row processes and column processes can request one or more computation memory areas C, first communication memory areas ML, and second communication memory areas MU, respectively, for data decomposition tasks in the current iteration.

[0175] Each process can determine the data block to be processed based on the correspondence between data blocks and processes in matrix A, and then request a computation memory area C, a first communication memory area ML, and a second communication memory area MU accordingly. The first and second communication memory areas ML and MU, serving as communication buffers, can be requested using a shared memory mechanism, with each process sharing its communication buffer. The near-end memory areas requested by each process can be reused in each iteration; that is, after a process has processed the data block in the current iteration using the requested near-end memory areas, it can continue to use these near-end memory areas to process the data block in the next iteration.

[0176] During the matrix factorization task, each process can set a data identifier tag for the requested communication buffer and determine the associated memory of the communication buffer based on the data identifier tag. Specifically, when setting the data identifier tag, the first communication memory area (ML) of the same row process can be set to the same data identifier tag, and the second communication memory area (MU) of the same column process can be set to the same data identifier tag. That is, the first communication memory area (ML) of the same row process can be used as the associated memory of each other, and the second communication memory area (MU) of the same column process can be used as the associated memory of each other.

[0177] In other words, for any process, the associated memory of the process's first communication memory area ML may include: the first communication memory areas ML of one or more line processes corresponding to the process; the associated memory of the process's second communication memory area MU may include: the second communication memory areas MU of one or more column processes corresponding to the process.

[0178] For example, the row processes corresponding to process R00 include processes R01 and R02, the column processes corresponding to process R00 include processes R10 and R20, the associated memory of the first communication memory area ML00 of process R00 may include: the first communication memory area ML of process R01 and / or process R02; the associated memory of the second communication memory area MU00 of process R00 may include: the second communication memory area MU of process R10 and / or process R20.

[0179] Each process can set a data identifier tag for each communication buffer based on the data blocks stored in each communication buffer, or it can set a data identifier tag for each communication buffer based on the business operation stage. For example, the matrix decomposition task includes multiple iterations, and each process can set a data identifier tag for each communication buffer based on the current iteration round.

[0180] For any process, the associated memory of the process's compute memory area can include remote memory requested by the process. For example, the associated memory of the compute memory area C00 of process R00 can include remote memory requested by process R00, specifically the address space in the remote memory requested by process R00 that stores the data related to data block A11.

[0181] For the near-end memory area requested by each process, the status flag Sta can be set to D.

[0182] Suppose process R00 encounters a UCE error in the first communication memory area ML00. The associated memory areas of ML00 include the first communication memory areas ML01 of process R01 and ML02 of process R02. Process R00 can set the status flag Sta of the first communication memory area ML11 to the second value E, and can determine a target memory area from ML01 and ML02. The current data flag Ta of the target memory area is the same as the data flag Ta of the first communication memory area ML00, and the current status flag Sta of the target memory area is either the first value D or the third value M. For example, if the target memory area is the first communication memory area ML01 of process R01, then process R00 can set the status flag Sta of the first communication memory area ML01 to the third value M through a shared memory mechanism, and can use CPU operations or DMA operations to copy data from the first communication memory area ML01 to the first communication memory area ML00. After the copy is complete, process R00 can set the status flag Sta of both the first communication memory area ML01 and the first communication memory area ML00 to the first value D. Process R00 can then use the restored first communication memory area ML00 to continue executing the matrix decomposition task.

[0183] In the above embodiments, the matrix decomposition task of HPL service adopts the data recovery method described in the foregoing embodiments, which does not require blocking the process to write checkpoints and can recover data from the point of service interruption, thereby improving computational efficiency.

[0184] The above example illustrates the application of the data recovery solution to matrix factorization tasks in HPL services. It is understood that this data recovery solution can also be applied to matrix factorization tasks in other services, such as model training, signal processing, and image reconstruction. The specific implementation principles are similar and will not be elaborated here.

[0185] The above describes the data recovery method provided by the embodiments of this application. In the various embodiments of this application, unless otherwise specified or logically conflicting, the terms and / or descriptions between the various embodiments are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.

[0186] The apparatus provided in the embodiments of this application will now be described in detail with reference to FIG9. It should be understood that the description of the apparatus embodiments corresponds to the description of the method embodiments. For ease of reading, the apparatus embodiments here will not repeat the details of the foregoing method embodiments one by one, but it should be clear that the apparatus in this embodiment can implement all the contents of the foregoing method embodiments.

[0187] Figure 9 is a schematic diagram of the data recovery device provided in the embodiment of this application. The device can be the above-mentioned computing device or a component in the computing device, used to implement the method involved in the above-mentioned method embodiment.

[0188] As shown in Figure 9, the device may include a processing module 210 and a data recovery module 220. The processing module 210 is used to determine the associated memory of the first memory area. The first memory area is the near memory requested by the first process. The associated memory includes a second memory area. Both the second memory area and the first memory area store the first data. The second memory area includes the near memory requested by the second process or the far memory requested by the first process.

[0189] The data recovery module 220 is used to recover data from the first memory area based on the first data stored in the associated memory area in the event of a memory error in the first memory area.

[0190] In one possible implementation, the second memory area is the near-end memory requested by the second process, and the first process and the second process share the first memory area and the second memory area.

[0191] In one possible implementation, the associated memory also includes a third memory region that stores the first data. The third memory region is near-end memory requested by the third process.

[0192] In one possible implementation, the near-end memory of the computing device includes multiple memory regions, including a first memory region and a second memory region. Each of the multiple memory regions has a corresponding data identifier, which is used to identify the data stored in the corresponding memory region. The data identifier values ​​of the first memory region and the second memory region are the same.

[0193] In one possible implementation, the data identifier is a checksum of the data stored in the corresponding memory area.

[0194] In one possible implementation, the near-end memory of the computing device includes multiple memory regions, including a first memory region and a second memory region. Each of the multiple memory regions has a corresponding status flag, which includes any one of the following:

[0195] The first value indicates whether the memory area is readable and writable;

[0196] The second value is used to indicate that the memory area is not readable;

[0197] The third value is used to indicate that the memory area is not writable.

[0198] In one possible implementation, after a memory error occurs in the first memory area, the status flag of the first memory area is updated from a first value to a second value, and the status flag of the second memory area is updated from a first value to a third value; after the first memory area completes data recovery, the status flag of the first memory area is updated from a second value to a first value, and the status flag of the second memory area is updated from a third value to a first value.

[0199] In one possible implementation, the second memory region is the address space in the remote memory requested by the first process that stores the first data.

[0200] In one possible implementation, the first process is any one of multiple business processes, each of which is used to execute the matrix decomposition task of the HPL business.

[0201] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The functional characteristics of the integrated unit can be implemented in hardware, software, or a combination of hardware and software. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0202] This application also provides a computing device cluster, which may include one or more computing devices, and the one or more computing devices can implement the methods described in the above method embodiments.

[0203] This application also provides a readable storage medium (also known as a computer-readable storage medium) storing a program thereon, which, when executed by a processor, implements the method described in the above-described method embodiments.

[0204] This application also provides a program product (also known as a computer program product) that, when run on a device, causes the device to implement the method described in the above-described method embodiments.

[0205] This application also provides a chip system including a processor coupled to a memory. The processor executes a program stored in the memory to implement the method described in the above embodiments. The chip system may be a single chip or a chip module composed of multiple chips.

[0206] In the above embodiments, each processing step or functional feature can be implemented, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a program product. The program product includes one or more instructions. When the instructions are loaded and executed on the device, the process or function described in accordance with the embodiments of this application is generated, in whole or in part. The instructions can be stored in a readable storage medium or transmitted through the readable storage medium.

[0207] The naming or numbering of steps in this application does not imply that the steps in the method flow must be executed in the time / logical order indicated by the naming or numbering. The execution order of the named or numbered process steps can be changed according to the technical purpose to be achieved, as long as the same or similar technical effect can be achieved.

[0208] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0209] In the embodiments provided in this application, it should be understood that the disclosed apparatus / devices and methods can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0210] It should be understood that in the description of this application and the appended claims, the terms "comprising," "including," "having," and any variations thereof are intended to cover a non-exclusive inclusion and mean "including but not limited to," unless otherwise specifically emphasized. For example, a process, method, system, product, or apparatus that includes a series of steps or modules is not necessarily limited to those steps or modules that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to such process, method, product, or apparatus.

[0211] In the description of this application, unless otherwise stated, " / " indicates that the objects before and after are in an "or" relationship. For example, A / B can mean A or B. "And / or" in this application is used to describe the relationship between the related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural.

[0212] Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0213] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0214] Furthermore, in the description of this application and the appended claims, the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein; features defined as "first" or "second" may explicitly or implicitly include at least one of those features.

[0215] In the embodiments of this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of the words "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.

[0216] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.

[0217] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A data recovery method, characterized in that, Applied to a computing device having near-end memory and far-end memory, the method includes: Determine the associated memory of the first memory area; the first memory area is the near memory requested by the first process, and the associated memory includes a second memory area. Both the second memory area and the first memory area store the first data. The second memory area includes the near memory requested by the second process or the far memory requested by the first process. In the event of a memory error in the first memory area, data recovery is performed on the first memory area based on the first data stored in the associated memory.

2. The method according to claim 1, characterized in that, The second memory area is the near-end memory requested by the second process, and the first process and the second process share the first memory area and the second memory area.

3. The method according to claim 2, characterized in that, The associated memory also includes a third memory area, which stores the first data. The third memory area is the near-end memory requested by the third process.

4. The method according to claim 2 or 3, characterized in that, The near-end memory of the computing device includes multiple memory regions, including a first memory region and a second memory region. Each of the multiple memory regions has a corresponding data identifier, which is used to identify the data stored in the corresponding memory region. The data identifier values ​​of the first memory region and the second memory region are the same.

5. The method according to claim 4, characterized in that, The data identifier is the checksum of the data stored in the corresponding memory area.

6. The method according to any one of claims 2-5, characterized in that, The near-end memory of the computing device includes multiple memory regions, including the first memory region and the second memory region. Each of the multiple memory regions has a corresponding status identifier, which includes any one of the following: The first value is used to indicate that the memory area is readable and writable; The second value is used to indicate that the memory area is unreadable; The third value is used to indicate that the memory area is not writable.

7. The method according to claim 6, characterized in that, After a memory error occurs in the first memory area, the status flag of the first memory area is updated from a first value to a second value, and the status flag of the second memory area is updated from a first value to a third value. After the first memory area completes data recovery, the status identifier of the first memory area is updated from the second value to the first value, and the status identifier of the second memory area is updated from the third value to the first value.

8. The method according to claim 1, characterized in that, The second memory area is the address space in the remote memory requested by the first process that stores the first data.

9. The method according to any one of claims 1-8, characterized in that, The first process and the second process are used to perform matrix decomposition tasks.

10. A data recovery device, characterized in that, Applied to a computing device having near-end memory and far-end memory, the device includes: The processing module is used to determine the associated memory of the first memory area; the first memory area is the near memory requested by the first process, and the associated memory includes a second memory area. Both the second memory area and the first memory area store the first data. The second memory area includes the near memory requested by the second process or the far memory requested by the first process. The data recovery module is used to recover data from the first memory area based on the first data stored in the associated memory area in the event of a memory error in the first memory area.

11. The apparatus according to claim 10, characterized in that, The second memory area is the near-end memory requested by the second process, and the first process and the second process share the first memory area and the second memory area.

12. The apparatus according to claim 11, characterized in that, The associated memory also includes a third memory area, which stores the first data. The third memory area is the near-end memory requested by the third process.

13. The apparatus according to claim 11 or 12, characterized in that, The near-end memory of the computing device includes multiple memory regions, including a first memory region and a second memory region. Each of the multiple memory regions has a corresponding data identifier, which is used to identify the data stored in the corresponding memory region. The data identifier values ​​of the first memory region and the second memory region are the same.

14. The apparatus according to any one of claims 11-13, characterized in that, The near-end memory of the computing device includes multiple memory regions, including the first memory region and the second memory region. Each of the multiple memory regions has a corresponding status identifier, which includes any one of the following: The first value is used to indicate that the memory area is readable and writable; The second value is used to indicate that the memory area is unreadable; The third value is used to indicate that the memory area is not writable.

15. The apparatus according to claim 14, characterized in that, After a memory error occurs in the first memory area, the status flag of the first memory area is updated from a first value to a second value, and the status flag of the second memory area is updated from a first value to a third value. After the first memory area completes data recovery, the status identifier of the first memory area is updated from the second value to the first value, and the status identifier of the second memory area is updated from the third value to the first value.

16. A computing device, characterized in that, include: A memory and a processor, the memory being used to store a computer program; the processor being used to execute the method as described in any one of claims 1-9 when the computer program is invoked.

17. A program product, characterized in that, When the program product is run on the device, the device performs the method as described in any one of claims 1-9.

18. A chip system, characterized in that, The chip system includes a processor coupled to a memory, the processor executing a program stored in the memory to implement the method as described in any one of claims 1-9.