Memory management method and related device

The processor core executes program code to manage memory data, which solves the problem that existing ECC technology cannot identify DRAM multi-bit errors, and realizes efficient data management without hardware changes, reducing costs and improving data reliability.

WO2025140156A1PCT designated stage expired Publication Date: 2025-07-03HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/141685
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-12-24
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The existing ECC technology has limitations in identifying data changes in DRAM, which cannot effectively resolve multi-bit errors, and requires changing the hardware circuit and increasing memory costs.

Method used

Through software, the processor core executes program code to manage memory data, and identify changes in the memory by comparing the reading of target data and verification data, avoiding hardware changes, and expanding usage scenarios.

Benefits of technology

It realizes the identification of memory data changes without changing the hardware circuit, reduces memory management costs, reduces the risk of abnormal device restarts, and improves data reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024141685_03072025_PF_FP_ABST
    Figure CN2024141685_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides a memory management method. The method is implemented by a processor core executing program code, so as to manage data in a memory. In this way, it is conducive to identifying, by means of software, whether the data stored in the memory changes due to memory reliability. Compared with ECC techniques that need controllers and / or DDR chips to support ECC characteristics, the memory management method provided in the present application does not need to modify the design of a hardware circuit, which is conducive to reducing requirements of memory management on DDR hardware, and expanding usage scenarios for the memory management method. In some scenarios, ECC devices of a CPU subsystem can be removed to achieve hardware cost reduction. The present application further provides a management apparatus corresponding to the management method, a computer device, a computer-readable storage medium, and a computer program product.
Need to check novelty before this filing date? Find Prior Art

Description

Memory management method and related equipment

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 29, 2023, with application number 202311863931.X and invention name “Memory management method and related equipment”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of storage technology, and in particular to a storage management method and related equipment. Background Art

[0003] Dynamic random access memory (DRAM) is commonly used as memory in computer system architecture. DRAM, as a volatile storage medium, uses the amount of electricity stored in a capacitor to represent data 0 and 1, making it a volatile memory. Due to leakage in capacitors, they can only maintain charge for a very short time. Capacitor leakage causes charge drift, which can lead to errors in stored data and failures such as bit jumps. With the evolution of manufacturing processes, especially in the context of diversification and localization, the problem of data failure caused by bit jumps is becoming increasingly serious. Therefore, it is particularly important to identify whether the data stored in DRAM has changed. Summary of the Invention

[0004] The present application provides a memory management method and related devices for determining whether data stored in the memory has changed.

[0005] On the first aspect, the present application provides a method for managing a memory, wherein the memory is connected to a processing unit via a memory controller, and a first address and a second address in the memory are used to store target data and verification data of the target data, respectively. The processing unit implements the method by running an executable file, and the method may include: reading the data of the first address and the data of the second address from the memory via the memory controller; using the data of the second address to determine the reliability information of the data of the first address, wherein the reliability information is used to indicate whether the data of the first address is consistent with the target data or inconsistent. In this way, it is helpful to identify whether the data stored in the memory has changed due to memory reliability issues by software means. Compared with ECC technology that requires the controller and / or DDR chip to support ECC features, the memory management method provided by the present application does not require changes to the design of the hardware circuit, which is helpful to reduce the requirements of memory management on DDR hardware and expand the use scenarios of the memory management method.

[0006] Optionally, reading the data at the first address and the data at the second address from the memory respectively through the memory controller includes: periodically reading the data at the first address and the data at the second address from the memory respectively through the memory controller.

[0007] Optionally, the target data is data in one or more executable files other than the executable file. This facilitates the core to check other executable files that have not been called when running the executable file, and facilitates early identification of errors in other executable files that have not been called, thereby facilitating pre-processing (such as reloading) before they are called. Compared with the ECC method that only checks other executable files that are being called, this helps avoid abnormal device restarts due to multi-bit errors in the executable file being called.

[0008] Optionally, the target data is a program code segment, a program static data segment, or a program dynamic data segment in the other one or more executable files. Because different types of segmentation faults in executable files generally have different response measures, by examining a certain type of segment in the executable file, when a segmentation fault of that type is identified, more precise response measures can be implemented.

[0009] Optionally, the target data includes data in the copy of the executable file, which is helpful in ensuring the data integrity of the executable file.

[0010] Optionally, the executable file is stored at a third address in the memory, and the executable file run by the core when executing the management method of the present application and the data to be managed are stored in the same memory. This helps to reduce the amount of memory that the core relies on to execute the method of the present application and expand the applicable scenarios of the method of the present application.

[0011] Optionally, the verification data of the target data is a characteristic value of the target data, which is conducive to saving storage resources of the memory. Optionally, the verification data of the target data is the target data, so that the verification data can be used as a backup of the target data, which is conducive to improving the security of the target data.

[0012] Optionally, the method further includes: calculating verification data for the target data; and writing the verification data for the target data to the second address via the storage controller. In this way, even if the verification data for the target data is not stored in the memory, the core can generate the verification data for the target data by running the executable file and write it to the memory, which helps expand the applicable scenarios of the method of the present application.

[0013] Optionally, the reliability information indicates that the data at the first address is inconsistent with the target data. The method further includes: reloading the target data into the memory; and / or reporting the reliability information to a host computer connected to the processing unit. This facilitates correcting errors in the target data before calling the target data, and helps reduce or avoid major issues such as abnormal device restarts.

[0014] Optionally, the memory is a memory, a flash memory or a hard disk.

[0015] In a second aspect, the present application provides a memory management device, wherein the memory is connected to a processing unit via a memory controller, wherein a first address and a second address in the memory are respectively used to store target data and verification data of the target data, and the processing unit generates the management device by running an executable file, and the management device includes a reading module and a determination module. The reading module is used to read the data at the first address and the data at the second address from the memory via the memory controller, respectively. The determination module is used to determine the reliability information of the data at the first address using the data at the second address, wherein the reliability information is used to indicate whether the data at the first address is consistent with the target data or not.

[0016] Optionally, the reading module can be used to read the data at the first address and the data at the second address from the memory through the memory controller after the first data and the verification data of the first data are written to the memory. Alternatively, the reading module can periodically read the data at the first address and the data at the second address from the memory through the memory controller, for example, by managing the data in the executable file through a timer task. Specifically, the core can add a memory management task to the timer and set the corresponding time slice size. When the time slice is exhausted, the core can manage the data in the executable file already stored in the DDR.

[0017] Optionally, the target data is data in one or more executable files other than the executable file.

[0018] Optionally, the target data is a program code segment or a program static data segment or a program dynamic data segment in the other one or more executable files.

[0019] Optionally, the target data includes data in a copy of the executable file.

[0020] Optionally, the executable file is stored in the memory, for example, at a third address in the memory.

[0021] Optionally, the verification data of the target data is the target data or a characteristic value of the target data.

[0022] Optionally, the management device further includes a calculation module and a writing module, wherein the calculation module is used to calculate the verification data of the target data, and the writing module is used to write the verification data of the target data to the second address through the storage controller.

[0023] Optionally, the management device also includes an exception handling module, which is used to reload the target data into the memory when the reliability information indicates that the data at the first address is inconsistent with the target data, and / or report the reliability information to the host computer connected to the processing unit.

[0024] In a third aspect, the present application provides a computer device, which may include a processing unit and a storage controller, wherein the processing unit is connected to the storage controller, the storage controller is used to connect to the memory, and the processing unit is used to execute the method described in the first aspect or any possible implementation method of the first aspect by running an executable file.

[0025] Optionally, the processing unit and the storage controller are integrated into the same chip. Alternatively, the processing unit and the storage controller are independently configured.

[0026] Optionally, the processing unit may be a processor or one or more processor cores in a processor.

[0027] Optionally, the computer device also includes the memory.

[0028] Optionally, the computer device may further include a host computer connected to the processing unit.

[0029] Optionally, the executable file is stored in a memory external to the computer device. When these instructions are decoded and executed by the processor of the computer device, part or all of the contents of the above instructions are temporarily stored in the memory within the computer device. Optionally, part of the contents of these instructions are stored in the memory external to the computer device, and the rest of the contents of these instructions are stored in the memory within the computer device.

[0030] The fourth aspect of the present application provides a computer-readable storage medium, which stores instructions (or computer-readable instructions or computer program instructions or functional programs or program codes). When these instructions are executed on a computer device, the computer device executes the method in the first aspect of the embodiment of the present application or any possible implementation of the first aspect.

[0031] The fifth aspect of the present application provides a computer program product, which, when the instructions (or computer-readable instructions or computer program instructions or functional programs or program codes) contained in the computer program product are executed by a computer device, implements the method in the first aspect of the embodiment of the present application or any possible implementation method of the first aspect.

[0032] Since the various devices provided in the embodiments of the present application can be used to execute the corresponding embodiment methods mentioned above, the technical effects that can be obtained by the various device embodiments of the present application can refer to the technical effects obtained by the corresponding method embodiments mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] FIG1 is a schematic diagram of a computer device provided in an embodiment of the present application;

[0034] Figure 2-1 schematically shows the principle diagram of sideband ECC;

[0035] FIG2-2 schematically shows the principle diagram of inline ECC;

[0036] Figure 2-3 schematically illustrates the principle of on-chip ECC;

[0037] Figure 2-4 schematically shows the principle diagram of link ECC;

[0038] FIG3 schematically illustrates a possible process of the method of the present application;

[0039] FIG4 schematically illustrates a possible process for the core to manage APP i through the operation review APP;

[0040] FIG5 schematically illustrates another process for the core to manage APP i by running the review APP;

[0041] FIG6 schematically illustrates another possible process of the method of the present application;

[0042] FIG7 schematically shows a possible structure of the application management device. DETAILED DESCRIPTION

[0043] Figure 1 is a schematic diagram of a computer device provided in an embodiment of the present application. As shown in Figure 1, the computer device includes a processor and DDR. The processor can be connected to the DDR via a DDR bus. DDR is the abbreviation for double data rate (DDR) synchronous dynamic random-access memory (SDRAM). Figure 1 takes DDR as an example of memory. DDR can be replaced with other types of memory. Accordingly, different memories may use different data buses to communicate with the processor. Therefore, the DDR bus can also be replaced with other types of data buses. The embodiment of the present application does not limit the bus type. In addition, the computer device also includes various I / O devices, and the processor can access these I / O devices via the PCIe bus.

[0044] The processor is the computing and control core of a computer device. The processor may include one or more processor cores. The processor may be a very large-scale integrated circuit. An operating system and other software programs are installed in the processor, so that the processor can access the memory and various PCIe devices. It is understandable that in the embodiment of the present invention, the core in the processor may be, for example, a central processing unit (CPU) or other specific integrated circuits (ASIC). The processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. In actual applications, a computer device may also include multiple processors.

[0045] A memory controller is a bus circuit controller that controls memory inside a computer device and is used to manage and plan data transmission from memory to the processor core (referred to as the core). Data can be exchanged between the memory and the core through the memory controller. The memory controller used to control DDR is also called a DDR controller. As shown in Figure 1, the memory controller can be integrated into the processor. Those skilled in the art will know that the memory controller can be a separate chip and connected to the core through the system bus and connected to the memory through the DDR bus, or the memory controller can also be built into the processor. The embodiments of the present invention do not limit the specific location and existence form of the memory controller. In actual applications, the memory controller can control the necessary logic to write data to the memory or read data from the memory. The memory controller can be a memory controller in a processor system such as a general-purpose processor, a dedicated accelerator, a GPU, an FPGA, or an embedded processor.

[0046] Memory is the main storage of a computer device. Memory is typically used to store various running software in the operating system, input and output data, and information exchanged with external memory. To improve processor access speed, memory needs to have a high access speed. In traditional computer system architectures, dynamic random access memory (DRAM) is commonly used as memory. The processor can access memory at high speed through a memory controller, performing read and write operations on any storage unit in the memory. In addition to DRAM, memory can also be other random access memories, such as static random access memory (SRAM). Memory can also be read-only memory (ROM). For example, read-only memory can be programmable read-only memory (PROM) or erasable programmable read-only memory (EPROM). This embodiment does not limit the amount and type of memory. In addition, the memory can be configured to have a power-saving function. This power-saving function ensures that data stored in the memory will not be lost when the system is powered off and then powered on again. Memory with power retention is called non-volatile memory.

[0047] Input / output (I / O) devices refer to hardware that enables data transmission and can also be understood as devices that interface with I / O interfaces. Common I / O devices include network cards, printers, keyboards, and mice. The processor accesses various I / O devices via the PCIe bus. A processor can be connected to no I / O devices or a larger number of them.

[0048] Serial ATA (SATA) is a type of computer mechanical hard disk, where the full name of ATA is Advanced Technology Attachment. The processor can be connected to SATA through a bus. SATA can also be used as an I / O device. The processor can access SATA through the PCIe bus. This application does not limit the type of external memory connected to the processor, for example, the external memory can be a hard disk, floppy disk, or optical disk. The processor can be connected to no external memory, or can be connected to a greater number and / or type of external memory. This application does not limit the type of SATA, and the following text takes SATA as a solid-state disk (SSD) as an example.

[0049] It should be noted that the PCIe bus mentioned above is only an example and can be replaced by other buses, such as a unified bus (UB) bus.

[0050] The host computer or baseboard management controller (BMC) can upgrade the firmware of the device, manage the operating status of the device, and troubleshoot faults. The processor can access the baseboard management controller through a management channel, which can be a PCIe bus or a bus such as USB or I2C. The baseboard management controller can maintain the program code in the memory, including upgrading or restoring it. The baseboard management controller can also control the power circuit or clock circuit in the computer device. In short, the baseboard management controller can manage the computer device in the above manner. However, the baseboard management controller is only an optional device. In some embodiments, the processor can communicate directly with the sensor to directly manage and maintain the computer device.

[0051] RAM can be the main memory of a computer device. RAM is connected to the main memory via a double data rate (DDR) bus. RAM is typically used to store various running software in the operating system, input and output data, and information exchanged with external memory. To improve processor access speed, RAM must have a high access speed. In traditional computer system architectures, dynamic random access memory (DRAM) is commonly used as RAM. The processor can access RAM at high speed through a memory controller, performing read and write operations on any storage location within the RAM.

[0052] In practical applications, memory can include one or more dual-inline memory modules (DIMMs). A DIMM can be considered a memory module, with memory chips (or memory dies) arranged on its surface. A memory rank is a group of memory chips that are arranged in parallel to meet the data bit width requirements of the CPU's memory controller. A DIMM can have one or more ranks.

[0053] Those skilled in the art will appreciate that data is stored in memory, specifically, in storage cells within a memory chip. In embodiments of the present invention, a storage cell refers to the smallest storage unit (cell) used to store data. Typically, a storage cell can store one bit of data. Of course, some storage cells can also implement multi-value storage. When DRAM is used as memory, the storage cells within the DRAM (also referred to as DRAM cells) are arranged and distributed in a matrix, which we call a memory bank or DRAM bank. In this manner, the storage cells within a memory chip can be logically divided into multiple banks, each of which can be considered a storage array consisting of multiple storage cells. Each memory cell within a bank is identified by its row address and column address, and the memory controller can locate any memory cell within the bank using corresponding row and column decoders. In embodiments of the present invention, a memory bank may also be referred to simply as a bank.

[0054] A memory chip may include a control and refresh circuit, multiple memory banks, a row address buffer, a column address buffer, a row decoder, and a column decoder. The control and refresh circuit is used to control the refresh operation of the memory cells.

[0055] During memory access, after the memory controller receives a memory access request from the core, it generates address signals and control signals based on the received memory access request and sends the generated address signals and control signals to the memory to perform memory access. Address signals may include row address signals and column address signals, and control signals may include chip select signals (CS), write enable signals (WE), column access strobe signals (CAS), and row access strobe signals (RAS).

[0056] Those skilled in the art will recognize that DRAM is a volatile memory that uses the amount of charge stored in capacitors to represent data 0s and 1s. Due to capacitor leakage, capacitors can only retain charge for a very short time. This capacitor leakage causes charge drift, which can lead to stored data errors and failures such as bit jumps. Currently, ECC is commonly used to improve this problem.

[0057] Currently, there are four methods for implementing error checking and correcting (ECC): side-band ECC, inline ECC, on-die ECC, and link ECC. The following describes each of these four methods.

[0058] Figure 2-1 schematically illustrates sideband ECC. As shown in Figure 2-1, sideband ECC is typically implemented using standard DDR (such as DDR4 and DDR5) in applications. With sideband ECC, ECC data is sent to the memory as sideband data along with the actual data. Typically, the actual data occupies 64 bits of data width, and the sideband data occupies an additional 8 bits of data width for ECC checksum storage, for a total of 72 bits of data width. The controller uses the same write (WR) and read (RD) commands to simultaneously write and read the actual data and the ECC checksum. Therefore, the sideband ECC solution does not require additional WR / RD command overhead. However, sideband ECC requires memory controller support and does not have any special requirements for the DDR core itself.

[0059] Figure 2-2 schematically illustrates inline ECC. As shown in Figure 2-2, inline ECC is commonly used in low-power (LP) DDR DRAM (LPDDR) or small systems with only one or two standard DRAM chips. In these scenarios, using sideband ECC wastes bandwidth and increases DRAM costs. For example, LPDDR has a fixed channel width (for example, LPDDR5 / 4 / 4X has a 16-bit channel width). If "64+8" sideband ECC is used, 16 bits of data width are required to provide a channel for the 8-bit ECC check code, resulting in bandwidth waste. For scenarios with low DRAM capacity requirements, such as when the memory controller controls a single DRAM chip, sideband ECC cannot be used. Inline ECC provides a suitable solution for these scenarios. Inline ECC allocates dedicated storage space for the ECC check code within the same DRAM chip. Both data and the ECC check code are stored within the same DRAM chip. The ECC check code and data use the same channel, so the WR / RD commands for the ECC check code and the WR / RD commands for the data are typically separate. Therefore, this solution incurs additional WR / RD command overhead. Furthermore, inline ECC requires support from the memory controller and has no special requirements for the DDR / LPDDR chips themselves.

[0060] Figure 2-3 schematically illustrates on-chip ECC. As shown in Figure 2-3, on-chip ECC is currently primarily used in DDR5. On-chip ECC is an advanced registration, admission, and status (RAS) function supported by DDR5 systems. During writes, the core transmits the data to the DRAM via the memory controller. The DRAM internally calculates the ECC of the data and stores it. During reads, the DRAM performs error detection and correction internally, and only the corrected data is sent to the core via the memory controller. Therefore, on-chip ECC can only correct single-bit errors in the memory array and cannot correct errors within the DDR channel. It is typically used in conjunction with sideband ECC. On-chip ECC is implemented internally in the DDR / LPDDR chip and is closed-loop, requiring no memory controller support.

[0061] Figure 2-4 schematically illustrates Link ECC. As shown in Figure 2-4, Link ECC is currently primarily used in LPDDR5 to protect LPDDR5 links or channels from faulty cells. During a write, the memory controller transmits the data to be written and a generated ECC check code (denoted as Check Code 1) to the DRAM. The DRAM internally generates an ECC check code (denoted as Check Code 2) based on the received data and compares ECC Check Code 1 with ECC Check Code 2. If the two are identical, the DRAM writes the data. If they differ, the DRAM performs error correction on the data before writing it (writing the ECC check code is not required at this time). During a read, the DRAM regenerates a new ECC check code (denoted as Check Code 3) based on the data to be read and transmits the data and Check Code 3 to the memory controller. The memory controller generates an ECC check code (denoted as Check Code 4) based on the received data and compares Check Code 3 with Check Code 4. If the two differ, the memory controller can perform error correction on the data before passing it to the core. Link ECC is often used in conjunction with inline ECC to provide more complete protection. Link ECC requires support from both the memory controller and DDR / LPDDR to be implemented.

[0062] An analysis of existing memory management methods reveals that the aforementioned ECC has some limitations. For example, one limitation of existing ECC (referred to as limitation 1) is that ECC is an in-place error correction technology. The memory controller only checks the reliability of the corresponding data after receiving a read request from the core. When data errors occur, the memory controller must perform in-place error correction. Furthermore, due to ECC's limited repair capabilities, it can only resolve single-bit errors, not multi-bit errors. When the data indicated by the read request contains multi-bit errors, the read will fail. For example, another limitation of existing ECC (referred to as limitation 2) is that sideband ECC requires additional memory chips, increasing memory costs and restricting its implementation on some cost-sensitive devices. For example, another limitation of existing ECC (referred to as limitation 3) is that ECC technology requires the memory controller and / or memory chips to support ECC features.

[0063] In order to solve the above technical problems, the present application proposes that the core manage the data in the memory by executing program code. This is conducive to identifying whether the data stored in the memory has changed due to memory reliability issues through software means. Compared with ECC technology that requires the controller and / or DDR chip to support ECC features, the memory management method provided by this application does not require changes to the design of the hardware circuit, which is conducive to reducing the requirements of memory management on DDR hardware and expanding the use scenarios of the memory management method. In some scenarios, the ECC particles of the CPU small system can be removed to reduce hardware costs.

[0064] For ease of description, the program code will be referred to as the review application (APP) in the following text. The application can refer to an executable file. The executable file can also be called a target file, which generally includes a code segment and a data segment. Among them, the code segment mainly contains the instructions of the program, which is generally readable and executable, but generally not writable. The data segment mainly stores various global variable data (called the program dynamic data segment) and / or static data (called the program static data segment) to be used in the code segment. The data segment is generally readable, writable, and executable.

[0065] FIG3 schematically illustrates a possible process of the method of the present application. As shown in FIG3 , the method may include:

[0066] S301, compile the device to generate a file package;

[0067] During the version building phase, the compiler can compile the source code into n apps, where n is a positive integer. The compiler can then calculate the hash value of each app to form a hash list. The compiler can then package the n apps and the hash list into a file package.

[0068] It should be noted that the hash value of the APP is used as a baseline (or verification data) to verify whether the APP has changed. This application does not limit the type of verification data of the APP. For example, the hash value of the APP can be replaced with other types of characteristic values, such as a cyclic redundancy check (CRC).

[0069] S302, the core writes the file package into the SSD;

[0070] The core may obtain the file package and then write the file package to the SSD. For example, the core may receive the file package from the host computer and then write the file package to the SSD. Alternatively, the SSD may already store the file package before being inserted into the computer device. Accordingly, the method may not include S302.

[0071] S303, the core writes n APPs and hash list in the file package into DDR;

[0072] During the device startup phase or the device operation phase, the core can read the file package in the SSD and write the n APPs and hash list in the file package into the DDR.

[0073] S304, the core manages n APPs in the DDR by running the review APP stored in the DDR;

[0074] After n apps and hash lists are written to the DDR, the core can manage n apps by running the audit app stored in the DDR. For example, the app's hash value can be used to determine whether the app stored in the DDR is consistent with the original app or whether there are any differences.

[0075] S304 takes the core managing n APPs as an example. Optionally, the core can manage a part of the n APPs.

[0076] The following describes a specific method in which the core manages APP i by running the review APP stored in the DDR, where i is a positive integer less than or equal to n.

[0077] Figure 4 schematically illustrates a possible process for the core to manage APP i by running the review APP. In S304 described above, the core's management of n APPs can be understood with reference to the process shown in Figure 4. The core can execute the process shown in Figure 4 for each of the n APPs in serial or parallel manner.

[0078] Assume that APP i and the hash value of APP i are stored at the first and second addresses in the DDR, respectively, and the review APP is stored at the third address in the DDR. Since the data stored in the DDR may change due to a DDR failure (e.g., a bit jump), in order to identify whether APP i stored at the first address in the DDR has changed, referring to FIG. 4 , the core can execute S401 to S402 by running the review APP.

[0079] S401, reading data at a first address and data at a second address from a memory respectively through a memory controller;

[0080] When the core is running the review APP, the memory controller can read data at a first address and data at a second address from the memory respectively. The first address and the second address can be consecutive addresses or discontinuous addresses.

[0081] This application does not limit the specific manner in which the core reads the data at the first address and the data at the second address from the memory via the memory controller. For example, the core may send a read request to the memory controller, the read request may carry the logical address of the first address and the logical address of the second address, respectively. The memory controller may perform address translation on the logical addresses in the read request to determine the first address and the second address, and then read the data at the first address and the data at the second address from the DDR.

[0082] This application does not limit the number of read requests that a core may send to a memory controller. For example, a core may send a single read request to a memory controller, where the read request is used to instruct the read of data at a first address and a second address in the DDR. Alternatively, a core may send multiple read requests to a memory controller, where at least one read request is used to instruct the read of data at a first address in the DDR, and the other read requests are used to instruct the read of data at a second address in the DDR.

[0083] This application does not limit the manner in which the core determines the hash value of APP i from the hash list. For example, each hash value recorded in the hash list is associated with the file location and / or file name of the corresponding APP. Alternatively, the core calculates a characteristic value of the identification of the APP corresponding to each hash value in the hash list, and determines the storage location of the hash value in the storage space allocated for the hash list according to the characteristic value corresponding to the hash value. When APP i needs to be reviewed, the characteristic value of the identification of APP i can be calculated, and the hash value of APP i can be read from the storage space based on the characteristic value.

[0084] S402. Determine reliability information of the data at the first address using the data at the second address, where the reliability information indicates whether the data at the first address is consistent with or inconsistent with APP i.

[0085] After the core reads the data at the first address and the data at the second address, it can use the data at the second address to determine reliability information of the data at the first address, where the reliability information indicates whether the data at the first address is consistent with or inconsistent with APP i.

[0086] For example, the core may calculate a hash value of the data at the first address, then compare the hash value with the data at the second address, and determine reliability information for the data at the first address based on the comparison result. Optionally, if the hash value is the same as the data at the second address, the reliability information may indicate that the data at the first address is consistent with APP i. If the hash value is different from the data at the second address, the reliability information may indicate that the data at the first address is inconsistent with APP i.

[0087] Optionally, after the core uses the data in the second address to determine that the data in the first address is inconsistent with APP i, the information of the first address can be recorded in the reliability information. For example, the information of the first address can be used to indicate at least one of the channel identifier, rank identifier, bank identifier, row identifier and column identifier of the storage unit corresponding to the first address.

[0088] Optionally, when the reliability information indicates that the data at the first address is inconsistent with APP i, the core may further execute S403 to S404.

[0089] S403, reload APP i into the memory;

[0090] When the reliability information indicates that the data at the first address is inconsistent with APP i, the core can reload APP i into the memory. For example, if APP i is stored in a storage device other than the memory (e.g., an SSD), the core can read APP i from the SSD and then write APP i to the memory. Since the inconsistency between the data at the first address and APP i may be due to a fault in the first address and / or the second address (e.g., a bit jump), the core can write APP i to an address other than the first address and the second address. Optionally, the core can also reload the hash value of APP i into the memory.

[0091] S404: Report reliability information to a host computer.

[0092] When the reliability information indicates that the data at the first address is inconsistent with APP i, the core can report the reliability information to the host computer. Upon receiving the reliability information, the host computer can alert the user based on the reliability information, indicating a DDR failure and / or APP i anomaly. Optionally, upon receiving the reliability information, the host computer can restart the computer device to reload the APPs into the DDR.

[0093] The core may execute all or part of the steps in S403 to S404 by running the review APP or other APPs other than the review APP, or execute all or part of the steps in S403 to S404 through hardware logic circuits.

[0094] In some examples, when the reliability information indicates that the data at the first address is inconsistent with APP i, the core may execute S403 instead of S404. In some examples, when the reliability information indicates that the data at the first address is inconsistent with APP i, the core may execute S404 instead of S403.

[0095] In the above, the hash value of APP i is calculated by other devices other than the core (such as a compilation device) as an example. Optionally, the hash value of APP i can be calculated by the core and then written to the second address through the storage controller.

[0096] Figure 5 schematically illustrates another process of managing APP i by running an audit APP. Referring to Figure 5 , the process may include S501 to S507.

[0097] S501, generating a baseline value of APP i;

[0098] The core can calculate the characteristic value of APP i stored in the DDR, which can be used as a baseline value for measuring whether a bit jump error occurs in the DDR. This application does not limit the method used by the core to calculate the characteristic value. For example, the core can choose a hash algorithm or a CRC algorithm to calculate the characteristic value of APP i. Algorithms with hardware acceleration capabilities (such as hash algorithms) are beneficial to improving computing performance. The following article takes the baseline value of APP i calculated by the core using a hash algorithm as an example.

[0099] S502, importing the baseline value into DDR;

[0100] After the core calculates the baseline value of APP i, it can write the baseline value into DDR.

[0101] After both APP i and APP i's baseline value are written to the DDR, the core can use APP i's baseline value to manage APP i, for example, by executing S503-S507. This application does not limit the specific timing when the core manages APP i by running the review APP stored in the DDR. The triggering mechanism of step S503 will be described later. As an example, the core can trigger step S503 on a regular basis. For example, the core can add a scheduled task to a timer, set the duration of the time slice, and then execute S503 when the time slice expires.

[0102] S503. Calculate the hash value of the data in the memory space where APP i is located;

[0103] After both APP i and the baseline values ​​of APP i are written into DDR, the core can read the data from the memory space where APP i is located and calculate the hash value of the data.

[0104] S504: Determine whether the calculated hash value is the same as the baseline value. If so, execute S505; if not, execute S506.

[0105] The core may read the baseline value from the DDR and compare the calculated hash value with the baseline value. If the calculated hash value is the same as the baseline value, the core may execute S505; if not, the core may execute S506.

[0106] S505: Determine whether verification of APP i stored in DDR is successful;

[0107] When the calculated hash value is the same as the baseline value, it can be considered that APP i stored in the DDR has not changed, it can be determined that there is no error in APP i stored in the DDR, and it can be considered that the verification of APP i is successful.

[0108] Optionally, the scheduled task may be a multi-cycle task. When the time slice is exhausted, the time slice may be set again. After the time slice is exhausted, the core may execute S503 again.

[0109] S506: Determine that verification of APP i stored in DDR fails;

[0110] When the calculated hash value is different from the baseline value, it can be considered that APP i stored in the DDR has changed, it can be determined that an error has occurred in APP i stored in the DDR, and it can be considered that the verification of APP i has failed.

[0111] S507: Execute subsequent actions.

[0112] After S506, the core recognizes that a DDR error has occurred and can perform subsequent actions based on the recognition result, for example, at least one of an alarm, program reload, and restart. Step S507 can be understood with reference to S403 and / or S404.

[0113] Alternatively, the core may calculate the hash value of APP i by examining other APPs or hardware logic circuits other than the APP.

[0114] In the above, the verification data used to verify the APP is the characteristic value of the APP. Optionally, the verification data of the APP can be the APP itself. That is, in addition to writing APP i to the first address of the DDR, the core also writes APP i to the second address of the DDR, and uses the data of the second address as the verification data of the data of the first address. This application does not limit the method by which the core uses the data of the second address to determine whether the data of the first address is consistent with APP i. For example, the core can calculate the characteristic value of the data of the first address and the characteristic value of the data of the second address respectively, and determine whether the data of the first address is consistent with APP i by comparing the two characteristic values.

[0115] The core manages the apps in memory by executing the review app. This helps identify, through software means, whether the data stored in the memory has changed due to memory reliability issues. Compared to ECC technology, which requires the controller and / or DDR chip to support ECC features, the memory management method provided by this application does not require changes to the hardware circuit design, which helps to reduce the requirements of memory management on DDR hardware and expand the use scenarios of the memory management method. In some scenarios, the ECC particles of the CPU small system can be removed to reduce hardware costs.

[0116] In existing ECC implementations, the memory controller and / or DDR only verifies data based on a read request when the core reaches APP i. A multi-bit error in APP i, as requested by a read request, can cause a read failure or even an abnormal device reboot.

[0117] The APP i managed by the core during the process of running the review APP is an APP other than the review APP. This is beneficial for the core to check the code segments that have not been called when running the review APP, and is beneficial for identifying errors in the code segments that have not been called in advance, so as to facilitate pre-processing (such as reloading, etc.) before they are called. Compared with the ECC method that only checks the code segment being called, it is beneficial to avoid abnormal restart of the device due to multi-bit errors in the code segment being called.

[0118] Optionally, when the review APP is in the running state, the core can select an APP in the DDR that is not in the running state for management by running the review APP. Accordingly, when the core manages APP i by running the review APP, APP i is not in the running state and can be in the ready state or the blocked state.

[0119] This application does not limit the specific timing when the core manages APP i by running the review APP stored in DDR.

[0120] Optionally, while the core is writing the n APPs and the hash list in the file package into the DDR, it can manage the APPs (eg, APP i) that have been written into the DDR and perform a check on the APP i that has been stored in the DDR.

[0121] Optionally, the core can periodically manage and check apps already stored in the DDR. For example, the core can manage apps through a scheduled task. Specifically, the core can add a memory management task to the timer and set a corresponding time slice. When the time slice expires, the core manages apps already stored in the DDR.

[0122] APP i stored in DDR may change due to DDR reliability issues. During the process of running and reviewing APP, the core manages the APP stored in the memory regularly / timely, which is helpful for identifying changes caused by DDR reliability issues before calling APP i. This gives the computer equipment a certain operating space to resolve DDR anomalies, thereby helping to reduce or avoid major problems such as abnormal device restarts.

[0123] Optionally, the core can set time slices of different sizes for different apps. For example, for two apps of different importance, the core can set a shorter time slice for the more important app and a longer time slice for the less important app. In this way, the core will check the more important app more frequently, which is conducive to ensuring the reliability of the app. This application does not limit the size of the time slice, and users can flexibly set the size of the time slice according to their needs.

[0124] In the above, the hash value in the hash list is used as an example to represent the hash value of all the data in the APP. Optionally, the hash list can include the hash value of a portion of the data in the APP. Accordingly, in S304, the core can use the hash value of this portion of data to determine whether the portion of data stored in the DDR is consistent with the portion of data in the original APP. Accordingly, in the process shown in Figure 4 or Figure 5, APP i can be replaced with a portion of the data in APP i, and the hash value of APP i can be replaced with the hash value of this portion of data.

[0125] As previously mentioned, executable files generally include code segments and data segments. The data segment further includes a program dynamic data segment and a program static data segment. Table 1 schematically illustrates the impact of line problems in each segment on APP i and the current response measures.

[0126] Table 1

[0127] Optionally, the hash list may include hash values ​​of program code segments in APP i. Accordingly, the core may use the hash values ​​to determine whether the program code segments in APP i stored at the first address are incorrect, i.e., to identify errors in scenario 1. If an error occurs, APP i may be reloaded, which helps avoid CPU resets due to program code segment errors during APP i's operation. Optionally, the hash list may include hash values ​​of program static data segments and / or program dynamic data segments in APP i. Accordingly, the core may use the hash values ​​to determine whether the corresponding data segments in APP i stored at the first address are incorrect, i.e., to identify errors in scenario 2 and / or scenario 3. If an error occurs, APP i may be reloaded, which helps avoid re-execution of APP i due to data segment errors during its operation.

[0128] To ensure the data integrity of the review app, a backup of the review app can also be stored in the DDR. For ease of distinction, the review app mentioned above is referred to as review app_A, and the backup of the review app is referred to as review app_B. Referring to the existing A / B mechanism, when review app_A fails, the core can manage the data stored in the DDR by running review app_B. The core's management of data stored in the DDR by running review app_B is similar to the method described above for managing data stored in the DDR by running review app_A.

[0129] Before the review APP_A fails, the core may not run the review APP_B, or the core may run the review APP_A and the review APP_B in parallel to manage the data stored in the DDR.

[0130] When the core is running the review APP_A and the review APP_B in parallel, it can use a load balancing method to manage different data stored in the DDR. Alternatively, when the core is running the review APP_A and the review APP_B in parallel, it can manage the same data stored in the DDR, and the upper computer can make decisions based on the reliability information reported by the review APP_A and the review APP_B respectively. For example, when the reliability information reported by both indicates that the APP i stored in the DDR has an error, the upper computer determines that the APP i stored in the DDR has an error, restarts the device, or reloads APP i. When only one reliability information reported by the review APP_A is received, the upper computer determines that the APP i stored in the DDR is suspected to have an error, and can alert the user.

[0131] Since the core manages APP i during the execution of the review APP other than the review APP itself, it can optionally manage review APP_B while running review APP_A. The core's management of review APP_B can refer to the aforementioned process for managing APP i. For example, the core can execute the process shown in Figure 4 or Figure 5 for review APP_B. This allows the core to identify errors in review APP_B stored in the DDR and take appropriate action before running review APP_B, thus ensuring the proper operation of review APP_B.

[0132] Figure 6 schematically illustrates another possible process of the method of the present application. The difference between this process and the process shown in Figure 3 is that two review APPs (i.e., review APP_A and review APP_B) can be stored in the DDR, and the core can manage review APP_B during the operation of review APP_A, and the core can manage review APP_A during the operation of review APP_B. Accordingly, in addition to the hash values ​​in the hash list in the file package, the hash list ' can also save the hash value of review APP_A and the hash value of review APP_B. The process of the core managing review APP_B can refer to the process of the core managing APP i mentioned above. For example, the core can execute the process shown in Figure 4 or Figure 5 on review APP_B. The process of the core managing review APP_A during the operation of review APP_B can refer to the process of the core managing APP i mentioned above. For example, the core can execute the process shown in Figure 4 or Figure 5 on review APP_A.

[0133] In the above, the hash values ​​of different APPs are stored in the hash list as an example. This application does not limit the storage method of the hash values ​​of multiple APPs, as long as the core can read the hash value of APP i.

[0134] In the above, the example of the core managing the APP or part of the data in the APP stored in the DDR by running the review APP is taken. Optionally, the core can manage other types of data other than the APP stored in the DDR by running the review APP.

[0135] In the above, taking the example of the core managing the data in the DDR by running the review APP, this application does not limit the type of memory managed by the core by running the review APP. For example, the core can manage the data in the flash memory (flash) and / or hard disk by running the review APP.

[0136] In the above, the review APP is stored in the managed memory as an example. Optionally, the review APP can be stored in other memories other than the managed memory.

[0137] After the core reads the executable file (e.g., the review app), it can generate the management device 7 shown in Figure 7. This management module 7 can be used to execute S304 or all or part of the steps shown in Figure 4 or all or part of the steps shown in Figure 5. Referring to Figure 7, the management device 7 may include a reading module 701 and a determination module 702. The functions of each software functional module are described below.

[0138] The reading module 701 can be used to read data at a first address and data at a second address from a memory through a memory controller, wherein the first address of the memory is used to store target data and the second address is used to store verification data of the target data.

[0139] The determination module 702 may be configured to determine reliability information of the data at the first address using the data at the second address, where the reliability information indicates whether the data at the first address is consistent with or inconsistent with the target data.

[0140] Taking the memory as DDR, the first address for storing APP i, and the second address for storing the hash value of APP i as an example, the reading module 701 can be used to execute S401, and the determining module 702 can be used to execute S402.

[0141] Optionally, the read module 701 can be used to read the data at the first address and the data at the second address from the memory through the storage controller after the first data and the verification data of the first data are written to the memory. Alternatively, the read module 701 can periodically read the data at the first address and the data at the second address from the memory through the storage controller, for example, by managing the APP through a scheduled task. Specifically, the core can add a memory management task to the timer and set the corresponding time slice size. When the time slice is exhausted, the core manages the APP already stored in the DDR.

[0142] Optionally, the target data is data in one or more executable files other than the executable file, for example, data in APP i.

[0143] Optionally, the target data is a program code segment or a program static data segment or a program dynamic data segment in one or more other executable files.

[0144] Optionally, the target data includes data in a copy of the executable file.

[0145] Optionally, the executable file is stored in the memory, for example, at a third address in the memory.

[0146] Optionally, the verification data of the target data is the target data or a characteristic value of the target data.

[0147] Optionally, the management device further includes a calculation module 703 and a writing module 704, wherein the calculation module 703 is used to calculate the verification data of the target data, and the writing module 704 is used to write the verification data of the target data to the second address through the storage controller.

[0148] Optionally, the management device further includes an exception handling module 705, which is used to reload the target data into the memory when the reliability information indicates that the data at the first address is inconsistent with the target data, and / or report the reliability information to the host computer connected to the processing unit.

[0149] Those skilled in the art will understand that when software is used to implement the various aspects of the embodiments of the present application, or the possible implementation of each aspect, the above-mentioned various aspects, or the possible implementation of each aspect can be implemented in whole or in part in the form of a computer program product. A computer program product refers to instructions (or computer-readable instructions or computer program instructions or functional programs or program codes) stored in a computer-readable medium. When these instructions are loaded and executed on a computer, the process or function described in the embodiments of the present application are generated in whole or in part.

[0150] The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. Computer-readable storage media include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatuses, or any suitable combination thereof. For example, the computer-readable storage medium may be a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or a portable compact disc read-only memory (CD-ROM).

[0151] The terms "first", "second", "third", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances. This is merely a way of distinguishing objects with the same properties when describing the embodiments of this application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device that includes a series of units is not necessarily limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or devices. The term "plurality" appearing in the embodiments of this application refers to two or more.

[0152] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the scope of the present invention. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present invention, the present invention is also intended to include these modifications and variations.

Claims

1. A management method for a memory, characterized in that, The memory is connected to the processing unit through a storage controller. A first address and a second address in the memory are respectively used to store target data and check data of the target data. The processing unit implements the method by running an executable file. The method includes: Reading the data at the first address and the data at the second address from the memory respectively through the storage controller; Determining reliability information of the data at the first address by using the data at the second address, where the reliability information is used to indicate whether the data at the first address is consistent or inconsistent with the target data.

2. The method according to claim 1, wherein The step of reading the data at the first address and the data at the second address from the memory respectively through the storage controller includes: Periodically reading the data at the first address and the data at the second address from the memory respectively through the storage controller.

3. The method according to claim 1 or 2, characterized in that, The target data is data in one or more other executable files other than the executable file.

4. The method according to claim 3, wherein The target data is a program code segment or a program static data segment or a program dynamic data segment in the one or more other executable files.

5. The method according to claim 3 or 4, characterized in that, The target data includes data in a copy of the executable file.

6. The method according to claim 5, characterized in that, The executable file is stored at a third address in the memory.

7. The method according to any one of claims 1-6, characterized in that, The check data of the target data is the target data or a feature value of the target data.

8. The method according to claim 7, wherein The method further includes: Calculating the check data of the target data; Writing the check data of the target data to the second address through the storage controller.

9. The method according to any one of claims 1 to 8, characterized in that When the reliability information indicates that the data at the first address is inconsistent with the target data, the method further includes: Reloading the target data into the memory; and / or Reporting the reliability information to a host computer connected to the processing unit.

10. The method according to any one of claims 1-9, characterized in that, The memory is a memory or a flash memory or a hard disk.

11. A management device for a memory, characterized in that, The memory is connected to the processing unit through a storage controller. A first address and a second address in the memory are respectively used to store target data and check data of the target data. The processing unit generates the management device by running an executable file. The management device includes: A reading module, configured to read the data at the first address and the data at the second address from the memory respectively through the storage controller; A determining module, configured to determine reliability information of the data at the first address by using the data at the second address, where the reliability information is used to indicate whether the data at the first address is consistent or inconsistent with the target data.

12. A computer device, characterized in that, It includes a processing unit and a storage controller. The processing unit is connected to the storage controller, and the storage controller is used to connect to a memory. The processing unit is configured to execute the method according to any one of claims 1 to 10 by running an executable file.

13. The computer device according to claim 12, wherein, The processing unit and the storage controller are integrated on the same chip.

14. The computer device according to claim 12 or 13, characterized in that, The computer device further includes the memory.

15. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium. When the program code is executed by the computer device, the method according to any one of claims 1 to 10 is implemented.

16. A computer program product, characterized in that, When the program code included in the computer program product is executed by a computer device, the method described in any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Memory management method and related equipment

    CN120233939A

  • Safety verification method and device, electronic equipment and storage medium

    CN112182584A

  • Test method and device based on memory, electronic equipment and medium

    CN116631489A

  • ECC verification management method and device of vehicle-mounted chip and storage medium

    CN116954985A

  • Independent Management of Data and Parity Logical Block Addresses

    US20140365821A1