Bad block detection method and device of memory, equipment and storage medium

By utilizing idle peripheral registers to detect bad blocks in memory during server operation, the problem of the inability to dynamically detect bad blocks in existing technologies is solved, enabling real-time processing of bad blocks and ensuring stable server operation and data security.

CN121075397APending Publication Date: 2025-12-05EDGELESS SEMICON CO LTD OF ZHUHAI +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511031855.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing technologies cannot dynamically detect and handle bad blocks in memory during server operation, leading to potential system instability and data loss risks.

Method used

By using idle peripheral registers as temporary data storage carriers during server operation, the running data of the memory area is transferred to the peripheral registers in a predetermined detection order to perform bad block detection. Bad block areas are identified through data verification and read/write tests. After repair, the data is returned to the storage area.

Benefits of technology

It enables real-time detection and handling of bad blocks without interrupting normal server operation, preventing data loss and system crashes, optimizing detection efficiency, extending the stable operation cycle of the server, and improving the system's resilience and data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121075397A_ABST
    Figure CN121075397A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a bad block detection method, device and equipment for a memory and a storage medium. An idle peripheral register is used as a temporary data storage carrier in the running process of a server, so that the data storage efficiency is improved on the premise of not interrupting the normal running of the server; and according to the determined detection sequence, sequentially transmitting the running data of each storage area to an idle peripheral register so as to complete bad block detection on each storage area. According to the invention, the server is ensured to work continuously and normally during the detection period, and system shutdown or performance reduction caused by detection is avoided; by reasonably planning the detection sequence of the storage area and utilizing the idle peripheral resources, the efficiency of the detection process is optimized, the influence of the detection operation on the normal performance of the server is minimized, and both the system stability and the operation efficiency are considered.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a bad block detection method, device and equipment of memory and storage medium. BACKGROUND

[0002] In the server, the memory as a key storage module, its reliability directly affects the stability and performance of the system; the bad block in the memory may cause data loss or system crash, the existing solution usually carries out comprehensive test and records the bad block at the system startup, but cannot dynamically detect and handle the new bad block in the real-time running, resulting in potential instability of the system. SUMMARY

[0003] In view of the above problems, the present application is proposed to provide a bad block detection method, device and equipment of memory and storage medium to overcome the above problems or at least partially solve the above problems.

[0004] In order to solve the above problems, the present application discloses a bad block detection method of memory, applied to a server, the server comprising a peripheral register and a memory; the memory comprises a plurality of storage areas, and the method comprises:

[0005] During the running of the server, determining the idle peripheral register in the server;

[0006] Determining the detection order of the plurality of storage areas;

[0007] According to the detection order, the running data of the plurality of storage areas is transmitted to the peripheral register in sequence;

[0008] After the running data of each storage area is transmitted to the peripheral register, detecting whether the each storage area is a bad block area.

[0009] Optionally, the determination of the detection order of the plurality of storage areas comprises:

[0010] Obtaining the historical bad block address of the memory;

[0011] According to the historical bad block address of the memory, the detection order of the plurality of storage areas is determined.

[0012] Optionally, the detection of whether the each storage area is a bad block area comprises:

[0013] Writing verification data in the storage area, and reading the data of the storage area;

[0014] determining whether the data of the storage area is consistent with the verification data, and if the data of the storage area is not consistent with the verification data, determining that the storage area is a bad block area.

[0015] Optionally, the server further comprises:

[0016] storing the detection results of the plurality of storage areas into a memory management bitmap.

[0017] allocating memory to each storage area in the memory according to the memory management bitmap when memory needs to be allocated to the memory.

[0018] Optionally, the allocating memory to each storage area in the memory according to the memory management bitmap comprises:

[0019] determining a bad block area in the memory according to the memory management bitmap, and not allocating memory to the bad block area.

[0020] Optionally, the server further comprises a direct memory access controller, and the sequentially transmitting the running data of the plurality of storage areas to the peripheral registers according to the detection sequence comprises:

[0021] sequentially transmitting the running data of the plurality of storage areas to the peripheral registers according to the detection sequence by the direct memory access controller.

[0022] Optionally, the determining the idle peripheral register comprises:

[0023] obtaining bitmap values and interrupt states of each peripheral register in the server;

[0024] determining the idle peripheral register from the peripheral registers according to the bitmap values and the interrupt states of the peripheral registers.

[0025] Optionally, the server further comprises a timer, and the method further comprises:

[0026] setting a time interval by the timer;

[0027] the determining the idle peripheral register in the server comprises:

[0028] periodically determining the idle peripheral register in the server according to the time interval.

[0029] Optionally, the server further comprises:

[0030] If it is determined that the target storage area is a bad block area, the target storage area is repaired, and the data of the target storage area stored in the peripheral register is returned to the target storage area after the target storage area is repaired.

[0031] The application further discloses a bad block detection device of a memory, which is applied to a server.

[0032] The first determination module is configured to determine an idle peripheral register in the server during running of the server.

[0033] The second determination module is configured to determine a detection sequence of the plurality of storage areas.

[0034] The transmission module is configured to sequentially transmit running data of the plurality of storage areas to the peripheral register according to the detection sequence.

[0035] The detection module is configured to detect whether each storage area is a bad block area after the running data of the each storage area is transmitted to the peripheral register.

[0036] Optionally, the second determination module comprises:

[0037] The acquisition sub-module is configured to acquire a historical bad block address of the memory.

[0038] The first determination sub-module is configured to determine the detection sequence of the plurality of storage areas according to the historical bad block address of the memory.

[0039] Optionally, the detection module comprises:

[0040] The writing sub-module is configured to write verification data into the storage area and read data of the storage area.

[0041] The judgment sub-module is configured to judge whether the data of the storage area is consistent with the verification data, and if the data of the storage area is not consistent with the verification data, it is determined that the storage area is a bad block area.

[0042] Optionally, the application further comprises:

[0043] The storage module is configured to store detection results of the plurality of storage areas into a memory management bitmap.

[0044] The allocation module is configured to allocate memory to each storage area in the memory according to the memory management bitmap when it is necessary to allocate memory to the memory.

[0045] Optionally, the allocation module comprises:

[0046] The second determining sub-module is configured to determine a bad block area in the memory according to the memory management bitmap, and not allocate memory to the bad block area.

[0047] Optionally, the server further comprises a direct memory access controller, and the transmission module comprises:

[0048] The transmission sub-module is configured to sequentially transmit the running data of the plurality of storage areas to the peripheral registers according to the detection sequence through the direct memory access controller.

[0049] Optionally, the first determining module comprises:

[0050] The third determining sub-module is configured to obtain bitmap values and interrupt states of each peripheral register in the server.

[0051] The fourth determining sub-module is configured to determine an idle peripheral register in the each peripheral register according to the bitmap values and the interrupt states of the each peripheral register.

[0052] Optionally, the server further comprises a timer, and further comprises:

[0053] The setting module is configured to set a time interval through the timer.

[0054] The first determining module comprises:

[0055] The fifth determining sub-module is configured to periodically determine an idle peripheral register in the server according to the time interval.

[0056] Optionally, the server further comprises:

[0057] The repairing module is configured to, if the target storage area is determined to be a bad block area, repair the target storage area, and return data of the target storage area stored in the peripheral register to the target storage area after the target storage area is repaired.

[0058] Correspondingly, an electronic device is disclosed, which comprises a processor, a memory, and a computer program stored in the memory and capable of running on the processor, and the computer program is executed by the processor to implement each step of the bad block detection method for the memory.

[0059] Correspondingly, a computer readable storage medium is disclosed, which stores a computer program, and the computer program is executed by a processor to implement each step of the bad block detection method for the memory.

[0060] Embodiments of the present application include the following advantages:

[0061] The application discloses a bad block detection method, device and equipment of a memory and a storage medium, and the application uses an idle peripheral register as a temporary data storage carrier during server operation, and transmits the operation data of each storage area to the idle peripheral register in a determined detection sequence without interrupting the normal operation of the server, so that the bad block detection of each storage area is completed; the application ensures the continuous normal operation of the server during detection, avoids system downtime or performance decline caused by detection, can discover and process newly added bad blocks in the operation process in real time, prevents the bad blocks from causing data loss, system crash and other risks due to untimely detection, dynamically maintains the health status of the memory, prolongs the stable operation period of the server, and improves the anti-risk ability and data security of the overall system; the application optimizes the efficiency of the detection process by reasonably planning the detection sequence of the storage area and using the idle peripheral resources, ensures that the influence of the detection operation on the normal performance of the server is minimized, and balances the system stability and operation efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0062] Figure 1 is a step flow chart of a bad block detection method embodiment of the application;

[0063] Figure 2 is a flowchart of a bad block detection method embodiment of the application;

[0064] Figure 3 is a structural block diagram of a bad block detection device embodiment of the application. DETAILED DESCRIPTION

[0065] In order to make the above-mentioned purposes, features and advantages of the application more obvious and easy to understand, the application will be further described in detail below with reference to the drawings and specific embodiments.

[0066] One of the core ideas of the embodiments of the present application is that the present application can use the idle peripheral register as a temporary data storage carrier during the server running process, and under the premise of not interrupting the normal operation of the server, the running data of each storage area is transmitted to the idle peripheral register in a determined detection order, so that the bad block detection of each storage area is completed; the present application ensures the continuous normal work of the server during the detection period, avoids the system downtime or performance degradation caused by detection, and can also discover and process the newly added bad blocks in the running process in real time, prevent the bad blocks from causing data loss, system crash and other risks due to not being detected in time, thereby dynamically maintaining the health status of the memory, prolonging the stable operation period of the server, and improving the anti-risk ability and data security of the overall system; the present application optimizes the efficiency of the detection process by reasonably planning the detection order of the storage area and using the idle peripheral resources, and ensures that the detection operation minimizes the impact on the normal performance of the server, and balances the system stability and operation efficiency.

[0067] Referring to Figure 1 , a step flow chart of a bad block detection method of a memory of the present application is shown, the method is applied to a server, the server includes a peripheral register and the memory; the memory includes a plurality of storage areas, and the method can include:

[0068] Step 101, during the server running process, determine the idle peripheral register in the server.

[0069] In the embodiments of the present application, when the server is running, the system resources are in a dynamic allocation state. To perform memory bad block detection, it is necessary to ensure that the peripheral register used for data transmission and detection is not occupied by other tasks, at this time, the idle peripheral register can be identified according to the usage state identifier of the register or by sending a probe signal to all registers and analyzing the response, to prepare a data processing channel for subsequent detection work.

[0070] It should be noted that the memory is divided according to the storage medium, including semiconductor memory, magnetic surface memory, optical memory, etc.; the semiconductor memory uses semiconductor devices as the storage medium, such as random access memory (RAM) and read-only memory (ROM), among which RAM can be divided into dynamic random access memory (DRAM) and static random access memory (SRAM), DRAM needs to be refreshed periodically to maintain data, SRAM does not need to be refreshed, and has faster speed but higher cost; the magnetic surface memory stores information by using the magnetization state of magnetic materials, and the common one is a hard disk (HDD), which reads and writes data by a magnetic head on a rotating disk surface; the optical memory uses optical signals as the storage and reading medium, such as optical discs (CD, DVD, Blu-ray disc, etc.), which stores and reads data by irradiating the concave and convex points on the surface of the optical disc with a laser beam.

[0071] According to the role in the computer system, it can be divided into main memory (memory) and auxiliary memory (external storage). The main memory exchanges data directly with the CPU, which is fast but has a relatively small capacity, and can be directly accessed by the CPU, such as DRAM and SRAM commonly used in main memory; the auxiliary memory is used as a supplement to the main memory, which has a large capacity but is slower, and cannot be directly accessed by the CPU, and needs to exchange data with the main memory through input / output operation, hard disk, optical disk, U disk and the like belong to auxiliary memory.

[0072] According to the access mode, it includes random access memory, sequential access memory, direct access memory and the like. The random access memory can randomly access any storage unit, and the access time is independent of the storage location, and the RAM is a typical random access memory; the access of the sequential access memory must be carried out in the order of the storage unit, such as a magnetic tape, which needs to be found step by step from the beginning; the direct access memory combines the characteristics of random access and sequential access, first finds the storage area by random access, and then sequentially accesses in the area, and the hard disk belongs to this category.

[0073] In the embodiment of the application, the memory is taken as an example of SRAM (Static Random-Access Memory, static random access memory) for introduction.

[0074] Step 102, determining the detection order of a plurality of storage areas.

[0075] In the embodiment of the application, the memory of the server generally includes a plurality of storage areas, in order to efficiently and orderly complete the detection work, it is necessary to determine the detection order of these storage areas, and the determination of this detection order has a plurality of bases, such as the physical address order of the storage area, from low address to high address or from high address to low address in turn; or according to the use frequency of the storage area, the frequently used area is detected preferentially; or combined with the historical bad block record, the area prone to problems before is detected first; the specific one is not limited here.

[0076] Step 103, according to the detection order, the running data of the plurality of storage areas is transmitted to the peripheral register in turn.

[0077] In the embodiment of the application, after the detection order is determined, the running data of the storage area can be transmitted to the idle peripheral register determined previously according to the detection order, and the process of data transmission needs to ensure the integrity and accuracy of the data; the peripheral register as a temporary storage unit for data processing can provide stable data source for subsequent bad block detection.

[0078] Step 104, detecting whether each storage area is a bad block area after the running data of each storage area is transmitted to the peripheral register.

[0079] In the embodiment of the present application, when the running data of each storage area is successfully transmitted to the peripheral register, detection can be immediately performed on whether the storage area is a bad block area. The detection can be performed in various ways, for example, by comparing the checksum of the data, comparing the checksum of the data transmitted to the register with the correct checksum stored in advance; or analyzing the specific features of the data to check whether the data conforms to the normal storage mode; or performing read-write testing, reading and writing the data in the peripheral register, and checking whether the data can be normally read and written. If an abnormality is found in the data during the detection, for example, the checksum does not match, the data cannot be normally read and written, and the like, the corresponding storage area is determined as a bad block area.

[0080] The present application discloses a bad block detection method for a memory. The method uses an idle peripheral register of a server during operation as a temporary data storage carrier, and transmits the running data of each storage area to the idle peripheral register in a determined detection sequence without interrupting the normal operation of the server, thereby completing the bad block detection of each storage area. The present application ensures the continuous normal operation of the server during detection, avoids system downtime or performance degradation caused by detection, and can also discover and process newly added bad blocks in real time, preventing the risks of data loss, system crash and the like caused by the bad blocks not being detected in time, thereby dynamically maintaining the health status of the memory, prolonging the stable operation period of the server, and improving the risk resistance and data security of the overall system. The present application optimizes the efficiency of the detection process by reasonably planning the detection sequence of the storage areas and using the idle peripheral resources, and ensures that the detection operation minimally affects the normal performance of the server, balancing the system stability and operation efficiency.

[0081] In an embodiment of the present application, the detection sequence of the plurality of storage areas is determined by obtaining the historical bad block addresses of the memory and determining the detection sequence of the plurality of storage areas according to the historical bad block addresses of the memory.

[0082] In the embodiment of the present application, when determining the detection sequence of the plurality of storage areas, the historical bad block addresses of the memory are first obtained. The bad block record log of the memory or the database specially storing historical fault information can be accessed to extract the specific address information of the storage areas that have been determined as bad block areas in the past. These historical bad block addresses can include the physical addresses or logical addresses of the storage areas that have once appeared data errors, read-write failures and the like. The addresses can be sorted and classified to form a complete list of historical bad block addresses, which provides a basis for determining the detection sequence.

[0083] After obtaining the historical bad block addresses of the memory, the detection order of the plurality of storage areas can be determined according to the addresses. Specifically, the distribution characteristics of the historical bad block addresses can be analyzed, such as checking where the bad block addresses are concentrated or which storage areas have frequently appeared bad blocks. For the storage areas where the historical bad block addresses are located and the storage areas adjacent to the historical bad block areas, they can be arranged at the front end of the detection order. For the storage areas that have never appeared bad block records or have a very low historical bad block occurrence rate, they can be arranged at the rear end of the detection order. In addition, the storage areas that have recently appeared bad blocks can be given a higher detection priority according to the time sequence of the historical bad blocks, so as to ensure that the areas that may appear problems again can be found in time.

[0084] In an example, the storage areas are R0, R1, R2,..., R15 respectively,

[0085] The historical bad block addresses recorded in the bad block record log are:

[0086] R3: appeared 3 times (2025-01-15, 2025-05-08, 2025-07-01)

[0087] R7: appeared 1 time (2025-03-22)

[0088] R8: appeared 1 time (2025-06-15)

[0089] It can be seen that R3 has appeared 3 times of bad blocks in three months in succession, and the latest occurrence is on July 1st;

[0090] R7 and R8 each appear 1 time of bad blocks, and the occurrence time of R8 (June 15th) is later than that of R7 (March 22nd);

[0091] The adjacent areas of R3 are R2 and R4, the adjacent areas of R7 are R6 and R8, and the adjacent areas of R8 are R7 and R9;

[0092] Therefore, the detection priority can be generated according to the above information, the first priority: R3;

[0093] The second priority: R8 (recent occurrence + adjacent high-risk area R7)

[0094] The third priority: R7 (historical bad block + adjacent high-risk area R8)

[0095] The fourth priority: R2, R4 (adjacent areas of R3)

[0096] The fifth priority: R6, R9 (adjacent areas of R7 and R8)

[0097] The sixth priority: the rest of the area (R0, R1, R5, R10-R15) without appearing bad block, the order of the rest of the area without appearing can be random.

[0098] The application can improve the pertinence and efficiency of bad block detection, discover potential bad block risk in time, reduce the possibility of data loss or system failure caused by bad block not being detected in time, and reasonably arrange the detection order to avoid unnecessary priority detection on the area without historical bad block record, save system resources and detection time, make the memory bad block detection more intelligent and practical, and ensure the stable operation of the server.

[0099] In an embodiment of the application, detecting whether each storage area is a bad block area comprises: writing verification data in the storage area, and reading data of the storage area; judging whether the data of the storage area is consistent with the verification data, and if the data of the storage area is inconsistent with the verification data, determining that the storage area is a bad block area.

[0100] In the embodiment of the application, when detecting whether each storage area is a bad block area, specific verification data needs to be written in the storage area first, the verification data is usually a binary sequence with explicit characteristics, the writing process needs to follow the read-write protocol of the memory to ensure that the data is accurately written in each storage cell of the target storage area, and after the writing is completed, the data of the storage area can be read immediately, the verification data just written is read out from the storage area and temporarily stored in the peripheral register or temporary cache for subsequent comparison.

[0101] Then, consistency judgment can be performed on the read data and the original written verification data, and the process is realized through byte-by-byte and bit-by-bit comparison, if the comparison result shows that the two are completely consistent, it indicates that the storage area can normally complete the data writing and reading operation, the storage function is normal, and it is determined as a non-bad block area, if the comparison result shows that the two are inconsistent, it indicates that the storage area has physical damage or logical failure and cannot reliably store data, and therefore it is determined as a bad block area.

[0102] The application can realize bad block detection through read-write operation at the software level without relying on a complex hardware diagnosis module, reduce the occupation of additional resources of the server, ensure the detection accuracy while taking into account the system running efficiency, and effectively reduce the risk of data loss or server downtime caused by bad block not being discovered in time.

[0103] In an embodiment of the application, the method further comprises: storing the detection results of the plurality of storage areas in a memory management bitmap; and when it is necessary to allocate memory to the memory, allocating memory to each storage area in the memory according to the memory management bitmap.

[0104] In the embodiment of the present application, the bitmap data is stored in an array, assuming that each element is an unsigned integer, each element can represent the state of 16 memory pools, each bit of each element corresponds to the state of a storage area, the 0th bit corresponds to the bad block state of storage area 0, the 1st bit corresponds to the bad block state of storage area 1, and the 15th bit represents the bad block state of storage area 15.

[0105] Storing the detection results of the plurality of storage areas into the memory management bitmap refers to allocating a corresponding binary bit in the memory management bitmap for each storage area of the memory, if a storage area is detected as a bad block area, the binary bit is marked as "1", if it is a normal area, it is marked as "0", the state of each storage area is recorded in such a bitmap form, which is intuitive and efficient; when it is necessary to allocate memory to the memory, the memory management bitmap can be queried, and the appropriate area is selected for allocation from the normal storage area marked as "0", if there is a demand for continuous normal area, the memory area corresponding to the continuous "0" sequence in the bitmap is searched to meet the allocation requirement, for example, assuming that the memory has 10 storage areas, the memory management bitmap is "0010010001" (from left to right, corresponding to 1 to 10 areas), among which 3, 6 and 10 are bad blocks (marked as "1"), when 2 continuous memory areas are needed, the bitmap is queried to find that 1-2, 4-5, 7-8 and 8-9 are continuous normal areas, which can be selected for allocation.

[0106] The memory management bitmap in the present application can store a large amount of state information of storage areas in a compact form, so that the query and management of the state of each area by the system are more efficient, and the time cost of state query is reduced; in memory allocation, the normal storage area can be quickly located according to the bitmap, so as to avoid allocating memory to the bad block area, improve the accuracy and reliability of memory allocation, and also meet the specific allocation demand of continuous memory area and the like more quickly, improve the efficiency of memory allocation, and ensure the stability of server memory use.

[0107] In an embodiment of the present application, the memory in the memory is allocated according to the memory management bitmap, comprising: determining the bad block area in the memory according to the memory management bitmap, and not allocating memory to the bad block area.

[0108] In the embodiment of the present application, when allocating memory in each storage area in the memory according to the memory management bitmap, each bit in the memory management bitmap can be traversed first, and the storage area corresponding to the bit marked as "1" is identified as a bad block area. This identification process is efficiently completed through bit operation. For example, for a bitmap unit of 32-bit integer type, the status of 32 storage areas can be checked at one time. After identifying the bad block area, when a memory allocation request arrives, these bad block areas are automatically excluded from the candidate list of the available memory pool. In specific implementation, a dynamic available memory area linked list is maintained. When initializing or updating the detection result, the node corresponding to the bad block area is removed from the linked list or marked as unavailable. When an application program requests to allocate memory, the memory allocator only selects normal areas from the linked list for allocation, so as to ensure that data is not written into the bad block area.

[0109] Suppose that the server memory is divided into 64 storage areas (numbered R0-R63), and the memory management bitmap is 8 bytes (64 bits), wherein the 3rd, 7th, 29th and 45th bits are marked as "1", and R3, R7, R29 and R45 are bad block areas. When the system is initialized, an available memory linked list containing R0-R2, R4-R6, R8-R28, R30-R44, R46-R63 is constructed. When an application program requests to allocate 3 continuous storage areas, the memory allocator traverses the linked list to find the first continuous area (such as R0-R2) that meets the condition, and allocates it to the application program, while updating the linked list state. If there is a subsequent request for 5 continuous areas, the system will continue to search from the linked list, such as allocating the R4-R8 area. In the entire process, the bad block areas marked are completely avoided, so as to ensure the safety and reliability of data storage.

[0110] The present application quickly identifies the bad block through hardware-level bit operation, greatly improves the allocation efficiency, completely avoids the permanent loss risk caused by writing data into the bad block, and makes the memory allocator not need to scan the entire memory at each request, thereby improving the allocation efficiency.

[0111] In an embodiment of the present application, the server further comprises a direct memory access controller configured to sequentially transmit the running data of the plurality of storage areas to the peripheral register according to the detection sequence.

[0112] In the embodiment of the present application, the server can further comprise a DMA (Direct Memory Access) controller, and the running data of the plurality of storage areas can be transmitted to the peripheral register in sequence according to the detection sequence through the DMA controller. Specifically, the DMA controller is first initialized and configured, including specifying the source address of data transmission, i.e. the start address of the storage area to be detected, determining the target address, i.e. the address of the idle peripheral register, and the data length of single transmission (matching the capacity of a single storage area) in sequence according to the detection sequence.

[0113] After the configuration is completed, the DMA controller does not need to be continuously intervened by the CPU, and can automatically read the running data from the current storage area to be detected according to the set detection sequence, and directly transmit the data to the specified idle peripheral register. In the transmission process, the DMA controller acquires the use right of the system bus through the bus arbitration mechanism, and independently completes the address increment, data transfer and transmission count operations. When the data transmission of a single storage area is completed, the DMA controller sends an interrupt signal to the CPU to inform the CPU that the current region transmission is completed, and then the CPU can trigger the transmission configuration of the next storage area, so that the DMA controller continues to perform the next round of data transmission in sequence.

[0114] Suppose that the server memory has five storage areas (A, B, C, D, E), the detection sequence is B→D→A→E→C, and the peripheral register R is in an idle state. First, a configuration instruction is sent to the DMA controller: the source address is set as the start address of the B region, the target address is set as the address of the register R, and the transmission length is the fixed capacity of a single region (such as 1MB). After receiving the instruction, the DMA controller automatically reads 1MB running data from the B region, directly transmits the data to R through the bus, and sends an interrupt to the CPU after the transmission is completed. After receiving the interrupt, the CPU reconfigures the DMA controller, changes the source address to the start address of the D region, and keeps the other parameters unchanged. When the detection of the storage area B is completed, the DMA controller performs the transmission again and notifies the CPU after the transmission is completed. In this way, the DMA controller transmits the data of each region to the register R in sequence according to the sequence of B→D→A→E→C. During the transmission, the CPU only performs the next round of configuration after each transmission is completed, and does not need to participate in the data transfer process.

[0115] The application can significantly reduce the resource occupation of the CPU in the data carrying process, so that the CPU can process other server tasks during the data transmission, and improve the overall operation efficiency; meanwhile, the DMA controller reduces the instruction interaction delay between the CPU and the peripheral register and the memory through the hardware level data transmission mechanism, improves the speed and stability of the data transmission, ensures that the storage area operation data can be quickly and accurately transmitted to the peripheral register, provides efficient data support for subsequent bad block detection, and further optimizes the overall process of the memory bad block detection.

[0116] In an embodiment of the application, the idle peripheral register is determined by obtaining the bitmap value and the interrupt state of each peripheral register in the server, and determining the idle peripheral register according to the bitmap value and the interrupt state of each peripheral register.

[0117] In the embodiment of the application, when determining the idle peripheral register, the bitmap value and the interrupt state of each peripheral register in the server are first obtained, wherein the bitmap value of the peripheral register is a state identifier represented by a binary bit, each bit corresponds to a peripheral register, and a specific bit such as '0' represents unoccupied and '1' represents occupied, which directly reflects whether the register is in the state of being used by other tasks; and the interrupt state is used to judge whether the register is processing data transmission, error feedback or other interrupt requests, which is usually obtained by querying the interrupt flag bit of the register, such as '0' representing no interrupt and '1' representing having an unprocessed interrupt; after obtaining these two types of information, the bitmap value and the interrupt state of each peripheral register are comprehensively judged: if the bitmap value corresponding bit of a certain peripheral register is '0' (unoccupied) and the interrupt state is '0' (no unprocessed interrupt), it is determined that the register is currently in an idle state; if the bitmap value corresponding bit is '1' (occupied) or the interrupt state is '1' (having an unprocessed interrupt), it is determined that the register is in a non-idle state, so that all idle peripheral registers meeting the conditions are screened out.

[0118] Suppose that the server has 8 peripheral registers: numbers 1-8, and the corresponding bitmap value is '01001010' from left to right corresponding to the 1-8 registers, '0' represents unoccupied and '1' represents occupied, that is, the bitmap values of the 1st, 3rd, 6th and 8th registers are '0': unoccupied, and the 2nd, 4th, 5th and 7th are '1': occupied; at the same time, it is found through the interrupt state that the interrupt flag bits of the 1st and 6th registers are '0' (no interrupt), and the interrupt flag bits of the 3rd and 8th registers are '1' (having an unprocessed interrupt), and after comprehensive judgment, the 1st and 6th registers with bitmap value '0' and interrupt state '0' can be screened out, and it is determined that these two are the current idle peripheral registers, which can be used for subsequent transmission of storage area operation data.

[0119] The application can comprehensively and accurately identify the registers in the available state by combining the bitmap value of the peripheral register and the interrupt state, avoid determining the registers with unprocessed interrupts as idle, and thus guarantee the stability of subsequent data transmission; meanwhile, the binary bit representation of the bitmap value makes the state query process efficient and convenient, reduces the consumption of system resources, and improves the efficiency of the idle register identification, thereby providing reliable hardware support for the smooth data transmission in the memory bad block detection.

[0120] In an embodiment of the application, the server further comprises a timer, and the method further comprises: setting a time interval by the timer; and determining the idle peripheral registers in the server, comprising: periodically determining the idle peripheral registers in the server according to the time interval.

[0121] In the embodiment of the application, the server can further comprise a timer, and the timer in the server can be used to set a fixed time interval, which is pre-configured according to the running load of the server, the detection requirement of the memory, and other factors, for example, 10 seconds, 30 seconds, etc., and after the time interval is set by the timer, the operation of determining the idle peripheral registers in the server can be periodically performed with the time interval as a period, and the state of all peripheral registers can be detected every time the period comes, so as to determine the idle peripheral registers in the server. This periodic detection ensures that the dynamic change of the peripheral registers can be grasped in real time, and the registers released by other tasks or no longer idle due to being occupied can be found in time.

[0122] Suppose that the time interval set by the server is 20 seconds, and the idle peripheral registers determined by the detection at the initial moment are R2 and R5; after 20 seconds, the timer triggers the first periodic detection, at this moment, R3 originally occupied changes to the bitmap value of "0" and the interrupt state of "0" due to the end of the related task, and R5 originally idle is occupied by a new task, and the bitmap value changes to "1", so that the idle peripheral registers determined by this detection are updated to R2 and R3; after another 20 seconds, the timer triggers the second periodic detection, and it is found that R1 is released to be idle and R3 is occupied, at this moment, the idle peripheral registers are updated to R1 and R2, and so on, and the system updates the information of the idle peripheral registers every 20 seconds.

[0123] The application can dynamically track the state change of the peripheral register, timely find the newly released idle register, avoid the resource waste caused by the static state of the register, timely eliminate the occupied register, ensure that the peripheral register used for data transmission is always in the available state, thereby improving the utilization rate of the peripheral register resource and the continuity and reliability of the memory bad block detection process.

[0124] In an embodiment of the application, the timer can be used to periodically perform the bad block detection task, and timely find the bad block area to provide convenience for the user.

[0125] In an embodiment of the application, the method further comprises: if it is determined that the target storage area is a bad block area, repairing the target storage area, and after the repair of the target storage area is completed, returning the data of the target storage area stored in the peripheral register to the target storage area.

[0126] In the embodiment of the application, when it is determined that the target storage area is a bad block area, the repair process of the bad block area can be started, and the repair method is different according to the type of the memory and the nature of the bad block. For example, for a flash memory, the error data in the bad block area can be cleared by performing a block erase operation, and the damaged storage unit can be replaced by a redundant spare storage unit. For a DRAM or other volatile memory, the state of the storage unit can be reset by a refresh circuit, or the storage area corresponding to the physically damaged address line / data line can be shielded.

[0127] During the repair process, the repair progress is continuously monitored, and whether the repair is completed is determined by verifying whether the repaired storage area can normally read and write data. Once the repair is completed, the data of the target storage area previously temporarily stored in the peripheral register is transmitted back to the storage area. The return process also needs to pass through the verification mechanism to ensure the integrity of the data, so as to avoid the data loss caused by the return error.

[0128] Suppose that the R6 area of the memory is detected as a bad block area, and the running data thereof has been temporarily stored in the peripheral register R. First, the repair of R6 is performed: for the R6 of the flash type, the block erase instruction is called to clear the error data therein, and then the preset spare block in the chip is enabled to replace the damaged physical block in R6, and the data "0xAA55" is written and successfully read out, so as to confirm that the repair is completed. Then, the data return process is started, the original running data of R6 such as "0x12345678" stored in the register R is written into the repaired R6, and after the return is completed, the CRC check value of the data is calculated, which is consistent with the check value before the return, so as to confirm that the data return is correct, and R6 returns to the available state.

[0129] The application can maximize the available storage space of the memory by repairing the bad block area and returning data, reduce the waste of storage resources caused by bad blocks, and has important significance for server environments with tight storage capacity. At the same time, the data return ensures the integrity of the original running data after the repair of the bad block area, avoids data loss caused by the bad block detection and repair process, guarantees the continuity and data consistency of the server operation, and reduces the impact of storage area failure on system business.

[0130] As Figure 2 , a flowchart of a bad block detection method of a memory is shown, which can trigger a bad block detection task by a timer, find an idle peripheral register when the time interval arrives, then backup the data in the SRAM to the idle peripheral register by DMA, and then perform the SRAM read-write test, for example, the bad block detection of the 6th storage area of the SRAM. If the detection is normal, the data of the 6th storage area stored in the peripheral register is returned to the 6th storage area. After the detection of each storage area is completed, the detection result can be stored in the memory management bitmap, and the memory management bitmap is dynamically updated after periodic detection.

[0131] The application discloses a bad block detection method of a memory, which uses the idle peripheral register of a server during operation as a temporary data storage carrier, and transmits the running data of each storage area to the idle peripheral register in a determined detection order without interrupting the normal operation of the server, so as to complete the bad block detection of each storage area. The application ensures the continuous normal work of the server during detection, avoids system downtime or performance degradation caused by detection, and can also discover and process newly added bad blocks in real time, prevent data loss, system crash and other risks caused by the bad blocks that are not detected in time, dynamically maintain the health status of the memory, prolong the stable operation period of the server, and improve the risk resistance and data security of the whole system. The application optimizes the efficiency of the detection process by reasonably planning the detection order of the storage area and using the idle peripheral resource, and ensures that the influence of the detection operation on the normal performance of the server is minimized, and the system stability and operation efficiency are considered.

[0132] It should be noted that, for the method embodiments, in order to simply describe, they are all described as a series of action combinations, but those skilled in the art should know that the application embodiments are not limited by the action order described, because according to the application embodiments, certain steps can be performed in other order or at the same time. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions involved are not necessarily necessary for the application embodiments.

[0133] ReferenceFigure 3 The application discloses a memory bad block detection device, and relates to the technical field of servers.

[0134] The first determining module 201 is used for determining an idle peripheral register in the server during server running.

[0135] The second determining module 202 is used for determining a detection sequence of the plurality of storage areas.

[0136] The transmission module 203 is used for sequentially transmitting running data of the plurality of storage areas to the peripheral register according to the detection sequence.

[0137] The detection module 204 is used for detecting whether each storage area is a bad block area after the running data of the each storage area is transmitted to the peripheral register.

[0138] The application discloses a memory bad block detection device, and relates to the technical field of servers.

[0139] In an embodiment of the application, the second determining module comprises:

[0140] The acquisition sub-module is used for acquiring a historical bad block address of the memory.

[0141] The first determining sub-module is used for determining a detection sequence of the plurality of storage areas according to the historical bad block address of the memory.

[0142] In an embodiment of the application, the detection module comprises:

[0143] The writing sub-module is used for writing verification data in the storage area and reading data of the storage area.

[0144] The judging sub-module is configured to judge whether the data of the storage area is consistent with the verification data, and if the data of the storage area is not consistent with the verification data, it is determined that the storage area is a bad block area.

[0145] In an embodiment of the present application, the server further comprises:

[0146] The storage module is configured to store the detection results of the plurality of storage areas into the memory management bitmap.

[0147] The allocation module is configured to allocate memory to each storage area in the memory according to the memory management bitmap when it is necessary to allocate memory to the memory.

[0148] In an embodiment of the present application, the allocation module comprises:

[0149] The second determining sub-module is configured to determine the bad block area in the memory according to the memory management bitmap, and not allocate memory to the bad block area.

[0150] In an embodiment of the present application, the server further comprises a direct memory access controller, and the transmission module comprises:

[0151] The transmission sub-module is configured to transmit the running data of the plurality of storage areas to the peripheral register according to the detection sequence through the direct memory access controller.

[0152] In an embodiment of the present application, the first determining module comprises:

[0153] The third determining sub-module is configured to obtain the bitmap value and the interrupt state of each peripheral register in the server.

[0154] The fourth determining sub-module is configured to determine the idle peripheral register in each peripheral register according to the bitmap value and the interrupt state of each peripheral register.

[0155] In an embodiment of the present application, the server further comprises a timer, and the server further comprises:

[0156] The setting module is configured to set the time interval through the timer.

[0157] The first determining module comprises:

[0158] The fifth determining sub-module is configured to periodically determine the idle peripheral register in the server according to the time interval.

[0159] In an embodiment of the present application, the server further comprises:

[0160] The repair module is configured to, if the target storage area is determined to be a bad block area, repair the target storage area, and after the repair of the target storage area is completed, return the data of the target storage area stored in the peripheral register to the target storage area.

[0161] The application discloses a bad block detection device for a memory, which uses an idle peripheral register in a server running process as a temporary data storage carrier, and sequentially transmits running data of each storage area to the idle peripheral register according to a determined detection sequence without interrupting the normal running of the server, so that the bad block detection of each storage area is completed; the application ensures the continuous normal work of the server during the detection, avoids system shutdown or performance decline caused by the detection, and can also discover and process newly added bad blocks in the running process in real time, so as to prevent the bad blocks from causing risks such as data loss and system crash due to untimely detection, thereby dynamically maintaining the health status of the memory, prolonging the stable running period of the server, and improving the anti-risk ability and data security of the overall system; the application optimizes the efficiency of the detection process by reasonably planning the detection sequence of the storage area and using the idle peripheral resources, and ensures that the influence of the detection operation on the normal performance of the server is minimized, and the system stability and running efficiency are considered.

[0162] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the related parts refer to the part of the method embodiment.

[0163] The application further provides an electronic device, comprising:

[0164] The computer program is stored in the memory and can be run on the processor, and when the computer program is executed by the processor, each process of the bad block detection method embodiment of the memory is realized, and the same technical effect is achieved.

[0165] The application further provides a computer readable storage medium, and a computer program is stored in the computer readable storage medium, and when the computer program is executed by the processor, each process of the bad block detection method embodiment of the memory is realized, and the same technical effect is achieved.

[0166] Each embodiment in the specification is described in a progressive manner, and each embodiment mainly describes the difference from other embodiments, and the same and similar parts of each embodiment can be referred to.

[0167] Those skilled in the art will appreciate that embodiments of the present application can be readily used as a method, apparatus, or computer program product. Accordingly, embodiments of the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, embodiments of the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, and the like) embodying computer program instructions.

[0168] Embodiments of the present application are described herein with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing terminal apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal apparatus, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.

[0169] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing terminal apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.

[0170] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal apparatus to cause a series of operational steps to be performed on the computer or other programmable terminal apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable terminal apparatus provide steps for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.

[0171] While preferred embodiments of the present application have been described, modifications and alterations thereto will occur to those skilled in the art upon reading the preceding description. In particular, it will be apparent to those skilled in the art that parts can be added to, or substituted for, parts of the described embodiments of the present application. Accordingly, the application is intended to be

[0172] Finally, it is to be understood that the phraseology or terminology such as "first" and "second" etc. used herein is merely intended to differentiate one entity or operation from another entity or operation, without necessarily requiring or implying any actual such relationship or order between such entities or operations. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0173] The memory bad block detection method, device, equipment and medium provided by the present application are described in detail above, and the principles and implementation manners of the present application are described by applying specific examples. The above description of the embodiments is only used to help understand the method of the present application and its core idea. Meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manners and application ranges will be changed, and the above description of the present application should not be understood as a limitation.

Claims

1. A method of bad block detection for a memory, the method comprising: The application is applied to a server, the server comprises a peripheral register and a memory, the memory comprises a plurality of storage areas, and the method comprises the following steps: determining an idle peripheral register in the server during running of the server; determining a detection sequence of the plurality of storage areas; transmitting running data of the plurality of storage areas to the peripheral register in sequence according to the detection sequence; detecting whether each storage area is a bad block area after running data of each storage area is transmitted to the peripheral register.

2. The method of claim 1, wherein, The step of determining the detection sequence of the plurality of storage areas comprises the following steps: obtaining a historical bad block address of the memory; determining the detection sequence of the plurality of storage areas according to the historical bad block address of the memory.

3. The method of claim 1, wherein the step of determining the bad block is performed by a memory controller. The step of detecting whether each storage area is a bad block area comprises the following steps: writing verification data into the storage area and reading data of the storage area; judging whether the data of the storage area is consistent with the verification data, and if the data of the storage area is not consistent with the verification data, determining that the storage area is a bad block area.

4. The method of claim 1, wherein the step of determining the bad block is performed by a memory controller. The method further comprises the following steps: storing detection results of the plurality of storage areas into a memory management bitmap; allocating memory to each storage area in the memory according to the memory management bitmap when it is necessary to allocate memory to the memory.

5. The method of claim 4, wherein the step of determining the number of bad blocks is performed by a method comprising: The step of allocating memory to each storage area in the memory according to the memory management bitmap comprises the following steps: determining a bad block area in the memory according to the memory management bitmap, and not allocating memory to the bad block area.

6. The method of claim 1, wherein the step of determining the bad block of the memory is performed by a memory controller. The server further comprises a direct memory access controller, and the step of transmitting running data of the plurality of storage areas to the peripheral register in sequence according to the detection sequence comprises the following steps: transmitting running data of the plurality of storage areas to the peripheral register in sequence according to the detection sequence through the direct memory access controller.

7. The method of claim 1, wherein the step of determining the bad block is performed by a memory controller. The step of determining an idle peripheral register comprises the following steps: obtaining bitmap values and interrupt states of each peripheral register in the server; determining an idle peripheral register among the peripheral registers according to the bitmap values and the interrupt states of the peripheral registers.

8. The method of claim 1, wherein, The server further comprises a timer, and the method further comprises the following steps: setting a time interval through the timer; The step of determining an idle peripheral register in the server comprises the following steps: periodically determining an idle peripheral register in the server according to the time interval.

9. The method of claim 1, wherein, The method further comprises the following steps: if a target storage area is determined to be a bad block area, repairing the target storage area and returning data of the target storage area stored in the peripheral register to the target storage area after the target storage area is repaired.

10. A bad block detection apparatus of a memory, characterized by, The application is applied to a server, the server comprises a peripheral register and a memory, the memory comprises a plurality of storage areas, and the device comprises the following steps: a first determining module, configured to determine an idle peripheral register in the server during running of the server; a second determining module, configured to determine a detection sequence of the plurality of storage areas; a transmitting module, configured to transmit running data of the plurality of storage areas to the peripheral register in sequence according to the detection sequence; a transmission module, configured to sequentially transmit the running data of the plurality of storage areas to the peripheral register according to the detection sequence; a detection module, configured to detect whether each storage area is a bad block area after the running data of the each storage area is transmitted to the peripheral register.

11. An electronic device, comprising: comprising: a processor, a memory, and a computer program stored on the memory and capable of running on the processor, which, when executed by the processor, implements the steps of the bad block detection method of the memory according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, a computer readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the bad block detection method of the memory according to any one of claims 1-9.