Flash memory failure early warning method and device, computer equipment and storage medium

By obtaining the failure rules of flash memory and the read and write status of each naked chip, dividing the status of the naked chip and sending warning information, the problem of flash memory being unable to be early warning is solved, and the effect of reducing the load of solid-state hard disk and improving reliability is achieved.

CN120179448APending Publication Date: 2025-06-20成都芯忆联信息技术有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510327250.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing flash has no early warning when the die fails, resulting in a sharp increase in the load of the solid-state drive, an increase in the risk of data loss and read and write errors, affecting the reliability and stability of the solid-state drive.

Method used

By obtaining the failure rules of flash memory and the read and write state of each die, the die is divided into a healthy state and a dead state, and sending warning information to the host when the number of dies in the dead state reaches the preset threshold.

Benefits of technology

It realizes early warning of the failure risk of the naked chip in flash memory, reduces the load of the solid-state drive, reduces the risk of data loss and read and write errors, and improves the reliability and stability of the solid-state drive.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179448A_ABST
    Figure CN120179448A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of storage, and discloses a flash memory failure early warning method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring a failure rule of the flash memory and a read-write state of each bare chip in the flash memory; dividing the bare chip into a healthy state and a failure state according to the failure rule and the read-write state; and judging whether the number of the bare chips in the failure state reaches a preset threshold value or not, and if yes, sending early warning information to a host. According to the method, the failure rule of the flash memory and the read-write state of each bare chip are obtained, the health state and the failure state of the bare chips are divided according to the failure rule, the early warning information can be sent to the host when the failure number of the bare chips reaches the preset threshold value, early warning of the failure risk of the bare chips in the flash memory is achieved, the host can master the storage state in time, and the reliability of the flash memory is improved. Therefore, the host can sense the abnormity of the solid state disk, the business of the host can be dynamically adjusted, the fault recovery capability of the whole system in the scene can be improved, and the influence on the subsequent business of the host can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of storage technology, and in particular to a flash memory failure early warning method, device, computer equipment and storage medium. Background Art

[0002] In the field of storage technology, solid-state drives are widely used in various electronic devices due to their advantages such as high-speed reading and writing and low power consumption. The core storage component of solid-state drives is flash memory, which is composed of multiple bare chips. As the basic storage unit of flash memory, the bare chip is responsible for storing and reading data. Many bare chips work together to form the storage capacity of flash memory.

[0003] During long-term use, the bare chip in the flash memory will inevitably experience performance degradation or even failure due to frequent read and write operations. When the bare chip in the flash memory fails, the SSD not only has to deal with garbage collection, but also needs to process the business read and write requests sent by the host, which causes a sharp increase in the load of the SSD, which can easily lead to serious problems such as data loss and increased read and write errors, greatly affecting the reliability and stability of the SSD, and further causing adverse effects on the systems and businesses that rely on the storage device, such as system crashes and business interruptions. In addition, the SSD cannot warn the host of the risk of bare chip failure in advance, making it difficult for the host to accurately judge the current storage status, resulting in a significant reduction in the performance of the SSD and extremely slow fault recovery, which seriously limits the overall performance and user experience of the SSD. Summary of the invention

[0004] The present invention provides a flash memory failure warning method, device, computer equipment and storage medium to solve the technical problem that the existing flash memory failure cannot be warned in advance.

[0005] In a first aspect, a flash memory failure early warning method is provided, comprising:

[0006] Obtain the failure pattern of flash memory and the read and write status of each die in the flash memory;

[0007] According to the failure rules and the read / write status, the bare chips are divided into a healthy state and a failure state;

[0008] It is determined whether the number of the bare chips in the failure state reaches a preset threshold, and if so, a warning message is sent to the host.

[0009] In a second aspect, a flash memory failure warning device is provided, comprising:

[0010] An acquisition module, used to acquire the failure rule of the flash memory and the read and write status of each bare chip in the flash memory;

[0011] A classification module, used for classifying the bare chips into healthy states and failed states according to the failure rules and read / write states;

[0012] An early warning module, configured to determine whether the number of die chips in the failure state reaches a preset threshold, and if so, send an early warning message to the host.

[0013] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above flash memory failure early warning method are implemented.

[0014] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above flash memory failure early warning method are implemented.

[0015] The beneficial effects of the present invention compared with the prior art are as follows: By obtaining the failure law of the flash memory and the read / write status of each die chip, the present invention divides the health and failure states of the die chips. When the number of failed die chips reaches the preset threshold, an early warning message can be sent to the host, realizing early warning of the failure risk of the die chips in the flash memory, enabling the host to timely master the storage status, making the host perceive the abnormality of the solid-state drive, helping to dynamically adjust the host's services, improving the fault recovery ability of the entire system in this scenario, reducing the impact on the subsequent services of the host, effectively avoiding the sharp increase in the load of the solid-state drive caused by the failure of the die chips, thereby reducing the risks of data loss and read / write errors, improving the reliability and stability of the solid-state drive, also improving the system operation status, avoiding system crashes or service interruptions, and enhancing the overall performance and user experience of the solid-state drive.

[0016] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following preferred embodiments are specifically described in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a flowchart of a flash memory failure early warning method in an embodiment of the present invention;

[0018] Figure 2 is Figure 1 a flowchart of a specific implementation manner of step S10 in;

[0019] Figure 3 is Figure 1 a flowchart of a specific implementation manner of step S20 in;

[0020] Figure 4 is Figure 1 a flowchart of a specific implementation manner of step S30 in;

[0021] Figure 5 It is a schematic structural diagram of a flash memory failure warning device in an embodiment of the present invention;

[0022] Figure 6 It is a schematic structural diagram of a computer device in an embodiment of the present invention;

[0023] Figure 7 It is another schematic structural diagram of a computer device in an embodiment of the present invention. Detailed implementation manners

[0024] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0025] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprise" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.

[0026] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0027] It should be further understood that the term " / and / " used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0028] Please refer to Figure 1 as shown in Figure 1 which is a schematic flowchart of the flash memory failure warning method provided by the embodiment of the present invention. The flash memory failure warning method includes the following steps:

[0029] S10: Obtain the failure pattern of the flash memory and the read / write status of each die in the flash memory.

[0030] For step S10, the solid-state drive performs a read operation on the flash memory to obtain information about the flash memory and the dies in the flash memory. Among them, a solid-state drive (SSD) is a non-volatile storage device based on flash chips, with extremely fast read and write speeds. Under normal conditions, the sequential read speed can reach several thousand MB per second, and the sequential write speed can easily exceed 1000 MB per second, greatly shortening the system startup time and the file access time. In addition, the solid-state drive has low power consumption and generates less heat during operation, which can effectively reduce the overall power consumption of the device and extend the device's battery life. Flash Memory belongs to a non-volatile semiconductor storage medium and is the core storage component of the solid-state drive. Using electrically erasable programmable read-only memory technology, it can retain the stored data even after power-off and has high-speed read and write capabilities. Compared with traditional mechanical hard drives, the read and write latency is greatly reduced, enabling fast data access. The flash memory in this embodiment is NAND Flash, which has high storage density and low cost and is widely used in the field of large-capacity storage. A die is the basic storage unit of the flash memory, responsible for the actual storage and reading operations of data. The performance of the die, such as storage density, read and write speed, durability, etc., directly affects the performance of the entire flash memory and the solid-state drive based on this flash memory.

[0031] More specifically, the flash memory failure warning method in this embodiment obtains the failure law of the flash memory, such as failing after reaching a certain range of read failure times, etc., and based on the read and write status of the die combined with the failure law, infers the health status of the die, which helps to lay the foundation for subsequent flash memory failure warning. It is beneficial to inform the user of the die health problem before the die actually fails, avoid data loss caused by sudden die failure, and ensure the integrity and security of the data.

[0032] Among them, as Figure 2 shown, step S10, that is, obtaining the failure law of the flash memory and the read and write status of each die in the flash memory, includes the following steps S11 - S12.

[0033] S11: Obtain the type of the flash memory and the failure law corresponding to the type.

[0034] It can be understood that the failure laws corresponding to different types of flash memories are not the same. Through classification algorithms, the failure laws of various types of flash memories are statistically analyzed. For example, when the ratio of the number of read failures to the number of write / erase operations reaches a:b respectively, the failure probability of this type of flash memory reaches 80%, and then the relationship between the number of read failures and the number of write / erase operations can be fitted, and the failure law of this type of flash memory can also be obtained.

[0035] S12: Detect the historical number of read failures and the historical number of write / erase operations of each die in the flash memory.

[0036] The number of read failures directly reflects the reliability of the die during data reading. If the number of read failures increases, it means that the die's ability to read data decreases, and a failure may be imminent. The number of write / erase cycles is closely related to the durability of the flash memory. Each time a write / erase operation is performed on the memory cells of the flash memory, its performance will be somewhat degraded. When the number of write / erase cycles approaches or reaches the limit designed for the flash memory, the risk of die failure increases significantly. By detecting these two key indicators of the die, the actual usage status of the die can be accurately grasped, providing key data support for inferring the health of the die based on the failure pattern, thereby detecting potential failure risks in advance, ensuring data storage security, and avoiding data loss or storage device failure caused by die failure.

[0037] Specifically, the detection of the historical number of read failures and the historical number of write / erase cycles of the die is periodic. Periodically detecting the die can continuously track the changing trend of the die's performance. During the long-term use of the flash memory, the number of read failures and the number of write / erase cycles will accumulate continuously. Regular detection can promptly discover abnormal growth situations. For example, if it is found in a certain detection that the number of read failures has increased significantly compared to the previous detection, it indicates that the die's reading function may be about to fail, and measures can be taken in advance to avoid sudden reading errors during device operation, ensuring stable data output and maintaining stable device operation.

[0038] S20: Divide the die into a healthy state and a failed state according to the failure pattern and the read / write status.

[0039] For step S20, compare the obtained read / write status of the die with the failure pattern to predict whether the die is about to reach or has reached the failed state. When the read / write status of the die reaches or exceeds the failure pattern, the die is classified into the failed state. The read / write ability of the die in the failed state is limited, and the probability of read failure increases significantly. When the read / write status of the die does not reach the failure pattern, the die is classified into the healthy state.

[0040] Among them, as Figure 3 shown, step S20, that is, divide the die into a healthy state and a failed state according to the failure pattern and the read / write status, includes the following steps S21 - S22.

[0041] S21: According to the type of flash memory and the failure pattern, successively fit the historical number of read failures and the historical number of write / erase cycles of the die, and calculate the equivalent pre-failure number of times and the equivalent failure number of times for each die.

[0042] Step S21 fits the historical read failure times and historical write / erase times of a single die according to the failure law, and predicts its future equivalent pre-failure times and equivalent failure times. The historical read failure times and historical write / erase times of each die are fitted according to the corresponding failure law. For example, if the failure law is a fitting function of the read failure times and write / erase times, substitute the historical read failure times and historical write / erase times into the fitting function to obtain the equivalent pre-failure times and equivalent failure times of the current die. The equivalent pre-failure times and equivalent failure times are the supporting indicators for predicting the failure state of the die.

[0043] S22: Define the die with the historical read failure times less than the equivalent pre-failure times as the healthy state, the die with the historical read failure times greater than the equivalent pre-failure times and less than the equivalent failure times as the first failure state, and the die with the historical read failure times greater than the equivalent failure times as the second failure state.

[0044] Specifically, the first failure state is used to warn the host that the die is about to fail, and the second failure state is used to warn the host that the die has failed, so as to facilitate the host to reduce business operations.

[0045] In step S22, the historical read failure times are compared with the equivalent pre-failure times and equivalent failure times, so as to clearly define the die as the healthy state, the first failure state and the second failure state, and classify the die state clearly and meticulously, providing intuitive and easy-to-understand die state information for the system and users. The definition of the healthy state allows users to know which dies are operating normally, and data storage and reading are safe and reliable; the first failure state indicates that the die has potential risks. Although it has not completely failed, it needs to be closely monitored, and countermeasures such as data backup or planning for replacement time can be prepared in advance; the second failure state directly indicates that the die has been severely damaged and may cause data loss or equipment failure at any time, and immediate action must be taken.

[0046] S30: Judge whether the number of dies in the failure state reaches a preset threshold. If so, send a warning message to the host.

[0047] Step S30 determines whether an early warning needs to be issued regarding the overall health of the current flash memory by counting the number of die in a failed state and comparing it with a preset threshold. When the number of die in a failed state in the flash memory reaches a certain proportion, i.e., touches the preset threshold, an early warning message is sent to the host in a timely manner, enabling the host to quickly become aware of the abnormality of the flash memory, giving the user or system administrator the opportunity to take corresponding measures, such as backing up important data in a timely manner, arranging equipment maintenance, or replacing the flash memory, etc., before the flash memory failure has a serious impact on data and system operation, effectively avoiding serious consequences such as data loss and system crashes caused by large-scale flash memory failures, ensuring the reliability and stability of the entire storage system, and ensuring business continuity; at the same time, reducing the processing load of the host's business, such as temporarily restricting the operation of some non-critical services and giving priority to ensuring the storage and reading of core data, can give the flash memory a certain buffer time, which helps to maintain the integrity and consistency of data storage. At the same time, it also buys more time for technicians to deal with the flash memory failure problem, thereby avoiding a complete paralysis of the host's business due to a complete flash memory crash and ensuring the continuous and stable operation of the entire system.

[0048] Among them, as Figure 4 shown, step S30, that is, determining whether the number of die in the failed state reaches the preset threshold. If so, an early warning message is sent to the host, including the following steps S31 - S34.

[0049] S31: Preset a first failure quantity threshold and a second failure quantity threshold in advance.

[0050] S32: Count the number of die in the first failed state and the second failed state in the flash memory respectively.

[0051] S33: Compare the number of die in the first failed state with the first failure quantity threshold. If the number of die in the first failed state reaches the first failure quantity threshold, a first early warning message is sent to the host.

[0052] S34: Compare the number of die in the second failed state with the second failure quantity threshold. If the number of die in the second failed state reaches the second failure quantity threshold, a second early warning message is sent to the host.

[0053] For steps S31 - S34, by setting different failure quantity thresholds, namely the first failure quantity threshold and the second failure quantity threshold, and counting the number of die in the first failed state and the second failed state respectively, it is possible to perform hierarchical early warning for different degrees of die failure situations, which is conducive to the host taking more precise response strategies according to the different early warning messages received; through the three-step logic of threshold setting, data statistics, and hierarchical early warning, the transformation from passive response to faults to active prevention of failures is achieved.

[0054] Specifically, in steps S33 and S34, the first warning message or the second warning message is sent to the host through SMART information.

[0055] SMART information, full name: Self-Monitoring, Analysis and Reporting Technology, that is, self-monitoring, analysis and reporting technology, is built into the solid-state drive. The SMART technology monitors various parameters of the storage device, such as the number of flash read failures and write-erase times, and intuitively reflects the operating state of the device. Once the flash memory fails, the SMART information will issue a warning signal according to the preset rules, enabling the solid-state drive and the host to cooperate together to accelerate the flash memory exception handling, and ultimately reducing the impact on the subsequent host services.

[0056] Specifically, after step S33, that is, comparing the number of die in the first failure state with the first failure number threshold. If the number of die in the first failure state reaches the first failure number threshold, after sending the first warning message to the host, it includes the steps of: controlling the host service to switch to the first low level.

[0057] The host service switching to the first low level means reducing the host's service operations, such as pausing or reducing the running priority of services that have less impact on the system core functions and are non-immediate, specifically, for example, background data synchronization tasks, non-urgent data analysis operations, etc. Although these services help to improve the richness of the system functions under normal circumstances, when there are potential risks in the flash memory, continuing to run them will increase the flash memory read and write burden. Pausing them can reduce unnecessary access to the flash memory, thereby reducing the possibility of further damage to the flash memory.

[0058] Specifically, after step S33, that is, comparing the number of die in the first failure state with the first failure number threshold. If the number of die in the first failure state reaches the first failure number threshold, after sending the first warning message to the host, it further includes the steps of: increasing the detection frequency of the read and write status of the die; ignoring the flash memory retry operation.

[0059] If the number of die in the first failure state in the flash memory exceeds the first failure number threshold, the cycle of detecting the read and write status of the die is shortened to increase the detection frequency of the die, so as to more timely capture the change of the read and write status of the die and prevent the die from deteriorating rapidly, resulting in data loss.

[0060] A die that has reached the first failure state has a high risk of read operation failure. After a read failure, the system will perform a flash retry operation, i.e., a flash retry operation, to attempt to continue the read operation. However, when the flash memory is already in an unstable state at this time, the flash retry operation may be a persistent error caused by the performance problem of the die, rather than a transient exception. Ignoring the flash retry operation can prevent the system from wasting a large amount of time and resources on repeatedly attempting read and write tasks that may not succeed, thereby improving the overall operating efficiency of the system. For example, when reading a file, if a certain die has a read error, the system may originally perform multiple retries. However, after ignoring the retry operation, the system can quickly skip that die and attempt to read data from other healthy dies, saving the time overhead caused by retries. Moreover, by reducing invalid retry operations, the system can concentrate resources on the processing of critical services and ensure the normal operation of core functions. When there are problems with the flash memory, giving priority to ensuring the operation of critical services is crucial for maintaining the availability of the system and the user experience.

[0061] Specifically, after step S34, that is, comparing the number of dies in the second failure state with the second failure number threshold. If the number of dies in the second failure state reaches the second failure number threshold, after sending the second warning message to the host, it includes the steps of: controlling the host service to switch to the second lower level.

[0062] The first lower level corresponds to the situation where there are potential risks in the flash memory, that is, the number of dies in the first failure state reaches the first failure number threshold. While the second lower level is for the case where the flash memory is already in a serious crisis state, the number of dies in the second failure state reaches the second failure number threshold, and the flash memory may completely fail at any time. Therefore, the second lower level is a countermeasure taken in a more severe risk situation, and the restrictions and adjustments on the service are more radical. The purpose is to do our best to ensure the security of core data and the operation of basic system functions at the edge of the flash memory about to completely collapse.

[0063] When the number of dies in the second failure state reaches the threshold, it means that the flash memory is already in a serious crisis state and may completely collapse at any time. At this time, switching the host service to the second lower level can minimize the read and write operations on the flash memory and concentrate resources to fully ensure the secure storage and reading of core data. For example, the core business data of an enterprise, such as order information, customer profiles, etc. By switching the service level, the system can suspend all non-essential services, such as secondary data analysis, report generation, etc., and use limited resources to ensure the stable storage of core data, preventing the loss of core data caused by the complete failure of the flash memory and bringing huge losses to the enterprise.

[0064] It can be seen that in the above solution, by obtaining the failure law of the flash memory and the read / write status of each die, the health and failure status of the die are divided accordingly. When the number of failed dies reaches the preset threshold, a warning message can be sent to the host, realizing early warning of the failure risk of the die in the flash memory, enabling the host to timely grasp the storage status, so that the host can sense the abnormality of the solid-state drive, which helps to dynamically adjust the host's services, improve the fault recovery ability of the entire system in this scenario, reduce the impact on the subsequent services of the host, effectively avoid the sharp increase in the load of the solid-state drive caused by die failure, and then reduce the risks of data loss and read / write errors, improve the reliability and stability of the solid-state drive, also improve the system operation status, avoid system crashes or service interruptions, and improve the overall performance and user experience of the solid-state drive.

[0065] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0066] In one embodiment, the present invention provides a flash memory failure warning device, which corresponds one-to-one with the flash memory failure warning method in the above embodiment. As Figure 5 shown, the flash memory failure warning device includes an acquisition module 101, a division module 102, and a warning module 103. The detailed descriptions of each functional module are as follows:

[0067] The acquisition module 101 is used to acquire the failure law of the flash memory and the read / write status of each die in the flash memory.

[0068] The division module 102 is used to divide the die into a healthy state and a failure state according to the failure law and the read / write status.

[0069] The warning module 103 is used to determine whether the number of dies in the failure state reaches the preset threshold. If so, a warning message is sent to the host.

[0070] In one embodiment, the acquisition module 101 is specifically used for:

[0071] Acquire the type of the flash memory and the failure law corresponding to the type;

[0072] Detect the historical read failure times and historical write / erase times of each die in the flash memory.

[0073] In one embodiment, the division module 102 is specifically used for:

[0074] According to the type of the flash memory and the failure law, successively fit the historical read failure times and historical write / erase times of the die, and calculate the equivalent pre-failure times and equivalent failure times of each die;

[0075] Define the die with the number of historical read failures less than the equivalent pre-failure number as the healthy state, the die with the number of historical read failures greater than the equivalent pre-failure number and less than the equivalent failure number as the first failure state, and the die with the number of historical read failures greater than the equivalent failure number as the second failure state.

[0076] In one embodiment, the warning module 103 is specifically configured to:

[0077] Preset a first failure quantity threshold and a second failure quantity threshold in advance;

[0078] Count the quantities of the dies in the first failure state and the second failure state in the flash memory respectively;

[0079] Compare the quantity of the dies in the first failure state with the first failure quantity threshold. If the quantity of the dies in the first failure state reaches the first failure quantity threshold, send a first warning message to the host;

[0080] Compare the quantity of the dies in the second failure state with the second failure quantity threshold. If the quantity of the dies in the second failure state reaches the second failure quantity threshold, send a second warning message to the host.

[0081] For the specific limitations of the flash memory failure warning device, reference can be made to the limitations on the flash memory failure warning method in the above text, which will not be elaborated here. Each module in the above flash memory failure warning device can be implemented in whole or in part through software, hardware and their combination. The above modules can be embedded in the processor in the computer device in the form of hardware or be independent of it, or be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0082] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a flash memory failure warning method server.

[0083] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be asFigure 7 As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a flash memory failure warning method.

[0084] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0085] Obtain the failure law of the flash memory and the read / write status of each die in the flash memory;

[0086] Divide the dies into a healthy state and a failure state according to the failure law and the read / write status;

[0087] Judge whether the number of dies in the failure state reaches a preset threshold. If so, send a warning message to the host.

[0088] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the following steps are implemented:

[0089] Obtain the failure law of the flash memory and the read / write status of each die in the flash memory;

[0090] Divide the dies into a healthy state and a failure state according to the failure law and the read / write status;

[0091] Judge whether the number of dies in the failure state reaches a preset threshold. If so, send a warning message to the host.

[0092] It should be noted that for the functions or steps that the above computer-readable storage medium or computer device can achieve, reference can be made to the relevant descriptions in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0093] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0094] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0095] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included in the protection scope of the present invention.

Claims

1. A flash memory failure early warning method, characterized in that: include: Obtain the failure pattern of flash memory and the read and write status of each die in the flash memory; According to the failure rules and the read / write status, the bare chips are divided into a healthy state and a failure state; It is determined whether the number of the bare chips in the failure state reaches a preset threshold, and if so, a warning message is sent to the host.

2. The flash memory failure warning method according to claim 1, characterized in that: The obtaining of the failure rule of the flash memory and the read / write status of each bare chip in the flash memory includes: Get the type of flash memory and the failure rule of the corresponding type; Check the historical read failure count and historical write and erase count of each die in the flash memory.

3. The flash memory failure warning method according to claim 2, characterized in that: The step of classifying the die into a healthy state and a failed state according to the failure rule and the read / write state includes: According to the type and failure law of flash memory, the historical read failure times and historical write and erase times of the bare chip are fitted in turn, and the equivalent pre-failure times and equivalent failure times of each bare chip are calculated; A die whose historical read failure number is less than the equivalent pre-failure number is defined as a healthy state, a die whose historical read failure number is greater than the equivalent pre-failure number and less than the equivalent failure number is defined as a first failure state, and a die whose historical read failure number is greater than the equivalent failure number is defined as a second failure state.

4. The flash memory failure warning method according to claim 3, characterized in that: The determining whether the number of the bare chips in the failure state reaches a preset threshold, and if so, sending a warning message to the host, includes: Presetting a first failure quantity threshold and a second failure quantity threshold; Counting the number of bare chips in a first failure state and a second failure state in the flash memory respectively; Comparing the number of bare chips in a first failure state with a first failure number threshold, and sending a first warning message to a host if the number of bare chips in the first failure state reaches the first failure number threshold; The number of dies in the second failure state is compared with a second failure number threshold, and if the number of dies in the second failure state reaches the second failure number threshold, a second warning message is sent to the host.

5. The flash memory failure warning method according to claim 4, characterized in that: After comparing the number of dies in the first failure state with the first failure number threshold, and sending the first warning information to the host if the number of dies in the first failure state reaches the first failure number threshold, the method further includes: Control host services to switch to the first lowest level.

6. The flash memory failure warning method according to claim 4, characterized in that: After comparing the number of dies in the first failure state with the first failure number threshold, and sending the first warning information to the host if the number of dies in the first failure state reaches the first failure number threshold, the method further includes: Increase the frequency of checking the read and write status of the die; Ignore the flash retry operation.

7. The flash memory failure warning method according to claim 4, characterized in that: After comparing the number of dies in the first failure state with the first failure number threshold, and sending the first warning information to the host if the number of dies in the first failure state reaches the first failure number threshold, the method further includes: Control host services to switch to the second lowest level.

8. Flash memory failure warning device, characterized in that: include: An acquisition module, used to acquire the failure rule of the flash memory and the read and write status of each bare chip in the flash memory; A classification module, used for classifying the bare chips into healthy states and failed states according to the failure rules and read / write states; The early warning module is used to determine whether the number of the bare chips in the failure state has reached a preset threshold, and if so, send an early warning message to the host.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the flash memory failure warning method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the flash memory failure warning method according to any one of claims 1 to 7 are implemented.