Error prediction and predictive memory maintenance through real time machine learning methods

Real-time machine learning analytics predict flash memory failures in IoT devices, enabling proactive maintenance and reducing costs by identifying high-risk devices and executing timely corrective actions.

WO2025120336A1PCT designated stage expired Publication Date: 2025-06-12ATHENS UNIVERSITY OF ECONOMICS & BUSINESS (AUEB) E L K E +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/GR2024/000039
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-06
Filing Date
2024-12-02
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Flash memory technology in IoT and embedded devices is prone to errors due to endurance, read disturb, write disturb, and data retention issues, exacerbated by high temperatures, frequent writing, electromagnetic fields, and long periods of inactivity, which current methodologies struggle to predict and mitigate effectively.

Method used

A method and system utilizing real-time machine learning techniques to predict failures in non-volatile data storage of embedded devices by analyzing data streams from these devices, enabling predictive maintenance, error mitigation, and cost reduction through targeted interventions.

Benefits of technology

The system effectively predicts failures and reduces maintenance costs by identifying high-risk devices and executing timely corrective actions, thereby extending the lifespan of flash memory chips and ensuring reliable operation of critical infrastructure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GR2024000039_12062025_PF_FP_ABST
    Figure GR2024000039_12062025_PF_FP_ABST
Patent Text Reader

Abstract

The invention disclosed herein concerns a method and a system to predict, in a collaborative manner, failures of the non-volatile data storage of embedded devices. These predictions allow the predictive and preventive maintenance of the devices and also provide an estimation of the impact of maintenance or the failure (e.g. downtime, service interruption). This is achieved through the process, analysis, and correlation in real time, through machine learning methods, of data streams that contain messages regarding the status of flash storage in devices.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DESCRIPTION

[0002] Error predietton and predictive memory maintenance through real time machine learning methods

[0003] The present invention is related with a) the risk management and analysis of critical infrastructure, b) techniques for error correction and data refresh in flash memory, c) embedded devices and d) real-time data stream analysis.

[0004] The vast majority of loT and embedded devices employ integrated flash circuits (hereafter “flash memory chips”) for their storage needs. At the same time, it is very common for these devices to be interconnected through a network connection in their deployments, as a part of their functionality during their op eration . Additionally, they may be considered as or included in critical infrastructure.

[0005] The flash memory technology, despite being much faster, more robust and resilient than other data storage technologies, such as magnetic (hard drives) and optical, is not error free. They include: ® Endurance issues, in which the ability of the building blocks of the memories (hereafter “the cells”), to store data, is lost;

[0006] • Read disturb issues, in which retrieval of data from some cells causes corruption of the data stored in nearby (topologically) cells;

[0007] • Write or program disturb issues, in which writing data to some cells causes corruption of the data stored in nearby (topologically) cells;

[0008] • Data retention issues, in which the lack ofusage of cells (mainly retrieval of data) for a large period of time causes corruption of their data

[0009] The problems of flash memory technology arc related to the following factors: • Operation at high temperatures

[0010] • Writing new data at high frequency

[0011] • Usage in environments with strong electromagnetic fields

[0012] • Lack ofusage for a long period of time The exact definition for the magnitude of the above factors depends on the supplier of the flash memory chips and their optimal operation specifications. For example, industrial-grade flash memory chips can withstand higher temperatures and larger numbers of data writes than the corresponding consumer grade flash integrated circuits (flash memory chips) .

[0013] At the same time, as technology advances, there is a need to store ever- increasing amounts of data in devices. The requirement to store a larger amount of data in smaller p hysical s izes of memory ch ip s, combined with the larger number of embedded devices that contain flash memory chips, render the technical effects above more prevalent and critical for the infrastructure they are part of, and demand their efficient and effective management.

[0014] The current methodologies that attempt to resolve these technical problems mainly include retrieving the state of storage, i.e. the identified from the device technical problems of the flash memory chip for each device, and dealing with errors after they occur, usually in conjunction with backup devices or storage. Description of Flash Memory and Refresh Methods

[0015] A flash memory chip is a type of non-volatile data storage mean. At a very high-level, flash memory works by trapping electrons in a gate of a MOSFET transistor (cell). Even though flash memory has nowadays become ubiquitous, from mobile phones and other every-day gadgets to high end servers and spacecrafts, it is still susceptible to effects that corrupt the stored data, including the errors of write and read disturb and data corruption, which are amplified by high temperatures, strong electromagnetic fields and long periods of inactivity. Error correction codes (ECCs) are usually stored along with the data, to mitigate some of these effects. However, frequent refreshing of the data is required to prevent the accumulation of errors, which renders their correction impossible.

[0016] Description of Machine Learning Methods Machine learning models are algorithms and methodologies that enable computers to learn from input data and perform predictions or make decisions without being explicitly programmed to do so. These models are key enablers of artificial intelligence and have revolutionized various fields, from image recognition and natural language processing to autonomous vehicles and personalized recommendations.

[0017] At their core, machine learning models are mathematical algorithms that identify and learn patterns and relationships within the data they are trained on. They analyze vast amounts of information, identifying relevant features and patterns that humans cannot easily discern. By leveraging statistical techniques, these models can make accurate predictions or classify new, unseen data based on the patterns they have already identified and learned.

[0018] There are various types of machine learn ing models, including supervised learning, unsupervised learning, and reinforcement learning. In supervised learning, models learn from labeled data, where they are provided with input examples and corresponding desired outputs. Unsupervised learning models, on the other hand, find patterns and structure in unlabeled data without explicit guidance. Unsupervised models mainly utilize clustering algorithms. A semi-supervised clustering methodology is one that incorporates a-priori known information to accelerate the clustering process. Reinforcement learning models learn through trial and error by interacting with an environment and receiving feedback in the form of rewards or penalties.

[0019] Machine learning models ha ve demonstrated their ability to solve complex problems and automate tasks that were once considered difficult or impossible for computers. They continue to advance rapidly, with ongoing research and development driving innovation in the respective field and they have been used for predictive and preventive maintenance of machinery.

[0020] The document EP2077559A2 presents a methodology that can be employed for refreshing data stored in flash memory chips, in order to avoid read disturb errors. The specific invention employs a refresh management table that monitors the refresh actions performed on each cell of the chip. The specific invention protects from read disturb errors, does not predict future errors and does not take under account the state of near and / or similar devices.

[0021] The document US 7079422B1 describes a periodic flash refresh methodology that aims to retain the threshold voltages of the cells (which determine the datum stored in each cell), and, at the same time, extending the life of the flash by avoiding repetitive erasing and writingof the same value in the same memory cell, through moving data to different physical storage locations. This invention does not predict future errors and does not take under account the state of near and / or similar devices.

[0022] The document US20090161466 A l describes extending flash memory data retention via rewrite refresh. This document discloses an invention that retains the threshold voltages of the cells without erasing these first. This can prevent degrading the storage capacity of the cells, which typically occurs when erasingthem. This invention does notpred set errors and does not take under account the state of near and / or similar devices.

[0023] The document US 10648735 B2 describes the predictive maintenance of a dryer based on machine learning. According to this document, the failures of a dryer are predicted through information transmitted over a communication network. Anomalies and failures during the functionality O f the dry er are predicted using machine learning using data retrieved from other dryers that exist within a communication network. This invention does not predict errors or allow the predictive maintenance of flash memory chips and does not support automa tic mitigation of errors.

[0024] The document US1 1307570B2 presents a machine learning based predictive maintenance of equipment. A general predictive maintenance methodology is described, which is achieved through data received from sensors of equipment. The specific invention does not predict errors or allow the predictive maintenance of flash memory chips, does not support, automatic mitigation of these errors and the specific methodology limits proactive maintenance through data received through sensors of equipment.

[0025] The document CNl 15576778 A presents a server predictive maintenance model method based on machine learning. The system contains a data collection module, a data analysis module, a feature extraction module and a model fusion module. Its output includes predictions for the failures of servers in a computer room. This invention does notpredict errors or allow the predictive maintenance of flash memory chips and does not support automatic mitigation of these errors.

[0026] The document CN 1 14237194A describes a predictive machine maintenance system. Specifically, it discloses a method that can detect abnormalities in machines that can indicate that maintenance is required. The processing requires collection of readings from sensors, which is not performed on servers that require complex communication networks to be reached, but instead it is performed near the sensors. The specific invention does not predict errors or allow the predictive maintenance of flash memory chips, does not support automatic mitigation of these errors and the specific methodology limits the proactive maintenance on data received through sensors.

[0027] In the document CN1 1 1598251 A a predictive maintenance system and method based on machine learning for Computer Numerical Control (CNC) systems is presented. Concisely, the state of the CNC machines is monitored and sent to a cen tral processing server through a communication network, while th e processing server employs mach inc learn ing techniques on these data to predict when maintenance will be required. The specific invention docs not predict errors of flash memory chips and docs not support automatic mitigation of these errors .

[0028] Followingly, a brief description of the invention is presented, in order to provide a basic understanding of certain elements of technology described in the present invention.

[0029] In short, this in vention concerns a method and a system that utilize machine learning methods to predict, in a collaborative manner, lai lures of the nonvolatile data storage in embedded devices. These predictions allow the predictive maintenance ofthc devices and also provide an estimation of the impact of maintenance or damage (e.g. downtime, service interruption). This is achieved through the collection of data that characterize the flash memory chip-based storage of thc devices. The aforementioned data, which characterize the flash memory chip -based storage of the devices, are sent by the devices to a server in the form of a data stream, where they are processed, analyzed and correlated in real time. Real-time analysis (as opposed to classification against already stored and analyzed data) is required in order to:

[0030] • identify groups of devices that consistently fail more often than others

[0031] • detect rising failure trends

[0032] • enable timely and optimal response to unprecedented circumstances Also, real-time analytics make sense to embedded devices, where failure data are very rare and difficult to collect.

[0033] In the context of this invention, the term 'predictive maintenance'' denotes maintenance based on targeted and specific p redictions that are related with the occurrence of functional errors, as opposed to the term "proactive maintenance", which denotes maintenance performed according to general funct ion al sp ccificat ion s .

[0034] The 12 drawings that accompany the invention, depict, concisely, the following: Figure 1 : a high-level abstract overview of the system architecture, in which the devices are considered as peers that are all able to communicate through a network with the Controller or the Controllers.

[0035] Figure 2: the Controller subsystem that comprises the Data receptors. Sample collector and deviation detector. Clustering, Classification, Reporting, and Device action units.

[0036] Figure 3: a high-level representation of an loT device that contains a number of sensors, interfaces (communication network interfaces or other interlaces) and flash memory.

[0037] Figure 4: the Command executor and the Flash status transmission units, which are Included within the loT devices and that, along the Controller subsystem, are included in the invention.

[0038] Figure 5: analysis of the Flash status transmission unit that resides within the loT devices. F igure 6 : th e Data recep tor, which is the uri it of the Controller that receives and performs an initial lightweight processing of the data stream of messages.

[0039] Figure 7; the Sample collector and deviation detector (hereafter referred to as SCoDe2), which is the unit of the Controller that monitors the stream and calculates and maintains its statistical properties.

[0040] Figure 8: the Clustering unit, which detects the clusters that exist in the samp le collected by the SCoDe2,

[0041] Figure 9: the Classification unit, which classifies the messages to the clusters detected by the Clustering unit, accelerated through the findings of the SCoDe2.

[0042] Figure 10: the Reporting unit, which crea tes reports based on the findings of the Clustering and the Classification units and can route mitigation actions through the Device action unit.

[0043] Figure 1 1 : the Device action unit, which communicates with the devices, in order to perform specific actions, according to the instructions of the Reporting unit.

[0044] Figure 12: the CluNN algorithm, which separates a dataset on clusters.

[0045] The invention is defined by the claims and disclosed below, with references to the drawings.

[0046] The general overview of the invention is depicted in Figure 1.

[0047] The invention includes the Controller (m 1 ) and the devices (s 1 ), which are all known to the Controller and can communicate with it over a comm unication network, employing a message transmission interface (11). Every device contains at least one flash memory chip (m2) as its data storage unit [Figure 31. The operating system of the devices (s2) can detect the state of its storage unit, which is able to fix a small number of errors. When this number of errors starts to become larger, the storage unit develops errors that cannot be corrected and requires maintenance, which represents a cost, both for the vendor and the owner of the devices. In the context of the invention, the operating system of the devices rep oils to the Controller its state through messages (p 1). The Controller collects the data stream (p2) from all the devices (12), processes a sample of the collected data and has the ability to predict failures of the non-volatile data storage of embedded devices. These predictions are used for predictive maintenance, and, additiona Uy, mil igation and corrective actions. Overall, through these p redictions, the subsequent actions are able to significantly reduce the maintenance cost. [Figure 2]

[0048] More specifically, the invention involves the synergy of the mechanisms included in the devices and the Controller, implemented on a computer. As such, the mechanisms that reside in each device that is desired to be monitored are the Flash status transmission unit (m3) and the Command executor unit (m4) [Figure 4], which are described in detail below. The Flash status transmission unit [Figure 5] can either be implemented as a standalone software unit or be embedded into the hardware controller of the flash memory chips (e.g. flash memory chip types eMMC, UFS). In the case of the Flash status transmission unit as a standalone software unit, it retrieves in formation from the drivers that control the flash memory chips, through its interface with them (s3), or other sensors (s4), through a corresponding interface (s5). Otherwise, in the case of its embeddedimplementation, the software driver of the flash memory chip p rovides the necessary information through its interface with the operating system (s6).

[0049] Regarding the information that will be collected by the Flash status transmission unit, they include at least (a) the number of faulty or malfunctioning blocks (hereafter referred to as “bad blocks”), (b) the blocks / second read, (c) the blocks / second written, (d) some device identification and (e) the time that the memory has been functional

[0050] However, additional metrics that can be also collected indicatively include (1) the number of errors corrected through the ECC, (2) the number of uncorrectable errors that have been detected, (3) the identity of the flash memory chip in wh ich the errors were detected, (4) the physical addresses of the chip where the faults are located, (5) the files that contain the data in which the faults were detected, (6 ) the number o f times that these files have been accessed, (7) the estimated lifespan of the internal flash memory chip, (8) the number of read requests, (9) the number of write requests, (10) the memory refresh interval, (1 1 ) the temperature of the device and (12) the physical location of the device. Those skilled in the art may be able to specify additional metrics. The larger the number of collected parameters, the more accurate predictions and recommendations will be produced by the invention.

[0051] The metrics above are accumulated into a message and are sent to the registered Data receptor of the Controller or stored in a local database (s7) when connectivity with the Data receptor is not available. The transmission of the messages is performed through a suitable for the communication network message transmission mechanism (s8). Parameter (d) uniquely identifies the device and allows the Controller to register the availability time, i.e. the defined response time as this is delimited through thehclp of predefined points of exchange of messages, in a way that allows the data receptor to populate and maintain the scoreboard, as well as identify response gaps. The availability time is defined as the number of consecutive time slots that the device has responded. If the availability time is smaller than the previous availability time, a service interruption would have occurred.

[0052] As far as the Command executor (m4) [Figure 4] is concerned, it is an optional mechanism, which receives the appropriate messages (p3) from the Device action unit that instruct it to calibrate the Hash refreshing interval (i,c. increase or decrease the period of time between each flash block refresh) or to repair files. Those skilled in the art could implement additional commands. In an implementation that does not contain an Command executor, the present invention will only be able to produce predictions for predictive maintenance and periods of service interruption, while an implementation including the Command executor allows the invention to also perform corrective actions.

[0053] Regarding the Controller, it is defined as a unit that includes and implements the machine learning procedure hereby referred to as Best Effort Real Time Clustering & classification - BEReTiC, which is implemented by the steps described below [Figure 2]. The description will abideby the conceptual flow of data, as they enter theController from the devices, to the output of results, butalsoto the transmission of settings and commands from the Controller to the devices. STEP 1

[0054] The messages from the devices to the Controller are received by the Data receptor (m5) [Figure 6]. In particular, due to the various protocols that can be used for the communication with the devices, the data received from the Data receptor may require adaptation and normalization , in order to transform the messages that will be processed by the later stages. Additionally, the Data receptor also retains a scoreboard (s 10) that contains the device identifier and the average number of time slots that the device failed to send data. The information of the scoreboard for the device is also included in the message. These messages can be compared usingdistance metrics, such as the Cosine or the Jaccard similarity coefficients, and the comparison should yield a number from 0 to 1 , inclusive, where 1 denotes identical messages. Many instances of the Data Receptor can exist simultaneously, in order to optimally receive data streams from a very large number of devices.

[0055] STEP 2

[0056] The stream ofmessages emitted from the Data receptor is monitored by the Sample collector and deviation detector (SCoDe2) (m6) [Figure 7] through a special interface (s 1 1 ). In the context of this step, SCoDe2monitors the stream and stores a sample of the received messages at the sample memory that it maintains (p4). The size of* the sample depends on the total number and rate of received messages and it should be sufficiently big to draw meaningful results. The calculations performed on the sample are: a) calculation of the standard deviation (p5) and b) the prevailing value (hereafter referred to as “mode”) of the minimum distance (p6). The output is employed by the clustering and the classification steps.

[0057] STEP 3 As far as the Clustering unit (m7) [Figure 8] is concerned, it performs clustering on the sample collected by SCoDe2, aided by a) the calculated mode of the minimum distances as an additional input and b) the retained Summary structures of the clusters (p7). The Summary structure of the clusters is an essential structure of the Clustering unit and it is calculated io through the average of each numerical metric contained in each sample message collected in each cluster. The Summary structure of the clusters allows, in essence, the Reporting unit to perform predictions and the Device action unit to perform predictive actions. The Summary structures are being retained and stored by the Clustering unit in a designated memory (p8). The time that they will remain stored is determined through the employment of various policies, as, for example, a time decaying policy, by which they are discarded after a specified time, or an attenuation-based policy, by which they are discarded according to the number of messages they contain. The discovered clusters are retained in a special memory (p 10) until the next time that clustering is executed. The Clustering unit can use any clustering algorithm (p9) that does not require a predefined number of clusters into which the set of messages contained in the sample must be separated . An example o f a clustering algor ithm that achieves good results is a clustering algorithm hereby referred as Clustering Nearest Neighbor - Clu’NN (aO) [Figure 12], which accepts the mode of the minimum distances as an input and works as follows:

[0058] Begin (al ). Read the mode of the minimum distances (a2). Read the next message ofthc sample (a3), ifthemessage already belongs to a cluster (a4), con tinue to the next message. Else, search for a message in the sample, lor whom the distance between the two messages is less than the mode of the minimum distance (a5). On a6, if such a message is found, then these messages are assumed to belong in the same cluster and, if the other message is not already contained in a cluster, a new cluster is created. If a message whose distance is smaller than the mode ofthc minimum distance does not exist and the unmatched message does not belong to a cluster, a new cluster that con tains only the unmatched message is created. If more messages exist within the sample, continue with the next message (a7), else clustering has been completed (a8).

[0059] STEP 4

[0060] The Classification unit (m8 ) [Figure 9] receives the whole message stream from the Data receptor, through the appropriate interface (s i 2). This way, every message, whether it was employed in the context of the previous steps or not, is also classified by the Classification unit in one of the n detected clusters, by employing the cluster hypothesis. The classification (p 1 1 ) can be performed either using the messages contained in the clusters of the sample, the Summary structures, the recently already classified Messages, which are stored in the Message cache (p 12) or a combination of these. As any clustering algorithm or a combination of them can be used during the clustering, any classification algorithm or a combination of algorithms can be used during the classi fication. In case a combination of algorithms gets employed, the result ofthe classification is decided through weighted voting. One classificationalgorithm that has been employed is a version o f the K -Nearest Neighbor (KNN) algorithm, which takes as an inp ul the mode of the minimum distance and classifies the message to the cluster that contains at least one message whose distance from it is smaller than the mode of the minimum distance. It is important to note that associations of messages with minority classes, i.e. clusters with a few members, are able to be detected and they are not discriminated against. Their detection is bound, though, on the quality of the sample, which is determined by the number of the messages that each time constitute a sample representative. The messages that could not be classified into any one of the clusters can cither form a new cluster or be considered as outliers.

[0061] STEP 5

[0062] From the Reporting unit (m9) [Figure 10], reports, recommendations and mitigation actions are produced, according to the findings of the Classification (p 13). Through the Summary structures of the clusters derived from the Clustering step, it is possible, in the context of this step, to calculate the evolution of the clusters, which is the main structure that allows deriving the metrics for operational the reporting step. The derival metrics include (a) the long-term effect of the mitigation of the consequences and the predictive maintenance actions (through the evolution ofthe clusters described above), (b) the discrepancy between the predictions and the actual results, (c) the state that the devices are converging to (through the ex istence of long standing Summary structures that consistently increase in size), as well as (d) rising trends through the emergence of relatively young Summary' structures that gain size quickly.

[0063] These metrics can p ermit the timely scheduling of predictive maintenance, the optimal reaction to unprecedented events and provide the ability to locate groups of devices that are consistently damaged more often than others. They are stored at the metric memory (p 14), in order to be used to produce reports from the report creation engine, with the purpose to be read by support staff (si 3). Clusters whose Summary structures contain a larger than zero number of errors corrected through the ECC or uncorrectable errors are considered as High-risk clusters. For devices contained in high- risk clusters (a) mitigation actions are requested thro ugh the interface with the Device action unit (s 14) and, if they contain uncorrectable errors, (b) predictive maintenance is recommended and the time required for the repairs is estimated (taking under account the average out-of-service time contained in the Summary structure of the cluster, given that the average out-of-service time has been included into the messages by the Data receptor and the Summary structure contains as average out-of-service time their average) by the report creation engine. The predictive maintenance recommendations for the high-risk clusters depend on the state of the metrics in the cluster. Those skilled in the art may choose the recommendations that can be contained in the reports. Some examples of predictive maintenance recommendations are: device Hashing (format), reduction of writes, replacement of the device, reduction of the temperature.

[0064] STEP 6

[0065] The interface of the Command executor that is contained in the devices is utilized by the Dev ice action unit (m 10) [Figure 1 1]. Thus, during this step, the mitigation actions are processed through the command analysis software (p 15), appropriate messages are formed, which are sent to the devices contained in High-risk clusters, in order to modify various parameters that have been shown to be efl'ective in mitigating similar consequences before they manifest themselves. Concurrently, at devices that are not contained in High-risk clusters, all-clear messages are prepared and sent from the command analysis software, which cause the devices to gradually return any modified parameters to the default values. The actions that must be performed are generated during the Reporting step and arc different from the recommendations (which arc intended for humans and arc more focused on predictive maintenance), as these are messages that are transmitted to the devices, wh ich either contain consequence mitigation actions (e.g. modification of the refresh rate of the flash memory or other device specific actions) performed automatically by the devices themselves, or direct corrective actions, such as the transmission of replacement files, which should be either accessible by the Device action unit or they will be included within the messages by the command analysis software, utilizing the storage space for these files (p 16).

[0066] The discussion above has presented in detail the method and the system that predicts, in a collaborative manner and in real time the failures of the non-volatile data storage of embedded devices. These predictions allow (a) the predictive maintenance of the devices and also offer an estimation of the impact of the maintenance or the failure (e.g. downtime, service interruption) and (b) the execution of actions that prevent the failures and also mitigate their effects. This is achieved through the process, analysis and correlation in real time, via machine learning methods, of data streams containing messages, which are transmitted by the devices and contain their state in relation with their flash memory chip based storage.

[0067] Embodiment examples of the invention

[0068] Two indicative examples ofpoten tial embodiments ofthe invention will be presented below, which differ on the flash memory technology used by the monitored devices, the machine learning algorithms that have been selected for classification and the number of data recep tors emp loyed.

[0069] The first example embodiment of the invention involves the implementation that monitors telematics devices that reside on bus stations in a city. Every device contains eMMC flash memory storage with firmware that has been modified appropriately, in order to include the Flash status transmission unit and, in cooperation with the device driver of the flash (that has been also modified), can transmit to the Controller the following parameters within each message: (a) the number of bad blocks, (b) the blocks / second that have been read , (c) the blocks / second that have been written, (d) the serial number of the device, (c) the uptime of the memory, (f) the number of flash memory chip errors that have been corrected through the ECC, (g) the number of non -correctable errors that have been detected, (h) the serial number and the model of the memory, (i) the lifetime estimation of the internal Hash device, (j) the refresh rate of the memory, (k) the temperature of the device and (1) the physical location of the device. The devices send the messages t hrough a proprietary wireless radio network towards two Data receptor units. Specifically , the devices that are located atthesouthempart ofthe city transmitthe messages to the southern Data receptor and, similarly, the northern Data receptor receives the messages of the devices that are located at the northern part. Within the

[0070] Data receptor the messages are converted from the format received from the devices (binary structures) to structures that can be processed by the rest of the system (JSON messages), the scoreboard is updated and the relevant information of the scoreboard is added at the appropriate messages . Then, the messages from the two Data receptors are forwarded to the rest of the BEReTiC steps and specifically the clustering step. Concurrently with the clustering, SCoDe- monitors the data stream, collects the sample, and, at the same time, calculates the mode of the distance between its messages and the standard deviation. For the clustering the CluNN algorithm is implemented. Thus, until the final mode o f minimum distances of the sample will be calculated, the running result of SCoDe2for the clustering is used. For the classification in the detected clusters, 3 convolutional neural networks are used with different parameters, where their training is performed against the groups detected on the sample, and the K-Nearest Neighbor algorithm, KNN. For the time required for the neural networks to be trained, KNN is used. After the training is completed, each message is classified based on the classification by the majority of neural networks, using weighted voting, in which the K.NN algorithm has a larger priority . At the devices that, have sent messages that have been classified into h igh risk clusters, the appropriate commands are sent, in order for the Command execu tor to mitigate the consequences, or their predictive maintenance is scheduled, if the Summary structure that corresponds to the cluster that they have been classified in to contains non- correctable errors. To the devices that have not sent high risk messages, all-clear messages are sent. The reports contain a summary of the actions p erformed , statistics regarding the state o f the memory units of the devices and, additionally, a correlation of the state of the memory with the location of the devices, their uptime and their temperature, is presented.

[0071] The second embodiment focuses on smart electric devices that are interconnected through the Internet and exchange information regarding energy consumption. Every device contains bare NAND flash memory, whose driver have been ap propriately modified, in order to transmit to the Flash status transmission unit the parameters that include (a) the number of bad blocks, (b) the blocks / second that have been read, (c) the blocks / second that have been written, (d) the model and the serial number of the device, (e) the uptime of the memory, (f) the number of errors that have been corrected through the ECC, (g) the number of non -correctable errors that have been detected, (h) the p hysical addresses o f the chip where the errors reside, (i) the files that contain the data where the errors have been detected, (j) the number of times these files have been accessed, (k) the number of read requests, (I) the number of write requests, (m) the refresh rate of the memory and (n) the temperature of the device. The messages are sent via the Internet to the Data recep tor, who does not need to p er form an adaptation, since the messages arrive at the format that can be forwarded to the next step s, but the scoreboard is up dated accordingly, and its information is included in the messages. The messages are then forwarded to the clustering step.

[0072] Concurrently with the clustering, SCoDc2monitors the message stream, collects the sample and at the same time calculates the mode ofthe distance between its messages and the standard deviation. For the clustering the CluNN algorithm is implemented. Thus, until the final mode of minimum distances ofthe sample will be calculated, the running result of SCoDe2for the clustering is used. Regarding the classification to the discovered clusters, the KNN algorithm is employed. At the devices that have sent messages that have been classified into high risk clusters, the appropriate commands are sent, in order for the Command executor to mitigate the consequences, or the formatting ofthe memory ofthe device is scheduled, if the Summary structure th at corresponds to the cluster that they have been classified into contains non-correctable errors. To the devices that have not sent high risk cluster messages, all-clear messages are sent. The reports contain a summary ofthe actions performed, statistics regarding the stale of the memory units of the devices and, additionally, a correlation of the state of the memory with (a) the model of the device, (b) the files that contain the data where the errors have been detected and (c) temperature of the device, is presented.

[0073] The embodiments described herein are intended to be illustrative, not limiting the invention. Those skilled in the art shall be able to easily recognize various modifications and alterations that can be performed at the present invention, which lie within the scope of the disclosure and application of the present invention.

Claims

AMENDED CLAIMS received by the International Bureau on 13 May 2025 (13.05.2025)1. Method for mitigating flash memory failures [BEReTiC, ml], implemented on a computer system, comprising the following stepsSTEP 1: receiving status messages, from a set of flash memory units, included in embedded or Internet of Things (loT) devices, from a message reception unit [m5], formatting of the status messages in a form suitable for further processing at the next steps [s9] setting up and updating a scoreboard [slO] with essential statistics for each device related with the device identity and its communication failures forwarding the messages from the message reception unit to the next stepsSTEP 2: selection of a representative sample [p4] from the received messages by a sample collector unit, [SCoDe2, m6] detection of deviations by a deviation detector unit [SCoDe2, m6], which calculates and monitors a minimum distance metric, along with the statistical properties of the sample [p5, p6]STEP 3: clustering of the collected sample [m7] and formation of clusters, based on a calculation of the summary structure of the clusters, said calculation based on said minimum distance metricsSTEP 4: classification [pl 1] of every message included in the data stream [m8] of the data reception step, whether it was employed in the context of the selection of the representative sample or not, in one of the clusters detected during the clusteringSTEP 5: creation of reports, generation of recommendations and estimation of the time required for repairs [m9], according to the findings of the clustering and classification steps and identification of high-risk clustersSTEP 6: (a) transmission of messages to devices that correspond to high- risk clusters [mlO], in order to modify functional parameters and / or execute commandsSTEP 6: (b) transmission of all-clear messages at the devices that do not correspond to high-risk clusters [mlO], which cause the devices to gradually return any modified parameters to their initial values.

2. A method, according to claim 1 , where the modification of the parameters and / or the execution of commands are performed through action executor units that reside within the embedded or loT devices [m4].

3. A method, according to claim 1 or 2, where the clustering and the classification procedures of the status messages are performed through machine learning methods [p9] [pl 1].

4. A system for the mitigation of failures and the predictive maintenance of data storage devices included in interconnected embedded devices or Internet of Things (loT) devices, which implements the method of claims 1-3 and comprising the following units a data receptor unit that accepts status messages from a set of flash memory chips that reside in embedded or loT devices [m5], each of said status messages containing a minimum distance metric a unit operative to adapt the status messages to the proper structure required for processing [s9], a unit operative to set up and update a scoreboard that includes essential statistics and information for each device regarding device identity and communication failures of this device[slO], a message forwarding unit from the data receptor unit to the next steps [sl l], a representative sample selection unit from the messages to the sample collection unit [p4],a deviation detection unit, which calculates and monitors the statistical properties of the sample [p5, p6], a clustering unit operative on the collected sample, in order to create clusters [m7], said unit operative to calculate a summary structure of the clusters, said calculation based on said minimum distance metrics a classification unit, operative to classify every message included in the data stream of the data receptor step, regardless of whether the message was employed within the representative sample collection unit or not, to the clusters detected during the clustering [m8], a unit for report creation, recommendation generation and estimation of the time required for maintenance, according to the findings of the clustering and classification steps, and the identification of high-risk clusters [m9], a message transmission unit [mlO] to the devices that have been classified into high-risk clusters, in order to modify parameters and / or execute commands, an all-clear message transmission unit [mlO] to the devices that have not been classified as high risk, causing them to gradually return any modified functional parameters to their default values.

5. A system, according to claim 4, which additionally includes an action executor unit, integrated or within the embedded or loT devices [m4].

6. A system, according to any of claims 4 or 5, where the unit responsible for the clustering of the collected sample and the creation of the clusters operate through machine learning methods [CluNN, aO],

Citation Information

Patent Citations

  • Apparatus, method, and system for providing a sample representation for event prediction

    EP3848859A1

  • Systems and / or methods for dynamic anomaly detection in machine sensor data

    US20160342903A1

  • Ensemble risk assessment method for networked devices

    US20200042370A1