Method and system for recognizing patterns in failures of hardware drives
The method synchronizes and analyzes data from multiple sources with varying formats to identify hardware drive failures, enhancing predictive maintenance and reducing human intervention.
Patent Information
- Application Number
- PCT/IB2024/050207
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-07-17
AI Technical Summary
Conventional methods lack a common platform to combine data in different formats for analyzing hardware drive failures and fail to identify or predict the cause of failures effectively.
A method and system that synchronizes input data from multiple sources with varying formats and sampling intervals, using a processor to generate and recognize data patterns by comparing with historical data and ground truth data to identify failure events.
Enables efficient recognition of failure patterns, facilitating predictive maintenance, root cause analysis, and minimal human intervention in analyzing hardware drive failures.
Smart Images

Figure IB2024050207_17072025_PF_FP_ABST
Abstract
Description
“METHOD AND SYSTEM FOR RECOGNIZING PATTERNS IN FAILURES OF HARDWARE DRIVES”TECHNICAL FIELD
[0001] The present disclosure generally relates to failure management in electrical components. In particular, the present disclosure relates to a method and system for recognizing a pattern in failures of the hardware drives.BACKGROUND
[0002] The following description includes information that may be useful in understanding the present disclosure. It is not an admission that any of the information provided herein is prior art or relevant to the presently claimed disclosure, or that any publication specifically or implicitly referenced is prior art.
[0003] Electrical components including, but not limited to, hardware drives may fail due to various factors such as aging of components, sudden spike in voltage, power issues, overheating, or alike. A human expert or a service engineer may collect data including signal data, event logs, data from data loggers, switching patterns etc., to analyze the root cause of the failures. Conventionally, the data may be present in different formats and the human expert may analyze each data individually to identify the root cause of the failures in the hardware drives. However, conventional approaches do not provide a common platform to combine the data present in different formats to analyze the cause of the failure in the hardware drives. Also, currently there exists no mechanism to identify or predict the cause of failure by combining the data present in different formats. Moreover, the existing solutions also lack ability to label the patterns of data that led to the failure of hardware components.
[0004] Therefore, there exists a need for an improved method and system for recognizing patterns of failures in the hardware drives, which overcomes the above-mentioned limitations of the conventional methods.
[0005] The above-mentioned drawbacks / difficulties / disadvantages of the conventional techniques are explained just for exemplary purpose and this disclosure and description mentioned below would never limit its scope only such problem. A person skilled in the art may understand that this disclosure and below mentioned description may also solve other problems or overcome the above-mentioned drawbacks / disadvantages of the conventional arts which are not explicitly captured above.SUMMARY
[0006] The present disclosure overcomes one or more shortcomings of the prior art and provides additional advantages discussed throughout the present disclosure. Additional features and advantages are realized through the techniques of the present disclosure. Other embodiments and aspects of the disclosure are described in detail herein and are considered a part of the claimed disclosure.
[0007] In one non-limiting embodiment of the present disclosure, a method of recognizing patterns in failures of hardware drives is disclosed. The method comprises receiving input data related to a plurality of parameters from a plurality of input sources upon detecting occurrence of a failure event. The input may be captured in a plurality of data formats with different sampling intervals. Further, the method comprises synchronizing the input data related to each of the plurality of parameters with reference to a timestamp of occurrence of the failure event. Thereafter, the method may comprise step of generating one or more data patterns by comparing the input data with historical data of each of the plurality of parameters. Finally, the method comprises recognizing the one or more data patternsindicating the failure events in the hardware drives based on mapping each of the one or more data patterns with a ground truth data.
[0008] In one non-limiting embodiment of the present disclosure, a failure analysis system for recognizing patterns in failures of hardware drives is disclosed. The failure analysis system comprises a memory, a communication interface, and a processor coupled to the memory to the communication interface. The communication interface is configured to establish communication with a control system associated with the hardware drives. The processor is configured to receive input data related to a plurality of parameters from a plurality of input sources associated with the control system upon detecting occurrence of a failure event. The input data is captured in a plurality of data formats with different sampling intervals. Further, the processor is configured to synchronize the input data related to each of the plurality of parameters with reference to a timestamp of occurrence of the failure event. Thereafter, the processor is configured to generate one or more data patterns by comparing the input data with historical data of each of the plurality of parameters. Alternatively, the processor is configured to recognize the one or more data patterns indicating the failure events in the hardware drives based on mapping each of the one or more data patterns with a ground truth data.
[0009] Furthermore, the present disclosure relates to a non-transitory computer readable medium including instructions stored thereon that when processed by at least one processor, cause a sequence designing system to perform operations comprising receiving input data related to a plurality of parameters from a plurality of input sources upon detecting occurrence of a failure event. The input data is captured in a plurality of data formats with different sampling intervals. Further, the instructions cause the processor to synchronize the input data related to each of the plurality of parameters with reference to a timestamp of occurrence of the failure event. Thereafter, the instructions cause the processor to generate one or more data patterns by comparing the input datawith historical data of each of the plurality of parameters. Finally, the instructions cause the processor to recognize the one or more data patterns indicating the failure events in the hardware drives based on mapping each of the one or more data patterns with ground truth data.
[0010] The foregoing summary is illustrative only and is not intended to be in any way limiting. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description.BRIEF DESCRIPTION OF THE ACCOMPANYING DRAWINGS
[0011] The embodiments of the disclosure itself, as well as a preferred mode of use, further objectives and advantages thereof, will best be understood by reference to the following detailed description of an illustrative embodiment when read in conjunction with the accompanying drawings. One or more embodiments are now described, by way of example only, with reference to the accompanying drawings in which:
[0012] Fig. 1 shows an exemplary architecture configured for recognizing patterns of failures in hardware drives, in accordance with an embodiment of the present disclosure;
[0013] Fig. 2 shows a detailed block diagram of a failure analysis system for recognizing patterns of failures in hardware drives, in accordance with an embodiment of the present disclosure;
[0014] Fig. 3a shows a flowchart depicting a method of recognizing patterns of failures in hardware drives, in accordance with an embodiment of the present disclosure;
[0015] Fig. 3b shows an exemplary flowchart depicting synchronization of data in different data formats in hardware drives, in accordance with an embodiment of the present disclosure; and
[0016] Fig. 4 illustrates a block diagram of an exemplary computer system for implementing embodiments consistent with the present disclosure.
[0017] The figures depict embodiments of the disclosure for purposes of illustration only. One skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods illustrated herein may be employed without departing from the principles of the disclosure described herein.DETAILED DESCRIPTION
[0018] The foregoing has broadly outlined the features and technical advantages of the present disclosure in order that the detailed description of the disclosure that follows may be better understood. It should be appreciated by those skilled in the art that the conception and specific embodiment disclosed may be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure.
[0019] The novel features which are believed to be characteristic of the disclosure, both as to its organization and method of operation, together with further objects and advantages will be better understood from the following description when considered in connection with the accompanying figures. It is to be expressly understood, however, that each of the figures is provided for the purpose of illustration and description only and is not intended as a definition of the limits of the present disclosure.
[0020] In an embodiment, the present disclosure provides a method and system for recognizing patterns in failures of hardware drives. Further, the present disclosure helps in synchronizing data present in different data formats to recognize data patterns that lead to the failures in the hardware drives. In one embodiment, the present disclosure helps in generating labelled data for prediction offailures in the hardware drives. To recognize the failure patterns and generate the labelled data, the method comprises receiving input data related to a plurality of parameters from a plurality of input sources upon detecting occurrence of a failure event. The input data may be captured in a plurality of data formats with different sampling intervals. Further, the method comprises synchronizing the input data of each of the plurality of parameters with reference to a timestamp of occurrence of the failure event. Thereafter, the method comprises generating one or more data patterns by comparing the input data of each of the plurality of parameters with historical data of each of the plurality of parameters. Finally, the method comprises recognizing the one or more data patterns indicating the failure events in the hardware drives based on mapping each of the one or more data patterns with a ground truth data.
[0021] The present disclosure provides an improved method for recognizing patterns of failures in hardware drives to identify the root cause of the failures. Also, the present disclosure provides a platform for performing predictive maintenance, root cause analysis, failure detection, anomaly detection, drive health prediction, age prediction etc., of the hardware drives. Additionally, the present disclosure provides a method and system for efficiently synchronizing data collected in different data formats with minimal human intervention.
[0022] Fig. 1 shows an exemplary architecture 100 configured for recognizing patterns of failures in one or more hardware drives, in accordance with an embodiment of the present disclosure. The exemplary architecture 100 may comprise, without limiting to, a failure analysis system 102, a control system 104 having the one or more hardware drives 106 and input sources 107, a user device 110a, a user device 110b, communication network 112a, 112b. In some embodiments, the architecture 100 may include other components, in addition to the components shown in Fig. 1, which may be configured to implement the subject matter disclosed in the present disclosure. According to theaspects of the present disclosure, the failure analysis system 102 may be used as a cloud-based platform for recognizing patterns of failures occurring in one or more hardware drives 106 associated with the control system 104. The failure analysis system 102 may recognize the patterns of failures by analyzing the data collected from different input sources 107 associated with the control system 104. The data collected from the input sources 107 may be in different data formats and may be collected at different time intervals. In some embodiments, the cloud-based platform on which the failure analysis system 102 may be implemented may include, without limiting to, a public cloud infrastructure, a private cloud infrastructure, and a hybrid cloud infrastructure.
[0023] In one implementation, the user device 110a present at a remote location may access the failure analysis system 102 using the communication network 112b. As an example, a user of the user device 110a may be a service expert. According to the aspects of the present disclosure, the user of a user device 110b may access and analyze the data collected from the input sources 107 to generate ground truth data 108 related to the failures occurring in the one or more hardware drives 106 of the control system 104. As an example, the of the user device 110b may be a service engineer or a service expert operating locally at the location of the control system 104. In some embodiments, the failure analysis system 102 may be configured to collect the ground truth data 108 and the data from the input sources 107 to recognize a pattern in the failures occurring in the one or more hardware drives 106.
[0024] In a non-limiting embodiment, the failure analysis system 102 may be configured to communicate with the control system 104 using the communication network 112a. In an aspect of the present disclosure, the control system 104 may comprise, but not limited to, the one or more hardware drives 106, the input sources 107 along with other devices, systems, networks, electrical components which are configured to operate in an industrial process or industrial control systems. However, in some implementations, aspects disclosed in the present disclosure may be equally applicable to othersystems, components, or devices, where root cause of failures needs to be identified. In some embodiments, the failure analysis system 102 may communicate with the control system 104 via the communication network 112a. In some embodiments the communication network 112a may be a wireless means or a wired means.
[0025] In an alternative implementation, the failure analysis system 102 may be configured within the control system 104 for recognizing the pattern of failure in the one or more hardware drives 106. The failure analysis system 102 may be configured to receive input data related to a plurality of parameters from the plurality of input sources 107 of the control system 104. In some embodiments, the plurality of input sources 107 may include, but not limited to, sensors, controllers, actuators, or any other components used for measuring readings before and after occurrence of failure in the one or more hardware drives 106. In some embodiments, the plurality of parameters may include, but not limited to, black box data, data from data loggers, settings data, signals data and events data etc. In some embodiments, the input data may include values, readings, etc., related to each of the plurality of parameters. The input data may be captured in a plurality of data formats with different sampling intervals.
[0026] The method and system disclosed in the present disclosure plays an important role in assisting a service expert located at one location for analyzing the root cause of failure or recognizing data patterns of failures occurring in the one or more hardware drives 106 present at another location. The failure analysis system 102 collects data from the input sources 107 via the communication network 112a. The service engineer located in a different location may login to the failure analysis system 102 for accessing the data and may continue with the analysis of the failures occurring in the one or more hardware drives 106. In this manner, the service engineers located in different locations may not even require accessing the control system or the one or more hardware drives 106 physically and mayperform all the analysis with the help of the remote failure analysis system 102. In an exemplary embodiment, suppose a failure has occurred in one of the hardware drives 106 of the control system 104 deployed in Germany. The service expert or service engineer located in Germany obtains the failure data related to the failed hardware drive via the communication network 112a. If the service expert located in Germany is unable to resolve the failure or identify the root cause of the failure, a service engineer located in Singapore may continue with the analysis by remotely connecting to the failure analysis system 102. The ground truth data 108 may be populated, and data patterns of failure may be identified by the service expert located in Singapore by accessing the cloud-based failure analysis system 102. The detailed explanation of the failure analysis system 102 is explained in detail in the paragraphs below.
[0027] Fig. 2 shows a detailed block diagram of the failure analysis system 102 for recognizing patterns of failures in the one or more hardware drives 106, in accordance with an embodiment of the present disclosure. In some implementations, the failure analysis system 102 may include a processor 204, an Input / Output interface 206, one or more modules 208, a memory 210, and data 212. The failure analysis system 102 may comprise additional components which may be configured to implement the subject matter disclosed in the present disclosure. As used herein, the term processor 204 may refer to an Application Specific Integrated Circuit (ASIC), an electronic circuit, a hardware processor (shared, dedicated, or group) and memory that execute one or more software or firmware programs, a combinational logic circuit, and / or other suitable components that provide the described functionality. In some embodiments, the processor 204 may be configured to perform one or more functions of the failure analysis system 102 for recognizing pattern of failures in the one or more hardware drives 106.
[0028] In an embodiment, the I / O interface 206 may include a variety of software and hardware interfaces, for example, a web interface, a graphical user interface, an input device, an output device, and the like for establishing an interconnection between the failure analysis system 102, the control system 104, and the user device 110a associated with the user located at a remote location. In some embodiments, the I / O interface 206 may be configured to transmit / receive input data to / from the control system 104 and user device 110a over the communication network 112a and the communication network 112b respectively.
[0029] Further, as used herein, the one or more modules 208 may comprise, but not limited to, a receiving module 214, a synchronization module 216, a generation module 218, a recognition module 220, a machine learning model 222, and other modules 224. The term one or more modules 208 may refer to an Application Specific Integrated Circuit (ASIC), an electronic circuit, a hardware processor (shared, dedicated, or group) and memory that execute one or more software or firmware programs, a combinational logic circuit, and / or other suitable components that provide the described functionality. In one implementation, each of the modules 208 may be configured as stand-alone hardware computing units. In an embodiment, the other modules 224 may include one or more modules that may be used to perform various miscellaneous functionalities of the failure analysis system 102. It shall be appreciated that such module may be represented as a single module or a combination of different modules.
[0030] In a non-limiting embodiment, the memory 210 may be an external memory chip or an inbuilt Electrically Erasable Programmable Read-only Memory (EEPROM) memory, within the failure analysis system 102. In an embodiment, the memory 210 may be a computer-readable medium known in the art including, for example, volatile memory, such as Static Random-Access Memory (SRAM) and Dynamic Random-Access Memory (DRAM), and / or Synchronous Dynamic Random-AccessMemory (SDRAM) and / or non-volatile memory, such as Read Only Memory (ROM), erasable programmable ROM, flash memories, hard disks, optical disks, and magnetic tapes. In some embodiments, the memory 210 may be communicatively coupled to the processor 204 and may be configured to store various data processed by the processor 204.
[0031] In an embodiment, the data 212 stored in the memory 210 may include, without limitation, input data 226, a plurality of parameters 228, one or more data patterns 230, historical data 232, and labelled dataset 234. In some embodiments, the data 212 may comprise other data apart from the above-mentioned data. The other data may include, but not limiting to, fault groups, drive trip instances, timestamp of occurrence of a failure event, etc. In some implementations, the data 212 may be stored within the memory 210 in the form of various data structures. Additionally, the data 212 may be organized using data models, such as relational or hierarchical data models. The other data may include various temporary data and files generated by the processor 204 while performing various functions of the failure analysis system 102.
[0032] According to aspects of the present disclosure, the receiving module 214 may be configured to receive input data 226 related to the plurality of parameters 228 from a plurality of input sources 107 upon detecting occurrence of the failure event. In some embodiments, the input data 226 may be captured or obtained from sensors, actuators, or alike which are configured in the control system 104. In some embodiments, the plurality of parameters 228 may comprise, but not limited to, event data and related warnings, settings data, blackbox or switching data, data loggers’ data, and signals data. In some embodiments, the input data 226 related to the event data and warnings may include, without limiting to, current or voltage values, temperature values, or any other equivalent measured readings. In some embodiments, the user associated with the user device 110b may capture the input data 226 in a plurality of data formats with different sampling intervals. As an example, the user device 110bmay capture the events data and related warnings at periodic intervals. In an embodiment, the user device 110b may collect the event data and related warnings for a period of 2 months. In another embodiment, the user device 110b may obtain the event data and related warnings as a plurality of samples of observations. In yet another embodiment, the user device 110b may record or capture the event data and related warnings when an event occurs with a timestamp precision up to ‘seconds’. In some embodiments, the user device 110b may capture the event data and related warnings from the sensors with respect to Universal Time Coordinated (UTC) timestamp. In some embodiments, the user device 110b may upload the collected event data and related warnings to the failure analysis system 102 or a cloud storage, where it may be stored as a ‘.csv’ file.
[0033] Subsequently, the user of the user device 110a operating from a remote location may fetch the ‘.csv’ file from the cloud whenever required. In some embodiments, the receiving module 214 may be configured to identify one or more drive trip instances comprised in the input data 226. Upon identifying the drive trip instances, the receiving module 214 may be configured to segregate each of the one or more drive instances into predefined fault groups. In some embodiments, the fault groups may comprise but not limited to, short circuit, over voltage, capacitor failure etc. Thereafter, the receiving module 214 may be configured to detect the occurrence of the failure event by identifying one of the predefined fault groups with deviated frequency distribution. In an exemplary embodiment, the receiving module 214 may identify frequency distribution in certain conditions such as over voltage condition, short circuit, and capacitor failure. To identify frequency distribution, the receiving module 214 may be configured to identify changes in values or readings for the over voltage condition, the short circuit condition and the capacitor failure condition. Further, the receiving module 214 may be configured to study frequency distribution of the over voltage condition, the short circuit condition and the capacitor failure condition. Further, the receiving module 214 may be configured to identifythe frequency distribution with highest deviation for one fault group, among the over voltage condition, the short circuit condition and the capacitor failure condition. In this manner, the receiving module 214 may select a target fault, which may have led to the failure in the one or more hardware drives 106 based on the frequency deviation.
[0034] In some embodiments, the frequency of the failure event may be less than a predefined threshold but may still be the major reason behind the failure in the one or more hardware drives 106. In an alternative embodiment, the frequency of the failure event may exceed the predefined threshold and may be considered as the major reason behind the failure in the one or more hardware drives 106. The input data 226 from the event data and the settings data may be used to generate a first set of patterns.
[0035] In some embodiments, the receiving module 214 may be further configured to receive input data 226 related to settings data. The input data 226 related to settings data may comprise, but not limited to, limits, threshold value, nominal values, and other settings. In some embodiments, the receiving module 214 may be configured to store the settings data as a text file on the cloud. The settings data may be a snapshot of signals and settings available in the one or more hardware drives 106. As an example, the receiving module 214 may be configured to collect the data every day at 12 am and during the instance of triggering event occurring in the one or more hardware drives 106. In some embodiments, whenever the failure event occurrence may be identified, the receiving module 214 may be configured to compare the settings data occurred on the day of failure occurrence with the settings data of previous day to determine any change in the nominal values or the settings in the one or more hardware drives 106.
[0036] Further, the receiving module 214 may be configured to receive input data 226 related to switching data 228c or blackbox data. In some embodiments, the switching data 228c may be morereliable because it may be directly collected from the one or more hardware drives 106 instead of sensors. In some embodiments, the switching data 228c may comprise fault information, warnings, notifications, switching patterns, and parameters related to the one or more hardware drives 106, which may be recorded for a duration of 6 months to 1 year before the occurrence of the failure event. The input data 226 from the switching data may be used to generate a second set of patterns.
[0037] In some embodiments, the receiving module 214 may be further configured to receive signals data with different sampling rates, i.e., in seconds, minutes, or hours. The receiving module 214 may be configured to store signals data as parquet files on a cloud storage. Further, the receiving module 214 may be configured to map the signal data around the target root fault for identifying the signals having major impact in causing the fault or failure event. In some embodiments, the signals data may help in identifying and studying the patterns or changes in the values of the signals for analyzing failures in the one or more hardware drives 106.
[0038] In some embodiments, the receiving module 214 may be configured to receive data loggers’ data. In some embodiments, the data logger’s data comprises two files, namely a text format file having fault occurrent details and a ‘.csv’ file having the signal data. The receiving module 214 may be configured to store data loggers’ data on the cloud storage. Further, unlike the other data mentioned above, the timestamp of the data logger’s data provides a precision in microseconds, seconds, minutes, or hours. Further, the signals around the target root fault may help in identifying signals that lead to the cause of failure in the one or more hardware drives 106 with finer details of change in values of the signals. Thus, the receiving module 214 may be configured to map the combination signals including the data logger’s data and signals with the target root fault for obtaining a third set of patterns.
[0039] In some embodiments, a service engineer or service expert may record ground truth data 108 which may be stored, for example, in a ‘.pdf’ file. The service engineer may record the cause of failure by performing an on-site analysis of the one or more hardware drives 106. In some embodiments, the ground truth data 108 may comprise, without limiting to, a serial number of the hardware drive, date, symptoms, details of components failed etc.
[0040] In some embodiments, the synchronization module 216 may be configured to synchronize the input data 226 of each of the plurality of parameters 228 with reference to a timestamp of occurrence of the failure event. Once the input data 226 related to the plurality of parameters 228 and the ground truth data 108 may be collected, changes in the input data 226 may be identified before the occurrence of the failure event and after the occurrence of the failure event. Thereafter, the synchronization module 216 may be configured to identify changes in the input data 226 based on the ground truth data 108 and signal data associated with the failure event of the one or more hardware drives 106. In an exemplary embodiment, the synchronization module 216 may be configured to identify peak values of signals and minimum values of signals to determine root cause of failure. Subsequently, the generation module 218 may be configured to generate the one or more data patterns 230 by comparing the input data 226 of each of the plurality of parameters 228 with the historical data 232 of each of the plurality of parameters 228. In some embodiments, the historical data 232 of each of the plurality of parameters 228 may comprise data collected before occurrence of the failure event. In some embodiments, the historical data 232 may comprise data collected one hour or even a few seconds before the occurrence of the failure event. In some embodiments, the historical data 232 may comprise, but not limited to, data collected for a period of 1 year or more before the occurrence of the failure event.
[0041] In some embodiments, the recognition module 220 may be configured to recognize the one or more data patterns 230 indicating the failure events in the one or more hardware drives 106 based on mapping of each of the one or more data patterns 230 with the ground truth data 108. In some embodiments, the recognition module 220 may be further configured to generate the labelled dataset 234 by assigning a unique fault label to the one or more data patterns 230 associated with a particular failure event based on the pattern recognition. In an exemplary embodiment, whenever a failure occurrence may be identified, with the help of the labelled dataset 234, the service engineer may identify the type of failure by analysing the data pattern that has led to the cause of the failure in the one or more hardware drives 106.
[0042] Thus, the present disclosure provides an improved method for recognizing one or more data patterns 230 of failures in the one or more hardware drives 106 to identify the root cause of the failure. Further, the aspects disclosed in the present disclosure provide the method and system for predicting the occurrences of failures in the one or more hardware drives 106 in future. Also, the present disclosure provides a platform for predictive maintenance of the one or more hardware drives 106, root cause analysis, failure detection, anomaly detection, drive health prediction, age prediction of the one or more hardware drives 106. Additionally, the present disclosure provides a method and system for synchronizing different data formats of failures of the one or more hardware drives 106 with minimal human intervention.
[0043] Fig. 3a depicts a method 300 of pattern recognition of failures in the one or more hardware drives 106, in accordance with an embodiment of the present disclosure. The method 300 may be described in the general context of computer executable instructions. Generally, computer executable instructions may include routines, programs, objects, components, data structures, procedures, modules, and functions, which perform specific functions or implement specific abstract data types.
[0044] The order in which the method 300 is described is not intended to be construed as a limitation, and any number of the described method blocks may be combined in any order to implement the method. Additionally, individual blocks may be deleted from the methods without departing from the spirit and scope of the subject matter described. Further, it may be noted that the method 300 may be performed by the processor 204 or any other modules of the failure analysis system 102, as shown in Fig. 2. The description of Fig. 3a is provided with reference to Figs. 1-2.
[0045] At block 302, the method 300 may comprise receiving input data 226 related to a plurality of parameters 228 from a plurality of input sources 107 upon detecting occurrence of a failure event. In some embodiments, the method at block 302 may be performed by the processor 204 or the receiving module 214. In some embodiments, the plurality of parameters 228 may comprise, but not limited to, switching data, data loggers’ data, settings data, signals data and events data. In some embodiments, the plurality of input data 226 may be captured in a plurality of data formats with different sampling intervals.
[0046] Further, the method 300 at block 304 may comprise synchronizing the input data 226 of each of the plurality of parameters 228 with reference to a timestamp of occurrence of the failure event. In some embodiments, for synchronizing the input data 226 of each of the plurality of parameters 228, the method may comprise identifying one or more drive trip instances based on event data and related warnings comprised in the input data 226 of the plurality of parameters 228. Further, the method 300 may comprise segregating each of the one or more drive trip instances into predefined fault groups. Thereafter, the method 300 may detect the occurrence of the failure event by identifying one of the predefined fault groups with deviated frequency distribution. Additionally, the method 300 may comprise identifying changes in the input data 226 before the occurrence of the failure event and after the occurrence of the failure event. Finally, the method 300 synchronizes the changes in the input data226 based on the ground truth data 108 and signals data associated with the failure event. The method 300 at block 304 may be performed by the synchronization module 216 or the processor 204.
[0047] After synchronizing the input data 226, the method 300 at block 306 may comprise generating one or more data patterns 230 by comparing the input data 226 of each of the plurality of parameters 228 with the historical data 232 of each of the plurality of parameters 228. In some embodiments, the one or more data patterns 230 may represent changes in the readings or values or changes in the data related to the plurality of parameters 228 of at the time of failure occurrence. In some embodiments, the historical data 232 of each of the plurality of parameters 228 may comprise data collected before occurrence of the failure event. In some embodiments, the method 300 at block 306 may be performed by the generation module 218 or the processor 204.
[0048] In some embodiments, at block 308, the method 300 may comprise recognizing the one or more data patterns 230 indicating the failure events in the one or more hardware drives 106 based on mapping of each of the one or more data patterns 230 with a ground truth data 108. In some embodiments, the ground truth data 108 may indicate an actual root cause of the failure event determined by a human expert associated with the failure analysis system. The method 300 at block 308 may be performed by the recognition module 220.
[0049] In some embodiments, the method 300 may include generating a labelled dataset 234 by assigning a unique fault label to the one or more data patterns 230 associated with a particular failure event based on the pattern recognition. Additionally, the method 300 may include training the machine learning model 222 with the labelled dataset 234 for performing at least a predictive maintenance of the one or more hardware drives 106.
[0050] Fig. 3b shows an exemplary flowchart depicting synchronization of data in different data formats 230 in the one or more hardware drives 106 for recognizing patterns of failure event in the one or more hardware drives 106, in accordance with an embodiment of the present disclosure. Initially, the input data 226 related to a plurality of parameters 228, which are in different data formats, may be received upon detecting occurrence of the failure event. In an exemplary embodiment, the plurality of parameters 228 may comprise but not limited to, event data and related warnings 228a, settings data 228b, switching data 228c, data loggers’ data 228d. In a non-limiting embodiment, the input data 226 for event data and related warnings 228a may be captured from the plurality of input sources 107 and stored as csv file on a cloud. In the event data and related warnings 228a, one or more drive trip instances may be identified for a short duration of period. In some embodiments, the event data and related warnings 228a may be recorded twice or thrice in a day and approximately 10-12 samples may be considered for the analysis. Further, at step 312, each of the one or more drive trip instances may be segregated into predefined fault groups in step 314. For example, the observed drive trip may be segregated into groups of faults such as short circuit, spike in current or voltage, capacitor failure for a predefined period. Based on identifying at least one of the predefined fault groups with deviated frequency distribution, as shown in step316, the occurrence of the failure event is detected. In an exemplary embodiment, sometimes frequency of a failure event may be less but could be considered as major reason behind the failure in the one or more hardware drives 106. In another exemplary embodiment, the frequency distribution of at least three faults may be studied and the one with a maximum number of frequencies may be selected for a target root fault. Thereafter, the target root fault may be selected in step 318 and various other parameters changes around the target root fault may be analyzed in step 320.
[0051] In another non -limiting embodiment, the input data 226 for settings data 228b may be captured and stored as a text file on the cloud storage. The settings data may comprise a snapshot of all signals and drive settings available in the one or more hardware drives 106 that may comprise limits, threshold values, nominal values, and other settings related to the one or more hardware drives 106. For example, the settings data may be collected and recorded every day at 12am and during the instance of triggering in the one or more hardware drives 106. Whenever the failure event occurrence is identified, the settings data on the day of failure occurrence may be compared with the settings data of previous day to determine any change in the nominal values or settings. The data from the event data and data 228c and the settings data may be used to generate a first set of patterns (i.e., pattern 1 shown in Fig. 3b).
[0052] In another non-limiting embodiment, the input data 226 for switching data 228c or blackbox data may include events from blackbox and may be same as events data and related warnings. Additionally, the switching data 228c may be more reliable because it is directly collected from the one or more hardware drives 106 instead of sensors. In some embodiments, in the switching data 228c fault, warnings, notifications and parameters may be recorded from 6 months to 1 year before the occurrence of the failure event. The data from the switching data 228c may be used to generate a second set of patterns (i.e., pattern 2 shown in Fig. 3b).
[0053] In some embodiments, the input data 226 may comprise signals data, which may be stored as parquet filed. The various signals in this data have different sampling rates in either seconds, minutes, or hours. The signals data helps in identifying and studying the patterns or changes in the values of the signals for analyzing the failure of the one or more hardware drives 106.
[0054] In another non-limiting embodiment, the input data 226 for data loggers’ data 228d data may include signal data. In some embodiments, the data logger’s data 228d comprises two files namely a‘text’ file having fault occurrent details and a ‘.csv’ file having the signal data. Unlike the other mentioned data, the timestamp of the data logger’s data 228d provides a precision in microseconds, seconds, minutes, or hours. Further, the signals around the target root fault may help in identifying signals that led to the cause of failure of the one or more hardware drives 106 with finer details of change in values of the signals. Thus, the combination signals including the data logger’s data 228d and signals mapped with the target root fault may help in obtaining a third set of patterns (i.e., pattern 3 as shown in Fig. 3b).
[0055] In another embodiment, a service engineer or service expert may record ground truth data 108 which may be stored in pdf format. The service engineer may record the cause of failure by an on-site analysis of the one or more hardware drives 106. In some embodiments, the ground truth data 108 may comprise, without limiting to, a serial number of the one or more hardware drives 106, date, symptoms, components failed, etc.
[0056] Once the data related to the plurality of parameters 228 and ground truth data 108 may be collected, changes in the input data 226 may be identified before the occurrence of the failure event and after the occurrence of the failure event. Thereafter, the changes in the input data 226 may be synchronized based on the ground truth data 108 and signal data associated with the failure event of the one or more hardware drives 106. Subsequently, one or more data patterns 230 indicating the failure events in the one or more hardware drives 106 may be recognized based on mapping of each of the one or more data patterns 230 with the ground truth data 108 as shown in Fig. 3b.
[0057] In some embodiments, the labelled dataset 234 may be generated by assigning a unique fault label to the one or more data patterns 230 associated with a particular failure event based on the pattern recognition. Accordingly, with the use of the labelled dataset 234, a machine learning model 222 may be trained for performing a predictive maintenance of the one or more hardware drives 106.
[0058] Thus, the present disclosure provides an improved method for recognizing the one or more data patterns 230 of failures in the one or more hardware drives 106 to identify the root cause of the failure. Further, the aspects disclosed in the present disclosure provide the method and system for predicting the occurrences of failures in the one or more hardware drives 106 in future. Also, the present disclosure provides a platform for predictive maintenance, root cause analysis, failure detection, anomaly detection, drive health prediction, age prediction etc., of the one or more hardware drives 106. Additionally, the present disclosure provides a method and system for synchronizing different data formats of failures of the one or more hardware drives 106 with minimal human intervention.
[0059] Computer System
[0060] FIG. 4 illustrates a block diagram of an exemplary computer system 400 for implementing embodiments consistent with the present disclosure. In an embodiment, the computer system 400 may be a failure analysis system 102 illustrated in Fig. 1, which may be used for recognizing patterns of failures in the one or more hardware drives 106. The computer system 400 may include a central processing unit (“CPU” or “processor”) 402. The processor 402 may comprise at least one data processor for executing program components for executing user- or system-generated business processes. The processor 402 may include specialized processing units such as integrated system (bus) controllers, memory management control units, floating point units, graphics processing units, digital signal processing units, and the like. In an embodiment, the processor 402 may be used to realize the processor 204 described in Fig. 2.
[0061] The processor 402 may be disposed in communication with one or more Input / Output (I / O) devices (404 and 406) via I / O interface 408. In some embodiments, the processor 402 may be disposed in communication with a communication network 112 via a network interface 410. The networkinterface 410 may communicate with the communication network 112. Using the network interface 410, the computer system 400 may connect with a control system 104 and a user device 110a via the communication network 112 for recognizing pattern of failure in the one or more hardware drives 106.
[0062] In an implementation, the communication network 112 may be implemented as one of the several types of networks, such as intranet or Local Area Network (LAN) and such within the organization. The communication network 112 may either be a dedicated network or a shared network, which represents an association of several types of networks that use a variety of protocols. In some embodiments, the processor 402 may be disposed in communication with a memory 418 (e.g., RAM 414, ROM 416 via a storage interface 412.
[0063] The memory 418 may store a collection of program or database components, including, without limitation, user / application interface 420, an operating system 422, a web browser 424, and the like. In some embodiments, computer system 400 may store user / application data 420, such as the input data 226, values, records, and the like, as described in this disclosure. In an embodiment, the memory 418 may be used to realize the memory 210 described in Fig. 2.
[0064] The operating system 422 may operation of the computer system 400. The user interface 420 may facilitate display, execution, interaction, manipulation, or operation of program components through textual or graphical facilities. For example, the user interface 420 may provide computer interaction interface elements on a display system operatively connected to the computer system 400.
[0065] The web browser 424 may be a hypertext viewing application. Secure web browsing may be provided using Secure Hypertext Transport Protocol (HTTPS), Secure Sockets Layer (SSL), Transport Layer Security (TLS), and the like.
[0066] A description of an embodiment with several components in communication with each other does not imply that all such components are required. On the contrary, a variety of optional components are described to illustrate the wide variety of possible embodiments of the disclosure.
[0067] When a single device or article is described herein, it will be clear that more than one device / article (whether they cooperate) may be used in place of a single device / article. Similarly, where more than one device or article is described herein (whether they cooperate), it will be clear that a single device / article may be used in place of the more than one device or article, or a different number of devices / articles may be used instead of the shown number of devices or programs. The functionality and / or the features of a device may be alternatively embodied by one or more other devices which are not explicitly described as having such functionality / features. Thus, other embodiments of the disclosure need not include the device itself.
[0068] Finally, the language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. It is therefore intended that the scope of the disclosure be limited not by this detailed description, but rather by any claims that issue on an application based here on. Accordingly, the embodiments of the present disclosure are intended to be illustrative, but not limiting, of the scope of the disclosure, which is set forth in the following claims.
[0069] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and are not intended to be limiting, with the true scope and spirit being indicated by the following claims.Reference Numerals:
Claims
The Claims:
1. A method of recognizing patterns in failures of hardware drives, the method comprising: receiving, by a failure analysis system, input data related to a plurality of parameters from a plurality of input sources, wherein the input data is captured in a plurality of data formats with different sampling intervals; synchronizing, by the failure analysis system, the input data related to each of the plurality of parameters with reference to a timestamp of occurrence of a failure event; generating, by the failure analysis system, one or more data patterns by comparing the input data with historical data of each of the plurality of parameters; and recognizing, by the failure analysis system, the one or more data patterns indicating the failure events in the hardware drives based on mapping each of the one or more data patterns with a ground truth data.
2. The method of claim 1, wherein the plurality of parameters comprises at least one of black box data, data loggers’ data, settings data, signals data and events data.
3. The method of claim 1, further comprising: identifying one or more drive trip instances based on event data and related warnings comprised in the input data of the plurality of parameters; segregating each of the one or more drive trip instances into predefined fault groups; and detecting the occurrence of the failure event by identifying one of the predefined fault groups with highest frequency distribution.
4. The method of claim 1, further comprising: generating a labelled dataset by assigning a unique fault label to the one or more data patterns associated with a particular failure event based on the pattern recognition.
5. The method of claim 4, further comprising training a machine learning model with the labelled dataset for performing at least a predictive maintenance of the hardware drives.
6. The method of claim 1, wherein the historical data of each of the plurality of parameters comprises data collected before occurrence of the failure event.
7. The method of claim 1, wherein the ground truth data indicates an actual root cause of the failure event determined by a human expert associated with the failure analysis system.
8. A failure analysis system for recognizing patterns in failures of hardware drives, the failure analysis system comprising: a memory; a communication interface configured to establish communication with a control system associated with the hardware drives; and a processor, coupled to the memory and the communication interface, wherein the processor is configured to: receive input data related to a plurality of parameters from a plurality of input sources associated with the control system, wherein the input data is captured in a plurality of data formats with different sampling intervals; synchronize the input data related to each of the plurality of parameters with reference to a timestamp of occurrence of a failure event; generate one or more data patterns by comparing the input data with historical data of each of the plurality of parameters; and recognize the one or more data patterns indicating the failure events in the hardware drives based on mapping each of the one or more data patterns with a ground truth data.
9. The failure analysis system of claim 8, wherein the plurality of parameters comprises at least one of black box data, data loggers’ data, settings data, signals data and events data.
10. The failure analysis system of claim 8, wherein the processor is further configured to: identify one or more drive trip instances based on event data and related warnings comprised in the input data of the plurality of parameters; segregate each of the one or more drive trip instances into predefined fault groups; anddetect the occurrence of the failure event by identifying one of the predefined fault groups with highest frequency distribution.
11. The failure analysis system of claim 8, wherein the processor is further configured to: generate a labelled dataset by assigning a unique fault label to the one or more data patterns associated with a particular failure event based on the pattern recognition.
12. The failure analysis system of claim 11, wherein the processor is further configured to: train a machine learning model with the labelled dataset for performing at least a predictive maintenance of the hardware drives.
13. The failure analysis system of claim 8, wherein the historical data of each of the plurality of parameters comprises data collected before occurrence of the failure event.
14. The failure analysis system of claim 8, wherein the ground truth data indicates an actual root cause of the failure event determined by a human expert associated with the failure analysis system.
15. A non -transitory computer readable medium including instructions stored thereon that when processed by at least one processor, cause a failure analysis system to perform operations comprising: receiving input data related to a plurality of parameters from a plurality of input sources upon detecting occurrence of a failure event, wherein the input data is captured in a plurality of data formats with different sampling intervals; synchronizing the input data of each of the plurality of parameters with reference to a timestamp of occurrence of the failure event; generating one or more data patterns by comparing the input data of each of the plurality of parameters with historical data of each of the plurality of parameters; and recognizing the one or more data patterns indicating the failure events in the hardware drives based on mapping each of the one or more data patterns with a ground truth data.
Citation Information
Patent Citations
Predicting failures in electrical submersible pumps using pattern recognition
US20190317488A1