Novelty Detection Method and Device for Data
By calculating the separation difficulty value and local density value of the data to be detected, using the R-tree or K-D tree algorithm and the exception index value formula, the existing novel detection methods are solved, and novel detection that simplifies operations and improves accuracy is achieved.
Patent Information
- Application Number
- CN202210240366.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-30
- Filing Date
- 2022-03-10
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-03-10
AI Technical Summary
The existing novel detection methods have complex calculation processes and poor accuracy, making it difficult to effectively identify abnormal data from industrial equipment.
By calculating the separation difficulty value and local density value of the data to be detected, using the R-tree or K-D tree algorithm, combined with the calculation formula of the abnormal index value, f(se, ld) = a - se * ld / c, simplifying the calculation process and improving the accuracy of novel detection.
The calculation process of novelty detection is simplified, the accuracy of novelty detection is improved, and abnormal data of industrial equipment can be more accurately identified, ensuring the stability of the production line.
Smart Images

Figure CN115146795B_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the field of industrial digitization, and particularly relates to a novel detection method and device for data. Background Art
[0002] In the industrial field, predictive maintenance of industrial equipment can predict and handle problems before the industrial equipment fails, avoiding the failure of industrial equipment from affecting the entire production line, thereby improving the stability of the entire production line.
[0003] Generally, a model is trained based on the historical operation data of industrial equipment, and then the trained model is used to evaluate or determine the current state of industrial equipment and predict the probability of industrial equipment failure. In theory, R & D personnel will collect positive samples (normal state) and negative samples (abnormal / fault state) to train the model. However, in most scenarios, only positive samples can be collected. The key method of predictive maintenance is to perform novelty detection on the operation data generated by industrial equipment. If the data generated by industrial equipment is novel data, it is presumed that the operation state of industrial equipment is abnormal, otherwise it is presumed that the operation state of industrial equipment is normal.
[0004] Currently, there are mainly three types of methods for novelty detection of data. The first type is the distribution-based method, mainly including the Elliptic Envelope method and the Gaussian Mixture Models method. The second type is the density-based method, mainly including the Local Outlier Factor method. The third type is the decision function-based method, mainly including the One-class SVM method and the Isolation Forest method. The current novelty detection methods have complex calculation processes and poor accuracy. Summary of the Invention
[0005] In order to solve the above technical problems, the present invention provides a novel detection method and device for data, so as to simplify the calculation process of the novelty detection method and improve the accuracy of novelty detection.
[0006] To achieve the above object, the present invention proposes a novel detection method for data, the method comprising: receiving data to be detected; calculating a separation difficulty value and a local density value of the data to be detected according to historical data and the data to be detected; calculating an anomaly index value of the data to be detected according to the separation difficulty value and the local density value; determining whether the data to be detected is novel data according to the anomaly index value. Therefore, by calculating the anomaly index value of the data to be detected through the separation difficulty value and the local density value, and determining whether the data to be detected is novel data according to the anomaly index value, the operation process is simplified and the accuracy of novelty detection is improved.
[0007] In an embodiment of the present invention, calculating the separation difficulty value of the data to be detected according to historical data and the data to be detected includes: using the R-tree algorithm to calculate the separation difficulty value of the data to be detected according to historical data and the data to be detected. For this purpose, a calculation method for the separation difficulty value is provided, which simplifies the calculation process of the separation difficulty value and improves the calculation speed of the separation difficulty value.
[0008] In an embodiment of the present invention, calculating the separation difficulty value of the data to be detected according to historical data and the data to be detected includes: using the K-D tree algorithm to calculate the separation difficulty value of the data to be detected according to historical data and the data to be detected. For this purpose, a calculation method for the separation difficulty value is provided, which improves the calculation accuracy of the separation difficulty value.
[0009] In an embodiment of the present invention, the following formula is used to calculate the local density value of the data to be detected according to historical data and the data to be detected:
[0010] ld(o) = 1 / d(o, M(X))
[0011] where ld(o) represents the local density value of the data to be detected, M(X) represents the centroid of all data in the local area X, and d(o, M(X)) represents the distance between the data to be detected and the centroid. For this purpose, a calculation method for the local density value is provided, which simplifies the calculation process of the local density value and improves the calculation speed of the local density value.
[0012] In an embodiment of the present invention, the separation difficulty value and the local density value are respectively negatively correlated with the anomaly index value.
[0013] In an embodiment of the present invention, the following formula is used to calculate the anomaly index value of the data to be detected according to the separation difficulty value and the local density value:
[0014] f(se, ld) = a -se*ld / c
[0015] where f(se, ld) represents the anomaly index value of the data to be detected, se represents the separation difficulty value, ld represents the local density value, and a and c are two constants, a > 1, c > 0. For this purpose, a calculation method for the anomaly index value is provided, which simplifies the calculation process of the anomaly index value and improves the calculation speed of the anomaly index value.
[0016] In an embodiment of the present invention, determining whether the data to be detected is novel data according to the anomaly index value includes: determining an anomaly threshold, and determining that the data to be detected is novel data when the anomaly index value of the data to be detected is greater than the anomaly threshold. For this purpose, it is possible to automatically determine whether the data to be detected is abnormal through the anomaly threshold.
[0017] The present invention also provides a novelty detection device for data, which comprises: a receiving module for receiving the data to be detected; a first calculation module for calculating a separation difficulty value and a local density value of the data to be detected according to historical data and the data to be detected; a second calculation module for calculating an anomaly index value of the data to be detected according to the separation difficulty value and the local density value; and a determination module for determining whether the data to be detected is novel data according to the anomaly index value.
[0018] In an embodiment of the present invention, when the first calculation module calculates the separation difficulty value of the data to be detected according to historical data and the data to be detected, it includes: calculating the separation difficulty value of the data to be detected according to historical data and the data to be detected by using the R-tree algorithm.
[0019] In an embodiment of the present invention, when the first calculation module calculates the separation difficulty value of the data to be detected according to historical data and the data to be detected, it includes: calculating the separation difficulty value of the data to be detected according to historical data and the data to be detected by using the K-D tree algorithm.
[0020] In an embodiment of the present invention, the first calculation module calculates the local density value of the data to be detected according to historical data and the data to be detected by using the following formula:
[0021] ld(o) = 1 / d(o, M(X))
[0022] where ld(o) represents the local density value of the data to be detected, M(X) represents the centroid of all data within the local region X, and d(o, M(X)) represents the distance between the data to be detected and the centroid.
[0023] In an embodiment of the present invention, the separation difficulty value and the local density value are respectively negatively correlated with the anomaly index value.
[0024] In an embodiment of the present invention, the second calculation module calculates the anomaly index value of the data to be detected according to the separation difficulty value and the local density value by using the following formula:
[0025] f(se, ld) = a -se*ld / c
[0026] where f(se, ld) represents the anomaly index value of the data to be detected, se represents the separation difficulty value, ld represents the local density value, and a and c are two constants, where a > 1 and c > 0.
[0027] In an embodiment of the present invention, the determining module determines whether the data to be detected is novel data according to the anomaly index value, including: determining an anomaly threshold, and determining that the data to be detected is novel data when the anomaly index value of the data to be detected is greater than the anomaly threshold.
[0028] The present invention also provides an electronic device, including a processor, a memory, and instructions stored in the memory, where when the instructions are executed by the processor, the method as described is implemented.
[0029] The present invention also provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions execute the method when running. Description of the Drawings
[0030] The following drawings are only intended to illustrate and explain the present invention schematically and do not limit the scope of the present invention. Among them,
[0031] Figure 1 is a flowchart of a method for detecting novelty of data according to an embodiment of the present invention;
[0032] Figure 2 is a schematic diagram of operation data of an industrial device according to an embodiment of the present invention;
[0033] Figure 3A and 3B is a schematic diagram of calculating the separation difficulty value using the R-tree algorithm according to an embodiment of the present invention;
[0034] Figure 4 is a schematic diagram of calculating the separation difficulty value using the K-D tree algorithm according to an embodiment of the present invention;
[0035] Figure 5 is a schematic diagram of an anomaly index value calculation function according to an embodiment of the present invention;
[0036] Figure 6 is a schematic diagram of result verification of a method for detecting novelty of data according to an embodiment of the present invention;
[0037] Figure 7 is a schematic diagram of a device for detecting novelty of data according to an embodiment of the present invention;
[0038] Figure 8 is a schematic diagram of an electronic device according to an embodiment of the present invention.
[0039] Description of the Reference Numerals
[0040] 100 Method for Detecting Novelty of Data
[0041] 110 - 140 Steps
[0042] Novelty detection device for 700 data
[0043] 710 Receiving module
[0044] 720 First calculation module
[0045] 730 Second calculation module
[0046] 740 Determination module
[0047] 800 Electronic device
[0048] 810 Processor
[0049] 820 Memory Detailed implementation manners
[0050] In order to have a clearer understanding of the technical features, objectives, and effects of the present invention, the specific implementation manners of the present invention will now be described with reference to the accompanying drawings.
[0051] In the following description, many specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0052] As shown in the present application and the claims, unless the context clearly indicates otherwise, the words "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "including" and "comprising" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0053] As introduced in the background, the key method of predictive maintenance is to perform novelty detection on the operation data generated by industrial equipment. On the premise that the operating conditions have not changed, if the data generated by the industrial equipment is novelty data, it is presumed that the operating state of the industrial equipment is abnormal; otherwise, it is presumed that the operating state of the industrial equipment is normal. Figure 2 It is a schematic diagram of the operation data of an industrial equipment according to an embodiment of the present invention, as Figure 2 shown, the operation data of the industrial equipment is two-dimensional data, and the abscissa and ordinate represent two different dimensions. For example, the abscissa represents temperature and the ordinate represents humidity. Figure 2The circular dots represent the historical operation data of industrial equipment. Usually, these dots are positive samples, that is, normal data. The non-circular dots represent the newly generated operation data of industrial equipment. By judging the relationship between the newly generated data and the historical operation data, it can be determined whether these newly generated data are abnormal, that is, novelty detection of data. By performing novelty detection on the newly generated data of industrial equipment, predictive maintenance of industrial equipment can be achieved, and the stability of the production line can be improved.
[0054] Figure 2 Four data to be detected are shown. Judging from experience, it can be determined that the data to be detected 1 is not novel data, the data to be detected 4 is novel data, and the data to be detected 2 and 3 are between novel data and non-novel data. An embodiment of the present invention provides a method for novelty detection of data to simplify the calculation process of novelty data detection and improve the accuracy of novelty data detection.
[0055] Figure 1 is a flowchart of a method 100 for novelty detection of data according to an embodiment of the present invention. As Figure 1 shown, the method for novelty detection of data in the embodiment of the present invention includes:
[0056] Step 110, receiving the data to be detected.
[0057] During the operation of industrial equipment, a large amount of operation data will be continuously generated. These operation data are used as the data to be detected to determine whether they are novel data. The data to be detected can be one-dimensional data, two-dimensional data, three-dimensional data or data of more dimensions. The dimension of the data can be defined by the user himself or set automatically by the system. Taking two-dimensional data as an example, the two dimensions of the two-dimensional data can be temperature and humidity. The two-dimensional data generated by industrial equipment can be located in a two-dimensional coordinate system composed of temperature and humidity. Optionally, the data to be detected also has a timestamp to determine the time when the data to be detected is generated, so as to facilitate subsequent data processing and analysis.
[0058] Step 120, calculating the separation difficulty value and the local density value of the data to be detected according to the historical data and the data to be detected.
[0059] Whether the data to be detected is novel data is determined by the relationship between the data to be detected and the historical data. In the embodiment of the present invention, the relationship between the data to be detected and the historical data is represented by the separation difficulty value and the local density value. Among them, the separation difficulty value represents the difficulty of separating the data to be detected from the historical data, and the separation difficulty value can be calculated by using the R-tree algorithm, the R* -tree algorithm, and the B-tree algorithm. The local density value represents the density value of the local area where the data to be detected is located, and the local density value can be calculated by using the local outlier factor (LOF) algorithm.
[0060] In some embodiments, calculating the separation difficulty value of the data to be detected based on historical data and the data to be detected includes: using the R-tree algorithm to calculate the separation difficulty value of the data to be detected based on historical data and the data to be detected. Figure 3A and 3B FIG. is a schematic diagram of calculating the separation difficulty value using the R-tree algorithm according to an embodiment of the present invention. Among them, Figure 3A is a schematic diagram of the position of the data, Figure 3B is for Figure 3A corresponding data layer structure diagram. As Figure 3A and 3B shown, data block A is the first level, data blocks B, C, and D are the second level, data blocks E, F, and G are the third level, the crosses and triangles represent the data to be detected, the cross-shaped data to be detected and the triangular data to be detected belong to data block D and data block G respectively. When separating and positioning from data block A at the first level from top to bottom, sinking one level to the second level can locate the data block D where the cross-shaped data to be detected is located, and sinking two levels to the third level can locate the data block G where the triangular data to be detected is located. Thus, the separation difficulty value of the cross-shaped data to be detected calculated using the R-tree algorithm is 1, and the separation difficulty value of the triangular data to be detected is 2.
[0061] In some embodiments, calculating the separation difficulty value of the data to be detected based on historical data and the data to be detected includes: using the K-D tree algorithm to calculate the separation difficulty value of the data to be detected based on historical data and the data to be detected. Figure 4 FIG. is a schematic diagram of calculating the separation difficulty value using the K-D tree algorithm according to an embodiment of the present invention. Figure 4 In, the dots represent historical data, and the crosses and triangles represent the data to be detected. The cross-shaped data to be detected can be separated from other data along the dotted line No. 1 and the dotted line No. 2. Therefore, the separation difficulty value of the cross-shaped data to be detected calculated using the K-D tree algorithm is 2. The triangular data to be detected can be separated from other data along the dotted line No. 1, the dotted line No. 2 until the dotted line No. 6. Therefore, the separation difficulty value of the triangular data to be detected calculated using the K-D tree algorithm is 6.
[0062] It can be understood that for a set of data to be detected, the same algorithm is adopted to calculate the separation difficulty value, for example, one is selected from the R-tree algorithm and the K-D tree algorithm to unify the calculation standard.
[0063] In some embodiments, the following formula is used to calculate the local density value of the data to be detected based on historical data and the data to be detected:
[0064] ld(o) = 1 / d(o, M(X))
[0065] Among them, ld(o) represents the local density value of the data to be detected, M(X) represents the centroid of all data within the local region X, and d(o, M(X)) represents the distance between the data to be detected and the centroid.
[0066] As Figure 3A shown, taking the triangular data to be detected as an example, the triangular data to be detected is located in the data block G. The local region can be determined as the data block G. First, calculate the centroid of the data block G, then calculate the distance between the triangular data to be detected and the centroid of the data block G, and finally take the reciprocal of the distance to obtain the local density value of the triangular data to be detected. The larger the local density value, the denser the data in the region where the data to be detected is located, and vice versa.
[0067] Step 130, calculate the anomaly index value of the data to be detected according to the separation difficulty value and the local density value.
[0068] The foregoing steps calculate the separation difficulty value and the local density value. This step calculates the anomaly index value of the data to be detected according to the separation difficulty value and the local density value. It can be understood that the separation difficulty value and the local density value are negatively correlated with the anomaly index value, that is, the larger the separation difficulty value and the local density value, the lower the anomaly index value. That is to say, the more difficult it is for the data to be detected to be separated from the historical data, or the greater the density of the local region where it is located, the lower the possibility that the data to be detected is a novel data. On the contrary, the possibility that the data to be detected is a novel data is higher. The embodiments of the present invention calculate the anomaly index value of the data to be detected according to the separation difficulty value and the local density value, realizing the quantification of the anomaly index value.
[0069] In some embodiments, the separation difficulty value and the local density value are respectively negatively correlated with the anomaly index value. That is, while the separation difficulty value is negatively correlated with the anomaly index value, the local density value is also negatively correlated with the anomaly index value.
[0070] In some embodiments, the following formula is used to calculate the anomaly index value of the data to be detected according to the separation difficulty value and the local density value:
[0071] f(se, ld) = a -se*ld / c
[0072] Among them, f(se, ld) represents the anomaly index value of the data to be detected, se represents the separation difficulty value, ld represents the local density value, and a and c are two constants, a > 1, c > 0.
[0073] Figure 5 is a schematic diagram of an anomaly index value calculation function according to an embodiment of the present invention. In Figure 5 it, the abscissa represents the product of the separation difficulty value and the local density value, and the ordinate represents the anomaly index value. From Figure 5It can be seen that the larger the product of the separation difficulty value and the local density value, the smaller the anomaly index value.
[0074] Step 140: Determine whether the data to be detected is novel data according to the anomaly index value.
[0075] After calculating the anomaly index value, this step determines whether the data to be detected is novel data according to the anomaly index value. Whether the data to be detected is novel data can be used as a basis for predictive maintenance. Intervention before the industrial equipment fails can avoid the occurrence of failures and improve the stability of the industrial equipment and production line.
[0076] In some embodiments, determining whether the data to be detected is novel data according to the anomaly index value includes: determining an anomaly threshold, and determining that the data to be detected is novel data when the anomaly index value of the data to be detected is greater than the anomaly threshold. The anomaly threshold can be determined by user input. For example, if the user inputs the anomaly threshold as 0.8 through the human-machine interface, then the data to be detected is novel data when its anomaly index value is greater than 0.8, otherwise it is not novel data. Another example is that the user can also adjust the anomaly threshold to 0.9 through the human-machine interface, then the data to be detected is novel data when its anomaly index value is greater than 0.9, otherwise it is not novel data. The anomaly threshold can also be automatically set by the system. For example, the anomaly threshold is default set to 0.8 when designing the system.
[0077] An embodiment of the present invention provides a method for detecting the novelty of data. The anomaly index value of the data to be detected is calculated through the separation difficulty value and the local density value, and whether the data to be detected is novel data is determined according to the anomaly index value, which simplifies the operation process and improves the accuracy of novelty detection.
[0078] Next, the method for detecting the novelty of data provided in the embodiments of the present invention is compared with three existing methods. Figure 6 It is a schematic diagram of the result verification of the method for detecting the novelty of data according to an embodiment of the present invention. Among them, the circular dots represent historical data, and the cross-shaped dots represent the data to be detected. There are a total of 7 pieces of data to be detected. According to the positions of the 7 pieces of data to be detected,
[0079] The anomaly index value of the data to be detected 1 < the anomaly index value of the data to be detected 4 < the anomaly index value of the data to be detected 2;
[0080] The anomaly index value of the data to be detected 1 < the anomaly index value of the data to be detected 4 < the anomaly index value of the data to be detected 3;
[0081] The anomaly index value of the data to be detected 1 < the anomaly index value of the data to be detected 4 < the anomaly index value of the data to be detected 5 < the anomaly index value of the data to be detected 6 < the anomaly index value of the data to be detected 7
[0082] Next, the existing three methods are used to calculate the outlier index values of 7 data to be detected respectively, and the results are shown in the following table. Among them, the embodiment of the present invention uses the k-d tree algorithm and the local density value to calculate the outlier index value.
[0083]
[0084]
[0085] Table 1 Verification Results of Different Novelty Detection Methods
[0086] For the Elliptic Envelope method, the outlier index value of the data 1 to be detected is greater than that of the data 4 to be detected, and the outlier index value of the data 4 to be detected is greater than that of the data 3 to be detected, and misjudgment obviously occurs.
[0087] For the Isolation Forest method, the outlier index values of the data 2 and 3 to be detected are less than that of the data 4 to be detected, and the data 6 and 7 to be detected have the same outlier index value, and misjudgment also occurs.
[0088] For the Local Outlier Factor method, the outlier index values of the data 1 and 4 to be detected are close, and the outlier degrees of the data 1 and 4 to be detected cannot be distinguished.
[0089] In contrast, the method of the embodiment of the present invention obviously conforms to the above rules and has higher accuracy and precision.
[0090] Figure 7 It is a schematic diagram of a novelty detection device 700 for data according to an embodiment of the present invention. As Figure 7 shown, the novelty detection device 700 includes:
[0091] A receiving module 710 for receiving the data to be detected.
[0092] A first calculation module 720 for calculating the separation difficulty value and the local density value of the data to be detected according to the historical data and the data to be detected.
[0093] A second calculation module 730 for calculating the outlier index value of the data to be detected according to the separation difficulty value and the local density value.
[0094] A determination module 740 for determining whether the data to be detected is novel data according to the outlier index value.
[0095] In some embodiments, the first calculation module 720 calculates the separation difficulty value of the data to be detected according to the historical data and the data to be detected, including: using the R-tree algorithm to calculate the separation difficulty value of the data to be detected according to the historical data and the data to be detected.
[0096] In some embodiments, the first calculation module 720 calculates the separation difficulty value of the data to be detected based on historical data and the data to be detected, including: using the K-D tree algorithm to calculate the separation difficulty value of the data to be detected based on historical data and the data to be detected.
[0097] In some embodiments, the first calculation module 720 calculates the local density value of the data to be detected according to the following formula based on historical data and the data to be detected:
[0098] ld(o) = 1 / d(o, M(X))
[0099] Where ld(o) represents the local density value of the data to be detected, M(X) represents the centroid of all data within the local area X, and d(o, M(X)) represents the distance between the data to be detected and the centroid.
[0100] In some embodiments, the separation difficulty value and the local density value are respectively negatively correlated with the anomaly index value.
[0101] In some embodiments, the second calculation module 730 calculates the anomaly index value of the data to be detected according to the following formula based on the separation difficulty value and the local density value:
[0102] f(se, ld) = a -se*ld / c
[0103] Where f(se, ld) represents the anomaly index value of the data to be detected, se represents the separation difficulty value, ld represents the local density value, and a and c are two constants, a > 1, c > 0.
[0104] In some embodiments, the determination module 740 determines whether the data to be detected is novel data according to the anomaly index value, including: determining an anomaly threshold, and determining that the data to be detected is novel data when the anomaly index value of the data to be detected is greater than the anomaly threshold.
[0105] The present invention also provides an electronic device 800. Figure 8 is a schematic diagram of an electronic device 800 according to an embodiment of the present invention. As Figure 8 shown, the electronic device 800 includes a processor 810 and a memory 820. Instructions are stored in the memory 820, and when the instructions are executed by the processor 810, the method 100 described above is implemented.
[0106] The present invention also provides a computer-readable storage medium, on which computer instructions are stored, and when the computer instructions are run, the method 100 described above is executed.
[0107] Some aspects of the methods and apparatuses of the present invention may be executed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The above-mentioned hardware or software may all be referred to as "data blocks", "modules", "engines", "units", "components" or "systems". The processor may be one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DAPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors or combinations thereof. In addition, aspects of the present invention may be embodied as a computer product located on one or more computer-readable media, which product includes computer-readable program code. For example, the computer-readable media may include, but are not limited to, magnetic storage devices (such as hard disks, floppy disks, magnetic tapes...), optical disks (such as compact discs (CDs), digital versatile discs (DVDs)...), smart cards, and flash memory devices (such as cards, sticks, key drives...).
[0108] Flowcharts are used herein to illustrate the operations performed by the methods according to embodiments of the present application. It should be understood that the foregoing operations are not necessarily performed precisely in order. Instead, various steps may be processed in reverse order or simultaneously. Also, one or more other operations may be added to these processes, or one or more steps may be removed from these processes.
[0109] It should be understood that although this specification is described according to various embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment may also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
[0110] The above are only illustrative specific embodiments of the present invention and are not intended to limit the scope of the present invention. Any equivalent changes, modifications and combinations made by those skilled in the art without departing from the concept and principles of the present invention shall fall within the scope of protection of the present invention.
Claims
1. A novel detection method (100) for data, characterized in that, The method (100) includes: Receiving data to be detected (110); wherein, the operation data generated during the operation of the industrial equipment is the data to be detected; Calculating the separation difficulty value and the local density value of the data to be detected according to the historical data and the data to be detected (120); wherein, the separation difficulty value represents the difficulty of separating the data to be detected from the historical data; Calculating the anomaly index value of the data to be detected according to the separation difficulty value and the local density value (130); wherein, the separation difficulty value and the local density value are respectively negatively correlated with the anomaly index value; Wherein, the following formula is used to calculate the anomaly index value of the data to be detected according to the separation difficulty value and the local density value: f(se, ld) = a -se*ld / c Wherein, f(se, ld) represents the anomaly index value of the data to be detected, se represents the separation difficulty value, ld represents the local density value, a and c are two constants, a>1, c>0; Determining whether the data to be detected is novel data according to the anomaly index value (140); wherein, an anomaly threshold is determined by input through a human-machine interface or automatically set by the system, and when the anomaly index value of the data to be detected is greater than the anomaly threshold, it is determined that the data to be detected is novel data; on the premise that the operating condition has not changed, it is presumed that the operating state of the industrial equipment is abnormal.
2. The method (100) according to claim 1, characterized in that, Calculating the separation difficulty value of the data to be detected according to the historical data and the data to be detected includes: using the R-tree algorithm to calculate the separation difficulty value of the data to be detected according to the historical data and the data to be detected.
3. The method (100) according to claim 1, characterized in that, Calculating the separation difficulty value of the data to be detected according to the historical data and the data to be detected includes: using the K-D tree algorithm to calculate the separation difficulty value of the data to be detected according to the historical data and the data to be detected.
4. The method (100) according to any one of claims 1-3, characterized in that, The following formula is used to calculate the local density value of the data to be detected according to the historical data and the data to be detected: ld(o) = 1 / d(o, M(X)) Wherein, ld(o) represents the local density value of the data to be detected, M(X) represents the centroid of all data in the local area X, and d(o, M(X)) represents the distance between the data to be detected and the centroid.
5. A novelty detection device (700) for data, characterized in that, The device (700) includes: A receiving module (710) for receiving data to be detected; wherein, the operation data generated during the operation of the industrial equipment is the data to be detected; A first calculation module (720) for calculating the separation difficulty value and the local density value of the data to be detected according to the historical data and the data to be detected; wherein, the separation difficulty value represents the difficulty of separating the data to be detected from the historical data; A second calculation module (730) for calculating the anomaly index value of the data to be detected according to the separation difficulty value and the local density value; wherein, the separation difficulty value and the local density value are respectively negatively correlated with the anomaly index value; The second calculation module (730) uses the following formula to calculate the anomaly index value of the data to be detected according to the separation difficulty value and the local density value: f(se, ld) = a -se*ld / c Among them, f(se, ld) represents the anomaly index value of the data to be detected, se represents the separation difficulty value, ld represents the local density value, a and c are two constants, a > 1, c > 0; A determination module (740) determines whether the data to be detected is novel data according to the anomaly index value; among them, an anomaly threshold is determined by input through a human-machine interface or automatically set by the system, and when the anomaly index value of the data to be detected is greater than the anomaly threshold, it is determined that the data to be detected is novel data; on the premise that the operating condition has not changed, it is presumed that the operating state of the industrial equipment is abnormal.
6. The device (700) according to claim 5, characterized in that, The first calculation module (720) calculates the separation difficulty value of the data to be detected according to historical data and the data to be detected, including: calculating the separation difficulty value of the data to be detected according to historical data and the data to be detected by using the R-tree algorithm.
7. The device (700) according to claim 5, characterized in that, The first calculation module (720) calculates the separation difficulty value of the data to be detected according to historical data and the data to be detected, including: calculating the separation difficulty value of the data to be detected according to historical data and the data to be detected by using the K-D tree algorithm.
8. The device (700) according to any one of claims 5 to 7, characterized in that, The first calculation module (720) calculates the local density value of the data to be detected according to historical data and the data to be detected by using the following formula: ld(o) = 1 / d(o, M(X)) Among them, ld(o) represents the local density value of the data to be detected, M(X) represents the centroid of all data in the local area X, and d(o, M(X)) represents the distance between the data to be detected and the centroid.
9. An electronic device (800) includes a processor (810), a memory (820), and instructions stored in the memory (820), where when the instructions are executed by the processor (810), the method described in any one of claims 1-4 is implemented.
10. A computer-readable storage medium stores computer instructions, and when the computer instructions are run, the method described in any one of claims 1-4 is executed.
Citation Information
Patent Citations
Data novelty detection method and apparatus
EP4068026A1