Abnormal data detection method based on deep learning

Through the deep learning-based anomaly data detection method, combined with multi-dimensional data and decision tree model, dynamic binning and feedback optimization, the problems of high computational complexity and incomplete detection in the existing technology are solved, and fast and accurate abnormal identification and response are achieved, which improves the efficiency and accuracy of safety monitoring in industrial plants.

CN120337064AInactive Publication Date: 2025-07-18严亚伟

Patent Information

Application Number
CN202510403926.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the abnormal detection algorithm has high computational complexity, cannot respond quickly in a real-time environment, and fails to fully combine multi-dimensional data to make abnormal judgments, resulting in insufficient comprehensiveness of detection and difficult to adapt to the situation of a small number of abnormal samples.

Method used

The abnormal data detection method based on deep learning is adopted, and multi-dimensional data is captured and preprocessed in real time, and event tuple exception detection and decision tree model are used, dynamic binning and feedback optimization are achieved to achieve rapid identification and accurate alarm of abnormal data.

Benefits of technology

This method takes into account real-time and dynamic adaptability, can quickly and accurately identify potential risks, reduce false alarms and missed reports, improve the accuracy and response speed of industrial plant safety monitoring, and provide strong technical support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337064A_ABST
    Figure CN120337064A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to an abnormal data detection method based on deep learning. The method comprises the steps; s1, capturing and preprocessing real-time data; s2, event tuple and anomaly judgment; s3, performing dynamic binning and decision tree modeling; s4, selecting an optimal detection path; s5, performing feedback optimization and continuous monitoring; according to the dynamic binning construction rule and the decision tree, the detection overhead is reduced, the real-time detection data is preprocessed, the abnormal state of the real-time data is comprehensively recognized, multi-level alarm timely response is achieved, a self-adaptive mechanism is fed back, real-time data monitoring is achieved, multi-dimensional data are collected in real time and preprocessed, and the real-time detection efficiency is improved. According to the method, the data exception risk is determined through dynamic binning and decision tree construction, the abnormal data is alarmed according to the risk level, meanwhile, the standard dependence intensity variable quantity is adjusted according to the dynamic feedback result, the detection accuracy and stability are improved, and the requirement for efficient and accurate exception detection of service data is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to an abnormal data detection method based on deep learning. Background Art

[0002] Abnormal data detection is an important data preprocessing technology, whose goal is to identify and process outliers or abnormal situations in a dataset to improve data quality and the accuracy of analysis results.

[0003] In the 1960s, with the development of database technology, abnormal data detection gradually became an important database management technology. Entering the 21st century, with the rapid development of machine learning and artificial intelligence technologies, the application scope of abnormal data detection has gradually expanded and become one of the key technologies in the fields of data analysis and data mining. The patent document with the Chinese patent publication number CN117935357A discloses a method for detecting abnormal behaviors of factory personnel based on self-supervised learning. Its technical point is to construct four graph structures and extract human skeleton features through a graph convolutional neural network. After further extracting action frame features, contrast learning is used for pre-training. The GCN and 3D-CNN features are spliced and input into a temporal convolutional network for action modeling to achieve abnormal behavior detection, and the detection accuracy is continuously optimized through a feedback mechanism to improve the real-time performance and efficiency of vulnerability detection; at the same time, the data source of this system is single: only relying on the skeleton action data obtained by the camera, without combining multi-dimensional data such as temperature, gas, and vibration, resulting in insufficient comprehensiveness of abnormal detection and difficulty in adapting to the situation of a small number of abnormal samples; the abnormal judgment dimension is less, rather than comprehensively monitoring more extensive risk factors such as the state of equipment and environmental safety, and there is no mention of a mechanism for optimizing the model based on long-term data feedback, and the feedback mechanism is weak. Summary of the Invention

[0004] Therefore, the present invention provides an abnormal data detection method based on deep learning to overcome the problem that the calculation complexity of the existing abnormal detection algorithm is relatively high and it cannot respond quickly in a real-time environment.

[0005] To achieve the above object, the present invention provides an abnormal data detection method based on deep learning, including,

[0006] Real-time capturing the input data in the interaction process;

[0007] Preprocessing the input data into unit input data, and performing event tuple abnormal detection on the unit input data to obtain a first detection result and a second detection result;

[0008] When the first detection result is obtained, execute a first data detection program to obtain abnormal data;

[0009] Determine whether to execute the second data detection program based on the real-time abnormal intensity change amount and the actual dependence degree when obtaining the second detection result;

[0010] Among them, when the real-time abnormal intensity change amount is within the range of the mild abnormal intensity change amount and the actual dependence degree is less than the standard dependence degree, execute the second data detection program and determine the abnormal data based on the decision tree model.

[0011] Furthermore, the event tuple anomaly detection for the unit input data includes,

[0012] Obtain the activity ID of the unit input data and compare it with the predefined activity list stored in the database.

[0013] If the activity ID is in the predefined activity list, obtain the first detection result;

[0014] If the activity ID is not in the predefined activity list, obtain the second detection result.

[0015] Furthermore, obtaining the first detection result includes,

[0016] Directly determine the abnormal situation of the unit input data according to the logical expression, obtain the abnormal data, and alarm the abnormal data.

[0017] Furthermore, obtaining the second detection effect includes judging the real-time abnormal intensity change amount and the actual dependence degree, and the judgment method is,

[0018] Determine that the input data is suspected abnormal data and obtain the real-time abnormal intensity change amount, and compare it with the range of the mild abnormal dependence intensity change amount.

[0019] If the real-time abnormal intensity change amount is outside the range of the mild abnormal intensity change amount, obtain the abnormal data and alarm the abnormal data;

[0020] If the real-time dependence intensity change amount is within the range of the mild abnormal intensity change amount, obtain the actual dependence degree and compare it with the standard dependence degree.

[0021] If the actual dependence degree is greater than the standard dependence degree, execute the second data detection program to determine the abnormal data;

[0022] If the actual dependence degree is less than or equal to the standard dependence degree, execute the third data detection program to determine the abnormal data.

[0023] Furthermore, executing the second data detection program includes dynamic binning, and among them, the process of the dynamic binning is,

[0024] Slice the unit input data into several input data subsets with a preset field length;

[0025] Obtain the historical hit frequencies of each of the input data subsets respectively, and sort them in descending order according to the historical hit frequencies;

[0026] Calculate the real-time correlation degree between two adjacent input data subsets;

[0027] Link each of the input data subsets based on the real-time correlation degree, obtain the real-time correlation degree, and compare it with the standard correlation degree.

[0028] If the real-time correlation degree is greater than or equal to the standard correlation degree, the input data subsets are linked to form several input data sets;

[0029] If the real-time correlation degree is less than the standard correlation degree, the input data subsets are not linked.

[0030] Furthermore, executing the second data detection program includes constructing a decision tree model at the same time, where

[0031] The process of constructing the decision tree model is to use each of the input data sets as training data to construct a decision tree model.

[0032] Furthermore, executing the second data detection program also includes detecting path optimization, where the process of the detecting path optimization is

[0033] Starting from the root node, traverse the splitting points of the splittable fields and calculate the splitting gain of each splitting point;

[0034] Determine the optimal splitting and the optimal detection path of each node based on the splitting gain;

[0035] Among them, the nodes include leaf nodes and non-leaf nodes.

[0036] Furthermore, executing the third data detection program includes

[0037] Screen all suspected abnormal data under the actual high-dependency degree, select the suspected abnormal data with the highest radiation influence proportion on other suspected abnormal data per unit time as the missing parameter, set the missing parameter as the continuous monitoring target, and trigger a potential risk alarm;

[0038] Among them, the calculation formula of the radiation influence proportion is

[0039]

[0040] Among them, R i is the influence proportion of data D i on the overall abnormal data set;

[0041] W i,j is the influence weight of data D i on data D j ;

[0042] A j is the degree of abnormality of data D j ;

[0043] ∑ k A k is the total abnormal value of all abnormal data points;

[0044] The formula calculates the influence ratio of a certain abnormal data point D i on the set of all abnormal data points to determine the key factor, i.e., the "missing parameter", whether this data can cause other data to be abnormal.

[0045] Furthermore, the calculation formula of the splitting gain is

[0046] ΔC = C node + P L * C(D L ) + P R * C(D R ) - C current

[0047] where

[0048] C node is the node judgment overhead;

[0049] C(D L ) is the left node detection overhead, and C(D R ) is the right node detection overhead;

[0050] P L * is the probability of entering the left child node, and P R * is the probability of entering the right child node;

[0051] C current is the current subtree overhead of this node.

[0052] Furthermore, executing the second data detection program to determine abnormal data includes

[0053] judging whether the non-abnormal data sample is known data and incorporating the non-abnormal data into the data dependency information library; where the data dependency information library is a set of historical input data, and the binning rule parameters and decision tree model fitting parameters are feedback-adjusted.

[0054] Compared with the prior art, the beneficial effects of the present invention are as follows. This method takes into account both real-time performance and has the capabilities of dynamic adaptability and feedback optimization. It can quickly and accurately identify potential risks, reduce false alarms and missed alarms, integrate multi-dimensional data, and achieve full-chain monitoring from data collection, preprocessing, binning, decision tree optimization to alarm response, greatly improving the accuracy and response speed of industrial plant safety monitoring, and providing strong technical support for preventing accidents and ensuring production safety.

[0055] Furthermore, the comparison mechanism for determining whether the activity ID exists in the list helps to distinguish normal and abnormal behaviors, achieve preliminary screening, and provide basic data for subsequent dynamic binning and decision tree detection, thereby improving the detection efficiency and accuracy of the entire system.

[0056] Furthermore, by determining the data priority through the real-time change amount of the dependence intensity, it is ensured that only data at the same level is binned, and then the abnormality can be judged more accurately.

[0057] Furthermore, in this embodiment, the joint correlation coefficient of three variables is used to calculate the association level between variables, and based on the strategy of combining and training high-correlation data and training low-correlation data separately, the anomaly detection system is optimized, improving the detection accuracy while reducing the calculation cost and enhancing the economic benefits in the practical application of the method.

[0058] Furthermore, compared with traditional single-point anomaly detection, this method can find out the core data points that have the greatest impact on the overall anomaly, accurately identify the key anomaly points, thereby optimizing the detection strategy, and at the same time avoiding all abnormal data being treated equally, focusing on the key data, and improving the efficiency of the detection system.

[0059] Furthermore, the split gain calculation can ensure that each split can minimize the detection overhead to the greatest extent, thereby reducing the occupancy of computing resources. At the same time, by dynamically selecting the split point with the largest gain, the decision tree can form the shortest and optimal detection path, improve the data classification accuracy, and avoid unnecessary calculations and complexities, maintaining the generalization ability of the model.

[0060] Furthermore, the historical input data set can be directly compared in the future, avoiding repeated calculations for known normal data, improving the system response speed. The accumulation of data in the data dependency library can more accurately adjust the binning boundaries and decision tree structures, making the detection method more in line with the actual data distribution and still maintaining a high accuracy rate when facing new types of abnormal data. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 is a flowchart of the anomaly data detection method based on deep learning according to an embodiment of the present invention;

[0062] Figure 2It is the logic decision diagram of the first data detection program in the embodiment of the present invention;

[0063] Figure 3 It is the logic decision diagram for judging the subsequent execution mode according to the actual dependence degree in the embodiment of the present invention;

[0064] Figure 4 It is the logic decision diagram of feedback adjustment in the embodiment of the present invention. Detailed implementation manners

[0065] In order to make the objectives and advantages of the present invention clearer, the present invention will be further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0066] The preferred implementation manners of the present invention will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these implementation manners are only used to explain the technical principle of the present invention and do not limit the protection scope of the present invention.

[0067] It should be noted that in the description of the present invention, the terms indicating directions or positional relationships such as "upper", "lower", "left", "right", "inner", "outer", etc. are based on the directions or positional relationships shown in the drawings. This is only for convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0068] In addition, it should also be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0069] Please refer to Figure 1 As shown, it is the flowchart of the abnormal data detection method based on deep learning in the embodiment of the present invention. The present invention provides an abnormal data detection method based on deep learning, including,

[0070] Real-time capture of input data in the interaction process;

[0071] Preprocess the input data into unit input data, and perform event tuple abnormal detection on the unit input data to obtain a first detection result and a second detection result;

[0072] When the first detection result is obtained, execute the first data detection program to obtain abnormal data;

[0073] Determine whether to execute the second data detection program based on the real-time abnormal intensity change amount and the actual dependence degree when obtaining the second detection result;

[0074] Among them, when the real-time abnormal intensity change amount is within the mild abnormal intensity change amount range and the actual dependence degree is less than the standard dependence degree, execute the second data detection program, and determine the abnormal data based on the decision tree model.

[0075] This system preprocesses by collecting factory safety data in real time, including multi-dimensional information such as temperature, vibration, gas leakage, noise, video, energy consumption, humidity, pressure, infrared thermal imaging, and smoke, and divides the original data into several units for input; after data preprocessing, the system uses the event tuple anomaly detection method to divide the input data into two detection paths: known and unknown: the known data directly matches the predefined rules in the database for judgment, while the unknown data further enters the dynamic binning process. In the binning stage, adaptive binning strategies such as equal-width and equal-frequency are used according to the preset fields to divide the data into subsets that cover all historical data, and a qualified rule subset for each subset is generated according to the historical rule hit frequency. The system uses the decision tree model to perform anomaly detection on the binned data, selects the optimal splitting scheme by calculating the splitting gain of each candidate splitting point, thereby optimizing the detection path and reducing the calculation overhead. At the same time, the system introduces the concept of missing parameters, reclassifies the abnormal data with a large radiation impact on other data, and once its impact ratio exceeds the predetermined threshold, starts the intermediate alarm and continuously monitors. This method not only takes into account real-time performance but also has dynamic adaptability and feedback optimization capabilities, can quickly and accurately identify potential risks, reduce false alarms and missed alarms, integrate multi-dimensional data and achieve full-chain monitoring from data collection, preprocessing, binning, decision tree optimization to alarm response, greatly improving the accuracy and response speed of industrial plant safety monitoring, and providing strong technical support for preventing accidents and ensuring production safety.

[0076] Refer to Figure 2 and Figure 3 as shown, Figure 2 is the logic determination diagram of the first data detection program of the embodiment of the present invention, Figure 3 is the logic determination diagram of the embodiment of the present invention for judging the subsequent execution method according to the actual dependence degree;

[0077] Specifically, performing event tuple anomaly detection on unit input data includes,

[0078] Obtain the activity ID of the unit input data and compare it with the predefined activity list stored in the database,

[0079] If the activity ID is in the predefined activity list, obtain the first detection result;

[0080] If the activity ID is not in the predefined activity list, a second detection result is obtained;

[0081] If the activity ID exists in the list, it indicates that the data unit belongs to a known and normal activity pattern. The system directly outputs the first detection result and determines it as normal data. Conversely, if the activity ID is not in the predefined list, the data unit is considered a suspected abnormal data and needs further processing. This comparison mechanism helps to distinguish normal and abnormal behaviors, achieve preliminary screening, and provide basic data for subsequent dynamic binning and decision tree detection, thereby improving the detection efficiency and accuracy of the entire system.

[0082] Specifically, obtaining the first detection result includes,

[0083] Directly determine the abnormal situation of the input data of this unit according to the logical expression, obtain the abnormal data, and alarm the abnormal data;

[0084] In this embodiment, temperature and gas concentration are used as key indicators, and the logical expression is defined as:

[0085]

[0086] Among them, T and G respectively represent temperature and gas concentration;

[0087] T th 、G th Are the high-risk thresholds of their respective;

[0088] The real-time dependence intensity change amount is used to determine the priority of the data. The higher the priority of the data, the higher the importance of affecting the production safety in the factory;

[0089] Only when all key indicators exceed the preset threshold, the unit data is determined as high-risk data, thus directly triggering an alarm;

[0090] This determination method uses multi-index cross-validation, greatly reducing the false alarm rate. At the same time, when a single data with high priority shows obvious abnormalities, the system can respond immediately, effectively preventing the spread of accidents, and greatly improving the real-time performance and accuracy of factory safety monitoring, thereby providing a timely and reliable basis for emergency handling.

[0091] Specifically, obtaining the second detection effect includes the real-time abnormal intensity change amount and the actual dependence degree judgment. The judgment method is,

[0092] Determine that the input data is suspected abnormal data and obtain the real-time abnormal intensity change amount, and compare it with the mild abnormal dependence intensity change amount interval,

[0093] If the real - time abnormal intensity change amount is outside the mild abnormal intensity change amount range, abnormal data is obtained, and an alarm is given for the abnormal data;

[0094] If the real - time dependence intensity change amount is within the mild abnormal intensity change amount range, the actual dependence degree is obtained and compared with the standard dependence degree.

[0095] Among them, based on the three - sigma principle, the distribution of data is divided into three intervals. Data falling within the mean ± 1σ has no abnormal intensity change amount. The range between 1σ and 2σ is the mild abnormal intensity change amount range. The fluctuation range in this interval exceeds the normal fluctuation range but is a minor deviation. Data outside the mild abnormal intensity change amount range has an obvious deviation and belongs to an extreme abnormal situation;

[0096] When the actual dependence degree is greater than the standard dependence degree, the second data detection program is executed to determine the abnormal data;

[0097] When the actual dependence degree is less than or equal to the standard dependence degree, the third data detection program is executed to determine the abnormal data;

[0098] Among them, the dependence degree is used to judge the influence relationship between data, and the standard dependence degree is used to judge the relationship of vertical control or horizontal classification between data;

[0099] Vertical control: If the detected data directly affects other data, such as the proportion of the increase in the temperature of the working equipment leading to a significant decrease in production efficiency exceeding the standard dependence degree, then the two kinds of data are in a tight vertical control relationship;

[0100] Horizontal classification: If the probability of data affecting each other is lower than the standard dependence degree and the dispersion is strong, hierarchical training and management are carried out on them;

[0101] When the dependence degree of a certain data on other data is higher than 0.3, it indicates that the influence of this data on j is significant and strong. At this time, the two are classified into a tight vertical control relationship;

[0102] If the dependence degree of a certain data on other data is less than or equal to 0.3, it indicates that the direct influence between data is weak and belongs to the horizontal classification relationship. Binning, hierarchical training and management can be carried out on each data;

[0103] The data priority is determined through the real - time dependence intensity change amount to ensure that only data at the same level is binned, thereby more accurately judging abnormalities;

[0104] Among them, for any two indicators i and j, we can define the dependence degree D(i,j) between them as:

[0105]

[0106] Among them:

[0107] W(i) is the influence weight of index i;

[0108] ΔA(i→j) represents the abnormal influence degree of the change of index i on index j (which can be determined based on historical data or experiments);

[0109] The denominator is the total weighted influence of all indexes that may affect index j.

[0110] In this embodiment, the monitoring system collects the following data in real time: toxic gas concentration, equipment vibration noise, air humidity, and equipment production efficiency. Through historical data statistics and training, the influence weights of each index are obtained as follows: toxic gas concentration 0.1; vibration noise 0.01; air humidity 0.05; production efficiency 0.08;

[0111] By traversing all combinations, it is obtained that in this embodiment, the dependence degree between indexes in all combination modes is significantly lower than 0.3, belonging to the horizontal classification relationship, and is applicable to discrete classification training and management modes;

[0112] In this embodiment, this step can finely distinguish the risk levels of different abnormal data, and focus on monitoring the key abnormal data, that is, the missing parameters. By dynamically adjusting the detection program, the system can respond to abnormal data in a timely manner, reduce the risks of false alarms and missed alarms, and improve the detection accuracy.

[0113] Specifically, executing the second data detection program includes dynamic binning, where,

[0114] The process of the dynamic binning is as follows,

[0115] The unit input data is segmented into several input data subsets with a preset field length;

[0116] The historical hit frequencies of each input data subset are obtained respectively, and sorted in descending order according to the historical hit frequencies;

[0117] Calculate the real-time correlation degree between two adjacent input data subsets;

[0118] Based on the real-time correlation degree, the input data subsets are linked to obtain the real-time correlation degree, and compared with the standard correlation degree,

[0119] If the real-time correlation degree is greater than or equal to the standard correlation degree, the input data subsets are linked to form several input data sets;

[0120] If the real-time correlation degree is less than the standard correlation degree, the input data subsets are not linked.

[0121] This embodiment comprehensively considers the actual production mode and defines the joint correlation coefficient with three real-time input data:

[0122]

[0123] Among them,

[0124] X, Y, Z: Three kinds of input data;

[0125] X i , Y i , Z i : The measured value at the i-th time point;

[0126] The mean value of the variable;

[0127] W: The weight factor, set according to the historical input data set;

[0128] When the correlation degree is high, the linked data subset is used to train the model with the input data set. When the correlation degree is low, they are trained separately to avoid the decrease in detection accuracy caused by data mixing.

[0129] In this embodiment, the joint correlation coefficient of three variables is used to calculate the correlation level between variables, and based on the strategy of combining and training high-correlation data and training low-correlation data separately, the anomaly detection system is optimized, which improves the detection accuracy while reducing the calculation cost and increases the economic benefits in the practical application of the method.

[0130] Specifically, executing the second data detection program also includes constructing a decision tree model. Among them,

[0131] The process of constructing the decision tree model is to use each of the input data sets as training data to construct a decision tree model;

[0132] In this embodiment, the sample size of the leaf node, the decision path length, and the Bayesian probability are used as the decision tree judgment methods. If the number of samples of a certain leaf node is much lower than the historical normal level, the data is listed as suspected abnormal data for in-depth analysis; then, according to the number of split layers passed by the sample from the root node to the leaf node, it is judged whether the data deviates from the normal state.

[0133] Specifically, executing the second data detection program also includes optimizing the detection path. Among them,

[0134] The process of optimizing the detection path is,

[0135] Starting from the root node, traverse the split points of the splittable fields and calculate the split gain of each split point;

[0136] If the split gain is positive, select the step with the least split consumption to execute. If the split gain is negative, stop executing;

[0137] Based on the split gain, determine the optimal split and the optimal detection path for each node;

[0138] Among them, the nodes include leaf nodes and non-leaf nodes.

[0139] Specifically, executing the third data detection program includes

[0140] screening all suspected abnormal data under the actual high-dependency level, selecting the suspected abnormal data with the highest proportion of radiation impact on other suspected abnormal data within a unit time as the missing parameter, setting the missing parameter as the continuous monitoring target, and triggering a potential risk alarm;

[0141] Among them, the calculation formula for the radiation impact proportion is

[0142]

[0143] Among them, the radiation impact proportion is the influence proportion of the abnormal data on the overall abnormal set, and it is judged whether this abnormal data can directly cause other data to be abnormal;

[0144] R i is the influence proportion of data D i on the overall abnormal data set;

[0145] W i,j is the influence weight of data D i on data D j ;

[0146] A j is the abnormal degree of data D j ;

[0147] ∑ k A k is the total abnormal value of all abnormal data points;

[0148] This formula calculates the influence proportion of a certain abnormal data point D i on the set of all abnormal data points, which is the key factor for judging whether this data can cause other data to be abnormal, that is, the "missing parameter"; in actual production applications, the numerical display is adjusted dynamically according to the data. The following gives an example calculation process:

[0149] Read the abnormal values: When a certain factory is in operation, the real-time temperature of the production equipment is 30 degrees Celsius higher than the working standard temperature, the real-time concentration of toxic gas has increased by 0.1% compared with the toxic gas standard concentration, the real-time noise of equipment vibration has increased by 10% compared with the equipment vibration standard noise, the real-time air humidity has decreased by 1% compared with the air standard humidity, and the real-time production efficiency of the equipment has decreased by 15% compared with the equipment standard production efficiency:

[0150] Set the influence weight of temperature to 0.8, the influence weight of toxic gas concentration to 0.3, the influence weight of equipment vibration noise to 0.05, the influence weight of air humidity to 0.15, and the influence weight of equipment production efficiency to 0.08:

[0151] The radiation impact percentages of each data obtained by the above calculation method are as follows:

[0152] Temperature: approximately 92.7%; concentration of toxic gases: approximately 0.12%; vibration noise: approximately 1.93%; air humidity: approximately 0.58%; production efficiency: approximately 4.64%.

[0153] When the radiation impact percentage of a certain abnormal data point reaches more than 30%, this data will be classified as a "missing parameter". In this embodiment, the radiation impact percentage of temperature is as high as 92.7%. Assuming that the temperature is the continuously monitored target, the detection strategy is further optimized to improve the accuracy of risk alarm;

[0154] Compared with the traditional single-point anomaly detection, this method can find out the core data points that have the greatest impact on the overall anomaly, accurately identify the key anomaly points, thereby optimizing the detection strategy, and at the same time avoiding all abnormal data being equally processed, focusing on the key data, and improving the efficiency of the detection system.

[0155] Specifically, the calculation formula of the splitting gain is,

[0156] ΔC = C node +P L *C(D L )+P R *C(D R )-C current

[0157] Wherein,

[0158] C node is the node judgment overhead;

[0159] C(D L ) is the left node detection overhead, and C(D R ) is the right node detection overhead;

[0160] P L * is the probability of entering the left child node, and P R * is the probability of entering the right child node;

[0161] C current is the current subtree overhead of this node;

[0162] The splitting gain calculation can ensure that each split can minimize the detection overhead to the greatest extent, thereby reducing the occupation of computing resources. At the same time, by dynamically selecting the split point with the largest gain, the decision tree can form the shortest and optimal detection path, improve the accuracy of data classification, and avoid unnecessary calculations and complexities, maintaining the generalization ability of the model.

[0163] Refer to Figure 4 shown, which is the logical decision diagram of the feedback adjustment in the embodiment of the present invention;

[0164] Specifically, when executing the second data detection program to determine abnormal data, it includes

[0165] judging whether the non-abnormal data samples are known data, and incorporating the non-abnormal data into the data dependency information library; wherein, the data dependency information library is a historical input data set, which feeds back and adjusts the binning rule parameters and the decision tree model fitting parameters; when there is a false alarm, narrow the binning boundary and prune the decision tree, and when there is a missed alarm, expand the binning boundary and refine the decision tree.

[0166] The historical input data set can be directly compared in the future, avoiding repeated calculations for known normal data and improving the system response speed. The accumulation of data in the data dependency library can more accurately adjust the binning boundary and the decision tree structure, making the detection method more in line with the actual data distribution and still maintaining a high accuracy rate when facing new types of abnormal data.

[0167] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.

[0168] The above are only the preferred embodiments of the present invention and are not used to limit the present invention; for those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An abnormal data detection method based on deep learning, characterized in that, including, capturing input data in the interactive process in real time; preprocessing the input data into unit input data, and performing event tuple anomaly detection on the unit input data to obtain a first detection result and a second detection result; executing a first data detection program when the first detection result is obtained to obtain abnormal data; determining whether to execute a second data detection program based on the real-time anomaly intensity change amount and the actual dependency when the second detection result is obtained; wherein, when the real-time anomaly intensity change amount is within the mild anomaly intensity change amount range and the actual dependency is less than the standard dependency, execute the second data detection program, and determine the abnormal data based on the decision tree model.

2. The anomaly data detection method based on deep learning according to claim 1, wherein Performing event tuple anomaly detection on the unit input data includes obtaining the activity ID of the unit input data and comparing it with a predefined activity list stored in the database, if the activity ID is in the predefined activity list, obtaining a first detection result; if the activity ID is not in the predefined activity list, obtaining a second detection result.

3. The anomaly data detection method based on deep learning according to claim 2, wherein Obtaining the first detection result includes directly determining the abnormal situation of the unit input data according to the logical expression, obtaining abnormal data, and alarming the abnormal data.

4. The anomaly data detection method based on deep learning according to claim 2, wherein, Determining whether to execute the second data detection program based on the real-time anomaly intensity change amount and the actual dependency includes determining that the input data is suspected abnormal data and obtaining the real-time anomaly intensity change amount, and comparing it with the mild anomaly dependency intensity change amount range, if the real-time anomaly intensity change amount is outside the mild anomaly intensity change amount range, obtaining abnormal data and alarming the abnormal data; if the real-time dependency intensity change amount is within the mild anomaly intensity change amount range, obtaining the actual dependency and comparing it with the standard dependency, if the actual dependency is greater than the standard dependency, execute the second data detection program to determine the abnormal data; if the actual dependency is less than or equal to the standard dependency, execute the third data detection program to determine the abnormal data.

5. The anomaly data detection method based on deep learning according to claim 4, wherein, Executing the second data detection program includes dynamic binning, wherein the process of the dynamic binning is segmenting the unit input data into several input data subsets with a preset field length; respectively obtaining the historical hit frequencies of each input data subset, and sorting them in descending order according to the historical hit frequencies; calculating the real-time correlation degree between two adjacent input data subsets; linking each input data subset based on the real-time correlation degree, obtaining the real-time correlation degree, and comparing it with the standard correlation degree, wherein if the real-time correlation degree is greater than or equal to the standard correlation degree, linking the input data subsets to form several input data sets; if the real-time correlation degree is less than the standard correlation degree, the input data subsets are not linked.

6. The anomaly data detection method based on deep learning according to claim 4, wherein Executing the second data detection program also includes constructing a decision tree model, wherein the process of constructing the decision tree model is to use each input data set as training data to construct a decision tree model.

7. The anomaly data detection method based on deep learning according to claim 4, wherein Executing the second data detection program also includes detecting path optimization, wherein the process of the detecting path optimization is starting from the root node, traversing the splitting points of the splittable fields, and calculating the splitting gain of each splitting point; determining the optimal split and the optimal detection path of each node based on the splitting gain. Among them, the nodes include leaf nodes and non-leaf nodes.

8. The anomaly data detection method based on deep learning according to claim 4, wherein Performing the third data detection program includes screening all suspected abnormal data under the actual high dependence degree, selecting the suspected abnormal data with the highest proportion of radiation influence on other suspected abnormal data within a unit time as the missing parameter, setting the missing parameter as the continuous monitoring target, and triggering a potential risk alarm; Among them, the calculation formula for the radiation influence proportion is Among them, R i is the influence ratio of data D i on the overall abnormal data set; W i,j is the influence weight of data D i on data D j ; A j is the degree of abnormality of data D j ; ∑ k A k is the total outlier value of all outlier data points.

9. The anomaly data detection method based on deep learning according to claim 6, characterized in that The calculation formula for the splitting gain is ΔC = C node + P L * C(D L ) + P R * C(D R ) - C current Among them, C node is the node judgment overhead; C(D L ) is the left node detection overhead, and C(D R ) is the right node detection overhead; P L * is the probability of entering the left child node, P R * is the probability of entering the right child node; C current Is the cost of the current subtree of this node.

10. The anomaly data detection method based on deep learning according to claim 4, wherein Performing the second data detection program to determine abnormal data includes judging whether the non-abnormal data sample is known data, and incorporating the non-abnormal data into the data dependence information library; Among them, the data dependence information library includes a historical input data set, feedback adjustment binning rule parameters, and the number of branches of the decision tree model.

Citation Information

Patent Citations

  • Factory personnel abnormal behavior detection method based on self-supervised learning

    CN117935357A

Cited By

  • Electrical equipment monitoring method, system and equipment based on deep learning and medium

    CN121256418A

  • Production line exception processing method and device, equipment and medium

    CN121581357A

  • Production line exception processing method, device, equipment and medium

    CN121581357B