Sensor monitoring data anomaly detection auxiliary filler training method, device, equipment and storage medium
By performing abnormal detection and iterative training on sensor monitoring data, the problem of missing data filling in the prior art is solved, and an improved data filling accuracy is achieved.
Patent Information
- Application Number
- CN202510207424.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-25
AI Technical Summary
When filling the missing data in the data set, the prior art fails to effectively consider the situation where abnormal data may exist in the data before and after the missing data, resulting in poor accuracy of filling the data.
Through a sensor monitoring data abnormality detection auxiliary filler training method, it includes training the untrained filler based on the original data, performing abnormality detection, processing abnormal data, iterative training of the filler until the preset iteration stop condition is met, and the target filler is obtained.
Through iterative training, the target filler obtained by the acquisition of iterative training can more accurately determine the abnormal data in the data set, thereby improving the filling effect of missing data.
Smart Images

Figure CN119691372B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a sensor monitoring data anomaly detection auxiliary filler training method, device, equipment and storage medium. Background Art
[0002] There are often data anomalies and missing data in time series datasets. Since the datasets are generally used in subsequent downstream tasks, such as context recognition and predictive maintenance, in order to prevent the missing data from having an adverse impact on downstream tasks, the missing data in the dataset needs to be filled.
[0003] At present, the method for filling the missing data in the data set is: analyze the data rules of the data before and after the missing data, determine the data to be filled according to the data rules, and train the sensor monitoring data anomaly detection auxiliary filler to the corresponding position in the data set.
[0004] However, the current method of filling missing data in a data set does not take into account the possibility that there may be abnormal data in the data before and after the missing data. If there are abnormal data in the data before and after the missing data, the accuracy of the data to be filled is poor. It can be seen that the effect of filling missing data in a data set using existing technology is poor. Summary of the invention
[0005] In order to improve the effect of filling missing data in a data set, the present application provides a sensor monitoring data anomaly detection auxiliary filler training method, device, equipment and storage medium.
[0006] In a first aspect, the present application provides a sensor monitoring data anomaly detection auxiliary filler training method, comprising:
[0007] The untrained filler is trained based on the original data to obtain a pre-trained filler;
[0008] Performing anomaly detection on the original data to obtain abnormal data and a first label;
[0009] Processing the abnormal data based on the pre-trained filler to obtain filled data, and performing anomaly detection on the filled data to obtain a second label;
[0010] The pre-trained filler is iteratively trained based on the original data, the first label, and the second label until a preset iteration stop condition is met to obtain a target filler.
[0011] In a second aspect, the present application provides a sensor monitoring data anomaly detection auxiliary filler training device, comprising:
[0012] A pre-training module, used for training an untrained filler based on original data to obtain a pre-trained filler;
[0013] A detection module, configured to perform anomaly detection on the original data to obtain abnormal data and a first label;
[0014] A processing module, configured to process the abnormal data based on the pre-trained filler to obtain filled data, and perform anomaly detection on the filled data to obtain a second label;
[0015] An iteration module is used to iteratively train the pre-trained filler based on the original data, the first label and the second label until a preset iteration stop condition is met to obtain a target filler.
[0016] In a third aspect, the present application provides a computer device, the computer device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the steps in the above method when executing the computer program.
[0017] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the above method when executed by a processor.
[0018] In a fifth aspect, the present application further provides a computer program product, wherein the computer program product comprises a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.
[0019] The above-mentioned sensor monitoring data anomaly detection auxiliary filler training method, device, equipment and storage medium obtain a pre-trained filler by training an untrained filler based on original data; perform anomaly detection on the original data to obtain abnormal data and a first label; process the abnormal data based on the pre-trained filler to obtain filled data, and perform anomaly detection on the filled data to obtain a second label; iteratively train the pre-trained filler based on the original data, the first label and the second label until a preset iteration stop condition is met to obtain a target filler; through the above implementation, the original data can be processed by the obtained first label and second label to determine the accuracy of the abnormal data existing in the original data, and then the pre-trained filler is iteratively trained with the processed original data to obtain the target filler, which can improve the accuracy of the target filler in selecting the data to be filled, thereby improving the effect of the target filler in filling the missing data in the data set.
[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 This is a diagram of the application environment of the sensor monitoring data anomaly detection auxiliary filler training method in one embodiment of the present application;
[0023] Figure 2 A flow chart of a sensor monitoring data anomaly detection auxiliary filler training method provided in Example 1 of the present application;
[0024] Figure 3 A schematic diagram for reflecting first abnormal data provided in Embodiment 1 of the present application;
[0025] Figure 4 This is a schematic diagram for reflecting filled data provided in Example 1 of the present application;
[0026] Figure 5 A schematic diagram for reflecting second abnormal data provided in Embodiment 1 of the present application;
[0027] Figure 6 This is a structural schematic diagram of a sensor monitoring data anomaly detection auxiliary filler training device provided in one embodiment of the present application;
[0028] Figure 7 A schematic diagram of the structure of a computer device provided in one embodiment of the present application;
[0029] Figure 8 An internal structure diagram of a computer-readable storage medium provided in one embodiment of the present application. DETAILED DESCRIPTION
[0030] In order to make the purpose, technical solution and advantages of the present disclosure more clear, the present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.
[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of this article and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of this article described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, device, product or equipment that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0032] In this article, the term "and / or" is only a description of the association relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the related objects before and after are in an "or" relationship.
[0033] To solve the above problems, the present disclosure provides a sensor monitoring data anomaly detection auxiliary filler training method, which can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. Among them, the terminal 102 can be but is not limited to various personal computers, laptops, smart phones, tablets, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented with an independent server or a server cluster consisting of multiple servers.
[0034] Embodiment 1
[0035] Figure 2 A flow chart of a sensor monitoring data anomaly detection auxiliary filler training method provided in Example 1 of the present application, refer to Figure 2 The method may be performed by a device for performing the method, and the device may be implemented by software and / or hardware. The method includes:
[0036] S110, training an untrained filler based on original data to obtain a pre-trained filler.
[0037] Among them, the original data contains a certain number of abnormal data, and the abnormal data may be data with abnormal values and / or data with missing values. The original data is used to obtain training data after subsequent processing, and the untrained filler is trained according to the training data. The untrained filler is also called an untrained filler, and the pre-trained filler is a filler preliminarily trained with the processed original data. It should be noted that the pre-trained filler has only completed the preliminary training at this time and has not yet met the requirements for data filling. The pre-trained filler needs to be further trained later.
[0038] Specifically, the original data is processed to obtain training data, and the untrained filler is trained according to the training data, thereby obtaining a pre-trained filler.
[0039] It should be noted that the network structure used by the untrained filler in this embodiment may be other complex network structures such as a multi-layer perceptron (MLP), RNN or Transformer.
[0040] S120: Perform anomaly detection on the original data to obtain abnormal data and a first label.
[0041] Among them, anomaly detection is performed on the original data using an anomaly detector that has been trained in advance. Exemplarily, the anomaly detector can be specifically a data anomaly detection model, and the abnormal data is the abnormal data in the original data detected by the anomaly detector. After the original data is anomaly detected, a corresponding label is set for each data in the original data, that is, a first label. The first label includes two situations: normal and abnormal. The first label of the data detected as abnormal in the original data is abnormal, and the first label of the data not detected as abnormal in the original data is normal.
[0042] Specifically, an anomaly detector that has been trained in advance performs an anomaly detection on the original data to obtain abnormal data in the original data, and the abnormal data in the original data is recorded as the first abnormal data, and the data of the non-abnormal data in the original data is recorded as the first normal data; Figure 3 , the blue dots are the original data, the orange dots are the output data of the anomaly detector, wherein the red dots are the anomaly data in the output data, that is, the first anomaly data; a first label is set for the first anomaly data, and the first label is abnormal, and a first label is set for the first normal data, and the first label is normal.
[0043] It should be noted that the specific method in which the anomaly detector in the present embodiment detects abnormal data in the original data is: the original data is input into the anomaly detector, the anomaly detector calculates the reconstructed data corresponding to the original data, and then calculates the reconstruction error between the original data and the reconstructed data. Taking the reconstruction error corresponding to a data in the original data as an example, it is determined whether the reconstruction error is greater than a preset error threshold. If so, the data is determined to be abnormal data.
[0044] In addition, in other embodiments, the algorithm of the anomaly detector for detecting abnormal data in the original data can also be based on 3sigma, box plot, isolation forest (IF), kernel density estimation (KDE), one-class SVM, deep learning algorithms such as autoencoder (AutoEncoder), prediction-type networks (Transformer, LSTM, RNN, etc.), generative networks (GAN) and Anomaly Transformer based on Transformer class, etc.
[0045] In addition, in this embodiment, the network structure used by the anomaly detector is specifically a multi-layer perceptron (MLP). In other embodiments, the network structure used by the anomaly detector is not specifically limited.
[0046] S130, processing the abnormal data based on the pre-trained filler to obtain filled data, and performing anomaly detection on the filled data to obtain a second label.
[0047] Among them, the original data with the first label setting is input into the pre-trained filler, and the pre-trained filler can fill the abnormal data in the original data, and replace the abnormal data in the original data with the corresponding filling value, so as to obtain the filled data; there may still be some abnormal data in the filled data, and the anomaly detector can also perform anomaly detection on the filled data, so as to determine the abnormal data in the filled data. After completing the anomaly detection of the filled data, a corresponding label is set for each data in the filled data, that is, the second label, and the second label includes two situations: normal and abnormal; the second label set for the abnormal data in the filled data is abnormal, and the label set for the non-abnormal data in the filled data is normal.
[0048] Specifically, the original data is input into the pre-trained filler, and the abnormal data in the original data is replaced with the corresponding filling value by the pre-trained filler, so as to obtain the filled data; Figure 4 , the orange dot in the red box is the filled data; further, the filled data is input into the anomaly detector for anomaly detection, so as to determine the abnormal data in the filled data, and the abnormal data in the filled data is recorded as the second abnormal data, referring to Figure 5, the red dot is the second abnormal data; the data of the non-abnormal data in the filled data is also recorded as the second normal data, a second label is set for the second abnormal data, and the second label is abnormal, and a second label is set for the second normal data, and the second label is normal.
[0049] S140, iteratively training the pre-trained filler based on the original data, the first label, and the second label until a preset iteration stop condition is met to obtain a target filler.
[0050] Among them, through the implementation of steps S120-S130, the first label and the second label can be set for each data in the original data, and the original data can be processed and new training data can be obtained through the first label and the second label of each data in the original data. The new training data is used to iteratively train the pre-trained filler to obtain a new pre-trained filler. The iterative training is to apply the new pre-trained filler to the above-mentioned S130 step to replace the original pre-trained filler, and then execute the S120-S140 steps again to obtain the new pre-trained filler again, and then execute the above steps in a loop to realize iterative training of the pre-trained filler until the preset iteration stop condition is met, and the pre-trained filler obtained when the preset iteration stop condition is met is recorded as the target filler, and the target filler is also the filler that finally completes the training.
[0051] Specifically, the original data is processed by the first label and the second label of each data in the original data to obtain new training data, and then the pre-trained filler is iteratively trained with the new training data until a preset iteration stop condition is met to obtain a target filler.
[0052] It should be noted that, in this embodiment, a pre-trained filler is obtained by training an untrained filler based on the original data; anomaly detection is performed on the original data to obtain abnormal data and a first label; the abnormal data is processed based on the pre-trained filler to obtain filled data, and anomaly detection is performed on the filled data to obtain a second label; the pre-trained filler is iteratively trained based on the original data, the first label and the second label until a preset iteration stop condition is met to obtain a target filler; through the above implementation, the original data can be processed by the obtained first label and the second label to determine the accuracy of the abnormal data existing in the original data, and then the pre-trained filler is iteratively trained by the processed original data to obtain the target filler, which can improve the accuracy of the target filler in selecting the data to be filled, thereby improving the effect of the target filler in filling the missing data in the data set.
[0053] Embodiment 2
[0054] Embodiment 2 of the present application provides a sensor monitoring data anomaly detection auxiliary filler training method, which optimizes the "training an untrained filler based on raw data to obtain a pre-trained filler" in Embodiment 1; it should be noted that for the parts not described in detail in this embodiment, reference can be made to the descriptions of other embodiments, and the method includes:
[0055] S211, preprocessing the original data to obtain preprocessed data.
[0056] Among them, preprocessing includes one or more of deduplication operations, standardization operations and normalization operations on the original data; in this embodiment, the original data is first deduplicated to obtain deduplicated data, and then the deduplicated data is standardized to obtain standardized data, and the standardized data is used as preprocessed data. In another embodiment, the original data is first deduplicated to obtain deduplicated data, and then the deduplicated data is normalized to obtain normalized data, and the normalized data is used as preprocessed data; in other embodiments, the specific implementation form of preprocessing the original data is not limited.
[0057] Specifically, the original data is preprocessed to obtain preprocessed data.
[0058] S212: Process the preprocessed data based on a preset missing rate to obtain primary missing data.
[0059] Among them, before training the filler, the user can preset the hyperparameters for controlling the filler training. In this embodiment, the hyperparameters include the missing rate, and the missing rate is set to 25%. The preprocessed data is processed based on the preset missing rate, that is, 25% of the data in the preprocessed data is randomly selected, and then the 25% of the data is replaced by the value 0; in another embodiment, 25% of the data in the preprocessed data can be selected according to the binomial distribution, and then the 25% of the data is replaced by the value 0; in other embodiments, the method of selecting 25% of the data in the preprocessed data is not specifically limited; the primary missing data is the data obtained after processing the preprocessed data based on the preset missing rate.
[0060] Specifically, data satisfying a preset missing rate in the preprocessed data are selected, and the values of the selected data are replaced with 0, thereby obtaining primary missing data.
[0061] S213, dividing the primary missing data into a training set and a validation set.
[0062] Among them, the primary missing data is subsequently used to train the untrained filler. For this purpose, the primary missing data needs to be divided. The user presets the hyperparameters for controlling the filler training, including the division parameters. The missing data can be divided by the division parameters. By dividing the primary missing data, the primary missing data can be divided into a training set and a validation set. Both the training set and the validation set are used to train the untrained filler.
[0063] Specifically, the primary missing data is divided according to the division parameters preset by the user in the form of hyperparameters, thereby obtaining a training set and a validation set.
[0064] S214. Train an untrained filler based on the training set and the validation set to obtain a pre-trained filler.
[0065] Among them, the training set and the validation set are derived from the primary missing data. It is known that some data in the primary missing data are selected to be replaced by 0, and the data selected to be replaced by 0 in the primary missing data is recorded as missing data, and the other data in the primary missing data except the missing data is recorded as non-missing data; therefore, there are missing data and non-missing data in both the training set and the validation set. The missing loss can be constructed by processing the missing data in the training set through the untrained filler , the non-missing data in the training set can be processed by the untrained filler to construct the non-missing loss It should be noted that users can preset the missing loss in the form of hyperparameters The loss weight α, so that the missing loss can be further calculated The loss weight is (1-α), further, according to the missing loss , missing loss The loss weight is (1-α), the non-missing loss , No missing loss The loss weight α constructs the initial training loss function , where the initial training loss function is The expression is as follows: ; Through the initial training loss function The untrained filler can be trained with the validation set to optimize the parameters of the untrained filler, thereby obtaining an untrained filler that has completed preliminary training, namely, a pre-trained filler.
[0066] Specifically, the training set is input into the untrained filler, which further constructs the missing loss by processing the missing data in the training set. The untrained filler further constructs the non-missing loss by processing the non-missing data in the training set. , then, according to the missing loss , preset missing loss The loss weight is (1-α), the non-missing loss , the default non-missing loss The loss weight α constructs the initial training loss function , and then according to the initial training loss function The untrained filler is trained with the validation set to optimize the parameters of the untrained filler to obtain a pre-trained filler.
[0067] S220: Perform anomaly detection on the original data to obtain abnormal data and a first label.
[0068] S230: Process the abnormal data based on the pre-trained filler to obtain filled data, and perform anomaly detection on the filled data to obtain a second label.
[0069] S240, iteratively training the pre-trained filler based on the original data, the first label, and the second label until a preset iteration stop condition is met to obtain a target filler.
[0070] Embodiment 3
[0071] Embodiment 3 of the present application provides a sensor monitoring data anomaly detection auxiliary filler training method, which optimizes the "dividing the primary missing data to obtain a training set and a validation set" in Embodiment 2; it should be noted that for the parts not described in detail in this embodiment, reference can be made to the descriptions of other embodiments, and the method includes:
[0072] S311, preprocessing the original data to obtain preprocessed data.
[0073] S312: Process the preprocessed data based on a preset missing rate to obtain primary missing data.
[0074] S313A: Divide the primary missing data based on a preset time window to obtain a data set.
[0075] Among them, the user presets the hyperparameters for controlling the filler training, including the time window division parameters, which specifically include the window length windows and the moving step step; the time window can be constructed through the window length windows and the moving step step, and the primary missing data can be divided through the time window, thereby dividing the primary missing data into multiple data sets.
[0076] Specifically, a corresponding time window is constructed according to the window length windows preset by the user with a hyperparameter, and the time window is moved on the primary missing data based on the moving step step, thereby realizing the division of the primary missing data to obtain multiple data sets.
[0077] S313B, dividing the data set into a training set and a validation set based on a preset division ratio.
[0078] Among them, the user-preset hyperparameters for controlling filler training also include the data set partition ratio a:b:c, which is used to divide the data in a single data set, thereby dividing the single data set into a training set, a validation set, and a test set.
[0079] Specifically, taking a data set divided according to primary missing data as an example, the data set is divided into a training set, a validation set, and a test set in sequence according to the data set division ratio a:b:c.
[0080] S314. Train an untrained filler based on the training set and the validation set to obtain a pre-trained filler.
[0081] S320: Perform anomaly detection on the original data to obtain abnormal data and a first label.
[0082] S330: Process the abnormal data based on the pre-trained filler to obtain filled data, and perform anomaly detection on the filled data to obtain a second label.
[0083] S340, iteratively training the pre-trained filler based on the original data, the first label, and the second label until a preset iteration stop condition is met to obtain a target filler.
[0084] Embodiment 4
[0085] Embodiment 4 of the present application provides a sensor monitoring data anomaly detection auxiliary filler training method, which optimizes "iteratively training the pre-trained filler based on the original data, the first label and the second label until a preset iteration stop condition is met to obtain a target filler" in Embodiment 1; it should be noted that for the parts not described in detail in this embodiment, reference may be made to the descriptions of other embodiments, and the method includes:
[0086] S410, training an untrained filler based on original data to obtain a pre-trained filler.
[0087] S420: Perform anomaly detection on the original data to obtain abnormal data and a first label.
[0088] S430: Process the abnormal data based on the pre-trained filler to obtain filled data, and perform anomaly detection on the filled data to obtain a second label.
[0089] S441. Process the original data based on the first label and the second label to obtain secondary missing data.
[0090] Among them, the first label corresponding to each data in the original data can be determined through steps S420-S430 With the second label , and the first label corresponding to each data With the second label Meet one of the following conditions:
[0091] First: abnormal, normal;
[0092] second: abnormal, abnormal;
[0093] third: normal, normal;
[0094] fourth: normal, abnormal;
[0095] It should be noted that if the first label corresponding to a data in the original data With the second label conform to" abnormal, Abnormal" indicates that the data is highly likely to be truly abnormal data. In order to prevent the data from adversely affecting the subsequent training of the pre-selected filler, the data is constructed with missing values, that is, the data is replaced with the value 0; abnormal, The "abnormal" data are all constructed with missing values to obtain secondary missing data.
[0096] Specifically, the first label corresponding to each data in the original data is obtained With the second label , and then determine the first label corresponding to the data With the second label Whether it meets the " abnormal, If it is abnormal, the data is constructed with missing values to obtain secondary missing data.
[0097] S442. Based on the secondary missing data, the pre-trained filler is trained to obtain a new pre-trained filler. Based on the new pre-trained filler, the step of obtaining the second label is cyclically executed until a preset number of single iteration cycles is met, thereby obtaining a new first label and a new second label.
[0098] Among them, by inputting the secondary missing data into the pre-trained filler for processing, a corresponding training loss can be constructed according to the processing result obtained after the processing and the secondary missing data, and the pre-trained filler can be trained according to the training loss to obtain a new pre-trained filler; the step of obtaining the second label is also the above-mentioned steps S420-S430, and the step of cyclically executing the second label based on the new pre-trained filler is also to replace the pre-trained filler in the first execution of steps S420-S430 with the new pre-trained filler obtained in step S442, and executing steps S420-S430 again to obtain a new first label and a new second label, and then executing the step of obtaining the second label based on the new pre-trained filler again, and so on, so as to realize the step of cyclically executing the second label based on the new pre-trained filler; when the number of cycles of the step of cyclically executing the second label based on the new pre-trained filler reaches the preset number of single iteration cycles, the step of cyclically executing the second label based on the new pre-trained filler can be terminated.
[0099] It should be noted that, in the process of cyclically executing the step of cyclically executing the step of obtaining the second label based on the new pre-trained filler, each time the step of obtaining the second label is executed based on the new pre-trained filler, it includes a process of training the pre-trained filler obtained last time to obtain a new pre-trained filler, and the process specifically uses the training loss to train the pre-trained filler obtained last time, wherein the construction process of the training loss is as follows: Determine the first label in the secondary missing data With the second label conform to" abnormal, The first data is obtained from the "normal" data, and the first label in the secondary missing data is determined With the second label conform to" abnormal, The second data is obtained by using the "abnormal" data to determine the first label in the secondary missing data With the second label conform to" normal, The third data is obtained by using the "normal" data to determine the first label in the secondary missing data With the second label conform to" normal, The fourth data is obtained by processing the first data through the pre-trained filler obtained last time to construct a first loss corresponding to the first data , the second data is processed by the pre-trained filler obtained last time to construct a second loss corresponding to the second data , the third data is processed by the pre-trained filler obtained last time to construct the third loss corresponding to the third data , the fourth data is processed by the pre-trained filler obtained last time to construct the fourth loss corresponding to the fourth data ; It should be noted that the user preset hyperparameters for controlling filler training also include: abnormal, Normal" sets the first loss weight weight_1, for " abnormal, The second loss weight weight_2 set by "abnormal" is for " normal, Normal" setting of the third loss weight weight_3, for " normal, The fourth loss weight weight_4 set by "abnormal"; training loss Based on the first loss Second loss , third loss 4. Loss , the first loss weight weight_1, the second loss weight weight_2, the third loss weight weight_3 and the fourth loss weight weight_4 are established, and the training loss The expression is as follows: .
[0100] Specifically, the first data, the second data, the third data, and the fourth data in the secondary missing data are processed by the pre-trained filler to establish the training loss , and then according to the training loss The pre-trained filler is trained to obtain a new pre-trained filler; then the new pre-trained filler replaces the pre-trained filler in the first execution of steps S420-S430, and steps S420-S430 are executed again to obtain a new first label and a new second label, and then the secondary missing data is processed according to the new first label and the new second label to obtain new first data, second data, third data and fourth data again, and then the new first data, second data, third data and fourth data are processed according to the new pre-trained filler to establish a new training loss. , and then according to the new training loss Train the new pre-trained filler to obtain a new pre-trained filler again; then execute the step of obtaining the second label based on the new pre-trained filler again, and so on, so as to realize the step of cyclically executing the step of obtaining the second label based on the new pre-trained filler; when the number of cycles of the step of cyclically executing the step of obtaining the second label based on the new pre-trained filler reaches the preset number of single iteration cycles, the step of cyclically executing the step of obtaining the second label based on the new pre-trained filler can be ended, so as to obtain a new first label and a new second label.
[0101] It should be noted that the above steps can be performed to implement single-iteration training of the pre-trained filler, and multiple rounds of training of the pre-trained filler are implemented in the single-iteration training, and the number of multiple rounds of training is also the preset number of single-iteration cycles.
[0102] S443, iteratively training the new pre-trained filler based on the new first label, the new second label and the original data until a preset iteration stop condition is met to obtain a target filler.
[0103] Among them, the new first label and the new second label obtained after completing a single iterative training of the pre-trained filler can be used to process the original data, so as to obtain the training data required for the second iterative training of the new pre-trained filler, and then the new pre-trained filler is trained for the second time through the training data. After completing the second iterative training, another set of new first labels and new second labels can be obtained, and then the original data is processed by the set of new first labels and new second labels, so as to obtain the training data required for the third iterative training of the new pre-trained filler, and so on, to achieve iterative training of the new pre-trained filler; the iteration stop condition is that the new pre-trained filler is iteratively trained to meet the preset iteration number threshold, or the first label and the second label corresponding to each data in the original data no longer change.
[0104] Specifically, taking one of the training of the new pre-trained filler as an example, the new first label and the new second label obtained from the last iterative training are obtained, and the original data is processed according to the new first label and the new second label to obtain the training data for the current iterative training of the new pre-trained filler, and the current iterative training of the new pre-trained filler is performed according to the training data, and so on, to achieve iterative training of the new pre-trained filler until the preset iteration stop condition is met and the target filler is obtained.
[0105] Embodiment 5
[0106] Embodiment 5 of the present application provides a sensor monitoring data anomaly detection auxiliary filler training method, which optimizes the "training the pre-trained filler based on the secondary missing data to obtain a new pre-trained filler" in Embodiment 4; it should be noted that for the parts not described in detail in this embodiment, reference can be made to the descriptions of other embodiments, and the method includes:
[0107] S510: Train the untrained filler based on the original data to obtain a pre-trained filler.
[0108] S520: Perform anomaly detection on the original data to obtain abnormal data and a first label.
[0109] S530: Process the abnormal data based on the pre-trained filler to obtain filled data, and perform anomaly detection on the filled data to obtain a second label.
[0110] S541. Process the original data based on the first label and the second label to obtain secondary missing data.
[0111] S542A, generating training data based on the secondary missing data, and processing the training data based on the pre-trained filler to obtain a processing result.
[0112] Among them, in this embodiment, training data is generated based on secondary missing data. It is necessary to first divide the secondary missing data according to the window length windows preset by the user with hyperparameters and the time window determined by the moving step step, so as to obtain multiple data sets. Taking one of the data sets as an example, the data set is divided according to the data set division ratio a:b:c preset by the user with hyperparameters to obtain a training set, a validation set and a test set, where the training data is also the training set; the processing result is the result output by the pre-training filler after processing the training data.
[0113] Specifically, the time window is determined according to the window length windows and the moving step step preset by the user with hyperparameters, and the secondary missing data is divided into multiple data sets according to the time window, and then the data sets are divided according to the preset data set division ratio. The data set is divided into training set, validation set and test set, and then the training data is processed by the pre-training filler to obtain the processing results.
[0114] S542B, constructing losses corresponding to different combinations of the first label and the second label based on the training data and the processing results.
[0115] The training data comes from the secondary missing data, and each data in the secondary missing data has its first label With the second label , first label With the second label The different combinations are as follows:
[0116] First: abnormal, normal;
[0117] second: abnormal, abnormal;
[0118] third: normal, normal;
[0119] fourth: normal, abnormal;
[0120] And the first label in the training data With the second label conform to" abnormal, The "normal" data is recorded as the first data, and the first label in the training data With the second label conform to" abnormal, The data with the first label in the training data is recorded as the second data. With the second label conform to" normal, The "normal" data is recorded as the third data, and the first label in the training data is With the second label conform to" normal, The "abnormal" data is recorded as the fourth data; the pre-trained filler processes the first data to obtain a processing result corresponding to the first data, which is recorded as the first processing result, and the corresponding first loss is constructed according to the first data and the first processing result. ; The pre-trained filler processes the second data to obtain a processing result corresponding to the second data, which is recorded as the second processing result, and a corresponding second loss is constructed according to the second data and the second processing result The pre-trained filler processes the third data to obtain a processing result corresponding to the third data, which is recorded as the third processing result. The corresponding third loss is constructed according to the third data and the third processing result. The pre-trained filler processes the fourth data to obtain a processing result corresponding to the fourth data, which is recorded as the fourth processing result. The corresponding fourth loss is constructed according to the fourth data and the fourth processing result. .
[0121] Specifically, a first loss is constructed according to the first data and the first processing result corresponding to the first data. ; Construct a second loss according to the second data and the second processing result corresponding to the second data ; Construct a third loss according to the third data and the third processing result corresponding to the third data ; Construct a fourth loss according to the fourth data and the fourth processing result corresponding to the fourth data .
[0122] S542C. Construct a model training loss based on the losses corresponding to different combinations of the first label and the second label, and the preset loss weights corresponding to different combinations of the first label and the second label.
[0123] Among them, the user preset hyperparameters for controlling filler training also include: abnormal, Normal" sets the first loss weight weight_1, for " abnormal, The second loss weight weight_2 set by "abnormal" is for " normal, Normal" setting of the third loss weight weight_3, for " normal, The fourth loss weight weight_4 set by "abnormal" is based on the above first loss Second loss , third loss 4. Loss , the first loss weight weight_1, the second loss weight weight_2, the third loss weight weight_3 and the fourth loss weight weight_4 can construct the corresponding model training loss , model training loss The expression is as follows: .
[0124] Specifically, based on the above first loss Second loss , third loss 4. Loss , the first loss weight weight_1, the second loss weight weight_2, the third loss weight weight_3 and the fourth loss weight weight_4 construct the corresponding model training loss , .
[0125] S542D. Train the pre-trained filler based on the model training loss to obtain a new pre-trained filler.
[0126] Specifically, according to the model training loss The pre-trained filler is trained to obtain a new pre-trained filler.
[0127] S542E: loop through the steps of obtaining the second label based on the new pre-trained filler until a preset number of single iteration cycles is met, thereby obtaining a new first label and a new second label.
[0128] S543, iteratively training the new pre-trained filler based on the new first label, the new second label and the original data until a preset iteration stop condition is met to obtain a target filler.
[0129] Embodiment 6
[0130] Embodiment 6 of the present application provides a sensor monitoring data anomaly detection auxiliary filler training method, which optimizes the "iterative training of the new pre-trained filler based on the new first label, the new second label and the original data until a preset iteration stop condition is satisfied to obtain a target filler" in Embodiment 4; it should be noted that for the parts not described in detail in this embodiment, reference may be made to the descriptions of other embodiments, and the method includes:
[0131] S610: Train the untrained filler based on the original data to obtain a pre-trained filler.
[0132] S620: Perform anomaly detection on the original data to obtain abnormal data and a first label.
[0133] S630: Process the abnormal data based on the pre-trained filler to obtain filled data, and perform anomaly detection on the filled data to obtain a second label.
[0134] S641. Process the original data based on the first label and the second label to obtain secondary missing data.
[0135] S642. Based on the secondary missing data, the pre-trained filler is trained to obtain a new pre-trained filler. Based on the new pre-trained filler, the step of obtaining the second label is cyclically executed until a preset number of single iteration cycles is met, thereby obtaining a new first label and a new second label.
[0136] S643A: Process the original data based on the new first label and the new second label to obtain cyclic missing data.
[0137] Among them, each data in the original data has a corresponding first label With the second label , if the first label corresponding to a data in the original data With the second label for" abnormal, Abnormal", it means that the data is likely to be truly abnormal data, and the data needs to be subsequently constructed with missing values, that is, the value of the data is replaced with 0. The data after the above missing value construction is performed on the original data is also called circular missing data.
[0138] Specifically, obtain the first label corresponding to each data in the original data With the second label , and then determine the first label corresponding to the data With the second label Is it " abnormal, If it is abnormal, the data is constructed with missing values to obtain cyclic missing data.
[0139] S643B, iteratively train the new pre-trained filler based on the loop missing data until a preset iteration stop condition is met to obtain a target filler.
[0140] Among them, the cyclic missing data can be input into the new pre-trained filler for processing to obtain the corresponding processing result, and the corresponding training loss can be constructed through the cyclic missing data and the corresponding processing result, and then the new pre-trained filler can be trained according to the training loss to obtain a new pre-trained filler again, and then the S642 step is re-executed to realize the second iterative training of the pre-trained filler, and so on, to realize the third iterative training of the pre-trained filler..., thereby realizing the iterative training of the pre-trained filler; the iteration stop condition is to iteratively train the new pre-trained filler to meet the preset iteration number threshold, or the first label and the second label corresponding to each data in the original data no longer change.
[0141] Specifically, the loop missing data is input into a new pre-trained filler for processing to obtain a corresponding processing result, a corresponding training loss is constructed according to the loop missing data and the corresponding processing result, the new pre-trained filler is trained according to the training loss to obtain a new pre-trained filler again, and then step S642 is re-executed to achieve the second iterative training of the pre-trained filler, and so on, to achieve the third iterative training of the pre-trained filler..., thereby achieving iterative training of the pre-trained filler, in response to reaching a preset iteration stop condition, the iterative training of the pre-trained filler is stopped, and the pre-trained filler that has finally completed the training, that is, the target filler, is obtained.
[0142] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0143] Embodiment 7
[0144] Based on the same inventive concept, the embodiment of the present disclosure also provides a sensor monitoring data anomaly detection auxiliary filler training device for implementing the sensor monitoring data anomaly detection auxiliary filler training method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more sensor monitoring data anomaly detection auxiliary filler training device embodiments provided below can refer to the above limitations on the sensor monitoring data anomaly detection auxiliary filler training method, and will not be repeated here.
[0145] In this embodiment, Figure 6 As shown, a sensor monitoring data anomaly detection auxiliary filler training device is provided, comprising:
[0146] A pre-training module, used for training an untrained filler based on original data to obtain a pre-trained filler;
[0147] A detection module, configured to perform anomaly detection on the original data to obtain abnormal data and a first label;
[0148] A processing module, configured to process the abnormal data based on the pre-trained filler to obtain filled data, and perform anomaly detection on the filled data to obtain a second label;
[0149] An iteration module is used to iteratively train the pre-trained filler based on the original data, the first label and the second label until a preset iteration stop condition is met to obtain a target filler.
[0150] Each module in the sensor monitoring data anomaly detection auxiliary filler training device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.
[0151] It should be noted that, in this embodiment, a pre-trained filler is obtained by training an untrained filler based on the original data; anomaly detection is performed on the original data to obtain abnormal data and a first label; the abnormal data is processed based on the pre-trained filler to obtain filled data, and anomaly detection is performed on the filled data to obtain a second label; the pre-trained filler is iteratively trained based on the original data, the first label and the second label until a preset iteration stop condition is met to obtain a target filler; through the above implementation, the original data can be processed by the obtained first label and the second label to determine the accuracy of the abnormal data existing in the original data, and then the pre-trained filler is iteratively trained by the processed original data to obtain the target filler, which can improve the accuracy of the target filler in selecting the data to be filled, thereby improving the effect of the target filler in filling the missing data in the data set.
[0152] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a sensor monitoring data anomaly detection auxiliary filler training method is implemented.
[0153] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the scheme of the present disclosure, and does not constitute a limitation on the computer device to which the scheme of the present disclosure is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0154] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0155] In one embodiment, a computer readable storage medium is provided. Figure 8 As shown, a computer program is stored thereon, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0156] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0157] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.
[0158] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided by the present disclosure can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided by the present disclosure may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided by the present disclosure may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, etc., but are not limited to this.
[0159] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0160] The above-described embodiments only express several implementation methods of the present disclosure, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present disclosure, and these all belong to the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the attached claims.
Claims
1. A sensor monitoring data anomaly detection auxiliary filler training method, characterized in that: include: The untrained filler is trained based on the original data to obtain a pre-trained filler; Performing anomaly detection on the original data to obtain abnormal data and a first label; Processing the abnormal data based on the pre-trained filler to obtain filled data, and performing anomaly detection on the filled data to obtain a second label; Iteratively training the pre-trained filler based on the original data, the first label, and the second label until a preset iteration stop condition is met to obtain a target filler; The iterative training of the pre-trained filler based on the original data, the first label, and the second label until a preset iteration stop condition is met to obtain a target filler includes: Processing the original data based on the first label and the second label to obtain secondary missing data; Training the pre-trained filler based on the secondary missing data to obtain a new pre-trained filler, and cyclically executing the step of obtaining the second label based on the new pre-trained filler until a preset number of single iteration cycles is met, thereby obtaining a new first label and a new second label; The new pre-trained filler is iteratively trained based on the new first label, the new second label and the original data until a preset iteration stop condition is met to obtain a target filler.
2. The method according to claim 1, characterized in that The step of training the untrained filler based on the original data to obtain the pre-trained filler includes: Preprocessing the raw data to obtain preprocessed data; Processing the preprocessed data based on a preset missing rate to obtain primary missing data; Dividing the primary missing data into a training set and a validation set; An untrained filler is trained based on the training set and the validation set to obtain a pre-trained filler.
3. The method according to claim 2, characterized in that The dividing the primary missing data to obtain a training set and a validation set includes: Dividing the primary missing data based on a preset time window to obtain a data set; The data set is divided into a training set and a validation set based on a preset division ratio.
4. The method according to claim 1, characterized in that: The step of training the pre-trained filler based on the secondary missing data to obtain a new pre-trained filler comprises: Generate training data based on the secondary missing data, and process the training data based on the pre-trained filler to obtain a processing result; Constructing losses corresponding to different combinations of the first label and the second label based on the training data and the processing result; Constructing a model training loss based on losses corresponding to different combinations of the first label and the second label, and preset loss weights corresponding to different combinations of the first label and the second label; The pre-trained filler is trained based on the model training loss to obtain a new pre-trained filler.
5. The method according to claim 1, characterized in that The iterative training of the new pre-trained filler based on the new first label, the new second label and the original data until a preset iteration stop condition is met to obtain a target filler includes: Processing the original data based on the new first label and the new second label to obtain cyclic missing data; The new pre-trained filler is iteratively trained based on the cycle missing data until a preset iteration stop condition is met to obtain a target filler.
6. A sensor monitoring data anomaly detection auxiliary filler training device, characterized in that: The device comprises: A pre-training module, used for training an untrained filler based on original data to obtain a pre-trained filler; A detection module, used for performing anomaly detection on the original data to obtain abnormal data and a first label; A processing module, configured to process the abnormal data based on the pre-trained filler to obtain filled data, and perform anomaly detection on the filled data to obtain a second label; An iteration module, configured to iteratively train the pre-trained filler based on the original data, the first label, and the second label until a preset iteration stop condition is met to obtain a target filler; The iterative training of the pre-trained filler based on the original data, the first label, and the second label until a preset iteration stop condition is met to obtain a target filler includes: Processing the original data based on the first label and the second label to obtain secondary missing data; Training the pre-trained filler based on the secondary missing data to obtain a new pre-trained filler, and cyclically executing the step of obtaining the second label based on the new pre-trained filler until a preset number of single iteration cycles is met, thereby obtaining a new first label and a new second label; The new pre-trained filler is iteratively trained based on the new first label, the new second label and the original data until a preset iteration stop condition is met to obtain a target filler.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Method and device for controlling traffic system
CN114399901A
Method and device for anomaly detection and parameter filling of time series data
CN114826988A