An end-to-end anomaly detection method, system and medium based on positive and negative sample fusion

By adopting an end-to-end abnormality detection method based on positive and negative sample fusion in substation inspection, and using multi-stage training and improved YoloV8 network, the problem of low detection accuracy is solved, and the abnormality detection effect with high accuracy and high recall is achieved.

CN119557817BActive Publication Date: 2025-05-27STATE GRID SICHUAN ELECTRIC POWER CORP ELECTRIC POWER RES INST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510125530.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-27
Estimated Expiration
2045-01-27

AI Technical Summary

Technical Problem

During the current substation inspection, the positive and negative samples of abnormal detection are insufficiently integrated, resulting in low detection accuracy, making it difficult to reduce false alarms and improve recall rates at the same time.

Method used

Using an end-to-end anomaly detection method based on positive and negative sample fusion, an end-to-end anomaly detection network is built, and a comparison language-image pre-trained model and an improved YoloV8 network are used to perform multi-stage training to fuse the advantages of positive and negative samples, improving the accuracy and recall of detection.

Benefits of technology

By fusing positive and negative samples at the deep semantic feature level of the image, the accuracy and recall of abnormal detection are improved, effectively solving the problem that both accuracy and recall in traditional methods are difficult to take into account.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119557817B_ABST
    Figure CN119557817B_ABST
Patent Text Reader

Abstract

The present invention discloses an end-to-end anomaly detection method, system and medium based on positive and negative sample fusion; it relates to the technical field of anomaly detection; this solution improves the method on the basis of traditional charging detection technology, makes full use of the advantages of positive and negative samples, and improves the accuracy and recall rate of anomaly detection; by constructing a generated anomaly data set, a real anomaly data set and a real anomaly graphic and text description pair, the training set A, the training set B and the training set C are correspondingly constructed, and the end-to-end anomaly detection network is trained in stages with the training set A, the training set B and the training set C; ensuring that the end-to-end anomaly detection network can effectively learn the features of anomalies and reduce false alarms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of anomaly detection, and specifically relates to an end-to-end anomaly detection method, system and medium based on positive and negative sample fusion. Background Art

[0002] With the development of deep learning technology, anomaly detection is increasingly widely used in various fields, especially in the anomaly detection of substation inspections; currently, during the substation inspection process, positive sample anomaly detection and negative sample anomaly detection usually adopt different technical routes in substation inspection anomaly detection; positive sample anomaly detection usually adopts a contrast learning route, which has a high detection rate but more false alarms, and there is no category division for the detected anomalies; while negative sample anomaly detection adopts an object detection route, which has fewer false alarms, the detected anomalies have categories, but the detection rate is low; this difference makes it difficult to truly improve the recall and reduce the false alarms after the two algorithms run independently and then fuse the positive and negative samples.

[0003] Currently, there are two common schemes for fusing positive and negative samples during the substation inspection process; the first is to perform limited merging and filtering on the output results of the two methods after both positive sample anomaly detection and negative sample anomaly detection are completed; this method can effectively reduce false alarms, but the anomaly recall cannot be guaranteed. The second scheme is to use the output of the positive samples as prior knowledge for the negative samples, and divide the fusion of positive and negative samples into two stages; however, due to the large number of false alarms in the output of positive sample anomaly detection, using the results with more false alarms as the prior knowledge for negative sample detection will have a greater impact on the detection accuracy of negative samples.

[0004] As described above, currently during the substation inspection process, the method of fusing positive and negative samples does not fully utilize the advantages of positive sample anomaly detection and negative sample anomaly detection, but only fuses on the surface of the two anomaly detections, without fusing at a deeper level of the two anomaly detection algorithms. This makes it difficult to both reduce the false alarms of anomaly detection and maintain a high recall for anomalies. Summary of the Invention

[0005] The technical problem to be solved by the present invention is: the problem of insufficient fusion of positive and negative samples in anomaly detection during the current substation inspection process, resulting in low detection accuracy; the purpose of the present invention is to provide an end-to-end anomaly detection method, system and medium based on positive and negative sample fusion, which improves the method on the basis of traditional charging detection technology, fully utilizes the advantages of positive and negative samples, and improves the accuracy and recall rate of anomaly detection.

[0006] The present invention is achieved through the following technical solutions:

[0007] This solution provides an end-to-end anomaly detection method based on positive and negative sample fusion, including:

[0008] Construct an end-to-end anomaly detection network based on the fusion of positive and negative samples;

[0009] Construct a generated anomaly dataset, a real anomaly dataset, and a real anomaly graphic-text description pair, and correspondingly construct training set A, training set B, and training set C; both training set A and training set B include a test map set and a base map set;

[0010] Train the end-to-end anomaly detection network in stages with training set A, training set B, and training set C respectively;

[0011] Perform anomaly detection based on the trained end-to-end anomaly detection network.

[0012] The working principle of this solution: In the current substation inspection process, the problem of insufficient fusion of positive and negative samples for anomaly detection leads to low detection accuracy; the purpose of the present invention is to provide an end-to-end anomaly detection method, system, and medium based on the fusion of positive and negative samples, which improves the method on the basis of traditional charging detection technology, makes full use of the advantages of positive and negative samples, and improves the accuracy and recall rate of anomaly detection; by constructing a generated anomaly dataset, a real anomaly dataset, and a real anomaly graphic-text description pair, and correspondingly constructing training set A, training set B, and training set C, and training the end-to-end anomaly detection network in stages with training set A, training set B, and training set C; ensure that the end-to-end anomaly detection network can effectively learn the features of anomalies and reduce false alarms.

[0013] A further optimized solution is that the construction of the end-to-end anomaly detection network based on the fusion of positive and negative samples includes the following methods:

[0014] Use the Text Encoder branch in the contrastive language-image pre-training model (clip model) as the anomaly prompt text encoding network, and use the improved yolov8 network as the positive and negative sample fusion network;

[0015] Configure a detection head including a predicted class feature vector, a predicted anomaly location box, and an anomaly confidence level;

[0016] Compare the class feature vector with the text feature vector output by the anomaly prompt text encoding network, and determine the anomaly class according to the maximum cosine distance.

[0017] A further optimized solution is that the improved yolov8 network includes branch A, branch B, and a feature fusion network;

[0018] Branch A extracts features from a test map based on the traditional yolov8 network; branch B extracts features from multiple base maps based on 3D convolution and 2D convolution;

[0019] The feature fusion network is used to fuse the features extracted by branch A and the features extracted by branch B.

[0020] A further optimization solution is that the abnormal data set, the real abnormal data set, and the real abnormal graphic and text description pairs are constructed, and the training sets A, B, and C are correspondingly constructed; the method includes:

[0021] Select N pictures at the same point at different times to generate abnormal pictures to be detected and abnormal pictures not to be detected, and construct an abnormal data set from the abnormal pictures to be detected and the abnormal pictures not to be detected as the training set A;

[0022] Select pictures with real abnormalities as test pictures and pictures without abnormalities as background pictures to generate a real abnormal data set as the training set B;

[0023] Cut out the abnormal part from the pictures with real abnormalities and describe the abnormalities in words, and generate real abnormal graphic and text description pairs based on the cut-out abnormal part and the abnormal description as the training set C.

[0024] A further optimization solution is that the end-to-end abnormal detection network is trained in stages with the training sets A, B, and C respectively; the method includes:

[0025] Train the positive and negative sample fusion network part of the end-to-end abnormal detection network with the training set A, and perform box prediction and abnormal confidence prediction;

[0026] Train the abnormal prompt text encoding network part of the end-to-end abnormal detection network with the training set C;

[0027] Train the entire end-to-end abnormal detection network with the training set B, and perform box prediction, abnormal confidence prediction, and category feature vector prediction.

[0028] A further optimization solution is that the predicted box loss is L CIou + DFL, where:

[0029] ;

[0030] ;

[0031] Where IoU is the intersection over union of the calibration box and the detection box, b and b gt are the center points of the detection box and the calibration box respectively, p 2 (*, #) is the square of the Euclidean distance between point * and point #, c is the diagonal distance of the closed area of the detection box and the calibration box, ν is to measure the relative proportional consistency of the detection box and the calibration box, ais the weight coefficient; S i 、S i+1 are the predicted value and adjacent predicted value output by the network, y, y i 、y i+1 are the actual value of the label, label integral value, and adjacent label integral value; L CIou represents the loss of the first prediction box; DFL() represents the loss of the second prediction box;

[0032] During the abnormal confidence prediction process, the abnormal confidence loss is:

[0033] ;

[0034] where y i is the sample label, p i is the predicted probability; N is the total number of detection boxes; i is the i-th detection box; L i is the abnormal confidence loss of the i-th detection box.

[0035] A further optimization solution is that during the process of training the abnormal prompt text encoding network of the end-to-end abnormal detection network with the training set B, the loss function is:

[0036]

[0037] where A and B are the image feature vector and text feature vector; n is the length of the feature vector; θ is the angle between the two vectors; i is the number of the i-th feature; A i is the i-th image feature; B i is the i-th text feature.

[0038] A further optimization solution is that the part of the abnormal prompt text encoding network for training the end-to-end abnormal detection network with the training set C includes the method:

[0039] Adjust the contrastive language-image pre-training model corresponding to the abnormal prompt text encoding network based on the training set C to align the image features and text features of the abnormal images and abnormal description texts in the training set C.

[0040] This solution also provides an end-to-end abnormal detection system based on positive and negative sample fusion for implementing the above-mentioned end-to-end abnormal detection method based on positive and negative sample fusion. The system includes:

[0041] A network construction module for constructing an end-to-end abnormal detection network based on positive and negative sample fusion;

[0042] The dataset construction module is used to construct an abnormal dataset, a real abnormal dataset, and real abnormal image-text description pairs, and correspondingly construct training set A, training set B, and training set C; both training set A and training set B include a test map set and a bottom map set;

[0043] The training module is used to train the end-to-end anomaly detection network in stages with training set A, training set B, and training set C respectively;

[0044] The detection module is used to perform anomaly detection based on the trained end-to-end anomaly detection network.

[0045] This solution also provides a computer-readable medium, on which a computer program is stored. It is characterized in that the computer program, when executed by a processor, can implement the above-mentioned end-to-end anomaly detection method based on positive and negative sample fusion.

[0046] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0047] 1. An end-to-end anomaly detection method, system, and medium based on positive and negative sample fusion provided by the present invention; by fusing positive and negative samples at the deep semantic feature level of images, making full use of the respective advantages of positive and negative samples, and improving the accuracy and recall rate of anomaly detection;

[0048] 2. An end-to-end anomaly detection method, system, and medium based on positive and negative sample fusion provided by the present invention constructs an end-to-end anomaly detection network based on an improved YoloV8 model. By introducing 3D convolution to process multiple bottom maps and fusing them with the features of the test maps extracted by 2D convolution, the recognition ability of the end-to-end anomaly detection network for anomaly features is enhanced;

[0049] 3. An end-to-end anomaly detection method, system, and medium based on positive and negative sample fusion provided by the present invention ensure that the model can effectively learn the features of anomalies and reduce false alarms by generating a dataset containing anomalies that really need to be detected and anomalies that do not need to be detected, as well as a real abnormal dataset and image-text description pairs;

[0050] 4. An end-to-end anomaly detection method, system, and medium based on positive and negative sample fusion provided by the present invention, based on a phased training strategy, first uses the generated abnormal dataset to train the model to reduce false alarms, then uses the real abnormal image-text description pairs for fine-tuning to achieve image-text feature alignment, and finally uses the real abnormal dataset for fine-tuning to accurately locate and discriminate anomalies; it can maintain both high accuracy and high recall rate, effectively solving the problem that it is difficult to balance accuracy and recall rate in traditional methods. Description of the Drawings

[0051] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings. In the drawings:

[0052] Figure 1 It is a schematic flowchart of an end-to-end anomaly detection method based on positive and negative sample fusion;

[0053] Figure 2 It is a schematic diagram of the positive and negative sample fusion network structure;

[0054] Figure 3 It is a schematic diagram of the anomaly prompt text encoding network structure;

[0055] Figure 4 It is a schematic diagram of the end-to-end anomaly detection system structure based on positive and negative sample fusion. Detailed implementation manners

[0056] To make the purpose, technical solutions, and advantages of the present invention clearer, the following will further elaborate on the present invention in combination with the embodiments and drawings. The illustrative embodiments of the present invention and their descriptions are only used to explain the present invention and do not limit the present invention.

[0057] During the current substation inspection process, the problem of insufficient fusion of positive and negative samples for anomaly detection leads to low detection accuracy. In view of this, the following embodiments are provided to solve this technical problem:

[0058] Embodiment 1: This embodiment provides an end-to-end anomaly detection method based on positive and negative sample fusion, as Figure 1 shown, including:

[0059] Step 1, construct an end-to-end anomaly detection network based on positive and negative sample fusion; this step specifically includes the method:

[0060] S11, use the Text Encoder branch in the multi-modal CLIP network as the anomaly prompt text encoding network (the anomaly prompt text encoding network structure is as Figure 3 shown), and use the improved yolov8 network as the positive and negative sample fusion network, and the positive and negative sample fusion network structure is as Figure 2 shown;

[0061] The improved yolov8 network includes branch A, branch B, and a feature fusion network;

[0062] The A branch extracts features from a test image based on the traditional YOLOv8 network; the B branch extracts features from multiple base images based on 3D convolution and 2D convolution;

[0063] The feature fusion network is used to fuse the features extracted by the A branch and the features extracted by the B branch.

[0064] In this step, 2D convolution is used to extract the features of the second feature layer of the test image, and 3D convolution is used to extract the features of the first feature layer of the test image, where the first feature layer is smaller than the second feature layer.

[0065] S12, Configure a detection head including a predicted class feature vector, a predicted abnormal location box, and an abnormal confidence level;

[0066] S13, Compare based on the class feature vector and the text feature vector output by the abnormal prompt text encoding network, and determine the abnormal class according to the maximum cosine distance.

[0067] Step 2, Construct a generated abnormal data set, a real abnormal data set, and a real abnormal graphic description pair, and correspondingly construct training set A, training set B, and training set C; both training set A and training set B include a test image set and a base image set; the specific methods of this step include:

[0068] S21, Select N images at the same point at different times, generate abnormal images that need to be detected and abnormal images that do not need to be detected, and construct a generated abnormal data set from the abnormal images that need to be detected and the abnormal images that do not need to be detected as training set A;

[0069] Specifically, select N images at the same point at different times, where one is used as the test image and N - 1 are used as the base images; randomly generate an abnormality on the test image as the abnormal image that needs to be detected, and mark the location of the abnormality with an outer bounding rectangle; randomly generate the same abnormality at the same position on the detection image and the base image as the abnormal image that does not need to be detected; construct a generated abnormal data set from the abnormal images that need to be detected and the abnormal images that do not need to be detected as training set A, and training set A contains a test image set and a base image set;

[0070] S22, Select images with real abnormalities as the test images, and images without abnormalities as the base images, and generate a real abnormal data set as training set B;

[0071] Specifically, select one image with a real abnormality as the test image; then select N - 1 images without abnormalities at this point as the base images, and mark the location of the abnormality with an outer bounding rectangle; generate a real abnormal data set from the test image and the base images as training set B, and training set B contains a test image set and a base image set.

[0072] S23. Crop out the abnormal part from the picture with real anomalies, describe the anomalies in words, and generate a pair of real anomaly picture and text descriptions based on the cropped abnormal part and the anomaly description, which is used as training set C.

[0073] Step 3. Train the end-to-end anomaly detection network in stages with training set A, training set B, and training set C respectively. This step specifically includes the following methods:

[0074] S31. Train the positive and negative sample fusion network part of the end-to-end anomaly detection network with training set A, and perform bounding box prediction and anomaly confidence prediction, enabling the end-to-end anomaly detection network to have the ability to reduce false alarms caused by random interference. It specifically includes the following steps:

[0075] S311. Train the positive and negative sample fusion network part of the end-to-end anomaly detection network with training set A.

[0076] S312. Take the test images in training set A as the single image to be detected part, and send them to the 2D convolutional part for feature extraction.

[0077] S313. Take the multiple base images corresponding to the current test image in training set A as the multiple base images part, and send them to the 3D convolutional part for feature extraction.

[0078] S314. After the base image features extracted by 3D convolution are transformed from 3D convolution to 2D convolution, they are fused with the features extracted by 2D convolution of the test image, so that the deep features of the positive samples are integrated into the 2D features of the detection image.

[0079] S315. After the finally fused features pass through 2D convolution, perform bounding box prediction and anomaly confidence prediction. In this step of training, specific category feature vector prediction is not performed.

[0080] The purpose of this round of training is to enable the model to have the following ability: when a certain area of the test image is different from all corresponding areas of the corresponding base images, it is determined that the anomaly is a real anomaly; otherwise, it is not an anomaly. Such an ability can reduce false alarms caused by random interference.

[0081] S32. Train the anomaly prompt text encoding network part of the end-to-end anomaly detection network with training set C. This step specifically includes the following method: adjust the contrastive language-image pre-training model corresponding to the anomaly prompt text encoding network based on training set C to align the graphic and text features of the abnormal images and abnormal description texts in training set C.

[0082] S33. Train the entire end-to-end anomaly detection network with training set B, and perform bounding box prediction, anomaly confidence prediction, and category feature vector prediction.

[0083] S331. Incorporate the entire end-to-end anomaly detection network as a pre-trained model into this model training, freeze the clip model, and do not participate in parameter adjustment;

[0084] S332. Use the test images in training set B as the single image to be detected, and send them to the 2D convolutional part for feature extraction;

[0085] S333. Use the multiple base images corresponding to the test images in training dataset B as the multiple base images part, and send them into the 3D convolutional part for feature extraction;

[0086] S334. After the base image features extracted by the 3D convolution are transformed from 3D convolution to 2D convolution, fuse them with the features extracted from the test images by the 2D convolution, so that the deep features of the positive samples are incorporated into the 2D features of the detection images;

[0087] S335. After the finally fused features pass through the 2D convolution, perform bounding box prediction, anomaly confidence prediction, and class feature vector prediction;

[0088] S336. The loss function calculates the loss of the predicted bounding box and anomaly confidence, as well as the cosine loss between the class feature vector and the text feature vector output by the Text Encoder branch (anomaly prompt text encoding network) of the adjusted clip model for the anomaly text description.

[0089] The loss of the predicted bounding box is L CIou + DFL, where:

[0090]

[0091] ;

[0092] Among them, IoU is the intersection over union of the calibrated bounding box and the detected bounding box, b and b gt are the center points of the detected bounding box and the calibrated bounding box respectively, p 2 (*, #) is the square of the Euclidean distance between point * and point #, c is the diagonal distance of the closed area of the detected bounding box and the calibrated bounding box, ν is to measure the relative proportion consistency of the detected bounding box and the calibrated bounding box, a is the weight coefficient; S i 、S i+1 are the predicted values and adjacent predicted values output by the network, y, y i 、y i+1For the actual value of the label, the label integral value, and the adjacent label integral value; L CIou Represents the loss of the first prediction box; DFL() represents the loss of the second prediction box;

[0093] During the abnormal confidence prediction process, the abnormal confidence loss is:

[0094]

[0095] Where y i Is the sample label, p i Is the predicted probability; N is the total number of detection boxes; i is the i-th detection box; L i Is the abnormal confidence loss of the i-th detection box;

[0096] During the process of training the entire end-to-end abnormal detection network with the training set B, the loss function is:

[0097]

[0098] Where A and B are the image feature vector and the text feature vector; n is the length of the feature vector; θ is the angle between the two vectors; i is the number of the i-th feature; A i Is the i-th image feature; B i Is the i-th text feature.

[0099] The purpose of this round of training is to finely tune a network that can accurately locate anomalies, accurately distinguish anomalies, and accurately give the anomaly category based on the above steps with a relatively small actual abnormal dataset of substations (the data volume is compared with the randomly generated abnormal dataset), and this network has a very high anomaly recall and a very high accuracy.

[0100] Step 4, perform anomaly detection based on the trained end-to-end anomaly detection network. The specific process includes:

[0101] Feature extraction is performed on the test image using the trained model. The test image is input into the 2D convolution part of the model, and the features of the test image are extracted through 2D convolution. Feature extraction is performed on the base image using the trained model. The base image is input into the 3D convolution part of the model, and the features of the base image are extracted through 3D convolution. The features of the base image and the features of the test image are fused. The fused features are used for bounding box prediction, anomaly confidence prediction, and class feature vector prediction. The fused features are passed through 2D convolution again for feature fusion. The fused features are respectively fed into the bounding box prediction head, anomaly confidence prediction head, and class feature vector prediction head for prediction. By comparing the predicted class feature vector with the text feature vector output by the Text Encoder branch in CLIP, the anomaly class is determined. The predicted class feature vector is compared with the text feature vector output by the Text Encoder branch in CLIP. The cosine distance between the two is calculated. Finally, the anomaly class corresponding to the class feature vector with the largest cosine distance is selected as the final anomaly class.

[0102] Embodiment 2: This embodiment provides an end-to-end anomaly detection system based on positive and negative sample fusion, as Figure 4 shown, for implementing the end-to-end anomaly detection method based on positive and negative sample fusion described in Embodiment 1. The system includes:

[0103] A network construction module for constructing an end-to-end anomaly detection network based on positive and negative sample fusion;

[0104] A dataset construction module for constructing and generating an anomaly dataset, a real anomaly dataset, and real anomaly text-image description pairs, and correspondingly constructing a training set A, a training set B, and a training set C. Both the training set A and the training set B include a test image set and a base image set;

[0105] A training module for training the end-to-end anomaly detection network in stages with the training set A, the training set B, and the training set C respectively;

[0106] A detection module for performing anomaly detection based on the trained end-to-end anomaly detection network.

[0107] Embodiment 3: This embodiment provides a computer-readable medium with a computer program stored thereon. The computer program, when executed by a processor, can implement the end-to-end anomaly detection method based on positive and negative sample fusion described in Embodiment 1. Specifically, the following steps are performed:

[0108] Step 1, construct an end-to-end anomaly detection network based on positive and negative sample fusion;

[0109] Step 2: Construct an abnormal dataset, a real abnormal dataset, and real abnormal image-text description pairs, and correspondingly construct training set A, training set B, and training set C; both training set A and training set B include a test image set and a background image set;

[0110] Step 3: Train the end-to-end anomaly detection network in stages using training set A, training set B, and training set C respectively;

[0111] Step 4: Perform anomaly detection based on the trained end-to-end anomaly detection network.

[0112] The above specific implementation manners further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are only the specific implementation manners of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An end-to-end anomaly detection method based on positive and negative sample fusion, characterized in that: include: Construct an end-to-end anomaly detection network based on positive and negative sample fusion; Included methods: The Text Encoder branch in the contrastive language-image pre-training model is used as the abnormal prompt text encoding network, and the improved yolov8 network is used as the positive and negative sample fusion network; the improved yolov8 network includes an A branch, a B branch and a feature fusion network; The A branch extracts features from a test image based on the traditional yolov8 network; the B branch extracts features from multiple base images based on 3D convolution and 2D convolution; The feature fusion network is used to fuse the features extracted by branch A and the features extracted by branch B; The configuration includes a detection head that predicts the category feature vector, the predicted anomaly location box, and the anomaly confidence; Compare the category feature vector with the text feature vector output by the abnormal prompt text encoding network, and determine the abnormal category according to the maximum value of the cosine distance; Constructing an abnormal data set, a real abnormal data set and a real abnormal image-text description pair, and correspondingly constructing a training set A, a training set B and a training set C; the training set A and the training set B both include a test image set and a base image set; The end-to-end anomaly detection network is trained in stages using training set A, training set B, and training set C respectively; specifically, the method includes: Use training set A to train the positive and negative sample fusion network part of the end-to-end anomaly detection network, and perform box prediction and anomaly confidence prediction; The abnormal prompt text encoding network part of the end-to-end anomaly detection network is trained with the training set C; Train the entire end-to-end anomaly detection network with training set B, and perform box prediction, anomaly confidence prediction, and category feature vector prediction; Anomaly detection is performed based on the trained end-to-end anomaly detection network.

2. According to claim 1, the end-to-end anomaly detection method based on positive and negative sample fusion is characterized in that: The construction generates an abnormal data set, a real abnormal data set and a real abnormal graphic description pair, and correspondingly constructs a training set A, a training set B and a training set C; including the method: Select N pictures of the same point at different times, generate abnormal images that need to be detected and abnormal images that do not need to be detected, and construct an abnormal data set from the abnormal images that need to be detected and the abnormal images that do not need to be detected as training set A; Select pictures with real anomalies as test pictures and pictures without anomalies as base pictures to generate a real anomaly dataset as training set B; The abnormal part is cut out from the picture with real abnormality, and the abnormality is described in text. Based on the cut-out abnormal part and the abnormal description, a real abnormal image and text description pair is generated as the training set C.

3. The end-to-end anomaly detection method based on positive and negative sample fusion according to claim 1, characterized in that: In the box prediction process, the prediction box loss is L CIou + DFL, where: ; ; in, IoU is the intersection-over-union ratio of the calibration frame and the detection frame, b and b gt are the center points of the detection frame and the calibration frame respectively, p 2 (*, #) is the square of the Euclidean distance between point * and point #, c is the diagonal distance between the detection frame and the closed area of ​​the calibration frame, ν To measure the relative proportion consistency between the detection frame and the calibration frame, a is the weight coefficient; S i 、S i+1 is the predicted value and the near-prediction value output by the network, y,y i 、y i+1 is the actual value of the label, the integral value of the label, and the integral value of the adjacent label; L CIou represents the first prediction box loss; DFL() represents the second prediction box loss; In the anomaly confidence prediction process, the anomaly confidence loss is: ; in y i is the sample label, p i is the prediction probability; N represents the total number of detection boxes; i represents the i-th detection box; L i is the abnormal confidence loss of the i-th detection box.

4. The end-to-end anomaly detection method based on positive and negative sample fusion according to claim 1, characterized in that: In the process of training the entire end-to-end anomaly detection network with training set B, the loss function is: ; Where A and B are image feature vectors and text feature vectors; n is the length of the feature vector; θ is the angle between two vectors; i is the number of the i-th feature; A i is the i-th image feature; B i is the i-th text feature.

5. The end-to-end anomaly detection method based on positive and negative sample fusion according to claim 1, characterized in that: The method of training the abnormal prompt text encoding network part of the end-to-end abnormal detection network with the training set C includes: Based on the training set C, the contrastive language-image pre-training model corresponding to the abnormal prompt text encoding network is adjusted to align the image and text features of the abnormal images and abnormal description texts in the training set C.

6. An end-to-end anomaly detection system based on positive and negative sample fusion, characterized in that: For implementing the end-to-end anomaly detection method based on positive and negative sample fusion according to any one of claims 1 to 5, the system comprises: Network construction module, used to build an end-to-end anomaly detection network based on positive and negative sample fusion; A data set construction module is used to construct a generated anomaly data set, a real anomaly data set, and a real anomaly image-text description pair, and correspondingly construct a training set A, a training set B, and a training set C; the training set A and the training set B both include a test image set and a base image set; A training module, used for training the end-to-end anomaly detection network in stages using training set A, training set B and training set C respectively; The detection module is used to perform anomaly detection based on the trained end-to-end anomaly detection network.

7. A computer readable medium having a computer program stored thereon, characterized in that: The computer program is executed by a processor to implement an end-to-end anomaly detection method based on positive and negative sample fusion as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Septic tank area abnormal behavior detection method based on convolutional neural network

    CN117351574A