Target detection model training method, anti-external damage monitoring and early warning method and equipment

In the power transmission line anti-outbreak monitoring, the Internimage visual big model, adapter model and auxiliary head model are adopted, combined with multi-image fusion data enhancement and multi-head hybrid collaboration training methods, the problem of low recognition accuracy of lightweight object detection models is solved, and high-precision external breaking object detection is achieved.

CN118365983BActive Publication Date: 2025-05-16SANLI VIDEO FREQUENCY SCI & TECH SHENZHEN
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410452140.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-30
Publication Date
2025-05-16
Estimated Expiration
2044-01-30

AI Technical Summary

Technical Problem

In the prior art, in the prevention of external breakage monitoring of transmission lines, the recognition accuracy of the lightweight target detection model is low, making it difficult to meet the requirements of high-precision monitoring.

Method used

The Internimage visual big model is used as the backbone network, combining the adapter model and the auxiliary head model, and the detection accuracy of the object detection model is improved through multi-image fusion data augmentation and multi-head hybrid collaboration training methods.

Benefits of technology

It improves the detection accuracy of the targets that break outside, reduces the possibility that the targets are affected by breaking outside, avoids the problem of catastrophic forgetting, and effectively solves the problems of small target detection and long-tail distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118365983B_ABST
    Figure CN118365983B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for training a target detection model, a method and equipment for monitoring and early warning against external damage, and the method comprises: obtaining a sample data set; constructing a target detection model, including a detection network model and an auxiliary head model, wherein the detection network model comprises a backbone network, a feature enhancement network and a detection head, the backbone network comprises a large visual model and an adapter model, and the auxiliary head model is parallel to the detection head; training the target detection model according to the sample data set and a preset target loss function, wherein the target loss function is constructed according to the original loss function corresponding to the detection network model and the auxiliary loss function corresponding to the auxiliary head model; obtaining a real-time monitoring image, and performing target detection through the trained detection network model to obtain a detection frame of a target to be detected; and performing monitoring and early warning against external damage according to the detection frame of the target to be detected and a preset protection target area. The present invention can effectively reduce the possibility of the protected target being affected by external damage.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This case is a divisional application based on the invention patent with application date of January 30, 2024, application number 202410122175.3, and name “Training method of target detection model, anti-external damage monitoring and early warning method and device” as the parent case. Technical Field

[0002] The present invention relates to the field of target detection technology, and in particular to a target detection model training method, an anti-external damage monitoring and early warning method and device. Background Art

[0003] Transmission lines are important carriers for power transmission tasks, focusing on reliably transmitting large amounts of electricity over long distances. Due to the large land area of ​​my country, there are many transmission systems under construction, and most of the transmission lines are built in mountainous areas. This situation has led to external destructive factors such as construction equipment (tower cranes, excavators, etc.), foreign objects on wires (plastic bags, kites, etc.), and wildfires becoming the main causes of transmission line failures, requiring real-time warning of external failures. However, manual inspection of transmission lines in complex areas is very time-consuming and labor-intensive, so high-precision automated monitoring through deep learning methods has become an urgent need.

[0004] Most of the existing methods for preventing power transmission lines from being damaged externally adopt lightweight target detection models with fewer parameters. Although they can be easily deployed in the ISP (Image Signal Processing) chip of the camera, the above models have low accuracy in identifying damaged targets due to the diverse types, small sizes and large differences in data volume, making it difficult to meet the requirements of accurate monitoring.

[0005] In recent years, the backbone networks of computer vision basic large models such as ViT, Swin Transformer, Internimage, etc. have developed rapidly. These networks have a large number of parameters and have been pre-trained with big data self-supervision. They have very good semantic understanding capabilities and have achieved very good results on various general data sets. However, if the scale of the industrial external damage data set is small, or there is a large distribution difference between the external damage data and the pre-training data of the large model, training the large model through simple direct fine-tuning may cause deviations and catastrophic forgetting problems. In addition, compared with the general data sets used in the pre-training of the large model, the industrial external damage data has a wide monitoring perspective, and many of the external damage source objects are small, which makes it difficult to detect small targets. At the same time, since construction machinery is relatively common among external damage objects, while wildfires and wire foreign objects are very rare, the training data has a long-tail distribution problem. Summary of the invention

[0006] The technical problem to be solved by the present invention is: a training method for a target detection model, an anti-external damage monitoring and early warning method and device, which can effectively reduce the possibility of the protected target being affected by external damage.

[0007] In a first aspect, the present invention provides a method for training a target detection model, comprising:

[0008] Acquire a sample data set, wherein the sample data set includes a preset number of sample scene images and a category label and a detection label corresponding to each sample scene image, wherein the detection label includes the coordinates of a labeling box of each target to be detected in the sample scene image;

[0009] Constructing a target detection model, wherein the target detection model includes a detection network model and an auxiliary head model, wherein the detection network model includes a backbone network, a feature enhancement network and a detection head, wherein the backbone network includes a large visual model and an adapter model, wherein the adapter model is parallel to the stem layer, the stage1 layer and the stage2 layer in the large visual model, and the auxiliary head model is parallel to the detection head in the detection network model;

[0010] According to the sample data set and the preset target loss function, the detection network model in the target detection model is trained, wherein the target loss function is constructed according to the original loss function corresponding to the detection network model and the auxiliary loss function corresponding to the auxiliary head model.

[0011] In a second aspect, the present invention also provides an anti-external damage monitoring and early warning method, comprising:

[0012] Acquire a real-time monitoring image, and perform target detection through a detection network model to obtain a detection frame of the target to be detected, wherein the detection network model is a detection network model in the target detection model trained by the training method provided in the first aspect, and the target to be detected is an external target;

[0013] According to the target detection frame and the preset protection target area, anti-external damage monitoring and early warning are performed.

[0014] In a third aspect, the present invention further provides an electronic device, the electronic device comprising:

[0015] one or more processors;

[0016] A storage device for storing one or more programs;

[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the target detection model training method provided in the first aspect or the anti-external damage monitoring and early warning method provided in the second aspect.

[0018] In a fourth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the target detection model training method provided in the first aspect or the anti-external damage monitoring and early warning method provided in the second aspect.

[0019] The beneficial effects of the present invention are as follows: by adopting the large visual model Internimage as the backbone network to extract features, the detection accuracy of the external target is improved by leveraging the strong semantic understanding ability of the large model; by setting an adapter in the backbone network and running it in parallel with the stem layer, stage1 layer and stage2 layer in the Internimage model, during the training process, the parameters of the stem layer, stage1 layer and stage2 layer in the Internimage model are not updated, and the model parameters after pre-training are always maintained, thereby reducing computing resources and allowing the large model to retain the original basic knowledge while fine-tuning the training, thereby preventing catastrophic forgetting problems; by setting an auxiliary head model to assist the training of the detection head, the training effectiveness of the detection head can be improved, thereby reducing the difficulty of training and reducing the probability of catastrophic forgetting problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A flowchart of a method for training a target detection model provided by the present invention;

[0021] Figure 2 A flow chart of an anti-external damage monitoring and early warning method provided by the present invention;

[0022] Figure 3 This is a flow chart of a method according to Embodiment 1 of the present invention;

[0023] Figure 4 A schematic diagram of a target detection model according to a first embodiment of the present invention;

[0024] Figure 5 A schematic diagram of a backbone network according to a first embodiment of the present invention;

[0025] Figure 6 It is a schematic diagram of multi-image fusion data enhancement according to the first embodiment of the present invention;

[0026] Figure 7 Schematic diagram of an improved Mosaic data enhancement method according to Embodiment 1 of the present invention;

[0027] Figure 8 Schematic diagram of sliding window detection according to the first embodiment of the present invention;

[0028] Fig. 9 Schematic diagram of calculation of overlap rate between the outer breach target detection frame and the protected target area according to the first embodiment of the present invention;

[0029] Fig.10 A schematic diagram of the structure of a training device for a target detection model provided by the present invention;

[0030] Fig.11 A schematic diagram of the structure of an anti-external damage monitoring and early warning device provided by the present invention;

[0031] Fig.12 The present invention provides a schematic structural diagram of an electronic device. DETAILED DESCRIPTION

[0032] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.

[0033] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the steps as sequential processes, many of the steps therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, sub-computer program, etc.

[0034] In addition, the terms "first", "second", etc. can be used in this article to describe various directions, actions, steps or elements, but these directions, actions, steps or elements are not limited by these terms. These terms are only used to distinguish a first direction, action, step or element from another direction, action, step or element. For example, without departing from the scope of the present application, the first information can be referred to as the second information, and similarly, the second information can be referred to as the first information. Both the first information and the second information are information, but they are not the same information. The terms "first", "second", etc. cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Thus, the features defined as "first" and "second" can expressly or implicitly include one or more of the features. In the description of the present invention, the meaning of "multiple" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0035] like Figure 1 As shown, a method for monitoring and early warning against external damage includes:

[0036] Acquire a sample data set, wherein the sample data set includes a preset number of sample scene images and a category label and a detection label corresponding to each sample scene image, wherein the detection label includes the coordinates of a labeling box of each target to be detected in the sample scene image;

[0037] Constructing a target detection model, wherein the target detection model includes a detection network model and an auxiliary head model, wherein the detection network model includes a backbone network, a feature enhancement network and a detection head, wherein the backbone network includes a large visual model and an adapter model, wherein the adapter model is parallel to the stem layer, the stage1 layer and the stage2 layer in the large visual model, and the auxiliary head model is parallel to the detection head in the detection network model;

[0038] According to the sample data set and the preset target loss function, the detection network model in the target detection model is trained, wherein the target loss function is constructed according to the original loss function corresponding to the detection network model and the auxiliary loss function corresponding to the auxiliary head model.

[0039] From the above description, it can be seen that by using the large visual model Internimage as the backbone network to extract features, the detection accuracy of external targets can be improved with the help of the strong semantic understanding ability of the large model; by setting an adapter in the backbone network and running it in parallel with the stem layer, stage1 layer and stage2 layer in the Internimage model, the parameters of the stem layer, stage1 layer and stage2 layer in the Internimage model are not updated during the training process, and the model parameters after pre-training are always maintained, thereby reducing computing resources and allowing the large model to retain the original basic knowledge while fine-tuning the training to prevent catastrophic forgetting problems; by setting an auxiliary head model to assist in the training of the detection head, the training effectiveness of the detection head can be improved, thereby reducing the difficulty of training and the probability of catastrophic forgetting problems.

[0040] In an optional embodiment, the adapter model includes a backbone adaptation module, a first adaptation module, a downsampling module, and a second adaptation module;

[0041] The backbone adaptation module is used to sequentially downsample the image input to the backbone network to obtain feature maps of different dimensions of a preset first number, stack the feature maps of a preset second number therein to obtain a stacked feature map, and output the stacked feature map to the first adapter module;

[0042] The first adapter module is used to use the stacked feature map as a key and a value, and the feature map output by the stem layer in the visual large model as a query, to generate a first adapted feature map, and output the first adapted feature map to the downsampling module; at the same time, the first adapted feature map is superimposed with the first dimensional feature map output by the stage1 layer in the visual large model, and the superimposed first dimensional feature map is output to the first downsampling layer between the stage1 layer and the stage2 layer in the visual large model;

[0043] The downsampling module is used to downsample the first adaptation feature map according to a preset downsampling multiple, and output the downsampled first adaptation feature map to the second adapter module;

[0044] The second adapter module is used to use the downsampled first adapted feature map as the key and value and the feature map output by the first downsampling layer in the visual large model as the query to generate a second adapted feature map; superimpose the second adapted feature map with the second dimensional feature map output by the stage2 layer in the visual large model, and output the superimposed second dimensional feature map to the second downsampling layer between the stage2 layer and the stage3 layer in the visual large model.

[0045] From the above description, we can see that by introducing multi-scale features and adding multi-scale information to the Internimage model in a cross-attention manner, the target detection effect can be better improved; at the same time, only adding adapters to the first two stage layers of the Internimage model can protect the shallow features of the model, that is, the learned general basic texture, color, position, and detail information of the image, and only using the adapter to fine-tune the anti-external destruction scene, while better learning the semantic information of the anti-external destruction scene at the deep features.

[0046] In an optional embodiment, the auxiliary head model includes a first auxiliary head module and a second auxiliary head module, and the auxiliary loss function includes a first auxiliary loss function corresponding to the first auxiliary head module and a second auxiliary loss function corresponding to the second auxiliary head module;

[0047] The objective loss function is L train =L fast-rcnn +L atss +L DETR , where L train Represents the overall loss value of the target detection model, L fast-rcnn represents the first auxiliary loss value calculated by the first auxiliary loss function, L atss represents the second auxiliary loss value calculated by the second auxiliary loss function, L DETR Represents the original loss value calculated by the original loss function.

[0048] From the above description, it can be seen that by adopting a multi-head hybrid assisted training method and setting an auxiliary head to assist the training of the detection head, the training effectiveness of the DETR decoder can be greatly improved, forcing it to have sufficient discrimination to support the training convergence of these auxiliary heads, thereby reducing the difficulty of training and reducing the probability of catastrophic forgetting problems.

[0049] In an optional embodiment, the first auxiliary loss function, the second auxiliary loss function and the original loss function all include a classification loss function and a border regression loss function; the classification loss function in the first auxiliary loss function is a cross entropy loss function, the classification loss function in the second auxiliary loss function is a Focal loss function, and the border regression loss function is a GIoU loss function; the classification loss function in the original loss function is a Quality Focal loss function, and the border regression loss function in the original loss function includes a GIoU loss function and an L1 loss function.

[0050] From the above description, we can see that Focal loss and Quality Focal loss are used in the classification loss of the auxiliary head and the detection head respectively, which can solve the long-tail distribution problem caused by class imbalance. It reduces the loss of easily classified categories and focuses the training on more difficult to classify categories, that is, the tail classes with fewer samples. At the same time, Quality Focal loss combines the quality prediction score and the classification score to prevent inconsistency problems in training and testing.

[0051] In an optional embodiment, in the first preset number of rounds of training, the input data of the target detection model is a fused image after data enhancement is performed on the sample scene images in the sample data set by a multi-image fusion data enhancement method; in the remaining rounds of training, the input data of the target detection model is an image after data enhancement is performed on the sample scene images in the sample data set by a random radial transformation method.

[0052] From the above description, it can be seen that by adding a multi-image fusion data enhancement method on the basis of the basic data enhancement method, the model can acquire more small target and long-tail data knowledge during training while learning the real distribution of external destruction scenes.

[0053] In an optional embodiment, data enhancement is performed on the sample scene images in the sample data set by a multi-image fusion data enhancement method, specifically:

[0054] Four sample scene images are randomly selected from the sample data set, and the four sample scene images are fused by a Mixup data enhancement method and / or a Mosaic data enhancement method to obtain a fused image.

[0055] From the above description, we can see that by constructing a new input image based on Mixup and improved Mosaic, the amount of tail class data is increased, and small targets are enlarged by scaling, thereby solving the long-tail distribution problem and small target detection problem in some application scenarios (such as anti-external damage scenarios).

[0056] In an optional embodiment, the four sample scene images are fused by a Mixup data enhancement method to obtain a fused image, specifically:

[0057] Pairing the four sample scene images in pairs;

[0058] Fusing two sample scene images belonging to the same pair to obtain a fused image, and determining a category label and a detection label corresponding to the fused image according to the category labels and detection labels corresponding to the two sample scene images belonging to the same pair;

[0059] Wherein, the fused image I mix =xI i+j +yI i+k , the category label Y corresponding to the fused image mix =xY i+j +yY i+k , the detection label B corresponding to the fused image mix =concat(B i+j ,B i+k ), I i+j and I i+k Represents two sample scene images belonging to the same pair, Y i+j and B i+j Represents the sample scene image I i+j The corresponding category label and detection label, Y i+k and B i+k Represents the sample scene image I i+k The corresponding category labels and detection labels, x and y represent the preset fusion weights, and concat() represents the connection function.

[0060] In an optional embodiment, the four sample scene images are fused by a Mosaic data enhancement method to obtain a fused image, specifically:

[0061] Determining the cropping regions corresponding to the four sample scene images according to the marked frames of the external targets in the four sample scene images respectively, and cropping the four sample scene images according to the corresponding cropping regions respectively to obtain four cropped images;

[0062] After rotating, scaling and changing the color of the four cropped images respectively, they are spliced ​​in the form of four grids to obtain a fused image;

[0063] Among them, the coordinates of the four corners of the cropped area corresponding to a sample scene image are (min n (b x1 )-z, min n (b y1 )-z, max n (b x2 )+z,max n (b y2 )+z), min n (b x1 ) and min n (b y1 ) are the minimum values ​​of the upper left corner coordinates of the labeled box of each target to be measured in the sample scene image in the X-axis and Y-axis directions, respectively, n (b x2 ) and max n (b y2 ) are respectively the maximum values ​​in the X-axis and Y-axis directions of the lower right corner coordinates of the marked boxes of each external target to be measured in the sample scene image, and z is a preset margin constant.

[0064] From the above description, it can be seen that by adding a cropping step before stitching, the image input to the stitching step includes as much of the target area as possible, so that the finally generated image includes the target as much as possible rather than background noise.

[0065] In an optional embodiment, after determining the cropping areas corresponding to the four sample scene images according to the annotation boxes of the respective targets to be measured in the four sample scene images, the method further includes:

[0066] If the width of the cropped area is less than the preset width threshold and the height is less than the preset height threshold, the cropped area is enlarged, and the coordinates of the four corners of the enlarged cropped area are (min n (b x1 )-zW gap / 2, min n (b y1 )-zH gap / 2, max n (b x2 )+z+W gap / 2, max n (b y2 )+z+H gap / 2), where W gap and H gap The width and height of the cropping area, respectively.

[0067] From the above description, it can be seen that if the cropping area is too small, the cropping area is appropriately expanded to avoid losing too much image information.

[0068] like Figure 2 As shown, the present invention also provides an anti-external damage monitoring and early warning method, comprising:

[0069] Acquire a real-time monitoring image, and perform target detection through a detection network model to obtain a detection frame of the target to be detected, wherein the detection network model is a detection network model in the target detection model trained by the training method described above, and the target to be detected is an external target;

[0070] According to the target detection frame and the preset protection target area, anti-external damage monitoring and early warning are performed.

[0071] From the above description, it can be seen that the detection accuracy of externally damaged targets can be improved, and the possibility of protected targets being affected by external damage can be effectively reduced.

[0072] In an optional embodiment, the real-time monitoring image is acquired, and target detection is performed through a detection network model to obtain a detection frame of the target to be detected, specifically:

[0073] Acquire a real-time monitoring image, and obtain a plurality of image blocks in the real-time monitoring image according to a sliding window of a preset size and a preset step size;

[0074] The trained detection network model is used to perform target detection on the real-time monitoring image and each image block, and the detection results are deduplicated using a non-maximum suppression method to obtain the final detection result.

[0075] From the above description, it can be seen that the small targets hidden in the original image are enlarged through the window, so that the small targets can be effectively detected.

[0076] In an optional embodiment, the anti-external damage monitoring and early warning is performed according to the target detection frame to be detected and the preset protection target area, specifically:

[0077] According to the maximum overlap rate calculation formula, the maximum overlap rate between each target detection frame and the preset protection target area is calculated. The maximum overlap rate calculation formula is:

[0078]

[0079] Among them, P over represents the maximum overlap rate, M represents the number of target detection frames to be tested, and B di represents the i-th target detection box, B s Indicates the target area for protection;

[0080] If the maximum overlap rate is greater than or equal to a preset warning threshold, a warning signal is sent.

[0081] In an optional embodiment, if the maximum overlap rate is greater than or equal to a preset warning threshold, sending a warning signal is specifically:

[0082] If the maximum overlap rate is greater than or equal to a preset first warning threshold and less than a preset second warning threshold, a secondary warning signal is sent;

[0083] If the maximum overlap rate is greater than or equal to a preset second warning threshold, a first-level warning signal is sent.

[0084] From the above description, it can be seen that the hidden danger of external damage can be effectively prevented.

[0085] In an optional embodiment, after performing anti-external damage monitoring and early warning according to the target detection frame to be detected and the preset protection target area, the method further includes:

[0086] According to a preset iterative training cycle, the real-time monitoring image in which the detection frame of the target to be tested is detected in the current cycle is used as a sample scene image and added to the sample data set;

[0087] According to the sample data set, the detection network model in the target detection model is iteratively trained.

[0088] From the above description, we can see that by continuously increasing the amount of data in the sample data set, we can gradually alleviate the catastrophic forgetting problem of large models caused by too little training data, making training gradually simpler; at the same time, through continuous iterative training, the detection accuracy of the model will continue to increase over time, thereby enhancing the detection performance of the model.

[0089] In an optional embodiment, during the iterative training process, a preset proportion of sample scene images are randomly selected from the sample data set, and multiple image blocks are obtained in the selected sample scene images according to a sliding window of a preset size and a preset step size.

[0090] From the above description, it can be seen that the performance of small target detection can be further improved.

[0091] The present invention further provides an electronic device, comprising:

[0092] one or more processors;

[0093] A storage device for storing one or more programs;

[0094] When the one or more programs are executed by the one or more processors, the one or more processors implement the target detection model training method or the anti-external damage monitoring and early warning method as described above.

[0095] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the training method of the target detection model or the anti-external damage monitoring and early warning method as described above is implemented.

[0096] Embodiment 1

[0097] Please refer to Figure 3-9 , Embodiment 1 of the present invention is: a method for training a target detection model and applying the target detection model to perform an anti-external damage monitoring and early warning method, which can effectively reduce the possibility of the protected target being affected by external damage. In this embodiment, the monitoring of external damage objects around the transmission line is used as an example for explanation. This embodiment aims to utilize the strong semantic understanding ability of the large visual model to improve the accuracy of external damage detection, while solving the problem of small target detection difficulty in the anti-external damage scene, and the problem of poor tail category detection performance caused by the long-tail distribution of data, and solve the catastrophic forgetting problem in large model training.

[0098] like Figure 3 As shown, the following steps are included:

[0099] S1: Obtain the external damage dataset, that is, build a large-scale external damage prevention dataset by collecting and preprocessing the monitoring videos of the transmission line site and uploading them to the cloud in real time.

[0100] In this embodiment, the external damage dataset is a collection of scene images containing 12 types of objects that may cause external damage to the transmission line (cranes, tower cranes, pump trucks, forklifts, bulldozers, excavators, dump trucks, pile drivers, foreign objects in wires, dust nets, wildfires, and smoke). All images are acquired by surveillance cameras at fixed points in multiple different transmission line scenes, including images at different times and under different lighting conditions to meet data diversity.

[0101] For example, we collected videos in 75 different scenes, including scenes of external breaches in multiple provinces collected from the monitoring platform and scenes of external breaches disclosed on some networks. After data cleaning preprocessing operations such as framing, resizing, and deduplication of the monitoring videos, a total of 143,466 images of external breaches with 500 resolutions (500×375 to 5280×3000) were collected. After careful annotation by professionals, each image was labeled with i Get the corresponding category label Y i and detection tag B i =(b n x1 , b n y1, b n x2 , b n y2 ), where n = 1, 2, ..., N i , N i For the i-th picture I i The number of annotation boxes of the external broken targets in (b n x1 , b n y1 ) and (b n x2 , b n y2 ) represent the coordinates of the upper left corner and lower right corner of the nth annotation box respectively.

[0102] S2: Build an object detection model.

[0103] like Figure 4 As shown, the target detection model includes a detection network model and an auxiliary head model. The detection network model includes a backbone network (Backbone), a feature enhancement network (Neck) and a detection head (Head). The auxiliary head model (Auxiliaryhead) is parallel to the detection head in the detection network model.

[0104] In this embodiment, the detection network model is built based on the Internimage visual large model. The Internimage-large version model is used as the backbone network (Backbone) for feature extraction, the FPN feature pyramid is used as the feature enhancement network (Neck) for feature multi-scale attention enhancement, and DETR (Detection Transformer) is used as the detection head (Head) to generate external target detection results.

[0105] According to the theory proposed in the Co-DETR work, although DETR is a large-parameter target detection head designed based on the Transformer encoder-decoder structure, it regards the target detection task as a one-to-one set prediction problem, but because it has too few positive sample queries (queries), its training difficulty is very high and its semantic understanding ability is poor. In other words, directly training the DETR detection head in an external scene is prone to sparse supervision of the encoder caused by fewer positive queries in the decoder, causing catastrophic forgetting or poor training of the large model. Therefore, in order to alleviate the sparse supervision of the encoder caused by fewer positive queries in the decoder, this embodiment adopts a multi-head hybrid collaborative training method, that is, the commonly used one-to-many method detection head is used as an auxiliary head to assist the training of the detection head DETR. In this embodiment, the auxiliary head is only used to optimize the training of the detection head DETR, and is not used in the application stage (i.e., when performing target detection on real-time monitoring images).

[0106] Specifically, in this embodiment, the auxiliary head model includes a first auxiliary head module (Faster-RCNN) and a second auxiliary head module (ATSS), that is, Faster-RCNN and ATSS are used as auxiliary heads, and supervision is performed through one-to-many different label assignment forms to help the encoder learn sufficient discrimination to support the convergence of multi-head training.

[0107] Combination Figure 5 As shown in Figure 1, the Internimage model includes a stem layer and a combination of four stage layers and a downsampling layer (downsample). Assuming that the size of the image input to the backbone network is H×W×3, after feature extraction in the initial stem stage and four stage layers in Internimage-large, four multi-scale features of different scales H / 4×W / 4×160, H / 8×W / 8×320, H / 16×W / 16×640 and H / 32×W / 32×1280 are obtained. After feature multi-scale attention enhancement by FPN, five feature pyramids consisting of features with the same number of channels H / 4×W / 4×256, H / 8×W / 8×256, H / 16×W / 16×256, H / 32×W / 32×256 and H / 64×W / 64×256 are obtained, which are then input into DETR and the auxiliary head, and finally the detection results are output.

[0108] Furthermore, since Internimage is a benchmark large model in the field of vision, if all parameters are fine-tuned and trained directly, a large amount of GPU computing resources are required, and it is easy to cause deviations and catastrophic forgetting problems caused by training on downstream tasks. In order to reduce the parameters that need to be trained while achieving better downstream tasks, namely, external target detection, this embodiment designs an adapter training technique, that is, the initial and first two network blocks of the backbone network Internimage-large are frozen, and an adapter model (Adapter) is added in a parallel position. In other words, the parameters in the stem layer, stage1 layer, and stage2 layer in the Internimage model are not updated during subsequent training, and the parameters of these three layers are maintained as the model parameters after pre-training of the large model, thereby avoiding the training deviation and catastrophic forgetting problems that may occur when there is a large distribution difference between the external data set and the pre-training data set of the large model.

[0109] Specifically, Figure 5 As shown, the adapter model (Adapter) includes a stem adapter module (stem-adapter), a first adapter module (adapter1), a downsampling module (downsample) and a second adapter module (adapter2).

[0110] The stem-adapter module consists of a preset first number of stem modules (the same as the stem layer in the Internimage model), a fusion module (concat) and a convolution module (1×1conv), which is used to sequentially downsample the image input to the backbone network to obtain feature maps of different dimensions of the preset first number, and then stack the feature maps of the preset second number to obtain the stacked feature map F adp , and finally the stacked feature map F adp Output to the first adapter module (adapter1).

[0111] In this embodiment, the backbone adaptation module includes four stem modules, namely stem1, stem2, stem3 and stem4. These four stem modules output feature maps that are downsampled by 4 times, 8 times, 16 times and 32 times respectively. Then, the feature maps output by stem2, stem3 and stem4 are stacked, and the size of the stacked feature map is (HW / 8 2 +HW / 16 2 +HW / 32 2)×Dim, Dim represents the image depth. In this embodiment, this value is the same as the depth of the feature map output by the stage1 layer, that is, 160. In other optional embodiments, the feature maps output by stem1, stem2, and stem3 can also be stacked, or the feature maps output by the four stem modules can be stacked.

[0112] The first adapter module (adapter1) and the second adapter module (adapter2) both include a preset number of multi-head cross-attention layers and a feed-forward neural network FFN. In this embodiment, 4 layers of 16-head cross-attention layers are set.

[0113] The first adapter module (adapter1) is used to stack the feature map F adp As the key (Key) and value (Value), the feature map output by the stem layer in the Internimage model is used as the query (Query) to generate the first adaptation feature map F adp1 , and output to the downsampling module (downsample), and at the same time the first adaptation feature map F adp1 It is superimposed with the first-dimensional feature map output by the stage1 layer in the Internimage model, and the superimposed first-dimensional feature map is output to the first downsampling layer between the stage1 layer and the stage2 layer in the Internimage model.

[0114] The downsampling module (downsample) is used to downsample the first adaptation feature map F according to a preset downsampling multiple (2 times in this embodiment). adp1 Downsample and convert the first adapted feature map F after downsampling adp1 Output to the second adapter module (adapter2).

[0115] The second adapter module (adapter2) is used to convert the downsampled first adaptation feature map F adp1 As the key and value, the feature map output by the first downsampling layer in the Internimage model is used as the query to generate the second adaptation feature map F adp2 ; Then the second adaptation feature map F adp2 It is superimposed with the second dimensional feature map output by the stage2 layer in the Internimage model, and the superimposed second dimensional feature map is output to the second downsampling layer between the stage2 layer and the stage3 layer in the Internimage model.

[0116] The subsequent processing flow of the Internimage model is consistent with the original flow.

[0117] By setting up an adapter, the large model can retain the original basic knowledge (freezing the original pre-trained weights) while training and fine-tuning, preventing catastrophic forgetting problems. Unlike some existing adapter solutions, this embodiment introduces multi-scale features in the initial adapter stage (i.e., the backbone adaptation module), and uses cross-attention to add multi-scale information to the backbone feature extraction network (i.e., the Internimage model), thereby better improving the effect of intensive prediction tasks such as target detection. At the same time, this embodiment only adds adapters to the first two stage layers of the Internimage model, with the aim of protecting the shallow features of the model, that is, the learned general basic texture, color, position, and detail information of the image, and only uses the adapter for fine-tuning the anti-external damage scene, while the semantic information of the anti-external damage scene is better learned at the deep features.

[0118] S3: According to the external data set and a preset target loss function, the detection network model in the target detection model is trained.

[0119] In this embodiment, the target detection model is trained for 12 rounds, and the model parameters in the trained target detection model are saved and output and deployed on a cloud server.

[0120] Furthermore, the external destruction scene images in the external destruction dataset are subjected to data enhancement before being input into the target detection model.

[0121] The data enhancement methods adopted by the prior art for training large models are basically a combination of random affine transformations and simple copy and paste methods, etc., which can increase the richness of training samples to a certain extent, but the effect is weak in the face of more challenging small target detection and long-tail distribution scenes. In this embodiment, a multi-image fusion data enhancement method is used to increase the number of tail class data in terms of input data enhancement. However, according to the theory proposed by YOLOX work, although multi-image fusion data enhancement can increase data diversity, the pictures it generates are far from the data distribution of real external broken images. Therefore, this embodiment adopts an early stopping technique, that is, in the first preset number of rounds of training (the first 70% of the rounds in this embodiment), data enhancement is performed by a multi-image fusion data enhancement method, and in the remaining rounds of training (that is, the last 30% of the rounds), basic random affine transformation technology is used for data enhancement. By adding a multi-image fusion data enhancement method on the basis of the basic data enhancement method, the model can acquire more small targets and long-tail data knowledge during training while learning the real distribution of external broken scenes.

[0122] For the multi-image fusion data enhancement method, specifically, Figure 6 As shown, four external destruction scene images are randomly selected from the external destruction dataset, assuming that {Ii , I i+1 , I i+2 , I i+3}, and then the images are fused through the Mixup data enhancement method (method 1) and the improved Mosaic data enhancement method (method 2) to construct a new input image, thereby increasing the amount of tail class data and enlarging small targets by scaling, alleviating the long-tail distribution problem and small target detection problem.

[0123] Among them, the first method is the Mixup data enhancement method, which pairs the four external scene images in pairs. Assume that one of the pairs is (I i+j , I i+k ), j, k∈{0,1,2,3}, and fuse the two to obtain the fused image I mix =xI i+j +yI i+k , the fused image I mix The corresponding category label Y mix =xY i+j +yY i+k , detect label B mix =concat(B i+j ,B i+k ), where x and y represent the preset fusion weights, and concat() represents the connection function used to connect arrays.

[0124] Method 2: Improved Mosaic data enhancement method, such as Figure 7 As shown, first, according to the labeling boxes of the external destruction targets in the four external destruction scene images, the cropping regions corresponding to the four external destruction scene images are determined, and the four external destruction scene images are cropped according to the corresponding cropping regions to obtain four cropped images. i The coordinates of the four corners of the clipping area in are (min n (b x1 )-z, min n (b y1 )-z, max n (b x2 )+z,max n (b y2 )+z), min n (b x1 ) and min n (b y1 ) are respectively the external scene images I i The minimum value of the X-axis and Y-axis directions in the upper left corner coordinates of the annotation box of each external target, max n (b x2 ) and maxn (b y2 ) are respectively the external scene images I i The maximum value of the X-axis and Y-axis coordinates of the lower right corner of the label box of each external target in the image, z is the preset margin constant. The cropped image can contain more external targets and reduce the proportion of the background area.

[0125] Furthermore, if the width of the cropping region W gap =max n (b x2 )-min n (b x1 )+2z is less than the preset width threshold W lim , and the height H gap =max n (b y2 )-min n (b y1 )+2z is less than the preset height threshold H lim , it is considered that the cropping area is too small and too much image information is lost. At this time, the cropping area is appropriately expanded. The coordinates of the four corners of the expanded cropping area are (min n (b x1 )-zW gap / 2, min n (b y1 )-zH gap / 2, max n (b x2 )+z+W gap / 2, max n (b y2 )+z+H gap / 2).

[0126] Then the four cropped images are randomly rotated, scaled, and discolored, and then spliced ​​in the form of a four-square grid to obtain the fused image I mix .

[0127] Since in the surveillance video scene training data, the external broken targets are usually small and not dense, and there is often only one or a few external broken targets in a picture, so directly using the existing Mosaic technology will cause the enhanced image to be mostly background, and the generated image is mostly noise, which will have a negative impact on model training. Compared with the existing Mosaic data enhancement technology, this embodiment adds a cropping step before splicing, so that the image input to the splicing step contains as much target area as possible, so that the final generated image contains the target as much as possible, rather than background noise.

[0128] Furthermore, the target loss function is constructed according to the original loss function corresponding to the detection network model and the auxiliary loss function corresponding to the auxiliary head model, and the auxiliary loss function includes the first auxiliary loss function corresponding to the first auxiliary head module (Faster-RCNN) and the second auxiliary loss function corresponding to the second auxiliary head module (ATSS). In this embodiment, the overall loss value L of the external target detection model is train =L fast-rcnn +L atss +L DETR , where L fast-rcnn represents the first auxiliary loss value calculated by the first auxiliary loss function, L atss represents the second auxiliary loss value calculated by the second auxiliary loss function, L DETR Represents the original loss value calculated by the original loss function.

[0129] The first auxiliary loss function and the second auxiliary loss function both include a classification loss function and a bounding box regression loss function; the classification loss function in the first auxiliary loss function is a cross entropy loss function, the classification loss function in the second auxiliary loss function is a Focal loss function, and the bounding box regression loss function is a GIoU loss function. The specific formula is as follows:

[0130] L fast-rcnn =λ1L CE +λ2L GIoU

[0131] L atss =λ3L focal +λ4L GIoU

[0132]

[0133]

[0134]

[0135] Among them, λ1, λ2, λ3, and λ4 represent preset weights, which are hyperparameters for adjusting the weights of the auxiliary loss function. In this embodiment, λ1=1, λ2=10, λ3=1, and λ4=2. In the cross entropy loss function and the Focal loss function, p is the classification score predicted by the network, y is the actual classification label, and γ is the preset attenuation coefficient used to control the loss weight of difficult samples. In the GIoU loss function, b p 、b g They represent the predicted box and the real label box (i.e., the annotation box), and IoU represents b p and b g The intersection-and-union ratio of p and b gThe area of ​​the intersection, A c Indicates b p and b g The area of ​​the minimum enclosing rectangle of .

[0136] The training of the detection head in the detection network model adopts the training technique proposed by the existing solution DINO, and the DETR series constructs target detection as a bipartite graph matching problem. The network training first uses the Hungarian matching method to match the elements of the set of prediction results and the set of true labels one by one to find the combination with the minimum matching loss. Then, the loss function L is constructed for the matched combination. DETR The original loss function also includes a classification loss function and a bounding box regression loss function. In this embodiment, the classification loss function uses the Quality Focal loss function to alleviate the long-tail distribution problem, and the bounding box regression loss function uses the GIoU loss function and the L1 loss function to further improve the model performance. The specific formula is as follows:

[0137] L DETR =λ5L qfocal +λ6L GIoU +λ7L l1

[0138] L qfocal (σ)=-|y-σ| β ((1-y))log(1-σ)+ylog(σ))

[0139] L l1 (b p , b g )=|b g -b p |

[0140] Among them, λ5, λ6, and λ7 represent preset weights, which are hyperparameters for adjusting the weights of the detection head loss function. In this embodiment, λ5=1, λ6=2, and λ7=5. σ represents the predicted quality-classification joint score, y is the quality-classification joint representation label of 0 to 1, and β is similar to γ ​​in the Focal loss function, which is a preset attenuation coefficient. In the L1 loss function, b p and b g Represent the coordinates of the predicted box and the true label box respectively.

[0141] This embodiment uses Focal loss and QualityFocal loss in the classification loss of the auxiliary head and the detection head respectively, aiming to solve the long-tail distribution problem caused by class imbalance. It reduces the loss of easily classified categories and focuses the training on more difficult-to-classify categories, i.e., the tail categories with fewer samples. At the same time, QualityFocal loss combines the quality prediction score with the classification score to prevent inconsistency problems in training and testing.

[0142] At the same time, in order to further increase the number of positive queries in the detection head DETR, this embodiment encodes the coordinates of the positive sample anchor box obtained by the positive and negative sample allocation method in the auxiliary head as the positive query input of the decoder. At this time, since the positive and negative samples have been allocated by the auxiliary head, the query can be directly matched with the allocated category label and detection label without Hungarian matching. Specifically, for a positive sample anchor box B allocated by an auxiliary head, pos and the corresponding label, which is embedded after passing through the position encoding layer of the Transformer, and then the embedding is passed through a fully connected layer to keep the dimension consistent with the query dimension to obtain Q pos , and then Q pos After inputting the decoder of DETR, the prediction result P is obtained pos , calculate P pos and the corresponding label L DETR That's it.

[0143] By setting the auxiliary heads, the training effectiveness of the DETR decoder can be greatly improved, forcing it to have sufficient discrimination to support the training convergence of these auxiliary heads, thereby alleviating the training difficulty and reducing the probability of catastrophic forgetting problems.

[0144] S4: Obtain real-time monitoring images, and perform target detection through the trained detection network model to obtain the external target detection frame.

[0145] This embodiment adopts the original + sliding window test method. Specifically, Figure 8 As shown, the cloud server regularly collects video frames uploaded by remote monitoring cameras as real-time monitoring images, and win ×H win The sliding window and the preset step size are used to monitor the image I in real time. t Multiple image blocks are obtained, that is, the real-time monitoring image is divided into K overlapping image blocks {Win t1 ,Win t2 ,……,Win tKThen, the trained detection network model is used to perform target detection on the real-time monitoring image and each image block, that is, {I t ,Win t1 ,Win t2 ,……,Win tK} Target detection is performed, and then all test results are deduplicated through the non-maximum suppression method (NMS) to obtain the final detection result. The small targets hidden in the original image are enlarged through the window, so that the small targets can be effectively detected.

[0146] S5: performing anti-external damage monitoring and early warning according to the external damage target detection frame and the preset protection target area.

[0147] Specifically, Fig. 9 As shown, assume that M external target detection frames B are detected d = {B d1 ,B d2 ,……,B dM}, in the real-time monitoring image, the protection target area is obtained by manual box selection or lightweight target detection method, that is, the approximate area frame of the transmission line, and then the detection frame B of each external target is calculated according to the maximum overlap rate calculation formula d With the preset protection target area B s The maximum overlap ratio P over , where the maximum overlap rate calculation formula is:

[0148]

[0149] P over That is, it represents the maximum overlap rate of the M external target detection frames and the protected target area. The early warning system sets two early warning thresholds, namely the first early warning threshold T1 and the second early warning threshold T2. When T1≤P over <T2, send a first-level warning signal (first-level alert); when P over ≥T2, send a second-level warning signal (second-level alert); when P over When <T1, no warning signal is sent, and the next round of detection and judgment is waited for to realize pseudo real-time warning.

[0150] In this embodiment, the cloud server regularly collects the video frames uploaded by the remote monitoring camera at a time interval of 10s, T1 is set to 0.5, and T2 is set to 0.75. Through pseudo-real-time monitoring and early warning, external damage hazards can be effectively prevented, making up for the disadvantage that it cannot be deployed on the actual camera end.

[0151] S6: According to the real-time monitoring image in which the target detection frame is detected to be broken, the detection network model in the target detection model is iteratively trained regularly.

[0152] Specifically, according to the preset iterative training cycle (set to 1 month in this embodiment), the real-time monitoring image in which the external breach target detection frame is detected in the current cycle is used as the external breach scene image and added to the external breach data set. The category label and detection label of the real-time detection image are determined according to the detection result, and can also be determined by manually adjusting the detection result. Then, according to the latest external breach data set, the detection network model in the target detection model is iteratively trained, that is, the detection network model is incrementally learned and the parameters are fine-tuned to obtain the model parameters of the new version, and the latest version of the detection network model is subsequently used to perform target detection on the real-time monitoring image.

[0153] By continuously increasing the amount of data in the external dataset, the catastrophic forgetting problem caused by too little training data in large models can be gradually alleviated, making training gradually simpler; at the same time, through continuous iterative training, the detection accuracy of the model will continue to increase over time, thereby enhancing the detection performance of the model.

[0154] Furthermore, in the iterative training process, a preset proportion (30% in this embodiment) of external destruction scene images is randomly selected from the external destruction data set, and a plurality of image blocks are obtained in the selected external destruction scene images according to a sliding window of a preset size and a preset step size, and the size of each image block is W win ×H win , and then train the model to further improve the performance of small target detection.

[0155] This embodiment builds a network structure based on the Internimage visual large model and DETR, and uses the strong semantic understanding ability of the large model to improve the detection accuracy of external damage targets. By adopting a multi-image fusion data enhancement method, and adopting the original + sliding window test method in the application stage, and introducing multi-scale features in the network architecture design, the small target and long-tail distribution problems in the anti-external damage scenario are effectively solved. By adopting multi-head hybrid collaborative training and adapter training methods during model training, and by regularly expanding the external damage data set and iteratively training the target detection model, catastrophic forgetting problems in large model training can be prevented. The detection accuracy of this embodiment is improved by about 50% compared with the existing solution, which can effectively reduce the possibility of transmission lines being affected by external damage, protect the safety of long-distance power transmission, and reduce monitoring and inspection costs.

[0156] Embodiment 2

[0157] Please refer to Fig.10 , Embodiment 2 of the present invention is: a training device for a target detection model, which can execute the training method for the target detection model described above, and has functional modules and beneficial effects corresponding to the execution method. The device can be implemented by software / or hardware, and specifically includes:

[0158] An acquisition module 201 is used to acquire a sample data set, wherein the sample data set includes a preset number of sample scene images and a category label and a detection label corresponding to each sample scene image, wherein the detection label includes the coordinates of the annotation box of each target to be detected in the sample scene image;

[0159] A construction module 202 is used to construct a target detection model, wherein the target detection model includes a detection network model and an auxiliary head model, wherein the detection network model includes a backbone network, a feature enhancement network and a detection head, wherein the backbone network includes a large visual model and an adapter model, wherein the adapter model is parallel to the stem layer, the stage1 layer and the stage2 layer in the large visual model, and the auxiliary head model is parallel to the detection head in the detection network model;

[0160] The training module 203 is used to train the detection network model in the target detection model according to the sample data set and the preset target loss function, wherein the target loss function is constructed according to the original loss function corresponding to the detection network model and the auxiliary loss function corresponding to the auxiliary head model.

[0161] In an optional embodiment, the adapter model includes a backbone adaptation module, a first adaptation module, a downsampling module, and a second adaptation module;

[0162] The backbone adaptation module is used to sequentially downsample the image input to the backbone network to obtain feature maps of different dimensions of a preset first number, stack the feature maps of a preset second number therein to obtain a stacked feature map, and output the stacked feature map to the first adapter module;

[0163] The first adapter module is used to use the stacked feature map as a key and a value, and the feature map output by the stem layer in the visual large model as a query, to generate a first adapted feature map, and output the first adapted feature map to the downsampling module; at the same time, the first adapted feature map is superimposed with the first dimensional feature map output by the stage1 layer in the visual large model, and the superimposed first dimensional feature map is output to the first downsampling layer between the stage1 layer and the stage2 layer in the visual large model;

[0164] The downsampling module is used to downsample the first adaptation feature map according to a preset downsampling multiple, and output the downsampled first adaptation feature map to the second adapter module;

[0165] The second adapter module is used to use the downsampled first adapted feature map as the key and value and the feature map output by the first downsampling layer in the visual large model as the query to generate a second adapted feature map; superimpose the second adapted feature map with the second dimensional feature map output by the stage2 layer in the visual large model, and output the superimposed second dimensional feature map to the second downsampling layer between the stage2 layer and the stage3 layer in the visual large model.

[0166] In an optional embodiment, the auxiliary head model includes a first auxiliary head module and a second auxiliary head module, and the auxiliary loss function includes a first auxiliary loss function corresponding to the first auxiliary head module and a second auxiliary loss function corresponding to the second auxiliary head module;

[0167] The objective loss function is L train =L fast-rcnn +L atss +L DETR , where L train Represents the overall loss value of the target detection model, L fast-rcnn represents the first auxiliary loss value calculated by the first auxiliary loss function, L atss represents the second auxiliary loss value calculated by the second auxiliary loss function, L DETR Represents the original loss value calculated by the original loss function.

[0168] In an optional embodiment, the first auxiliary loss function, the second auxiliary loss function and the original loss function all include a classification loss function and a border regression loss function; the classification loss function in the first auxiliary loss function is a cross entropy loss function, the classification loss function in the second auxiliary loss function is a Focal loss function, and the border regression loss function is a GIoU loss function; the classification loss function in the original loss function is a Quality Focal loss function, and the border regression loss function in the original loss function includes a GIoU loss function and an L1 loss function.

[0169] In an optional embodiment, in the first preset number of rounds of training, the input data of the target detection model is a fused image after data enhancement is performed on the sample scene images in the sample data set using a multi-image fusion data enhancement method; in the remaining rounds of training, the input data of the target detection model is an image after data enhancement is performed on the sample scene images in the sample data set using a random radial transformation method.

[0170] In an optional implementation, the multi-image fusion data enhancement method includes a Mosaic data enhancement method; data enhancement is performed on the sample scene images in the sample data set by the Mosaic data enhancement method, specifically:

[0171] Randomly select four sample scene images from the sample data set;

[0172] Determining the cropping regions corresponding to the four sample scene images according to the marked frames of the respective targets to be measured in the four sample scene images, and cropping the four sample scene images according to the corresponding cropping regions to obtain four cropped images;

[0173] After rotating, scaling and changing the color of the four cropped images respectively, they are spliced ​​in the form of four grids to obtain a fused image;

[0174] Among them, the coordinates of the four corners of the cropped area corresponding to a sample scene image are (min n (b x1 )-z, min n (b y1 )-z, max n (b x2 )+z,max n (b y2 )+z), min n (b x1 ) and min n (b y1 ) are the minimum values ​​of the upper left corner coordinates of the labeled box of each target to be measured in the sample scene image in the X-axis and Y-axis directions, respectively, n (b x2 ) and max n (b y2 ) are respectively the maximum values ​​in the X-axis and Y-axis directions of the coordinates of the lower right corner of the annotation box of each target to be measured in the sample scene image, and z is a preset margin constant.

[0175] In an optional implementation, after determining the cropping areas corresponding to the four sample scene images according to the annotation boxes of the respective targets to be measured in the four sample scene images, the method further includes:

[0176] If the width of the cropped area is less than the preset width threshold and the height is less than the preset height threshold, the cropped area is enlarged, and the coordinates of the four corners of the enlarged cropped area are (min n (b x1 )-zW gap / 2, min n (b y1 )-zHgap / 2, max n (b x2 )+z+W gap / 2, max n (b y2 )+z+H gap / 2), where W gap and H gap The width and height of the cropping area, respectively.

[0177] Embodiment 3

[0178] Please refer to Fig.11 The third embodiment of the present invention is: an anti-external damage monitoring and early warning device, which can execute the anti-external damage monitoring and early warning method as described above, and has the corresponding functional modules and beneficial effects of the execution method. The device can be implemented by software / or hardware, specifically including:

[0179] The target detection module 204 is used to obtain a real-time monitoring image and perform target detection through a detection network model to obtain an externally broken target detection frame, wherein the detection network model is a detection network model in the target detection model trained by the training method described above, and the target to be detected is an externally broken target;

[0180] The monitoring and early warning module 205 is used to perform anti-external damage monitoring and early warning according to the target detection frame to be detected and the preset protection target area.

[0181] In an optional implementation, the target detection module 204 includes:

[0182] A sliding window unit, used to obtain a real-time monitoring image, and obtain a plurality of image blocks in the real-time monitoring image according to a sliding window of a preset size and a preset step size;

[0183] The detection unit is used to perform target detection on the real-time monitoring image and each image block through the trained detection network model, and to deduplicate the detection results through the non-maximum suppression method to obtain the final detection result.

[0184] In an optional implementation, the monitoring and early warning module 205 includes:

[0185] A calculation unit is used to calculate the maximum overlap rate between each target detection frame and the preset protection target area according to the maximum overlap rate calculation formula. The maximum overlap rate calculation formula is:

[0186]

[0187] Among them, P over represents the maximum overlap rate, M represents the number of target detection frames to be tested, and B direpresents the i-th target detection box, B s Indicates the target area for protection;

[0188] A sending unit is used to send a warning signal if the maximum overlap rate is greater than or equal to a preset warning threshold.

[0189] In an optional embodiment, the sending unit is specifically used to send a secondary warning signal if the maximum overlap rate is greater than or equal to a preset first warning threshold and less than a preset second warning threshold; if the maximum overlap rate is greater than or equal to the preset second warning threshold, then send a first warning signal.

[0190] In an optional embodiment, the anti-external damage monitoring and early warning device further includes:

[0191] An adding module is used to add the real-time monitoring image in which the detection frame of the target to be detected is detected in the current cycle as a sample scene image to the sample data set according to a preset iterative training cycle;

[0192] The iterative training module is used to iteratively train the detection network model in the target detection model according to the sample data set.

[0193] Embodiment 4

[0194] Please refer to Fig.12 , Embodiment 4 of the present invention is: an electronic device, the electronic device comprising:

[0195] One or more processors 301;

[0196] Storage device 302, used to store one or more programs;

[0197] When the one or more programs are executed by the one or more processors 301, the one or more processors 301 implement the various processes in the embodiments described above and can achieve the same technical effect. To avoid repetition, they are not repeated here.

[0198] Embodiment 5

[0199] Embodiment 5 of the present invention provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the various processes in the embodiments described above are implemented and the same technical effects can be achieved. To avoid repetition, they will not be described here.

[0200] In summary, the present invention provides a method for training a target detection model, a method and device for monitoring and early warning against external damage, which construct a network structure based on the Internimage visual large model and DETR, and use the strong semantic understanding ability of the large model to improve the detection accuracy of external damage targets. By adopting a multi-image fusion data enhancement method, and adopting an original + sliding window test method in the application stage, and introducing multi-scale features in the network architecture design, the problems of small targets and long-tail distribution in the anti-external damage scenario are effectively solved. By adopting multi-head hybrid collaborative training and adapter training methods during model training, and by regularly expanding the external damage data set and iteratively training the target detection model, catastrophic forgetting problems in large model training can be prevented.

[0201] Through the above description of the implementation methods, the technicians in the relevant field can clearly understand that the present invention can be implemented by means of software and necessary general hardware, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (FLASH), a hard disk or an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.

[0202] It is worth noting that in the embodiment of the above-mentioned device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.

[0203] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent transformations made using the contents of the present invention's specification and drawings, or directly or indirectly applied in related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A method for training a target detection model, characterized in that: include: Acquire a sample data set, wherein the sample data set includes a preset number of sample scene images and a category label and a detection label corresponding to each sample scene image, wherein the detection label includes the coordinates of a labeling box of each target to be detected in the sample scene image; Constructing a target detection model, wherein the target detection model includes a detection network model and an auxiliary head model, wherein the detection network model includes a backbone network, a feature enhancement network and a detection head, wherein the backbone network includes a large visual model and an adapter model, wherein the adapter model is parallel to the stem layer, the stage1 layer and the stage2 layer in the large visual model, and the auxiliary head model is parallel to the detection head in the detection network model; According to the sample data set and the preset target loss function, the detection network model in the target detection model is trained, wherein the target loss function is constructed according to the original loss function corresponding to the detection network model and the auxiliary loss function corresponding to the auxiliary head model; During training, the input data of the target detection model is a fused image obtained by performing data enhancement on the sample scene images in the sample data set; The sample scene images in the sample data set are enhanced by: Randomly selecting four sample scene images from the sample data set, and fusing the four sample scene images by a Mixup data enhancement method and / or a Mosaic data enhancement method to obtain a fused image; The four sample scene images are fused by the Mixup data enhancement method to obtain a fused image, specifically: the four sample scene images are paired in pairs; two sample scene images belonging to the same pair are fused to obtain a fused image, and the category label and detection label corresponding to the fused image are determined according to the category labels and detection labels corresponding to the two sample scene images belonging to the same pair; The four sample scene images are fused by the Mosaic data enhancement method to obtain a fused image, specifically: the cropping areas corresponding to the four sample scene images are determined according to the annotation boxes of the targets to be measured in the four sample scene images, and the four sample scene images are cropped according to the corresponding cropping areas to obtain four cropped images; the four cropped images are rotated, scaled and color-changed respectively, and then spliced ​​in the form of four grids to obtain a fused image.

2. The target detection model training method according to claim 1, characterized in that: The adapter model includes a backbone adaptation module, a first adaptation module, a downsampling module and a second adaptation module; The backbone adaptation module is used to sequentially downsample the image input to the backbone network to obtain feature maps of different dimensions of a preset first number, stack the feature maps of a preset second number therein to obtain a stacked feature map, and output the stacked feature map to the first adapter module; The first adapter module is used to use the stacked feature map as a key and a value, and the feature map output by the stem layer in the visual large model as a query, to generate a first adapted feature map, and output the first adapted feature map to the downsampling module; at the same time, the first adapted feature map is superimposed with the first dimensional feature map output by the stage1 layer in the visual large model, and the superimposed first dimensional feature map is output to the first downsampling layer between the stage1 layer and the stage2 layer in the visual large model; The downsampling module is used to downsample the first adaptation feature map according to a preset downsampling multiple, and output the downsampled first adaptation feature map to the second adapter module; The second adapter module is used to use the downsampled first adapted feature map as the key and value and the feature map output by the first downsampling layer in the visual large model as the query to generate a second adapted feature map; superimpose the second adapted feature map with the second dimensional feature map output by the stage2 layer in the visual large model, and output the superimposed second dimensional feature map to the second downsampling layer between the stage2 layer and the stage3 layer in the visual large model.

3. The method for training a target detection model according to claim 1, characterized in that: The auxiliary head model includes a first auxiliary head module and a second auxiliary head module, and the auxiliary loss function includes a first auxiliary loss function corresponding to the first auxiliary head module and a second auxiliary loss function corresponding to the second auxiliary head module; The objective loss function is L train =L fast-rcnn +L atss +L DETR , where L train Represents the overall loss value of the target detection model, L fast-rcnn represents the first auxiliary loss value calculated by the first auxiliary loss function, L atss represents the second auxiliary loss value calculated by the second auxiliary loss function, L DETR Represents the original loss value calculated by the original loss function.

4. The method for training a target detection model according to claim 1, characterized in that: In the Mixup data enhancement method, the fused image I mix =xI i+j +yI i+k , the category label Y corresponding to the fused image mix =xY i+j +yY i+k , the detection label B corresponding to the fused image mix =concat(B i+j ,B i+k ), I i+j and I i+k Represents two sample scene images belonging to the same pair, Y i+j and B i+j Represents the sample scene image I i+j The corresponding category label and detection label, Y i+k and B i+k Represents the sample scene image I i+k The corresponding category labels and detection labels, x and y represent the preset fusion weights, and concat() represents the connection function.

5. The method for training a target detection model according to claim 1, characterized in that: In the Mosaic data enhancement method, the coordinates of the four corners of the cropped area corresponding to a sample scene image are (min n (b x1 )-z, min n (b y1 )-z, max n (b x2 )+z,max n (b y2 )+z)、min n (b x1 ) and min n (b y1 ) are the minimum values ​​of the upper left corner coordinates of the labeled box of each target to be measured in the sample scene image in the X-axis and Y-axis directions, respectively, n (b x2 ) and max n (b y2 ) are respectively the maximum values ​​in the X-axis and Y-axis directions of the coordinates of the lower right corner of the annotation box of each target to be measured in the sample scene image, and z is a preset margin constant.

6. The method for training a target detection model according to claim 5, characterized in that: After determining the cropping areas corresponding to the four sample scene images according to the marked frames of the respective targets to be measured in the four sample scene images, the method further includes: If the width of the cropped area is less than the preset width threshold and the height is less than the preset height threshold, the cropped area is enlarged, and the coordinates of the four corners of the enlarged cropped area are (min n (b x1 )-zW gap / 2, min n (b y1 )-zH gap / 2, max n (b x2 )+z+W gap / 2, max n (b y2 )+z+H gap / 2), where W gap and H gap The width and height of the cropping area, respectively.

7. A method for monitoring and early warning against external damage, characterized in that: include: Acquire a real-time monitoring image, and perform target detection through a detection network model to obtain a detection frame of the target to be detected, wherein the detection network model is a detection network model in a target detection model trained by the training method according to any one of claims 1 to 6, and the target to be detected is an external target; According to the target detection frame and the preset protection target area, anti-external damage monitoring and early warning are performed.

8. The anti-external damage monitoring and early warning method according to claim 7 is characterized in that: The real-time monitoring image is acquired, and target detection is performed through the detection network model to obtain the target detection frame to be detected, specifically: Acquire a real-time monitoring image, and obtain a plurality of image blocks in the real-time monitoring image according to a sliding window of a preset size and a preset step size; The trained detection network model is used to perform target detection on the real-time monitoring image and each image block, and the detection results are deduplicated using a non-maximum suppression method to obtain the final detection result.

9. The anti-external damage monitoring and early warning method according to claim 8 is characterized in that: The anti-external damage monitoring and early warning is performed according to the target detection frame to be detected and the preset protection target area, specifically: According to the maximum overlap rate calculation formula, the maximum overlap rate between each target detection frame and the preset protection target area is calculated. The maximum overlap rate calculation formula is: , Among them, P over represents the maximum overlap rate, M represents the number of target detection frames to be tested, and B di represents the i-th target detection box, B s Indicates the target area for protection; If the maximum overlap rate is greater than or equal to a preset warning threshold, a warning signal is sent.

10. The anti-external damage monitoring and early warning method according to claim 9 is characterized in that: If the maximum overlap rate is greater than or equal to a preset warning threshold, a warning signal is sent, specifically: If the maximum overlap rate is greater than or equal to a preset first warning threshold and less than a preset second warning threshold, a secondary warning signal is sent; If the maximum overlap rate is greater than or equal to a preset second warning threshold, a first-level warning signal is sent.

11. The anti-external damage monitoring and early warning method according to claim 7, characterized in that: After the anti-external damage monitoring and early warning are performed according to the target detection frame to be detected and the preset protection target area, the method further includes: According to a preset iterative training cycle, the real-time monitoring image in which the detection frame of the target to be tested is detected in the current cycle is used as a sample scene image and added to the sample data set; According to the sample data set, the detection network model in the target detection model is iteratively trained.

12. An electronic device, characterized in that: The electronic device comprises: one or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the target detection model training method as described in any one of claims 1 to 6 or the anti-external damage monitoring and early warning method as described in any one of claims 7 to 11.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of the target detection model training method as described in any one of claims 1 to 6 or the anti-external damage monitoring and early warning method as described in any one of claims 7 to 11.

Citation Information

Patent Citations

  • Lightweight target detection system for resource-constrained equipment and construction method of lightweight target detection system

    CN115375962A

  • Small target detection method based on enhanced feature extraction

    CN115984172A