Small target detection method of Internet of Things video and related device
By combining a global-local target detection network and a target-dense region extraction network, along with super-resolution processing and progressive scale fusion, the problems of missed detection and false detection in dense small target detection are solved, achieving high-precision small target detection in IoT videos.
Patent Information
- Application Number
- CN202511071079.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-07
AI Technical Summary
Existing target detection methods suffer from missed detections and false detections in dense small target detection scenarios, resulting in low detection accuracy.
A global-local target detection network is used for global detection, combined with a target-dense region extraction network and super-resolution processing. Fine detection is performed through the global-local target detection network, and the detection results are fused in a progressive scale fusion network to improve detection accuracy.
It effectively avoids missed or false detections, significantly improves the accuracy of dense small target detection, and enhances the accuracy and practicality of IoT video analysis.
Smart Images

Figure CN120912869A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of Internet of Things and big data, and particularly relates to a small target detection method for Internet of Things video and a related device. BACKGROUND
[0002] With the rapid development of Internet of Things technology, Internet of Things video is widely used in smart city, industrial monitoring, smart home and other scenarios. Target detection, as a core technology of Internet of Things video analysis, directly affects the intelligent level and application value of the system. However, the existing target detection method often appears missed detection or false detection in the dense small target detection scene, thereby resulting in low detection accuracy. SUMMARY
[0003] In view of the above problems, the present application provides a small target detection method for Internet of Things video and a related device to realize the purpose of high-precision target detection for Internet of Things video. The specific scheme is as follows:
[0004] The first aspect of the present application provides a small target detection method for Internet of Things video, comprising:
[0005] inputting the Internet of Things video to be detected into a global-local target detection network for global detection to obtain a global detection result;
[0006] inputting the global detection result into a target dense area extraction network to obtain a target dense area, performing super-resolution processing on the target dense area to obtain a high-resolution local image;
[0007] inputting the high-resolution local image into the global-local target detection network for fine detection to obtain a local detection result;
[0008] inputting the global detection result and the local detection result into a progressive scale fusion network for fusion to obtain a high-precision small target detection result.
[0009] In one possible implementation, the Internet of Things video to be detected is input into the global-local target detection network for global detection to obtain a global detection result, comprising:
[0010] extracting image features from the Internet of Things video to be detected by using a classic backbone network of a fusion cross-stage local network;
[0011] mapping the image features to a recursive feature pyramid to obtain multi-scale image features;
[0012] inputting the multi-scale image features into a region candidate network for target positioning to obtain target candidate region feature information;
[0013] The target candidate region feature information is input into a Transformer detection head for end-to-end set prediction and two-part matching to obtain a global detection result.
[0014] In a possible implementation, a cross-stage local network of a classic backbone network fused with a cross-stage local network extracts image features from the to-be-detected video of Internet of Things, including:
[0015] The to-be-detected video of Internet of Things is input into the classic backbone network to obtain low-level image features;
[0016] The low-level image features are input into the cross-stage local network, and the low-level image features are evenly divided into first low-level feature images and second low-level feature images;
[0017] The first low-level feature images are propagated through a straight-through path to obtain straight-through path image features; and the second low-level feature images are propagated through a complex path to obtain complex path image features;
[0018] The straight-through path image features and the complex path image features are spliced to obtain image features including different gradient combinations.
[0019] In a possible implementation, the multi-scale image features are input into a region candidate network for target positioning to obtain target candidate region feature information, including:
[0020] An initial candidate region is determined according to the multi-scale image features;
[0021] An intersection-over-union between the initial candidate region and a preconfigured real box is calculated;
[0022] The initial candidate region whose intersection-over-union reaches a preset overlap threshold is determined as a target candidate region;
[0023] The multi-scale image features are subjected to accurate alignment feature pooling processing to extract global prediction box information and position encoding information of the target candidate region; and the target candidate region, the global prediction box information and the position encoding information of the target candidate region are integrated to obtain the target candidate region feature information.
[0024] In a possible implementation, the global detection result is input into a target dense region extraction network to obtain a target dense region, and the target dense region is subjected to super-resolution processing to obtain a high-resolution local image, including:
[0025] An aggregation score of each global detection box in the global detection result is calculated;
[0026] If the aggregation score reaches a preset score threshold, coordinates of the global detection box are determined as effective coordinates;
[0027] The effective coordinates are subjected to density clustering processing to obtain target dense sub-regions of different scales;
[0028] The target dense sub-regions of different scales are subjected to scale adjustment to obtain standard target dense regions through cropping;
[0029] The target dense regions are input into a super-lightweight super-resolution processing network to obtain high-resolution local images.
[0030] In a possible implementation, the global detection result and the local detection result are input into a progressive scale fusion network for fusion to obtain a high-precision small target detection result, including:
[0031] The global detection result and the local detection result are subjected to weighted box fusion to obtain the high-precision small target detection result.
[0032] The second aspect of the present application provides a small target detection device for Internet of Things video, including:
[0033] A global detection unit is configured to input a to-be-detected Internet of Things video into a global-local target detection network for global detection to obtain a global detection result;
[0034] A super-resolution processing unit is configured to input the global detection result into a target dense region extraction network to obtain a target dense region, and perform super-resolution processing on the target dense region to obtain a high-resolution local image;
[0035] A local detection unit is configured to input the high-resolution local image into the global-local target detection network for fine detection to obtain a local detection result;
[0036] A fusion unit is configured to input the global detection result and the local detection result into a progressive scale fusion network for fusion to obtain a high-precision small target detection result.
[0037] The third aspect of the present application provides a computer program product, including computer readable instructions, when the computer readable instructions run on an electronic device, the electronic device implements the small target detection method for Internet of Things video of the first aspect or any implementation manner of the first aspect.
[0038] The fourth aspect of the present application provides a small target detection device for Internet of Things video, including at least one processor and a memory connected with the processor, wherein:
[0039] The memory is configured to store a computer program;
[0040] The processor is configured to execute the computer program, so that the electronic device can implement the small target detection method for Internet of Things video of the first aspect or any implementation manner of the first aspect.
[0041] The fifth aspect of the present application provides a computer storage medium, the storage medium carries one or more computer programs, when the one or more computer programs are executed by an electronic device, the electronic device can execute the small target detection method of the Internet of Things video of the first aspect or any implementation manner of the first aspect.
[0042] By the above technical solution, the small target detection method and related device of the Internet of Things video provided by the present application first uses the global-local target detection network to perform global detection on the to-be-detected Internet of Things video to obtain a global detection result; uses the target dense area extraction network to determine a target dense area from the global detection result, performs super-resolution processing on the target dense area to obtain a high-resolution local image; again uses the global-local target detection network to perform fine detection on the high-resolution local image to obtain a local detection result; and finally fuses the global detection result and the local detection result to obtain a high-precision small target detection result. The present application performs fine local detection on the result of global detection on the basis of performing global detection on the to-be-detected Internet of Things video, and finally fuses the two detection results, effectively avoiding the situation of missed detection or false detection, and can greatly improve the detection precision of the dense small target detection in the to-be-detected Internet of Things video. BRIEF DESCRIPTION OF DRAWINGS
[0043] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals designate identical or similar elements. It is to be understood that the drawings are schematic, and the sizes of the components and elements are not necessarily drawn to scale.
[0044] Figure 1 A flowchart of the small target detection method of the Internet of Things video provided by the present application;
[0045] Figure 2 A structure composition diagram of the overall network architecture of the backbone network CSPResNet101 and RFP provided by the present application;
[0046] Figure 3 An example diagram of the network framework structure of the global-local target detection network provided by the present application;
[0047] Figure 4 An example diagram of the training and testing flow of the ultra-lightweight super-resolution processing network provided by the present application;
[0048] Figure 5 An example diagram of the small target detection method of the Internet of Things video provided by the present application;
[0049] Figure 6 A structure diagram of the small target detection device of the Internet of Things video provided by the present application;
[0050] Figure 7 A structure diagram of a small target detection device of an Internet of Things video provided in the present application is shown. DETAILED DESCRIPTION
[0051] The embodiments of the present application are described below in conjunction with the accompanying drawings of the embodiments of the present application. The terms used in the embodiment part of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0052] The embodiments of the present application are described below in conjunction with the accompanying drawings. It is known to those of ordinary skill in the art that, as technology develops and new scenarios appear, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0053] The terms “first”, “second”, and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, and this is only a distinguishing way used in the description of the embodiments of the present application to describe the objects with the same attributes. In addition, the terms “include” and “have” and any variations thereof are intended to cover non-exclusive inclusion, so that the processes, methods, systems, products or devices containing a series of units do not have to be limited to those units, but can include other units that are not clearly listed or inherent to these processes, methods, products or devices.
[0054] Internet of Things video analysis is mainly a process of analyzing and processing video data using Internet technology, and usually involves video acquisition, data transmission, video analysis and intelligent decision-making. Target detection is a core technology of Internet of Things video analysis, which is a core task in the field of computer vision, and can automatically identify the position and category of a specific target from an image or video.
[0055] The existing target detection methods mainly include traditional methods and deep learning-based target detection methods:
[0056] Traditional target detection methods rely on hand-designed rules for feature extraction, such as detection methods based on visual contrast mechanism, which use the characteristics that the target has a higher radiation energy in a local range and has a significant difference with the adjacent background, to design corresponding operators to amplify the pixel difference between the target and the background region; by selecting an appropriate threshold, the target is extracted from the background.
[0057] The traditional target detection method has insufficient feature expression capability, and the manual features are difficult to adapt to the dynamic changes of target shape, texture and background clutter in a complex scene. The robustness is poor, and when the contrast between the target and the background is reduced, such as in low light or under occlusion, or the target is densely distributed, the traditional target detection method cannot effectively distinguish the target from the interference, and the target detection accuracy sharply decreases.
[0058] The target detection method based on deep learning completes different tasks such as feature extraction, fusion and classification in the same network, directly takes the image as the input of the network, and autonomously trains and learns the features representing the deep nature of the data. The target detection method based on deep learning: although good detection performance has been achieved in the field of natural light target detection, in the scene with few target scales, uneven distribution and high local density, the detection accuracy sharply decreases, resulting in that most of the existing target detection networks cannot be directly used in the field of dense small target detection.
[0059] In summary, the existing target detection method has low detection accuracy for small targets in the Internet of Things video scene, which specifically shows that:
[0060] High missing detection rate: dense small targets have overlapping features or background interference, and the traditional method cannot effectively distinguish them, and the deep learning model also has insufficient multi-scale feature fusion, resulting in missing detection;
[0061] High false detection rate: manual features are easily affected by noise, and the deep learning model is easy to misjudge the background noise as a target in a low resolution or occlusion scene.
[0062] To solve the above problems, the present application provides a small target detection method for Internet of Things video and related devices.
[0063] Detecting dense small targets in Internet of Things video analysis can improve the accuracy and efficiency of video analysis. Among them, dense small targets usually refer to a large number of small objects in a video, such as pedestrians in a crowd, vehicles in traffic, or crops in agriculture, etc. It is not difficult to understand that detecting these dense small targets helps to achieve more fine scene understanding, such as in an intelligent traffic system, detecting dense small vehicles can optimize traffic flow; in agricultural monitoring, detecting dense small crops can help detect the growth of crops. As can be seen, dense small target detection can also improve the practicality of Internet of Things video analysis.
[0064] The small target detection method for Internet of Things video and related devices provided by the present application can improve the accuracy and practicality of Internet of Things video analysis, effectively avoid missing detection or false detection of dense small targets, improve the intelligent level of the Internet of Things system, and provide more reliable data support for subsequent decision-making.
[0065] Reference Figure 1A flowchart of a small target detection method of an Internet of Things video is provided in the present application, as shown in Figure 1 The small target detection method of the Internet of Things video comprises the following steps:
[0066] Step 101: input the Internet of Things video to be detected into a global-local target detection network for global detection to obtain a global detection result.
[0067] It should be noted that a multi-scale dense small target detection framework is designed for dense small target detection in the present application, which can be divided into two parts: one part is the global-local target detection network in the present step, and the other part is the target dense region extraction network mentioned in step 102.
[0068] The Golbal-Local Target Detection Network (G-LTD) is applied twice in the present application: firstly, it is used for global preliminary detection of the Internet of Things video to be detected in the present step to obtain the outline and distribution of the target in the Internet of Things video to be detected; secondly, it is used for fine detection of the target dense region extracted by the target dense region extraction network in step 103 to further obtain more accurate positioning information of the dense small target, so as to further improve the accuracy of the small target detection of the Internet of Things video.
[0069] The network framework of the global-local target detection network mainly comprises three parts: a Backbone (main network), a Neck (neck), and a Head (detection head). Among them, the main network is a classic backbone network fused with a cross-stage local network, the neck is a recursive feature pyramid, and the detection head is a Transformer detection head capable of realizing end-to-end detection.
[0070] The classic backbone network fused with the cross-stage local network can be a CSPResNet101 obtained by improving ResNet101 with a cross-stage local network (CrossStage Partial Network, CSPNet) to realize the extraction of image features from the Internet of Things video to be detected and provide feature information for subsequent detection tasks.
[0071] The design principle of the network architecture of the classic backbone network fused with the cross-stage local network is as follows:
[0072] Considering that the global-local target detection network needs to extract detailed feature information from the to-be-detected video of the Internet of Things, that is, the classic backbone network of the cross-stage local network needs to extract more detailed feature information from the to-be-detected video of the Internet of Things. Network architecture designers associate: deepening the network depth is beneficial to the extraction of image features. However, with the increase of the level of extracted features, gradient dispersion or explosion becomes an obstacle to training deep networks, and in extreme cases, network degradation even occurs, causing the network to fail to converge.
[0073] In order to solve the above problems, first, the classic backbone network ResNet is introduced, which mainly breaks through the training bottleneck of deep networks by residual structure in image feature extraction, replicates the features of shallow networks to deep networks, and ensures the training performance of deep networks, so as to extract more complex and hierarchical image features.
[0074] Further, in order to further improve the accuracy and efficiency of image feature extraction of the classic backbone network ResNet, network architecture designers consider starting from enriching the gradient combination of the classic backbone network ResNet, and fusing the cross-stage partial network (CrossStage Partial Network, CSPNet) into the classic backbone network ResNet to obtain CSPResNet101 as the backbone network in the global-local target detection network.
[0075] The backbone network CSPResNet101 fused with CSPNet realizes the problem of relieving ResNet from needing a large amount of inference calculation from the perspective of network architecture. The process of obtaining rich gradient combination can be as follows:
[0076] The to-be-detected video of the Internet of Things is input into the classic backbone network to obtain low-level image features. The low-level image features are input into the cross-stage local network, and the low-level image features are evenly divided into first low-level feature images and second low-level feature images; the first low-level feature images are propagated through a straight-through path to obtain straight-through path image features; the second low-level feature images are propagated through a complex path to obtain complex path image features; and the straight-through path image features and the complex path image features are spliced to obtain image features including different gradient combinations.
[0077] Specifically, the cross-stage local network receives low-level feature maps from the classic backbone network, divides the low-level feature maps into two parts, and then merges the two parts of feature maps through the cross-stage hierarchical structure of the cross-stage local network. Figure 1
[0078] It should be noted that fusing cross-stage local networks can enhance the feature extraction capabilities of each convolutional layer in the backbone network, and can also evenly distribute the computational load of each layer, overcoming computational bottlenecks and ensuring sufficient accuracy while improving computational speed.
[0079] The Global-Local Object Detection Network introduces a Recursive Feature Pyramid (RFP) network as a feature hub connecting the backbone network and the Transformer detection head in the Global-Local Object Detection Network. The RFP mainly achieves feature reprocessing and rational utilization through multi-scale feature fusion, which can further improve the diversity of features extracted by the Global-Local Object Detection Network and the robustness of the network.
[0080] Optionally, the backbone network CSP ResNet101 maps its extracted image features to a recursive feature pyramid to obtain multi-scale image features.
[0081] In addition, additional feedback connections are set in the recursive feature pyramid to bring semantically rich high-level features back to the lower-level feature layers of the backbone network, thereby enhancing the feature extraction performance of the backbone network and accelerating network training.
[0082] Specifically, in the backbone network CSP ResNet101, Indicates the first from bottom to top Layered convolution operations, in a recursive feature pyramid structure This indicates a top-down RFP operation. After adding a feedback connection to the RFP, it uses... This represents the feature transformation before feeding high-level features back to the backbone network CSPResNet101 for processing. Let the series of feature maps output by the backbone network CSPResNet101 be... The RFP outputs a series of feature maps. ,in, It represents the number of convolutional layers in the backbone network CSP ResNet101.
[0083] for The feature maps of the backbone network CSP ResNet101 and RFP are finally output. , They are defined as follows:
[0084] ;
[0085] ;
[0086] for Set the number of iterations to 2, and use superscript. Indicates the first The operation and features of the step iteration, the above feedback operation is expanded into a sequential network, the mathematical formula is as follows:
[0087] ;
[0088] ;
[0089] For example, referring to Figure 2 , the backbone network CSPResNet101 provided in the present application and the overall network architecture structure composition diagram of RFP.
[0090] As Figure 2 shown, in the backbone network CSPResNet101, the present application adds a convolution layer with a kernel size of 1 in the first block of each stage of the classical backbone network ResNet, which can simultaneously accept images and as input, and sets the initial weight to 0 to ensure that the convolution layer does not have any actual impact when loading the weight from the pre-trained checkpoint.
[0091] Finally, the network framework of the global-local target detection network is introduced, and the detection head part adopts the Transformer which can realize end-to-end detection, and the end-to-end set prediction and two-part matching are carried out, and the two-part matching mode is used to force one-to-one prediction, which avoids the operation of NMS (Non-Maximum Suppression) in the global-local target detection network and ensures the accuracy of detection.
[0092] The input of the Transformer detection head is the target candidate region feature information obtained after the region candidate network performs target positioning on the multi-scale image features output by the recursive feature pyramid.
[0093] It should be noted that the region candidate network can be an RPN (Region Proposal Network) network, which is a key component in the network framework of the global-local target detection network, and shares the multi-scale image features extracted by the backbone network with the entire global-local target detection network. It is mainly used to automatically generate candidate regions, judge whether there is a target in the candidate region, realize almost cost-free region suggestion, and further improve the detection accuracy of the global-local target detection network.
[0094] Optionally, the multi-scale image features output by the recursive feature pyramid are input into the RPN network, initial candidate regions are determined by the RPN network, the intersection-over-union between the initial candidate regions and the pre-configured real boxes is calculated, the initial candidate regions whose intersection-over-union reaches a preset overlap threshold are determined as target candidate regions, feature pooling processing for accurate alignment of the multi-scale image features is performed, global prediction box information and position encoding information of the target candidate regions are extracted therefrom, the target candidate regions, the global prediction box information and the position encoding information of the target candidate regions are integrated, and target candidate region feature information is obtained.
[0095] Specifically, the initial candidate regions can be anchor boxes of different scales generated by the RPN network covering the image, a binary classification label is assigned to each candidate region, the IoU between the candidate region and the real box is calculated, when the IoU overlap between the initial candidate region and the real box is higher than a preset overlap threshold, a positive label is given to the initial candidate region, for example, the positive label is 1, which represents that the initial candidate region is a target candidate region; correspondingly, when the IoU overlap between the initial candidate region and the real box is lower than the preset overlap threshold, the label given to the initial candidate region is 0, and the label 0 represents that the initial candidate region is not a target candidate region, all initial candidate regions assigned with the label 1 constitute the target candidate region; then RolAlign (Region of Interest Align) is used to extract global prediction box information and position encoding information of the target candidate region from the multi-scale image features; finally, the obtained target candidate region, global prediction box information and position encoding information of the target candidate region are integrated into target candidate region feature information; finally, the target candidate region feature information is input into the Transformer encoder.
[0096] For example, the position information of the candidate region is defined by four variables, is the upper left corner coordinate of the prediction box (target candidate region), is the width and height of the prediction box. The position encoding is defined as wherein, represents connection, and the function PE is defined as:
[0097] ;
[0098] ;
[0099] The target candidate region feature information is input into the Transformer encoder, aggregated through the self-attention mechanism, and finally reaches the shared feedforward network (FFN, Feed Forward Networks). The FFN is composed of a 3-layer perceptron with ReLU activation function and a linear layer. The shared feedforward network predicts the class label and the bounding box of each target candidate region. The final detection result, i.e., the global prediction result, can be represented as: wherein, represents a low-resolution original image, is a predicted frame code, is a target class, is a predicted frame information.
[0100] It should be noted that the Hungarian loss is used to train and supervise the Transformer detection head, and the prediction of the global-local target detection network is assigned a real label.
[0101] For example, the Hungarian loss is represented as represents a real target set, and represents a prediction set. The Hungarian loss is represented as follows:
[0102] ;
[0103] wherein, because there is a case where the predicted frame corresponds to no target object, generally , is a classification loss, is a bounding box regression loss, is a matching loss between and . The specific representation can be as follows:
[0104] ;
[0105] wherein, is an N-prediction frame set, is a pair matching loss.
[0106] From this, the network framework of the global-local target detection network is completed.
[0107] For example, referring to Figure 3 , the network framework structure example diagram of the global-local target detection network provided by the application is provided.
[0108] As Figure 3As shown, the network framework adopts a backbone network including multiple convolutional layers, extracts image features from the to-be-detected video of Internet of Things by the backbone network, maps the image features to a recursive feature pyramid to obtain multi-scale image features, inputs the multi-scale image features to a region proposal network and RolAlign respectively, extracts global prediction frame information and position encoding information of target candidate regions from the multi-scale image features by using RolAlign, performs target positioning on the multi-scale image features by using the region proposal network to determine the target candidate regions in the form of region suggestion frames, and finally inputs target candidate region feature information obtained by integrating the target candidate regions, the global prediction frame information and the position encoding information of the target candidate regions into a Transformer detection head, aggregates by using a self-attention mechanism of the Transformer detection head, and finally reaches an FFN to obtain a global detection result after FFN processing.
[0109] Optionally, the global-local target detection network is first used for global detection of the to-be-detected video of Internet of Things, image features are extracted from the to-be-detected video of Internet of Things by inputting the to-be-detected video of Internet of Things into a classic backbone network of a fusion cross-stage local network, the image features are mapped to a recursive feature pyramid to obtain multi-scale image features, the multi-scale image features are input into a region proposal network for target positioning to obtain target candidate region feature information, and finally the target candidate region feature information is input into a Transformer detection head for end-to-end set prediction and two-part matching to obtain a global detection result.
[0110] Step 102, inputting the global detection result into a target dense area extraction network to obtain a target dense area, performing super-resolution processing on the target dense area to obtain a high-resolution local image.
[0111] Next, another part of the multi-scale dense small target detection framework, the target dense area extraction network, is introduced.
[0112] The target dense area extraction network (T-DAE) is mainly used for adaptive clipping according to the global detection result to obtain a target dense area, and performing super-resolution processing on the clipped target dense area to obtain a high-resolution local image.
[0113] The global-local target detection network is used to complete global detection on the to-be-detected IOT image. In order to further improve the robustness and accuracy of the detection of dense small targets, a target dense region extraction network is used to realize the interception of the target dense region in the video image, and through super-resolution processing, the features of the intercepted target dense region are strengthened to obtain a high-resolution local image. The global-local target detection network is used for the second fine detection, so as to improve the detection accuracy of the to-be-detected IOT video and obtain more accurate detection results.
[0114] Optionally, the process of the target dense region extraction network processing the global detection result can be as follows:
[0115] Step one, calculate the aggregation score of each global detection frame in the global detection result.
[0116] An aggregation score calculation model is used to calculate the aggregation score of each global detection frame in the global detection result. The aggregation score calculation model is mainly used to measure the density of the region where the bounding box of the global detection frame obtained in the global detection result is located.
[0117] Let the bounding box be: ;
[0118] Wherein, represents a low-resolution original image, is a prediction frame code, is a target category, is a prediction frame information.
[0119] The related model is represented as follows: ;
[0120] Wherein, x and y are horizontal and vertical coordinates respectively, and G(x, y) is the aggregation score.
[0121] Step two, if the aggregation score reaches a preset score threshold, the coordinates of the global detection frame are determined as effective coordinates.
[0122] A preset score threshold is set, and the coordinates of the global detection frame whose aggregation score reaches the preset score threshold are set as effective coordinates. The effective coordinates can also be called high-score coordinates, and the global detection frame corresponding to the high-score coordinates can also be called an adaptive selection region. All effective coordinates can form an effective coordinate set.
[0123] For example, the preset score threshold can be set as Let be an effective coordinate set, and the effective coordinate set can be represented as:
[0124] ;
[0125] where x, y are horizontal and vertical coordinates respectively, and G(x, y) is the aggregation score.
[0126] Step three, performing density clustering on the effective coordinates to obtain target dense sub-regions of different scales.
[0127] It should be noted that the density clustering process can specifically adopt DBSCAN (Density-Based Spatial Clustering of Application with Noise), a clustering region selection algorithm, which takes the set of effective coordinates and the boundary coordinates of each global detection frame as input to obtain target dense sub-regions of different scales.
[0128] Specifically, each coordinate in the target dense sub-region is assigned to a specific class, and the boundary of the target dense sub-region can be easily obtained, and the intercepted target dense sub-region contains all the global frames of the target, which can avoid object truncation.
[0129] Step four, performing scale adjustment on the target dense sub-regions of different scales to obtain standard target dense regions.
[0130] Further, considering that the scales of the target dense sub-regions after adaptive interception are different, some may have the problem of excessively large or small height / width ratio, which cannot be directly sent to the local detector, so scale adjustment is needed. To keep the size and scale of the target dense sub-region within a reasonable range, an adaptive scale adjustment method for the target dense sub-region is designed, which can be as follows:
[0131] Let the boundary box of the target dense sub-region be , where is the top-left corner coordinate of the boundary box, is the bottom-right corner coordinate of the boundary box, the center coordinate is , the scale standard is , the width-height ratio is , and the scale adjustment formula can be as follows:
[0132] ;
[0133] ;
[0134] ;
[0135] where the scale standard is limited to , the width-height ratio is , and the boundary box of the target dense sub-region that needs to be cropped can be represented as:
[0136] ;
[0137] ;
[0138] ;
[0139] ;
[0140] wherein, and respectively represent the final cropping height and width of the target dense sub-region, represents a scale standard.
[0141] Based on the above formula, the target dense sub-region of different scales is adaptively scaled to a reasonable range, and a standard target dense region is obtained, which can also be referred to as a final cropped image .
[0142] Step five, input the target dense region into the super-lightweight super-resolution processing network to obtain a high-resolution local image.
[0143] After obtaining the standard target dense region in the above step, the final cropped image obtained will inevitably have problems such as image quality blur, image detail loss, and resolution reduction. In view of these problems, a super-lightweight super-resolution processing network is used in the present application to obtain a high-resolution local image.
[0144] It should be noted that the super-lightweight super-resolution processing network can be s-LWSR (SuperLightweight Super-Resolution Network), which can achieve similar performance to other advanced but cumbersome super-resolution methods under limited parameters and operations, and realizes efficient super-resolution processing through a lightweight network structure.
[0145] For example, referring to Figure 4 , the present application provides an example diagram of the training and testing process of the super-lightweight super-resolution processing network.
[0146] As shown in Figure 4 , a single image super-resolution method is used to restore the lost target edge information of the low-resolution image based on the prior knowledge of the image, so as to obtain a high-definition image and provide more semantic information for the target dense region extraction network.
[0147] For the training process: bicubic interpolation is used on the original image to downsample the original image to 0.5 times the previous size, and it is input as the training set of the super-lightweight super-resolution processing network.
[0148] For the test process: considering that the target dense area size intercepted by the adaptive selection algorithm in the previous step is not uniform, it is unnecessary to perform super-resolution processing on a large-area target dense area, and therefore a scale threshold can be set , which is less than the scale threshold before the proportion adjustment The image is input into the super-lightweight super-resolution processing network for processing to achieve resolution improvement.
[0149] For example, assuming that a certain target dense area obtained by cropping After processing by the super-lightweight super-resolution processing network, the final expression obtained can be:
[0150] ;
[0151] wherein, the function represents a super-resolution image generated using the super-lightweight super-resolution processing network, is the scale of the target dense area, and based on this, the final output is obtained.
[0152] Step 103: inputting the high-resolution local image into the global-local target detection network for fine detection to obtain a local detection result.
[0153] It should be noted that this is the second application of the global-local target detection network, and the purpose is to perform fine detection on the target dense area extracted by the target dense area extraction network to further obtain more accurate positioning information of the dense small target, thereby further improving the accuracy of the entire Internet of Things video small target detection.
[0154] The global-local target detection network has been described in detail in step 101, and therefore will not be described again here.
[0155] Optionally, the high-resolution local image obtained after processing by the target dense area extraction network is subjected to fine detection by the global-local target detection network to obtain a local detection result.
[0156] It is not difficult to understand that the global-local target detection network is used to extract more detailed features in the high-resolution local image.
[0157] Step 104: inputting the global detection result and the local detection result into the progressive scale fusion network for fusion to obtain a high-precision small target detection result.
[0158] It should be noted that in the progressive scale fusion network, the global detection result and the local detection result can be fused in a weighted fusion box (WBF, Weighted Boxes Fusion) manner. Specifically, the global detection result can be a global detection box, and the local detection result can be a local fine detection box. The WBF can fuse detection boxes of different scales.
[0159] Optionally, the global detection box and the local fine detection box are fused in a weighted box fusion manner to obtain a high-precision small target detection result.
[0160] Exemplarily, the formula for fusing the detection results of the two scales to obtain the final detection result can be:
[0161] ;
[0162] wherein, represents the final high-precision small target detection result, represents the video to be detected, represents a high-resolution local image generated by the T-DAE network. represents network detection, represents a WBF weighted box fusion operation.
[0163] In summary, the small target detection method for the video of things provided in the present application first uses a global-local target detection network to perform global detection on the video to be detected, to obtain a global detection result. A target dense region extraction network is used to determine a target dense region from the global detection result, and the target dense region is subjected to super-resolution processing to obtain a high-resolution local image. The global-local target detection network is used again to perform fine detection on the high-resolution local image to obtain a local detection result. Finally, the global detection result and the local detection result are fused to obtain a high-precision small target detection result. The present application performs fine local detection on the result of the global detection on the basis of the global detection on the video to be detected, and finally fuses the two detection results, effectively avoiding the situation of missed detection or false detection, and can greatly improve the detection precision of the dense small target detection in the video to be detected.
[0164] Exemplarily, referring to Figure 5 , the flowchart of the dense small target detection method for the video of things provided in the present application is shown.
[0165] As Figure 5As shown, the to-be-detected object is input into the global-local target detection network G-LTD to obtain a global detection result. Specifically, in the global-local target detection network framework, the recursive feature pyramid structure can fuse the abstract semantic information extracted by the high layer with the detail information such as the contour texture of the low layer from top to bottom, so as to achieve the purpose of feature enhancement. Moreover, the output result can be connected to the backbone network through the feedback connection for secondary feature extraction, which can avoid the loss of small target information and strengthen the feature extraction performance of the backbone network. Secondly, the extraction of the candidate region is realized through the RPN, and since the network structure is simple and embedded conveniently, the region proposal can be realized almost at no cost, which can improve the accuracy and speed up the operation of the target detection model. The Transfomer as the detection head can omit the NMS operation, and the one-to-one prediction can greatly simplify the entire architecture.
[0166] Next, the global detection result is processed by the target dense region extraction network T-DAE. First, the clustering score of each global detection frame in the global detection result is calculated, and the effective coordinate set is determined according to the clustering score. The adaptive region (target dense sub-region) corresponding to the effective coordinate is selected. The selected adaptive region is adaptively adjusted in scale to obtain a scale-standardized sub-region (standard target dense sub-region). Finally, the scale-standardized sub-region is subjected to super-resolution processing to obtain a high-resolution local image. It should be noted that when the target dense region is adaptively extracted, the design of the clustering score calculation model can measure the density of each region. On this basis, the clustering algorithm can extract the dense region in an unsupervised manner. The high-resolution image can be obtained by performing super-resolution processing on the cropped image.
[0167] Again, the global-local target detection network G-LTD is used to process the high-resolution local image to obtain more accurate local detection results, thereby improving the accuracy and robustness of the dense small target detection.
[0168] Finally, the global detection result and the local detection result are fused by WBF. Finally, the global detection result and the local detection result are fused, which can greatly improve the detection accuracy of the entire network.
[0169] In summary, the target detection framework based on Transfomer is designed, CSPResNet101 is used as the backbone network, the recursive FPN is introduced, the high-level semantic information is brought back to the lower-level feature layer of the backbone network through the additional feedback link, the feature extraction performance of the backbone network is enhanced, and the training is accelerated. The Transformer is used as the detection head to realize one-to-one matching between the real box and the predicted box, and manual operations such as NMS are avoided.
[0170] The application designs an adaptive region selection algorithm, which can identify and intercept the region where the target distribution is dense. The algorithm is unsupervised and can be easily inserted into the network for training. By performing super-resolution processing on the intercepted dense region, more feature information can be provided for the local detection network.
[0171] The application fuses the multi-scale detection results as the idea, constructs a gradual scale fusion network, fuses the global detection result and the local fine detection result through the WBF operation to obtain the final output result, and can overcome the problem of high variance of the deep learning model due to the flexibility of the random training algorithm.
[0172] The above introduces a small target detection method for Internet of Things video provided by the embodiment of the application. The device for executing the small target detection method for Internet of Things video will be introduced below.
[0173] Please refer to Figure 6 , Figure 6 The structure diagram of the small target detection device for Internet of Things video provided by the application is shown in FIG. 1. Figure 6 As shown in the figure, the device comprises:
[0174] a global detection unit 10, a super-resolution processing unit 20, a local detection unit 30 and a fusion unit 40; wherein:
[0175] The global detection unit 10 is configured to input the Internet of Things video to be detected into a global-local target detection network for global detection, and obtain a global detection result.
[0176] The super-resolution processing unit 20 is configured to input the global detection result into a target dense region extraction network to obtain a target dense region, and perform super-resolution processing on the target dense region to obtain a high-resolution local image.
[0177] The local detection unit 30 is configured to input the high-resolution local image into the global-local target detection network for fine detection, and obtain a local detection result.
[0178] The fusion unit 40 is configured to input the global detection result and the local detection result into a gradual scale fusion network for fusion, and obtain a high-precision small target detection result.
[0179] In an embodiment, the global detection unit 10 is specifically configured to:
[0180] extract image features from the Internet of Things video to be detected by using a classic backbone network of a fusion cross-stage local network;
[0181] map the image features to a recursive feature pyramid to obtain multi-scale image features;
[0182] Input the multi-scale image features into a region proposal network for target positioning to obtain target candidate region feature information;
[0183] Input the target candidate region feature information into a Transformer detection head for end-to-end set prediction and two-part matching to obtain a global detection result.
[0184] In an embodiment, the global detection unit 10 is specifically configured to:
[0185] Input the to-be-detected video into a classic backbone network to obtain low-level image features;
[0186] Input the low-level image features into a cross-stage local network to divide the low-level image features into first low-level feature images and second low-level feature images;
[0187] Propagate the first low-level feature images through a straight-through path to obtain straight-through path image features; propagate the second low-level feature images through a complex path to obtain complex path image features;
[0188] Splice the straight-through path image features and the complex path image features to obtain image features including different gradient combinations.
[0189] In an embodiment, the global detection unit 10 is specifically configured to:
[0190] Determine initial candidate regions according to the multi-scale image features;
[0191] Calculate the intersection over union between the initial candidate regions and pre-configured real boxes;
[0192] Determine the initial candidate regions whose intersection over union reaches a preset overlap threshold as target candidate regions;
[0193] Perform accurate alignment feature pooling processing on the multi-scale image features to extract global prediction box information and position encoding information of the target candidate regions; integrate the target candidate regions, the global prediction box information, and the position encoding information of the target candidate regions to obtain target candidate region feature information.
[0194] In an embodiment, the super-resolution processing unit 20 is specifically configured to:
[0195] Calculate the aggregation score of each global detection box in the global detection result;
[0196] If the aggregation score reaches a preset score threshold, determine the coordinates of the global detection box as effective coordinates;
[0197] Perform density clustering processing on the effective coordinates to obtain target dense sub-regions of different scales;
[0198] The target dense sub-regions of different scales are proportionally adjusted, and a standard target dense region is obtained by cropping;
[0199] The target dense region is input into a super-lightweight super-resolution processing network to obtain a high-resolution local image.
[0200] In an embodiment, the fusion unit 40 is specifically configured to:
[0201] The global detection result and the local detection result are weighted frame fused to obtain a high-precision small target detection result.
[0202] In an embodiment of the present application, a computer program product is provided, which includes computer readable instructions, when the computer readable instructions are run on an electronic device, the electronic device implements any one of the small target detection methods for Internet of Video provided by the embodiments of the present application.
[0203] In an embodiment of the present application, a small target detection device for Internet of Video is provided. Referring to Figure 7 Fig. 1 shows a structural schematic diagram of a small target detection device for Internet of Video according to an embodiment of the present application. The small target detection device for Internet of Video in the embodiments of the present application can include, but is not limited to, fixed terminals such as mobile phones, notebook computers, PDAs (Personal Digital Assistants), PADs (Tablet Personal Computers), desktop computers, and the like. Figure 7 The small target detection device for Internet of Video shown is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present application.
[0204] As shown in Figure 7 The small target detection device for Internet of Video can include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or programs loaded from a storage device 608 to a random access memory (RAM) 603. In the state that the small target detection device for Internet of Video is powered on, the RAM 603 also stores various programs and data required for the operation of the small target detection device for Internet of Video. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0205] Generally, the following devices can be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 can allow the Internet of Video Small Target Detection Device to communicate wirelessly or wired with other devices to exchange data. Although Figure 7 The Internet of Video Small Target Detection Device with various devices is shown, but it should be understood that all the shown devices are not required to be implemented or possessed. More or fewer devices can be alternatively implemented or possessed.
[0206] The application also provides a computer storage medium carrying one or more computer programs, which can enable an electronic device to implement any Internet of Video Small Target Detection Method provided by the application when the one or more computer programs are executed by the electronic device.
[0207] In addition, it should be noted that the device embodiments described above are only schematic, and the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e. they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment. In addition, the connection relationship between the modules in the device embodiment provided by the application indicates that there is a communication connection between them, which can be implemented as one or more communication buses or signal lines.
[0208] Through the description of the above embodiments, those skilled in the art can clearly understand that the application can be realized by means of software and necessary general hardware, of course, it can also be realized by special hardware including special integrated circuit, special CPU, special memory, special components, etc. Generally, functions completed by computer programs can be easily realized by corresponding hardware, and specific hardware structures for realizing the same function can also be various, such as analog circuit, digital circuit or special circuit, etc. However, for the application, software program implementation is a better embodiment. Based on such understanding, the technical solutions of the application can be embodied in the form of software product, which is stored in a readable storage medium, such as a computer floppy disk, U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., including a plurality of instructions to make a computer device (which can be a personal computer, training device or network device, etc.) execute the methods described in various embodiments of the application.
[0209] In the embodiments described above, the entire or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, the entire or part of the embodiments can be implemented in the form of a computer program product.
[0210] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the entire or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general purpose computer, a special purpose computer, a computer network, or other programmable apparatus. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website site, computer, training device or data center to another website site, computer, training device or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer readable storage medium can be any available medium that can be stored by the computer or data storage device such as training device, data center, etc. integrated with one or more available media sets. The available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD), or a semiconductor medium (for example, solid state disk (SSD)) and the like.
Claims
1. A method for detecting small targets in an internet video, characterized in that, The method comprises the steps of: inputting the to-be-detected video into a global-local target detection network for global detection to obtain a global detection result; inputting the global detection result into a target dense area extraction network to obtain a target dense area, performing super-resolution processing on the target dense area to obtain a high-resolution local image; inputting the high-resolution local image into the global-local target detection network for fine detection to obtain a local detection result; inputting the global detection result and the local detection result into a progressive scale fusion network for fusion to obtain a high-precision small target detection result. 2.The method of claim 1, wherein, The method comprises the steps of: extracting image features from the to-be-detected video by using a classic backbone network fused with a cross-stage local network; mapping the image features to a recursive feature pyramid to obtain multi-scale image features; inputting the multi-scale image features into a region candidate network for target positioning to obtain target candidate region feature information; inputting the target candidate region feature information into a Transformer detection head for end-to-end set prediction and two-part matching to obtain the global detection result. 3.The method of claim 2, wherein, The method comprises the steps of: inputting the to-be-detected video into the classic backbone network to obtain low-level image features; inputting the low-level image features into the cross-stage local network, and dividing the low-level image features into first low-level feature images and second low-level feature images; propagating the first low-level feature images through a straight-through path to obtain straight-through path image features, and propagating the second low-level feature images through a complex path to obtain complex path image features; splicing the straight-through path image features and the complex path image features to obtain the image features including different gradient combinations. 4.The method of claim 2, wherein, The method comprises the steps of: determining initial candidate regions according to the multi-scale image features; calculating the intersection-over-union between the initial candidate regions and pre-configured real boxes; determining the initial candidate regions whose intersection-over-union reaches a preset overlap threshold as target candidate regions; performing accurate alignment feature pooling processing on the multi-scale image features to extract global prediction box information and position encoding information of the target candidate regions; and integrating the target candidate regions, the global prediction box information, and the position encoding information of the target candidate regions to obtain the target candidate region feature information. 5.The method of claim 1, wherein, The method comprises the steps of: calculating the aggregation score of each global detection box in the global detection result; if the aggregation score reaches a preset score threshold, determining the coordinates of the global detection box as effective coordinates; The effective coordinates are subjected to density clustering processing to obtain target dense sub-regions of different scales; The target dense sub-regions of different scales are subjected to scale adjustment to obtain standard target dense regions through cropping; The target dense regions are input into a super-lightweight super-resolution processing network to obtain the high-resolution local image. 6.The small target detection method of the video of things according to claim 1, characterized in that, The global detection result and the local detection result are input into a progressive scale fusion network for fusion to obtain a high-precision small target detection result, including: The global detection result and the local detection result are subjected to weighted box fusion to obtain the high-precision small target detection result.
7. A small target detection device for internet of video, characterized in that, including: a global detection unit configured to input a to-be-detected video into a global-local target detection network for global detection to obtain a global detection result; a super-resolution processing unit configured to input the global detection result into a target dense region extraction network to obtain a target dense region, and perform super-resolution processing on the target dense region to obtain a high-resolution local image; a local detection unit configured to input the high-resolution local image into the global-local target detection network for fine detection to obtain a local detection result; a fusion unit configured to input the global detection result and the local detection result into a progressive scale fusion network for fusion to obtain a high-precision small target detection result.
8. A computer program product, characterised in that, The computer readable instructions, when executed on an electronic device, cause the electronic device to implement the small target detection method for the video of the Internet of Things according to any one of claims 1 to 6. 9.A small target detection device of an internet of vehicles, characterized in that, The memory is configured to store computer programs; The processor is configured to execute the computer programs to enable the small target detection device for the video of the Internet of Things to implement the small target detection method for the video of the Internet of Things according to any one of claims 1 to 6. The storage medium carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the small target detection method for the video of the Internet of Things according to any one of claims 1 to 6.
10. A computer storage medium, characterized in that,
Citation Information
Cited By
Target detection method and system with hierarchical scanning perception fused with super-resolution enhancement
CN121921499A
Hierarchical scanning perception fusion super-resolution enhanced target detection method and system
CN121921499B