Intelligent Detection Method for Wide-Field-of-View High-Resolution Video Targets Based on Elastic Sparsification

By extracting multi-scale features in wide field of view high-resolution video detection, performing feature fusion and target density analysis, dynamic sparse processing, the problem of insufficient detection accuracy and calculation efficiency in the prior art is solved, and efficient and accurate target detection in different density scenarios is achieved.

CN120071226BActive Publication Date: 2025-07-22TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510551655.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-22
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The existing wide-field high-resolution video detection methods have insufficient adaptability and low processing efficiency in different densities scenarios, which leads to insufficient detection accuracy and calculation efficiency.

Method used

By extracting the distribution characteristics of video images at multiple image scales, performing feature fusion and target density analysis, dynamically estimate the degree of sparseness, adopting elastic sparseness strategy, generating compressed data for object detection and annotation, and using multi-scale feature extraction and fusion, fusion data compression and target detection and annotation, improving the adaptability and performance of the detection method.

Benefits of technology

High-accurate detection results can be obtained in high-density and low-density scenarios, which significantly improves the computing efficiency and detection accuracy, solves the problem of insufficient overall efficiency and accuracy in the existing technology, and achieves efficient and accurate target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071226B_ABST
    Figure CN120071226B_ABST
Patent Text Reader

Abstract

The present application discloses an intelligent detection method for wide-field high-resolution video targets based on elastic sparsification, as well as a training method, device, storage medium, equipment, and computer program product for an image intelligent detection model, including: in response to the input of a video image, determining the distribution characteristics of the video image at multiple image scales, and determining the fusion data of the video image according to the multiple distribution characteristics; determining the compressed data of the fusion data of the video image according to the target density of the target object in the video image determined from the fusion data and the target distribution characteristics in the multiple distribution characteristics; and determining the detection result of the image target in the video image according to the compressed data. Through multi-scale feature extraction and fusion, as well as target detection and annotation, the present application improves the adaptability and performance of the detection method in a changing environment, achieves efficient and accurate target detection in high-resolution imaging applications, and solves the problems of insufficient detection accuracy and computational efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of image annotation, and specifically relates to an intelligent detection method for wide-field high-resolution video targets based on elastic sparsification, as well as a training method, device, storage medium, equipment, and computer program product for an image intelligent detection model. Background Art

[0002] With the rapid development of high-resolution imaging technology, the application of wide-field high-resolution video has become increasingly widespread in many fields. These fields include urban surveillance, autonomous driving, smart city construction, astronomical observation, and drone inspection. In these application scenarios, video images usually have extremely high resolutions and wide field-of-view ranges, covering a large amount of details and target objects. Therefore, how to efficiently process and analyze these massive amounts of data while ensuring detection accuracy has become an important technical challenge.

[0003] Currently, there are already a variety of intelligent detection technologies for processing wide-field high-resolution video. Common methods include object detection methods based on a fixed sparsity ratio and efficient intelligent detection methods for wide-field high-resolution video. These technical means have achieved certain effects in specific application scenarios.

[0004] However, the object detection method based on a fixed sparsity ratio cannot adaptively process images in different density scenarios, thus reducing the detection accuracy; while the efficient intelligent detection method for wide-field high-resolution video has poor adaptability and performance in variable environments. In addition, existing methods also have problems of insufficient overall efficiency and accuracy. Summary of the Invention

[0005] This application aims to provide an intelligent detection method for wide-field high-resolution video targets based on elastic sparsification, as well as a training method, device, storage medium, equipment, and computer program product for an image intelligent detection model, which at least solves the problems of insufficient adaptation accuracy and low processing efficiency of existing image annotation schemes in different density scenarios.

[0006] In a first aspect, an embodiment of this application discloses an intelligent detection method for wide-field high-resolution video targets based on elastic sparsification, including:

[0007] In response to the input of a video image, determining the distribution characteristics of the video image at multiple image scales, and determining the fusion data of the video image according to the multiple distribution characteristics; the distribution characteristics are used to characterize the distribution of image targets in the video image at each image scale;

[0008] Determining the compressed data of the fusion data of the video image according to the target density of the target object in the video image determined from the fusion data and the target distribution characteristics among the multiple distribution characteristics;

[0009] Determine the detection result of the image target in the video image according to the compressed data; the detection result is used to label the image target in the video image.

[0010] In a second aspect, an embodiment of the present application also discloses an intelligent detection method for wide-field high-resolution video targets based on elastic sparsification, including:

[0011] In response to the input of the video image, determine the distribution characteristics of the video image at multiple image scales, and determine the fusion data of the video image according to the multiple distribution characteristics; the distribution characteristics are used to characterize the distribution of the image targets in the video image at each image scale;

[0012] Determine the compressed data of the fusion data of the video image according to the target density of the target object in the video image determined from the fusion data and the target distribution characteristics among the multiple distribution characteristics;

[0013] Determine the detection result of the image target in the video image according to the compressed data; the detection result is used to label the image target in the video image;

[0014] The determining the compressed data of the fusion data of the video image according to the target density of the target object in the video image determined from the fusion data and the target distribution characteristics among the multiple distribution characteristics includes:

[0015] Divide the fusion data into multiple image blocks, and generate data tokens corresponding to each image block;

[0016] Input each data token into a scoring model respectively to obtain an evaluation value of each data token; the evaluation value is used to characterize the target density of the target object in the data token;

[0017] Determine the distribution characteristic with the smallest corresponding image scale among the multiple distribution characteristics as the target distribution characteristic, and input the target distribution characteristic into the elastic selection sub-model in the image intelligent detection model to obtain the selection ratio of the multiple data tokens;

[0018] Determine multiple target data tokens from the token sequence of the video image according to the selection ratio, and generate the compressed data according to the determined multiple target data tokens; the token sequence is the sequence of the data tokens obtained after being arranged in descending order according to the evaluation value;

[0019] The determining the detection result of the image target in the video image according to the compressed data includes:

[0020] Input the compressed data into the annotation sub-model in the image intelligent detection model to annotate the image blocks in the video image;

[0021] Determine the detection result of the image target in the video image according to the deletion of the target image block in the image block; the target image block is the image block determined to be misannotated.

[0022] In a third aspect, an embodiment of the present application also discloses a training method for an image intelligent detection model, which is used to train the image intelligent detection model in the wide-field high-resolution video target intelligent detection method based on elastic sparsification as described in the second aspect, including:

[0023] Obtain training images for training, and input the training images into the image intelligent detection model to obtain the training selection ratio of the elastic selection sub-model of the image intelligent detection model for the training images, and the detection training result of the annotation sub-model of the image intelligent detection model for the training image targets in the training images;

[0024] Establish a density consistency loss function of the image intelligent detection model according to the training selection ratio, and establish a target detection loss function of the image intelligent detection model according to the training result;

[0025] Jointly train the image intelligent detection model through the density consistency loss function and the target detection loss function.

[0026] In a fourth aspect, an embodiment of the present application also discloses a wide-field high-resolution video target intelligent detection device based on elastic sparsification, including:

[0027] A fusion module, configured to respond to the input of a video image, determine the distribution characteristics of the video image at multiple image scales, and determine the fusion data of the video image according to the multiple distribution characteristics; the distribution characteristics are used to characterize the distribution of the image targets in the video image at each image scale;

[0028] A compression module, configured to determine the compressed data of the fusion data of the video image according to the target density of the target object in the video image determined from the fusion data and the target distribution characteristics in the multiple distribution characteristics;

[0029] A detection module, configured to determine the detection result of the image target in the video image according to the compressed data; the detection result is used to annotate the image target in the video image.

[0030] Fifth aspect, the embodiments of the present application also disclose a training device for an image intelligent detection model, which is used to train the image intelligent detection model in the intelligent detection method for wide-angle high-resolution video targets based on elastic sparsification as described in the second aspect, including:

[0031] An input module, configured to obtain training images for training, and input the training images into the image intelligent detection model, so as to obtain the training selection ratio of the elastic selection sub-model of the image intelligent detection model for the training images, and the detection training result of the annotation sub-model of the image intelligent detection model for the training image targets in the training images;

[0032] A loss module, configured to establish a density consistency loss function of the image intelligent detection model according to the training selection ratio, and establish an object detection loss function of the image intelligent detection model according to the training result;

[0033] A training module, configured to jointly train the image intelligent detection model through the density consistency loss function and the object detection loss function.

[0034] Sixth aspect, the embodiments of the present application also disclose a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps described in the first aspect or the second aspect or the third aspect are implemented.

[0035] Seventh aspect, the embodiments of the present application also disclose an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, the steps described in the first aspect or the second aspect or the third aspect are implemented.

[0036] Eighth aspect, the embodiments of the present application also disclose a computer program product, on which a computer program is stored, and when the computer program is executed by a processor, the steps described in the first aspect or the second aspect or the third aspect are implemented.

[0037] In summary, in the embodiments of the present application, by extracting the distribution features of video images at multiple image scales and enhancing the expression ability of multi-scale features through a feature fusion method, it is ensured that targets of different sizes in video images can be effectively captured and detected, the detection accuracy is improved, and it is applicable to the detection requirements of targets of different sizes. Furthermore, according to the target density of the target object in the video image determined from the fusion data and the target distribution features among the multiple distribution features, the compressed data of the fusion data of the video image is determined. By dynamically estimating the sparsity degree, the elastic sparsity strategy realizes flexible sparsification of the feature map, thereby while ensuring the target detection accuracy, significantly improving the detection speed, and solving the problem in the prior art that images in different density scenarios cannot be adaptively processed. Finally, according to the compressed data, the detection result of the image target in the video image is determined, and then the image target in the video image is labeled. Using the sparsified feature map for target detection significantly improves the computational efficiency and detection accuracy, ensuring high-accuracy detection results in both high-density and low-density scenarios, and solving the problem of insufficient overall efficiency and accuracy in the prior methods. Thus, based on the method of the embodiments of the present application, through multi-scale feature extraction and fusion, fusion data compression, and target detection and labeling, the problems of insufficient detection accuracy and computational efficiency in the prior art are solved, the adaptability and performance of the detection method in a changing environment are improved, and thus efficient and accurate target detection is achieved in high-resolution imaging applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0039] Figure 1 is a flowchart of the steps of an intelligent detection method for wide-field high-resolution video targets based on elastic sparsification provided by an embodiment of the present application;

[0040] Figure 2 is a flowchart of the steps of another intelligent detection method for wide-field high-resolution video targets based on elastic sparsification provided by an embodiment of the present application;

[0041] Figure 3 is a flowchart of the steps of a training method for an image intelligent detection model provided by an embodiment of the present application;

[0042] Figure 4 is a complete data flow process provided according to an embodiment of the present application;

[0043] Figure 5It is a schematic structural diagram of an intelligent detection device for wide-field high-resolution video targets based on elastic sparsification provided by an embodiment of the present application;

[0044] Figure 6 It is a schematic structural diagram of a training device for an image intelligent detection model provided by an embodiment of the present application;

[0045] Figure 7 It is a block diagram of an electronic device provided by an embodiment of the present application;

[0046] Figure 8 It is a block diagram of another electronic device provided by an embodiment of the present application. Detailed implementation manners

[0047] Hereinafter, exemplary embodiments of the present application will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be fully conveyed to those skilled in the art.

[0048] Figure 1 It is an intelligent detection method for wide-field high-resolution video targets based on elastic sparsification provided by this embodiment, specifically including the following steps:

[0049] Step 101, in response to the input of a video image, determine the distribution characteristics of the video image at multiple image scales, and determine the fusion data of the video image according to the multiple distribution characteristics.

[0050] Among them, the distribution characteristics are used to characterize the distribution of image targets in the video image at each image scale.

[0051] In some embodiments of the present application, in response to the input of a video image, determine the distribution characteristics of the video image at multiple image scales, and determine the fusion data of the video image according to the multiple distribution characteristics. This is because through multi-scale feature extraction, different-sized targets in the video image can be effectively captured and expressed, making the detection method applicable to the detection requirements of various-sized targets. To execute the corresponding steps, a high-resolution video image can be first input, sent to the backbone network (Backbone Network, BN), and multiple feature maps (Feature Map, FM) of different scales are generated through multi-layer convolution operations, and the feature maps of different scales are fused. The fusion data refers to the data integrating the feature information at multiple image scales and can more comprehensively characterize the target objects in the video image. After determining the distribution characteristics of the video image and performing feature fusion (Feature Fusion, FF), the accuracy of target detection is significantly improved.

[0052] In a specific example, a wide - field - of - view and high - resolution video image applied to an unmanned driving scenario is input into a detection system. First, the system extracts multi - scale feature maps through a backbone network and fuses these feature maps. Network structures such as the Residual Network (ResNet) or the Efficient Network (EffNet) can be selected for feature extraction. Feature maps at different scales have different representation capabilities in terms of spatial resolution and semantic levels. After feature fusion, the generated fused data can comprehensively capture the target information in the image, enabling the detection system to accurately identify and locate targets of different sizes, and ultimately realizing more reliable target detection and navigation assistance functions during unmanned driving.

[0053] Step 102: Determine the compressed data of the fused data of the video image according to the target density of the target object in the video image determined from the fused data and the target distribution feature among multiple distribution features.

[0054] In some embodiments of the present application, determining the compressed data of the fused data of the video image according to the target density of the target object in the video image determined from the fused data and the target distribution feature among multiple distribution features is because by analyzing the target density of the target object and the distribution features in the image, the sparsification of the video image data can be effectively achieved, thereby optimizing the utilization efficiency of computing resources. To execute the corresponding step, the target density information of the target object can be first extracted from the fused data, and then combined with the target distribution feature determined among multiple distribution features to estimate the degree of sparsification, and the compressed data is generated accordingly. The target density refers to the distribution density of the target object in the video image, and the multiple distribution features refer to the target distribution situations in the image at different scales. Through this process, the compressed data can significantly reduce redundant information while retaining key feature information, improving the detection efficiency.

[0055] In a specific example, in the scenario of urban monitoring, the system extracts the target density information of pedestrians and vehicles from the fused data and combines the target distribution features of the image at different scales. The experimenter inputs this information into a sparsification algorithm to estimate the optimal degree of sparsification. The sparsification algorithm generates compressed data according to the target density and distribution features. This compressed data contains the key target object information in the urban monitoring video without redundant data. Finally, the compressed data is used for target detection to ensure that pedestrians and vehicles in different density scenarios can be accurately identified and tracked during urban monitoring.

[0056] Step 103: Determine the detection result of the image target in the video image according to the compressed data.

[0057] Among them, the detection results are used to label the image targets in the video image.

[0058] In some embodiments of the present application, the detection results of the image targets in the video image are determined according to the compressed data. This is because by using the compressed data for object detection, the computational efficiency can be significantly improved, and high-accuracy detection results can be obtained in both high-density and low-density scenarios. To execute the corresponding steps, the compressed data can be input into the image intelligent detection model, and through the processing of the model, the detection results of the image targets are generated. These detection results include the location information and classification labels of the target objects, which are used to label the image targets in the video image. The image intelligent detection model is a trained deep learning model that can perform efficient and accurate object detection on the sparsified feature map. After generating the detection results, the detection system can effectively label various target objects in the video image, improving the detection accuracy and efficiency.

[0059] In a specific example, in the scenario of drone patrol, the system inputs the compressed data after sparsification processing into a pre-trained image intelligent detection model. The experimenter selects a model structure of a Convolutional Neural Network (CNN) for processing, and the model generates detection results according to the compressed data. The detection results include the location information and classification labels of various infrastructures in the drone patrol video, such as wires, towers, buildings, etc. According to these detection results, the system labels the infrastructures in the video image, enabling the drone patrol task to quickly identify and locate the target objects, improving the patrol efficiency and accuracy.

[0060] In summary, in the embodiments of the present application, by extracting the distribution features of video images at multiple image scales and enhancing the expression ability of multi-scale features through a feature fusion method, it is ensured that targets of different sizes in video images can be effectively captured and detected, improving the detection accuracy and making it applicable to the detection requirements of targets of different sizes. Furthermore, according to the target density of the target object in the video image determined from the fused data and the target distribution features among the multiple distribution features, the compressed data of the fused data of the video image is determined. By dynamically estimating the sparsity degree, the elastic sparsity strategy realizes flexible sparsification of the feature map, thereby significantly improving the detection speed while ensuring the target detection accuracy, and solving the problem in the prior art that images in different density scenarios cannot be adaptively processed. Finally, according to the compressed data, the detection result of the image target in the video image is determined, and then the image target in the video image is labeled. Using the sparsified feature map for target detection significantly improves the calculation efficiency and detection accuracy, ensuring high-accuracy detection results in both high-density and low-density scenarios, and solving the problem of insufficient overall efficiency and accuracy in the prior methods. Thus, based on the method of the embodiments of the present application, through multi-scale feature extraction and fusion, fused data compression, and target detection and annotation, the problems of insufficient detection accuracy and calculation efficiency in the prior art are solved, improving the adaptability and performance of the detection method in a changing environment, and thus achieving efficient and accurate target detection in high-resolution imaging applications.

[0061] Figure 2 Another intelligent detection method for wide-field high-resolution video targets based on elastic sparsification provided by this embodiment specifically includes the following steps:

[0062] Step 201, in response to the input of a video image, determine the distribution features of the video image at multiple image scales, and determine the fused data of the video image according to the multiple distribution features.

[0063] Among them, the distribution features are used to characterize the distribution of image targets in the video image at each image scale.

[0064] The method shown in this step has been described in step 101 and will not be elaborated here.

[0065] Optionally, step 201 includes the following sub-steps:

[0066] Sub-step 2011, in response to the input of a video image, extract multiple feature maps from the video image.

[0067] Among them, each feature map has a different image scale.

[0068] In some embodiments of the present application, in response to the input of a video image, multiple feature maps are extracted from the video image. This is because by extracting multiple feature maps, feature information of different sizes and levels in the video image can be captured, providing rich feature representations for subsequent feature fusion and object detection. To perform the corresponding steps, the input high-resolution video image can be fed into a backbone network, and multiple feature maps with different image scales are generated through multi-layer convolution operations. Feature maps are image feature representations generated through multi-layer convolution processing in a convolutional neural network. By extracting multiple feature maps, the target objects in the video image can be more comprehensively characterized, thereby improving the accuracy of object detection.

[0069] In a specific example, in the scenario of autonomous driving, the system receives a high-resolution video image and inputs it into the backbone network. For example, the ResNet network structure can be selected, and multiple feature maps of different scales are generated through multi-layer convolution operations. These feature maps respectively represent low-level and high-level feature information in the video image. By extracting these feature maps, the system can capture information about target objects of different sizes and levels, providing rich feature representations for subsequent feature fusion and object detection.

[0070] Sub-step 2012: Fuse the determined multiple feature maps to obtain fused data.

[0071] In some embodiments of the present application, the determined multiple feature maps are fused to obtain fused data. This is because by fusing multiple feature maps, feature information of different scales and levels can be effectively integrated, improving the accuracy and robustness of object detection. To perform the corresponding steps, multiple feature maps in the video image can be extracted first, and these feature maps have different representation capabilities in terms of spatial resolution and semantic level; then, a feature fusion method (such as a Feature Pyramid Network (FPN) or an adaptive feature fusion mechanism) is used to fuse these feature maps to generate fused data. The fused data integrates the information of multiple feature maps and can more comprehensively characterize the target objects in the video image. Through this process, it can be ensured that the object detection method can capture and recognize target objects of various sizes and levels, improving the detection accuracy.

[0072] In a specific example, in the scenario of drone inspection, the system extracts multiple feature maps from a high-resolution video image, and these feature maps respectively represent feature information of different spatial resolutions and semantic levels. The experimenter uses the FPN method to fuse these feature maps. Through feature fusion, fused data is generated, and the fused data can comprehensively reflect various target objects in the image.

[0073] Step 202: Divide the fused data into multiple image patches and generate data tokens corresponding to each image patch.

[0074] In some embodiments of the present application, dividing the fused data into multiple image patches and generating data tokens corresponding to each image patch is because by dividing the fused data into multiple smaller image patches, the target objects in the image can be processed and analyzed more efficiently, improving the detection accuracy. To perform the corresponding steps, the fused data can first be divided into several image patches according to certain rules, and then data tokens corresponding to each image patch are generated. An image patch refers to a smaller image area divided according to a certain size, and a data token is a high-dimensional vector representing the feature information of these image patches. After generating the data tokens, subsequent object detection and annotation work can be carried out to achieve accurate recognition of the target objects in the video image.

[0075] In a specific example, in the scenario of smart city construction, the system divides the fused data into multiple image patches (for example, it can be divided according to a size of 7×7 pixels). A fixed-size grid division method is selected, and data tokens corresponding to each image patch are generated according to the feature information of each image patch. For example, each data token contains feature information such as the color, texture, and shape of the image patch. By dividing the fused data into image patches and generating data tokens, the system can process and analyze the pedestrian and vehicle targets in the urban surveillance video more efficiently, and finally achieve accurate detection and annotation of different target objects.

[0076] Step 203: Input each data token into a scoring model respectively to obtain an evaluation value for each data token.

[0077] Among them, the evaluation value is used to characterize the target density of the target object in the data token.

[0078] In some embodiments of the present application, inputting each data token into a scoring model respectively to obtain an evaluation value for each data token is because by scoring each data token, the density information of the target object contained in the data token can be quantified, thereby providing a basis for subsequent sparsification strategies. To perform the corresponding steps, each data token can be input into a specially designed scoring model. The model processes the input data token through several linear layers or convolutional layers and outputs a scalar evaluation value. The scoring model is a neural network model that can perform feature extraction and evaluation on the input data token. Through this process, the evaluation value of each data token reflects the density information of the target object it contains, thereby guiding the selection and retention of important data tokens in subsequent steps.

[0079] In a specific example, in the scenario of astronomical observation, the system inputs the data tokens generated from the image patches into the scoring model respectively. For example, a scoring network consisting of two fully connected layers can be designed, and an activation function, such as the Rectified Linear Unit (ReLU), is used between each layer. The scoring model processes each input data token and outputs an evaluation value. The evaluation value represents the density information of the celestial bodies in the corresponding image patch. Through this scoring process, the system can determine the importance of each data token, thereby providing a basis for the selection of key image patches in subsequent steps and ultimately achieving precise detection of celestial bodies in astronomical images.

[0080] Step 204: Among the multiple distribution features, determine the distribution feature with the smallest corresponding image scale as the target distribution feature, and input the target distribution feature into the elastic selection sub-model in the image intelligent detection model to obtain the selection ratio for multiple data tokens.

[0081] In some embodiments of the present application, determining the distribution feature with the smallest corresponding image scale among the multiple distribution features as the target distribution feature and inputting the target distribution feature into the elastic selection sub-model in the image intelligent detection model to obtain the selection ratio for multiple data tokens is because by selecting the distribution feature with the smallest image scale, the density information of the target object in the image can be more accurately reflected, providing a basis for the elastic sparsification strategy. To execute the corresponding step, first, among the multiple distribution features, find the distribution feature with the smallest corresponding image scale and determine it as the target distribution feature, and then input the target distribution feature into the elastic selection sub-model in the image intelligent detection model to obtain the selection ratio for multiple data tokens. The target distribution feature refers to the feature information that can best represent the distribution of the image target object at different scales. Through this process, the optimal sparsification ratio in different scenarios can be more accurately estimated, optimizing the utilization of computing resources and improving the detection efficiency and accuracy.

[0082] In a specific example, in the scenario of autonomous driving, the system obtains distribution features at multiple scales, including the distribution of pedestrians, vehicles, and obstacles. Find the distribution feature with the smallest corresponding image scale, determine it as the target distribution feature, and input it into the elastic selection sub-model. The elastic selection sub-model calculates the selection ratio for multiple data tokens according to the target distribution feature. The selection ratio reflects the sparsification requirements of different target objects in the current scenario.

[0083] Step 205: Determine multiple target data tokens from the token sequence of the video image according to the selection ratio, and generate compressed data based on the determined multiple target data tokens.

[0084] Among them, the token sequence is the sequence of data tokens obtained after arranging the evaluation values from largest to smallest.

[0085] In some embodiments of the present application, multiple target data tokens are determined from the token sequence of the video image according to the selection ratio, and compressed data is generated based on the determined multiple target data tokens. This is because by selecting important target data tokens according to the selection ratio, key feature information can be retained, redundant data can be reduced, and the detection efficiency can be improved. To perform the corresponding steps, the evaluation values can first be arranged from largest to smallest to form a token sequence; then, multiple target data tokens are determined from the token sequence according to the selection ratio, and compressed data is generated based on these target data tokens. The token sequence refers to the sequence containing target object information obtained after sorting according to the evaluation values. Through this process, the compressed data can effectively reflect the key feature information in the video image, thereby improving the accuracy and efficiency of target detection.

[0086] In a specific example, in the scenario of drone inspection, the system determines multiple target data tokens from the token sequence of the video image according to the selection ratio. The experimenter first arranges the evaluation values of the data tokens from largest to smallest to form a token sequence. Then, according to the selection ratio calculated by the elastic selection sub-model, the top ρ% of the data tokens are selected from the token sequence as the target data tokens. Next, the experimenter generates compressed data based on these target data tokens, retaining the key target object information in the video image.

[0087] Step 206, input the compressed data into the annotation sub-model in the image intelligent detection model to annotate the image blocks in the video image.

[0088] In some embodiments of the present application, the compressed data is input into the annotation sub-model in the image intelligent detection model to annotate the image blocks in the video image. This is because by using the compressed data for annotation, the efficiency and accuracy of the detection system in annotating the target objects in the video image can be improved. To perform the corresponding steps, the generated compressed data can be input into the annotation sub-model in the image intelligent detection model, and the annotation sub-model processes the input compressed data to generate the annotation results of the image blocks. The annotation sub-model is a trained deep learning model that can perform efficient and accurate image block annotation based on the compressed data. After generating the annotation results, the system can effectively annotate various target objects in the video image, improving the detection accuracy.

[0089] In a specific example, in an unmanned driving scenario, the system inputs the compressed data generated according to the sparsification strategy into the annotation sub-model in the image intelligent detection model. The experimenter selects an annotation sub-model based on a convolutional neural network for processing. The annotation sub-model extracts and analyzes the features of the input compressed data to generate the annotation results of the image blocks. The annotation results include the position information and classification labels of various target objects (such as pedestrians, vehicles, and obstacles) in the unmanned driving video. Through this annotation process, the system can accurately annotate the target objects in the video image, improving the detection accuracy and operation efficiency of the unmanned driving system in complex environments.

[0090] Step 207, determine the detection result of the image target in the video image according to the deletion of the target image block in the image block.

[0091] Among them, the target image block is the image block determined to be mis-annotated.

[0092] In some embodiments of the present application, the detection result of the image target in the video image is determined according to the deletion of the target image block in the image block. This is because by deleting the mis-annotated target image blocks, the false detections can be effectively reduced, improving the accuracy of the detection result. To execute the corresponding steps, first, through the previous annotation steps, identify and annotate the target image blocks in the video image. Then, verify these target image blocks to determine which are the mis-annotated image blocks. Finally, perform the deletion operation according to the mis-annotated image blocks. The target image block refers to the image area detected and annotated in the video image. By deleting the mis-annotated image blocks, it can be ensured that the finally generated detection result has high accuracy and reliability.

[0093] In a specific example, in an unmanned driving scenario, the system has detected and annotated pedestrians and vehicles in the video image through the previous steps. The annotator verifies the annotation results and finds that there are mis-annotations of the target objects in some image blocks. According to these verification results, the system performs a deletion operation on the mis-annotated target image blocks to remove the mis-detected pedestrian and vehicle information. By deleting the mis-annotated image blocks, the system finally generates accurate detection results, ensuring that the detection and positioning of pedestrians and vehicles are more reliable during unmanned driving, improving the safety and performance of the unmanned driving system.

[0094] In summary, in the embodiments of the present application, by extracting the distribution features of video images at multiple image scales and enhancing the expression ability of multi-scale features through a feature fusion method, it is ensured that targets of different sizes in video images can be effectively captured and detected, the detection accuracy is improved, and it is applicable to the detection requirements of targets of different sizes; furthermore, according to the target density of the target object in the video image determined from the fusion data and the target distribution features in multiple distribution features, the compressed data of the fusion data of the video image is determined. By dynamically estimating the sparsity degree, the elastic sparsity strategy realizes flexible sparsification of the feature map, thereby while ensuring the target detection accuracy, significantly improving the detection speed and solving the problem in the prior art that images in different density scenarios cannot be adaptively processed; finally, according to the compressed data, the detection result of the image target in the video image is determined, and then the image target in the video image is labeled; using the sparsified feature map for target detection significantly improves the calculation efficiency and detection accuracy, ensuring high-accuracy detection results in both high-density and low-density scenarios and solving the problem of insufficient overall efficiency and accuracy in the prior methods. Thus, based on the method of the embodiments of the present application, through multi-scale feature extraction and fusion, fusion data compression, and target detection and labeling, the problems of insufficient detection accuracy and calculation efficiency in the prior art are solved, the adaptability and performance of the detection method in a changing environment are improved, and thus efficient and accurate target detection is achieved in high-resolution imaging applications.

[0095] As Figure 3 shown, the embodiments of the present application further provide a training method for an image intelligent detection model, which is used to train the image intelligent detection model in the wide-field high-resolution video target intelligent detection method based on elastic sparsification mentioned in the above embodiments, and specifically includes the following steps:

[0096] Step 301, obtain training images for training, and input the training images into the image intelligent detection model to obtain the training selection ratio of the training images by the elastic selection sub-model of the image intelligent detection model, and the detection training results of the training image targets in the training images by the annotation sub-model of the image intelligent detection model.

[0097] In some embodiments of the present application, training images for training are obtained and input into an image intelligent detection model to obtain the training selection ratio of the elastic selection sub-model of the image intelligent detection model for the training images, and the detection training results of the annotation sub-model of the image intelligent detection model for the training image targets in the training images. This is because by training the image intelligent detection model using the training images, the model can learn and optimize its detection and annotation capabilities for target objects in different scenarios. To perform the corresponding steps, training images for training can be obtained first. The training images can be a labeled high-resolution image dataset. Then, the training images are input into the image intelligent detection model, and through the processing of the model, the training selection ratio of the elastic selection sub-model and the detection training results of the annotation sub-model are obtained respectively. Training images refer to the image data used for model training. The elastic selection sub-model and the annotation sub-model are two sub-modules in the image intelligent detection model, which are used to select important features and generate annotation results respectively. Through this process, the performance of the model can be optimized, and the accuracy of target detection and annotation can be improved.

[0098] In a specific example, in the scenario of autonomous driving, the system obtains a set of high-resolution road images for training as the training images. These images contain road scenes under different weather, lighting, and traffic conditions. The training images are input into the image intelligent detection model, and through the processing of the model, the training selection ratio of the elastic selection sub-model for each training image and the detection training results of the annotation sub-model for vehicles, pedestrians, and obstacles in each training image are obtained. Through this training process, the model can learn and optimize its detection and annotation capabilities for target objects in different road scenes, and ultimately improve the detection accuracy and safety performance of the autonomous driving system.

[0099] Step 302, establish a density consistency loss function of the image intelligent detection model according to the training selection ratio, and establish a target detection loss function of the image intelligent detection model according to the training results.

[0100] In some embodiments of the present application, a density consistency loss function of the image intelligent detection model is established according to the training selection ratio, and an object detection loss function of the image intelligent detection model is established according to the training results. This is because by establishing the density consistency loss function and the object detection loss function, the image intelligent detection model can be trained and optimized more effectively, thereby improving the detection accuracy and adaptability of the model. To execute the corresponding steps, the density consistency loss function can be designed first according to the training selection ratio to ensure that the model can accurately reflect the target density when selecting data tokens; then, according to the training results, the object detection loss function is designed to evaluate the detection effect of the model on the target object. The density consistency loss function and the object detection loss function are two key indicators for optimizing the model performance. Through this process, the image intelligent detection model can have higher detection capabilities and generalization performance in different scenarios.

[0101] In a specific example, in the scenario of autonomous driving, the system designs a density consistency loss function according to the training selection ratio obtained in the previous step to ensure that the model can accurately select target data tokens when dealing with different density scenarios. The experimenter designs an object detection loss function according to the actually detected vehicle and pedestrian targets in the training data to evaluate the detection accuracy of the model. Through this step, the system effectively trains and optimizes the image intelligent detection model, enabling it to accurately identify and detect target objects in different road and traffic environments, and ultimately improving the overall performance and safety of the autonomous driving system.

[0102] Step 303, jointly train the image intelligent detection model through the density consistency loss function and the object detection loss function.

[0103] In some embodiments of the present application, the image intelligent detection model is jointly trained through the density consistency loss function and the object detection loss function. This is because through joint training, the performance of the model in terms of sparsification selection and object detection can be optimized simultaneously, improving the detection accuracy and adaptability of the model in different scenarios. To execute the corresponding steps, the model can be first input into the training data, and the density consistency loss and the object detection loss are calculated; then, through the backpropagation algorithm, the model parameters are adjusted and optimized to minimize these two loss functions. The density consistency loss function is used to ensure that the data tokens selected by the model are consistent with the actual target density, and the object detection loss function is used to evaluate the object detection effect of the model. Through this joint training process, the model can better coordinate the sparsification strategy and the detection accuracy, improving the overall performance.

[0104] In a specific example, in the scenario of smart city monitoring, the system inputs high-resolution video image data with annotation information into the image intelligent detection model. The experimenter first calculates the density consistency loss and the object detection loss of the training data, and then adjusts the parameters of the model through the backpropagation algorithm to optimize these two loss functions. The density consistency loss function ensures that the model can accurately select the data tokens containing important information during the sparsification process, while the object detection loss function ensures that the model can achieve high accuracy during the annotation and detection processes. Through joint training, the model can accurately detect and annotate target objects such as pedestrians and vehicles in smart city monitoring, effectively improving the detection performance and adaptability of the system.

[0105] Optionally, the density consistency loss function is:

[0106] ,

[0107] where A union is the area covered by the boundaries output by the codec, H and W are the height and width of the input image respectively, ρ is the sparsification ratio predicted by the elastic selection module; λ cons is a manually set parameter.

[0108] Through this formula, the consistency between the sparsification selection ratio and the actual target density can be ensured. The advantage of the density consistency loss function is that by measuring the deviation between the sparsification selection ratio and the target density, the selection module of the image intelligent detection model can be effectively optimized, enabling it to have higher accuracy and robustness when dealing with different density scenarios. This formula ensures that the selected data tokens can accurately reflect the actual target density in the input image, thereby improving the accuracy and efficiency of object detection.

[0109] Optionally, the object detection loss function is:

[0110] ,

[0111] where, are the four boundary components of the predicted bounding box, are the four boundary components of the bounding box obtained by manual annotation, N represents the number of predicted bounding boxes, represents the predicted class probability, represents the class obtained by annotation.

[0112] For each predicted bounding box, calculate the cross-entropy loss between the predicted class probability and the class obtained by annotation:

[0113] ,

[0114] Among them, y i is the true category of the i-th object, and is the predicted category probability.

[0115] The above formula can better handle small and large errors by using smooth L1 loss, ensuring the stability of the regression loss. In addition, by evaluating the classification loss through cross-entropy loss, the accuracy of the model for class classification can be improved. By combining the regression loss and the classification loss, this formula can optimize the performance of the image intelligent detection model and improve the accuracy and robustness of object detection.

[0116] As Figure 4 shown, it is a complete data flow process under the method of the embodiment of the present application, specifically including the following processes:

[0117] R1: The video image is input into the system as the starting point of the entire processing flow. The main purpose of this image input process is to obtain high-resolution video image data for subsequent steps to process and analyze;

[0118] R2: The image is processed by the backbone network, and feature maps of different scales and levels are extracted through multiple convolutional layers. In this step, the system generates multiple feature maps (such as S1, S2, S3, S4), and these feature maps can respectively represent different feature information in the video image, including low-level features and high-level features;

[0119] R3: The system selects elastic Tokens according to density perception. The steps include multi-scale feature extraction, feature ranking, and scoring processing, and appropriate Tokens are selected through the elastic selection module. The purpose of this step is to adaptively select key features according to the density of the target object and optimize the utilization of computing resources;

[0120] R4: The selected Tokens enter the encoder and decoder for processing to generate the bounding box position information (such as 1254, 230 shown in the figure) and category information (such as "boy") of the target object. This step also involves the calculation of density and loss function for optimizing the accuracy of the detection result;

[0121] R5: The system annotates the target objects in the video image. The annotated image is used to verify and improve the detection performance of the model to ensure that the system can accurately identify and annotate various target objects.

[0122] Through these steps, the entire data flow process from image input to feature extraction, Token selection, encoding and decoding processing, and manual annotation comprehensively covers all aspects of image processing and object detection.

[0123] In summary, in the embodiments of the present application, by extracting the distribution features of video images at multiple image scales and enhancing the expression ability of multi-scale features through a feature fusion method, it is ensured that different-sized targets in the video images can be effectively captured and detected, improving the detection accuracy and making it applicable to the detection requirements of different-sized targets. Furthermore, according to the target density of the target object in the video image determined from the fusion data and the target distribution features among the multiple distribution features, the compressed data of the fusion data of the video image is determined. By dynamically estimating the sparsity degree, the elastic sparsity strategy realizes flexible sparsification of the feature map, thereby significantly improving the detection speed while ensuring the target detection accuracy, and solving the problem in the prior art that images in different density scenarios cannot be adaptively processed. Finally, according to the compressed data, the detection result of the image target in the video image is determined, and then the image target in the video image is labeled. Using the sparsified feature map for target detection significantly improves the calculation efficiency and detection accuracy, ensuring high-accuracy detection results in both high-density and low-density scenarios, and solving the problem of insufficient overall efficiency and accuracy in the prior methods. Thus, based on the method of the embodiments of the present application, through multi-scale feature extraction and fusion, fusion data compression, and target detection and labeling, the problems of insufficient detection accuracy and calculation efficiency in the prior art are solved, improving the adaptability and performance of the detection method in a changing environment, and thus achieving efficient and accurate target detection in high-resolution imaging applications.

[0124] As Figure 5 shown, the embodiments of the present application also disclose a wide-field high-resolution video target intelligent detection device 40 based on elastic sparsification, including:

[0125] A fusion module 401, configured to determine the distribution features of a video image at multiple image scales in response to the input of the video image, and determine the fusion data of the video image according to the multiple distribution features; the distribution features are used to characterize the distribution of the image targets in the video image at each image scale;

[0126] A compression module 402, configured to determine the compressed data of the fusion data of the video image according to the target density of the target object in the video image determined from the fusion data and the target distribution features among the multiple distribution features;

[0127] A detection module 403, configured to determine the detection result of the image target in the video image according to the compressed data; the detection result is used to label the image target in the video image.

[0128] Optionally, the fusion module 401 includes:

[0129] An extraction sub-module, configured to extract a plurality of feature maps from the video image in response to the input of the video image; each feature map has a different image scale;

[0130] A fusion sub-module, configured to fuse a plurality of determined feature maps to obtain fused data.

[0131] Optionally, the compression module 402 includes:

[0132] A token sub-module, configured to divide the fused data into a plurality of image blocks and generate data tokens corresponding to each image block;

[0133] An evaluation sub-module, configured to input each data token into a scoring model respectively to obtain an evaluation value of each data token; the evaluation value is used to characterize the target density of the target object in the data token;

[0134] A ratio sub-module, configured to determine the target distribution feature as the distribution feature with the smallest corresponding image scale among a plurality of distribution features, and input the target distribution feature into an elastic selection sub-model in the image intelligent detection model to obtain a selection ratio for a plurality of data tokens;

[0135] A fusion sub-module, configured to determine a plurality of target data tokens from a token sequence of a video image according to the selection ratio, and generate compressed data according to the determined plurality of target data tokens; the token sequence is a sequence of data tokens obtained after being sorted in descending order of the evaluation value.

[0136] Optionally, the detection module 403 includes:

[0137] A labeling sub-module, configured to input the compressed data into a labeling sub-model in the image intelligent detection model to label image blocks in the video image;

[0138] A detection sub-module, configured to determine a detection result of an image target in the video image according to the deletion of a target image block in the image block; the target image block is an image block determined to be mislabeled.

[0139] As Figure 6 shown, an embodiment of the present application also discloses a training device 50 for an image intelligent detection model, configured to train the image intelligent detection model in the wide-field high-resolution video target intelligent detection method based on elastic sparsification mentioned in the above embodiment, including:

[0140] An input module 501, configured to obtain training images for training, and input the training images into the image intelligent detection model to obtain a training selection ratio of the elastic selection sub-model of the image intelligent detection model for the training images, and a detection training result of the labeling sub-model of the image intelligent detection model for training image targets in the training images;

[0141] A loss module 502, configured to establish a density consistency loss function of the image intelligent detection model according to a training selection ratio, and establish an object detection loss function of the image intelligent detection model according to a training result;

[0142] A training module 503, configured to jointly train the image intelligent detection model by using the density consistency loss function and the object detection loss function.

[0143] In summary, in the embodiment of the present application, by extracting the distribution features of video images at multiple image scales and enhancing the expression ability of multi-scale features through a feature fusion method, it is ensured that targets of different sizes in video images can be effectively captured and detected, improving the detection accuracy and making it applicable to the detection requirements of targets of different sizes; furthermore, according to the target density of the target object in the video image determined from the fused data and the target distribution features in multiple distribution features, the compressed data of the fused data of the video image is determined, and by dynamically estimating the sparsity degree, the elastic sparsity strategy realizes flexible sparsification of the feature map, thereby while ensuring the target detection accuracy, significantly improving the detection speed, and solving the problem in the prior art that images with different density scenarios cannot be adaptively processed; finally, according to the compressed data, the detection result of the image target in the video image is determined, and then the image target in the video image is labeled; using the sparsified feature map for target detection significantly improves the calculation efficiency and detection accuracy, ensuring high-accuracy detection results in both high-density and low-density scenarios, and solving the problem of insufficient overall efficiency and accuracy in the prior methods. Thus, based on the method of the embodiment of the present application, through multi-scale feature extraction and fusion, fused data compression, and target detection and labeling, the problems of insufficient detection accuracy and calculation efficiency in the prior art are solved, improving the adaptability and performance of the detection method in a changing environment, and thus achieving efficient and accurate target detection in high-resolution imaging applications.

[0144] The embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above-mentioned embodiment of the method for intelligent detection of wide-field high-resolution video targets based on elastic sparsification or the training method of the image intelligent detection model, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0145] Figure 7It is a block diagram of an electronic device 700 provided by an embodiment of the present application. For example, the electronic device 700 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0146] Referring to Figure 7 , the electronic device 700 may include one or more of the following components: a processing component 702, a memory 704, a power component 706, a multimedia component 708, an audio component 710, an input / output (I / O) interface 712, a sensor component 714, and a communication component 716.

[0147] The processing component 702 generally controls the overall operation of the electronic device 700, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the above-mentioned wide-field high-resolution video target intelligent detection method based on elastic sparsification or the training method of the image intelligent detection model. In addition, the processing component 702 may include one or more modules to facilitate the interaction between the processing component 702 and other components. For example, the processing component 702 may include a multimedia module to facilitate the interaction between the multimedia component 708 and the processing component 702.

[0148] The memory 704 is used to store various types of data to support the operation of the electronic device 700. Examples of these data include instructions for any application or method operating on the electronic device 700, contact data, phone book data, messages, pictures, multimedia, etc. The memory 704 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0149] The power component 706 provides power to various components of the electronic device 700. The power component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 700.

[0150] The multimedia component 708 includes a screen that provides an output interface between the electronic device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 708 includes a front camera and / or a rear camera. When the electronic device 700 is in an operating mode, such as a shooting mode or a multimedia mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0151] The audio component 710 is used to output and / or input audio signals. For example, the audio component 710 includes a microphone (MIC) that is used to receive external audio signals when the electronic device 700 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 704 or transmitted via the communication component 716. In some embodiments, the audio component 710 further includes a speaker for outputting audio signals.

[0152] The I / O interface 712 provides an interface between the processing component 702 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.

[0153] The sensor component 714 includes one or more sensors for providing status assessments of various aspects of the electronic device 700. For example, the sensor component 714 can detect the on / off state of the electronic device 700, the relative positioning of components, such as the display and the keypad of the electronic device 700. The sensor component 714 can also detect a change in the position of the electronic device 700 or a component of the electronic device 700, the presence or absence of user contact with the electronic device 700, the orientation or acceleration / deceleration of the electronic device 700, and the temperature change of the electronic device 700. The sensor component 714 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 714 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 714 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0154] The communication component 716 is used to facilitate communication between the electronic device 700 and other devices in a wired or wireless manner. The electronic device 700 can access a communication standard-based wireless network, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 7G), or a combination thereof. In an exemplary embodiment, the communication component 716 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 716 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0155] In an exemplary embodiment, the electronic device 700 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for implementing the intelligent detection method for wide-field high-resolution video targets based on elastic sparsification or the training method for the image intelligent detection model provided in the embodiments of the present application.

[0156] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions. The above instructions can be executed by the processor 720 of the electronic device 700 to complete the above-mentioned intelligent detection method for wide-field high-resolution video targets based on elastic sparsification or the training method for the image intelligent detection model. For example, the non-transitory storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0157] Figure 8 is a block diagram of an electronic device 800 shown according to an exemplary embodiment. For example, the electronic device 800 can be provided as a server. Referring to Figure 8 , the electronic device 800 includes a processing component 822, which further includes one or more processors, and memory resources represented by a memory 832 for storing instructions executable by the processing component 822, such as application programs. The application programs stored in the memory 832 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 822 is configured to execute instructions to perform the intelligent detection method for wide-field high-resolution video targets based on elastic sparsification or the training method for the image intelligent detection model provided in the embodiments of the present application.

[0158] The electronic device 800 may further include a power supply component 826 configured to perform power management of the electronic device 800, a wired or wireless network interface 850 configured to connect the electronic device 800 to a network, and an input / output (I / O) interface 858. The electronic device 800 may operate based on an operating system stored in the memory 832, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSD TM or the like.

[0159] The embodiments of the present application also provide a computer program product, including a computer program, which when executed by a processor, implements a method for intelligent detection of wide-field high-resolution video targets based on elastic sparsification or a method for training an image intelligent detection model.

[0160] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the application disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are to be considered as exemplary only, and the true scope and spirit of the present application are pointed out by the following claims.

[0161] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

[0162] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.

[0163] It is easy for those skilled in the art to think that any combination application of the above various embodiments is feasible. Therefore, any combination among the above various embodiments is an embodiment of the present application. However, due to space limitations, this specification will not elaborate on each of them here.

[0164] The method for intelligent detection of wide-field high-resolution video targets based on elastic sparsification or the method for training an image intelligent detection model provided herein is not inherently related to any specific computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. Based on the above description, it is obvious to construct the required structure of a system with the solution of the present application. In addition, the present application is not directed to any specific programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the description of a specific language above is to disclose the best implementation mode of the present application.

[0165] In the specification provided herein, a number of specific details are set forth. However, it will be understood that embodiments of the present application may be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been shown in detail so as not to obscure an understanding of this description.

[0166] Similarly, it should be understood that, in order to streamline the present application and assist in understanding one or more of the various aspects thereof, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together in a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed present application requires more features than are expressly recited in each claim. Rather, as reflected by the claims, the aspects of the application lie in less than all the features of the single embodiments disclosed previously. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, where each claim stands on its own as a separate embodiment of the present application.

[0167] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract and drawings) can be replaced by an alternative feature that provides the same, equivalent or similar purpose.

[0168] In addition, those skilled in the art will be able to understand that, although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of the present application and forms different embodiments. For example, in the claims, any one of the claimed embodiments can be used in any combination.

[0169] Each component embodiment of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the method for intelligent detection of wide-field high-resolution video targets based on elastic sparsification or the method for training an image intelligent detection model according to the embodiments of the present application. The present application can also be implemented as a device or apparatus program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0170] In yet another embodiment provided by the present invention, there is also provided a computer program product containing instructions, which when run on a computer, causes the computer to execute the method for intelligent detection of wide-field high-resolution video targets based on elastic sparsification or the method for training an image intelligent detection model according to the embodiments of the present application.

[0171] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0172] It should be noted that the above embodiments are illustrative of the present application rather than restrictive thereof, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words may be interpreted as names.

[0173] It should be noted that for the method embodiments of the present application, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should understand that the embodiments of the present application are not limited by the described order of actions, because according to the embodiments of the present application, certain steps may be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present application.

[0174] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the system or device, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.

[0175] The above is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. An intelligent detection method for wide-field high-resolution video targets based on elastic sparsification, characterized in that, Including: In response to the input of a video image, determining the distribution characteristics of the video image at multiple image scales, and determining the fusion data of the video image according to the multiple distribution characteristics; the distribution characteristics are used to characterize the distribution of image targets in the video image at each image scale; Determining the compressed data of the fusion data of the video image according to the target density of the target object in the video image determined from the fusion data and the target distribution characteristics among the multiple distribution characteristics; Determining the detection result of the image target in the video image according to the compressed data; the detection result is used to label the image target in the video image; The determining the compressed data of the fusion data of the video image according to the target density of the target object in the video image determined from the fusion data and the target distribution characteristics among the multiple distribution characteristics includes: Dividing the fusion data into multiple image blocks and generating data tokens corresponding to each image block; Inputting each data token into a scoring model respectively to obtain an evaluation value of each data token; the evaluation value is used to characterize the target density of the target object in the data token; Determining the distribution characteristic with the smallest corresponding image scale among the multiple distribution characteristics as the target distribution characteristic, and inputting the target distribution characteristic into an elastic selection sub-model in the image intelligent detection model to obtain the selection ratio of the multiple data tokens; Determining multiple target data tokens from the token sequence of the video image according to the selection ratio, and generating the compressed data according to the determined multiple target data tokens; the token sequence is the sequence of the data tokens obtained after being arranged in descending order according to the evaluation value.

2. The intelligent detection method for wide-field high-resolution video targets based on elastic sparsification as claimed in claim 1, wherein, The responding to the input of a video image, determining the distribution characteristics of the video image at multiple image scales, and determining the fusion data of the video image according to the multiple distribution characteristics includes: In response to the input of the video image, extracting multiple feature maps from the video image; each feature map has a different image scale; Fusing the determined multiple feature maps to obtain the fusion data.

3. The intelligent detection method for wide-field high-resolution video targets based on elastic sparsification according to claim 1, characterized in that, The determining the detection result of the image target in the video image according to the compressed data includes: Inputting the compressed data into a labeling sub-model in the image intelligent detection model to label the image blocks in the video image; Determining the detection result of the image target in the video image according to the deletion of the target image blocks in the image blocks; the target image blocks are the image blocks determined to be labeled incorrectly.

4. A training method for an image intelligent detection model, characterized in that, For training the image intelligent detection model in the wide-field high-resolution video target intelligent detection method based on elastic sparsification as claimed in claim 3, including: Obtain training images for training, and input the training images into the image intelligent detection model to obtain the training selection ratio of the elastic selection sub-model of the image intelligent detection model for the training images, and the detection training results of the annotation sub-model of the image intelligent detection model for the training image targets in the training images; Establish the density consistency loss function of the image intelligent detection model according to the training selection ratio, and establish the target detection loss function of the image intelligent detection model according to the training results; Jointly train the image intelligent detection model through the density consistency loss function and the target detection loss function.

5. An intelligent detection device for wide-field high-resolution video targets based on elastic sparsification, characterized in that, Including: A fusion module, configured to, in response to the input of a video image, determine the distribution characteristics of the video image at multiple image scales, and determine the fusion data of the video image according to the multiple distribution characteristics; the distribution characteristics are used to characterize the distribution of the image targets in the video image at each image scale; A compression module, configured to determine the compressed data of the fusion data of the video image according to the target density of the target object in the video image determined from the fusion data and the target distribution characteristics among the multiple distribution characteristics; A detection module, configured to determine the detection results of the image targets in the video image according to the compressed data; The detection results are used to annotate the image targets in the video image; The compression module includes: A token generation sub-module, configured to divide the fusion data into multiple image blocks and generate data tokens corresponding to each image block; An evaluation sub-module, configured to input each data token into a scoring model respectively to obtain an evaluation value of each data token; the evaluation value is used to characterize the target density of the target object in the data token; A ratio sub-module, configured to determine the target distribution characteristic as the distribution characteristic with the smallest corresponding image scale among the multiple distribution characteristics, and input the target distribution characteristic into the elastic selection sub-model in the image intelligent detection model to obtain the selection ratio of the multiple data tokens; A fusion sub-module, configured to determine multiple target data tokens from the token sequence of the video image according to the selection ratio, and generate the compressed data according to the determined multiple target data tokens; the token sequence is the sequence of the data tokens obtained after being arranged in descending order according to the evaluation value.

6. A training device for an image intelligent detection model, characterized in that For training the image intelligent detection model in the wide-field high-resolution video target intelligent detection method based on elastic sparsification as claimed in claim 3, including: An input module, configured to obtain training images for training, and input the training images into the image intelligent detection model to obtain the training selection ratio of the elastic selection sub-model of the image intelligent detection model for the training images, and the detection training results of the annotation sub-model of the image intelligent detection model for the training image targets in the training images; A loss module, configured to establish a density consistency loss function of the image intelligent detection model according to the training selection ratio, and establish an object detection loss function of the image intelligent detection model according to the training result; A training module, configured to jointly train the image intelligent detection model by using the density consistency loss function and the object detection loss function.

7. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

8. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 4 are implemented.

9. A computer program product, characterized in that, A computer program is stored on the computer program product. When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.