Dense small target detection method and device, medium and equipment

By using the CSP structure and dense feature connection network structure of the YOLOv8 network in drone aerial photography, combined with feature refinement module, the problem of low detection accuracy of dense small targets in drone aerial photography is solved, and more efficient feature representation and detection accuracy is achieved.

CN119942089AActive Publication Date: 2025-05-06NANJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510347491.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-05-06
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

The prior art has low detection accuracy of dense small targets in drone aerial photography, mainly due to the presence of a lot of noise in shallow feature maps, making it difficult to effectively capture global context information.

Method used

By using the CSP structure in the YOLOv8 network in the backbone feature extraction network, and introducing dense feature connection network structure and feature refinement modules into the neck multi-scale feature fusion network, feature extraction and feature fusion are performed to reduce noise and enhance feature representation of small targets.

Benefits of technology

It effectively reduces noise in shallow feature maps, enhances feature representation capabilities, and improves the accuracy and robustness of dense small object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942089A_ABST
    Figure CN119942089A_ABST
Patent Text Reader

Abstract

The invention discloses a dense small target detection method and device, a medium and equipment in the technical field of computer vision and target detection, and aims to solve the optimization problem of aerial photography dense small target detection. The method comprises the following steps: carrying out image preprocessing on acquired aerial image data to obtain an input image, and inputting the input image into a trained dense small target detection model: carrying out feature extraction through a trunk feature extraction network to obtain a multi-scale feature map; through a neck multi-scale feature fusion network, feature fusion and feature refinement are carried out on the multi-scale feature map, and multi-scale fusion features are obtained; and the small target detection layer is used for carrying out target detection on the multi-scale feature map to obtain a detection result. According to the invention, the detection performance of small target detection of the aerial image of the unmanned aerial vehicle can be improved, and robustness is increased for an algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and target detection, and in particular to a method, device, medium and equipment for detecting dense small targets. Background Art

[0002] As an important way for drones to perceive ground information, drone aerial photography has been increasingly widely used in the fields of intelligent transportation and disaster detection due to its wide monitoring range, strong flexibility and low cost. Advanced technologies such as computer vision represented by target detection have given drones the ability to perceive, analyze and make decisions autonomously, making them play an increasingly important role in real life. Spectral imaging technology is an important means of target recognition. The RGB images it captures have the advantages of high color reproduction and a wide shooting range to achieve high-resolution imaging of land targets. However, due to the high flying altitude of drones, the size of small targets in the image is small, and the surface texture and appearance features are not obvious; at the same time, the small targets are blocked by a large number of dense objects, resulting in low detection accuracy of small targets.

[0003] The existing methods for improving the recognition accuracy of dense small targets include one-stage detection method and two-stage detection method. The two-stage detection method has high recognition accuracy, but due to the huge number of parameters, it has poor real-time performance and is difficult to adapt to drones. Therefore, in order to balance real-time performance and detection accuracy, a one-stage method is used for dense small target detection. The use of pyramid models to extract and enhance features of images is the core of improving the recognition accuracy of small targets in the one-stage method. However, when performing feature fusion in the pyramid network, the shallow features that have been downsampled a small amount are not subjected to feature noise removal, and contain more useless information, resulting in reduced feature fusion performance and affecting the final small target detection results. In order to solve the key problem of low recognition accuracy caused by feature noise, convolution pooling, feature denoising methods based on diffusion networks, spatial attention mechanisms, channel attention mechanisms, or a combination of the two are usually used to reduce the impact of noise by focusing on the target area. However, the pooling and attention operations in the above methods focus on the local information of the feature map, making it difficult to capture the long-range dependencies between the input feature maps, thus ignoring the understanding of global context information. The feature denoising method based on diffusion networks is computationally intensive and complex, making it difficult to deploy in drones.

[0004] Therefore, how to alleviate the problem of feature maps containing more noise and further improve the accuracy of aerial photography dense small target detection is a problem that technical personnel in this field still need to solve. Summary of the invention

[0005] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a method, device, medium and equipment for dense small target detection. By performing feature refinement operations on shallow features and performing interactive fusion on feature maps, it is possible to reduce noise while enhancing the feature representation of small targets, thereby improving the accuracy of dense small target detection in aerial photography.

[0006] In order to solve the above technical problems, the present invention is implemented by adopting the following technical solutions:

[0007] The present invention provides a dense small target detection method, comprising:

[0008] Performing image preprocessing on the acquired aerial image data to obtain an input image;

[0009] The input image is input into the trained dense small target detection model: feature extraction is performed through the backbone feature extraction network to obtain a feature map at multiple scales; feature fusion and feature refinement are performed on the feature map at multiple scales through the neck multi-scale feature fusion network to obtain multi-scale fusion features; the small target detection layer performs target detection on the feature map at multiple scales to obtain a detection result.

[0010] Optionally, the image preprocessing adopts Mosaic data enhancement, and the image size is unified through a scaling operation.

[0011] Optionally, the backbone feature extraction network adopts a CSP-based backbone network in a Yolov8 network, including a first CBS module, a Split module, n Bottleneck modules, a Concat module, and a second CBS module; the data processing flow of the backbone feature extraction network includes:

[0012] Input the image or feature map into the first CBS module consisting of a convolution layer, a batch normalization layer and a SiLU activation function to generate an intermediate feature map, and introduce a residual connection to pass the intermediate feature map to the Concat module for splicing processing;

[0013] The intermediate feature map is split into two parts by the Split module, and a residual connection is introduced again to pass one part to the Concat module, and the other part is input into n Bottleneck modules for processing;

[0014] The feature map input to the Bottleneck module is convolved, normalized, and activated, and the output result is concatenated with the features passed to the Concat module twice by residual connection to obtain the concatenated result;

[0015] Input the splicing result into the second CBS module for processing to obtain a feature map at the current scale;

[0016] The feature map at the current scale is pooled and down-sampled, and the above process is repeated to obtain feature maps at multiple scales.

[0017] Optionally, the neck multi-scale feature fusion network includes a dense feature propagation module and a feature refinement module; the data processing flow of the neck multi-scale feature fusion network includes:

[0018] Performing 1*1 convolution operations on the multi-scale features extracted by the backbone feature extraction network, and representing the convolution results as P2, P3, P4 and P5 respectively;

[0019] According to P2, P3, P4 and P5, construct a top-down feature dense connection:

[0020] Perform the first upsampling operation on P5, input the sampling result to the P4 node, and perform feature fusion through the feature refinement module to obtain ;

[0021] right Perform the first upsampling operation, input the sampling result to the P3 node, and perform feature fusion through the feature refinement module to obtain ;

[0022] Perform a second upsampling operation on P5 and input the sampling result into Node, obtained by feature fusion through splicing scaling function ;

[0023] right Perform upsampling operation, input the sampling result to P2 node, and obtain the feature fusion through feature refinement module ;

[0024] Respectively Perform the second upsampling operation, P5 performs the third upsampling operation, and inputs the two sampling results into Node, obtained by feature fusion through splicing scaling function ;

[0025] The upsampling operation is implemented by nearest neighbor interpolation;

[0026] After dense connection of top-down features , and , construct bottom-up feature dense connections:

[0027] right Perform the first pooling downsampling operation and input the sampling results into Node, obtained by feature fusion through splicing scaling function ;

[0028] Respectively Perform pooling downsampling operations, Perform a second pooling downsampling operation and input the two sampling results into Node, obtained by feature fusion through splicing scaling function .

[0029] Optionally, the feature fusion through the feature refinement module includes:

[0030] According to the preset query vector dimension and value vector dimension, the shallow features are flattened twice to obtain the feature tensor and , wherein the shallow feature represents the node where the sampling result is input;

[0031] According to the preset key vector dimension, a flattening operation is performed on the deep features to obtain the feature tensor , wherein the deep feature represents the sampling result;

[0032] The feature tensor With the preset weight matrix Multiply to get a vector matrix ,in, represents the query vector weight, represents the key vector weight, represents the value vector weight, represents the query vector, represents the key vector, represents a value vector;

[0033] The query vector With key vector After matrix multiplication and weight normalization, the cross weight matrix is ​​obtained, and the cross weight matrix is ​​combined with the value vector Multiply them together to get the attention head vector, which is calculated as follows:

[0034]

[0035] in, represents the attention head vector, represents the number of attention heads in the attention head vector, Indicates Attention head, represents the attention function, represents the activation function, represents the scaling parameter;

[0036] All attention heads are concatenated together and resized to the original size of the attention heads using the output transformation matrix to obtain the updated feature representation, which is calculated as follows:

[0037]

[0038] in, represents the updated feature representation, represents the concatenation function, Represents the output transformation matrix.

[0039] Optionally, the small target detection layer includes two parallel links, the first link includes two CBS layers, a two-dimensional convolution layer and a target regression loss function, which is used to predict the bounding box; the second link includes two CBS layers, a two-dimensional convolution layer and a target classification loss function, which is used to determine the category of the target;

[0040] The small target detection layer uses the target regression loss function and the target classification loss function to calculate the target regression loss and target classification loss , and the overall loss of the dense small target detection model is obtained , the overall loss The calculation formula is as follows:

[0041]

[0042] in, represents the target classification loss weight, represents the target regression loss weight;

[0043] The target classification loss function is as follows:

[0044]

[0045] in, Representation sample The real category, represents the positive class, represents the negative class, Representation sample The probability of being predicted as the positive class, represents the number of samples;

[0046] The target regression loss Including CioU loss and DFL loss , the target regression loss function is as follows:

[0047]

[0048] in, represents the CioU loss weight, represents the DFL loss weight;

[0049] The CioU loss The calculation formula is as follows:

[0050]

[0051] in, Represents the Euclidean distance between the center points of the target box and the predicted box, represents the diagonal distance of the target box, represents the parameter to measure the consistency of aspect ratio, Represents the intersection-over-union ratio of the target box and the predicted box;

[0052] The DFL loss The calculation formula is as follows:

[0053]

[0054] in, Representation sample The true value of the target box, Representation sample The target box prediction value.

[0055] In a second aspect, the present invention provides a dense small target detection device, comprising:

[0056] Preprocessing module: used to perform image preprocessing on the acquired aerial image data to obtain an input image;

[0057] The target detection module is used to: input the input image into the trained dense small target detection model; perform feature extraction through the backbone feature extraction network to obtain a feature map at multiple scales; perform feature fusion and feature refinement on the feature map at multiple scales through the neck multi-scale feature fusion network to obtain multi-scale fusion features; and perform target detection on the feature map at multiple scales in the small target detection layer to obtain a detection result.

[0058] In a third aspect, the present invention provides a computer-readable storage medium having a computer program / instruction stored thereon. When the computer program / instruction is executed by a processor, the steps of any of the dense small target detection methods described in the first aspect are implemented.

[0059] In a fourth aspect, the present invention provides a computer device / equipment / system, comprising:

[0060] Memory, for storing computer programs / instructions;

[0061] A processor is used to execute the computer program / instructions to implement the steps of the dense small target detection method described in any one of the first aspects.

[0062] In a fifth aspect, the present invention provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the dense small target detection method described in any one of the first aspects.

[0063] Compared with the prior art, the present invention has the following beneficial effects:

[0064] 1. The dense small target detection method provided by the present invention uses the CSP structure in the YOLOv8 network in the backbone feature extraction network, introduces a dense feature connection network structure in the neck multi-scale fusion network, and introduces a feature refinement module. While effectively extracting features, it strengthens multi-level feature fusion, enriches small target feature information, reduces shallow feature map noise, enhances feature representation capabilities, and improves small target detection performance;

[0065] 2. The dense small target detection device provided by the present invention realizes dense small target detection by setting a data preprocessing module and a target detection module, which can improve the detection performance of small target detection in drone aerial images and increase robustness, and has practical significance and good application prospects;

[0066] 3. The computer-readable storage medium, computer device / equipment / system and computer program product provided by the present invention can execute the steps of the dense small target detection method provided by the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 A flow chart of a dense small target detection method provided according to an embodiment of the present invention;

[0068] Figure 2 A schematic diagram of a CSP module provided according to an embodiment of the present invention;

[0069] Figure 3 A schematic diagram of a Bottleneck module provided according to an embodiment of the present invention;

[0070] Figure 4 A schematic diagram of a dense feature connection network structure provided according to an embodiment of the present invention;

[0071] Figure 5 A schematic diagram of a feature refinement module provided according to an embodiment of the present invention. DETAILED DESCRIPTION

[0072] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. In the absence of conflict, the embodiments of the present application and the technical features in the embodiments can be combined with each other.

[0073] It should be noted that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0074] Embodiment 1:

[0075] The embodiment of the present invention discloses a method for detecting dense small targets, referring to Figure 1 As shown, the specific steps include:

[0076] S1, performing image preprocessing on the acquired aerial image data to obtain an input image;

[0077] S2, input the input image into the trained dense small target detection model: perform feature extraction through the backbone feature extraction network to obtain a feature map at multiple scales; perform feature fusion and feature refinement on the feature map at multiple scales through the neck multi-scale feature fusion network to obtain multi-scale fusion features; small target detection layer, perform target detection on the feature map at multiple scales to obtain detection results.

[0078] Specifically,

[0079] In step S1, the image preprocessing adopts Mosaic data enhancement and combines a scaling operation to unify the image size. In this embodiment, the resolution of the input image data is uniformly set to 640*640.

[0080] In step S2, the training of the dense small target detection model includes:

[0081] S2.1, construct training set;

[0082] S2.2, build a dense small target detection model, including: backbone feature extraction network, neck multi-scale feature fusion network and small target detection layer;

[0083] S2.3, inputting the training set into a backbone feature extraction network for feature extraction to obtain a feature map at multiple scales;

[0084] S2.4, transferring the multi-scale feature map to the neck multi-scale feature fusion network to perform feature fusion and feature refinement to obtain multi-scale fusion features;

[0085] S2.5, transferring the multi-scale feature map to the small target detection layer for target detection to obtain a detection result;

[0086] S2.6. According to the detection results, the parameters of the dense small target detection model are optimized through a loss function to obtain a trained dense small target detection model.

[0087] In step S2.2, building a dense small target detection model includes:

[0088] S2.2.1, design the overall architecture of dense small object detection network;

[0089] S2.2.2, build a backbone feature extraction network: use the CSP network structure to extract features from the input image and obtain a multi-scale feature map;

[0090] S2.2.3, build a multi-scale fusion network for the neck: including a dense feature connection network structure and a feature refinement module. The dense feature connection network structure enriches the feature information of small targets by optimizing the feature propagation path. The feature refinement module reduces feature map noise through the self-attention mechanism and improves the detection accuracy of small targets.

[0091] S2.2.4, small target detection layer and loss function construction.

[0092] In this embodiment, the backbone network in the Yolov8 network is used as the backbone feature extraction network, and the CSP module is specifically used to extract features from the input image to obtain feature maps with scales of 160*160, 80*80, 40*40, and 20*20; Figure 2 As shown, the CSP module includes a first CBS module, a Split module, n Bottleneck modules, a Concat module and a second CBS module; Figure 3 Shown is a schematic diagram of a Bottleneck module, which includes two CBS modules connected in series.

[0093] In step S2.2.2, the CSP network structure is used to extract features from the input image, and the multi-scale feature map obtained includes:

[0094] Input the image or feature map into the first CBS module consisting of a convolution layer, a batch normalization layer and a SiLU activation function to generate an intermediate feature map, and introduce a residual connection to pass the intermediate feature map to the Concat module for splicing processing;

[0095] The intermediate feature map is split into two parts by the Split module, and a residual connection is introduced again to pass one part to the Concat module, and the other part is input into n Bottleneck modules for processing;

[0096] The feature map input to the Bottleneck module is convolved, normalized, and activated, and the output result is concatenated with the features passed to the Concat module twice by residual connection to obtain the concatenated result;

[0097] Input the splicing result into the second CBS module for processing to obtain a feature map at the current scale;

[0098] The feature map at the current scale is pooled and down-sampled, and the above process is repeated to obtain feature maps at scales of 160*160, 80*80, 40*40, and 20*20.

[0099] refer to Figure 4 As shown, in this embodiment, the neck multi-scale feature fusion network is based on the idea of ​​the mainstream BiFPN network, and is optimized to obtain a feature pyramid network architecture based on dense feature propagation. Multiple information flows are established between feature information propagation links to obtain richer information. In the process of feature propagation, additional branches are added to perform feature refinement operations to eliminate noise.

[0100] In step S2.2.3, building a neck multi-scale feature fusion network includes:

[0101] First, build a dense feature connection network structure:

[0102] (1) Construct a 1*1 convolution layer: perform 1*1 convolution operations on the multi-scale features extracted by the backbone feature extraction network, and represent the convolution results as P2, P3, P4, and P5. Changing the number of channels facilitates feature fusion operations. The scales of the multi-scale features are 160*160, 80*80, 40*40, and 20*20 respectively.

[0103] (2) Constructing top-down feature dense connections:

[0104] Perform the first upsampling operation on P5, input the sampling result to the P4 node, and perform feature fusion through the feature refinement module to obtain ;

[0105] right Perform the first upsampling operation, input the sampling result to the P3 node, and perform feature fusion through the feature refinement module to obtain ;

[0106] Perform a second upsampling operation on P5 and input the sampling result into Node, obtained by feature fusion through splicing scaling function ;

[0107] right Perform upsampling operation, input the sampling result to P2 node, and obtain the feature fusion through feature refinement module ;

[0108] Respectively Perform the second upsampling operation, P5 performs the third upsampling operation, and inputs the two sampling results into Node, obtained by feature fusion through splicing scaling function ;

[0109] The upsampling operation is implemented by nearest neighbor interpolation;

[0110] Among them, in the top-down feature propagation process, P2 obtains the fusion input of P3, P4 and P5 nodes, P3 node obtains the fusion input of P4 and P5 nodes, P4 obtains the fusion input of P5 node, and P5 does not obtain input from other levels;

[0111] (3) Constructing bottom-up feature dense connections:

[0112] After dense connection of top-down features , and , construct bottom-up feature dense connections:

[0113] right Perform the first pooling downsampling operation and input the sampling results into Node, obtained by feature fusion through splicing scaling function ;

[0114] Respectively Perform pooling downsampling operations, Perform a second pooling downsampling operation and input the two sampling results into Node, obtained by feature fusion through splicing scaling function The upsampling operation is implemented by nearest neighbor interpolation;

[0115] Among them, P5 is a sub-pixel area, which shows less information representation of small target features. The target area corresponding to the original image is not highlighted, so no operation is performed here. It can be obtained that in the bottom-up feature propagation process, Obtained and The connection input, Obtained The connection input, No other level input is obtained.

[0116] After building the dense feature connection network structure, build a shallow feature refinement module inside. The specific steps are as follows:

[0117] In the above dense feature propagation process, an additional branch is added to perform feature refinement operation to remove noise. The specific operation is as follows: in the above step (2), the deep features are used to guide the shallow features to perform feature refinement operation to remove noise in the feature map. It is necessary to add three branches to the top-down propagation network architecture and perform three feature refinement operations, namely, the P5 node performs an upsampling operation and then inputs it to P4, the P4 node performs an upsampling operation and inputs it to P3, and the P3 node performs an upsampling operation and inputs it to P2.

[0118] refer to Figure 5 As shown, the feature refinement steps are as follows:

[0119] The feature fusion through the feature refinement module includes:

[0120] According to the preset query vector dimension and value vector dimension, the shallow features are flattened twice to obtain the feature tensor and , wherein the shallow feature represents the node where the sampling result is input;

[0121] According to the preset key vector dimension, a flattening operation is performed on the deep features to obtain the feature tensor , wherein the deep feature represents the sampling result;

[0122] The feature tensor With the preset weight matrix Multiply to get a vector matrix ,in, represents the query vector weight, represents the key vector weight, represents the value vector weight, represents the query vector query, represents the key vector key, Represents the value vector value;

[0123] The query vector With key vector After matrix multiplication and weight normalization, the cross weight matrix is ​​obtained, and the cross weight matrix is ​​combined with the value vector Multiply them together to get the attention head vector, which is calculated as follows:

[0124]

[0125] in, represents the attention head vector, represents the number of attention heads in the attention head vector, Indicates Attention head, represents the attention function, represents the activation function, Represents the scaling parameter, which is used to scale the dot product to avoid the attention weight being too small or too large;

[0126] All attention heads are concatenated together and resized to the original size of the attention heads using the output transformation matrix to obtain the updated feature representation, which is calculated as follows:

[0127]

[0128] in, represents the updated feature representation, represents the concatenation function, Represents the output transformation matrix.

[0129] This embodiment uses P2, P3, and P4 as small target detection layers for regression and classification. Compared with the traditional YOLO series that uses P3, P4, and P5 as predictions, since the deep layer P5 is a sub-pixel area, it shows less information representation of small target features, and the target area corresponding to the original image is not highlighted. Therefore, using P2, P3, and P4 as detection heads is more conducive to comprehensive detection of dense small targets.

[0130] The neck multi-scale feature fusion network proposed in this embodiment uses residual connections to maintain the original features and prevent gradient disappearance in P2, P3 and P4 on the same-scale feature propagation path; adaptive weights are used to ensure effective fusion of multiple input feature maps during the fusion of multi-layer feature maps; feature refinement operations are performed to eliminate noise during the feature propagation process, and finally residual connections are used to merge information to enhance feature representation, thereby generating final output features.

[0131] In this embodiment, the small target detection layer includes two parallel links. The first link includes two CBS layers, a two-dimensional convolution layer and a target regression loss function, which is used to achieve accurate prediction of the bounding box; the second link also includes two CBS layers, a two-dimensional convolution layer and a target classification loss function, which is used to achieve target category determination;

[0132] The small target detection layer uses the target regression loss function and the target classification loss function to calculate the target regression loss and target classification loss , and the overall loss of the dense small target detection model is obtained , the overall loss The calculation formula is as follows:

[0133]

[0134] in, represents the target classification loss weight, represents the target regression loss weight;

[0135] The target classification loss function is as follows:

[0136]

[0137] in, Representation sample The real category, represents the positive class, represents the negative class, Representation sample The probability of being predicted as the positive class, represents the number of samples;

[0138] The target regression loss Including CioU loss and DFL loss , the target regression loss function is as follows:

[0139]

[0140] in, represents the CioU loss weight, represents the DFL loss weight;

[0141] The CioU loss The calculation formula is as follows:

[0142]

[0143] in, Represents the Euclidean distance between the center points of the target box and the predicted box, represents the diagonal distance of the target box, represents the parameter to measure the consistency of aspect ratio, Represents the intersection-over-union ratio of the target box and the predicted box;

[0144] The DFL loss The calculation formula is as follows:

[0145]

[0146] in, Representation sample The true value of the target box, Representation sample The target box prediction value.

[0147] In other embodiments, the dense small target detection model may further include an output terminal to visualize the detection results.

[0148] In summary, the dense small target detection method proposed in this embodiment first inputs an image with a resolution of 640*640 into the backbone feature extraction network; secondly, the obtained feature maps with resolutions of 160*160, 80*80, 40*40, and 20*20 are input into the neck multi-scale feature fusion (through feature denoising and dense feature transmission) to reduce feature noise, and enrich the semantic information and detail information of the shallow and deep feature maps; finally, the feature maps with resolutions of 160*160, 80*80, and 40*40 after the multi-scale feature fusion of the neck are input into the small target detection layer to obtain the detection results, which are visualized through the output end output; the detection performance of small target detection in drone aerial images is effectively improved, and the robustness of the algorithm is increased.

[0149] Embodiment 2:

[0150] Based on the same inventive concept as the first embodiment, the embodiment of the present invention discloses a dense small target detection device, including:

[0151] Preprocessing module: used to perform image preprocessing on the acquired aerial image data to obtain an input image;

[0152] The target detection module is used to: input the input image into the trained dense small target detection model; perform feature extraction through the backbone feature extraction network to obtain a feature map at multiple scales; perform feature fusion and feature refinement on the feature map at multiple scales through the neck multi-scale feature fusion network to obtain multi-scale fusion features; and perform target detection on the feature map at multiple scales in the small target detection layer to obtain a detection result.

[0153] The specific functional implementation of each of the above modules can be found in the relevant content of the method in Example 1 and will not be elaborated here.

[0154] Embodiment three:

[0155] This embodiment provides a computer-readable storage medium on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the steps of the dense small target detection method as described in any one of the first embodiments are implemented.

[0156] Embodiment 4:

[0157] This embodiment provides a computer device / equipment / system, including:

[0158] Memory, for storing computer programs / instructions;

[0159] A processor is used to execute the computer program / instructions to implement the steps of the dense small target detection method described in any one of the first aspects.

[0160] Embodiment five:

[0161] This embodiment provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the dense small target detection method as described in any one of the first embodiments.

[0162] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0163] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0164] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0165] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0166] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the enlightenment of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present invention and the claims, which all fall within the protection of the present invention.

Claims

1. A method for detecting dense small targets, characterized in that: include: Performing image preprocessing on the acquired aerial image data to obtain an input image; The input image is input into the trained dense small target detection model: feature extraction is performed through the backbone feature extraction network to obtain a feature map at multiple scales; feature fusion and feature refinement are performed on the feature map at multiple scales through the neck multi-scale feature fusion network to obtain multi-scale fusion features; the small target detection layer performs target detection on the feature map at multiple scales to obtain a detection result.

2. The method for detecting dense small targets according to claim 1, characterized in that: The image preprocessing adopts Mosaic data enhancement and unifies the image size through scaling operation.

3. The method for detecting dense small targets according to claim 1, characterized in that: The backbone feature extraction network adopts the CSP-based backbone network in the Yolov8 network, including a first CBS module, a Split module, n Bottleneck modules, a Concat module and a second CBS module; The data processing flow of the backbone feature extraction network includes: Input the image or feature map into the first CBS module consisting of a convolution layer, a batch normalization layer and a SiLU activation function to generate an intermediate feature map, and introduce a residual connection to pass the intermediate feature map to the Concat module for splicing processing; The intermediate feature map is split into two parts by the Split module, and a residual connection is introduced again to pass one part to the Concat module, and the other part is input into n Bottleneck modules for processing; The feature map input to the Bottleneck module is convolved, normalized, and activated, and the output result is concatenated with the features passed to the Concat module twice by residual connection to obtain the concatenated result. Input the splicing result into the second CBS module for processing to obtain a feature map at the current scale; The feature map at the current scale is pooled and down-sampled, and the above process is repeated to obtain feature maps at multiple scales.

4. The method for detecting dense small targets according to claim 1, characterized in that: The neck multi-scale feature fusion network includes a dense feature propagation module and a feature refinement module; The data processing flow of the neck multi-scale feature fusion network includes: Performing 1*1 convolution operations on the multi-scale features extracted by the backbone feature extraction network, and representing the convolution results as P2, P3, P4 and P5 respectively; According to P2, P3, P4 and P5, construct a top-down feature dense connection: Perform the first upsampling operation on P5, input the sampling result to the P4 node, and perform feature fusion through the feature refinement module to obtain ; right Perform the first upsampling operation, input the sampling result to the P3 node, and perform feature fusion through the feature refinement module to obtain ; Perform a second upsampling operation on P5 and input the sampling result into Node, obtained by feature fusion through splicing scaling function ; right Perform upsampling operation, input the sampling result to P2 node, and obtain the feature fusion through feature refinement module ; Respectively Perform the second upsampling operation, P5 performs the third upsampling operation, and inputs the two sampling results into Node, obtained by feature fusion through splicing scaling function ; The upsampling operation is implemented by nearest neighbor interpolation; After dense connection of top-down features , and , construct bottom-up feature dense connections: right Perform the first pooling downsampling operation and input the sampling results into Node, obtained by feature fusion through splicing scaling function ; Respectively Perform pooling downsampling operations, Perform a second pooling downsampling operation and input the two sampling results into Node, obtained by feature fusion through splicing scaling function .

5. The method for detecting dense small targets according to claim 3, characterized in that: The feature fusion through the feature refinement module includes: According to the preset query vector dimension and value vector dimension, the shallow features are flattened twice to obtain the feature tensor and , wherein the shallow feature represents the node where the sampling result is input; According to the preset key vector dimension, a flattening operation is performed on the deep features to obtain the feature tensor , wherein the deep feature represents the sampling result; The feature tensor With the preset weight matrix Multiply to get a vector matrix ,in, represents the query vector weight, represents the key vector weight, represents the value vector weight, represents the query vector, represents the key vector, represents a value vector; The query vector With key vector After matrix multiplication and weight normalization, the cross weight matrix is ​​obtained, and the cross weight matrix is ​​combined with the value vector Multiply them together to get the attention head vector, which is calculated as follows: ; in, represents the attention head vector, represents the number of attention heads in the attention head vector, Indicates A head of attention, represents the attention function, represents the activation function, represents the scaling parameter; All attention heads are concatenated together and resized to the original size of the attention heads using the output transformation matrix to obtain the updated feature representation, which is calculated as follows: ; in, represents the updated feature representation, represents the concatenation function, Represents the output transformation matrix.

6. The method for detecting dense small targets according to claim 1, characterized in that: The small target detection layer includes two parallel links. The first link includes two CBS layers, a two-dimensional convolution layer and a target regression loss function, which is used to predict the bounding box; the second link includes two CBS layers, a two-dimensional convolution layer and a target classification loss function, which is used to determine the category of the target. The small target detection layer uses the target regression loss function and the target classification loss function to calculate the target regression loss and target classification loss , and the overall loss of the dense small target detection model is obtained , the overall loss The calculation formula is as follows: ; in, represents the target classification loss weight, represents the target regression loss weight; The target classification loss function is as follows: ; in, Representation sample The real category, represents the positive class, represents the negative class, Representation sample The probability of being predicted as the positive class, represents the number of samples; The target regression loss Including CioU loss and DFL loss , the target regression loss function is as follows: ; in, represents the CioU loss weight, represents the DFL loss weight; The CioU loss The calculation formula is as follows: ; in, Represents the Euclidean distance between the center points of the target box and the predicted box, represents the diagonal distance of the target box, represents the parameter to measure the consistency of aspect ratio, Represents the intersection-over-union ratio of the target box and the predicted box; The DFL loss The calculation formula is as follows: ; in, Representation sample The true value of the target box, Representation sample The target box prediction value.

7. A dense small target detection device, characterized in that: include: Preprocessing module: Used to: perform image preprocessing on the acquired aerial image data to obtain an input image; The target detection module is used to: input the input image into the trained dense small target detection model; perform feature extraction through the backbone feature extraction network to obtain a feature map at multiple scales; perform feature fusion and feature refinement on the feature map at multiple scales through the neck multi-scale feature fusion network to obtain multi-scale fusion features; and perform target detection on the feature map at multiple scales in the small target detection layer to obtain a detection result.

8. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instruction is executed by a processor, the steps of the dense small target detection method described in any one of claims 1-6 are implemented.

9. A computer device / equipment / system, characterized in that: include: Memory, for storing computer programs / instructions; A processor, configured to execute the computer program / instructions to implement the steps of the dense small target detection method according to any one of claims 1 to 6.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the dense small target detection method described in any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Image pyramid feature guided multi-scale target detection method

    CN114612709A

  • Multi-scale attention-fused traffic helmet small target detection system and method

    CN116665156A

  • Improved YOLOv8 unmanned aerial vehicle aerial target detection method

    CN117557922A

  • Defect detection model training method, detection method, equipment, medium and product

    CN119168976A

  • Target detection method and system in complex environment and medium

    CN119399449A

Cited By

  • Target detection method and device based on anti-interference characteristic information compensation

    CN120219725A

  • Method for detecting package continuity of security inspection machine

    CN120747589A

  • Oil and gas pipeline intrusion detection method and system based on bionic vision

    CN120894571A

  • Method, device and equipment for detecting port state of optical cable cross connecting cabinet and storage medium

    CN121053138A

  • Visible light target detection method and device and storage medium

    CN121170241A