Small object detection method for commercial vehicle based on improved YOLO v9

By improving the YOLO v9 model, cross-layer bidirectional feature fusion and parallel multi-branch convolution processing are adopted to optimize the micro-object detection model, which solves the problems of high computing resource consumption and long training time, improves detection accuracy and adaptability, and enhances the ability to understand complex scenarios.

CN120472427APending Publication Date: 2025-08-12GUANGXI UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510583764.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the prior art, micro-object detection relies on complex convolutional neural network models, resulting in high computing resource consumption and increased model training time cost, making it difficult to run efficiently on embedded and mobile devices, and multi-scale feature fusion methods rely on manual adjustment, reducing detection accuracy.

Method used

The improved YOLO v9 model is adopted, through cross-layer bidirectional feature fusion and parallel multi-branch convolution processing, combined with lightweight attention mechanism, the micro-object detection model is optimized, including the fusion of low-level detail features and high-level semantic features, parallel multi-branch convolution and lightweight attention processing, and adaptive dynamic weighted fusion feature data.

Benefits of technology

It improves the accuracy and adaptability of micro-object detection, reduces computing resource consumption and model training time cost, enhances the understanding of complex scenarios, and reduces the dependence on external data enhancement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472427A_ABST
    Figure CN120472427A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of target detection, and discloses a commercial vehicle tiny object detection method based on improved YOLO v9, and the method comprises the steps: inputting collected to-be-detected image data into a tiny target detection model, and achieving the correction of a vehicle driving function through an obtained target detection result. The model construction process comprises the steps that an initial tiny target detection model is constructed according to a YOLO v9 model, a composite backbone network and a detection head of the initial tiny target detection model are optimized, and the optimized composite backbone network is used for conducting multi-level feature mapping processing and cross-branch bidirectional feature fusion on a first multi-scale feature map; the optimized detection head is used for performing parallel multi-branch convolution, first feature fusion, lightweight attention and adaptive dynamic weighted fusion on the first multi-scale feature map in sequence; and training the optimized tiny target detection model to obtain the tiny target detection model. According to the method, the detection precision of the tiny object is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection technology, and in particular to a small object detection method for commercial vehicles based on an improved YOLO v9. Background Art

[0002] With the rapid development of intelligent transportation systems, small object detection around vehicles has become a key technology for ensuring driving safety and improving logistics and transportation efficiency. For example, in logistics and transportation scenarios, accurately detecting small objects such as small traffic signs, small packages, and small animals around vehicles is crucial for ensuring driving safety and efficient logistics operations. Accurate small object detection can help vehicles make timely avoidance maneuvers, such as slowing down, to prevent collisions.

[0003] In existing technologies, small target detection during vehicle operation relies on convolutional neural network models, which improve target detection accuracy through feature extraction enhancement and multi-scale feature fusion. The feature extraction enhancement method relies on a complex network structure and consumes a large amount of computing resources during training and inference, making it difficult to run efficiently on resource-constrained platforms such as embedded and mobile devices, limiting its practical application. The multi-scale feature fusion method requires manual setting of branch weights during training, relying heavily on the technical experience and extensive experimental adjustments. This reduces the accuracy of small target detection and increases the time cost of model training.

[0004] It can be seen that how to improve the accuracy of small target detection results while reducing computing resource consumption and model training time cost has become a technical problem that technical personnel in this field need to solve urgently. Summary of the Invention

[0005] The present invention provides a small object detection method for commercial vehicles based on an improved YOLO v9 to solve the technical problem of how to improve the accuracy of small target detection results while reducing computing resource consumption and model training time costs, thereby achieving the effect of improving the detection accuracy of small objects and reducing computing resource consumption and model training time costs.

[0006] In a first aspect, the present invention provides a method for detecting small objects in commercial vehicles based on an improved YOLO v9. The method comprises: during vehicle driving, inputting collected image data to be detected into a small target detection model, and using the obtained target detection results to correct the vehicle driving function. The process of constructing the small target detection model comprises:

[0007] An initial small target detection model is constructed based on the YOLO v9 model. The composite backbone network of the initial small target detection model is set to perform multi-level feature mapping processing on a first multi-scale feature map, and based on the obtained multi-level feature mapping result, a cross-branch bidirectional feature fusion of low-level detail features and high-level semantic features is performed on the first multi-scale feature map to obtain first fused feature maps of different scales. The first multi-scale feature map is obtained by performing multi-scale feature extraction on the image data to be detected by the backbone network. The first fused feature map is used as an input to the backbone network;

[0008] Optimizing the detection head of the initial small target detection model to obtain a small target detection model after the detection head is optimized, wherein the optimized detection head is configured to sequentially perform parallel multi-branch convolution processing, first feature fusion, and lightweight attention processing on the first multi-scale feature map to obtain an attention weight matrix, and adaptively and dynamically weighted fusion is performed on spatial feature data of different scales obtained by the parallel multi-branch convolution processing according to the attention weight matrix to obtain a third fused feature map, wherein the third fused feature map is used for upsampling and splicing processing;

[0009] The small target detection model after the detection head is optimized is trained to obtain the small target detection model.

[0010] Preferably, the performing multi-level feature mapping processing on the first multi-scale feature map, and performing cross-branch bidirectional feature fusion of low-level detail features and high-level semantic features on the first multi-scale feature map according to the obtained multi-level feature mapping result to obtain first fused feature maps of different scales, includes:

[0011] arbitrarily selecting a first feature map to be fused and a second feature map to be fused from the first multi-scale feature map;

[0012] Using a linear projection method to map the first channel number of the first feature map to be fused to the second channel number of the second feature map to be fused, obtaining a first multi-level feature mapping result; using the linear projection method to map the second channel number of the second feature map to be fused to the first channel number of the first feature map to be fused, obtaining a second multi-level feature mapping result;

[0013] According to the first multi-level feature mapping result, the first feature map to be fused and the second feature map to be fused are subjected to bottom-up channel fusion, and according to the second multi-level feature mapping result, the first feature map to be fused and the second feature map to be fused are subjected to top-down channel fusion to obtain first fused feature maps of different scales.

[0014] Preferably, the step of sequentially performing parallel multi-branch convolution processing, first feature fusion, and lightweight attention processing on the first multi-scale feature map to obtain an attention weight matrix includes:

[0015] Using parallel branch convolution kernels of different sizes, perform convolution processing of different scales on the first multi-scale feature map to obtain multi-scale hierarchical feature data;

[0016] Performing batch normalization processing on the multi-scale hierarchical feature data, and performing element-by-element threshold processing on the multi-scale hierarchical feature data after the batch normalization processing using a ReLU activation function to obtain the spatial feature data of different scales;

[0017] Performing a first feature fusion on the spatial feature data of different scales to obtain a second fused feature map;

[0018] The second fused feature map is sequentially subjected to global average pooling, channel dimensionality reduction, and attention mapping to obtain an attention weight matrix.

[0019] Preferably, the method further comprises:

[0020] The backbone network of the small target detection model after the detection head is optimized is optimized to obtain the target detection model after the backbone network is optimized. The optimized backbone network is set to perform reparameterized convolution processing and multi-branch convolution processing on the first multi-scale feature map in sequence through the first convolution branch to obtain local spatial feature data, and the first multi-scale feature map is partially connected across stages through the second convolution branch to obtain a second multi-scale feature map and a third multi-scale feature map. After dense convolution processing, the second multi-scale feature map is directly jump-connected with the third multi-scale feature map to obtain a fourth fused feature map. The local spatial feature data and the fourth fused feature map are subjected to second feature fusion to obtain a fifth fused feature map. The fifth fused feature map is used to input into the detection head.

[0021] Preferably, before inputting the collected image data to be detected into the small target detection model to correct the vehicle driving function with the obtained target detection results, the method further includes:

[0022] The collected image data to be detected are sequentially subjected to resolution enhancement processing, denoising filtering processing, contrast enhancement processing and normalization processing to obtain standardized image data to be detected;

[0023] Saliency detection is performed on the standardized image data to be detected to obtain the image data to be detected including a potential target area.

[0024] Preferably, the step of training the small target detection model after the detection head is optimized to obtain the small target detection model includes:

[0025] Collecting sample image data, and annotating the sample image data with tiny target text to obtain label text data;

[0026] Constructing a training data set based on the sample image data and the label text data;

[0027] The training data set is used to train the small target detection model after the detection head is optimized to obtain the small target detection model.

[0028] In a second aspect, the present invention further provides a small object detection system for commercial vehicles based on an improved YOLO v9, which implements the small object detection method for commercial vehicles based on the improved YOLO v9 described above, and the system includes: a target detection module;

[0029] The target detection module is used to input the collected image data to be detected into the small target detection model during the vehicle driving process, and to correct the vehicle driving function with the target detection results obtained;

[0030] The target detection module includes: a first model optimization submodule, a second model optimization submodule and a model training submodule;

[0031] The first model optimization submodule is used to construct an initial small target detection model based on the YOLO v9 model. The composite backbone network of the initial small target detection model is set to perform multi-level feature mapping processing on the first multi-scale feature map, and based on the obtained multi-level feature mapping results, perform cross-branch bidirectional feature fusion of low-level detail features and high-level semantic features on the first multi-scale feature map to obtain first fused feature maps of different scales. The first multi-scale feature map is obtained by the backbone network performing multi-scale feature extraction on the image data to be detected. The first fused feature map is used to input the backbone network;

[0032] The second model optimization submodule is used to optimize the detection head of the initial small target detection model to obtain a small target detection model after the detection head is optimized. The optimized detection head is set to perform parallel multi-branch convolution processing, first feature fusion and lightweight attention processing on the first multi-scale feature map in sequence to obtain an attention weight matrix. According to the attention weight matrix, adaptive dynamic weighted fusion is performed on the spatial feature data of different scales obtained by the parallel multi-branch convolution processing to obtain a third fused feature map. The third fused feature map is used for upsampling and splicing processing;

[0033] The model training submodule is used to train the small target detection model after the detection head is optimized to obtain the small target detection model.

[0034] Preferably, the first model optimization submodule includes:

[0035] An image selection unit, configured to arbitrarily select a first feature map to be fused and a second feature map to be fused from the first multi-scale feature map;

[0036] A linear projection unit is configured to map the first number of channels of the first feature map to be fused to the second number of channels of the second feature map to be fused using a linear projection method to obtain a first multi-level feature mapping result, and map the second number of channels of the second feature map to be fused to the first number of channels of the first feature map to be fused using the linear projection method to obtain a second multi-level feature mapping result;

[0037] A first feature fusion unit is configured to perform bottom-up channel fusion on the first feature map to be fused and the second feature map to be fused according to the first multi-level feature mapping result, and to perform top-down channel fusion on the first feature map to be fused and the second feature map to be fused according to the second multi-level feature mapping result, to obtain first fused feature maps of different scales.

[0038] Preferably, the second model optimization submodule includes:

[0039] A parallel branch convolution processing unit, configured to use parallel branch convolution kernels of different sizes to perform convolution processing of different scales on the first multi-scale feature map to obtain multi-scale hierarchical feature data;

[0040] a normalization threshold processing unit, configured to perform batch normalization processing on the multi-scale hierarchical feature data, and perform element-by-element threshold processing on the multi-scale hierarchical feature data after batch normalization processing using a ReLU activation function to obtain the spatial feature data of different scales;

[0041] A second feature fusion unit is used to perform first feature fusion on the spatial feature data of different scales to obtain a second fused feature map;

[0042] An attention mapping processing unit is used to perform global average pooling processing, channel dimensionality reduction processing and attention mapping processing on the second fused feature map in sequence to obtain an attention weight matrix.

[0043] Preferably, the target detection module further includes a third model optimization submodule;

[0044] The third model optimization submodule is used to optimize the backbone network of the small target detection model after the detection head is optimized to obtain the target detection model after the backbone network is optimized. The optimized backbone network is set to perform reparameterized convolution processing and multi-branch convolution processing on the first multi-scale feature map in sequence through the first convolution branch to obtain local spatial feature data, and perform cross-stage partial connection on the first multi-scale feature map through the second convolution branch to obtain a second multi-scale feature map and a third multi-scale feature map. After dense convolution processing, the second multi-scale feature map is directly jump-connected with the third multi-scale feature map to obtain a fourth fused feature map. The local spatial feature data and the fourth fused feature map are subjected to second feature fusion to obtain a fifth fused feature map. The fifth fused feature map is used to input into the detection head.

[0045] This application provides a method for detecting small objects in commercial vehicles based on an improved YOLO v9. Compared with the existing technology, the embodiments of this application have the following beneficial effects:

[0046] By fusing low-level detail features and high-level semantic features, the detailed information of the image is retained, which helps to detect the contours of tiny objects, enhances the ability to understand complex scenes, and helps to distinguish targets of different categories; through weight adjustment, the contribution of low-level detail features and high-level semantic features can be flexibly balanced according to the needs of different scenarios, thereby improving the adaptability of the model; through parallel multi-branch convolution and lightweight attention mechanism, spatial feature data of different scales are dynamically weighted and fused, thereby improving the detection accuracy of tiny objects and reducing dependence on external data enhancement. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 Schematic diagram of the steps of a method for constructing a small target detection model provided by a preferred embodiment of the present invention;

[0048] Figure 2 The figure is a schematic structural diagram of a small object detection system for commercial vehicles based on an improved YOLO v9, provided in a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0049] The following is a detailed explanation of the embodiments of the present invention in conjunction with the accompanying drawings. The embodiments are provided for illustrative purposes only and cannot be understood as limitations on the present invention. The accompanying drawings are for reference and illustration purposes only and do not constitute a limitation on the scope of patent protection of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. In the description of the present invention, the terms "first", "second", "third", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first", "second", "third", etc. may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, the meaning of "multiple" is two or more.

[0050] In the description of the present invention, it should be noted that, unless otherwise expressly specified and limited, the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be a communication between the two components. The terms "vertical", "horizontal", "left", "right", "up", "down" and similar expressions used herein are for illustrative purposes only, and do not indicate or imply that the device or component referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0051] In describing the present invention, it should be noted that, unless otherwise defined, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. Those skilled in the art will understand the specific meanings of the above terms in the present invention in specific circumstances.

[0052] In intelligent transportation systems, commercial vehicles are typically large and navigate complex environments, including highways, city streets, and rural roads. On highways, vehicles travel at high speeds, requiring rapid and accurate detection of small objects such as small vehicles and obstacles ahead and around. On city streets, the system must cope with numerous small objects, including pedestrians and non-motorized vehicles, while also considering small objects such as traffic signs. In autonomous driving technology, efficient and accurate identification of small objects around the vehicle is crucial for safe operation.

[0053] In view of this, in an embodiment of the present invention, a method for detecting small objects in commercial vehicles based on an improved YOLO v9 is provided, the method comprising:

[0054] During the driving process, the collected image data to be detected is input into the small target detection model, and the target detection results are used to correct the vehicle driving function. For the construction process of the small target detection model, please refer to Figure 1 ,include:

[0055] S1. Construct an initial small target detection model based on the YOLO v9 model. The composite backbone network of the initial small target detection model is set to perform multi-level feature mapping on the first multi-scale feature map, and based on the obtained multi-level feature mapping result, perform cross-branch bidirectional feature fusion of low-level detail features and high-level semantic features on the first multi-scale feature map to obtain first fused feature maps of different scales. The first multi-scale feature map is obtained by the backbone network performing multi-scale feature extraction on the image data to be detected. The first fused feature map is used to input the backbone network; YOLO The v9 model (YOLO 9th generation model) is the latest generation of real-time target detection models in the YOLO series. It is composed of a backbone network, a neck network and a detection head. In this application, it also includes a composite backbone network (CBNet). The backbone network is mainly responsible for extracting features of different scales from the input image data to be detected, and obtaining basic features of different levels. In this application, it includes P1 feature map, P2 feature map, P3 feature map, P4 feature map and P5 feature map. The P1 feature map represents a 1 / 2 scale feature map with 64 channels, the P2 feature map represents a 1 / 4 scale feature map with 128 channels, the P3 feature map represents a 1 / 8 scale feature map with 256 channels, the P4 feature map represents a 1 / 16 scale feature map with 512 channels, and the P5 feature map represents a 1 / 32 scale feature map with 1024 channels. The neck network further processes and fuses the feature maps output by the backbone network. Through a series of convolution, upsampling and downsampling operations, it fuses feature maps of different scales to provide the detection head with richer and more representative feature information. The detection head predicts the category and position of the target based on the features output by the neck network. The detection head contains structures such as Upsample+Concat (upsampling and splicing processing) and RepNCSPELAN4 blocks (modules that combine reparameterization, cross-stage partial connections and efficient layer aggregation networks). These structures are used to refine and classify the features and finally output the target detection results. The composite backbone network uses multiple backbone networks of different types or with different parameter settings for feature extraction. By connecting multiple backbone networks in parallel and fusing features at different levels, it can fully utilize the advantages of each backbone network to extract richer and more representative features. In the existing YOLO v9 model, a feature pyramid is constructed and multi-layer feature maps are used to detect objects of different scales. However, the feature pyramid requires a large amount of computing resources to support the complex network structure and training process. In practical applications, especially scenarios with high real-time requirements such as autonomous driving, the feature pyramid cannot meet the requirements of computational efficiency.

[0056] In an embodiment of the present application, a cross-layer bidirectional feature fusion module is proposed. The cross-layer bidirectional feature fusion module includes a CBLinear module (channel balanced linear module) and a CBFuse module (cross-branch fusion module). The CBLinear module is a linear transformation module that combines cross-branch feature interaction and channel dimension optimization to enhance feature expression capabilities while reducing computational redundancy. The CBFuse module processes features from different levels (such as P3 feature maps, P4 feature maps, and P5 feature maps) simultaneously, automatically balances the contributions of each branch through learnable weights, and uses grouped convolution and channel compression to reduce the amount of computation. During the processing, the first multi-scale feature map extracted by the backbone network is subjected to multi-level feature mapping processing to obtain a multi-level feature mapping result, and the first multi-scale feature map is subjected to cross-branch bidirectional feature fusion based on the multi-level feature mapping result to obtain a first fused feature map of different scales.

[0057] Specifically, the P1 feature map and P2 feature map extracted by the backbone network mainly capture low-level detail features, such as edges and textures, and contain less semantic information. They are not used for small target detection in this application. This application mainly uses P3 feature maps, P4 feature maps, and P5 feature maps for correlation processing and small target detection. A first feature map to be fused and a second feature map to be fused are arbitrarily selected from the P3 feature map, P4 feature map, and P5 feature map output by the backbone network; a linear projection method is used to map the first channel number of the first feature map to be fused to the second channel number of the second feature map to be fused, thereby obtaining a first multi-level feature mapping result; a linear projection method is used to map the second channel number of the second fused feature map to the first channel number of the first fused feature map, thereby obtaining a second multi-level feature mapping result; based on the first multi-level feature mapping result, the first feature map to be fused and the second feature map to be fused are bottom-up channel fused, and based on the second multi-level feature mapping result, the first feature map to be fused and the second feature map to be fused are top-down channel fused to obtain first fused feature maps of different scales.

[0058] Specifically, the CBLinear module performs multi-level feature mapping processing on the P3 feature map, P4 feature map and P5 feature map output by the backbone network to obtain a multi-level feature mapping result. It mainly processes the P3 feature map, P4 feature map and P5 feature map, and can arbitrarily select the first feature map to be fused and the second feature map to be fused from the P3 feature map, P4 feature map and P5 feature map. When the first feature map to be fused is the P3 feature map and the second feature map to be fused is the P4 feature map, the CBLinear module projects the 256-channel P3 feature map to 512 channels so as to fuse it with the P4 feature map. When the first feature map to be fused is the P4 feature map and the second feature map to be fused is the P5 feature map, the CBLinear module projects the 512-channel P4 feature map to 1024 channels to achieve channel alignment with the P5 feature map. As a higher-level feature, the P5 feature map needs to project its own channel number to the low-level channel number when it needs to be fused with low-level features. For example, the 1024 channels of the P5 feature map are projected to 512 channels to facilitate fusion with the P4 feature map, and the 1024 channels of the P5 feature map are projected to 256 channels to facilitate fusion with the P3 feature map.

[0059] Furthermore, the CBFuse module is used to perform cross-branch bidirectional feature fusion of low-level detail features and high-level semantic features on the first multi-scale feature map to obtain first fused feature maps of different scales, thereby enhancing the initial small target detection model's ability to detect targets, especially in complex scenes and small targets. The P3 feature map includes rich low-level detail features but less semantic information. The P4 and P5 feature maps include rich semantic features but less low-level detail features. When performing feature fusion, the nearest neighbor interpolation algorithm is used to upsample the P3 feature map so that it is aligned with the P4 or P5 feature map in the spatial dimension. The maximum pooling algorithm is used to downsample the P4 or P5 feature map so that it is aligned with the P3 feature map in the spatial dimension. In the channel dimension, channel alignment is achieved based on the projection mapping results of the CBLinear module. Furthermore, the aligned low-level features and high-level features are added according to the feature map weights to achieve fusion of the low-level features and high-level features, thereby obtaining first fused features of different scales. The feature map weight is obtained by the CBFuse module through learning the first feature to be fused and the second feature to be fused to balance the contribution between low-level detail features and high-level semantic features.

[0060] This application retains the detailed information of the image by fusing low-level detail features and high-level semantic features, helps detect the contours of tiny objects, enhances the ability to understand complex scenes, and helps distinguish targets of different categories. By adjusting the feature map weights, it can flexibly balance the contributions of low-level detail features and high-level semantic features according to the needs of different scenarios, thereby improving the adaptability of the model.

[0061] S2. Optimize the detection head of the initial small target detection model to obtain a small target detection model after the detection head is optimized. The optimized detection head is set to perform parallel multi-branch convolution processing, first feature fusion and lightweight attention processing on the first multi-scale feature map in sequence to obtain an attention weight matrix. According to the attention weight matrix, the spatial feature data of different scales obtained by the parallel multi-branch convolution processing is adaptively and dynamically weighted to obtain a third fused feature map. The third fused feature map is used for upsampling and splicing processing. In a preferred embodiment of the present application, a selective kernel attention mechanism (SKConv) is proposed. SKConv is added to the detection head of the initial small target detection model, specifically between the Upsample+Concat(P4, P3) (P4 feature map upsampling and P3 feature map channel splicing fusion) module and the Upsample+Concat(P3, P5) (P3 feature map upsampling and P5 feature map channel splicing fusion) module to enhance the feature expression capability. The selective kernel attention mechanism includes parallel multi-branch convolution processing and lightweight attention processing, adaptively selects features of different scales, and improves the detection capability of objects of different scales. Specifically, the parallel multi-branch convolution processing uses parallel branch convolution kernels of different sizes to perform parallel convolution processing respectively. In the preferred real-time example of the present application, M parallel branch convolutions are included, and each branch convolution uses convolution kernels of different sizes, such as 3x3, 5x5, ┄, (3+2*(M-1))x(3+2*(M-1)). The selected first multi-scale feature map is subjected to convolution processing of different scales to obtain multi-scale hierarchical feature data. Furthermore, the multi-scale hierarchical feature data is batch normalized, and the ReLU activation function is used to perform element-by-element threshold processing on the batch-normalized multi-scale hierarchical feature data to obtain spatial feature data of different scales. Furthermore, the spatial feature data of different scales are subjected to a first feature fusion, and the outputs of each branch convolution are added element by element to obtain a second fused feature map. The lightweight attention processing adopts the lightweight attention mechanism, and performs global average pooling, channel dimensionality reduction and attention mapping on the second fusion feature map in sequence to obtain the attention weight matrix. The channel dimensionality reduction process compresses the number of channels to d through the first fully connected layer. The calculation formula of d is:

[0062] d=max(C / r,L)

[0063] Among them, C represents the number of channels of the second fusion feature map, r represents the compression ratio, and L represents the minimum dimension.

[0064] The attention mapping process batch generates M×C attention weight matrices through the second fully connected layer. The attention weight matrix represents the importance of each channel on different branch convolutions. Based on the attention weight matrix, the spatial feature data of different scales obtained by parallel multi-branch convolution processing are dynamically weighted and fused. Before dynamic weighted fusion, the attention weight matrix needs to be normalized based on the Softmax function. The spatial feature data of different scales obtained by each branch convolution processing are then weighted and summed channel by channel according to the weight values corresponding to the attention weight matrix to obtain the third fused feature. This third fused feature map is input into the subsequent Upsample+Concat(P3,P5) module for further upsampling and feature map channel splicing and fusion.

[0065] In a preferred embodiment of the present application, features of different scales are extracted from the first multi-scale feature map using parallel multi-branch convolution kernels of different sizes, replacing the cyclic splicing of the original YOLO v9 model. This enhances the network's ability to detect objects of different scales while reducing loop operations. A lightweight attention mechanism adaptively selects spatial feature data processed by different branch convolution kernels to improve feature expression. Dynamically weighted fusion of spatial feature data of different scales improves the detection accuracy of small objects, optimizing feature expression from within the model and reducing reliance on external data enhancement.

[0066] S3. Train the small target detection model after the detection head is optimized to obtain the small target detection model; after obtaining the small target detection model after the detection head is optimized, collect sample image data, and annotate the sample data image with small target text to form label text data. In this application, sample image data related to transportation are mainly collected, and the sample image data mainly includes traffic signs, small-sized goods, small animals, pedestrians and vehicles. A training data set is constructed based on the sample image data and the label text data, and the sample image data is used as the input of the small target detection model after the detection head is optimized, and the label text data is used as the output of the small target detection model after the detection head is optimized, so as to train the small target detection model after the detection head is optimized to obtain a small target detection model.

[0067] The trained small target detection model is embedded in the vehicle control system. During driving, real-time image data collected to be detected is fed into the small target detection model to perform small target detection on the image data, accurately identifying small targets in the image data that could affect the vehicle's safe driving. Furthermore, control signals are generated based on the target detection results to modify the vehicle's driving functions. For example, during autonomous driving, the control signal is used as an input to the autonomous driving control system to modify autonomous driving instructions.

[0068] Before inputting the collected image data to be detected into the small target detection model and using the obtained target detection results to correct the vehicle's driving function, the collected image data to be detected needs to be preprocessed. Specifically, the collected image data to be detected is sequentially subjected to resolution enhancement processing, denoising filtering processing, contrast enhancement processing, and normalization processing to obtain standardized image data to be detected. Resolution enhancement processing can improve the resolution of the image data to be detected, avoiding the loss of details of small targets due to low resolution; denoising filtering processing can use non-local mean denoising and wavelet denoising to retain the edge information of low-quality image data; contrast enhancement processing can highlight the edge features of small targets; and normalization processing uses a normalization method to scale the pixel values of the image data to a distribution of [0,1].

[0069] Furthermore, saliency detection is performed on the standardized image data to be detected to obtain the image data to be detected including the potential target area. The potential target area in the standardized image data to be detected is preliminarily located through saliency detection, thereby reducing the amount of calculation.

[0070] In a preferred embodiment of the present application, the method also includes: optimizing the backbone network of the small target detection model after the detection head is optimized to obtain the target detection model after the backbone network is optimized, and the optimized backbone network is set to perform reparameterized convolution processing and multi-branch convolution processing on the first multi-scale feature map in sequence through the first convolution branch to obtain local spatial feature data, and perform cross-stage partial connection on the first multi-scale feature map through the second convolution branch to obtain a second multi-scale feature map and a third multi-scale feature map, and perform dense convolution processing on the second multi-scale feature map and directly jump connect it with the third multi-scale feature map to obtain a fourth fused feature map, and perform second feature fusion on the local spatial feature data and the fourth fused feature map to obtain a fifth fused feature map, and the fifth fused feature map is used to input into the detection head.

[0071] The RepNCSPELAN4 module is a neural network module that combines reparameterization, cross-stage partial connections and efficient layer aggregation networks. It has the characteristics of reparameterization design, cross-stage feature fusion, efficient calculation and multi-scale feature extraction. In the YOLO v9 model, the RepNCSPELAN4 module is located in the neck network. In the preferred embodiment of the present application, it is improved and integrated into the backbone network, specifically inserted after the key downsampling layer of the backbone network. The improved RepNCSPELAN4 module includes a first convolution branch and a second convolution branch. In the first convolution branch, a reparameterized convolution structure is used to replace the standard convolution structure. The multi-branch convolution structure in the training phase is fused with the equivalent single branch in the inference phase to balance feature expression and computational efficiency. Each branch of the multi-branch convolution structure uses convolution kernels of different sizes to capture features of different scales, thereby enhancing the detection ability of the small target detection model for multi-scale targets. The first convolution branch is used to perform reparameterized convolution and multi-branch convolution on the first multi-scale feature map (such as P3 feature map, P4 feature map and P5 feature map) output by the key downsampling layer to obtain local spatial feature data. The second convolution branch divides the first multi-scale feature map output by the key downsampling layer into two parts through cross-stage partial connections, namely the second multi-scale feature map and the third multi-scale feature map. The second multi-scale feature map is densely convolved with the dense convolution block and then directly jump-connected with the third multi-scale feature map to obtain the fourth fused feature map, which enhances gradient propagation and reduces computational redundancy, thereby reducing computational cost and capturing richer feature information. Finally, the local spatial feature data and the fourth fused feature map are subjected to the second feature fusion to obtain the fifth fused feature map, which is input into the detection head for small target detection.

[0072] The improved RepNCSPELAN4 module improves the detection accuracy of multi-scale targets through re-parameterized convolution processing and multi-branch convolution processing. Direct skip connections enhance gradient propagation and reduce computational redundancy, thereby reducing computational costs and capturing richer feature information. While improving the accuracy of small target detection, it makes the network structure of the small target detection model more concise and can run efficiently even on resource-constrained platforms such as embedded and mobile devices.

[0073] In a preferred embodiment of the present invention, during the vehicle driving process, the collected image data to be detected is input into the small target detection model, and the target detection result obtained is used to correct the vehicle driving function. The construction process of the small target detection model includes: according to YOLO The v9 model constructs an initial small target detection model. The composite backbone network of the initial small target detection model is set to perform multi-level feature mapping processing on the first multi-scale feature map, and based on the obtained multi-level feature mapping results, the first multi-scale feature map is subjected to cross-branch bidirectional feature fusion of low-level detail features and high-level semantic features to obtain first fused feature maps of different scales. The first multi-scale feature map is obtained by the backbone network performing multi-scale feature extraction on the image data to be detected, and the first fused feature map is used to input the backbone network; the detection head of the initial small target detection model is optimized to obtain a small target detection model after the detection head is optimized. The optimized detection head is set to perform parallel multi-branch convolution processing, first feature fusion and lightweight attention processing on the first multi-scale feature map in sequence to obtain an attention weight matrix. According to the attention weight matrix, the spatial feature data of different scales obtained by the parallel multi-branch convolution processing are adaptively dynamically weighted and fused to obtain a third fused feature map. The third fused feature map is used for upsampling and splicing processing; the small target detection model after the detection head is optimized is trained to obtain a small target detection model. The tiny object detection method for commercial vehicles based on improved YOLO v9 disclosed in this application retains the detailed information of the image by fusing low-level detail features and high-level semantic features, helps to detect the contours of tiny objects, enhances the ability to understand complex scenes, and helps to distinguish targets of different categories; through weight adjustment, it can flexibly balance the contributions of low-level detail features and high-level semantic features according to the needs of different scenarios, thereby improving the adaptability of the model; through parallel multi-branch convolution and lightweight attention mechanism, spatial feature data of different scales are dynamically weighted and fused, thereby improving the detection accuracy of tiny objects and reducing dependence on external data enhancement.

[0074] Accordingly, if Figure 2 As shown, based on a small object detection method for commercial vehicles based on an improved YOLO v9, an embodiment of the present invention further provides a small object detection system for commercial vehicles based on an improved YOLO v9, which implements the small object detection method for commercial vehicles based on an improved YOLO v9 disclosed in an embodiment of the present invention. The system includes: a target detection module 1;

[0075] The target detection module 1 is used to input the collected image data to be detected into the small target detection model during the vehicle driving process, and to correct the vehicle driving function with the target detection results obtained;

[0076] The target detection module includes 1: a first model optimization submodule 11, a second model optimization submodule 12 and a model training submodule 13;

[0077] The first model optimization submodule 11 is used to construct an initial small target detection model based on the YOLO v9 model. The composite backbone network of the initial small target detection model is set to perform multi-level feature mapping processing on the first multi-scale feature map, and based on the obtained multi-level feature mapping result, perform cross-branch bidirectional feature fusion of low-level detail features and high-level semantic features on the first multi-scale feature map to obtain first fused feature maps of different scales. The first multi-scale feature map is obtained by the backbone network performing multi-scale feature extraction on the image data to be detected. The first fused feature map is used to input the backbone network;

[0078] The second model optimization submodule 12 is used to optimize the detection head of the initial small target detection model to obtain a small target detection model after the detection head is optimized. The optimized detection head is set to perform parallel multi-branch convolution processing, first feature fusion and lightweight attention processing on the first multi-scale feature map in sequence to obtain an attention weight matrix. According to the attention weight matrix, adaptive dynamic weighted fusion is performed on the spatial feature data of different scales obtained by the parallel multi-branch convolution processing to obtain a third fused feature map. The third fused feature map is used for upsampling and splicing processing;

[0079] The model training submodule 13 is used to train the small target detection model after the detection head is optimized to obtain the small target detection model.

[0080] Furthermore, the first model optimization submodule 11 includes:

[0081] An image selection unit, configured to arbitrarily select a first feature map to be fused and a second feature map to be fused from the first multi-scale feature map;

[0082] A linear projection unit is configured to map the first number of channels of the first feature map to be fused to the second number of channels of the second feature map to be fused using a linear projection method to obtain a first multi-level feature mapping result, and map the second number of channels of the second feature map to be fused to the first number of channels of the first feature map to be fused using the linear projection method to obtain a second multi-level feature mapping result;

[0083] A first feature fusion unit is configured to perform bottom-up channel fusion on the first feature map to be fused and the second feature map to be fused according to the first multi-level feature mapping result, and to perform top-down channel fusion on the first feature map to be fused and the second feature map to be fused according to the second multi-level feature mapping result, to obtain first fused feature maps of different scales.

[0084] Furthermore, the second model optimization submodule 12 includes:

[0085] A parallel branch convolution processing unit, configured to use parallel branch convolution kernels of different sizes to perform convolution processing of different scales on the first multi-scale feature map to obtain multi-scale hierarchical feature data;

[0086] a normalization threshold processing unit, configured to perform batch normalization processing on the multi-scale hierarchical feature data, and perform element-by-element threshold processing on the multi-scale hierarchical feature data after batch normalization processing using a ReLU activation function to obtain the spatial feature data of different scales;

[0087] A second feature fusion unit is used to perform first feature fusion on the spatial feature data of different scales to obtain a second fused feature map;

[0088] An attention mapping processing unit is used to perform global average pooling processing, channel dimensionality reduction processing and attention mapping processing on the second fused feature map in sequence to obtain an attention weight matrix.

[0089] Furthermore, the target detection module also includes a third model optimization submodule;

[0090] The third model optimization submodule is used to optimize the backbone network of the small target detection model after the detection head is optimized to obtain the target detection model after the backbone network is optimized. The optimized backbone network is set to perform reparameterized convolution processing and multi-branch convolution processing on the first multi-scale feature map in sequence through the first convolution branch to obtain local spatial feature data, and perform cross-stage partial connection on the first multi-scale feature map through the second convolution branch to obtain a second multi-scale feature map and a third multi-scale feature map. After dense convolution processing, the second multi-scale feature map is directly jump-connected with the third multi-scale feature map to obtain a fourth fused feature map. The local spatial feature data and the fourth fused feature map are subjected to second feature fusion to obtain a fifth fused feature map. The fifth fused feature map is used to input into the detection head.

[0091] For the specific definition of a small object detection system for commercial vehicles based on improved YOLO v9, please refer to the above-mentioned definition of a small object detection method for commercial vehicles based on improved YOLO v9, which will not be repeated here. Those of ordinary skill in the art will appreciate that the various modules and steps described in conjunction with the embodiments disclosed in the present invention can be implemented in hardware, software, or a combination of both. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0092] In summary, the embodiment of the present application provides a method and system for detecting small objects in commercial vehicles based on the improved YOLO v9, which solves the technical problem of how to improve the accuracy of small target detection results while reducing computing resource consumption and model training time costs. The method includes: during the vehicle driving process, the collected image data to be detected is input into the small target detection model, and the target detection results obtained are used to correct the vehicle driving function. The construction process of the small target detection model includes: according to YOLO The v9 model constructs an initial small target detection model. The composite backbone network of the initial small target detection model is set to perform multi-level feature mapping processing on the first multi-scale feature map, and based on the obtained multi-level feature mapping results, the first multi-scale feature map is subjected to cross-branch bidirectional feature fusion of low-level detail features and high-level semantic features to obtain first fused feature maps of different scales. The first multi-scale feature map is obtained by the backbone network performing multi-scale feature extraction on the image data to be detected, and the first fused feature map is used to input the backbone network; the detection head of the initial small target detection model is optimized to obtain a small target detection model after the detection head is optimized. The optimized detection head is set to perform parallel multi-branch convolution processing, first feature fusion and lightweight attention processing on the first multi-scale feature map in sequence to obtain an attention weight matrix. According to the attention weight matrix, the spatial feature data of different scales obtained by the parallel multi-branch convolution processing are adaptively dynamically weighted and fused to obtain a third fused feature map. The third fused feature map is used for upsampling and splicing processing; the small target detection model after the detection head is optimized is trained to obtain a small target detection model. The tiny object detection method for commercial vehicles based on improved YOLO v9 disclosed in this application retains the detailed information of the image by fusing low-level detail features and high-level semantic features, helps to detect the contours of tiny objects, enhances the ability to understand complex scenes, and helps to distinguish targets of different categories; through weight adjustment, it can flexibly balance the contributions of low-level detail features and high-level semantic features according to the needs of different scenarios, thereby improving the adaptability of the model; through parallel multi-branch convolution and lightweight attention mechanism, spatial feature data of different scales are dynamically weighted and fused, thereby improving the detection accuracy of tiny objects and reducing dependence on external data enhancement.

[0093] Each embodiment in this specification is described in a progressive manner, and the same or similar parts of each embodiment can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. It should be noted that the various technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0094] The above-described embodiments merely represent several preferred implementations of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person skilled in the art could make several improvements and substitutions without departing from the technical principles of the present application, and such improvements and substitutions should also be considered within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be based on the scope of protection of the claims.

Claims

1. A small object detection method for commercial vehicles based on improved YOLO v9, characterized in that: The method comprises: During vehicle driving, the collected image data to be detected is input into a small target detection model, and the target detection results obtained are used to correct the vehicle driving function. The construction process of the small target detection model includes: An initial small target detection model is constructed based on the YOLO v9 model. The composite backbone network of the initial small target detection model is set to perform multi-level feature mapping processing on a first multi-scale feature map, and based on the obtained multi-level feature mapping result, a cross-branch bidirectional feature fusion of low-level detail features and high-level semantic features is performed on the first multi-scale feature map to obtain first fused feature maps of different scales. The first multi-scale feature map is obtained by performing multi-scale feature extraction on the image data to be detected by the backbone network. The first fused feature map is used as an input to the backbone network; Optimizing the detection head of the initial small target detection model to obtain a small target detection model after the detection head is optimized, wherein the optimized detection head is configured to sequentially perform parallel multi-branch convolution processing, first feature fusion, and lightweight attention processing on the first multi-scale feature map to obtain an attention weight matrix, and adaptively and dynamically weighted fusion is performed on spatial feature data of different scales obtained by the parallel multi-branch convolution processing according to the attention weight matrix to obtain a third fused feature map, wherein the third fused feature map is used for upsampling and splicing processing; The small target detection model after the detection head is optimized is trained to obtain the small target detection model.

2. The method for detecting small objects in commercial vehicles based on the improved YOLO v9 as claimed in claim 1, characterized in that: The performing multi-level feature mapping processing on the first multi-scale feature map, and performing cross-branch bidirectional feature fusion of low-level detail features and high-level semantic features on the first multi-scale feature map according to the obtained multi-level feature mapping result to obtain first fused feature maps of different scales, including: arbitrarily selecting a first feature map to be fused and a second feature map to be fused from the first multi-scale feature map; Using a linear projection method to map the first channel number of the first feature map to be fused to the second channel number of the second feature map to be fused, obtaining a first multi-level feature mapping result; using the linear projection method to map the second channel number of the second feature map to be fused to the first channel number of the first feature map to be fused, obtaining a second multi-level feature mapping result; According to the first multi-level feature mapping result, the first feature map to be fused and the second feature map to be fused are subjected to bottom-up channel fusion, and according to the second multi-level feature mapping result, the first feature map to be fused and the second feature map to be fused are subjected to top-down channel fusion to obtain first fused feature maps of different scales.

3. The method for detecting small objects in commercial vehicles based on the improved YOLO v9 as claimed in claim 1, characterized in that: The step of sequentially performing parallel multi-branch convolution processing, first feature fusion, and lightweight attention processing on the first multi-scale feature map to obtain an attention weight matrix includes: Using parallel branch convolution kernels of different sizes, perform convolution processing of different scales on the first multi-scale feature map to obtain multi-scale hierarchical feature data; Performing batch normalization processing on the multi-scale hierarchical feature data, and performing element-by-element threshold processing on the multi-scale hierarchical feature data after the batch normalization processing using a ReLU activation function to obtain the spatial feature data of different scales; Performing a first feature fusion on the spatial feature data of different scales to obtain a second fused feature map; The second fused feature map is sequentially subjected to global average pooling, channel dimensionality reduction, and attention mapping to obtain an attention weight matrix.

4. The method for detecting small objects in commercial vehicles based on the improved YOLO v9 as claimed in claim 1, wherein: The method further comprises: The backbone network of the small target detection model after the detection head is optimized is optimized to obtain the target detection model after the backbone network is optimized. The optimized backbone network is set to perform reparameterized convolution processing and multi-branch convolution processing on the first multi-scale feature map in sequence through the first convolution branch to obtain local spatial feature data, and the first multi-scale feature map is partially connected across stages through the second convolution branch to obtain a second multi-scale feature map and a third multi-scale feature map. After dense convolution processing, the second multi-scale feature map is directly jump-connected with the third multi-scale feature map to obtain a fourth fused feature map. The local spatial feature data and the fourth fused feature map are subjected to second feature fusion to obtain a fifth fused feature map. The fifth fused feature map is used to input into the detection head.

5. The method for detecting small objects in commercial vehicles based on improved YOLO v9 as claimed in claim 1, wherein: Before inputting the collected image data to be detected into the small target detection model to correct the vehicle driving function with the obtained target detection results, the method further includes: The collected image data to be detected are sequentially subjected to resolution enhancement processing, denoising filtering processing, contrast enhancement processing and normalization processing to obtain standardized image data to be detected; Saliency detection is performed on the standardized image data to be detected to obtain the image data to be detected including a potential target area.

6. The method for detecting small objects in commercial vehicles based on improved YOLO v9 as claimed in claim 1, characterized in that: The step of training the small target detection model after the detection head is optimized to obtain the small target detection model includes: Collecting sample image data, and annotating the sample image data with tiny target text to obtain label text data; Constructing a training data set based on the sample image data and the label text data; The training data set is used to train the small target detection model after the detection head is optimized to obtain the small target detection model.

7. A small object detection system for commercial vehicles based on an improved YOLO v9, used to implement the small object detection method for commercial vehicles based on an improved YOLO v9 according to any one of claims 1 to 6, characterized in that: The system includes: a target detection module; The target detection module is used to input the collected image data to be detected into the small target detection model during the vehicle driving process, and to correct the vehicle driving function with the target detection results obtained; The target detection module includes: a first model optimization submodule, a second model optimization submodule and a model training submodule; The first model optimization submodule is used to construct an initial small target detection model based on the YOLO v9 model. The composite backbone network of the initial small target detection model is set to perform multi-level feature mapping processing on the first multi-scale feature map, and based on the obtained multi-level feature mapping results, perform cross-branch bidirectional feature fusion of low-level detail features and high-level semantic features on the first multi-scale feature map to obtain first fused feature maps of different scales. The first multi-scale feature map is obtained by the backbone network performing multi-scale feature extraction on the image data to be detected. The first fused feature map is used to input the backbone network; The second model optimization submodule is used to optimize the detection head of the initial small target detection model to obtain a small target detection model after the detection head is optimized. The optimized detection head is set to perform parallel multi-branch convolution processing, first feature fusion and lightweight attention processing on the first multi-scale feature map in sequence to obtain an attention weight matrix. According to the attention weight matrix, adaptive dynamic weighted fusion is performed on the spatial feature data of different scales obtained by the parallel multi-branch convolution processing to obtain a third fused feature map. The third fused feature map is used for upsampling and splicing processing; The model training submodule is used to train the small target detection model after the detection head is optimized to obtain the small target detection model.

8. The small object detection system for commercial vehicles based on the improved YOLO v9 as claimed in claim 7, characterized in that: The first model optimization submodule includes: An image selection unit, configured to arbitrarily select a first feature map to be fused and a second feature map to be fused from the first multi-scale feature map; A linear projection unit is configured to map the first number of channels of the first feature map to be fused to the second number of channels of the second feature map to be fused using a linear projection method to obtain a first multi-level feature mapping result, and map the second number of channels of the second feature map to be fused to the first number of channels of the first feature map to be fused using the linear projection method to obtain a second multi-level feature mapping result; A first feature fusion unit is configured to perform bottom-up channel fusion on the first feature map to be fused and the second feature map to be fused according to the first multi-level feature mapping result, and to perform top-down channel fusion on the first feature map to be fused and the second feature map to be fused according to the second multi-level feature mapping result, to obtain first fused feature maps of different scales.

9. The small object detection system for commercial vehicles based on the improved YOLO v9 as claimed in claim 7, characterized in that: The second model optimization submodule includes: A parallel branch convolution processing unit, configured to use parallel branch convolution kernels of different sizes to perform convolution processing of different scales on the first multi-scale feature map to obtain multi-scale hierarchical feature data; a normalization threshold processing unit, configured to perform batch normalization processing on the multi-scale hierarchical feature data, and perform element-by-element threshold processing on the multi-scale hierarchical feature data after batch normalization processing using a ReLU activation function to obtain the spatial feature data of different scales; A second feature fusion unit is used to perform first feature fusion on the spatial feature data of different scales to obtain a second fused feature map; An attention mapping processing unit is used to perform global average pooling processing, channel dimensionality reduction processing and attention mapping processing on the second fused feature map in sequence to obtain an attention weight matrix.

10. The small object detection system for commercial vehicles based on the improved YOLO v9 as claimed in claim 7, characterized in that: The target detection module also includes a third model optimization submodule; The third model optimization submodule is used to optimize the backbone network of the small target detection model after the detection head is optimized to obtain the target detection model after the backbone network is optimized. The optimized backbone network is set to perform reparameterized convolution processing and multi-branch convolution processing on the first multi-scale feature map in sequence through the first convolution branch to obtain local spatial feature data, and perform cross-stage partial connection on the first multi-scale feature map through the second convolution branch to obtain a second multi-scale feature map and a third multi-scale feature map. After dense convolution processing, the second multi-scale feature map is directly jump-connected with the third multi-scale feature map to obtain a fourth fused feature map. The local spatial feature data and the fourth fused feature map are subjected to second feature fusion to obtain a fifth fused feature map. The fifth fused feature map is used to input into the detection head.

Citation Information

Cited By

  • Real-time detection and tracking method and device for water athletes and medium

    CN120747171A

  • A real-time detection and tracking method, device and medium for water sports players

    CN120747171B