A multi-scale road vehicle detection method based on reparameterization visual transformer

By employing a reparameterized visual converter and a multi-scale feature fusion network, the problems of small target detection in complex environments and high model computation costs are solved, achieving efficient and accurate road vehicle detection.

CN120564164BActive Publication Date: 2025-11-21NANCHANG MONI SOFTWARE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511075354.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-21
Estimated Expiration
2045-08-01

AI Technical Summary

Technical Problem

Existing road vehicle detection algorithms suffer from low accuracy in detecting small targets in complex environments, high model computation costs, and insufficient multi-scale feature fusion capabilities, making it difficult to meet the requirements for real-time performance and robustness.

Method used

We replace the feature extraction network of the object detection neural network with a reparameterized visual converter, and combine a multi-branch structure and a multi-scale feature fusion network to enhance feature representation through spatial feature mixing and channel feature mixing. We also use bidirectional path feature fusion from bottom to top and from top to bottom, and combine an improved squeeze-excited attention module and SF-DIoU loss function to optimize model performance.

Benefits of technology

It improves the model's detection performance in complex scenes and small target detection, reduces inference time, enhances real-time detection capabilities and the ability to identify small targets, and strengthens the model's generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564164B_ABST
    Figure CN120564164B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-scale road vehicle detection methods based on reparameterization vision converter, comprising the following steps: selecting road vehicle dataset as the dataset of improved network model input end;Improved network model adopts series-parallel structure, and is divided into feature extraction network, multi-scale feature fusion network, detection layer;The initial feature map of processed road vehicle dataset is input to feature extraction network and is preprocessed to obtain extraction feature map: extraction feature map is input to multi-scale feature fusion network and outputs fusion feature map;Fusion feature map is input to detection layer and is processed, and the different size targets needing to be identified in initial feature map are predicted.This application has the beneficial effects that: reparameterization vision converter network more effectively captures long-distance dependence relationship, while maintaining the efficient processing of local receptive field, better adapts to the needs of target detection task, and is easier to adjust and optimize, improves the performance of model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a road vehicle detection technology, in particular to a multi-scale road vehicle detection method based on a reparameterization visual transformer. BACKGROUND

[0002] With the rapid progress of artificial intelligence technology, intelligent transportation systems are accelerating the reconstruction of modern traffic patterns. Autonomous driving technology, as a core technology in the intelligent transportation system, has gradually entered the public view and has made significant progress. Road vehicle detection and recognition plays a crucial role in autonomous driving technology and has a direct impact on intelligent driving decisions. Accurate vehicle detection can provide low-latency real-time road vehicle information for intelligent driving systems, improving vehicle driving safety. It can be seen that in practical applications, accurate and efficient road vehicle detection technology is an important component of modern transportation systems and provides an important guarantee for autonomous driving.

[0003] However, in the face of actual complex traffic scenarios, how to achieve accurate, efficient, and automatic road vehicle detection still faces many problems. Identifying vehicles in complex environments can face problems such as low light, occlusion, or complex background, increasing the difficulty of identification. And the variety and size of vehicles vary greatly, especially for long-distance blurred detection and recognition, existing methods are not sufficient to provide accuracy and robustness.

[0004] Currently, road vehicle detection and recognition algorithms mainly include traditional algorithms and deep learning-based algorithms. Under the background of rapid development of artificial intelligence, deep learning-based methods have gradually become the mainstream. Through automatic feature learning, end-to-end optimization, multi-scale processing, and other technologies, they have surpassed traditional methods in precision, speed, and adaptability, and continue to break through performance boundaries through architecture innovation and training strategy optimization. However, in actual application scenarios, many challenges still exist: first, the accuracy of small target detection and in complex environments is one of the difficulties; second, the increase in model training parameters increases the computational cost, resulting in a decrease in detection speed, especially on devices with limited computing resources, reducing deployment costs and improving model real-time performance are problems that need to be solved. In the case of resolution, the loss of target feature expression information will lead to a decrease in recognition accuracy in complex environments, so improving the ability of multi-scale feature fusion is also a task that should be completed.

[0005] In summary, current road vehicle detection algorithms need to be optimized and improved in terms of small target detection in complex environments, model lightweight, multi-scale feature fusion, and enhancement to meet the needs of model accuracy, efficiency, and robustness in complex environments. SUMMARY

[0006] The purpose of the present application is to provide a multi-scale road vehicle detection method based on a reparameterization visual converter, which replaces the feature extraction network of the target detection neural network with a reparameterization visual converter network, more effectively captures long-distance dependencies, while maintaining the efficient processing of local receptive fields, better adapts to the needs of the target detection task, and is easier to adjust and optimize, improving the performance of the model.

[0007] The technical scheme adopted by the present application is as follows: a multi-scale road vehicle detection method based on a reparameterization visual converter, comprising the following steps:

[0008] Step S1, data set cleaning and construction; select KITTI 2D road vehicle data set, clean, screen and detect KITTI 2D road vehicle data set, and obtain processed road vehicle data set, which is used as the input data set of the improved network model;

[0009] Step S2, improved network model; the improved network model adopts a series-parallel structure and is divided into a feature extraction network, a multi-scale feature fusion network and a detection layer; the feature extraction network is composed of an input adaptation layer and a multi-branch structure reparameterization extraction module; the multi-scale feature fusion network includes a 1x1 convolution module, a 3x3 convolution module, a dynamic up-sampling module, a region enhanced gradient pooling attention module, a splicing module, an improved squeeze and excitation attention module and a C3k2 module;

[0010] Step S3, feature extraction network processing process; the initial feature map of the road vehicle data set processed in step S1 is input into the input adaptation layer in the feature extraction network for preprocessing to obtain a first extracted feature map, which is input into the multi-branch structure reparameterization extraction module in the feature extraction network for preprocessing to obtain a third extracted feature map, a fifth extracted feature map and a seventh extracted feature map:

[0011] Step S4, multi-scale feature fusion network processing process; adopt a bottom-up and top-down dual-path feature fusion; input the third extracted feature map, the fifth extracted feature map and the seventh extracted feature map into the multi-scale feature fusion network to obtain an eleventh fused feature map, a thirteenth fused feature map and a fifteenth fused feature map;

[0012] Step S5, detection layer processing process; input the eleventh fused feature map, the thirteenth fused feature map and the fifteenth fused feature map into the detection layer composed of three detection heads respectively; the detection layer processes the eleventh fused feature map, the thirteenth fused feature map and the fifteenth fused feature map respectively, and completes the prediction of different size targets in the initial feature map that need to be recognized.

[0013] Further, in step S3, the initial feature map of the road vehicle data set processed in step S1 is input to the input adaptation layer in the feature extraction network for preprocessing to obtain a first extracted feature map after preprocessing; specifically:

[0014] In step S301, the input adaptation layer includes two convolution normalization modules, which are divided into a first convolution normalization module and a second convolution normalization module; wherein the convolution normalization module is composed of a convolution layer and a batch normalization layer, which is used to define the resolution and channel number of the input feature map adjusted by several convolution modules in the improved network model;

[0015] In step S302, the first convolution normalization module reduces the channel number of the initial feature map to half of the initial feature map, and down-samples to a resolution of 320x320, and uses the activation function GELU for nonlinear transformation;

[0016] In step S303, the second convolution normalization module restores the channel number of the feature map output from the first convolution normalization module to the channel number of the initial feature map, and down-samples to a resolution of 160x160; obtaining a first extracted feature map after preprocessing, input to a multi-branch structure reparameterization extraction module.

[0017] Further, in step S3, the first extracted feature map is input to the multi-branch structure reparameterization extraction module in the feature extraction network for preprocessing to obtain a third extracted feature map, a fifth extracted feature map and a seventh extracted feature map; specifically:

[0018] In step S304, the multi-branch structure reparameterization extraction module is improved by a reparameterization visual converter, which is composed of 3 reparameterization visual converter blocks, 2 stage down-sampling blocks and a preset built-in parameter table;

[0019] The reparameterization visual converter block is composed of a spatial feature mixing module and a channel feature mixing module;

[0020] The preset built-in parameter table is preset by the user, when the spatial feature mixing is performed, the spatial feature mixing module includes a reparameterization depth separable convolution and a depth separable convolution, and the spatial feature mixing module is provided with a parameter step stride, which will select different depth separable convolutions according to the parameter step stride given in the preset built-in parameter table;

[0021] The spatial feature mixing module is also provided with a channel attention control parameter use_se, which determines whether to use a channel attention module according to the parameter in the preset built-in parameter table;

[0022] Step S305, the first extracted feature map preprocessed by the input adaptation layer is input into the first reparameterized visual transformer block; the first reparameterized visual transformer block performs two-step mixed processing on the spatial features and channel features of the first extracted feature map;

[0023] Step S306, the spatial feature mixing module is used for spatial feature mixing of the first extracted feature map; the spatial feature mixing module includes a 3x3 normalized depth separable convolution defined by a convolution normalization module, a feature extraction channel attention module, a 1x1 normalized point convolution defined by a convolution normalization module, and a reparameterized depth separable convolution;

[0024] The reparameterized depth separable convolution is composed of a 3x3 depth separable convolution defined by a normal convolution layer, and a point convolution defined by a normal 1x1 convolution layer; the first extracted feature map after preprocessing by the input adaptation layer is processed by the reparameterized visual transformer block:

[0025] Step S307, when the parameter step length stride=1, the spatial feature mixing module uses the reparameterized depth separable convolution and the feature extraction channel attention module; the reparameterized depth separable convolution is used to enhance the local structural features of the first extracted feature map; then when the channel attention control parameter use_se is true, the feature extraction channel attention module is used to enhance the channel of the first extracted feature map; when the channel attention control parameter use_se is false, the first extracted feature map is not processed by the feature extraction channel attention module, and is directly output as the second extracted feature map; finally, the second extracted feature map after spatial feature mixing is obtained;

[0026] Step S308, when the parameter step length stride=2, the spatial feature mixing module will enable the 3x3 normalized depth separable convolution, the feature extraction channel attention module and the 1x1 normalized point convolution, which are used for adjusting the channel number of the extracted feature map; the 3x3 normalized depth separable convolution is used to downsample the first extracted feature map to 1 / 8 resolution, the feature extraction channel attention module is used for channel enhancement, and the 1x1 normalized point convolution is used to restore the channel dimension to the channel number of the first extracted feature map input;

[0027] Step S309, the channel feature mixing module is used for channel feature mixing of the second extracted feature map; the channel feature mixing module is composed of an up-sampling 1x1 convolution defined by a convolution normalization module, a down-sampling 1x1 convolution defined by a convolution normalization module, and an activation function GELU nonlinear activation function;

[0028] Step S310, the dimension increasing 1x1 convolution expands the second extracted feature map channel number to 2 times of the channel number of the second extracted feature map output by the spatial feature mixing module, the activation function GELU retains negative value information, the dimension reducing 1x1 convolution compresses the channel number of the second extracted feature map back to the channel number when the second extracted feature map is input, and then the second extracted feature map is connected with the first extracted feature map in residual, retains the features of the first extracted feature map which have not been mixed with spatial features and channel features, and obtains a third extracted feature map;

[0029] Step S311, the third extracted feature map is input into the first stage downsampling block to obtain a fourth extracted feature map with a resolution of 1 / 16, and the fourth extracted feature map is input into the second reparameterized visual transformer block to extract local features of the fourth extracted feature map to obtain a fifth extracted feature map;

[0030] Step S312, the fifth extracted feature map is input into the second stage downsampling block to obtain a sixth extracted feature map with a resolution of 1 / 32; the sixth extracted feature map is input into the third reparameterized visual transformer block to perform overall feature correlation calculation to obtain a seventh extracted feature map.

[0031] Further, in step S4, the multi-scale feature fusion network processing process is as follows:

[0032] Step S401, the multi-scale feature fusion network is composed of three 1x1 convolution modules, two 3x3 convolution modules, two dynamic upsampling modules, three regional enhanced gradient pooling attention modules, four splicing modules, one improved squeeze and excitation attention module, and one C3k2 module;

[0033] Step S402, the three 1x1 convolution modules respectively receive the third extracted feature map, the fifth extracted feature map and the seventh extracted feature map output by the three reparameterized visual transformer blocks of the feature extraction network, and unify the channel numbers of the third extracted feature map, the fifth extracted feature map and the seventh extracted feature map to 256; wherein the third extracted feature map is processed by the first 1x1 convolution module to obtain a high-resolution first fusion feature map carrying high-level semantic information, the fifth extracted feature map is processed by the second 1x1 convolution module to obtain a medium-resolution second fusion feature map balancing details and semantics, and the seventh extracted feature map is processed by the third 1x1 convolution module to obtain a low-resolution third fusion feature map carrying deep-level semantic information;

[0034] Step S403, bottom-up feature fusion is performed, the third fusion feature map is processed by the first dynamic upsampling module to a resolution of 1 / 16, and then spliced with the second fusion feature map by the first splicing module to obtain a fourth fusion feature map, and the fourth fusion feature map is input into the first regional enhanced gradient pooling attention module for cross-scale feature fusion;

[0035] The area enhanced gradient pooling attention module is composed of a fourth 1x1 convolution module, two perception patch attention modules, a fusion splicing module and a ReLU regularized activation function; wherein the perception patch attention module is composed of a spatial attention module, two local-global attention modules with different receptive fields, a feature fusion channel attention module, a four-branch 3x3 convolution and a four-branch 1x1 convolution;

[0036] In step S404, the fourth fusion feature map is input into the first perception patch attention module and the second perception patch attention module after being aligned in channels by the fourth 1x1 convolution module. The perception patch attention module has a four-branch structure, which includes a 2-receptive-field local-global attention module, a 4-receptive-field local-global attention module, a four-branch 3x3 convolution and a four-branch 1x1 convolution. The fourth fusion feature map is input into the first perception patch attention module and the second perception patch attention module for four-branch structure processing.

[0037] In step S405, the four-branch structure processing flow of the fourth fusion feature map is as follows: the local attention branches of the 2-receptive-field local-global attention module and the 4-receptive-field local-global attention module extract local features of different scales, which are fused with global features obtained by the global attention branches, and then the fourth fusion feature map is constructed by a four-branch 3x3 convolution to capture local-to-global detailed features and represent multi-level local features. A four-branch 1x1 convolution is used to retain global feature information and local feature information to obtain a multi-source heterogeneous feature map.

[0038] In step S406, the multi-source heterogeneous feature map is input into the feature fusion channel attention module, and then into the spatial attention module to generate a spatial weight mask to highlight small target regions. Finally, the first perception patch attention module processes the fourth fusion feature map to obtain a fifth fusion feature map. Similarly, the fourth fusion feature map is processed by the second perception patch attention module to obtain a sixth fusion feature map. The fifth fusion feature map and the sixth fusion feature map are spliced by the fusion splicing module to obtain a seventh fusion feature map. Finally, the seventh fusion feature map is input into the ReLU regularized activation function to suppress gradient vanishing, and an eighth fusion feature map is obtained. The eighth fusion feature map is output to the second dynamic upsampling module.

[0039] In step S407, the eighth fusion feature map is input into the second dynamic upsampling module, and the eighth fusion feature map is upsampled to 1 / 8 resolution. Then, the first fusion feature map is spliced with the eighth fusion feature map by the second splicing module to obtain a ninth fusion feature map.

[0040] Step S408, the ninth fusion feature map is input into the second region enhancement gradient pooling attention module for further fusion to obtain a tenth fusion feature map.

[0041] Further, in step S4, the top-down feature fusion in the multi-scale feature fusion network processing process is specifically:

[0042] Step S409, the tenth fusion feature map is input into the improved squeeze-and-excitation attention module to obtain an eleventh fusion feature map, and the eleventh fusion feature map is output to the first detection head of the detection layer.

[0043] Step S410, the eleventh fusion feature map obtained by processing the improved squeeze-and-excitation attention module is down-sampled through the first 3x3 convolution module, so that the resolution of the eleventh fusion feature map is reduced to 1 / 16, the channel number of the eleventh fusion feature map is increased to 256, and the eleventh fusion feature map is output. The eleventh fusion feature map is spliced with the eighth fusion feature map obtained by the first region enhancement gradient pooling attention module through the third splicing module to obtain a twelfth fusion feature map, and the twelfth fusion feature map is input into the third region enhancement gradient pooling attention module to obtain a thirteenth fusion feature map. The thirteenth fusion feature map is input into the second detection head in the detection layer.

[0044] Step S411, the thirteenth fusion feature map is down-sampled through the second 3x3 convolution module, so that the resolution of the thirteenth fusion feature map is reduced to 1 / 32, and the channel number of the thirteenth fusion feature map is increased to 512. The thirteenth fusion feature map and the seventh extraction feature map output in step S312 are spliced using the fourth splicing module to obtain a fourteenth fusion feature map, and the fourteenth fusion feature map is input into the C3k2 module to obtain a fifteenth fusion feature map, which is output to the third detection head in the detection layer.

[0045] Further, in step S5, in the detection layer processing process, the detection head loss function is provided with a positioning loss, the error between the predicted bounding box after prediction and the real bounding box is calculated through the SF-DIoU loss function, and is transmitted to the next training to optimize the bounding box coordinates.

[0046] The beneficial effects of the present application are: the present application combines the multi-branch structure reparameterization extraction module of spatial feature mixing and channel feature mixing as the feature extraction network, uses the spatial feature mixing module and the channel feature mixing module of the multi-branch structure to enhance the feature expression during training. The reparameterization deep separable convolution multi-branch extraction feature is used to improve the feature richness and does not increase the inference time. The operation speed of the extracted feature is improved, and the real-time detection capability is improved.

[0047] The application designs a multi-scale feature fusion network, adopts bidirectional path feature fusion from bottom to top and from top to bottom, and realizes the interaction of different scale features.

[0048] In view of the problems of information overload of the high-resolution feature layer of the model feature fusion network in small target detection and insufficient context-dependent modeling, the improved squeeze and excitation attention module is adopted to capture different dimensional channel relationships through four independent fully connected branches, and the weak signal of the small target is enhanced.

[0049] The SF-DIoU loss function is used instead of the original CIoU and DFL loss function to calculate the positioning loss. The SF-DIoU loss function uses the shape perception penalty of shape-IoU to improve the geometric extraction capability. Meanwhile, dynamic stages and scaled IoU values are introduced to further optimize the sample weight distribution; and through dynamic sample weighting, the loss contribution of different difficulty samples is adjusted to further improve the generalization ability of the model, and the combination of the two significantly improves the detection ability of the model for complex targets, especially in high-precision indicators and difficult scenes. BRIEF DESCRIPTION OF DRAWINGS

[0050] Figure 1 The flowchart of the application.

[0051] Figure 2 The regional enhanced gradient pooling attention module structure diagram of the application. DETAILED DESCRIPTION

[0052] As Figure 1 shown, a multi-scale road vehicle detection method based on a reparameterization visual converter includes the following steps:

[0053] Step S1, data set cleaning and construction; select KITTI 2D road vehicle data set, clean, screen and detect KITTI 2D road vehicle data set, and obtain a processed road vehicle data set, the road vehicle data set as an improved network model input data set;

[0054] Step S2, improved network model; the improved network model adopts a series-parallel structure and is divided into a feature extraction network, a multi-scale feature fusion network and a detection layer; the feature extraction network is composed of an input adaptive layer and a multi-branch structure reparameterization extraction module; the multi-scale feature fusion network includes a 1x1 convolution module, a 3x3 convolution module, a dynamic up-sampling module, a region enhanced gradient pooling attention module, a splicing module, an improved squeeze and excitation attention module and a C3k2 module;

[0055] Step S3, feature extraction network processing process; the initial feature map of the road vehicle data set processed in step S1 is input into the input adaptive layer in the feature extraction network for preprocessing to obtain a first extracted feature map, and the first extracted feature map is input into the multi-branch structure reparameterization extraction module in the feature extraction network for preprocessing to obtain a third extracted feature map, a fifth extracted feature map and a seventh extracted feature map;

[0056] Step S4, multi-scale feature fusion network processing process; a bottom-up and top-down bidirectional path feature fusion is adopted; the third extracted feature map, the fifth extracted feature map and the seventh extracted feature map are input into the multi-scale feature fusion network to obtain an eleventh fused feature map, a thirteenth fused feature map and a fifteenth fused feature map;

[0057] Step S5, detection layer processing process; the eleventh fused feature map, the thirteenth fused feature map and the fifteenth fused feature map are respectively input into a detection layer composed of three detection heads; the detection layer processes the eleventh fused feature map, the thirteenth fused feature map and the fifteenth fused feature map to complete the prediction of different size targets in the initial feature map that need to be recognized.

[0058] Further, in step S3, the initial feature map of the road vehicle data set processed in step S1 is input into the input adaptive layer in the feature extraction network for preprocessing to obtain a first extracted feature map; specifically:

[0059] Step S301, the input adaptive layer includes two convolution normalization modules, which are divided into a first convolution normalization module and a second convolution normalization module; wherein the convolution normalization module is composed of a convolution layer and a batch normalization layer, and is used to define the resolution and channel number of the input feature map adjusted by a plurality of convolution modules in the improved network model;

[0060] In step S302, the first convolution normalization module reduces the initial feature map channel number to half of the initial feature map channel number and down-samples to a resolution of 320x320, and performs nonlinear transformation by using an activation function GELU.

[0061] In step S303, the second convolution normalization module restores the feature map channel number output from the first convolution normalization module to the initial feature map channel number and down-samples to a resolution of 160x160, to obtain a first extracted feature map after preprocessing, which is input to a multi-branch structure reparameterization extraction module.

[0062] Further, in step S3, the first extracted feature map is input to the multi-branch structure reparameterization extraction module in the feature extraction network for preprocessing, to obtain a third extracted feature map, a fifth extracted feature map and a seventh extracted feature map; specifically:

[0063] In step S304, the multi-branch structure reparameterization extraction module is improved by a reparameterization visual transformer, which is composed of 3 reparameterization visual transformer blocks, 2 stage down-sampling blocks and a preset built-in parameter table.

[0064] The reparameterization visual transformer block is composed of a spatial feature mixing module, a channel feature mixing module and a residual connection.

[0065] The preset built-in parameter table is preset by a user, and when spatial feature mixing is performed, the spatial feature mixing module includes a reparameterization depth separable convolution and a depth separable convolution, and the spatial feature mixing module is provided with a parameter step stride, which will select different depth separable convolutions according to the parameter step stride given in the preset built-in parameter table.

[0066] The spatial feature mixing module is also provided with a channel attention control parameter use_se, which determines whether to use a channel attention module according to the parameter in the preset built-in parameter table.

[0067] In step S305, the first extracted feature map after preprocessing is input to the first reparameterization visual transformer block through an input adaptation layer; the first reparameterization visual transformer block performs two-step mixing processing on the spatial features and channel features of the first extracted feature map.

[0068] In step S306, the spatial feature mixing processing of the first extracted feature map is performed by a spatial feature mixing module, which includes a 3x3 normalized depth separable convolution defined by a convolution normalization module, a feature extraction channel attention module, a 1x1 normalized point convolution defined by a convolution normalization module and a reparameterized depth separable convolution.

[0069] Wherein the reparameterized depth separable convolution is composed of a 3x3 depth separable convolution defined by a normal convolution layer, and a point convolution defined by a normal 1x1 convolution layer; the first extracted feature map is processed by the reparameterized visual transformer block after the input adaptation layer is preprocessed:

[0070] In step S307, when the parameter step length stride=1, the spatial feature mixing module uses the reparameterized depth separable convolution and the feature extraction channel attention module; the local structural features of the first extracted feature map are enhanced by using the reparameterized depth separable convolution; then when the channel attention control parameter use_se is true, the feature extraction channel attention module is used to enhance the channel of the first extracted feature map; when the channel attention control parameter use_se is false, the first extracted feature map is not processed by the feature extraction channel attention module, and is directly output as the second extracted feature map; finally, the second extracted feature map after spatial feature mixing is obtained;

[0071] In step S308, when the parameter step length stride=2, the spatial feature mixing module will enable 3x3 normalized depth separable convolution, feature extraction channel attention module and 1x1 normalized point convolution for extracting feature map channel number adjustment; the first extracted feature map is down-sampled to 1 / 8 resolution by using 3x3 normalized depth separable convolution, the channel is enhanced by using the feature extraction channel attention module, and the channel dimension is restored to the channel number when the first extracted feature map is input by using 1x1 normalized point convolution;

[0072] In step S309, the channel feature mixing module is used to mix the channel features of the second extracted feature map; the channel feature mixing module is composed of an up-sampling 1x1 convolution defined by a convolution normalization module, a down-sampling 1x1 convolution defined by a convolution normalization module, and an activation function GELU nonlinear activation function;

[0073] In step S310, the up-sampling 1x1 convolution expands the channel number of the second extracted feature map to twice the channel number of the second extracted feature map output by the spatial feature mixing module, the activation function GELU retains the negative value information, the down-sampling 1x1 convolution compresses the channel number of the second extracted feature map back to the channel number when the second extracted feature map is input, and then the second extracted feature map is connected with the first extracted feature map through a residual connection, retaining the features of the first extracted feature map that have not been mixed with spatial features and channel features, to obtain a third extracted feature map;

[0074] In step S311, the third extracted feature map is input into the first stage down-sampling block to down-sample to 1 / 16 resolution to obtain a fourth extracted feature map, and the fourth extracted feature map is input into a second reparameterized visual transformer block to extract local features of the fourth extracted feature map to obtain a fifth extracted feature map;

[0075] Step S312, input the fifth extraction feature map to the second stage downsampling block to 1 / 32 resolution to obtain a sixth extraction feature map; the sixth extraction feature map is input to the third reparameterization visual transformer block to perform overall feature correlation calculation to obtain a seventh extraction feature map.

[0076] Further, in step S4, the multi-scale feature fusion network processing process is as follows:

[0077] Step S401, the multi-scale feature fusion network is composed of 3 1x1 convolution modules, 2 3x3 convolution modules, 2 dynamic upsampling modules, 3 region enhanced gradient pooling attention modules, 4 splicing modules, 1 improved squeeze and excitation attention module and a C3k2 module;

[0078] Step S402, the three 1x1 convolution modules respectively receive the third extraction feature map, the fifth extraction feature map and the seventh extraction feature map output from the three reparameterization visual transformer blocks of the feature extraction network, and the channel numbers of the third extraction feature map, the fifth extraction feature map and the seventh extraction feature map are unified to 256; wherein the third extraction feature map is processed by the first 1x1 convolution module to obtain a high-resolution first fusion feature map carrying high-level semantic information, the fifth extraction feature map is processed by the second 1x1 convolution module to obtain a medium-resolution second fusion feature map balancing details and semantics, and the seventh extraction feature map is processed by the third 1x1 convolution module to obtain a low-resolution third fusion feature map carrying deep-level semantic information;

[0079] Step S403, perform bottom-up feature fusion, process the third fusion feature map to 1 / 16 resolution by using the first dynamic upsampling module, and then splice the second fusion feature map by using the first splicing module to obtain a fourth fusion feature map, and input the fourth fusion feature map to the first region enhanced gradient pooling attention module for cross-scale feature fusion;

[0080] As shown in Figure 2 , the region enhanced gradient pooling attention module is composed of a fourth 1x1 convolution module, 2 perception patch attention modules, 1 fusion splicing module and a ReLU regularization activation function; wherein the perception patch attention module is composed of 1 spatial attention module, 2 local-global attention modules with different receptive fields, 1 feature fusion channel attention module, 1 four-branch 3x3 convolution and 1 four-branch 1x1 convolution;

[0081] Step S404, the fourth fused feature map is aligned in channel by a fourth 1x1 convolution module, and then is respectively input into a first perception patch attention module and a second perception patch attention module. The perception patch attention module is provided with a four-branch structure, and the four-branch structure includes a 2-receptive field local-global attention module, a 4-receptive field local-global attention module, a four-branch 3x3 convolution, and a four-branch 1x1 convolution. The fourth fused feature map is input into the first perception patch attention module and the second perception patch attention module for four-branch structure processing.

[0082] Step S405, the four-branch structure processing flow of the fourth fused feature map is as follows: the local attention branches of the 2-receptive field local-global attention module and the 4-receptive field local-global attention module extract local features of different scales, and are fused with global features obtained by the global attention branches, and a four-branch 3x3 convolution is used to capture local-to-global detail features to construct a multi-level local feature representation of the fourth fused feature map; and a four-branch 1x1 convolution is used to retain global feature information and local feature information to obtain a multi-source heterogeneous feature map.

[0083] Step S406, the multi-source heterogeneous feature map is input into a feature fusion channel attention module, and then is input into a spatial attention module to generate a spatial weight mask to highlight a small target region. Finally, the first perception patch attention module obtains a fifth fused feature map after processing the fourth fused feature map. Similarly, the fourth fused feature map is processed by the second perception patch attention module to obtain a sixth fused feature map. A fusion splicing module is used to splice the fifth fused feature map and the sixth fused feature map to obtain a seventh fused feature map. Finally, the seventh fused feature map is input into a ReLU regularization activation function to suppress gradient vanishing to obtain an eighth fused feature map, and the eighth fused feature map is output to a second dynamic upsampling module.

[0084] Step S407, the eighth fused feature map is input into the second dynamic upsampling module, and the eighth fused feature map is upsampled to 1 / 8 resolution. Then, the eighth fused feature map is spliced with the first fused feature map by a second splicing module to obtain a ninth fused feature map.

[0085] Step S408, the ninth fused feature map is input into a second region enhancement gradient pooling attention module for further fusion to obtain a tenth fused feature map.

[0086] Further, in step S4, the top-down feature fusion in the multi-scale feature fusion network processing process is as follows:

[0087] Step S409, the tenth fused feature map is input into an improved squeeze-and-excitation attention module to obtain an eleventh fused feature map, and the eleventh fused feature map is output to a first detection head of a detection layer.

[0088] Step S410, the eleventh fusion feature map obtained by processing the improved extrusion attention module is down-sampled by a first 3x3 convolution module, so that the resolution of the eleventh fusion feature map is reduced to 1 / 16, the channel number of the eleventh fusion feature map is increased to 256, and the eleventh fusion feature map is output. The eighth fusion feature map obtained by the first regional enhancement gradient pooling attention module is spliced through a third splicing module to obtain a twelfth fusion feature map, and the twelfth fusion feature map is input into a third regional enhancement gradient pooling attention module to obtain a thirteenth fusion feature map; the thirteenth fusion feature map is input into a second detection head in the detection layer;

[0089] Step S411, the thirteenth fusion feature map obtained is down-sampled by a second 3x3 convolution module, so that the resolution of the thirteenth fusion feature map is reduced to 1 / 32, and the channel number of the thirteenth fusion feature map is increased to 512. The thirteenth fusion feature map and the seventh extraction feature map output in step S312 are spliced by using a fourth splicing module to obtain a fourteenth fusion feature map, and the fourteenth fusion feature map is input into a C3K2 module to obtain a fifteenth fusion feature map, which is output to a third detection head in the detection layer.

[0090] Further, in step S5, during the processing of the detection layer, the detection head loss function is provided with a positioning loss, the error between the predicted bounding box after prediction and the real bounding box is calculated by using an SF-DIoU loss function, and is transmitted to the next training to optimize the bounding box coordinates.

[0091] Further, in step S3, the feature extraction network: the initial feature map of the road vehicle data set processed in step S1 is input into the input adaptation layer of the feature extraction network for preprocessing, which is expressed by the formula:

[0092] ;

[0093] In the formula, X is the initial feature map of the road vehicle data set processed in step S1; is the first extraction feature map output by the input adaptation layer after processing the input initial feature map X; is a convolution normalization module; is a nonlinear transformation activation function; the input adaptation layer processes the initial feature map X after two convolution normalization modules and an activation function GELU to obtain a feature map of 160x160;

[0094] The first extraction feature map processed by the input adaptation layer is input into the first reparameterization visual transformer block of the multi-branch structure reparameterization extraction module, which is expressed by the formula:

[0095] ;

[0096] In the formula, It is a channel feature mixing module; It is a spatial feature hybrid module; E output It is the output of the reparameterized vision transformer block; The parameter step size is given by the preset built-in parameter table;

[0097] When the stride parameter is 1, the use of the feature extraction channel attention module is determined based on the channel attention control parameter use_se obtained from the preset built-in parameter table. The formula is as follows:

[0098] ;

[0099] In the formula, It is a reparameterized depthwise separable convolution; SE() is the feature extraction channel attention module; it processes the first feature map using reparameterized depthwise separable convolution; when the channel attention control parameter use_se is true, the feature extraction channel attention module is used for channel enhancement; when the channel attention control parameter use_se is false, the second extracted feature map obtained after reparameterized depthwise convolution is directly output.

[0100] The local structural features of the first extracted feature map are enhanced by using reparameterized depthwise separable convolutions, as expressed by the formula:

[0101] );

[0102] In the formula, It is a 3x3 depthwise separable convolution defined by a normal convolutional layer, with a kernel size of 3x3; It is a point convolution defined by a normal 1x1 convolutional layer, with a kernel size of 1x1; BN() performs batch normalization on the three branches, merging the three branches;

[0103] When the parameter stride is 2, the calculation formula for the spatial feature mixing module is expressed as follows:

[0104] ;

[0105] In the formula, It is a 3x3 normalized depthwise separable convolution defined by the convolution normalization module, with a kernel size of 3x3; It is a 1x1 normalized point convolution defined by the convolution normalization module.

[0106] The second extracted feature map is mixed using a channel feature mixing module to obtain the third extracted feature map; the calculation formula is expressed as follows:

[0107] ;

[0108] ;

[0109] In the formula, is a channel feature mixing module; U is a second extracted feature map; GELU is a nonlinear activation function; is a dimensional 1x1 convolution defined by a convolution normalization module, which is used for convolution operation to expand the channel after nonlinear processing; is a dimensional 1x1 convolution defined by a convolution normalization module, which is used for second convolution operation to compress the channel processing; is a residual addition operation; is a third extracted feature map.

[0110] Further, in step S4, the multi-scale feature fusion network is used for bottom-up feature fusion, the first dynamic upsampling module is used to process the third fusion feature map to 1 / 16 resolution, and then the second fusion feature map is spliced with the first splicing module to obtain the fourth fusion feature map. The fourth fusion feature map is input into the first regional enhancement gradient pooling attention module for cross-scale feature fusion; the formula is expressed as:

[0111] ;

[0112] ;

[0113] In the formula, is the feature perception weight predicted by the third fusion feature map at the (i, j) position; g is a back propagation operation, which is used to extract local information in the third fusion feature map, , is the third fusion feature map at the (i, j) position, which is a horizontal and vertical adaptive dynamic offset; is the up-sampling rate, which is usually 2; is the second fusion feature map; is the splicing operation of the first splicing module; is the third fusion feature map is the fourth fusion feature map obtained by splicing the second fusion feature map with the first splicing module after processing by the dynamic upsampling module;

[0114] Further, the fifth fusion feature map is obtained, as shown in the formula:

[0115] ;

[0116] ;

[0117] In the formula, is a multi-source heterogeneous feature map obtained by fusing features of four branches; is the fourth fusion feature map, is a four-branch 3x3 convolution inside the perception patch attention module, used to build multi-level local feature representation; 、 is a local-global attention module with 2 receptive fields and a local-global attention module with 4 receptive fields; is the fifth fusion feature map; and SAM are feature fusion channel attention module and spatial attention module respectively, ;

[0118] Further, the top-down feature fusion in the process of the multi-scale feature fusion network, the tenth fusion feature map is input into the improved squeeze and excitation attention module to obtain the eleventh fusion feature map; as shown in the formula:

[0119] ;

[0120] ;

[0121] ;

[0122] In the formula, S is the generated channel statistics, H and W are the height and width of the key channel of the tenth fusion feature map respectively, is the pixel value of the (i, j) position in the tenth fusion feature map; is the channel information spliced along the channel dimension of the result of the four branches; is the channel dimension of the kth branch; is the kth branch channel parameter, is the eleventh fusion feature map, is the tenth fusion feature map, is the Sigmoid function.

[0123] Further, in step S5, in the process of the detection layer, the detection head loss function is provided with a positioning loss, the error between the predicted bounding box after prediction and the real bounding box is calculated through the SF-DIoU loss function, and is transmitted to the next training to optimize the bounding box coordinates; Specifically:

[0124] In step S51, the SF-DIoU loss function first uses shape-IoU dynamic sample weighting and shape-aware penalty to adjust the loss contribution of different difficulty samples; wherein shape-IoU measures the overlapping area of the predicted bounding box and the real bounding box based on IoU, and the formula is as follows:

[0125] ;

[0126] In the formula: is the shape-weighted center point distance penalty, and respectively are the center point coordinate difference of the predicted bounding box coordinates x, y of the vehicle in the initial feature map detected by the detection layer; 、 respectively are the width and height of the detection box controlled by the sensitivity scale scale and penalized by the detection box size; c is the minimum enclosing and diagonal length of the detection box;

[0127] Step S52, the shape-IoU loss function introduces a status difference penalty term, the formula is as follows:

[0128] ;

[0129] In the formula: is the status difference penalty term, e is the exponential function, is the width difference term, is the height difference term;

[0130] The width difference term and the height difference term make the improved network model make fine adjustment of the width-height ratio in the next training by putting the size difference through the exponential function and the fourth power; the difference term formula is as follows:

[0131] ;

[0132] In the formula: h, w are respectively the height and width of the intersection part of the real bounding box and the predicted bounding box of the vehicle position in the initial feature map after prediction; 、 respectively are the width of the real bounding box of the actual vehicle position in the initial feature map and the width of the predicted bounding box of the vehicle position in the initial feature map after prediction by the detection layer; respectively are the width of the real bounding box of the actual vehicle position in the initial feature map and the width of the predicted bounding box of the vehicle position in the initial feature map after prediction by the detection layer;

[0133] The shape-IoU loss function, the formula is as follows:

[0134] ;

[0135] In the formula, is the status difference penalty term; is the threshold value for measuring the overlapping area of the predicted bounding box and the real bounding box;

[0136] Step S53, the SF_DIoU loss function introduces a dynamic stage and scales the shape_IoU value to optimize the sample weight distribution; the final loss function SF_DIoU is obtained; the formula is as follows:

[0137] ;

[0138] In the formula, is the final loss function; u is the prediction accuracy of the vehicle in the initial feature map predicted by the detection layer; d is the prediction difficulty index of the vehicle in the initial feature map predicted by the detection layer;

[0139] Step S54, truncate and linearly scale the IoU value, adjust the prediction difficulty loss contribution of the vehicle in the initial feature map predicted by the different detection layers;

[0140] When the prediction difficulty index d of the vehicle in the initial feature map predicted by the detection layer is 0, ignore the extremely low sample;

[0141] When the prediction accuracy u of the vehicle in the initial feature map predicted by the detection layer is 0.95, stop optimizing the high index sample, that is, the target with an identification accuracy higher than 0.95 to prevent overfitting;

[0142] For complex samples, that is, , the loss contribution of the target identified by the detection layer is enhanced, forcing the next training to make fine adjustments;

[0143] For simple samples, that is, , the weight of the target identified by the detection layer is reduced to avoid overfitting;

[0144] For noise samples, that is, , the target identified by the detection layer is completely ignored to improve robustness.

Claims

1. A multi-scale road vehicle detection method based on reparameterized visual transformer, characterized in that: The method comprises the following steps: Step S1, data set cleaning and construction; a KITTI 2D road vehicle data set is selected, the KITTI 2D road vehicle data set is cleaned, screened and detected, and a processed road vehicle data set is obtained, and the road vehicle data set is used as an improved network model input data set; Step S2, an improved network model; the improved network model adopts a series-parallel structure and is divided into a feature extraction network, a multi-scale feature fusion network and a detection layer; the feature extraction network is composed of an input adaptive layer and a multi-branch structure reparameterization extraction module; the multi-scale feature fusion network comprises a 1x1 convolution module, a 3x3 convolution module, a dynamic up-sampling module, a region enhanced gradient pooling attention module, a splicing module, an improved squeeze and excitation attention module and a C3k2 module; Step S3, a feature extraction network processing process; The initial feature map of the road vehicle data set processed in step S1 is input into the input adaptive layer in the feature extraction network for preprocessing to obtain a first extracted feature map, and the first extracted feature map is input into the multi-branch structure reparameterization extraction module in the feature extraction network for preprocessing to obtain a third extracted feature map, a fifth extracted feature map and a seventh extracted feature map; Step S4, a multi-scale feature fusion network processing process; A bottom-up and top-down bidirectional path feature fusion is adopted; The third extracted feature map, the fifth extracted feature map and the seventh extracted feature map are input into the multi-scale feature fusion network to obtain an eleventh fused feature map, a thirteenth fused feature map and a fifteenth fused feature map; Step S5, a detection layer processing process; the eleventh fused feature map, the thirteenth fused feature map and the fifteenth fused feature map are respectively input into a detection layer composed of three detection heads; the detection layer processes the eleventh fused feature map, the thirteenth fused feature map and the fifteenth fused feature map respectively, and completes the prediction of different size targets that need to be recognized in the initial feature map.

2. The method of claim 1, wherein: In step S3, the initial feature map of the road vehicle data set processed in step S1 is input into the input adaptive layer in the feature extraction network for preprocessing to obtain a first extracted feature map; specifically: Step S301, the input adaptive layer comprises two convolution normalization modules, which are divided into a first convolution normalization module and a second convolution normalization module; wherein the convolution normalization module is composed of a convolution layer and a batch normalization layer, and is used to define a plurality of convolution modules in the improved network model to adjust the resolution and channel number of the input feature map; Step S302, the first convolution normalization module reduces the channel number of the initial feature map to half of the channel number of the initial feature map, and down-samples to a resolution of 320x320, and performs nonlinear transformation by using an activation function GELU; Step S303, the second convolution normalization module restores the channel number of the feature map output from the first convolution normalization module to the channel number of the initial feature map, and down-samples to a resolution of 160x160; a first extracted feature map after preprocessing is obtained and input into the multi-branch structure reparameterization extraction module.

3. The method of claim 2, wherein: In step S3, the first extracted feature map is input to the multi-branch structure reparameterization extraction module in the feature extraction network for preprocessing to obtain a third extracted feature map, a fifth extracted feature map and a seventh extracted feature map; specifically: In step S304, the multi-branch structure reparameterization extraction module is improved by a reparameterization visual transformer, which is composed of three reparameterization visual transformer blocks, two stage down-sampling blocks and a preset built-in parameter table; The reparameterization visual transformer block is composed of a spatial feature mixing module and a channel feature mixing module; The preset built-in parameter table is preset by the user, and when the spatial feature mixing is performed, the spatial feature mixing module includes a reparameterization depth separable convolution and a depth separable convolution, and the spatial feature mixing module is provided with a parameter step stride, which selects different depth separable convolutions according to the parameter step stride given in the preset built-in parameter table; The spatial feature mixing module is also provided with a channel attention control parameter use_se, which determines whether to use the channel attention module according to the parameter in the preset built-in parameter table; In step S305, the first extracted feature map preprocessed by the input adaptation layer is input to the first reparameterization visual transformer block; the first reparameterization visual transformer block performs two-step mixing processing on the spatial features and channel features of the first extracted feature map; In step S306, the spatial feature mixing processing of the first extracted feature map is performed by the spatial feature mixing module, which includes a 3x3 normalized depth separable convolution defined by a convolution normalization module, a feature extraction channel attention module, a 1x1 normalized point convolution defined by a convolution normalization module and a reparameterization depth separable convolution; The reparameterization depth separable convolution is composed of a 3x3 depth separable convolution defined by a normal convolution layer and a point convolution defined by a normal 1x1 convolution layer; the input adaptation layer preprocessed first extracted feature map is processed by the reparameterization visual transformer block; In step S307, when the parameter step stride is 1, the spatial feature mixing module uses the reparameterization depth separable convolution and the feature extraction channel attention module; the reparameterization depth separable convolution is used to enhance the local structure features of the first extracted feature map; then when the channel attention control parameter use_se is true, the feature extraction channel attention module is used to enhance the channel of the first extracted feature map; when the channel attention control parameter use_se is false, the first extracted feature map is not processed by the feature extraction channel attention module and is directly output as the second extracted feature map; finally, the second extracted feature map after spatial feature mixing is obtained. Step S308, when the parameter step length stride=2, the spatial feature mixing module will enable 3x3 normalized depth separable convolution, feature extraction channel attention module and 1x1 normalized point convolution, which are used to extract feature map channel number adjustment; the 3x3 normalized depth separable convolution is used to downsample the first extracted feature map to 1 / 8 resolution, the feature extraction channel attention module is used for channel enhancement, and the 1x1 normalized point convolution is used to restore the channel dimension to the channel number when the first extracted feature map is input; Step S309, the channel feature mixing module is used for channel feature mixing of the second extracted feature map; The channel feature mixing module is composed of an up-sampling 1x1 convolution defined by a convolution normalization module, a down-sampling 1x1 convolution defined by a convolution normalization module and an activation function GELU nonlinear activation function; Step S310, the up-sampling 1x1 convolution expands the channel number of the second extracted feature map to twice the channel number of the second extracted feature map output by the spatial feature mixing module, the activation function GELU retains negative value information, the down-sampling 1x1 convolution compresses the channel number of the second extracted feature map back to the channel number when the second extracted feature map is input, and then the second extracted feature map is connected with the first extracted feature map through residual connection, retaining the features of the first extracted feature map that have not undergone spatial feature and channel feature mixing, to obtain a third extracted feature map; Step S311, the third extracted feature map is input into the first stage down-sampling block to down-sample to 1 / 16 resolution to obtain a fourth extracted feature map, and the fourth extracted feature map is input into the second re-parameterized visual transformer block to extract local features of the fourth extracted feature map to obtain a fifth extracted feature map; Step S312, the fifth extracted feature map is input into the second stage down-sampling block to 1 / 32 resolution to obtain a sixth extracted feature map; the sixth extracted feature map is input into the third re-parameterized visual transformer block to calculate the overall feature correlation to obtain a seventh extracted feature map.

4. The multi-scale road vehicle detection method based on reparameterization visual transformer according to claim 3, characterized in that: In step S4, the multi-scale feature fusion network processing process is as follows: Step S401, the multi-scale feature fusion network is composed of 3 1x1 convolution modules, 2 3x3 convolution modules, 2 dynamic up-sampling modules, 3 regional enhancement gradient pooling attention modules, 4 splicing modules, 1 improved squeeze and excitation attention module and 1 C3k2 module; Step S402, the three 1x1 convolution modules respectively receive the third extracted feature map, the fifth extracted feature map and the seventh extracted feature map output by the three re-parameterized visual transformer blocks of the feature extraction network, and the channel numbers of the third extracted feature map, the fifth extracted feature map and the seventh extracted feature map are unified to 256; wherein the third extracted feature map is processed by the first 1x1 convolution module to obtain a high-resolution first fusion feature map carrying high-level semantic information, the fifth extracted feature map is processed by the second 1x1 convolution module to obtain a medium-resolution second fusion feature map balancing details and semantics, and the seventh extracted feature map is processed by the third 1x1 convolution module to obtain a low-resolution third fusion feature map carrying deep-level semantic information; Step S403, bottom-up feature fusion is performed, the third fused feature map is processed to 1 / 16 resolution by using a first dynamic upsampling module, and then the fourth fused feature map is obtained by splicing the second fused feature map and the third fused feature map by using a first splicing module; the fourth fused feature map is input into a first regional enhancement gradient pooling attention module for cross-scale feature fusion; The regional enhancement gradient pooling attention module is composed of a fourth 1x1 convolution module, two perception patch attention modules, a fusion splicing module and a ReLU regularization activation function; wherein the perception patch attention module is composed of a spatial attention module, two local-global attention modules with different receptive fields, a feature fusion channel attention module, a four-branch 3x3 convolution and a four-branch 1x1 convolution; Step S404, the fourth fused feature map is aligned in channel by the fourth 1x1 convolution module, and then is input into the first perception patch attention module and the second perception patch attention module respectively; the perception patch attention module is provided with a four-branch structure, the four-branch structure includes a 2-receptive field local-global attention module, a 4-receptive field local-global attention module, a four-branch 3x3 convolution and a four-branch 1x1 convolution; the fourth fused feature map is input into the first perception patch attention module and the second perception patch attention module respectively for four-branch structure processing; Step S405, the four-branch structure processing flow of the fourth fused feature map is as follows: the local attention branch of the 2-receptive field local-global attention module and the 4-receptive field local-global attention module extracts local features of different scales, and the global features obtained by the global attention branch are fused, and the fourth fused feature map is constructed by using the four-branch 3x3 convolution to capture local-to-global detailed features to represent multi-level local features; the four-branch 1x1 convolution is used to retain global feature information and local feature information to obtain a multi-source heterogeneous feature map; Step S406, the multi-source heterogeneous feature map is input into a feature fusion channel attention module, and then is input into a spatial attention module to generate a spatial weight mask to highlight a small target region; finally, the fourth fused feature map is processed by the first perception patch attention module to obtain a fifth fused feature map; Similarly, the fourth fused feature map is processed by the second perception patch attention module to obtain a sixth fused feature map; The fifth fused feature map and the sixth fused feature map are spliced by using a fusion splicing module to obtain a seventh fused feature map; finally, the seventh fused feature map is input into a ReLU regularization activation function to suppress gradient vanishing to obtain an eighth fused feature map, and the eighth fused feature map is output to a second dynamic upsampling module; Step S407, the eighth fused feature map is input into the second dynamic upsampling module, the eighth fused feature map is upsampled to 1 / 8 resolution, and then the first fused feature map is spliced by using a second splicing module to obtain a ninth fused feature map; Step S408, the ninth fused feature map is input into a second regional enhancement gradient pooling attention module for further fusion to obtain a tenth fused feature map.

5. The method of claim 4, wherein: In step S4, the multi-scale feature fusion network processes the top-down feature fusion in the process, specifically: Step S409, the tenth fusion feature map is input into the improved squeeze and excitation attention module to obtain the eleventh fusion feature map, and output to the first detection head of the detection layer; Step S410, the eleventh fusion feature map obtained by processing the improved squeeze and excitation attention module is down-sampled through the first 3x3 convolution module, so that the resolution of the eleventh fusion feature map is reduced to 1 / 16, the channel number of the eleventh fusion feature map is increased to 256, and the eleventh fusion feature map is output, the eleventh fusion feature map is spliced with the eighth fusion feature map obtained by the first regional enhancement gradient pooling attention module through the third splicing module to obtain the twelfth fusion feature map, the twelfth fusion feature map is input into the third regional enhancement gradient pooling attention module to obtain the thirteenth fusion feature map; the thirteenth fusion feature map is input into the second detection head in the detection layer; Step S411, the thirteenth fusion feature map is down-sampled through the second 3x3 convolution module, so that the resolution of the thirteenth fusion feature map is reduced to 1 / 32, and the channel number of the thirteenth fusion feature map is increased to 512 and output; The thirteenth fusion feature map and the seventh extraction feature map output in step S312 are spliced by using the fourth splicing module to obtain the fourteenth fusion feature map, and the fourteenth fusion feature map is input into the C3k2 module to obtain the fifteenth fusion feature map, which is output to the third detection head in the detection layer.

6. The multi-scale road vehicle detection method based on reparameterization visual transformer according to claim 5, wherein: In step S5, in the process of the detection layer, the detection head loss function is provided with a positioning loss, the error between the predicted bounding box after prediction and the real bounding box is calculated by using the SF-DIoU loss function, and is transmitted to the next training to optimize the bounding box coordinates.

Citation Information

Patent Citations

  • Target detection method and device

    CN110096960A

  • Pavement disease rapid detection method and system based on YOLOV7 algorithm

    CN117058459A