Traffic target detection model training method and device based on deep learning and double-flow feature fusion enhancement

By adopting a traffic target detection model based on deep learning and dual-stream feature fusion enhancement in the field of intelligent transportation, the problems of large changes in traffic target scales and overlapping occlusion are solved, and higher detection accuracy is achieved.

CN120125802APending Publication Date: 2025-06-10INST OF AUTOMATION CHINESE ACAD OF SCI +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510183413.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Traffic target detection has problems of large scale changes and overlapping occlusion in the field of intelligent transportation, resulting in low detection accuracy.

Method used

The traffic object detection model based on deep learning and dual-stream feature fusion enhancement is adopted, and multi-scale feature extraction, enhancement and fusion are carried out through dual-stream feature extraction, backbone network, feature enhancement network based on dynamic packet convolution, and multi-scale feature fusion network to improve detection accuracy.

Benefits of technology

It effectively solves the problems of large changes in traffic target scales and overlapping occlusion, and significantly improves the accuracy of traffic target detection by traffic system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125802A_ABST
    Figure CN120125802A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic target detection model training method and device based on deep learning and double-flow feature fusion enhancement. The training method comprises the following steps: acquiring traffic target image data; performing preprocessing including labeling and standardization on the traffic target image data to obtain an input data set including a training data set and a verification data set; inputting the input data set into a traffic target detection model based on double-flow feature fusion enhancement, and performing multi-scale feature extraction processing, multi-scale feature enhancement processing and multi-scale feature fusion processing on the input data set to obtain an output data set including traffic target type and position information; and according to the verification data set and the output data set, adjusting each parameter of the traffic target detection model based on double-flow feature fusion enhancement, and obtaining a trained traffic target detection model. According to the scheme provided by the invention, the detection precision of the traffic target can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of computer vision technology, and more specifically, to a training method and device for a traffic target detection model based on deep learning and dual-stream feature fusion enhancement. Background Art

[0002] With the acceleration of the urbanization process, problems such as traffic congestion, traffic accidents, and traffic environmental pollution have become increasingly serious, and the research and application of intelligent transportation systems have become particularly important. In related fields, target detection technology is widely used in intelligent transportation, including but not limited to traffic flow analysis, intelligent monitoring, anomaly detection, etc. As one of the key components in the intelligent transportation system, traffic target detection technology involves automatically identifying objects belonging to target categories in a large amount of data, and calibrating the positions and regions of the identified objects. Traditional multi-target detection algorithms are mainly based on manually designed feature extraction and machine learning methods.

[0003] In recent years, with the development of deep learning technology, multi-target detection algorithms based on deep learning have received extensive attention. Deep learning models can extract high-level deep features, better capture target features, and achieve efficient multi-motion target detection. However, there are still some challenges in the application of target detection in the field of intelligent transportation, such as complex environment adaptability, real-time performance, multi-target detection, data imbalance, algorithm robustness, computational resource limitations, and security and privacy protection. Summary of the Invention

[0004] Embodiments of the present disclosure provide a training method and device for a traffic target detection model based on deep learning and dual-stream feature fusion enhancement, and a traffic target detection method. By training a traffic target detection model based on deep learning and dual-stream feature fusion enhancement, problems such as large-scale changes and overlapping occlusions of traffic targets can be solved, and the detection accuracy of traffic targets by, for example, a traffic system can be effectively improved.

[0005] In one general aspect, a training method for a traffic target detection model enhanced by deep learning and dual-stream feature fusion is provided. The training method includes: obtaining traffic target image data; performing preprocessing including annotation and normalization on the traffic target image data to obtain an input data set including a training data set and a validation data set; inputting the input data set into a traffic target detection model enhanced by dual-stream feature fusion, and obtaining an output data set including traffic target types and location information by performing multi-scale feature extraction processing, multi-scale feature enhancement processing, and multi-scale feature fusion processing on the input data set; adjusting each parameter of the traffic target detection model enhanced by dual-stream feature fusion according to the validation data set and the output data set to obtain a trained traffic target detection model, where the traffic target detection model enhanced by dual-stream feature fusion includes a dual-stream feature extraction backbone network for parallel data processing, a feature enhancement network based on dynamic group convolution, and a multi-scale feature fusion network connected in sequence. The dual-stream feature extraction backbone network is used to perform the multi-scale feature extraction processing, the feature enhancement network based on dynamic group convolution is used to perform the multi-scale feature enhancement processing, and the multi-scale feature fusion network is used to perform the multi-scale feature fusion processing.

[0006] Optionally, the multi-scale feature extraction processing may include the following processing: dividing the training data set into two branch data sets and inputting the two branch data sets into two backbone networks and a dual multi-scale attention network included in the dual-stream feature extraction backbone network respectively to perform the following processing in parallel: performed by two backbone networks: inputting the two branch data sets into a first convolutional block including 2D convolution, normalization, and a preset activation function to obtain corresponding initial feature maps, inputting the initial feature maps into a second convolutional block for performing feature extraction downsampling to obtain two sets of updated first feature maps, and performed by the dual multi-scale attention network: combining the local features corresponding to the training data set with the corresponding global dependencies to obtain a combined second feature map; fusing the features of each level of the two backbone networks corresponding to the first feature map and the second feature map in the channel dimension to obtain a fused third feature map; compressing the third feature map to obtain a compressed fourth feature map as the multi-scale feature extraction result.

[0007] Optionally, the multi-scale feature enhancement process may include the following processes: dividing the multi-scale feature extraction result obtained through the multi-scale feature extraction process into multiple input features in the channel dimension; performing dynamic group convolution and channel shuffle processing on a part of the multiple input features to obtain a first feature; fusing another part of the multiple input features that have not been processed with the first feature in the channel dimension to obtain a fused second feature; inputting the second feature into a preset feed-forward network included in the feature enhancement network based on dynamic group convolution for enhancing the context understanding ability to obtain a processed third feature as the multi-scale feature enhancement result.

[0008] Optionally, the multi-scale feature fusion process may include the following processes: performing global feature information extraction processing and feature encoding processing on the multi-scale feature enhancement result obtained through the multi-scale feature enhancement process to obtain a processed first multi-scale feature map; connecting and splitting the first multi-scale feature map in the channel dimension to obtain two parts of multi-scale feature maps; inputting one part of the two parts of multi-scale feature maps into a reparameterization network in the multi-scale feature fusion network to perform convolution and several reparameterization processes to obtain a second multi-scale feature map; inputting the other part of the two parts of multi-scale feature maps into a preset cross-layer aggregation network in the multi-scale feature fusion network to perform cross-layer aggregation processing to obtain a third multi-scale feature map; connecting the third multi-scale feature map with the second multi-scale feature map and performing convolution processing on the connected multi-scale feature map to obtain a fused multi-scale feature map as the multi-scale feature fusion result.

[0009] Optionally, the following formula may be used to obtain the fused third feature map:

[0010]

[0011] where y out represents the third feature map, x initial represents the training data set, F k* represents a feature extraction process with k levels of processing and serves as the main backbone network, represents the i-th auxiliary backbone network with k levels, where k, l, and i are natural numbers, and the value of i is greater than or equal to 1 and less than or equal to l.

[0012] Optionally, the traffic target detection model based on dual-stream feature fusion enhancement may further include a target prediction network for performing prediction processing on the multi-scale feature fusion result obtained through the multi-scale feature fusion process to obtain the output data set.

[0013] In another general aspect, there is provided a training device for a traffic target detection model enhanced by deep learning and dual-stream feature fusion. The training device includes: a data acquisition module configured to acquire traffic target image data; a data preprocessing module configured to perform preprocessing including annotation and normalization on the traffic target image data to obtain an input data set including a training data set and a validation data set; a feature training module configured to input the input data set into a traffic target detection model enhanced by dual-stream feature fusion, and obtain an output data set including traffic target types and location information by performing multi-scale feature extraction processing, multi-scale feature enhancement processing, and multi-scale feature fusion processing on the input data set; a parameter adjustment module configured to adjust each parameter of the traffic target detection model enhanced by dual-stream feature fusion according to the validation data set and the output data set to obtain a trained traffic target detection model. Among them, the traffic target detection model enhanced by dual-stream feature fusion includes a dual-stream feature extraction backbone network for parallel data processing, a feature enhancement network based on dynamic group convolution, and a multi-scale feature fusion network connected in sequence. The dual-stream feature extraction backbone network is used to perform the multi-scale feature extraction processing, the feature enhancement network based on dynamic group convolution is used to perform the multi-scale feature enhancement processing, and the multi-scale feature fusion network is used to perform the multi-scale feature fusion processing.

[0014] Optionally, the multi-scale feature extraction processing performed by the feature training module may include the following processing: dividing the training data set into two branch data sets and respectively inputting the two branch data sets into two backbone networks and a dual multi-scale attention network included in the dual-stream feature extraction backbone network to perform the following processing in parallel: performed by the two backbone networks: inputting the two branch data sets into a first convolution block including 2D convolution, normalization, and a preset activation function to obtain corresponding initial feature maps, inputting the initial feature maps into a second convolution block for performing feature extraction downsampling to obtain two sets of updated first feature maps, and performed by the dual multi-scale attention network: combining the local features corresponding to the training data set with the corresponding global dependencies to obtain combined second feature maps; fusing the features of each level of the two backbone networks corresponding to the first feature maps and the second feature maps in the channel dimension to obtain a fused third feature map; compressing the third feature map to obtain a compressed fourth feature map as the multi-scale feature extraction result.

[0015] Optionally, the multi-scale feature enhancement process performed by the feature training module may include the following processes: dividing the multi-scale feature extraction result obtained through the multi-scale feature extraction process into multiple input features in the channel dimension; performing dynamic group convolution and channel shuffle processing on a part of the multiple input features to obtain a first feature; fusing another part of the multiple input features that have not been processed with the first feature in the channel dimension to obtain a second fused feature; inputting the second feature into a preset feed-forward network included in the feature enhancement network based on dynamic group convolution for enhancing the context understanding ability to obtain a processed third feature as the multi-scale feature enhancement result.

[0016] Optionally, the multi-scale feature fusion process performed by the feature training module may include the following processes: performing global feature information extraction processing and feature encoding processing on the multi-scale feature enhancement result obtained through the multi-scale feature enhancement process to obtain a processed first multi-scale feature map; connecting and splitting the first multi-scale feature map in the channel dimension to obtain two parts of multi-scale feature maps; inputting one part of the two parts of multi-scale feature maps into a reparameterization network in the multi-scale feature fusion network to perform convolution and several reparameterization processes to obtain a second multi-scale feature map; inputting the other part of the two parts of multi-scale feature maps into a preset cross-layer aggregation network in the multi-scale feature fusion network to perform cross-layer aggregation processing to obtain a third multi-scale feature map; connecting the third multi-scale feature map with the second multi-scale feature map and performing convolution processing on the connected multi-scale feature map to obtain a fused multi-scale feature map as the multi-scale feature fusion result.

[0017] Optionally, the following formula may be used to obtain the fused third feature map:

[0018]

[0019] where y out represents the third feature map, x initial represents the training dataset, F k* represents a feature extraction process with k levels of processing and serves as the main backbone network, represents the i-th auxiliary backbone network with k levels, where k, l, and i are natural numbers, and the value of i is greater than or equal to 1 and less than or equal to l.

[0020] Optionally, the traffic target detection model based on dual-stream feature fusion enhancement may further include a target prediction network for performing prediction processing on the multi-scale feature fusion result obtained through the multi-scale feature fusion process to obtain the output dataset.

[0021] In another general aspect, a traffic target detection method is provided. The traffic target detection method includes: obtaining traffic target image data to be detected; inputting the traffic target image data to be detected into a traffic target detection model based on dual-stream feature fusion enhancement to obtain corresponding traffic target detection data, where the traffic target detection model based on dual-stream feature fusion enhancement is trained by using the training method described above.

[0022] In another general aspect, a traffic target detection device is provided. The traffic target detection device includes: an obtaining module configured to obtain traffic target image data to be detected; a detection module configured to input the traffic target image data to be detected into a traffic target detection model based on dual-stream feature fusion enhancement to obtain corresponding traffic target detection data, where the traffic target detection model based on dual-stream feature fusion enhancement is trained by using the training method described above.

[0023] In another general aspect, a computer program product is provided. The computer program product includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the training method and the traffic target detection method of the traffic target detection model based on deep learning and dual-stream feature fusion enhancement described above are implemented.

[0024] In another general aspect, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device / server, the electronic device / server is enabled to execute the training method and the traffic target detection method of the traffic target detection model based on deep learning and dual-stream feature fusion enhancement described above.

[0025] In another general aspect, a computing device is provided. The computing device includes: at least one processor; at least one memory storing computer-executable instructions, where when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the training method and the traffic target detection method of the traffic target detection model based on deep learning and dual-stream feature fusion enhancement described above.

[0026] According to the training method and device of the traffic target detection model based on deep learning and dual-stream feature fusion enhancement and the traffic target detection method of the embodiments of the present disclosure, by training the traffic target detection model based on deep learning and dual-stream feature fusion enhancement, the problems of large-scale variation and overlapping occlusion of traffic targets can be solved, and the detection accuracy of traffic targets by, for example, a traffic system can be effectively improved. Description of the Drawings

[0027] The above and other objects and features of the embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings showing embodiments, wherein:

[0028] Figure 1 is a flowchart showing a training method of a traffic target detection model enhanced by deep learning and dual-stream feature fusion according to an embodiment of the present disclosure;

[0029] Figure 2 is a flowchart showing an example of a training method of a traffic target detection model enhanced by deep learning and dual-stream feature fusion according to an embodiment of the present disclosure;

[0030] Figure 3 is a schematic diagram showing a dual-stream fusion network structure according to an embodiment of the present disclosure;

[0031] Figure 4 is a schematic diagram showing a Shuffle Transformer module structure according to an embodiment of the present disclosure;

[0032] Figure 5 is a schematic diagram showing an Efficient RepGFPN module structure according to an embodiment of the present disclosure;

[0033] Figure 6 is a flowchart showing a traffic target detection method according to an embodiment of the present disclosure;

[0034] Figure 7 is a block diagram showing a training device of a traffic target detection model enhanced by deep learning and dual-stream feature fusion according to an embodiment of the present disclosure;

[0035] Figure 8 is a block diagram showing a traffic target detection device according to an embodiment of the present disclosure;

[0036] Figure 9 is a block diagram showing a computing device according to an embodiment of the present disclosure. Detailed Description of the Embodiments

[0037] The following detailed description is provided to assist the reader in obtaining a comprehensive understanding of the methods, apparatuses, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, apparatuses, and / or systems described herein will be apparent after understanding the disclosure of the present application. For example, the order of operations described herein is merely exemplary and is not limited to those set forth herein, but may be changed as will be apparent after understanding the disclosure of the present application, except for operations that must occur in a specific order. In addition, descriptions of features known in the art may be omitted for greater clarity and conciseness.

[0038] Reference will now be made in detail to the embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings, wherein like reference numerals always refer to like elements. The following will explain the embodiments by referring to the accompanying drawings so as to explain the present disclosure.

[0039] As described above, in the related art, in the intelligent transportation scenario, there are still some challenges in the application of target detection in the field of intelligent transportation. For example, as the distance between the vehicle and the camera varies, the scale of the vehicle will also change accordingly. Improving scale adaptability is one of the key points for the effective application of the target tracking algorithm in the intelligent transportation scenario. In addition, the crowded traffic flow will cause vehicles to block each other, resulting in the detection algorithm losing some information, which brings great difficulties to target detection.

[0040] To solve the above and other problems, the present disclosure proposes a traffic target detection method and system based on deep learning and dual-stream feature fusion enhancement. Specifically, a training method and device for a traffic target detection model based on deep learning and dual-stream feature fusion enhancement and a traffic target detection method and device are proposed. Through the technical solution proposed by the present disclosure, the problems of large-scale changes and overlapping occlusions of traffic targets can be overcome, and the detection accuracy of traffic targets by the traffic system can be effectively improved.

[0041] Next, with reference to Figures 1 to 9 the training method and device for a traffic target detection model based on deep learning and dual-stream feature fusion enhancement and the traffic target detection method and device according to the embodiments of the present disclosure will be described in detail.

[0042] First, with reference to Figures 1 to 5 the training method for a traffic target detection model based on deep learning and dual-stream feature fusion enhancement according to the embodiments of the present disclosure will be described.

[0043] Figure 1 FIG. is a flowchart showing a training method 100 for a traffic target detection model based on deep learning and dual-stream feature fusion enhancement according to the embodiments of the present disclosure. Figure 2 FIG. is a flowchart showing an example of a training method for a traffic target detection model based on deep learning and dual-stream feature fusion enhancement according to the embodiments of the present disclosure. Figure 3 FIG. is a schematic diagram showing the structure of a dual-stream fusion network according to the embodiments of the present disclosure.

[0044] Figure 4 FIG. is a schematic diagram showing the structure of a Shuffle Transformer module (a new model based on the vision Transformer architecture) according to the embodiments of the present disclosure. Figure 5 FIG. is a schematic diagram showing the structure of an Efficient RepGFPN (Efficient Reparameterized Generalized Feature Pyramid Network) module according to the embodiments of the present disclosure.

[0045] As an example, the traffic target detection model based on dual-stream feature fusion enhancement includes a dual-stream feature extraction backbone network for parallel data processing, a feature enhancement network based on dynamic group convolution, and a multi-scale feature fusion network connected in sequence. Here, the dual-stream feature extraction backbone network is used to perform multi-scale feature extraction processing, the feature enhancement network based on dynamic group convolution is used to perform multi-scale feature enhancement processing, and the multi-scale feature fusion network is used to perform multi-scale feature fusion processing.

[0046] Referring to Figure 1 , according to an embodiment of the present disclosure, in step S101, traffic target image data is acquired.

[0047] For example, as shown in step S21 of Figure 2 , traffic target image data can be collected in the following manner: use road sensors and vehicle-mounted sensors to collect traffic target images on urban roads and highways. For example, the specific targets are various traffic-related targets / objects such as pedestrians, bicycles, motorcycles, cars, buses, trucks, and vans. In addition, the images of traffic targets can be images from multiple angles.

[0048] In addition, for example, the specific acquisition method of the road sensor is: through the acquisition device installed on the roadside, obtain the images of various traffic targets passing on the road within a predetermined time. For example, the vehicle-mounted sensor can be mounted on a vehicle (such as in front of / behind the vehicle), and it obtains the images of surrounding traffic targets as the vehicle moves.

[0049] According to an embodiment of the present disclosure, in step S102, preprocessing including annotation and normalization is performed on the traffic target image data to obtain an input data set including a training data set and a validation data set.

[0050] For example, as shown in step S21 of Figure 2 , a data set can be made by performing annotation and data normalization on the collected traffic images as the above input data set. For example, the data normalization operation can include: unifying the input image size, normalizing the image pixels, and arranging them in the order of CHW to obtain the processed input image.

[0051] According to an embodiment of the present disclosure, in step S103, the input data set is input into the traffic target detection model based on dual-stream feature fusion enhancement, and by performing multi-scale feature extraction processing, multi-scale feature enhancement processing, and multi-scale feature fusion processing on the input data set, an output data set including traffic target types and location information is obtained.

[0052] As an example, the multi-scale feature extraction process may include the following processes: dividing the training data set into two branch data sets and inputting the two branch data sets into two backbone networks and a dual multi-scale attention network included in the dual-stream feature extraction backbone network respectively to perform the following processes in parallel: performed by the two backbone networks: inputting the two branch data sets into a first convolutional block including 2D convolution, normalization, and a preset activation function to obtain corresponding initial feature maps, inputting the initial feature maps into a second convolutional block for performing feature extraction downsampling to obtain two sets of updated first feature maps, and performed by the dual multi-scale attention network: combining the local features corresponding to the training data set with the corresponding global dependencies to obtain combined second feature maps; fusing the features of each level of the two backbone networks corresponding to the first feature maps and the second feature maps in the channel dimension to obtain a fused third feature map; compressing the third feature map to obtain a compressed fourth feature map as the multi-scale feature extraction result.

[0053] Through the feature fusion of the above dual backbone networks, richer feature expressions can be obtained, and through the compression processing of the third feature map, the subsequent calculation amount can be reduced, thereby improving the algorithm calculation efficiency and realizing the high efficiency of the model.

[0054] For example, the fused third feature map can be obtained using the following formula:

[0055]

[0056] where y out represents the third feature map, x initial represents the training data set, F k* represents the feature extraction process with k levels of processing and serves as the main backbone network, represents the i-th auxiliary backbone network with k levels, k, l, and i are natural numbers, and the value of i is greater than or equal to 1 and less than or equal to l.

[0057] For example, the above-mentioned auxiliary backbone networks are implemented in the same or different ways to provide richer feature representations for the model. As shown by the "auxiliary network" in Figure 3 , the auxiliary backbone network can be a network for stagewise enhancing features. In addition, the "auxiliary network" can also be connected to the dual multi-scale attention network to provide it with feature inputs.

[0058] In the dual-stream fusion network according to the present disclosure, multiple identical backbones are assembled through composite connections between adjacent backbones to form a more powerful backbone, which can be referred to as a composite backbone network, for example. Specifically, by gradually feeding the output features of the previous backbone network as part of the input features into the subsequent backbone network, the feature maps of the last backbone network are used for subsequent detection, enhancing the overall feature richness of the model and assisting in enhancing the key feature information in the object detection stream.

[0059] Through the above dual-stream fusion network, two sets of features with different expression forms are extracted through differential dual backbones, enhancing the overall feature richness of the model, thereby assisting in enhancing the key feature information in the object detection stream.

[0060] In addition, by feeding local features into a convolutional layer, two new feature maps B and C are generated, matrix multiplication is performed between C and the transpose of B, and a softmax (an activation function) layer is applied to calculate the fusion result.

[0061] Referring to Figure 2 and Figure 3 In step S22, a dual-stream fusion network based on a dual multi-scale attention network is constructed to extract multi-scale features of traffic images. Specifically, the data set is fed into the model for training, and the image features are extracted through the constructed dual-stream fusion network to obtain two sets of features. Here, each of the two sets of obtained features contains multiple feature maps of different sizes.

[0062] As an example, step S22 may further include the following steps S221 to S226:

[0063] In step S221, the input image is divided into two branches, and the data of these two branches is processed in parallel through two backbone networks.

[0064] In step S222, an initial feature map (including multiple sets of independent feature maps) is obtained through a convolutional block (which can be referred to as a Conv block) including 2D convolution (two-dimensional convolution), Batch Norm (batch normalization), and SiLU activation function.

[0065] In step S223, feature extraction downsampling is performed on the initial feature map through multiple Conv blocks to obtain two sets of new features (both including global features and local features).

[0066] Here, each of the two sets of new features contains several feature maps with decreasing sizes (for example, in the case where the initial feature map for each set of features includes 5 feature maps, the feature maps after size reduction here may include 4). The higher the layer, the smaller the size, and the more rich the semantic information it contains, but the spatial information cannot be well expressed.

[0067] In step S224, the local features are combined with their global dependencies in parallel and adaptively using a dual multi-scale attention network.

[0068] In step S225, the corresponding features at each level of the two backbone networks are fused in the channel dimension to obtain a richer feature representation.

[0069] In step S226, the fused features are compressed through a Conv block to reduce the subsequent computational amount, thereby obtaining the multi-scale feature extraction result.

[0070] In addition, in Figure 3 N 0 represents any number of repetitions, and the input to the auxiliary network in Figure 3 is also image data.

[0071] As an example, the multi-scale feature enhancement process may include the following processes: dividing the multi-scale feature extraction result obtained through the multi-scale feature extraction process into multiple input features in the channel dimension; performing dynamic group convolution and channel shuffle processing on a part of the multiple input features to obtain a first feature; fusing another part of the multiple input features that have not been processed with the first feature in the channel dimension to obtain a second fused feature; inputting the second feature into a preset feed-forward network included in the feature enhancement network based on dynamic group convolution for enhancing the context understanding ability to obtain a processed third feature as the multi-scale feature enhancement result.

[0072] Through the above multi-scale feature enhancement process based on group convolution, the enhancement of multi-scale features is achieved. For example, by introducing channel shuffle and dynamic group convolution, the diversity of features is increased, and the model's ability to understand context is improved.

[0073] Referring to Figure 2 and Figure 4 , in step S23, a module such as Group ShuffleTransformer that incorporates group convolution is used to enhance the features.

[0074] As an example, step S23 may further include the following steps S231 to S236:

[0075] In step S231, the input features are divided in the channel dimension.

[0076] In step S232, a part of the features pass through dynamic group convolution, and the number of groups of convolution kernels is adaptively selected according to different regions of the feature map.

[0077] In step S233, further channel shuffle is performed on the features after dynamic group convolution to increase the diversity of features.

[0078] In step S234, another part of the features is not processed but used as a residual channel.

[0079] In step S235, the above two parts of features are re - fused in the channel dimension.

[0080] In step S236, the fused features are processed by a Transformer feed - forward network to enhance the context understanding ability, and a multi - scale feature enhancement result is obtained.

[0081] In addition, in Figure 4 the blank boxes without labeled text explanations shown represent intermediate features.

[0082] As an example, the multi - scale feature fusion process may include the following processes: performing global feature information extraction processing and feature encoding processing on the multi - scale feature enhancement result obtained through multi - scale feature enhancement processing to obtain a processed first multi - scale feature map; connecting and splitting the first multi - scale feature map in the channel dimension to obtain two parts of multi - scale feature maps; inputting one part of the two parts of multi - scale feature maps into a re - parameterized network in a multi - scale feature fusion network to perform convolution and several re - parameterization processes to obtain a second multi - scale feature map; inputting the other part of the two parts of multi - scale feature maps into a preset cross - layer aggregation network in the multi - scale feature fusion network to perform cross - layer aggregation processing to obtain a third multi - scale feature map; connecting the third multi - scale feature map with the second multi - scale feature map and performing convolution processing on the connected multi - scale feature map to obtain a fused multi - scale feature map as the multi - scale feature fusion result.

[0083] For example, the first multi - scale feature map may include multiple feature maps of different scales.

[0084] Through the above multi - scale feature fusion process, by flexibly configuring the number of channels of different scales, multi - scale feature fusion can be efficiently completed. For example, through the above multi - scale feature enhancement process, the fusion of features at each scale can be quickly realized, while improving the target detection effect of the model, reducing the number of model parameters.

[0085] Referring to Figure 2 and Figure 5 , in step S24, for example, an Efficient RepGFPN network can be used for multi - scale feature fusion, and by flexibly configuring the number of channels of different scales, multi - scale feature fusion can be efficiently completed.

[0086] As an example, step S24 may further include the following steps S241 to S246:

[0087] In step S241, a lightweight MLP (Multi-Layer Perceptron) is used to capture global feature information, and the feature data is encoded using VNC (Virtual Network Computing).

[0088] In step S242, the multi-size feature maps are concatenated in the channel dimension and split.

[0089] In step S243, a part of the split data is introduced into the N - time reparameterization structure after, for example, a 1*1 convolution operation.

[0090] In step S244, another part of the split data is directly connected to the previous part in the way of ELAN (Efficient Layer Aggregation Network) connection.

[0091] In step S245, another convolution is performed to obtain the fused feature map.

[0092] Here, it should be noted that a two - branch structure is adopted in the training stage, and the two branches are fused together during inference as the multi - scale feature fusion result.

[0093] In addition, in Figure 4 , the part in the upper half dotted box corresponds to the reparameterization structure network, and the part outside the lower half dotted box corresponds to ELAN.

[0094] In addition, as an example, the traffic target detection model based on dual - stream feature fusion enhancement may further include a target prediction network for performing prediction processing on the multi - scale feature fusion result obtained through multi - scale feature fusion processing to obtain an output data set. Here, the prediction network can be implemented using various networks for prediction in the art, and no specific limitation is made in this disclosure.

[0095] The prediction result is output through the target prediction network to optimize and adjust the model parameters to obtain a trained model with adjusted parameters.

[0096] Referring back again to Figure 1 , in step S104, according to the validation data set and the output data set, each parameter of the traffic target detection model based on dual - stream feature fusion enhancement is adjusted to obtain the trained traffic target detection model.

[0097] Here, each parameter of the traffic target detection model may include the parameters of each network included in the model. For example, the parameters of the three networks (i.e., the dual - stream feature extraction backbone network, the feature enhancement network based on dynamic group convolution, and the multi - scale feature fusion network) included in the model.

[0098] The trained traffic target detection model obtained by the above training method according to an embodiment of the present disclosure can overcome the problems of large scale variations and overlapping occlusions of traffic targets, and improve the detection accuracy of traffic targets by the traffic system.

[0099] As an example, as Figure 2 shown in step S25 of

[0100] By inputting the fused feature map into the prediction network, the final target type and location information are obtained. For example, the Focal EIoU (a loss function that combines Focal Loss and EIoU (an improved IoU (Intersection over Union) loss)) can be used to optimize the calculation of the bounding box loss. The side length is used as a penalty term to avoid the influence of relative proportions, and at the same time, the problem of drastic fluctuations in the loss value caused by low-quality samples is solved. By using three decoupled detection heads, the input feature maps of three different scales are calculated to obtain the regression loss and classification loss, and the final prediction results are calculated, including the target location, bounding box, target category, and confidence.

[0101] Here, it should be noted that the above descriptions of the Figure 1 individual steps are merely exemplary, and the steps in the method according to the present disclosure are not limited thereto.

[0102] In addition, the present disclosure also provides a traffic target detection method 600. Referring to Figure 6 Figure 6 is a flowchart showing the traffic target detection method 600 according to an embodiment of the present disclosure. The traffic target detection method 600 includes steps S61 and S62:

[0103] According to the present disclosure, in step S61, traffic target image data to be detected is obtained.

[0104] According to the present disclosure, in step S62, the traffic target image data to be detected is input into the traffic target detection model based on dual-stream feature fusion enhancement to obtain corresponding traffic target detection data. Here, the traffic target detection model based on dual-stream feature fusion enhancement is trained by using the training method 100 as described above.

[0105] Next, referring to Figure 7 to describe the training apparatus for the traffic target detection model based on deep learning and dual-stream feature fusion enhancement according to an embodiment of the present disclosure.

[0106] Figure 7 is a block diagram showing the training apparatus 700 for the traffic target detection model based on deep learning and dual-stream feature fusion enhancement according to an embodiment of the present disclosure. ​

[0107] For example, a traffic target detection model based on dual-stream feature fusion enhancement includes a dual-stream feature extraction backbone network for parallel data processing, a feature enhancement network based on dynamic group convolution, and a multi-scale feature fusion network connected in sequence. Here, the dual-stream feature extraction backbone network is used to perform multi-scale feature extraction processing, the feature enhancement network based on dynamic group convolution is used to perform multi-scale feature enhancement processing, and the multi-scale feature fusion network is used to perform multi-scale feature fusion processing.

[0108] Referring to Figure 7 , according to an embodiment of the present disclosure, a training apparatus 700 for a traffic target detection model based on deep learning and dual-stream feature fusion enhancement may include a data acquisition module 710, a data preprocessing module 720, a feature training module 730, and a parameter adjustment module 740.

[0109] According to an embodiment of the present disclosure, the data acquisition module 710 may execute: acquiring traffic target image data.

[0110] According to an embodiment of the present disclosure, the data preprocessing module 720 may execute: performing preprocessing including annotation and normalization on the traffic target image data to obtain an input data set including a training data set and a validation data set.

[0111] According to an embodiment of the present disclosure, the feature training module 730 may execute: inputting the input data set into a traffic target detection model based on dual-stream feature fusion enhancement, and obtaining an output data set including traffic target types and location information by performing multi-scale feature extraction processing, multi-scale feature enhancement processing, and multi-scale feature fusion processing on the input data set.

[0112] As an example, the multi-scale feature extraction processing executed by the feature training module 730 may include the following processes 7311) to 7313):

[0113] In process 7311), the training data set is divided into two branch data sets, and the two branch data sets are respectively input into two backbone networks and a dual multi-scale attention network included in the dual-stream feature extraction backbone network to parallelly execute the following processes (1) and (2):

[0114] (1) Executed by two backbone networks: inputting the two branch data sets into a first convolutional block including 2D convolution, normalization, and a preset activation function to obtain corresponding initial feature maps, and inputting the initial feature maps into a second convolutional block for performing feature extraction downsampling to obtain two sets of updated first feature maps.

[0115] (2) Executed by the dual multi-scale attention network: combining the local features corresponding to the training data set with the corresponding global dependencies to obtain a combined second feature map.

[0116] In process 7312), the features at each level of the two backbone networks corresponding to the first feature map and the second feature map are fused in the channel dimension to obtain a fused third feature map.

[0117] In process 7313), the third feature map is compressed to obtain a compressed fourth feature map as the multi-scale feature extraction result.

[0118] According to an embodiment of the present disclosure, the fused third feature map can be obtained by using formula (1) as described above.

[0119] As an example, the multi-scale feature enhancement process performed by the feature training module 730 may include the following processes 7321) to 7324):

[0120] In process 7321), the multi-scale feature extraction result obtained by the multi-scale feature extraction process is divided into multiple input features in the channel dimension.

[0121] In process 7322), dynamic group convolution and channel shuffle processing are performed on a part of the multiple input features to obtain a first feature.

[0122] In process 7323), another part of the multiple input features that have not been processed is fused with the first feature in the channel dimension to obtain a fused second feature.

[0123] In process 7324), the second feature is input into a preset feed-forward network included in the feature enhancement network based on dynamic group convolution to enhance the context understanding ability, and a processed third feature is obtained as the multi-scale feature enhancement result.

[0124] As an example, the multi-scale feature fusion process performed by the feature training module 730 may include the following processes 7331) to 7334):

[0125] In process 7331), global feature information extraction processing and feature encoding processing are performed on the multi-scale feature enhancement result obtained by the multi-scale feature enhancement process to obtain a processed first multi-scale feature map.

[0126] In process 7332), two parts of multi-scale feature maps are obtained by connecting and splitting the first multi-scale feature map in the channel dimension.

[0127] In process 7333), a part of the two parts of multi-scale feature maps is input into the reparameterization network in the multi-scale feature fusion network to perform convolution and several reparameterization processes to obtain a second multi-scale feature map.

[0128] In process 7334), the other part of the two-part multi-scale feature map is input into a preset cross-layer aggregation network in the multi-scale feature fusion network to perform cross-layer aggregation processing, obtaining a third multi-scale feature map; the third multi-scale feature map is concatenated with the second multi-scale feature map, and the concatenated multi-scale feature map is subjected to convolution processing to obtain a fused multi-scale feature map as the multi-scale feature fusion result.

[0129] In addition, for example, the traffic target detection model based on dual-stream feature fusion enhancement may further include a target prediction network for performing prediction processing on the multi-scale feature fusion result obtained through multi-scale feature fusion processing to obtain an output data set.

[0130] Specifically, after process 7334), the above-mentioned target prediction network performs prediction processing on the multi-scale feature fusion result obtained through multi-scale feature fusion processing to obtain an output data set.

[0131] According to an embodiment of the present disclosure, the parameter adjustment module 740 may perform: adjusting each parameter of the traffic target detection model based on dual-stream feature fusion enhancement according to the validation data set and the output data set to obtain a trained traffic target detection model.

[0132] It should be noted that the operations performed by the above-mentioned respective structural blocks may be similar to the related content described with reference to Figure 1 and will not be elaborated here.

[0133] In addition, the present disclosure also provides a traffic target detection device 800. Figure 8 is a block diagram showing a traffic target detection device 800 according to an embodiment of the present disclosure. The traffic target detection device 800 includes an acquisition module 810 and a detection module 820.

[0134] According to the present disclosure, the acquisition module 810 may perform: acquiring traffic target image data to be detected.

[0135] According to the present disclosure, the detection module 820 may perform: inputting the traffic target image data to be detected into a traffic target detection model based on dual-stream feature fusion enhancement to obtain corresponding traffic target detection data. Here, the traffic target detection model based on dual-stream feature fusion enhancement is trained by using the training method 100 as described above.

[0136] In addition, as an example, the present disclosure also provides a traffic target detection system based on deep learning and dual-stream feature fusion enhancement. The traffic target detection system may include a data set acquisition and production mechanism, a feature extraction mechanism, a feature enhancement mechanism, a feature fusion mechanism, and a target detection mechanism.

[0137] Specifically, the data set acquisition and production institution is used to obtain images of traffic targets from multiple angles, and annotate and preprocess the images to obtain an input data set. The feature extraction institution is used to generate multiple groups of independent feature maps by splitting the initial input, and obtain a composite feature map after fusion. The feature enhancement institution is used to enhance the feature map. The feature fusion institution is used to fuse the information of feature maps at each scale. The target detection institution is used to optimize the calculation of the bounding box loss, obtain the regression loss and the classification loss, and thus calculate the final prediction result.

[0138] It should be noted that the operations performed by each of the above institutions may be similar to the relevant content described with reference to Figure 1 and will not be elaborated here.

[0139] Figure 9 is a block diagram showing a computing device 900 according to an embodiment of the present disclosure.

[0140] With reference to Figure 9 , the computing device 900 according to an embodiment of the present disclosure may include a processor 910 and a memory 920. The processor 910 may include (but is not limited to) a central processing unit (CPU), a digital signal processor (DSP), a microcomputer, a field programmable gate array (FPGA), a system on chip (SoC), a microprocessor, an application specific integrated circuit (ASIC), etc. The memory 920 may store computer executable instructions to be executed by the processor 910. The memory 920 includes high-speed random access memory and / or non-volatile computer-readable storage media. When the processor 910 executes the computer executable instructions stored in the memory 920, the training method of the traffic target detection model based on deep learning and dual-stream feature fusion enhancement and the traffic target detection method as described above can be implemented.

[0141] The training method and traffic target detection method of a traffic target detection model enhanced by deep learning and dual-stream feature fusion according to an embodiment of the present disclosure can be written as computer programs / instructions to form a computer program product and stored on a computer-readable storage medium. When the computer programs / instructions are executed by a processor, the training method and traffic target detection method of the traffic target detection model enhanced by deep learning and dual-stream feature fusion as described above can be implemented. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device / server, the electronic device / server can be enabled to execute the training method and traffic target detection method of the traffic target detection model enhanced by deep learning and dual-stream feature fusion as described above. Examples of computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), card memory (such as multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. In one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.

[0142] The training method and device of a traffic target detection model enhanced by deep learning and dual-stream feature fusion and the traffic target detection method according to an embodiment of the present disclosure can solve the problems of large scale variation and overlapping occlusion of traffic targets by training a traffic target detection model enhanced by deep learning and dual-stream feature fusion, and effectively improve the detection accuracy of traffic targets by, for example, a traffic system.

[0143] In addition, the training method and apparatus for a traffic object detection model based on deep learning and enhanced dual-stream feature fusion according to an embodiment of the present disclosure, and the traffic object detection method increase the diversity of features and improve the model's ability to understand context by introducing channel shuffle and dynamic group convolution.

[0144] In addition, the training method and apparatus for a traffic object detection model based on deep learning and enhanced dual-stream feature fusion according to an embodiment of the present disclosure, and the traffic object detection method can quickly achieve the fusion of features at various scales by using a novel multi-scale feature fusion network, while improving the object detection effect of the model and reducing the number of model parameters.

[0145] Although some embodiments of the present disclosure have been disclosed and described, those skilled in the art should understand that these embodiments can be modified and varied without departing from the concept and spirit of the present disclosure as defined by the claims and their equivalents.

Claims

1. A training method for a traffic target detection model based on deep learning and dual-stream feature fusion enhancement, characterized in that: The training method comprises: Acquire traffic target image data; Performing preprocessing including labeling and standardization on the traffic target image data to obtain an input data set including a training data set and a verification data set; The input data set is input into a traffic target detection model based on dual-stream feature fusion enhancement, and an output data set including traffic target type and location information is obtained by performing multi-scale feature extraction processing, multi-scale feature enhancement processing and multi-scale feature fusion processing on the input data set; According to the verification data set and the output data set, various parameters of the traffic target detection model based on dual-stream feature fusion enhancement are adjusted to obtain a trained traffic target detection model, Among them, the traffic target detection model based on dual-stream feature fusion enhancement includes a dual-stream feature extraction backbone network for parallel data processing, a feature enhancement network based on dynamic group convolution, and a multi-scale feature fusion network connected in sequence, wherein the dual-stream feature extraction backbone network is used to perform the multi-scale feature extraction processing, the feature enhancement network based on dynamic group convolution is used to perform the multi-scale feature enhancement processing, and the multi-scale feature fusion network is used to perform the multi-scale feature fusion processing.

2. The training method according to claim 1, characterized in that: The multi-scale feature extraction process includes the following processes: The training data set is divided into two tributary data sets and the two tributary data sets are respectively input into two backbone networks and a dual multi-scale attention network included in the dual-stream feature extraction backbone network to perform the following processing in parallel: It is performed by two backbone networks: The two tributary data sets are input into the first convolution block including 2D convolution, normalization and preset activation function to obtain the corresponding initial feature map, Inputting the initial feature map into a second convolution block for performing feature extraction downsampling to obtain two sets of updated first feature maps, and The dual multi-scale attention network performs: combining the local features corresponding to the training data set with the corresponding global dependencies to obtain a combined second feature map; Fusing the features of each level of the two backbone networks corresponding to the first feature map and the second feature map in the channel dimension to obtain a fused third feature map; The third feature map is compressed to obtain a compressed fourth feature map as a multi-scale feature extraction result.

3. The training method according to claim 1, characterized in that: The multi-scale feature enhancement process includes the following processes: Dividing a multi-scale feature extraction result obtained by the multi-scale feature extraction process into a plurality of input features in a channel dimension; Performing dynamic grouped convolution and channel shuffling processing on a portion of the multiple input features to obtain a first feature; Fusing another part of the unprocessed input features with the first feature in a channel dimension to obtain a fused second feature; The second feature is input into a preset feedforward network for enhancing context understanding capability included in the feature enhancement network based on dynamic grouping convolution, to obtain a processed third feature as a multi-scale feature enhancement result.

4. The training method according to claim 1, characterized in that: The multi-scale feature fusion process includes the following processes: Performing global feature information extraction processing and feature encoding processing on the multi-scale feature enhancement result obtained by the multi-scale feature enhancement processing to obtain a processed first multi-scale feature map; By connecting and splitting the first multi-scale feature map in the channel dimension, two multi-scale feature maps are obtained; Inputting one part of the two multi-scale feature maps into a re-parameterized network in the multi-scale feature fusion network to perform convolution and several re-parameterization processes to obtain a second multi-scale feature map; Inputting the other part of the two multi-scale feature maps into a preset cross-layer aggregation network in the multi-scale feature fusion network to perform cross-layer aggregation processing to obtain a third multi-scale feature map; The third multi-scale feature map is connected to the second multi-scale feature map, and a convolution process is performed on the connected multi-scale feature map to obtain a fused multi-scale feature map as a multi-scale feature fusion result.

5. The training method according to claim 2, characterized in that: The fused third feature map is obtained using the following formula: Among them, y out represents the third characteristic map, x initial represents the training data set, F k* represents the feature extraction process with k-level processing and serves as the main backbone network, represents the i-th auxiliary backbone network with k levels, k, l, i are natural numbers, and the value of i is greater than or equal to 1 and less than or equal to l.

6. The training method according to claim 1, characterized in that: The traffic target detection model based on dual-stream feature fusion enhancement also includes a target prediction network, which is used to perform prediction processing on the multi-scale feature fusion result obtained by the multi-scale feature fusion processing to obtain the output data set.

7. A training device for a traffic target detection model based on deep learning and dual-stream feature fusion enhancement, characterized in that: The training device comprises: The data acquisition module is configured to: acquire traffic target image data; The data preprocessing module is configured to: perform preprocessing including labeling and standardization on the traffic target image data to obtain an input data set including a training data set and a verification data set; The feature training module is configured to: input the input data set into the traffic target detection model based on dual-stream feature fusion enhancement, and obtain an output data set including traffic target type and location information by performing multi-scale feature extraction processing, multi-scale feature enhancement processing and multi-scale feature fusion processing on the input data set; The parameter adjustment module is configured to: adjust various parameters of the traffic target detection model based on dual-stream feature fusion enhancement according to the verification data set and the output data set to obtain a trained traffic target detection model, Among them, the traffic target detection model based on dual-stream feature fusion enhancement includes a dual-stream feature extraction backbone network for parallel data processing, a feature enhancement network based on dynamic group convolution, and a multi-scale feature fusion network connected in sequence, wherein the dual-stream feature extraction backbone network is used to perform the multi-scale feature extraction processing, the feature enhancement network based on dynamic group convolution is used to perform the multi-scale feature enhancement processing, and the multi-scale feature fusion network is used to perform the multi-scale feature fusion processing.

8. A traffic target detection method, characterized in that: The traffic target detection method comprises: Acquire image data of traffic targets to be detected; The traffic target image data to be detected is input into the traffic target detection model based on dual-stream feature fusion enhancement to obtain corresponding traffic target detection data. The traffic target detection model based on dual-stream feature fusion enhancement is trained using the training method described in any one of claims 1 to 6.

9. A computer program product, characterized in that The computer program product includes a computer program / instruction, which, when executed by a processor, implements the training method of the traffic target detection model based on deep learning and dual-stream feature fusion enhancement as described in any one of claims 1 to 6 and the traffic target detection method as described in claim 8.

10. A computing device, characterized in that: The computing device includes: at least one processor; at least one memory storing computer executable instructions, wherein the computer executable instructions, when executed by the at least one processor, cause the at least one processor to execute the training method of the traffic target detection model based on deep learning and dual-stream feature fusion enhancement as described in any one of claims 1 to 6 and the traffic target detection method as described in claim 8.

Citation Information

Cited By

  • Method for detecting objects in car window based on deep learning

    CN122223665A