Traffic scene multi-target-of-interest recognition method based on transformer-yolo

By improving the Transformer-Yolo-based method, feature extraction and robustness are enhanced, solving the problem of identifying specific targets of interest in traffic scenarios. This achieves high-precision and high-real-time multi-target detection, which is suitable for autonomous driving and traffic monitoring systems.

CN120411894BActive Publication Date: 2026-02-03CHANGAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510378605.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2026-02-03
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

Existing YOLO algorithms cannot effectively identify specific targets of interest in complex and ever-changing traffic scenarios. They lack feature extraction capabilities, real-time performance, and accuracy, and therefore cannot meet the multi-target detection needs in traffic scenarios.

Method used

We adopt a Transformer-Yolo-based approach, which enhances feature extraction capabilities and robustness by adding improved DfDC and VSST modules and combining deformable convolutions DCNv2 and Cf2. We also use the Focal SIoU loss function to optimize model performance and adapt to complex and diverse traffic scenarios.

Benefits of technology

It achieves accurate identification and early warning of specific targets of interest in complex traffic scenarios, improves the robustness and real-time performance of the model, adapts to different weather conditions and occlusion situations, and meets the needs of traffic safety assurance and emergency handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411894B_ABST
    Figure CN120411894B_ABST
Patent Text Reader

Abstract

The application discloses a traffic scene interested multi-target detection and recognition method based on a Transform-Yolo, which comprises the following steps: step 1, establishing a traffic scene image set; step 2, obtaining a traffic scene interested multi-target recognition data set; step 3, establishing a backbone network model and adding an improved DfDC module based on Cf2; step 4, establishing a neck network model and adding an improved attention mechanism VSST module; step 5, establishing a detection network model based on Transform-Yolo; step 6, reading a data set and using migration learning to train the network model; step 7, detecting and recognizing a test set; and step 8, inputting a traffic scene image to be recognized into the trained network model and outputting a recognition result. The network in the application is more suitable for shape changes of multi-targets with complex and various features, can solve the accurate recognition and early warning of specific interested targets in different traffic scenes, and maximally reduces road safety problems.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of image recognition, and particularly relates to a traffic scene interested multi-target recognition method based on a Transformer-Yolo. BACKGROUND

[0002] With the intelligent and integrated development of the comprehensive transportation system, the rapid advancement of urbanization, and the substantial growth of transportation demand, traffic management is facing increasingly complex challenges, and the multi-target recognition task has become one of the core tasks of intelligent transportation, and its application in traffic management, intelligent monitoring, and autonomous driving is becoming increasingly important. Multi-target detection in traffic scenes not only includes conventional targets such as vehicles, pedestrians, non-motor vehicles, and traffic roads, but also includes the detection and identification of specific interested targets, such as emergency vehicles, traffic accidents, traffic signs, traffic signal lights, and special foreign objects, and provides prompts. The recognition of interested multi-targets in traffic scenes plays a crucial role in traffic management, road safety, and emergency response.

[0003] Currently, traditional multi-target task recognition in traffic scenes is mostly based on the identification and detection of ordinary vehicles, pedestrians, roads, and other target tasks, and there are few specific interested multi-target recognition and detection methods, which face multiple challenges. First, the challenge of diversified targets, the target objects in traffic scenes are diverse, including different types of vehicles, pedestrians, traffic signs, and signal lights, and these targets have different characteristics in different scenes. Second, the challenge of complex background and environmental interference, the road environment is complex, with many background elements such as buildings, trees, and other vehicles, and sometimes the light conditions such as rain, snow, night, or strong light will interfere with the detection results. Third, the challenge of real-time requirements, real-time detection is particularly urgent in traffic scenes, whether it is autonomous driving, traffic flow monitoring, or emergency response, the detection system must have high speed and accuracy. Fourth, the challenge of occlusion and overlap, multi-target occlusion and overlap are very common in traffic scenes, which poses a great challenge to traditional target detection methods, especially when interested targets may be partially occluded by other objects, the accuracy of recognition will be greatly affected. In order to cope with these multiple challenges, the interested multi-target detection and recognition method in traffic scenes not only needs to identify conventional targets, but also needs to accurately detect specific interested targets, and adapt to different scene detection tasks. In order to achieve this goal, it is necessary to seek a detection and recognition method with high precision, high real-time performance, and robustness.

[0004] Obviously, the traditional method cannot meet this demand, compared with the deep learning technology, which has high processing performance and strong operability in solving the image multi-target processing task. At present, the target detection algorithm of deep learning is divided into two-stage algorithm and single-stage. The two-stage algorithm has high accuracy, but the speed is slightly slow. The single-stage recognition effect is poor, but the speed is faster, mainly including SSD, YOLO series algorithm. Among them, YOLO is widely used in practical engineering construction because of its powerful and efficient detection performance, and has become the mainstream recognition algorithm in the traffic scene. But if it is directly used in the multi-target recognition task of interest in the traffic scene, there are too many target types, and the background is complex, so the feature extraction ability is insufficient, and the specific multi-target of interest cannot be extracted more accurately and quickly, and the safety of the traffic and the handling of the emergency situation have higher requirements for real-time. The current YOLO algorithm cannot meet the requirements of the multi-target recognition task of interest in the traffic scene.

[0005] Therefore, in order to improve the performance of the multi-target detection of interest in the complex traffic background, the application provides a traffic scene multi-target recognition method of interest based on Transformer-Yolo. SUMMARY

[0006] The purpose of the application is to provide a traffic scene multi-target recognition method of interest based on Transformer-Yolo, which can accurately identify and warn specific targets of interest in different traffic scenes, and minimize road safety problems. The problems of poor feature extraction ability of multi-target in complex and changeable traffic scenes, slightly poor real-time accuracy of accurate judgment and lack of ability to adapt to complex scenes in the prior art are solved.

[0007] In order to achieve the above purpose, the technical scheme adopted by the application is as follows:

[0008] A traffic scene multi-target detection and recognition method of interest based on Transformer-Yolo, specifically comprising the following steps:

[0009] Step 1, collect traffic scene image data, establish a traffic scene image set containing multiple traffic target categories, the traffic target categories include ordinary traffic targets and specific traffic targets that may be of interest;

[0010] Step 2, screen the obtained traffic scene data image set, classify and label the ordinary traffic targets and specific traffic targets that may be of interest in the screened images, obtain a traffic scene multi-target recognition data set, and divide it into a training set, a validation set and a test set;

[0011] Step 3: Build a backbone network model based on Transformer-Yolo and add an improved DfDC module based on Cf2;

[0012] Step 4: Build a neck network model based on Transformer-Yolo and add the VSST module, an improved attention mechanism based on Swing Transformer.

[0013] Step 5: Build a detection network model based on Transformer-Yolo, using Focal SIoU as the loss function;

[0014] Step 6: Read the traffic scene multi-object recognition dataset obtained in Step 2, and use transfer learning to fully train the network model to obtain the trained network model;

[0015] Step 7: Use the weights obtained from training with transfer learning to detect and identify on the test set to complete the performance evaluation of the network model;

[0016] Step 8: Input the traffic scene image to be identified into the trained network model. The model detects and identifies multiple targets of interest in the image and outputs the identification results, including target category, location bounding box, and confidence score.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0018] This invention presents a Transformer-Yolo-based method for identifying multiple targets of interest (ROI) in traffic scenes. This method can extract key information of ROI from complex traffic scenes and effectively distinguish them from other background elements. It is robust, capable of handling target detection requirements under different weather conditions, lighting variations, and occlusion, maintaining high recognition accuracy and exhibiting high real-time performance. It maximizes traffic safety and emergency response capabilities, and can be widely applied to corresponding autonomous driving systems, traffic congestion detection and early warning systems, and traffic flow monitoring and early warning systems. This invention proposes a CfCD module, combining deformable convolutional modeling (DCNv2) and Cf2. This module adapts to the shape variations of multiple targets with complex and diverse features, enhancing its feature extraction and feature reuse capabilities. Furthermore, this invention improves the Swing Transformer module, specifically the VSST module, which combines the advantages of Swing Transformer and VSS. This module dynamically adjusts feature weights by combining local and global contextual features, making the model more focused on regions of interest, capturing multi-target performance at different scales, and enhancing the model's robustness. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating the method for detecting and recognizing multiple targets of interest in traffic scenes based on Transformer-Yolo according to the present invention.

[0020] Figure 2 This is a schematic diagram of the network structure of the Transformer-Yolo-based method for detecting and recognizing multiple targets of interest in traffic scenes according to the present invention;

[0021] Figure 3 This is a schematic diagram of the DfDC network structure;

[0022] Figure 4 This is an improved Swing Transormer module: VSST network structure diagram;

[0023] Figure 5 This invention, Transformer-Yolo, describes the detection and recognition method for multiple targets of interest in traffic scenes, specifically for identifying and recognizing multiple targets of interest within a traffic scene. Detailed Implementation

[0024] The following detailed description of the Transformer-Yolo method for detecting and recognizing multiple targets of interest in traffic scenes, in conjunction with the accompanying drawings and specific embodiments, provides a detailed explanation of the present invention.

[0025] like Figure 1 As shown, the method for detecting and recognizing multiple targets of interest in traffic scenes based on Transformer-Yolo is implemented according to the following steps:

[0026] Step 1: Collect traffic scene image data and establish a traffic scene image set containing multiple traffic target categories. The traffic target categories include common traffic targets: pedestrians and vehicles, as well as specific traffic targets that may be of interest: emergency vehicles, illegally parked vehicles, traffic accidents, traffic signs, and special foreign objects, etc.

[0027] Step 2: Filter the acquired traffic scene data image set, classify and label the ordinary traffic targets and potentially interesting specific traffic targets in the selected images to obtain the traffic scene multi-target recognition dataset, and divide it into training set, validation set and test set;

[0028] Step 3: Build a backbone network model based on Transformer-Yolo and add an improved DfDC module based on Cf2 to adapt to shape changes of multiple targets of interest and enhance their feature extraction capabilities.

[0029] Step 4: Build a neck network model based on Transformer-Yolo and add an improved attention mechanism VSST module based on Swing Transformer to enhance the ability to locate information of interest and capture the performance of the target at different scales, thereby enhancing the robustness of the model.

[0030] Step 5: Build a detection network model based on Transformer-Yolo. The loss function used is Focal SIoU to reduce the computational cost of the model and improve the overall optimization performance of the model.

[0031] Step 6: Read the traffic scene multi-object recognition dataset obtained in Step 2, and use transfer learning to fully train the network model to obtain the trained network model;

[0032] Step 7: Use the weights obtained from training with transfer learning to detect and identify on the test set, thus completing the performance evaluation of the network model.

[0033] Step 8: Input the traffic scene image to be identified into the trained network model. The model detects and identifies multiple targets of interest in the image and outputs the identification results, including target category, location bounding box, and confidence score.

[0034] This invention presents a Transformer-Yolo-based method for detecting and recognizing multiple targets of interest in traffic scenes. By proposing a DfDC module, an improved attention mechanism VSST module based on the Swing Transformer, and replacing the regression loss function with Focal SIoU, the network can adapt to the shape variations of multiple targets with complex and diverse features, enhancing its feature extraction and feature reuse capabilities. Furthermore, by combining local and global contextual features, the weights of features are dynamically adjusted, allowing the model to focus more on regions of interest and capture the performance of multiple targets at different scales, thus enhancing the model's robustness. In practical applications, due to the diversity and complexity of traffic scenes, it is necessary to collect images from multiple data sources, such as vehicle-mounted surveillance cameras, drones, and mobile phone photography, and then upload the captured images to a computer for multi-target detection and recognition. This solves the problem of accurately identifying and issuing warnings for specific targets of interest in different traffic scenarios, minimizing road safety issues and handling emergencies.

[0035] The present invention will be further described below with reference to specific embodiments.

[0036] Example 1;

[0037] This embodiment uses a Transformer-Yolo-based method for detecting and recognizing multiple objects of interest in traffic scenes, including the following steps:

[0038] In step 1, due to the diversity and complexity of traffic scenarios, it is necessary to collect data from multiple data sources, using vehicle-mounted surveillance cameras, drones, and mobile phones from three different directions to improve the diversity and representativeness of the data.

[0039] Step 2 is as follows:

[0040] Step 2.1: Import images from different traffic scenarios captured by vehicle-mounted surveillance cameras, drones, and mobile phones into the computer, and manually filter and label them into ordinary traffic targets and specific traffic targets of interest. Ordinary traffic targets include conventional targets such as vehicles, pedestrians, non-motorized vehicles, and traffic roads, while specific traffic targets of interest include special objects such as emergency vehicles, illegally parked vehicles, traffic accidents, traffic signs, and traffic lights.

[0041] In step 2.1, when filtering traffic scene images, images that are blurry, excessively noisy, or poorly exposed are first removed through image sharpness detection to ensure the quality of the input data. Then, images without obvious targets or with severely occluded targets are filtered out. Next, based on the requirements of scene diversity, data covering various scenes such as urban roads, highways, and intersections are retained, including data from daytime, nighttime, different weather conditions, and various acquisition devices. At the same time, sufficient samples are ensured for each type of traffic target to avoid class imbalance. Finally, highly similar images are redundantly removed, and only representative samples are retained, thereby improving the representativeness and diversity of the data and providing a high-quality data foundation for subsequent annotation and training.

[0042] Step 2.2: The 2570 selected images are labeled according to different target categories. The LabelImg annotation tool is used and the images are labeled with reference to the CoCo dataset. The annotation information of each image is saved in txt file format, which mainly includes: the overall image height and width information, and the specific position information of the four vertices of the target annotation rectangle.

[0043] Step 2.3: Divide the labeled images into sets with a ratio of 8:1:1 (training set: validation set: test set).

[0044] Step 3 specifically involves:

[0045] Step 3.1: Combine deformable convolution DCNv2 with Cf2 to form the CfDC module: Replace the Bottleneck block in Cf2 with deformable convolution DCNv2. The outer structure of the CfDC module is similar to that of Cf2, consisting of two Conv blocks, n deformable convolution DCNv2 layers, a Chunk block, and a Concat block. The deformable convolution DCNv2 consists of multiple deformable convolutional layers, multiple standard convolutional layers, and an offset prediction layer connected in sequence. Among them, the deformable convolutional layers dynamically adjust the position of the convolutional kernel by introducing offsets to adapt to the shape of the input feature map; the standard convolutional layers include standard convolution operations to extract important features; and the offset prediction layer is used to calculate the offset of the convolutional kernel to ensure the flexibility of the deformable convolution.

[0046] In this embodiment, a CfDC module is constructed: the feature map output by the convolutional module above the CfDC module is 1024×20×20 in size, as shown below. Figure 3 As shown, after the feature map enters the CfDC module, it first goes through a 1×1 Conv block, where the output feature map size remains unchanged at 1024×20×20. Then it passes through a Chunk block, which divides the feature map into two parts, both 512×20×20. The deformable convolution DCNv2 receives the 512×20×20 feature input, passes through three 1×1 deformable convolutional layers, three 3×3 standard convolutional layers, and an offset prediction layer, outputting a feature map with the same 512×20×20 dimension. This is then fed into the Concat module, which receives two 512×20×20 feature outputs from the Chunk operation and n×512×20×20 feature outputs from the deformable convolution DCNv2. The concatenated feature map has a dimension of (2+n)×512×20×20. Finally, it passes through a 1×1 Conv block, outputting a feature map with a size of 1024×20×20.

[0047] Step 3.2: Replace the last Cf2 in the backbone network, i.e. the layer above SPPF, with the proposed CfDC module to adapt to the shape changes of multi-targets with complex and diverse features, and enhance its feature extraction and feature reuse expression capabilities.

[0048] Step 4 specifically involves:

[0049] Step 4.1: Construct an improved attention mechanism VSST module based on the Swin transformer: Use VSSblock as a long-range low-level feature location extractor to extract long-range low-level features, and use SwinTransformer on top of the output for high-level short-range feature location extraction and task processing, such as... Figure 4As shown, a VSS block module is added above the Swin TransformerBlock in each of the four stages of the conventional Swin transformer module, so that each stage consists of a Patch expanding layer, a VSS block module, and several Swin Transformer Block modules connected in sequence.

[0050] In this embodiment, the VSST module replaces the two Cf2 modules used for downsampling in the Transformer-Yolo neck network. Taking the replacement of the first downsampling module in the neck network as an example, the improved attention mechanism VSST module's input feature map has a dimension of 512×40×40. The feature map is fed into the Patch partition layer, decomposed into 25×8×8×512 blocks, i.e., 25 small blocks of 8×8×512. Then, after passing through the Patch expanding layer, the block features are recombined into a complete feature map, restoring the dimension to 512×40×40. Subsequently, the feature map enters the VSS Block module, passing through a normalization layer LN (Layer). The output of the normalization layer is fed into two branches. The first branch consists of a linear layer (Linera), a 1×1 depthwise separable convolution (DWConv), an SS2D block, and a normalization layer (LN) connected in sequence to capture the multi-target performance features at different scales. The second branch consists of a normalization layer (LN). The outputs of the two branches are multiplied and fed into a linear layer (Linera), which is then added to the input of the VSS Block module to obtain the output of the VSS Block module. The number of feature map channels in the input and output remains unchanged at 512×40×40.

[0051] Next, the Swin Transformer Block receives the output from the VSS Block. The Swin Transformer Block is the core module of the Swin Transformer, and its structure includes a normalization layer (LayerNorm, LN), a shifted window multi-head self-attention (SW-MSA), a multi-head self-attention (W-MSA), a feedforward fully connected network (MLP), etc. This module first normalizes the input feature map through the LN layer to ensure a stable distribution across channels. Then, the normalized feature map is divided into multiple non-overlapping windows, each 7×7 in size. Multi-head self-attention (MSA) is independently computed within these windows, capturing feature correlations within the window by generating Query, Key, and Value. Subsequently, to enhance cross-window feature interaction, a sliding window mechanism is used, sliding the window position relative to each other by a certain step and recalculating self-attention, thereby achieving the modeling of both local and global features. After passing through the SW-MSA module, the feature map retains the original input information through residual connections and is then passed to the MLP. The MLP consists of two fully connected layers: the first layer extracts non-linear features using the GELU activation function, and the second layer reshapes the features. Finally, the output of the MLP is again fed back through residual connections, maintaining the feature map dimension and providing richer contextual information for the next module. The feature vectors in this module interact with the feature vectors of the two blocks above it to construct a global feature representation based on an attention mechanism. This process typically generates a 40×40 attention matrix of the same size as the input feature map, resulting in a final input dimension of 512×40×40. The subsequent three stages are as described above. After four stages of feature extraction, the final output feature map has a dimension of 512×40×40.

[0052] Step 4.2: Replace the two Cf2 modules used for downsampling in the neck network with the improved attention mechanism VSST module. This module combines local and global contextual features to dynamically adjust the feature weights, making the model pay more attention to the region of interest, capturing multi-target performance at different scales, and enhancing the model's robustness.

[0053] Step 5 specifically involves:

[0054] Step 5.1 replaces the original regression loss function CIoU with Focal SIoU (Smooth Intersection over Union with Focal Loss). This is a loss function that combines IoU and Focal Loss to enhance the performance of the object detection model in handling class imbalance and bounding box regression. This loss function, through IoU calculation and Focal Loss weight adjustment, better addresses the hard and uncertain sample problems in object detection, making the model more convergent and providing an advantage for model training.

[0055] Focal SIoU incorporates the concept of Focal Loss into the bounding box regression loss, combining IoU and Focal Loss to optimize bounding box overlap while paying more attention to hard-to-locate targets and uncertain predictions. The formula for calculating Focal SIoU is:

[0056] L Focal SIoU =(1-SIoU) γ +log (SIoU) (1)

[0057] Where SIoU is the calculated result of IoU, and γ is the adjustment factor in the focus loss, used to adjust the contribution of samples with different overlap rates to the loss.

[0058] γ = 2: This is the default value commonly used by Focal Loss in many tasks. It significantly reduces the weight of easily classified samples, encouraging the model to pay more attention to difficult-to-classify samples.

[0059] γ = 1: This is equivalent to directly weighting the loss function based on IoU, making the loss function more balanced and the weight differences among all samples smaller.

[0060] γ = 0: In this case, Focal Loss is equivalent to ordinary cross-entropy loss or IoU loss, and weight adjustment is completely ignored.

[0061] In practice, the specific value of γ needs to be adjusted experimentally based on the difficulty of the task and the data distribution. In this method, a larger γ = 3 is attempted to more significantly focus on the multi-target samples of interest.

[0062] Step 6 specifically involves:

[0063] A transfer learning approach was used to train a multi-object detection and recognition model for traffic scenarios. Model training consisted of two phases: coarse training and fine-tuning. In the coarse training phase, all convolutional layer weights were frozen, and training was performed only on fully connected layers. The number of iterations was set to 50, the batch size to 16, and the initial learning rate to 0.001. After the coarse training phase, the fine-tuning phase began. In this phase, all convolutional layers were unfrozen, allowing the model to train the entire network. The number of iterations was set to 150, the batch size to 8, and the initial learning rate to 0.0001. To prevent overfitting, early stopping was employed. Training was stopped when the error on the validation set worsened, and the parameters from the previous iteration were used as the final model parameters. The momentum parameter was set to 0.973, and the optimization algorithm used was stochastic gradient descent (SGD).

[0064] Step 7 specifically involves using the weights obtained from transfer learning training to detect and identify the test set, i.e., using the weights generated during the last training to detect the test set. Precision (P), Recall (R), Average Precision (AP), Mean Average Precision (mAP), Frames Per Second (FPS), and Inference time are selected as evaluation metrics for this experiment to complete the performance evaluation of the algorithm. Precision (P): Used to evaluate the model's ability to identify positive samples. Higher precision indicates more detected positive examples. Its calculation formula is shown in equation (2).

[0065]

[0066] Where TP is the number of true positives (samples that are both predicted and actually positive), and FP is the number of false positives (samples that are predicted to be positive but are actually negative). TP+FP represents the total number of samples that are predicted to be positive.

[0067] Recall: measures the proportion of positive samples correctly identified in the test set. A higher recall means that more positive samples are correctly identified. Its calculation formula is shown in equation (3):

[0068]

[0069] Where FN represents false negatives (the number of samples predicted as negative but actually positive). TP+FN represents the total number of actual positive samples.

[0070] Average precision: The value obtained by integrating the precision under different recall rates reflects the overall performance of the target detection algorithm. Its calculation formula is shown in Equation (4):

[0071]

[0072] Where P is precision, r is recall, and AP ranges from 0 to 1.

[0073] Mean Accuracy (AACC): The AP values ​​of multiple categories are averaged to represent the model's performance in the object detection task. The calculation formula is shown in Equation (5):

[0074]

[0075] Where N is the number of classes, and AP is the average accuracy per class. A higher mAP value indicates better model performance.

[0076] Frames Per Second (FPS): In object detection, it represents the number of images that can be detected per second. It is used to evaluate the speed of object detection algorithms. Its calculation formula is shown in Equation (6):

[0077]

[0078] Inference time refers to the time required for a model to process a single input sample, usually measured in milliseconds (ms). Inference time directly affects the model's performance in real-time applications; a shorter inference time means a faster response speed.

Claims

1. A method for detecting and recognizing multiple targets of interest in traffic scenes based on Transformer-Yolo, characterized in that, Specifically, the steps include the following: Step 1: Collect traffic scene image data and establish a traffic scene image set containing multiple traffic target categories, including ordinary traffic targets and traffic targets of interest; Step 2: Filter the acquired traffic scene data image set, classify and label the ordinary traffic targets and traffic targets of interest in the selected images to obtain the traffic scene multi-target recognition dataset, and divide it into training set, validation set and test set; Step 3: Establish a backbone network model based on Transformer-Yolo, and add an improved DfDC module based on Cf2; including the following sub-steps: Step 3.1, Constructing the CfDC Module: The feature map output by the convolutional module preceding the CfDC module is 1024×20×20 in size. After entering the CfDC module, the feature map first goes through a 1×1 Conv block, where the output feature map size remains unchanged at 1024×20×20. Then it passes through a Chunk block, which divides the feature map into two parts, both 512×20×20. The deformable convolution DCNv2 accepts 512×20×20 feature maps. After inputting the features, the output feature map still has a dimension of 512×20×20 after passing through three 1×1 deformable convolutional layers, three 3×3 standard convolutional layers, and an offset prediction layer. It then enters the Concat module, which receives two 512×20×20 feature outputs from the Chunk operation and n×512×20×20 feature outputs from n deformable convolutional DCNv2 layers. Finally, after passing through a 1×1 Conv block, the output feature map has a dimension of 1024×20×20. Step 3.2: Replace the last Cf2 in the backbone network with the proposed CfDC module; Step 4: Build a neck network model based on Transformer-Yolo and add an improved attention mechanism VSST module based on Swing Transformer; replace the two Cf2 modules used for downsampling in the neck network with the improved attention mechanism VSST module. Step 5: Build a detection network model based on Transformer-Yolo, using Focal SIoU as the loss function; Step 6: Read the traffic scene multi-object recognition dataset obtained in Step 2, and use transfer learning to fully train the network model to obtain the trained network model; Step 7: Use the weights obtained from training with transfer learning to detect and identify on the test set to complete the performance evaluation of the network model; Step 8: Input the traffic scene image to be identified into the trained network model. The model detects and identifies multiple targets of interest in the image and outputs the identification results, including target category, location bounding box, and confidence score.

2. The method for detecting and recognizing multiple targets of interest in traffic scenes based on Transformer-Yolo as described in claim 1, characterized in that, In step 1, the traffic scene image data is collected using vehicle-mounted surveillance cameras, drones, and mobile phones.

3. The method for detecting and recognizing multiple targets of interest in traffic scenes based on Transformer-Yolo as described in claim 1, characterized in that, Step 2 includes the following sub-steps: Step 2.1: Filter the images in the traffic scene image set obtained in Step 1; Step 2.2: Classify and label the ordinary traffic targets and traffic targets of interest in the selected images. The labeling information includes: image height and width, and the positions of the four vertices of the target labeling rectangle. Step 2.3: Divide the labeled data images into a ratio of training set: validation set: test set = 8:1:

1.

4. The method for detecting and recognizing multiple targets of interest in traffic scenes based on Transformer-Yolo as described in claim 1, characterized in that, Step 4 includes the following sub-steps: Step 4.1: Construct an improved attention mechanism VSST module based on the Swin transformer. The input feature map of the improved attention mechanism VSST module has a dimension of 512×40×40. The feature map is fed into the Patch partition layer and decomposed into 25×8×8×512, that is, there are 25 small blocks of 8×8×512. Then, after passing through the Patch expanding layer, the block features are recombined into a complete feature map, and the dimension is restored to 512×40×40. Subsequently, the feature map enters the VSS Block module. After passing through a normalized layer LN, the output enters two branches. The first branch includes a linear layer, a depthwise separable convolution with a dimension of 1×1, an SS2D block, and a normalized layer LN connected in sequence to capture the multi-target performance features at different scales. The second branch includes a normalized layer LN. The outputs of the two branches are multiplied and then enter a linear layer, which is added to the input of the VSS Block module to obtain the output of the VSS Block module. The number of channels of the input and output feature maps remains unchanged, still 512×40×40. Step 4.2: Replace the two Cf2 modules used for downsampling in the neck network with the improved attention mechanism VSST module.

5. The method for detecting and recognizing multiple targets of interest in traffic scenes based on Transformer-Yolo as described in claim 1, characterized in that, Step 5 includes the following sub-steps: Step 5.1, replace the original regression loss function CIoU with Focal SIoU, calculated as follows: Where SIoU is the result of IoU calculation. It is a moderating factor in focus loss.

6. The method for detecting and recognizing multiple targets of interest in traffic scenes based on Transformer-Yolo as described in claim 1, characterized in that, Step 6 specifically involves: A transfer learning approach was used to train a multi-object detection and recognition model for traffic scenarios. The model training was divided into a coarse training phase and a fine-tuning phase. In the coarse training phase, all convolutional layer weights were frozen, and only the fully connected layers were trained. The number of iterations was set to 50, the batch size to 16, and the initial learning rate to 0.

001. After completing the coarse training phase, the model entered the fine-tuning phase. In this phase, all convolutional layers were unfrozen, allowing the model to train the entire network. The number of iterations was set to 150, the batch size to 8, and the initial learning rate to 0.0001.