Traffic scene interested multi-target identification method based on Transform-Yolo

Through the Transformer-Yolo-based method, feature extraction and robustness are enhanced, and the problem of insufficient feature extraction and real-time real-time feature extraction of multi-object recognition in traffic scenarios is solved, and high-precision and high-reality detection effects are achieved. It is suitable for applications such as autonomous driving, traffic abnormality detection and early warning.

CN120411894AActive Publication Date: 2025-08-01CHANGAN UNIV +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510378605.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-08-01
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

In the multi-objective recognition task of interest in traffic scenarios, the existing YOLO algorithm has a wide variety of targets, complex backgrounds, insufficient feature extraction capabilities, unable to meet the needs of high precision and high real-time, and cannot adapt to the detection requirements of complex scenarios.

Method used

Using the Transformer-Yolo method, by adding Cf2 improved DfDC module and Swin Transformer's improved attention mechanism VSST module, combined with Focal SIoU loss function, enhance feature extraction ability and robustness, adapt to shape changes of multiple objects, combine local and global features, dynamically adjust feature weights, and improve detection accuracy and real-time performance.

Benefits of technology

It realizes accurate identification and early warning of the targets of interest in complex traffic scenarios, improves the accuracy and real-time detection, adapts to different weather conditions and occlusion conditions, and meets the needs of traffic safety assurance and emergency handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411894A_ABST
    Figure CN120411894A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic scene interested multi-target detection and identification method based on Transform-Yolo. The method comprises the following steps: step 1, establishing a traffic scene image set; step 2, obtaining an interested multi-target identification data set in a traffic scene; step 3, establishing a backbone network model, and adding an improved DfDC module based on Cf2; 4, establishing a neck network model, and adding an improved attention mechanism VSST module; step 5, a detection network model based on Transform-Yolo is established; 6, reading a data set, and carrying out network model training by using transfer learning; step 7, detecting and identifying the test set; and step 8, inputting a traffic scene image to be identified into the trained network model, and outputting an identification result. The network provided by the invention is more suitable for the shape change of multiple targets with complex and diverse characteristics, and can realize the accurate recognition and early warning of specific interested targets in different traffic scenes, thereby reducing the road safety to the greatest extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image recognition, and particularly relates to a method for multi-object recognition of interest in traffic scenarios based on Transformer-Yolo. Background Art

[0002] With the intelligent and integrated development of the comprehensive transportation system, the rapid advancement of the urbanization process, and the substantial growth of transportation demand, traffic management is facing increasingly complex challenges. And the multi-object recognition task has become one of the core tasks of intelligent transportation, and its application in the fields of traffic management, intelligent monitoring, autonomous driving, etc. is becoming increasingly important. Multi-object detection in traffic scenarios not only includes conventional objects such as vehicles, pedestrians, non-motor vehicles, and traffic roads that appear on the road, but also includes the detection and recognition of specific objects of interest, namely emergency vehicles, traffic accidents, traffic signs, traffic lights, special foreign objects, etc., and giving prompts. The recognition of multi-objects of interest in traffic scenarios plays a crucial role in traffic management, road safety, and emergency response.

[0003] Currently, the recognition of multi-object tasks in traditional traffic scenarios mostly focuses on the recognition and detection of ordinary objects such as vehicles, pedestrians, and roads. There are few methods for the recognition and detection of specific multi-objects of interest, and they face multiple challenges. First is the challenge of diverse objects. The types of target objects in traffic scenarios are numerous, including different types of vehicles, pedestrians, traffic signs, traffic lights, etc., and these objects have different characteristics in different scenarios. Second is the challenge of complex background and environmental interference. The road environment is complex, and there are many background elements, such as buildings, trees, other vehicles, etc. And sometimes, lighting conditions such as rain, snow, night, or strong light irradiation will interfere with the detection results. There is also the challenge of real-time requirements. The need for real-time detection in traffic scenarios is particularly urgent. Whether it is autonomous driving, traffic flow monitoring, or emergency response, the detection system must have high speed and accuracy. There is also the challenge of occlusion and overlap problems. The occlusion and overlap of multi-objects in traffic scenarios are very common, which poses a great challenge to traditional object detection methods. Especially when the object of interest may be partially occluded by other objects, the accuracy of recognition will be greatly affected. To address these multiple challenges, it is required that the method for detecting and recognizing multi-objects of interest in traffic scenarios can not only recognize conventional objects, but also accurately detect specific objects of interest and adapt to detection tasks in different scenarios. To achieve this goal, it is necessary to seek a detection and recognition method with high accuracy, high real-time performance, and robustness.

[0004] Obviously, traditional methods cannot meet this requirement. In contrast, deep learning technology shines in solving the task of multi-object image processing, with high processing performance and strong operability. Currently, the object detection algorithms of deep learning are divided into two-stage algorithms and single-stage algorithms. The two-stage algorithms have high accuracy but slightly slower speed; the single-stage algorithms have poor recognition effects but faster speed, mainly including SSD, YOLO series algorithms. Among them, YOLO has been widely applied to actual engineering construction due to its powerful and efficient detection performance, and has become the mainstream recognition algorithm in traffic scenarios. However, directly applying it to the task of identifying multiple objects of interest in traffic scenarios has problems such as too many object types and complex backgrounds, resulting in insufficient feature extraction ability, inability to extract specific objects of interest more accurately and quickly, and higher requirements for real-time performance in traffic safety guarantee and emergency handling. And the current YOLO algorithm cannot meet the requirements of this task of identifying multiple objects of interest in traffic scenarios.

[0005] Therefore, in order to improve the performance of detecting multiple objects of interest in complex traffic backgrounds, the present invention proposes a method for identifying multiple objects of interest in traffic scenarios based on Transformer-Yolo. Summary of the Invention

[0006] The purpose of the present invention is to provide a method for identifying multiple objects of interest in traffic scenarios based on Transformer-Yolo, which can solve the problems of accurately identifying and warning specific objects of interest in different traffic scenarios, minimizing road safety problems, and solving the problems in the prior art of poor ability to extract features of multiple object types in complex and changeable traffic scenarios, slightly poor real-time performance of accurate judgment, and lack of ability to adapt to complex scenarios.

[0007] In order to achieve the above purpose, the technical solutions adopted by the present invention are as follows:

[0008] A method for detecting and identifying multiple objects of interest in traffic scenarios based on Transformer-Yolo specifically includes the following steps:

[0009] Step 1, collect traffic scene image data, and establish a traffic scene image set containing multiple traffic object categories, where the traffic object categories include ordinary traffic objects and specific traffic objects that may be of interest;

[0010] Step 2, screen the obtained traffic scene data image set, classify and label the ordinary traffic objects and specific traffic objects that may be of interest in the screened images to obtain a data set for identifying multiple objects of interest in traffic scenarios, and divide it into a training set, a validation set, and a test set;

[0011] Step 3: Establish a backbone network model based on Transformer-Yolo and add an improved DfDC module based on Cf2;

[0012] Step 4: Establish a neck network model based on Transformer-Yolo and add an improved attention mechanism VSST module based on Swin Transformer;

[0013] Step 5: Establish a detection network model based on Transformer-Yolo and use Focal SIoU as the loss function;

[0014] Step 6: Read the multi-object recognition dataset of interest in the traffic scene obtained in Step 2, and use transfer learning to comprehensively train the network model to obtain a trained network model;

[0015] Step 7: Use the weights obtained by training with transfer learning to detect and recognize the test set to complete the performance evaluation of the network model;

[0016] Step 8: Input the traffic scene image to be recognized into the trained network model, detect and recognize the multi-objects of interest in the image through the model, and output the recognition results, including the target category, position bounding box, and confidence level.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0018] The multi-object recognition method of interest in traffic scenes based on Transformer-Yolo of the present invention can extract key information of the objects of interest from complex traffic scenes and effectively distinguish them from other background elements. And it has robustness, can meet the target detection requirements under different weather conditions, light changes, and occlusion situations, maintain a high recognition accuracy, and has high real-time performance. It can meet the safety guarantee of traffic roads and the handling of emergencies to the greatest extent, and can be widely applied to corresponding autonomous driving systems, traffic abnormal congestion detection and warning systems, traffic flow monitoring and warning systems, etc.; The present invention proposes a CfCD module formed by combining deformable convolution DCNv2 and Cf2, which is used to adapt to the shape changes of multi-objects with complex and diverse features, and enhance its feature extraction ability and feature reuse expression ability. The present invention improves the Swin Transormer module, that is, the joint design VSST module that integrates the advantages of SwinTransformer and VSS, which is used to combine local features and global context features, dynamically adjust the weights of features, make the model pay more attention to the regions of interest, capture the multi-object performance at different scales, and enhance the robustness of the model. Description of the Drawings

[0019] Figure 1 It is a schematic flowchart of the method for multi-object detection and recognition of traffic scenarios of interest based on Transformer-Yolo of the present invention;

[0020] Figure 2 It is a schematic diagram of the network structure of the method for multi-object detection and recognition of traffic scenarios of interest based on Transformer-Yolo of the present invention;

[0021] Figure 3 It is a schematic diagram of the DfDC network structure;

[0022] Figure 4 It is a schematic diagram of the network structure of the improved Swin Transormer module: VSST;

[0023] Figure 5 It is a recognition and detection diagram of some multi-objects of interest in the traffic scenario by the method for multi-object detection and recognition of traffic scenarios of interest based on Transformer-Yolo of the present invention. Specific embodiments

[0024] The following combines the accompanying drawings and specific embodiments to elaborate in detail on the method for multi-object detection and recognition of traffic scenarios of interest based on Transformer-Yolo of the present invention.

[0025] As Figure 1 shown, the method for multi-object detection and recognition of traffic scenarios of interest based on Transformer-Yolo is specifically executed according to the following steps:

[0026] Step 1: Collect traffic scenario image data, and establish a traffic scenario image set containing multiple traffic target categories. The traffic target categories include common traffic targets: pedestrians, vehicles, and specific traffic targets that may be of interest: emergency vehicles, illegally parked vehicles, traffic accidents, traffic signs, and special foreign objects, etc.;

[0027] Step 2: Screen the obtained traffic scenario data image set, classify and label the common traffic targets and specific traffic targets that may be of interest in the screened images to obtain a multi-object recognition data set of interest in the traffic scenario, and divide it into a training set, a validation set, and a test set;

[0028] Step 3: Establish a backbone network model based on Transformer-Yolo, and add an improved DfDC module based on Cf2 to adapt to the shape changes of multi-objects of interest and enhance its feature extraction ability;

[0029] Step 4: Establish a neck network model based on Transformer-Yolo, and add an improved attention mechanism VSST module based on Swin Transformer to enhance the positioning ability of the information of the position of interest and capture the performance of the target at different scales, thereby enhancing the robustness of the model.

[0030] Step 5: Establish a detection network model based on Transformer-Yolo, and use the Focal SIoU as the loss function to reduce the model calculation cost and improve the overall optimization performance of the model.

[0031] Step 6: Read the multi-object recognition dataset of interest in the traffic scene obtained in Step 2, and use transfer learning to comprehensively train the network model to obtain a trained network model.

[0032] Step 7: Use the weights obtained by training with transfer learning to detect and recognize the test set, and complete the performance evaluation of the network model.

[0033] Step 8: Input the traffic scene image to be recognized into the trained network model, detect and recognize the multi-objects of interest in the image through the model, and output the recognition results, including the target category, position bounding box, and confidence level.

[0034] The method for detecting and recognizing multi-objects of interest in a traffic scene based on Transformer-Yolo of the present invention enables the network to adapt to the shape changes of multi-objects with complex and diverse features, enhance its feature extraction ability and feature reuse and expression ability, and combine local features and global context features to dynamically adjust the weights of features, making the model pay more attention to the region of interest, capture the performance of multi-objects at different scales, and enhance the robustness of the model. In practical applications, due to the diversity and complexity of traffic scenes, it is necessary to collect images from multiple data sources, such as in-vehicle surveillance cameras, drones, and mobile phones, and then upload the captured images to a computer device for detecting and recognizing multi-objects of interest, so as to solve the problem of accurately recognizing and warning specific objects of interest in different traffic scenes, and minimize road safety problems and the handling of emergencies.

[0035] The present invention will be further described below in conjunction with specific embodiments.

[0036] Embodiment 1;

[0037] The method for detecting and recognizing multi-objects of interest in a traffic scene based on Transformer-Yolo in this embodiment includes the following steps:

[0038] In Step 1, due to the diversity and complexity of traffic scenarios, it is necessary to take pictures from multiple data sources in three different types of directions using in-vehicle monitoring cameras, drones, and mobile phones to improve the diversity and representativeness of the data.

[0039] Step 2 is specifically as follows:

[0040] Step 2.1: Import the images of different traffic scenarios taken by in-vehicle monitoring cameras, drones, and mobile phones into the computer and conduct manual screening and annotation, which are divided into ordinary traffic targets and specific traffic targets of interest. Among them, ordinary traffic targets include conventional targets such as vehicles, pedestrians, non-motor vehicles, and traffic roads, and specific traffic targets that may be of interest include special foreign objects such as emergency vehicles, illegally parked vehicles, traffic accidents, traffic signs, and traffic lights.

[0041] In Step 2.1, when screening traffic scenario images, first remove blurred, noisy, or poorly exposed images through image clarity detection to ensure the quality of the input data; then screen out images without obvious targets or with severely occluded targets; next, according to the requirements of scenario diversity, retain images covering various scenarios such as urban roads, highways, intersections, etc., and include data of day, night, different weather conditions, and multiple acquisition devices; at the same time, ensure that there are sufficient samples for each type of traffic target to avoid class imbalance; finally, remove redundant images that are highly similar and only retain representative samples, thereby enhancing the representativeness and diversity of the data and providing a high-quality data basis for subsequent annotation and training.

[0042] Step 2.2: Annotate the 2,570 selected images according to different target categories. Use the LabelImg annotation tool and refer to the CoCo dataset to annotate the images. The information annotated for each image is saved in the txt file format, mainly including: the height and width information of the overall image, and the specific positions of the four vertices of the target annotation rectangle.

[0043] Step 2.3: Divide the annotated images, and the specific ratio of the training set: validation set: test set is 8:1:1.

[0044] Step 3 is specifically as follows:

[0045] Step 3.1, Combine the deformable convolution DCNv2 and Cf2 to form the CfDC module: Replace the Bottleneck block in Cf2 with the deformable convolution DCNv2. The outer structure of the CfDC module is similar to that of Cf2 and consists of two Conv blocks, n deformable convolution DCNv2s, a Chunk block, and a Concat block. The deformable convolution DCNv2 consists of multiple deformable convolution layers, multiple standard convolution layers, and an offset prediction layer connected in sequence. Among them, the deformable convolution layer: Dynamically adjusts the position of the convolution kernel by introducing an offset to adapt to the shape of the input feature map; the standard convolution layer: Includes standard convolution operations to extract important features; the offset prediction layer: Calculates the offset of the convolution kernel to ensure the flexibility of the deformable convolution.

[0046] In this embodiment, construct the CfDC module: The size of the feature map output by the previous convolution module of the CfDC module is 1024×20×20. As Figure 3 shown, after the feature map enters the CfDC module, it first enters a 1×1 Conv block, and the size dimension of the output feature map remains unchanged, which is 1024×20×20. Then it passes through a Chunk block, which divides the feature map into two parts, and both the first part and the second part are 512×20×20. After receiving the 512×20×20 feature input, the deformable convolution DCNv2 outputs a feature map with a feature dimension of 512×20×20 after passing through 3 1×1 deformable convolution layers, 3 3×3 standard convolution layers, and an offset prediction layer, and then enters the Concat module, that is, the splicing module, which receives the two 512×20×20 feature outputs from the Chunk operation and the n×512×20×20 feature outputs of n deformable convolution DCNv2s. The spliced feature dimension is (2 + n)×512×20×20. Finally, after passing through a 1×1 Conv block, a feature map with an output dimension of 1024×20×20 is output.

[0047] Step 3.2, Replace the last Cf2 in the backbone network, that is, the layer above SPPF, with the proposed CfDC module to adapt to the shape changes of multi-targets with complex and diverse features, and enhance its feature extraction ability and feature reuse expression ability.

[0048] Step 4 is specifically as follows:

[0049] Step 4.1, Construct an improved attention mechanism VSST module based on the Swin transformer: Use the VSSblock as a long-distance low-level feature position extractor to extract long-distance low-level features, and use the SwinTransformer above the output for high-level close-range feature position extraction and task processing, as Figure 4As shown, a VSS block module is added to the upper layer of the Swin TransormerBlock in each of the four stages of the conventional Swin transformer module, so that the composition of each stage is a Patch expanding layer, a VSS block module, and several Swin Transormer Block modules connected in sequence.

[0050] In this embodiment, the VSST module is used to replace the 2 Cf2 modules for downsampling in the Transformer-Yolo neck network. Taking the replacement of the first downsampling in the neck network as an example, specifically: the input feature map of the improved attention mechanism VSST module has a dimension of 512×40×40. The feature map will be sent to the Patch partition layer and decomposed into 25×8×8×512, that is, there are 25 small blocks of 8×8×512. Then, through the Patch expanding layer, the block features are recombined into a complete feature map, and the dimension is restored to 512×40×40. Subsequently, the feature map enters the VSS Block module. The outputs of 1 layer normalization layer LN (Layer Normalization) enter two branches respectively. The first branch includes a linear layer (Linera), a depthwise separable convolution (DWConv) with a dimension of 1×1, an SS2D block, and 1 layer normalization layer LN connected in sequence to capture the multi-object performance features at different scales. The second branch includes 1 layer normalization layer LN. After the outputs of the two branches are multiplied, they enter 1 linear layer (Linera), and are added to the input of the VSS Block module to obtain the output of the VSS Block module. The number of channels of the input and output feature maps remains unchanged, still 512×40×40.

[0051] Next, the Swin Transformer Block module receives the output from the VSS Block. The Swin Transformer Block is the core module in Swin Transformer. Its structure includes a normalization layer (LayerNorm, LN), a shifted window multi-head self-attention (Shifted Window Multi-Head Self-Attention, SW-MSA), a windows multi-head self-attention (Windows Multi-Head Self-Attention, W-MSA), a feed-forward network (MLP), etc. This module first normalizes the input feature map through the LN layer to ensure stable distributions for each channel. Then, the normalized feature map is divided into multiple non-overlapping windows, each with a size of 7×7. Within these windows, multi-head self-attention (Multi-Head Self-Attention, MSA) is calculated independently. By generating Query, Key, and Value, the feature correlations within the windows are captured. Subsequently, to enhance cross-window feature interaction, a sliding window mechanism is adopted. The window positions are slid relative to each other by a certain step size, and self-attention calculation is performed again, thus realizing the modeling of local and global features. After passing through the SW-MSA module, the feature map retains the original input information through residual connections and is passed to the MLP. The MLP consists of two fully connected layers. The first layer extracts non-linear features through the activation function GELU, and the second layer is used to reshape the features. Finally, the output of the MLP is passed back through residual connections again, and the dimension of the feature map remains unchanged, providing richer context information for the next module. The feature vectors in this module will interact with the feature vectors of the two upper blocks of other blocks to construct a global feature representation based on the attention mechanism. This process usually generates a 40×40 attention matrix of the same size as the input feature map, and the final input dimension is 512×40×40. The subsequent three stages are as above. After feature extraction in four stages, the dimension of the final output feature map is 512×40×40.

[0052] Step 4.2: Replace the 2 Cf2 modules for downsampling in the neck network with the improved attention mechanism VSST module, which is used to combine local features and global context features, dynamically adjust the weights of features, enable the model to pay more attention to the regions of interest, capture the multi-object performance at different scales, and enhance the robustness of the model.

[0053] Step 5 is specifically as follows:

[0054] Step 5.1: Use Focal SIoU (Smooth Intersection over Union with Focal Loss) to replace the original regression loss function CIoU. This is a loss function that combines IoU and Focal Loss to enhance the performance of the object detection model in dealing with class imbalance and bounding box regression. This loss function better solves the problems of hard samples and uncertain samples in object detection through IoU calculation and weight adjustment of Focal Loss, making the model easier to converge and providing advantages for model training.

[0055] By integrating the concept of Focal Loss into the bounding box regression loss, Focal SIoU combines IoU and Focal Loss to optimize the overlap of bounding boxes, while paying more attention to difficult-to-locate objects and uncertain predictions. The calculation formula of Focal SIoU is:

[0056] L Focal SIoU =(1 - SIoU) γ + log(SIoU)(1)

[0057] where SIoU is the calculation result of IoU, and γ is the adjustment factor in the focal loss, which is used to adjust the contribution of samples with different overlap degrees to the loss.

[0058] γ = 2: This is the default value commonly used in Focal Loss for many tasks. It can significantly reduce the weight of easy-to-classify samples and encourage the model to pay more attention to difficult-to-classify samples.

[0059] γ = 1: It is equivalent to directly weighting on the basis of IoU, and the loss function becomes relatively balanced, with less weight difference for all samples.

[0060] γ = 0: In this case, Focal Loss is equivalent to ordinary cross-entropy loss or IoU loss, completely ignoring weight adjustment.

[0061] In practice, the specific value of γ needs to be adjusted through experiments according to the difficulty of the task and data distribution. In this method, a relatively large γ = 3 is tried to more significantly focus on multi-object samples of interest.

[0062] Step 6 is specifically as follows:

[0063] The training of the multi-object detection and recognition model of interest in the traffic scenario is carried out by using the method of transfer learning. The model training is divided into a rough training stage and a fine-tuning training stage. In the rough training stage, all convolutional layer weights are frozen, and only the fully connected layers are trained. The number of iterations is set to 50, the batch size is 16, and the initial learning rate is 0.001. After completing the rough training stage, enter the fine-tuning training stage. In this stage, all convolutional layers are unfrozen, so that the model can train the entire network. The number of iterations is set to 150, the batch size is 8, and the initial learning rate is set to 0.0001. And in order to prevent the model from overfitting, the early stopping method is used for training, that is, when the error on the validation set deteriorates, the training is stopped, and the parameters of the previous iteration are used as the final parameters of the model, and the momentum parameter is set to 0.973, and the optimization algorithm uses the stochastic gradient descent algorithm, that is, SGD.

[0064] Step 7 is specifically: use the weights obtained by training with transfer learning to detect and recognize the test set, that is, use the weights generated by the last training to detect the test set. And select the accuracy P (Precision), recall R (Recall), average accuracy AP (Average Precision), mean average precision mAP (mean Average Precision), frames per second FPS (Frame Per Second) and inference time (inference) as the evaluation indicators of this experiment to complete the performance evaluation of the algorithm. Accuracy P (Precision): It is used to evaluate the ability of the model to recognize positive samples. The higher the accuracy, the more positive examples are detected. Its calculation formula is shown in Equation (2):

[0065]

[0066] Among them, TP is the true positive example (the number of samples where the prediction and the actual are both positive examples), and FP is the false positive example (the number of samples where the prediction is a positive example but the actual is a negative example). TP + FP represents the total number of samples predicted as positive examples.

[0067] Recall: It measures the proportion of positive samples in the test set that are correctly recognized. The higher the recall, the more positive samples are correctly recognized. Its calculation formula is shown in Equation (3):

[0068]

[0069] Among them, FN is the false negative example (the number of samples where the prediction is a negative example but the actual is a positive example). TP + FN represents the total number of actual positive samples.

[0070] Average accuracy: By integrating the accuracy at different recall rates, the obtained value reflects the overall performance of the object detection algorithm. Its calculation formula is shown in Equation (4):

[0071]

[0072] Among them, P is precision, r is recall, and the value range of AP is from 0 to 1.

[0073] Mean Average Precision (mAP): It averages the AP values of multiple classes and represents the performance of the model in the object detection task. Its calculation formula is shown in Equation (5):

[0074]

[0075] Among them, N is the number of classes, and AP is the average accuracy of each class. The higher the mAP value, the better the model performance.

[0076] Frames Per Second (FPS): In object detection, it represents the number of images that can complete object detection per second and is used to evaluate the speed of the object detection algorithm. Its calculation formula is shown in Equation (6):

[0077]

[0078] Inference time: It refers to the time required for the model to process a single input sample, usually in milliseconds (ms). The inference time directly affects the performance of the model in real-time applications. A shorter inference time means a faster response speed of the model.

Claims

1. A multi-object detection and recognition method for traffic scenarios of interest based on Transformer-Yolo, characterized in that, Specifically, it includes the following steps: Step 1: Collect traffic scene image data, and establish a traffic scene image set containing multiple traffic target categories, where the traffic target categories include ordinary traffic targets and specific traffic targets that may be of interest; Step 2: Screen the obtained traffic scene data image set, classify and label the ordinary traffic targets and specific traffic targets that may be of interest in the screened images to obtain a multi-target recognition data set of interest in the traffic scene, and divide it into a training set, a validation set and a test set; Step 3: Establish a backbone network model based on Transformer-Yolo, and add an improved DfDC module based on Cf2; Step 4: Establish a neck network model based on Transformer-Yolo, and add an improved attention mechanism VSST module based on Swin Transformer; Step 5: Establish a detection network model based on Transformer-Yolo, and use Focal SIoU as the loss function; Step 6: Read the multi-target recognition data set of interest in the traffic scene obtained in Step 2, and use transfer learning to comprehensively train the network model to obtain a trained network model; Step 7: Use the weights obtained by training with transfer learning to detect and recognize the test set, and complete the performance evaluation of the network model; Step 8: Input the traffic scene image to be recognized into the trained network model, detect and recognize the multi-targets of interest in the image through the model, and output the recognition results, including target category, position bounding box, and confidence; 2. The multi-object detection and recognition method for traffic scene of interest based on Transformer-Yolo according to claim 1, characterized in that, In Step 1, the traffic scene image data is collected by using in-vehicle monitoring cameras, drones and mobile phones; 3. The multi-object detection and recognition method for traffic scenes of interest based on Transformer-Yolo according to claim 1, characterized in that, Step 2 includes the following sub-steps: Step 2.1: Screen the images in the traffic scene image set obtained in Step 1; Step 2.2: Classify and label the ordinary traffic targets and specific traffic targets that may be of interest in the screened images, and the labeled information includes: image height and width, and the positions of the four vertices of the target annotation rectangle; Step 2.3: Divide the labeled data images according to a certain ratio, and the division ratio is training set: validation set: test set = 8:1:1; 4. The multi-object detection and recognition method for traffic scene of interest based on Transformer-Yolo according to claim 1, characterized in that, Step 3 includes the following sub-steps: Step 3.1, construct the CfDC module: The size of the feature map output by the upper convolutional module of the CfDC module is 1024×20×20. After the feature map enters the CfDC module, it first enters a 1×1 Conv block, and the size dimension of the output feature map remains unchanged, which is 1024×20×20. Then it passes through a Chunk block, and this module divides the feature map into two parts, both the first part and the second part are 512×20×20; after the deformable convolution DCNv2 receives the 512×20×20 feature input, after 3 1×1 deformable convolution layers, 3 3×3 standard convolution layers and the offset prediction layer, it outputs a feature map with a feature dimension of still 512×20×20, and then enters the Concat module, that is, the splicing module, which receives the two 512×20×20 feature outputs from the Chunk operation and the n×512×20×20 feature outputs of n deformable convolutions DCNv2; finally, through a 1×1 Conv block, it outputs a feature map with a dimension of 1024×20×20; Step 3.2, replace the last Cf2 in the backbone network with the proposed CfDC module.

5. The method for multi-object detection and recognition of traffic scene of interest based on Transformer-Yolo according to claim 1, characterized in that, Step 4 includes the following sub-steps: Step 4.1, construct the improved attention mechanism VSST module based on the Swin transformer. The input feature map dimension of the improved attention mechanism VSST module is 512×40×40. The feature map will be sent to the Patch partition layer and decomposed into 25×8×8×512, that is, there are 25 small blocks of 8×8×512. Then it passes through the Patch expanding layer, and the block features are recombined into a complete feature map, and the dimension is restored to 512×40×40. Subsequently, the feature map enters the VSS Block module. The output of 1 normalization layer LN enters two branches respectively. The first branch includes a linear layer, a depthwise separable convolution with a dimension of 1×1, an SS2D block, and 1 normalization layer LN connected in sequence to capture the multi-object performance features at different scales. The second branch includes 1 normalization layer LN. After the outputs of the two branches are multiplied, they enter 1 linear layer and are added to the input of the VSS Block module to obtain the output of the VSS Block module. The number of channels of the input and output feature maps remains unchanged, still 512×40×40; Step 4.2, replace the 2 Cf2 modules for downsampling in the neck network with the improved attention mechanism VSST module.

6. The multi-object detection and recognition method for traffic scene of interest based on Transformer-Yolo according to claim 1, characterized in that, Step 5 includes the following sub-steps: Step 5.1, use Focal SIoU to replace the original regression loss function CIoU, and the calculation formula: L FocalSIoU = (1 - SIoU) γ + log(SIoU) where SIoU is the calculation result of IoU, and γ is the adjustment factor in the focal loss.

7. The method for multi-object detection and recognition of traffic scene of interest based on Transformer-Yolo according to claim 1, characterized in that, Step 6 is specifically: The training of the multi-object detection and recognition model of interest in traffic scenarios is carried out by using the method of transfer learning. The model training is divided into a rough training stage and a fine-tuning training stage. In the rough training stage, all convolutional layer weights are frozen, and only the fully connected layer is trained. The number of iterations is set to 50, the batch size is 16, and the initial learning rate is 0.

001. After completing the rough training stage, enter the fine-tuning training stage. In this stage, all convolutional layers are unfrozen, enabling the model to train the entire network. The number of iterations is set to 150, the batch size is 8, and the initial learning rate is set to 0.0001.

Citation Information

Patent Citations

  • Single-target tracking network based on combination of Vmamba and Transform

    CN118537369A

  • Deep learning-based depression angle sensitive small target remote identification method

    CN119131338A

  • Data assimilation method and device, equipment and medium

    CN119513515A

  • Real-time monitoring system for personal protective equipment compliance at worksites

    US12243317B1