An intelligent traffic-oriented domain adaptive target detection network optimization method
Patent Information
- Application Number
- CN202310164798.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-26
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2043-02-26
AI Technical Summary
本发明能很大程度解决目标场景变化引起的深度学习模型性能大幅下降的问题,不需要大量目标部署场景的标注数据,且设计的优化网络仅在训练阶段起作用,不影响检测网络推理阶段处理速度
[0019] This invention, starting from practical application needs, focuses on multiple feature dimensions, including global, inter-channel, and single-channel features, and designs a domain-adaptive target detection network optimization method for intelligent transportation. First, this method can fully utilize public datasets and unlabeled target scenes to optimize the detection network, solving the problem of significant performance degradation across domains and avoiding costly manual annotation. Second, this invention is versatile and can be applied to various traffic scenarios. Finally, this method is designed for the efficient YOLOX detector, which features high efficiency and performance, is easy to deploy on various platforms, and has high practical value.
Smart Images

Figure CN117079064B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of digital image / computer vision, and specifically relates to a domain adaptive target detection network optimization method for intelligent transportation. Background Technology
[0002] Object detection is a crucial component of visual image understanding. Real-time detection and recognition of obstacles in vehicle operating areas is a key focus and challenge for autonomous driving and assisted traffic control. Currently, autonomous driving primarily relies on sensor arrays consisting of LiDAR and industrial cameras to collect environmental data. Image object detection algorithms and semantic segmentation algorithms process this data to provide results for driver assistance. Therefore, object detection algorithms are a vital part of autonomous driving. In recent years, thanks to significant advancements in artificial intelligence, technologies related to autonomous driving for intelligent transportation have seen substantial development and application, leading to a marked improvement in the operational safety of transportation systems. In the field of vehicle autonomous driving, high-performance object detection algorithms are a core technology. Accurate identification of targets in vehicle operating areas using image object detection algorithms is fundamental for subsequent early warning decisions. Therefore, the efficient deployment and application of image object detection algorithms are of great significance for the practical implementation of autonomous driving technology.
[0003] Taking rail transit as an example, training high-performance object detection models often requires a large amount of labeled data from the target application scenario. However, due to operational safety concerns, it is impossible to manually place targets in tunnels, making data acquisition difficult in rail transit scenarios with obstacles. Furthermore, re-labeling the data incurs high manual costs. Road scenarios possess a large amount of labeled data that can be used to train object detection models, but due to differences in data distribution compared to rail transit scenarios, directly applying models trained in road scenarios to rail transit scenarios often results in a significant performance drop. Given the limited training data in the rail transit domain, which makes it difficult to support iterative training of high-performance object detection networks, methods for training high-performance object detection models without relying on large amounts of labeled data have significant research and application value.
[0004] Domain-adaptive object detection is a hot research topic in transfer learning, aiming to improve model performance on test datasets by mining general knowledge from training datasets. In this field, the data space of the training dataset is the source domain, and the data space of the test dataset is the target domain. Typically, the source domain possesses rich supervised learning information, while the target domain has only a small amount of labeled information or no labels at all. Applying domain-adaptive object detection network optimization algorithms can address the performance degradation caused by cross-domain detection. Some researchers have proposed using generative adversarial networks (GANs) to perform style transfer from the source domain to the target domain, thereby expanding the number of labeled samples in the target domain and solving the problem of insufficient labeled data for training deep neural networks. However, because the simulated images generated by this method lack real-world physical constraints, even if visually similar data to the target domain is generated, its data distribution differs from real images. Furthermore, this method cannot achieve end-to-end detector optimization, and the detection accuracy depends on the results of style transfer and intermediate fine-tuning. Other researchers have proposed rule-based cross-domain object detection algorithms, which mainly rely on manually established rules to constrain deep neural networks to learn common features across different domains, thereby achieving feature alignment. Subspace alignment methods are used to align features between the source and target domains, effectively projecting the source domain subspace onto the target domain subspace, thus achieving domain-adaptive object detection. While this approach mitigates data distribution differences between different domains to some extent, its real-time performance is poor due to its design for two-stage detectors, failing to meet the application requirements of practical systems. Therefore, designing a domain-adaptive object detection network optimization method for single-stage object detectors remains a pressing issue for achieving more efficient intelligent transportation systems. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of existing technologies and design a domain-adaptive target detection network optimization method for intelligent transportation. The overall process is shown in the attached figure. Figure 1 As shown, the invention mainly comprises three stages: dataset establishment, domain-adaptive object detection network training, and object detection network inference. The technologies involved primarily include adversarial domain adaptation, decoupled representation learning, channel attention, spatial attention, and instance regularization. This invention can largely solve the problem of significant performance degradation in deep learning models caused by changes in the target scene. It does not require a large amount of labeled data for the target deployment scene, and the optimized network only functions during the training phase, without affecting the processing speed of the detection network inference phase. A brief description of each part follows.
[0006] Dataset establishment: Public datasets are used as source domain data, and train operation data from train vehicle operation scenarios are collected as target domain data.
[0007] Domain Adaptive Object Detection Network: The network includes a YOLOX object detector, a multi-scale global feature alignment module, a domain-invariant feature decoupling module, and a multi-scale instance-level adaptive module. By training the network, the domain adaptive optimization of the YOLOX object detector is achieved, enabling it to accurately identify objects in images within a target scene.
[0008] Object detection network inference: After completing the domain adaptive optimization of the object detector, the auxiliary branches used for cross-domain optimization are removed, and a cross-domain detection model that can be efficiently deployed is derived.
[0009] The specific steps for creating the dataset are as follows:
[0010] The dataset creation process designed in this invention is as follows: Figure 2 As shown. For the collected train vehicle operation data, keyframe data is extracted from the video at specified frame intervals to obtain a coarsely filtered dataset. Then, blurry, duplicate, and unusable data are removed to obtain the target domain dataset. A large number of labeled public datasets are collected as the source domain dataset.
[0011] The design of a domain-adaptive object detection network follows these steps:
[0012] The domain-adaptive target detection network designed in this invention is shown in the appendix. Figure 3 As shown, it mainly consists of a YOLOX target detector (including a preprocessing module, a backbone network, and a detection head), a multi-scale global feature alignment module, a domain-invariant feature decoupling module, and a multi-scale instance-level adaptive module. The adversarial domain discriminator serves as the basic module for optimizing the weights of the domain adaptive network and is reused in multiple modules.
[0013] The structure of the adversarial domain discriminator is shown in the attached figure. Figure 4 The diagram shows a gradient reversal layer (GRL) and a domain discriminator. The domain discriminator contains three cascaded "Conv+GroupNorm+ReLU" operation blocks and an output convolutional layer (where GroupNorm is a regularization operation). The three operation blocks of the domain discriminator all have 1×1 kernels that do not change the number of channels in the features. The output convolutional layer reduces the number of channels in the features to 1 and outputs the domain prediction result for the features.
[0014] Multi-scale global feature alignment module. The structure diagram of the multi-scale global feature alignment module is attached. Figure 5As shown. Unlike most domain-adaptive object detection methods that only adapt the domain to the final scale of the feature extractor, this invention performs domain adaptation on all three scales of the feature extractor to address the vanishing gradient problem. The backbone network of the YOLOX object detector has outputs at three scales (feature sizes of 80×80×128, 40×40×256, and 20×20×512, respectively; in this invention, the output features at the three scales are named C128, C256, and C512 according to the number of channels). Due to the vanishing gradient problem, if only the last scale (C128) is adapted to the domain, the optimization effect on the preceding scales (C256 and C512) will be quite small. Therefore, this invention adopts a multi-scale strategy, connecting the three feature layers C128, C256, and C512 of the backbone network to a multi-scale global feature alignment module through three GRLs. As training progresses, the backbone network of the object detector will tend to extract domain-invariant features from the image data.
[0015] Domain-invariant feature decoupling module. The structure diagram of the domain-invariant feature decoupling module is attached. Figure 6-1 As shown. Existing end-to-end learning methods, represented by supervised deep networks, are a shortcut learning strategy. The abstract features learned by the network through iterative training often fail to generalize in the desired direction. To alleviate this problem, this invention employs a decoupled representation learning method to decouple the output features of the target detector's backbone network into domain-invariant features and domain-specific features. These domain-invariant and domain-specific features are independent of each other. As training progresses, the backbone network of the target detector should extract more domain-invariant features to achieve better detection performance on both source and target domain data. Specifically, this invention designs a feature encoder that divides the backbone network's output features into domain-invariant and domain-specific features along the feature channel dimension, sending only the domain-invariant features to the next stage of operation. The structure of the feature encoder is shown in the attached figure. Figure 6-2 As shown, it consists of three cascaded "Conv+IN+ReLU" operation blocks (where IN is the InstanceNorm regularization operation), "Conv+Sigmoid" operation blocks, and a global average pooling layer (GAP), with the input and output features having the same size.
[0016] Multi-scale instance-level adaptive module. The structure diagram of the multi-scale instance-level adaptive module is attached. Figure 7-1As shown. Object detection typically includes object localization and classification, and accurate object localization greatly aids in accurate object recognition. However, domain classifiers align the entire feature map without considering the distinction between foreground and background features, potentially leading to negative transfer. To address this issue, this invention designs a multi-scale instance-level adaptive module based on an attention mechanism to achieve data distribution alignment that is aware of foreground and background information. Unlike methods for two-stage object detectors that can directly extract foreground information from the RPN network output, one-stage networks lack this intermediate process. Therefore, this invention designs an attention map generator that uses the output information of the detection head to divide the features into foreground and background parts. The structure diagram of the attention map generator is attached. Figure 7-2 As shown, the input is the category prediction information output by the detection head, where H×W and C represent the size and number of channels of the feature map, respectively. After averaging by channel dimension, an intermediate feature of size H×W×1 is obtained. To further filter out regions that are more likely to be background, this invention takes the average of the intermediate features as a threshold, sets features with response values below the threshold to zero, and retains the original values of features with response values not lower than the threshold. After the above operations, an attention map can be obtained to achieve data distribution alignment that is aware of foreground and background information.
[0017] The inference process for the object detection network is as follows:
[0018] The object detection network inference process used in this invention is shown in the appendix. Figure 8 As shown. After completing the domain adaptive optimization of the target detector, it is not necessary to retain the domain adaptive branch during the inference stage. After deleting the relevant branches, the cross-domain detector model is exported. Real-time acquired image data is sent to the preprocessing module for scaling and data normalization. The backbone network completes the extraction of deep features, and the detection head further infers the deep features of the backbone network, outputting the target location, category, and prediction confidence information in the image. The output results can then be used to implement secondary tasks.
[0019] This invention, starting from practical application needs, focuses on multiple feature dimensions, including global, inter-channel, and single-channel features, and designs a domain-adaptive target detection network optimization method for intelligent transportation. First, this method can fully utilize public datasets and unlabeled target scenes to optimize the detection network, solving the problem of significant performance degradation across domains and avoiding costly manual annotation. Second, this invention is versatile and can be applied to various traffic scenarios. Finally, this method is designed for the efficient YOLOX detector, which features high efficiency and performance, is easy to deploy on various platforms, and has high practical value. Attached Figure Description
[0020] Figure 1 Overall Flowchart
[0021] Figure 2Dataset creation
[0022] Figure 3 Domain Adaptive Object Detection Network Structure
[0023] Figure 4 Adversarial domain discriminator module
[0024] Figure 5 Multi-scale global feature alignment module
[0025] Figure 6-1 Domain-invariant feature decoupling module
[0026] Figure 6-2 Feature encoder
[0027] Figure 7-1 Multi-scale instance-level adaptive module
[0028] Figure 7-2 Attention map generator
[0029] Figure 8 Object detection network inference Detailed implementation method:
[0030] The following detailed description of embodiments of the present invention is provided in conjunction with the accompanying drawings.
[0031] An optimization method for domain-adaptive object detection networks in intelligent transportation is presented in the attached flowchart. Figure 1 As shown in the attached diagram, the main stages include dataset creation, domain-adaptive object detection network training, and object detection network inference. The dataset creation process is illustrated in the attached diagram. Figure 2 As shown in the attached figure, the structure of the domain adaptive target detection network is as follows. Figure 3 As shown in the attached diagram, the inference process of the object detection network is as follows: Figure 8 As shown. The specific process of this invention is as follows: First, a dataset is created, including a large amount of public labeled source domain data and self-collected unlabeled target domain data of traffic scenarios; then, the domain adaptive object detection network optimization method designed in this invention is applied to optimize and train the detector; finally, the detector model is exported for model inference.
[0032] The specific steps for the dataset creation phase are as follows:
[0033] The data collection covers all operating sections, collecting train operation data during the operating hours. The obtained train operation video data, with a resolution of 1600×1200, is processed in multiple steps. First, a fixed frame extraction interval is set, and image data is extracted in 5-frame increments according to the video stream playback order, resulting in a coarsely extracted image dataset. Second, the coarsely extracted data is filtered, removing a large amount of duplicate data generated when the train is stationary, as well as some unusable data. After data filtering, data augmentation is performed on the dataset, mainly using methods such as random horizontal flipping, vertical and horizontal translation, color perturbation, and angle rotation to further expand the data volume.
[0034] The specific steps during the training phase of the domain-adaptive object detection network are as follows:
[0035] Domain-adaptive object detection network as shown in the appendix Figure 3 As shown, it consists of a YOLOX object detector (including a preprocessing module, a backbone network, and a detector head), a multi-scale global feature alignment module, a domain-invariant feature decoupling module, and a multi-scale instance-level adaptive module.
[0036] The YOLOX object detector mainly consists of a preprocessing module, a backbone network, and a detection head. The upper limit of the acceptable data resolution for the object detector is 640×640. The preprocessing module is mainly responsible for maintaining the aspect ratio of the image data by scaling the longer side to 640 resolution, and filling in any missing parts with data [128,128,128]. The backbone network is responsible for extracting useful features from the input image data and is the key to performing domain adaptive optimization. The detection head infers whether the image contains an object and information such as the object's location and size from the features output by the backbone network.
[0037] The multi-scale global feature alignment module is attached. Figure 5 As shown in the figure, three sets of adversarial discriminators were designed for the three scale output features of the YOLOX detector backbone network. The structures of the three sets of adversarial discriminators are identical except for the input feature size: 80×80×128, 40×40×256, and 20×20×512. The structure of a single adversarial discriminator is shown in the attached figure. Figure 4As shown, the system consists of a gradient inversion layer (GRL) and a domain discriminator. The GRL performs no operation during forward propagation but inverts the gradient values during backward propagation, optimizing the network in a direction where it becomes difficult to distinguish whether features originate from the source or target domain. The convolutional operation parameters of the domain discriminator are shown in Table 1. This structure is reused multiple times in this invention, with N being a configurable value determined by the dimension of the input data. For example, in the multi-scale global feature alignment module, the input feature dimensions for the three scales are 128, 256, and 512, respectively, so the corresponding N values for the adversarial domain discriminator are set to 128, 256, and 512. Specifically, the input data first passes through three "ConvGNReLU" layers, including convolutional operations with a kernel size of 1, a stride of 1, and zero padding, group regularization, and the ReLU activation function. It then passes through an output convolutional layer with a kernel size of 1, a stride of 1, and zero padding to obtain the domain classification result.
[0038] The loss function of the adversarial domain discriminator is shown in Equation 1.
[0039]
[0040] Where, x s and x t These represent source domain data and target domain data, respectively. and Let E represent the source domain and the target domain, respectively. E is the feature encoder, which is the backbone network of the target detector in this invention. D is the domain discriminator, and D(·) represents the probability that a feature is a feature from the source domain. θ E and θ D These are the neural network model parameters for E and D, respectively.
[0041] The domain-invariant feature decoupling module is attached. Figure 6-1 As shown in the figure. The convolutional layer structures of the two feature encoders are as follows: Figure 6-2 As shown in Tables 2-1 and 2-2, the parameters are identical in structure, differing only in model parameters. Input data first passes through three "ConvINReLU" layers, including convolution operations with a kernel size of 1, a stride of 1, and padding of 1, instance regularization, and the ReLU activation function. Then, it undergoes another convolution operation with a kernel size of 1, a stride of 1, and padding of 1, followed by a Sigmoid activation function to obtain features of the same size as the input. Finally, it passes through a global average pooling (GAP) layer to obtain a 1x1 channel attention vector with the same number of channels as the input. Multiplying this vector by the corresponding scale of the backbone network output features yields the decoupled features. The domain-invariant feature decoupling module, after obtaining the backbone network output features, splits into two parallel branches. One branch extracts domain-specific features F. ds The signal is then fed into a set of adversarial discriminators for cross-domain optimization, while another branch extracts domain-invariant features F. diThe result is summed with the input feature elements and then output to the next level module for further processing. This ensures that the domain-invariant feature F... di Domain-specific features F ds To ensure independence, a mutual information loss is set between the two features, and the two features are made independent by minimizing the mutual information.
[0042] The multi-scale instance-level adaptive module is attached. Figure 7-1 As shown, an attention map generator was designed to extract foreground information from the category information output by the detection head. In the figure, H is the height of the input feature map, W is the width of the input feature map, and C is the number of channels of the input feature map. The structure of the attention map generator is shown in the attached figure. Figure 7-2 As shown in Equation 2, firstly, the input features are averaged according to the channel dimension to obtain an intermediate feature M of size H×W×1. (h,w) To further filter out regions that are more likely to be background, this invention takes the average of the intermediate features as a threshold, as shown in Equation 3, and adds an indicator function as shown in Equation 4. Features with response values below the threshold are set to zero, while features with response values above the threshold retain their original values, resulting in an attention map that indicates background information. Finally, the domain discriminator output loss of the multi-scale instance-level adaptive module is element-wise multiplied with the attention map to obtain a spatially attention-weighted feature map.
[0043]
[0044]
[0045]
[0046] The computation process of the attention map generator can be represented by Equations 2, 3, and 4. Here, M represents the intermediate features, and T... C Thrd is the C-class category prediction information output by the detection head of the object detection network, thrd is the threshold of the indicator function, and A is the attention map output by the attention generator.
[0047] The specific steps of the object detection network inference phase are as follows:
[0048] After completing the domain adaptive optimization of the YOLOX object detector, the auxiliary branches used for cross-domain optimization, namely the multi-scale global feature alignment module, domain-invariant feature decoupling module, and multi-scale instance-level adaptive module designed in this invention, are removed. Only the network model file of the object detector part is exported, and the cross-domain optimized object detector can be deployed on the vehicle-mounted inference platform. The domain adaptive object detection network optimization method designed in this invention does not change the network structure of the object detector. The inference behavior of the object detector is consistent with that before optimization. It reads in image data and infers the category, location size, and prediction confidence information of the target object.
[0049] Table 1. Convolutional Structure of the Adversarial Domain Discriminator
[0050]
[0051]
[0052] Table 2-1 Domain-Invariant Feature Encoder Edi
[0053]
[0054] Table 2-2 Domain-Specific Feature Encoders Eds
[0055]
Claims
1. A domain-adaptive target detection network optimization method for intelligent transportation, characterized in that, It includes three stages: dataset creation, domain-adaptive object detection network training, and object detection network inference. The specific steps for creating the dataset are as follows: For the collected train vehicle operation data, key frame data is extracted from the video according to the specified frame interval to obtain a coarsely screened dataset. Then, blurry, duplicate, and unusable data are removed by filtering to obtain the target domain dataset. Annotated public datasets are collected as the source domain dataset. The design of the domain-adaptive target detection network involves the following steps: The designed domain-adaptive target detection network includes a YOLOX target detector, a multi-scale global feature alignment module, a domain-invariant feature decoupling module, and a multi-scale instance-level adaptive module. Three sets of adversarial discriminators are designed for the three-scale output features of the YOLOX detector backbone network. The adversarial discriminators serve as the basic module for optimizing the weights of the domain-adaptive network. The YOLOX target detector includes a preprocessing module, a backbone network, and a detection head. The adversarial domain discriminator consists of a gradient inversion layer (GRL) and a domain discriminator. The domain discriminator contains three cascaded "Conv+GroupNorm+ReLU" operation blocks and an output convolutional layer, where GroupNorm is a regularization operation. The kernel size of the three operation blocks of the domain discriminator is 1×1 and does not change the number of channels of the feature. The output convolutional layer reduces the number of channels of the feature to 1 and outputs the domain prediction result of the feature. Multi-scale global feature alignment module; Domain adaptation is performed on the three scales of the feature extractor; the backbone network of the YOLOX object detector has outputs at three scales, with feature sizes of 80×80×128, 40×40×256 and 20×20×512, respectively. The output features at the three scales are named C128, C256 and C512 according to the number of channels. A multi-scale strategy was adopted, and the three feature layers C128, C256 and C512 of the backbone network were connected to the multi-scale global feature alignment module through three gradient inversion layers (GRL). As training progressed, the backbone network of the object detector extracted domain-invariant features from the image data. Domain-invariant feature decoupling module; A feature encoder was designed that divides the output features of the backbone network into domain-invariant features and domain-specific features along the channel dimension of the features, and only sends the domain-invariant features to the next stage of operation; it consists of three cascaded "Conv+IN+ReLU" operation blocks, "Conv+Sigmoid" operation blocks, and a global average pooling layer GAP, with the input and output feature sizes being the same; where IN is the InstanceNorm regularization operation; Multi-scale instance-level adaptive module; An attention map generator was designed, which uses the output information of the detection head to divide the features into foreground and background parts; the input is the category prediction information output by the detection head, where H×W and C represent the size and number of channels of the feature map, respectively; after averaging by channel dimension, an intermediate feature of size H×W×1 is obtained. In order to further filter out regions that are more likely to be background, the average value of the intermediate features is taken as the threshold, and features with response values lower than the threshold are set to zero, while features with response values not lower than the threshold are retained. After the above operations, an attention map is obtained to achieve data distribution alignment that is aware of foreground and background information. The object detection network inference process involves the following steps: Real-time image data is collected and sent to the preprocessing module for scaling and data normalization. The backbone network extracts depth features, and the detection head further infers the depth features of the backbone network, outputting the location, category, and prediction confidence information of the target object in the image.
Citation Information
Patent Citations
Self-driving automobile cross-domain target detection method based on improved YOLOv3
CN112800906A
Abnormal image detection method and device
CN113673631A