A Feature Fusion Method Applied to Vehicle Re-identification in Foggy Weather

By designing a feature fusion method in a convolutional neural network, combining multi-stage feature extraction and loss fusion of the defogging network and the re-identification network, the problem of low accuracy of vehicle re-identification under foggy weather conditions is solved, and a higher accuracy of vehicle re-identification is achieved, and traffic safety management capabilities are improved.

CN119810609BActive Publication Date: 2025-05-30SOUTHWEST PETROLEUM UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510306419.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-05-30
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

The existing convolutional neural network has low accuracy in vehicle re-identification and insufficient generalization capabilities in foggy weather conditions.

Method used

A feature fusion method is designed. The overall framework consists of a defogging network and a re-identification network. Through the multi-stage feature extraction and loss fusion of the public feature extraction module, a defogging network and a re-identification network, combined with attention mechanism and pyramid enhancement, stable re-identification features are extracted.

Benefits of technology

In foggy environments, higher accuracy of vehicle re-identification is achieved, improved the identification and tracking capabilities of traffic management and autonomous driving systems in severe weather conditions, and enhanced traffic safety management and environmental perception capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810609B_ABST
    Figure CN119810609B_ABST
Patent Text Reader

Abstract

The present invention discloses a feature fusion method applied to vehicle re-identification in foggy weather. The method specifically includes the following steps: input a foggy vehicle image into a common feature extraction module to extract common features; use a two-stage restoration method for the common features to extract relevant features for defogging; further process using pyramid enhancement and calculate the defogging loss; use a re-identification network for further feature extraction of the common features, and use a convolutional block attention mechanism and a feature pyramid method to extract global and local features, fuse the global and local features, and calculate the triplet loss and vehicle identity classification loss; fuse the three losses to obtain an overall loss function, and finally identify the images taken by other cameras with the same identity as the input image in the gallery set. The present invention realizes a vehicle re-identification method with better re-identification effect under foggy conditions. Compared with the existing methods, the present invention has achieved more excellent results in terms of accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of re-identification in computer vision object detection, and particularly relates to a feature fusion method applied to vehicle re-identification in foggy weather. Background Art

[0002] Foggy weather is one of the important weather conditions that cause traffic accidents. Low visibility will limit the driver's field of vision and increase the risk of traffic accidents. With the rapid development of intelligent transportation systems and autonomous driving technologies, the research on vehicle re-identification technology in foggy weather has become particularly important. Vehicle re-identification in foggy environments is a challenging task. Traditional vision systems are easily interfered in foggy weather, resulting in a significant reduction in the accuracy of vehicle detection and recognition. Especially in intelligent transportation systems, vehicle re-identification is a key link for realizing traffic flow monitoring, violation behavior recognition, and accident warning. The complexity and uncertainty in foggy environments further increase the technical difficulty. With the acceleration of urbanization and the sharp increase in the number of motor vehicles, traffic management and road safety are facing huge challenges. In adverse weather conditions such as foggy days, the incidence of traffic accidents has increased significantly, posing a serious threat to people's lives and property safety. Therefore, how to achieve efficient and accurate vehicle re-identification in foggy environments and improve the reliability and safety of intelligent transportation systems is of great significance for protecting people's lives and property safety.

[0003] Early vehicle re-identification methods in foggy weather mainly relied on traditional image enhancement techniques and feature extraction methods based on manual design. Image enhancement-based methods, such as histogram equalization and contrast-limited adaptive histogram equalization (CLAHE), although they can improve the visibility of images to a certain extent, often lead to the loss of image details and the amplification of noise, and it is difficult to maintain stable performance in complex foggy environments. Feature extraction methods based on manual design, such as scale-invariant feature transform (SIFT) and histogram of oriented gradients (HOG), although they show good recognition effects in clear images, under foggy conditions, due to the reduced image contrast and blurred details, the performance of these methods drops significantly. In addition to using traditional image processing techniques, some scholars have also tried to use physical models to estimate the fog concentration and perform defogging, but these methods do not work well in the case of thick fog or non-uniform fog. With the development of deep learning technology, vehicle re-identification methods based on convolutional neural networks have gradually become the mainstream. These methods can automatically extract more discriminative features through an end-to-end learning method, significantly improving the performance of vehicle re-identification. However, in foggy environments, due to the degradation of image quality and blurred features, traditional convolutional neural network methods still face the problem that the re-identification accuracy is insufficient to meet actual needs. Summary of the Invention

[0004] To solve the problems of low accuracy and insufficient generalization ability of existing vehicle re-identification methods based on convolutional neural networks under foggy conditions, the present invention combines the advantages of existing methods and designs a feature fusion method for vehicle re-identification in foggy weather. The overall framework of this method consists of two parts: a defogging network and a re-identification network. First, input the foggy vehicle image to be re-identified, and send the picture to the common feature extraction module to extract the common features to be sent to the defogging network and the re-identification network. The defogging network uses semi-supervised learning, and the re-identification network uses supervised learning. Second, the extracted common features pass through the defogging network and the re-identification network to extract the key features of defogging and the key features of re-identification respectively. A feature fusion module and an attention mechanism are introduced into the re-identification network to extract more detailed re-identification features. Finally, extract stable re-identification features that are not affected by fog and perform the re-identification task on the input image.

[0005] Specifically, the present invention provides a feature fusion method for vehicle re-identification in foggy weather, which mainly includes the following steps:

[0006] S1. Input the foggy vehicle image into the overall network framework, and extract the common features required later through the common extraction module;

[0007] S2. Use a two-stage recovery method for the common features extracted in S1 to enhance the accuracy and pertinence of feature representation, and extract the relevant features of defogging;

[0008] S3. Use pyramid enhancement for the defogging features obtained in S2 to form a richer and more detailed feature representation, and calculate the defogging loss L D ;

[0009] S4. For the common features extracted in S1, use the re-identification network to extract features, take the features processed by the attention mechanism as global features, and take the multi-scale features extracted by the feature pyramid method as local features. Fuse the global features and local features to generate a feature map that retains rich local details and combines context information, and calculate the triplet loss L T and the vehicle identity classification loss L ID ;

[0010] S5. Fuse the losses obtained in S3 and S4 to obtain the overall loss function, and use the overall loss function to change the weights and parameters of the common feature extraction module; finally, use the re-identification branch to identify the images taken by other cameras with the same identity as the input image from the gallery set.

[0011] The common extraction module described in step S1 is composed of the first two stages of Resnet50 and is used to roughly extract the features of the image.

[0012] The two-stage restoration method described in step S2 first uses a physical model to remove haze, and then uses a convolutional network for further detail optimization to obtain the feature vector required for haze removal.

[0013] As described in step S3 L D According to the differences between supervised learning and unsupervised learning, it is divided into two. When it is supervised learning, it is L DFS When it is unsupervised learning, it is L DFU .

[0014] The attention mechanism described in step S4 is the convolutional block attention mechanism, which processes the features finally generated by Renset50. The multi-scale features are the features obtained after processing by the feature pyramid. The input of the feature pyramid is the intermediate features generated by the first four stages of Resnet50.

[0015] The fusion processing described in step S5 refers to the process of weighted summation using the three loss functions obtained from the two branches to obtain the overall loss function; using the overall loss function to change the parameters and weights of the common feature extraction module so that the network common feature extraction module can achieve the purpose of extracting stable features. Since the haze removal branch uses semi-supervised learning, there are two overall loss functions: when the haze removal branch performs supervised learning, it is L S , and the formula is as follows:

[0016] ;

[0017] Among them, L T represents the triplet loss of the re-identification branch; L ID represents the classification loss of the re-identification branch; L DFS represents the haze removal branch loss in the supervised learning stage; λ 1 is the weight coefficient used to balance the haze removal branch loss;

[0018] When the haze removal branch performs unsupervised learning, the overall loss is L U , and its formula is as follows:

[0019] ;

[0020] Among them, L DFUis the loss of the defogging branch in the unsupervised learning stage; L T and L ID are the triplet loss and classification loss of the re-identification branch respectively; when actually performing the re-identification task, the defogging branch does not participate in the re-identification process, and only the re-identification branch is used to complete the vehicle re-identification task.

[0021] Beneficial effects: It realizes more accurate re-identification of target vehicles in foggy environments. Using the present invention for foggy-weather vehicle re-identification can help traffic management departments and autonomous driving systems more efficiently identify and track target vehicles under adverse weather conditions, improve traffic safety management capabilities and environmental perception capabilities, and prevent traffic accidents caused by low visibility. By designing a multi-scale feature extraction module and combining an attention mechanism, the present invention can capture local details and global structure information of vehicles at different scales. This strategy of fusing multi-scale local features and global features can effectively address the problems of blurred image details and low contrast in foggy environments, ensuring that the model can obtain reliable feature representations under different fog concentrations, thereby improving the accuracy of vehicle re-identification under foggy conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a schematic flowchart of a feature fusion method for foggy-weather vehicle re-identification according to the present invention;

[0023] Figure 2 is a network structure diagram of generating multi-scale local features by a feature pyramid according to the present invention;

[0024] Figure 3 is a re-identification result diagram of foggy-weather vehicle images, including query images, the top ten re-identification results, and the correctness of the re-identification results. The green box indicates that the re-identification result is correct; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the embodiments and the drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.

[0026] As Figure 1 、 Figure 2 、 Figure 3 shown, the present invention provides a feature fusion method for foggy-weather vehicle re-identification, including the following steps:

[0027] S1. Input the foggy vehicle image to be re-identified into the overall network. After extracting the common features through the common feature extraction module, it will pass through two branches, the defogging network branch and the re-identification network branch: the defogging network branch is responsible for defogging the foggy image; the re-identification branch is responsible for re-identifying the input image to find other image data with the same identity as the input foggy vehicle image.

[0028] S2. Use a two-stage recovery method for the common features extracted in S1 to extract the relevant features for defogging.

[0029] S3. Perform pyramid enhancement operations on the defogging features extracted in S2 to further enhance the relevant features for defogging and obtain the defogging loss L D ;

[0030] S4. Use the subsequent three stages of Resnet50 to further process the common features extracted in S1. After feature extraction, use the attention mechanism to enhance the relevant features for re-identification to obtain a stronger feature representation and use it as the global feature; use the four feature maps of different sizes extracted in the first four stages of Resnet50 as the input of the feature pyramid to obtain multi-scale local features; finally, splice the global feature and the local feature to fuse them into the final re-identification feature and calculate the triplet loss L T and the classification loss of the vehicle identity L ID ;

[0031] S5. Finally, fuse the three losses obtained in S3 and S4 to obtain the overall loss function, change the weights and parameters of the common feature extraction module; and use the re-identification branch to find the vehicle images captured by other cameras with the same identity as the input foggy vehicle image.

[0032] All the input images in step S1 are reshaped to 384×384 for subsequent unified feature extraction; the defogging branch uses semi-supervised learning. In this example, the ResNet-50 model pre-trained for the ImageNet image classification task is used. The first two stages of this model are used as our common feature extraction module, and the feature maps of these two stages are saved so that they can be used as the input of the subsequent feature pyramid. The number of channels of the saved feature maps are 64 and 256 respectively.

[0033] In step S4, the feature maps generated in the last three stages of Resnet50 are saved accordingly, and their numbers of channels are 512, 1024, and 2048 respectively; the attention mechanism is the convolutional block attention mechanism, which combines the spatial attention mechanism and the channel attention mechanism; the feature pyramid uses the four feature maps generated in the first four stages of Resnet50 as inputs, and the multi-scale local feature generation process is as Figure 2 .

[0034] The fusion process described in step S5 refers to the operation of weighted summation using the three losses obtained from the dehazing branch and the re-identification branch during the training process to calculate the overall loss function; subsequently, the overall loss function is used to make corresponding changes to the weights and parameters of the common feature extraction module, so as to enable the common feature extraction module to extract stable and re-identification-friendly key features from vehicle images in foggy days.

[0035] In a specific implementation case of the present invention, the dataset used is the FVRID dataset, which is obtained by processing the VehicleID and VERI-Wild datasets. The FVRID dataset includes two parts: real data and synthetic data. Each dataset includes a training set (Train Set), a probe set (Probe Set), and a gallery set (Gallery Set). The probe set contains a set of specific query samples, which are used to test whether the system can correctly find the corresponding matching items in the gallery set. In the training stage, for the vehicle re-identification task, the basic Resnet50 network is selected as the basic network architecture, and the initial weights are loaded from the ImageNet pre-trained model. As for the dehazing branch, its parameters are initialized by the Kaiming normal initialization method. To avoid the overfitting problem that may be caused by the limited number of training samples, data augmentation techniques of horizontal flipping and random cropping are introduced during the training process. The entire model is trained for 120 rounds, the training batch size is set to 32, and a learning rate scheduling mechanism with gradually increasing temperature is adopted: the initial learning rate is 1.09e-5, and it is increased to 1e-4 after the 10th epoch. To optimize the model, in the specific implementation case of the present invention, the optimizer used to train the model is the Adam optimizer, in which a learning rate decay coefficient of 0.6 is set, the learning rate is 0.0001, the deep learning framework used is Pytorch1.8, and the hyperparameter λ 1 is set to 1. Both branches participate in the training stage, and the dehazing branch helps the re-identification branch extract stable re-identification features; while only the re-identification branch performs the re-identification task in the test stage.

[0036] The corresponding data re-identification result graph is as Figure 3Shown: The first column of the picture is the query image, and the following ten images are the top ten results of the re-identification of the input vehicle image by this method. Compared with previous methods, the present invention combines multi-scale feature fusion and global-local feature complementary mechanism, and has higher vehicle re-identification accuracy in foggy environments.

[0037] The technical means disclosed in the solution of the present invention are not limited to the technical means disclosed in the above embodiments, but also include technical solutions composed of any combination of the above technical features. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and retouches can be made, and these improvements and retouches are also regarded as the protection scope of the present invention.

Claims

1. A feature fusion method for vehicle re-identification in foggy weather, characterized by: The method comprises the following steps:

1. Input the foggy vehicle image to be identified into the overall network framework, use the first two stages of Resnet50 to extract common features, and save the feature maps of the first two stages; use the common features extracted in the first two stages of Resnet50 as the input of the defogging network branch and the re-identification network branch; the defogging network branch is used for defogging foggy images; the re-identification network branch is used to re-identify the input image and find other image data that are consistent with the input foggy vehicle image; Second, use a two-stage recovery method on the common features extracted in step 1 to obtain the relevant features for defogging; 3. Use pyramid enhancement to enhance the defogging features obtained in step 2 to form a richer and more detailed defogging feature map, and calculate the defogging loss; Fourth, the common features extracted in step 1 are further extracted using the three stages of Resnet50, and the three feature maps extracted in the three stages are saved; after the feature extraction is completed, the re-identification related features are enhanced using the attention mechanism to obtain a stronger feature representation, and the feature representation is used as the global feature; the multi-scale features extracted using the feature pyramid method are used as local features, the global features are concatenated with the local features, and the final re-identification feature map is fused, and the triplet loss and vehicle identity classification loss are calculated; 5. The defogging loss, triplet loss and vehicle identity classification loss obtained in step 3 and step 4 are fused. The fusion process refers to the weighted summation of the defogging loss, triplet loss and vehicle identity classification loss to obtain the overall loss function. The overall loss function is used to change the parameters and weights of the first two stages of the Resnet50, and finally the re-identification branch is used to identify images taken by other cameras that are consistent with the input image identity from the gallery set; the overall loss function is divided into the overall loss of the defogging branch during supervised learning L S And the overall loss of the dehazing branch during unsupervised learning L U ; Overall loss of supervised learning of the dehazing branch L S The formula is as follows: ; in, L T represents the triplet loss of the re-identification branch; L ID represents the classification loss of the re-identification branch; L DFS represents the dehazing branch loss in the supervised learning stage; λ 1 is the weight coefficient used to balance the loss of the defogging branch; Overall loss of unsupervised learning of the dehazing branch L U The formula is as follows: ; in, L DFU is the loss of the dehazing branch in the unsupervised learning stage; L T and L ID They are the triplet loss and classification loss of the re-identification branch respectively. When actually performing the re-identification task, the defogging branch does not participate in the re-identification process, and only the re-identification branch is used to complete the vehicle re-identification task.

2. According to claim 1, a feature fusion method for vehicle re-identification in foggy weather, characterized in that: The first two stages of Resnet50 in step 1 serve as the common feature extraction module of the entire framework, and the number of channels of the two feature maps saved is 64 and 256.

3. According to claim 1, a feature fusion method for vehicle re-identification in foggy weather, characterized in that: The number of channels of the feature maps of the three stages saved in step 4 are 512, 1024, and 2048 respectively; the input of the convolutional block attention mechanism is the features extracted in the last stage of Resnet50; The feature pyramid takes the features extracted by the first four stages of Resnet50 as input.

Citation Information

Patent Citations

  • Vehicle re-identification method and device based on roadside perception, and electronic equipment

    CN114170516A

  • Vehicle re-identification method based on foggy-day high-altitude shooting of depth camera

    CN117649535A