Visible light-infrared cross-modal vehicle re-identification method based on multi-scale deformation convolution
By introducing a multi-scale deformation convolution feature extraction module into the vehicle recognition network, the problem of poor performance of visible-infrared vehicle recognition in poor light environments is solved, and more efficient vehicle recognition capabilities are achieved.
Patent Information
- Application Number
- CN202510011095.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-27
AI Technical Summary
The existing vehicle re-identification methods have limited effects in poor light or night environments, and cross-modal differences lead to poor visible-infrared vehicle re-identification effects.
The feature extraction module of multi-scale deformation convolution is adopted to generate linear deformable convolution kernels of different scales to improve the network's ability to capture information at different levels and embed representations, and comprehensively improve the visible-infrared cross-modal vehicle re-identification capability.
It effectively solves the problem that vehicle target images are difficult to align under different perspectives and attitudes, improves the network's ability to capture information at different levels, and improves the accuracy and stability of visible-infrared cross-modal vehicle re-identification.
Smart Images

Figure CN120047904A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and vehicle re-identification, and particularly relates to a visible light-infrared cross-modal vehicle re-identification method based on multi-scale deformable convolution. Background Art
[0002] In the field of computer vision, vehicle re-identification has become an important and cutting-edge research direction. This task aims to identify vehicles with the same identity from an image library captured by non-overlapping cameras in different traffic environments. This technology plays a key role in the implementation of intelligent transportation systems and smart cities, and demonstrates great potential in various applications such as urban surveillance and social security. In order to obtain the global visual features of vehicle images, vehicle re-identification mainly uses convolutional neural networks for feature extraction. With the rapid development of convolutional neural networks and their variants in recent years, and the continuous emergence of large-scale vehicle re-identification datasets, the performance of vehicle re-identification technology has been significantly improved, and its performance in practical applications has become increasingly excellent.
[0003] Currently, most vehicle re-identification methods are based on visible light single-modal matching, but their effects are limited in poor light or night environments. In low-light or night-time surveillance, thermal infrared cameras can replace optical cameras to capture the appearance features of targets, which highlights the importance of visible light-infrared vehicle re-identification. Compared with single-modal vehicle re-identification, cross-modal vehicle re-identification is more challenging, mainly due to the significant modal differences between visible light and infrared images. To solve the cross-modal difference problem, researchers have proposed feature-level methods and image-level methods. Feature-level methods project vehicle features in different modalities onto a unified feature subspace to reduce the differences between modalities. Image-level methods use generative adversarial networks to convert visible light images into infrared images, or infrared images into visible light images, so as to achieve modal alignment. However, the generated images often have large noise, resulting in poor cross-modal matching effects. Summary of the Invention
[0004] In view of the above technical problems, the present invention proposes a visible light-infrared cross-modal vehicle re-identification method based on multi-scale deformable convolution, which can specifically address the problem that vehicle target images at different perspectives and poses are difficult to align, improve the network's ability to capture different levels of information and embedding representations, and comprehensively improve the visible light-infrared cross-modal vehicle re-identification ability.
[0005] The technical solution for achieving the object of the present invention is: a visible light-infrared cross-modal vehicle re-identification method based on multi-scale deformable convolution, comprising the following steps:
[0006] Step S1, constructing a visible light-infrared cross-modal vehicle re-identification dataset;
[0007] Step S2: Construct a visible light-infrared cross-modal vehicle re-identification basic network;
[0008] Step S3: Construct a feature extraction module based on multi-scale deformable convolution and add it to the basic network to obtain a vehicle re-identification model;
[0009] Step S4: Use the training set in the data to train the vehicle re-identification model;
[0010] Step S5: Use the trained vehicle re-identification model to complete visible light-infrared cross-modal vehicle re-identification.
[0011] According to a technical solution of the present invention, in the step S1, it specifically includes:
[0012] Step S11: Based on the category system of vehicle targets and the sample annotation specification, collect a qualified visible light-infrared cross-modal vehicle target detection data set with horizontal box annotations for the target positions and categories in the images;
[0013] Step S12: Slice the collected visible light-infrared vehicle target detection data set to form a visible light-infrared cross-modal vehicle re-identification data set;
[0014] Step S13: Split the data into a training set, a validation set, and a test set according to a preset ratio.
[0015] According to a technical solution of the present invention, in the step S2, a two-stream network is used as the visible light-infrared cross-modal vehicle re-identification basic network;
[0016] Adopt a non-local attention module to capture global information and long-range dependence relationships;
[0017] Use generalized mean pooling to capture fine-grained discriminative features in a specific domain;
[0018] On the features, a similarity measurement network is respectively adopted to obtain the cross-modal visible light-infrared vehicle re-identification result.
[0019] According to a technical solution of the present invention, in the step S3, a feature extraction module based on multi-scale deformable convolution is constructed, and by generating linear deformable convolution kernels of different scales, the ability of the network to capture information at different levels and embedding representations is improved.
[0020] According to a technical solution of the present invention, in the step S3, it further includes:
[0021] Use three deformable convolution layers with different numbers of convolution kernels to reduce the number t of feature maps to 1 / 4 of its own size and combine them into the same feature map;
[0022] Use the activation layer F ReLU Improve the non - linear representation ability of the multi - scale deformable convolution module;
[0023] Apply another convolutional layer θ to the obtained feature map 1×1 , whose kernel size is 1×1, making its dimension the same as the number t of feature maps, then the embedded representation generated by the branch is:
[0024]
[0025] According to a technical solution of the present invention, in step S4, the training set is used to train the visible - light infrared cross - modal vehicle re - identification network with multi - scale deformable convolution;
[0026] Use the validation set to observe the matching performance of the visible - light infrared cross - modal vehicle re - identification network, design a triplet loss function, adopt the stochastic gradient descent optimization algorithm, set the learning rate decay strategy, and update the parameters of the visible - light infrared cross - modal vehicle re - identification network until the network performance converges, obtaining a trained visible - light infrared cross - modal vehicle re - identification model based on multi - scale deformable convolution.
[0027] According to a technical solution of the present invention, before step S5, the test set is input into the trained visible - light infrared cross - modal vehicle re - identification model based on multi - scale deformable convolution to obtain the matching result on the test set, completing model verification.
[0028] According to an aspect of the present invention, there is provided an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the above - mentioned one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes a visible - light infrared cross - modal vehicle re - identification method according to any one of the above - mentioned technical solutions.
[0029] According to an aspect of the present invention, there is provided a computer - readable storage medium for storing computer instructions, and when the computer instructions are executed by a processor, a visible - light infrared cross - modal vehicle re - identification method according to any one of the above - mentioned technical solutions is implemented.
[0030] Compared with the prior art, the present invention has the following beneficial effects:
[0031] According to the concept of the present invention, a visible-light infrared cross-modal vehicle re-identification method based on multi-scale deformable convolution is proposed. Aiming at the problem that it is difficult to align vehicle target images under different perspectives and postures, by generating linear deformable convolution kernels of different scales, it is possible to learn richer feature representations, improve the network's ability to capture information at different levels and embedding representations, and comprehensively improve the visible-light infrared cross-modal vehicle re-identification ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 FIG. 6 schematically shows a flowchart of a visible-light infrared cross-modal vehicle re-identification method based on multi-scale deformable convolution according to an embodiment of the present invention;
[0033] FIGS. 2(a-d) schematically show a schematic diagram of a visible-light infrared cross-modal vehicle re-identification network based on multi-scale deformable convolution according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0035] The present invention will be described in detail below with reference to the drawings and specific embodiments. The embodiments cannot be enumerated one by one here, but the embodiments of the present invention are not limited to the following embodiments.
[0036] As Figure 1 shown in FIGS. 2 and 6, a visible-light infrared cross-modal vehicle re-identification method based on multi-scale deformable convolution of the present invention includes the following steps:
[0037] Step S1, constructing a visible-light infrared cross-modal vehicle re-identification data set;
[0038] Based on the category system and sample annotation specifications of vehicle targets, a qualified visible-light infrared cross-modal vehicle target detection data set with horizontal box annotations for the target positions and categories in the images is collected. The annotated images are sliced, and the slice size is set to 144*288 to form a visible-light infrared cross-modal vehicle re-identification data set, which is divided into a training set, a validation set, and a test set according to a ratio of 7:1:2.
[0039] By using a category system and sample annotation specifications based on vehicle targets, a visible light-infrared vehicle target detection dataset is collected and sliced to form a cross-modal vehicle re-identification dataset. This process ensures the diversity and representativeness of the dataset, which helps improve the effectiveness of model training. By dividing the training set, validation set, and test set, the generalization ability of the model can be effectively evaluated to ensure the accuracy and stability of the model in practical applications.
[0040] Step S2: Construct a visible light-infrared cross-modal vehicle re-identification basic network;
[0041] As shown in Figure 2(a), a two-stream network is used as the visible light-infrared cross-modal vehicle re-identification basic network, and this network uses a two-stream ResNet-50 network as the backbone. Visible light and infrared features are input into multiple non-local attention mechanism modules. The non-local attention mechanism module is shown in Figure 2(b), which can capture global information and long-range dependencies, better capture global context information, and enhance the expression ability of features. In the training stage, all features before and after the batch normalization layer are input into different losses, as shown in Figure 2(c), to jointly optimize the visible light-infrared cross-modal vehicle re-identification network.
[0042] The backbone network ResNet-50 is used to extract visible light-infrared two-modal feature maps of the input image, and a non-local attention module is adopted to capture global information and long-range dependencies, better capture global context information, and enhance the expression ability of features; generalized mean pooling is used to capture fine-grained discriminant features in a specific domain; a similarity metric network is respectively adopted on the features to obtain the cross-modal visible light-infrared vehicle re-identification result.
[0043] Step S3: Construct a feature extraction module based on multi-scale deformable convolution and add it to the basic network to obtain a vehicle re-identification model;
[0044] The visible light-infrared cross-modal vehicle re-identification basic network uses the backbone network Resnet-50 to extract feature maps. However, since traditional convolution operators are used to extract target features, the feature information of targets scattered at different perspectives and postures cannot be effectively captured, and the feature expression ability of the basic network is weak; as shown in Figure 2(a), in this step, a feature extraction module based on multi-scale deformable convolution is constructed and added to the feature extraction network to enhance the feature extraction and expression ability of the network.
[0045] To alleviate the single feature extraction ability of traditional convolution operators, as shown in Figure 2(b), in this step, a feature extraction module based on multi-scale deformable convolution is constructed. By generating linear deformable convolution kernels of different scales, more abundant feature representations are learned, the ability of the network to capture information at different levels and embedding representations is improved, and then feature maps with better discriminant ability are generated.
[0046] For each branch of the multi-scale deformable convolution, a deformable convolution layer with three different numbers of convolutional kernels (here 5, 7, and 9) is used Reduce the number t of feature maps to 1 / 4 of their own size, then combine them into the same feature map, and then use the activation layer F ReLU to improve the non-linear representation ability of the multi-scale deformable convolution module. Then, apply another convolutional layer θ 1×1 to the obtained feature map, with a kernel size of 1×1, to make its dimension t the same. Therefore, the embedding generated by this branch can be written as follows:
[0047]
[0048] where f represents the original embedding, f + represents the embedding generated by the multi-scale deformable convolution module, F ReLU represents the ReLU activation layer, represents the j-th deformable convolution layer with a deformable convolutional kernel size of i, and θ 1×1 represents the convolutional layer with a kernel size of 1×1.
[0049] By constructing a feature extraction module based on multi-scale deformable convolution, the model can generate deformable convolutional kernels of different scales, thereby learning richer feature representations. This design enhances the network's ability to capture information at different levels, improves the representation ability of feature embeddings, and makes the model perform more excellently when dealing with complex scenes and fine-grained features. In addition, the deformable convolution layers with different scales in multiple branches further enrich the feature embeddings and enhance the non-linear representation ability of the model.
[0050] Step S4: Use the training set in the data to train the vehicle re-identification model;
[0051] Use the training set in the above step S1 to train the visible-light infrared cross-modal vehicle re-identification network based on multi-scale convolution, and use the validation set in step S1 to observe the matching performance of the re-identification network. Use the cross-entropy loss function as the classification loss function, and add the weighted regularization triplet loss function. Initialize the backbone network with Resnet-50 pre-trained on the ImageNet dataset, and use the stochastic gradient descent method to iteratively train for 100 epochs, with a batch size of 4 and an initial learning rate set to 0.01. Update the parameters of the visible-light infrared cross-modal vehicle re-identification network until the network performance converges to obtain the visible-light infrared cross-modal vehicle re-identification model based on multi-scale deformable convolution.
[0052] By counting the number of instances of each category in the training set, calculating the category occurrence frequency, designing a category balance probability function, and constructing a category balance copy-paste data augmentation strategy, the number of target instances in the training set is balanced through copy-paste to alleviate the sub-optimization problem of the fine-grained object detection network caused by the often-existing problem of unbalanced target categories in fine-grained object detection.
[0053] Step S5: Use the trained vehicle re-identification model to complete visible-light to infrared cross-modal vehicle re-identification.
[0054] In some embodiments of the present invention, before the step S5, the test set is input into the trained visible-light to infrared cross-modal vehicle re-identification model based on multi-scale deformable convolution to obtain the matching results on the test set, and the model verification is completed.
[0055] Input the test set in step S1 into the trained visible-light to infrared cross-modal vehicle re-identification model based on multi-scale deformable convolution to obtain the matching results. On the one hand, rank-1 and mean average precision mAP are used as evaluation indicators to compare the numerical results of the basic network and the visible-light to infrared cross-modal vehicle re-identification method based on multi-scale deformable convolution. On the other hand, the visualization results of the basic network and the visible-light to infrared cross-modal vehicle re-identification method based on multi-scale deformable convolution are compared.
[0056] According to one aspect of the present invention, there is provided an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the above one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory so that the electronic device executes the visible-light to infrared cross-modal vehicle re-identification method based on multi-scale deformable convolution as described in any one of the above technical solutions.
[0057] According to one aspect of the present invention, there is provided a computer-readable storage medium for storing computer instructions, which when executed by a processor, implement a visible-light to infrared cross-modal vehicle re-identification method based on multi-scale deformable convolution as described in any one of the above technical solutions.
[0058] In summary, the present invention proposes a visible-light infrared cross-modal vehicle re-identification method based on multi-scale deformable convolution. The visible-light infrared cross-modal vehicle re-identification method based on multi-scale deformable convolution includes: Step S1, constructing a visible-light infrared cross-modal vehicle re-identification data set; Step S2, constructing a visible-light infrared cross-modal vehicle re-identification basic network; Step S3, constructing a feature extraction module based on multi-scale deformable convolution and adding it to the basic network to obtain a vehicle re-identification model; Step S4, training the vehicle re-identification model using the training set in the data; Step S5, using the trained vehicle re-identification model to complete visible-light infrared cross-modal vehicle re-identification. It can address the problem that vehicle target images are difficult to align under different perspectives and poses. By generating linear deformable convolution kernels of different scales, it can learn richer feature representations, enhance the network's ability to capture different levels of information and embedding representations, and comprehensively improve the visible-light infrared cross-modal vehicle re-identification ability.
[0059] In addition, it should be noted that the present invention can be provided as a method, device, or computer program product. Therefore, the embodiments of the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.
[0060] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0061] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1The functions specified in one or more boxes. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes and / or boxes. Figure 1 One process or more processes and / or boxes Figure 1 Steps for implementing the functions specified in one or more boxes.
[0062] It should also be noted that in this text, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such a process, method, article or terminal device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or terminal device including the said element.
[0063] Finally, it should be noted that the above is the preferred embodiment of the present invention. It should be pointed out that although the preferred embodiments of the present invention have been described, for those skilled in the art of this technology, once the basic creative concept of the present invention is known, without departing from the principle described in the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
Claims
1. A visible light-infrared cross-modal vehicle re-identification method based on multi-scale deformation convolution, comprising the following steps: Step S1, constructing a visible light-infrared cross-modal vehicle re-identification dataset; Step S2, constructing a visible light-infrared cross-modal vehicle re-identification basic network; Step S3: construct a feature extraction module based on multi-scale deformation convolution and add it to the basic network to obtain a vehicle re-identification model; Step S4, training a vehicle re-identification model using a training set in the data; Step S5: Use the trained vehicle re-identification model to complete visible light-infrared cross-modal vehicle re-identification.
2. The method according to claim 1, characterized in that In the step S1, it specifically includes: Step S11: Based on the vehicle target classification system and sample annotation specifications, collect a qualified visible light-infrared cross-modal vehicle target detection dataset that annotates the target position and category in the image with a horizontal box; Step S12: Slice the collected visible light-infrared vehicle target detection dataset to form a visible light-infrared cross-modal vehicle re-identification dataset; Step S13: split the data into a training set, a validation set and a test set according to a preset ratio.
3. The method according to claim 2, characterized in that In the step S2, a dual-stream network is used as a basic network for visible light-infrared cross-modal vehicle re-identification; Adopt non-local attention module to capture global information and long-distance dependencies; Use generalized average pooling to capture fine-grained discriminative features in specific areas; Similarity measurement networks are used on the features to obtain cross-modal visible light-infrared vehicle re-identification results.
4. The method according to claim 2, characterized in that: In step S3, a feature extraction module based on multi-scale deformable convolution is constructed, and the ability of the network to capture information at different levels and embed representation is improved by generating linear deformable convolution kernels of different scales.
5. The method according to claim 3, characterized in that: The step S3 further includes: Using three deformable convolutional layers with different numbers of convolution kernels, the number of feature maps t is reduced to 1 / 4 of its own size and combined into the same feature map; Use activation layer F ReLU Improve the nonlinear representation capability of the multi-scale deformable convolution module; Apply another convolutional layer θ to the resulting feature map 1×1 , whose kernel size is 1×1, making its dimension the same as the number of feature maps t, then the embedding generated by the branch is represented as: Where f represents the original embedding, f + represents the embedding generated by the multi-scale deformable convolution module, F ReLU represents the ReLU activation layer, represents the jth deformable convolution layer with a deformable convolution kernel size of i, θ 1×1 Represents a convolutional layer with a convolution kernel size of 1×1.
6. The method according to claim 3, characterized in that In the step S4, the training set is used to train a multi-scale deformable convolutional visible light-infrared cross-modal vehicle re-identification network; The validation set is used to observe the matching performance of the visible light-infrared cross-modal vehicle re-identification network, a triplet loss function is designed, a stochastic gradient descent optimization algorithm is adopted, a learning rate decay strategy is set, and the parameters of the visible light-infrared cross-modal vehicle re-identification network are updated until the network performance converges, thereby obtaining a trained visible light-infrared cross-modal vehicle re-identification model based on multi-scale deformation convolution.
7. The method according to claim 3, characterized in that Before step S5, the test set is input into the trained multi-scale deformation convolution-based visible light-infrared cross-modal vehicle re-identification model to obtain the matching result on the test set and complete the model verification.
8. An electronic device, characterized in that: include: One or more processors, one or more memories, and one or more computer programs; wherein the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device is running, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes the visible light-infrared cross-modal vehicle re-identification method of multi-scale deformation convolution as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that: Used to store computer instructions, which, when executed by a processor, implement the visible light-infrared cross-modal vehicle re-identification method based on multi-scale deformation convolution as described in any one of claims 1 to 7.