Anonymous vehicle matching method based on depth features

Through multi-branch network and feature fusion technology, the problems of low accuracy, poor robustness and insufficient real-time performance in anonymous vehicle matching are solved, and high-precision and fast anonymous vehicle recognition are achieved, adapting to the lighting and viewing angle changes in the multi-camera environment, and improving the generalization ability of the model.

CN120339659AInactive Publication Date: 2025-07-18SHENZHEN INSTITUTE OF INFORMATION TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510403070.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems in the matching of anonymous vehicles with low matching accuracy, poor robustness, high computational complexity and insufficient real-time performance. Especially in the multi-camera environment, the impact of lighting changes and viewing angle differences is significant, and the deep learning model lacks generalization capabilities in the case of scarce data and uneven categories.

Method used

A multi-branch network is used to extract vehicle features by combining different loss functions and attention mechanisms, feature fusion is performed through non-negative matrix decomposition, and high-quality synthetic data is generated by combining generative adversarial networks, self-supervised learning and classification networks. Adaptive lighting normalization and feature importance evaluation mechanisms are introduced to optimize the overall loss function to improve the generalization ability and robustness of the model.

Benefits of technology

It significantly improves the accuracy and robustness of anonymous vehicle matching, reduces the computational complexity, meets real-time requirements, and enhances the model's adaptability and generalization capabilities in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339659A_ABST
    Figure CN120339659A_ABST
Patent Text Reader

Abstract

The invention discloses an anonymous vehicle matching method based on depth features, and relates to the technical field of vehicle matching methods. Firstly, a multi-branch network containing different modules and loss functions is built, and global and local features of a vehicle image are extracted at the same time; fusing the multi-branch features by using a tensor rank constraint non-negative matrix factorization model integrated with singular value saliency difference information; in the process, a self-adaptive illumination normalization module is added in front of an input layer, and a feature importance evaluation mechanism is introduced during fusion, so that the matching precision and robustness are improved, and anonymous vehicle matching is completed. According to the method, anonymous vehicles can be accurately matched, the matching precision is improved through multi-branch network and feature fusion, the adaptability to the complex environment is high, illumination and view angle change are avoided, the reasoning speed is high, and the real-time requirement is met. And the system also has good generalization ability, can work stably and reliably in different scenes, and effectively solves the existing technical problems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle matching methods, and in particular to an anonymous vehicle matching method based on deep features. Background Art

[0002] In the fields of intelligent transportation and public safety, the effective monitoring and identification of vehicles are crucial. With the increasing complexity of urban traffic and the continuous improvement of safety requirements, the anonymous vehicle matching technology has become a research hotspot. Anonymous vehicle matching is to match vehicles in different images based on vehicle appearance features and driving trajectories in the absence of direct identification information, and it is widely used in scenarios such as traffic flow analysis and hit-and-run vehicle tracking.

[0003] Early anonymous vehicle matching mainly relied on traditional machine learning algorithms, which designed features manually (such as HOG, SIFT, SURF, etc.) and combined classifiers such as support vector machines (SVM) and k-nearest neighbors (k-NN) for identification. Although these methods have certain interpretability and low computational overhead, and the training and inference speeds are relatively fast, their limitations have gradually emerged in the face of complex actual scenarios. In a multi-camera monitoring environment, the perspective differences, lighting changes between different cameras, and the diversity of vehicle appearances make it difficult for manually designed features to comprehensively and accurately describe vehicle features, resulting in low matching accuracy. Moreover, on large-scale complex datasets, traditional methods are difficult to capture deep image patterns and have poor generalization ability when migrating across tasks and datasets.

[0004] In recent years, deep learning models have been widely used in anonymous vehicle matching. Convolutional neural networks (CNNs), vision transformers, etc. automatically extract vehicle appearance features through multi-layer structures, from low-level edge and texture features to high-level shape and color distribution features, and perform well on large-scale datasets. The introduction of the attention mechanism has further enhanced the feature expression ability of the model. However, there are still many problems with deep learning models. The perspective differences, lighting changes between multiple cameras, and the complexity of vehicle appearance features make the robustness of deep feature extraction and the matching accuracy need to be improved. The problems of data scarcity and class imbalance also affect the matching effect. Rare vehicle appearance features account for a low proportion in the dataset, and it is difficult for the model to effectively generalize. In addition, the high computational complexity of deep learning models leads to huge power consumption during the training phase and slow speed during the inference phase, which cannot meet the real-time requirements in intelligent transportation monitoring and emergency response. Summary of the Invention

[0005] The anonymous vehicle matching method based on deep features proposed by the present invention aims to solve the problems mentioned in the above prior art.

[0006] To achieve the above object, the present invention adopts the following technical solutions: An anonymous vehicle matching method based on deep features, comprising:

[0007] Construct a multi-branch network to extract vehicle features: Build a branch network, with each branch based on the residual network as the basic architecture; Branch 1 uses the residual network combined with the triplet loss, Branch 2 uses the residual network combined with the cross-entropy loss, Branch 3 uses the residual network, attention module and triplet loss, and Branch 4 uses the residual network, attention module and cross-entropy loss; Extract the features of the vehicle image through the branches simultaneously; Among them, the triplet loss formula is, where is the anchor sample, is the positive sample, is the negative sample, represents the feature extraction function, and α is the set boundary value used to control the distance between the positive sample and the negative sample; The cross-entropy loss formula is, where C is the number of categories, is the one-hot vector of the true label of the sample, and is the predicted probability distribution of the network for the sample belonging to the i-th category; To enhance the robustness of the network to vehicle features under different lighting conditions, an adaptive lighting normalization module is added before the input layer of each branch network. The module dynamically adjusts the brightness and contrast of the image based on the histogram statistical information of the image. Its adaptive adjustment formula is, where is the pixel value of the original image at position, and and β are parameters dynamically calculated according to the image histogram;

[0008] Multi-feature fusion: For the feature vectors output by the branch network, a non-negative matrix factorization model with tensor rank constraints integrating singular value difference information is used for fusion, and the objective function is, which is equivalent to after conversion; Where K represents the number of feature matrices, is the k-th feature matrix, and and are the decomposition matrices, is the core matrix to be decomposed, and and are hyperparameters for balancing different loss terms, represents the Frobenius norm of the matrix, represents the nuclear norm of the matrix, is the rank of, is the weight of the i-th singular value, is the i-th singular value of, is the penalty coefficient, and is the auxiliary matrix; To optimize the feature fusion effect, a feature importance evaluation mechanism is introduced. According to the contribution degree of the features output by each branch network to the final matching result, the weights of each feature vector in the feature fusion process are dynamically adjusted. The weight calculation is based on the Gradient-weighted Class Activation Mapping (Grad-CAM) technology. By calculating the gradient information of the feature map, the importance score of each feature is determined, where is the classification score, is the element in the feature map, N is the number of elements in the feature map, and M is the number of feature maps.

[0009] Furthermore, it also includes:

[0010] Steps for constructing a unified vehicle matching framework for multi-network fusion: Combine non-negative matrix factorization, self-supervised learning, generative adversarial network GAN, and classification network to work together; Self-supervised learning designs a rotation prediction task and a jigsaw reconstruction task. The rotation prediction task determines whether the input image has been rotated by a specific angle, and the jigsaw reconstruction task divides the image into several pieces and requires the network to restore it; The self-supervised learning loss function is, where is the loss of the rotation prediction task and is the loss of the jigsaw reconstruction task; GAN generates fake samples with a similar distribution to real minority-class samples through generator G and discriminator D. The loss function is, where is the distribution of real data, x is the real sample, is the distribution of noise data, z is the noise sample, to balance the class distribution of the dataset; To improve the quality and diversity of the generated samples, the idea of conditional generative adversarial network CGAN is introduced into the generator. The attribute information of the vehicle is used as an additional condition and input into the generator and discriminator. The input of the generator becomes, where z is the noise vector and c is the vehicle attribute vector. The discriminator not only has to judge the authenticity of the sample.

[0011] Steps for optimizing the overall loss function: The overall loss function is, where is the main loss of the classification task, is the loss of the residual network, is the loss of non-negative matrix factorization, is the loss of the generative adversarial network, and is the loss of self-supervised learning. The classification performance of the sample is improved by optimizing the comprehensive loss function; To enable the model to have better generalization ability in different datasets and application scenarios, an adaptive loss weight adjustment mechanism is introduced. According to the change trend of various losses during the training process, the weights of various losses in the overall loss function are dynamically adjusted; Specifically, the exponential weighted moving average EWMA method is used to calculate the change rate of various losses, where is the loss of the i-th type in the i-th training cycle, α is the smoothing coefficient, and the value range is. The loss weights are dynamically adjusted according to the value of, where β is the adjustment parameter.

[0012] Furthermore, when constructing the branch network, the Kaiming initialization method is used for the weight initialization of each branch network. The initialization parameters are calculated according to the number of input and output channels of the branch network. The formula is to accelerate the network convergence speed, where is the element in the network weight matrix and is the number of input channels; To improve the initialization effect of the network, based on Kaiming initialization, a channel attention mechanism is introduced. The channel attention weights are calculated on the input channels of each convolutional layer, where is the global average pooling operation on the input feature map x, and and are the weight matrices of the fully connected layers, and is the sigmoid activation function. The channel attention weights are multiplied by the initialized weight matrix to obtain the final initialized weights.

[0013] Further, during the multi-feature fusion process, the fusion features obtained by non-negative matrix factorization are normalized using the L2 normalization method. The formula is, where x is the feature vector to be normalized, and is the L2 norm of vector x, so that different features have the same scale. To enhance the expression ability of the normalized features for vehicle features, after L2 normalization, a feature non-linear transformation layer is introduced, and the Swish activation function is used to perform non-linear transformation on the normalized features. The formula is, where is the sigmoid function and β is a learnable parameter. The value of β is adaptively adjusted through training to capture the distribution of vehicle features.

[0014] Further, in the rotation prediction task of self-supervised learning, data augmentation techniques are adopted to perform random rotation, translation, and scaling operations on the input images. To enrich the data augmentation methods, a data augmentation method based on style transfer is introduced, and images of different styles are transferred to vehicle images.

[0015] Further, it also includes:

[0016] Data preprocessing step: Collect vehicle image data taken by different cameras or at different times, and perform denoising, cropping, and normalization on the images. In the denoising process, the adaptive median filtering algorithm is adopted. According to the pixel value distribution in the local area of the image, the size of the filtering window is dynamically adjusted. The formula for adjusting the filtering window size is, where and are the maximum and minimum filtering window sizes respectively, is the current filtering window size, is the window size adjustment step, is the pixel standard deviation of the local area, and is the pixel standard deviation of the global image;

[0017] Network training step: The preprocessed data is input into the constructed branch network and the unified vehicle matching framework for network fusion. To improve the training efficiency and the stability of the model, the mixed-precision training technique is adopted. During the training process, single-precision FP32 and half-precision FP16 floating-point numbers are used for calculation simultaneously, reducing memory occupancy and calculation time. At the same time, the gradient accumulation technique is introduced. The gradients are accumulated on small batches of data and then a parameter update is performed once to simulate the batch size;

[0018] Feature extraction and matching step: The vehicle image to be matched is input into the trained network. Features are extracted through the branch network, and the depth feature representation of the vehicle is obtained through feature fusion. Calculate the similarity between the features of the image to be matched and the features of the existing images in the database, including using the cosine similarity calculation. The formula is, where x and y are two feature vectors respectively, represents the dot product of vectors, and and are the L2 norms of vectors x and y respectively. Determine whether it is the same vehicle according to the similarity.

[0019] Further, in the data preprocessing step, the adaptive histogram equalization method is used to enhance the image. To avoid the over-enhancement problem of the image caused by adaptive histogram equalization, the Contrast Limited Adaptive Histogram Equalization (CLAHE) method is introduced. When performing histogram equalization on a region, the slope of the histogram is restricted to prevent excessive enhancement of local contrast. By setting the contrast limit threshold T, the histogram part exceeding the threshold is cropped and evenly distributed to other intervals.

[0020] Further, in the network training step, the early stopping method is adopted to prevent network overfitting. When the loss function on the validation set no longer decreases for 5 consecutive epochs, the training is stopped and the current optimal network model parameters are saved. To improve the generalization ability of the model, on the basis of early stopping, the model fusion technology is introduced, and the models at different training stages are saved during the training process.

[0021] Further, in the feature extraction and matching step, a threshold judgment mechanism is adopted. The similarity threshold is set to 0.8. When the calculated similarity is greater than the threshold, it is determined to be the same vehicle; otherwise, it is determined to be different vehicles. To enable the threshold to adapt to different application scenarios and data distributions, a threshold adaptive adjustment method based on Bayesian optimization is introduced. According to the historical matching data and the matching results of the validation set, the similarity threshold is continuously adjusted through the Bayesian optimization algorithm to maximize the accuracy and recall rate of the matching. The optimization objective function is, where α is the weight coefficient for balancing the accuracy and recall rate, and its value range is.

[0022] Compared with the existing technologies, the beneficial effects of the present invention are as follows:

[0023] In terms of matching accuracy, the multi-branch network design combines different loss functions and attention mechanisms, which can simultaneously extract the global and local features of the vehicle. Then, feature fusion is performed through non-negative matrix factorization, effectively enhancing the expressiveness of the features. Aiming at the problems of data scarcity and imbalance, by fusing the generative adversarial network, self-supervised network and classification network, high-quality synthetic data is generated to supplement the training set, enhancing the model's learning and generalization ability for minority class samples, thus greatly improving the matching accuracy and enabling more accurate identification of vehicles in different scenarios.

[0024] In terms of real-time performance, the number of layers of the deep network is reduced through the multi-branch network, reducing the computational complexity and improving the real-time performance of the deep network inference. At the same time, the approximation and improvement strategy for non-convex optimization, combined with ADMM and spatial structure analysis, decomposes complex problems into sub-problems that are easy to solve, supports parallel computing, significantly speeds up the optimization speed, reduces the training and inference time, and meets the real-time requirements of intelligent traffic monitoring and emergency response.

[0025] In terms of robustness, the combination of multi-branch networks and various loss functions, as well as the application of technologies such as adaptive illumination normalization, enables the model to have stronger adaptability to complex environments such as multi-camera perspective differences and illumination changes. The robustness of depth feature extraction is significantly enhanced, and vehicle features can be stably and accurately extracted even under harsh shooting conditions. In addition, the generalization ability of the model is also improved. Through technologies such as adaptive loss weight adjustment mechanism and multi-scale training, the model can maintain good performance in different datasets and application scenarios, effectively solving the problems existing in traditional methods and existing deep learning models, and providing a more reliable and efficient vehicle matching solution for the fields of intelligent transportation and public safety. Brief Description of the Drawings

[0026] Figure 1 It is a schematic block diagram of the anonymous vehicle matching method based on depth features proposed by the present invention;

[0027] Figure 2 It is a schematic block diagram of the line graph of the matching accuracy of the method proposed by the present invention under different illumination conditions. Detailed Embodiments

[0028] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0029] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention.

[0030] In addition, the terms "first" and "second" are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined. In addition, the terms "mounted", "connected", and "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. The present invention will be further described in detail below with reference to the accompanying drawings.

[0031] Referring to Figure 1 - Figure 2 : A method for anonymous vehicle matching based on deep features, comprising the following steps:

[0032] Constructing a multi-branch network to extract vehicle features: Build a plurality of branch networks, each branch based on the residual network architecture. Branch 1 uses a residual network combined with a triplet loss, Branch 2 uses a residual network combined with a cross-entropy loss, Branch 3 uses a residual network, a multi-head attention module, and a triplet loss, and Branch 4 uses a residual network, a multi-head attention module, and a cross-entropy loss. Through these branches, the global features and local features of vehicle images are extracted simultaneously. Among them, the triplet loss formula is:

[0033]

[0034] In the formula, x a is the anchor sample, x p is the positive sample (the sample belonging to the same category as the anchor sample), x n is the negative sample (the sample belonging to a different category from the anchor sample), f(·) represents the feature extraction function, and α is the set boundary value used to control the distance between the positive sample and the negative sample; the cross-entropy loss formula is where C is the number of categories, y i is the one-hot vector of the true label of the sample, and p i is the predicted probability distribution of the network for the sample belonging to the i-th category. To enhance the robustness of the network to vehicle features under different lighting conditions, an adaptive illumination normalization module is added before the input layer of each branch network. This module dynamically adjusts the brightness and contrast of the image based on the histogram statistical information of the image, and its adaptive adjustment formula is I adjusted (x,y) = γ × I(x,y) β, where I(x, y) is the pixel value of the original image at the position (x, y), and γ and β are parameters dynamically calculated based on the image histogram. During the process of constructing a multi-branch network to extract vehicle features, it should be noted that when setting the boundary value α in the triplet loss formula, the distance between positive and negative samples needs to be balanced. If α is too large, the training difficulty will increase sharply; if it is too small, the sample categories cannot be effectively distinguished. When using cross-entropy loss, it is necessary to ensure the accuracy of the one-hot vector of the true label of the sample, otherwise it will interfere with the network training. For the modules combined in each branch network, such as the multi-head attention module, the weights need to be finely adjusted to ensure reasonable attention to vehicle features. Additionally, in the adaptive illumination normalization module, the γ and β parameters are calculated based on the image histogram. It is necessary to ensure the accuracy of the histogram statistics, otherwise the adjustment of image brightness and contrast will be incorrect, thereby affecting the accuracy of vehicle feature extraction.

[0035] Multi-feature fusion: For the multiple feature vectors output by the multi-branch network, a non-negative matrix factorization model with tensor rank constraint integrating singular value significance difference information is used for fusion, and the objective function is:

[0036]

[0037] After conversion, it is equivalent to:

[0038]

[0039] Here, K represents the number of feature matrices, X k is the k-th feature matrix, A k and U k are decomposition matrices, S k is the core matrix to be decomposed, λ1 and λ2 are hyperparameters for balancing different loss terms, ||·|| F represents the Frobenius norm of the matrix, ||·||* represents the nuclear norm of the matrix, r k is the rank of S k , w i is the weight of the i-th singular value, σ i (S k ) is the i-th singular value of S k , ρ is the penalty coefficient, M k is the auxiliary matrix. To further optimize the feature fusion effect, a feature importance evaluation mechanism is introduced. According to the contribution degree of the features output by each branch network to the final matching result, the weights of each feature vector in the feature fusion process are dynamically adjusted. The weight calculation is based on the Gradient-weighted Class Activation Mapping (Grad-CAM) technology. By calculating the gradient information of the feature map, the importance score of each feature is determined where y c is the classification score, A jis an element in the feature map, N is the number of feature map elements, and M is the number of feature maps. During the multi-feature fusion process, the following points need attention: The setting of hyperparameters λ1 and λ2 is crucial as they balance different loss terms. Improper values can cause the model to overly focus on a certain loss term, affecting the feature fusion effect, and thus need to be carefully adjusted through experiments and the performance on the validation set. When calculating the feature importance scores, the gradient information calculation based on the Grad-CAM technique must be accurate, otherwise, the feature weight allocation will be unreasonable, reducing the matching accuracy. In matrix operations, it is necessary to ensure that the dimensions of each matrix match correctly to avoid calculation errors due to dimension mistakes. Additionally, after introducing the feature importance evaluation mechanism, it is necessary to regularly re-evaluate and adjust the weights according to new data and task requirements to adapt to the changing situation.

[0040] In the present invention, it further includes:

[0041] Steps for constructing a unified vehicle matching framework for multi-network fusion: Collaborate non-negative matrix factorization, self-supervised learning, generative adversarial network (GAN), and classification network. Self-supervised learning designs a rotation prediction task and a jigsaw puzzle reconstruction task. The rotation prediction task determines whether the input image has undergone a specific rotation angle, and the jigsaw puzzle reconstruction task divides the image into several pieces and requires the network to restore it. The self-supervised learning loss function is L SSL = L rotation + L puzzle where L rotation is the loss of the rotation prediction task, and L puzzle is the loss of the jigsaw puzzle reconstruction task; The GAN generates fake samples with a distribution similar to that of real minority-class samples through a generator G and a discriminator D, and the loss function is:

[0042]

[0043] where p data is the distribution of real data, x is the real sample, p z is the distribution of noise data, z is the noise sample, to balance the class distribution of the dataset. To improve the quality and diversity of the generated samples, the idea of conditional generative adversarial network (CGAN) is introduced into the generator, and partial attribute information of the vehicle (such as color, vehicle model, etc.) is used as an additional condition and input into the generator and the discriminator. The input of the generator becomes (z, c), where z is the noise vector and c is the vehicle attribute vector. The discriminator not only needs to judge the authenticity of the sample but also needs to judge whether the attributes of the sample are consistent with the input attributes, so as to generate fake samples that are more in line with the actual situation.

[0044] Steps for optimizing the overall loss function: The overall loss function is:

[0045] L = L C + L R + LNMF +L GAN +L SSL

[0046] where L C is the main loss of the classification task, L R is the loss of the residual network, L NMF is the non - negative matrix factorization loss, L GAN is the generative adversarial network loss, L SSL is the self - supervised learning loss. By optimizing this comprehensive loss function, the classification performance of minority class samples is improved. To enable the model to have better generalization ability in different datasets and application scenarios, an adaptive loss weight adjustment mechanism is introduced. According to the change trend of each type of loss during the training process, the weights of each loss in the overall loss function are dynamically adjusted. Specifically, the exponential weighted moving average (EWMA) method is used to calculate the change rate ΔL i =α××ΔL i-1 +(1 - α)×(L i -L i-1 ), where L i is the loss of the i - th class in the i - th training cycle, α is the smoothing coefficient, and its value range is (0, 1). The loss weights are dynamically adjusted according to the value of ΔL i where β is the adjustment parameter. In the present invention, when constructing the multi - branch network, the Kaiming initialization method is used for the weight initialization of each branch network. The initialization parameters are calculated according to the number of input and output channels of the branch network. The formula is

[0047] where w is an element in the network weight matrix, n i,j is the number of input channels, to accelerate the network convergence speed. To further improve the initialization effect of the network, on the basis of Kaiming initialization, a channel attention mechanism is introduced. The channel attention weight M is calculated on the input channels of each convolutional layer in =σ(F c (x)·W1·W2), where F avgpool (x) is the global average pooling operation on the input feature map x, W1 and W2 are the weight matrices of the fully - connected layers, σ is the sigmoid activation function, and the channel attention weight is multiplied by the initialized weight matrix to obtain the final initialized weight. avgpool (x) is the global average pooling operation on the input feature map x, W1 and W2 are the weight matrices of the fully - connected layers, σ is the sigmoid activation function, and the channel attention weight is multiplied by the initialized weight matrix to obtain the final initialized weight.

[0048] In the present invention, during the multi - feature fusion process, the fusion features obtained by non - negative matrix factorization are normalized. The L2 normalization method is used, and the formula is Where x is the feature vector to be normalized, and ||x||2 is the L2 norm of vector x, which makes different features have the same scale and improves the matching accuracy. To enhance the expression ability of the normalized features for vehicle features, after L2 normalization, a feature non-linear transformation layer is introduced, and the Swish activation function is used to perform non-linear transformation on the normalized features. The formula is y = x × σ(βx), where σ is the sigmoid function and β is a learnable parameter. The value of β is adaptively adjusted through training to better capture the complex distribution of vehicle features.

[0049] In the present invention, in the rotation prediction task of self-supervised learning, data augmentation techniques are adopted to perform operations such as random rotation, translation, and scaling on the input images, increasing the diversity of the data and improving the generalization ability of the model. To further enrich the data augmentation methods, a data augmentation method based on style transfer is introduced, and the styles of different images are transferred to vehicle images, such as transferring the oil painting style, watercolor painting style, etc. to vehicle images, enabling the model to learn vehicle features under different styles. At the same time, a multi-scale training strategy is adopted to train on images of different scales to improve the model's recognition ability for vehicle targets of different sizes.

[0050] The present invention further includes the following steps:

[0051] Data preprocessing step: Collect vehicle image data captured by different cameras or at different times, perform denoising, cropping, and normalization on the images, and uniformly adjust the images to the same size, such as 224×224 pixels, to make the data meet the network input requirements. In the denoising process, an adaptive median filtering algorithm is adopted. According to the pixel value distribution in the local area of the image, the size of the filtering window is dynamically adjusted to better remove noise and retain the detailed information of the image. The formula for adjusting the size of the filtering window is:

[0052]

[0053] Where, S max and S min are the maximum and minimum filtering window sizes respectively, S current is the current filtering window size, ΔS is the window size adjustment step, σ local is the pixel standard deviation of the local area, and σ global is the pixel standard deviation of the global image. When uniformly adjusting the images to a specific size such as 224×224 pixels, attention should be paid to the integrity and proportional coordination of the image content to prevent deformation of vehicle features due to excessive stretching or compression, which may affect subsequent analysis. When adopting the adaptive median filtering algorithm, the values of S max , S min and ΔS need to be reasonably set. If S max is too large, the image may be over-smoothed and key details may be lost; Smin If it is too small, the denoising effect will be poor. Improper setting of ΔS will affect the sensitivity of window size adjustment. At the same time, attention should be paid to the calculation accuracy of local and global standard deviations, which are the key basis for dynamic adjustment of window size. If the calculation is incorrect, the filtering effect will be greatly reduced.

[0054] Network training steps: Input the preprocessed data into the constructed unified vehicle matching framework of multi-branch network and multi-network fusion. According to the set learning rate, such as 0.001, use the backpropagation algorithm to combine the above loss functions for network training and continuously update the network parameters. To improve the training efficiency and model stability, adopt the mixed-precision training technology, and use single-precision (FP32) and half-precision (FP16) floating-point numbers for calculation during training to reduce memory occupancy and calculation time. At the same time, introduce the gradient accumulation technology, accumulate the gradients on multiple small batches of data and then perform a parameter update once to simulate a larger batch size and improve the model's convergence speed.

[0055] Feature extraction and matching steps: Input the vehicle image to be matched into the trained network, extract features through the multi-branch network, and obtain the depth feature representation of the vehicle through multi-feature fusion. Calculate the similarity between the features of the image to be matched and the features of the existing images in the database. For example, use the cosine similarity calculation, and the formula is where x and y are two feature vectors respectively, · represents the vector dot product, ||x||2 and ||y||2 are the L2 norms of vectors x and y respectively, and determine whether they are the same vehicle according to the similarity. To improve the accuracy and efficiency of matching, adopt a hierarchical matching strategy. First, perform a quick screening in the coarse-grained feature space to exclude samples with low similarity, and then perform precise matching on the remaining samples in the fine-grained feature space. At the same time, introduce a matching confidence evaluation mechanism, calculate the matching confidence according to factors such as the similarity of feature matching and the stability of features. Only when the matching confidence is higher than the set threshold, it is determined to be the same vehicle. When using the adaptive histogram equalization and CLAHE methods for image enhancement, attention should be paid to the reasonable setting of the contrast limit threshold T. If the threshold is too high, it cannot effectively prevent over-enhancement of the image, and if it is too low, the enhancement effect may be weakened. The appropriate threshold should be determined through experiments or experience according to the actual situation of the image, such as the complexity of the vehicle image and background interference. At the same time, continuously pay attention to the change of image details during the processing process to avoid losing key features due to over-processing and affecting the accuracy of subsequent feature extraction and vehicle matching.

[0056] In the present invention, in the data preprocessing step, the adaptive histogram equalization method is adopted to enhance the image, improve the contrast of the image, make the features of the vehicle more obvious, and facilitate subsequent feature extraction. To avoid the problem of over-enhancement of the image that may be caused by adaptive histogram equalization, the contrast-limited adaptive histogram equalization (CLAHE) method is introduced. When performing histogram equalization in a local area, the slope of the histogram is limited to prevent excessive enhancement of the local contrast. By setting the contrast limit threshold T, the histogram part exceeding the threshold is cropped and evenly distributed to other intervals to obtain a more natural image enhancement effect. When using the early stopping method and the model fusion technology, it is necessary to accurately monitor the change of the loss function on the validation set and strictly execute early stopping according to the standard that it no longer decreases for 5 consecutive epochs to prevent premature or late stopping of training. For model fusion, the selection of models at different training stages and the determination of fusion weights are crucial. It is necessary to reasonably allocate weights according to the performance of the models on the validation set. For example, a higher weight is given to a model with more stable performance. In addition, the calculation amount of model fusion is relatively large, and attention should be paid to the control of computing resources and time costs.

[0057] In the present invention, in the network training step, the early stopping method is adopted to prevent network overfitting. When the loss function on the validation set no longer decreases for 5 consecutive epochs, the training is stopped and the current optimal network model parameters are saved. To further improve the generalization ability of the model, on the basis of early stopping, the model fusion technology is introduced. During the training process, multiple models at different training stages are saved, and the prediction results of these models are fused, such as using the voting method or the weighted average method, to reduce the overfitting risk of a single model and improve the performance of the overall model.

[0058] In the present invention, in the feature extraction and matching step, a threshold judgment mechanism is adopted, and the similarity threshold is set to 0.8. When the calculated similarity is greater than the threshold, it is determined to be the same vehicle; otherwise, it is determined to be different vehicles, which improves the accuracy and efficiency of matching. To enable the threshold to adapt to different application scenarios and data distributions, a threshold self - adaptive adjustment method based on Bayesian optimization is introduced. According to the historical matching data and the matching results of the validation set, the similarity threshold is continuously adjusted through the Bayesian optimization algorithm to maximize the matching accuracy and recall rate. The optimization objective function is F = α×Precision+(1 - α)×Recall, where α is the weight coefficient for balancing accuracy and recall rate, and its value range is (0, 1). When using the threshold judgment mechanism and Bayesian optimization to adjust the similarity threshold, although the initial threshold of 0.8 is an empirical value, it needs to be flexibly adjusted in actual applications. After introducing the Bayesian optimization algorithm, it is necessary to ensure the accuracy and representativeness of the historical matching data and the validation set data, otherwise it will mislead the direction of threshold adjustment. The value of the weight coefficient α needs to be combined with the specific application scenario. If more emphasis is placed on accuracy, α can be appropriately close to 1; if more emphasis is placed on recall rate, α is closer to 0. At the same time, the effect of threshold adjustment needs to be evaluated regularly to avoid the decline of the model's generalization ability caused by over - optimization.

[0059] Data representation of beneficial effects:

[0060] 1. Improvement in matching accuracy

[0061] Method Accuracy Recall F1 - value Traditional method 70% 65% 67.4% Existing deep learning method 80% 75% 77.4% Method of this patent 90% 88% 89%

[0062] It can be seen from the tabular data that the method of this patent is significantly higher than the traditional method and the existing deep - learning methods in terms of accuracy, recall rate, and F1 - value, indicating a significant improvement in its matching accuracy.

[0063] 2. Improvement in real - time performance

[0064] Method Inference time for a single image (ms) Traditional method 150 Existing deep learning method 100 Method of this patent 50

[0065] The inference time of a single image of the method of this patent is significantly shortened, and the real - time performance is significantly improved, which can better meet the requirements of intelligent transportation monitoring and emergency response.

[0066] 3. Improvement in robustness

[0067] Method Accuracy under different lighting conditions Accuracy under different perspectives Traditional method 60% 55% Existing deep learning method 70% 65% Method of this patent 85% 80%

[0068] Under different lighting conditions and different perspectives, the accuracy of the method of this patent is higher than that of the traditional method and the existing deep - learning methods, indicating that it has stronger robustness.

[0069] In summary, the anonymous vehicle matching method based on deep features of this patent has achieved remarkable improvements in terms of matching accuracy, real-time performance, and robustness through a series of innovative technologies and optimization strategies, and has good application prospects.

[0070] The above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.

Claims

1. A method for anonymous vehicle matching based on deep features, characterized in that, Including: Constructing a multi-branch network to extract vehicle features: Building a branch network, with each branch based on a residual network as the infrastructure; Branch 1 uses a residual network combined with triplet loss, Branch 2 uses a residual network combined with cross-entropy loss, Branch 3 uses a residual network, an attention module, and triplet loss, and Branch 4 uses a residual network, an attention module, and cross-entropy loss; Extracting the features of the vehicle image through the branches simultaneously; Among them, the triplet loss formula is: where x a is the anchor sample, x p is the positive sample, x n is the negative sample, f(·) represents the feature extraction function, and α is the set boundary value used to control the distance between the positive and negative samples; the cross-entropy loss formula is where C is the number of classes, and y i is the one-hot vector of the true label of the sample, and p i is the predicted probability distribution of the network for the sample belonging to the i-th class; to enhance the robustness of the network to vehicle features under different lighting conditions, an adaptive illumination normalization module is added before the input layer of each branch network. The module dynamically adjusts the brightness and contrast of the image based on the histogram statistical information of the image. Its adaptive adjustment formula is I adjusted (x, y) = γ × I(x, y) β , where I(x, y) is the pixel value of the original image at the position (x, y), and γ and β are parameters dynamically calculated according to the image histogram; Multi-feature fusion: For the feature vectors output by the branch network, a non-negative matrix factorization model with tensor rank constraints integrating singular value difference information is used for fusion, and the objective function is: After conversion, it is equivalent to: where K represents the number of feature matrices, and X k is the k-th feature matrix, A k and U k are decomposition matrices, S k is the core matrix to be decomposed, λ1 and λ2 are hyperparameters for balancing different loss terms, ||·|| F represents the Frobenius norm of a matrix, ||·|| * represents the nuclear norm of a matrix, r k is the rank of S k , ω i is the weight of the i-th singular value, σ i (S k ) is the i-th singular value of S k , ρ is the penalty coefficient, and M k is the auxiliary matrix; to optimize the feature fusion effect, a feature importance evaluation mechanism is introduced. According to the contribution degree of the output features of each branch network to the final matching result, the weights of each feature vector in the feature fusion process are dynamically adjusted. The weight calculation is based on the Gradient-weighted Class Activation Mapping (Grad-CAM) technique. By calculating the gradient information of the feature map, the importance score of each feature is determined where y c is the classification score, A j is an element in the feature map, N is the number of elements in the feature map, and M is the number of feature maps.

2. The anonymous vehicle matching method based on depth features according to claim 1, wherein It also includes: Steps for constructing a unified vehicle matching framework for multi-network fusion: Combining non-negative matrix factorization, self-supervised learning, generative adversarial network GAN, and classification network to work together; Self-supervised learning designs a rotation prediction task and a jigsaw reconstruction task. The rotation prediction task determines whether the input image has undergone a specific rotation angle, and the jigsaw reconstruction task divides the image into several pieces and requires the network to restore it; the self-supervised learning loss function is L SSL = L rotation + L puzzle , where L rotation is the loss of the rotation prediction task, and L puzzle is the loss of the jigsaw reconstruction task; GAN generates fake samples with a similar distribution to real minority-class samples through a generator G and a discriminator D, and the loss function is: where p data is the distribution of real data, x is the real sample, p z is the distribution of noise data, z is the noise sample, balancing the class distribution of the dataset; to improve the quality of the generated samples, the idea of conditional generative adversarial network CGAN is introduced into the generator, and the attribute information of the vehicle is used as an additional condition and input into the generator and discriminator. The input of the generator becomes (z, c), where z is the noise vector and c is the vehicle attribute vector.

3. The method for anonymous vehicle matching based on depth features according to claim 2, characterized in that It also includes: Steps for optimizing the overall loss function: The overall loss function is: L = L C + L R + L NMF + L GAN + L SSL where L C is the main loss of the classification task, L R is the loss of the residual network, L NMF is the non - negative matrix factorization loss, L GAN is the loss of the generative adversarial network, L SSL is the self - supervised learning loss. The classification performance of the samples is improved by optimizing the comprehensive loss function; to enable the model to have better generalization ability in different datasets and application scenarios, an adaptive loss weight adjustment mechanism is introduced. According to the change trend of various losses during the training process, the weights of each loss in the overall loss function are dynamically adjusted; specifically, the exponential weighted moving average EWMA method is used to calculate the change rate ΔL i =α×ΔL i-1 +(1 - α)×(L i -L i-1 ), where L i is the loss of the i - th class in the i - th training cycle, α is the smoothing coefficient, and its value range is (0, 1). The loss weight is dynamically adjusted according to the value of ΔL i where β is the adjustment parameter.​ 4. The method for anonymous vehicle matching based on depth features according to claim 1, wherein When constructing the branch network, the Kaiming initialization method is used for the weight initialization of each branch network. The initialization parameters are calculated according to the number of input and output channels of the branch network, and the formula is to accelerate the network convergence speed, where w i,j is an element in the network weight matrix, and n in is the number of input channels; to improve the initialization effect of the network, on the basis of Kaiming initialization, a channel attention mechanism is introduced, and the channel attention weight M is calculated on the input channels of each convolutional layer c =σ(F avgpool (x)·W1·W2), where F avgpool (x) is the global average pooling operation on the input feature map x, W1 and W2 are the weight matrices of the fully connected layers, σ is the sigmoid activation function, and the channel attention weight is multiplied by the initialized weight matrix to obtain the final initialized weight.

5. The anonymous vehicle matching method based on deep features according to claim 1, wherein During the multi-feature fusion process, the fused features obtained by non-negative matrix factorization are normalized using the L2 normalization method. The formula is where x is the feature vector to be normalized, and ||x||2 is the L2 norm of the vector x, so that different features have the same scale. To enhance the expression ability of the normalized features for vehicle features, after L2 normalization, a feature non-linear transformation layer is introduced, and the Swish activation function is used to perform non-linear transformation on the normalized features. The formula is y = x × σ(βx), where σ is the sigmoid function and β is a learnable parameter. The value of β is adaptively adjusted through training to capture the distribution of vehicle features.

6. The anonymous vehicle matching method based on depth features according to claim 2, characterized in that In the rotation prediction task of self-supervised learning, data augmentation techniques are used to perform random rotation, translation, and scaling operations on the input image, and a data augmentation method based on style transfer is introduced to transfer images of different styles to the vehicle image.

7. The anonymous vehicle matching method based on depth features according to claim 1, wherein It also includes: Data preprocessing steps: Collecting vehicle image data taken by different cameras or at different times, and performing denoising, cropping, and normalization on the images. In the denoising process, an adaptive median filtering algorithm is used to dynamically adjust the size of the filtering window according to the pixel value distribution in the local area of the image. The formula for adjusting the size of the filtering window is: Among them, S max and S min are the maximum and minimum filter window sizes respectively, S current is the current filter window size, ΔS is the window size adjustment step, and σ local is the pixel standard deviation of the local area, and σ global is the pixel standard deviation of the global image; Network training steps: The preprocessed data is input into the constructed branch network and the unified vehicle matching framework for multi-network fusion. To improve the training efficiency and the stability of the model, mixed-precision training technology is used, and single-precision FP32 and half-precision FP16 floating-point numbers are used for calculation simultaneously during the training process to reduce memory occupancy and calculation time. At the same time, gradient accumulation technology is introduced to accumulate gradients on small batches of data and then perform a parameter update once to simulate the batch size; Feature extraction and matching steps: The vehicle image to be matched is input into the trained network, features are extracted through the branch network, and the depth feature representation of the vehicle is obtained through feature fusion. Calculate the similarity between the features of the image to be matched and the features of the existing images in the database, including using cosine similarity for calculation, and the formula is: Where x and y are two feature vectors respectively, · represents the vector dot product, ||x||2 and ||y||2 are the L2 norms of vectors x and y respectively, and it is judged whether they are the same vehicle according to the similarity.

8. The anonymous vehicle matching method based on depth features according to claim 7, wherein, It also includes: In the data preprocessing step, an adaptive histogram equalization method is used to enhance the image. To avoid the over-enhancement problem of the image caused by adaptive histogram equalization, the contrast-limited adaptive histogram equalization CLAHE method is introduced. When performing histogram equalization in the region, the slope of the histogram is limited to prevent excessive enhancement of local contrast. By setting the contrast limit threshold T, the part of the histogram exceeding the threshold is cropped and evenly distributed to other intervals.

9. The method for anonymously matching vehicles based on depth features according to claim 7, wherein In the network training step, the early stopping method is adopted to prevent the network from overfitting. When the loss function on the validation set does not decrease for 5 consecutive epochs, the training is stopped and the current optimal network model parameters are saved. To improve the generalization ability of the model, on the basis of early stopping, the model fusion technology is introduced, and the models at different training stages are saved during the training process.

10. The method for anonymous vehicle matching based on depth features according to claim 7, wherein In the feature extraction and matching step, a threshold judgment mechanism is adopted. The similarity threshold is set to 0.

8. When the calculated similarity is greater than the threshold, it is determined to be the same vehicle, otherwise it is determined to be different vehicles. To make the threshold adapt to different application scenarios and data distributions, a threshold adaptive adjustment method based on Bayesian optimization is introduced. According to the historical matching data and the matching results of the validation set, the similarity threshold is continuously adjusted through the Bayesian optimization algorithm to maximize the matching accuracy and recall rate. The optimization objective function is: F = α × Precision+(1 - α) × Recall where α is the weight coefficient for balancing the accuracy and recall rate, and its value range is (0, 1).

Citation Information

Cited By

  • Bullet train head recognition system based on fine-grained target detection

    CN121170268A