Vehicle re-identification model training method and device and vehicle re-identification method

By embedding non-local blocks and super-resolution technologies in the vehicle recognition model, global and local features are extracted and loss function optimization is performed, the problem of view angle changes in vehicle recognition is solved, and the recognition accuracy and robustness are improved.

CN120544003APending Publication Date: 2025-08-26SHANGHAI VCS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510623348.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the prior art, vehicle re-identification research has insufficient ability to extract discriminative features and cannot adapt to actual traffic scenarios where vehicle perspectives change drastically.

Method used

By embedding non-local blocks in the first network model to obtain the global features of the vehicle, and using the second network model to extract local features from different perspectives, combining super-resolution technology to enhance features, calculate the loss functions of global and local features for optimization, and build a vehicle re-identification model.

Benefits of technology

The vehicle re-identification model's local feature extraction capability and robustness for different vehicle perspectives is improved, the recognition accuracy is improved, and computing time and resources are saved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The invention discloses a training method for a vehicle re-identification model, and the method comprises the steps: inputting a randomly selected target vehicle image into a first network model and a second network model, and obtaining corresponding vehicle global features and vehicle local features at different visual angles; performing feature enhancement on the vehicle local features to obtain fine local features of the different visual angles; calculating a first loss function corresponding to the vehicle global feature and a second loss function corresponding to the different visual angles; and optimizing the first network model and the second network model based on the first loss function and the second loss function to obtain a vehicle re-identification model. Meanwhile, the invention further discloses a vehicle re-identification model device and a vehicle re-identification method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of vehicle re-identification, and in particular to a vehicle re-identification model training method, device and vehicle re-identification method. Background Art

[0002] In recent years, with the development of society and the improvement of living standards, cars have become an indispensable means of transportation for people, and the automobile and related industries have experienced rapid development in my country. This influx of vehicles has posed significant challenges to urban management. Illegal modifications, license plate obstructions, and other crimes have caused significant harm to society, seriously disrupting traffic safety and public order. However, current research on vehicle re-identification generally suffers from insufficient ability to extract discriminative features, resulting in an inability to adapt to the drastic changes in vehicle perspectives in real-world traffic scenarios. Summary of the Invention

[0003] In view of this, the embodiments of the present application hope to provide a vehicle re-identification model training method, device and vehicle re-identification method to at least solve the above technical problems.

[0004] To achieve the above objectives, the technical solution of this application is implemented as follows:

[0005] According to one aspect of an embodiment of the present application, a method for training a vehicle re-identification model is provided, the method comprising:

[0006] Input the randomly selected target vehicle image into the first network model and the second network model respectively to obtain the corresponding vehicle global features and vehicle local features from different perspectives;

[0007] Performing feature enhancement on the local features of the vehicle to obtain fine local features of the different perspectives;

[0008] Calculating a first loss function corresponding to the global feature of the vehicle and a second loss function corresponding to the different perspectives;

[0009] The first network model and the second network model are optimized based on the first loss function and the second loss function to obtain a vehicle re-identification model.

[0010] In the above solution, obtaining the corresponding vehicle global features includes:

[0011] Using the non-local blocks embedded in the first network model, capturing long-range dependencies between features in the target vehicle image in a preset image library;

[0012] The corresponding vehicle global features are obtained based on the long-distance dependency.

[0013] In the above solution, obtaining the local features of the vehicle from different perspectives includes:

[0014] Performing image segmentation on the target vehicle image using the second network model to obtain multiple perspective images of the target vehicle image;

[0015] Extracting multiple perspective features corresponding to the multiple perspective images from different data sets respectively;

[0016] The multiple viewing angle features are determined as vehicle local features of the target vehicle image.

[0017] In the above solution, the feature enhancement of the local features of the vehicle to obtain the fine local features of different viewing angles includes:

[0018] Using super-resolution technology to enhance all features of the vehicle local features to obtain fine local features of different perspectives;

[0019] Alternatively, a super-resolution technique is used to enhance the first target feature among the local features of the vehicle that meets the preset low-resolution requirement, thereby obtaining the fine local features at different viewing angles;

[0020] Alternatively, super-resolution technology is used to enhance the first target feature among the local features of the vehicle that meets the preset low-resolution requirements, and the enhanced first target feature is fused with the second target feature among the local features of the vehicle that meets the preset high-resolution requirements to obtain the fine local features of the different perspectives.

[0021] In the above solution, the step of calculating the first loss function corresponding to the global feature of the vehicle includes:

[0022] Calculating an ID loss function corresponding to the vehicle global feature based on a feature distance between the target vehicle image and each vehicle image;

[0023] Calculating a triplet loss function corresponding to the vehicle global feature based on the feature distance between the target vehicle image and different vehicle images;

[0024] Calculating a central loss function corresponding to the vehicle global feature based on the number of vehicle image samples selected in each iterative training of the first network model;

[0025] The ID loss function, the triplet loss function and the center loss function are used as the first loss function corresponding to the vehicle global feature.

[0026] In the above solution, the second loss functions corresponding to different perspectives include:

[0027] Performing image alignment on the fine local features of different viewing angles with the viewing angle image features in the vehicle image sample set, and calculating the discriminant score of the fine local features corresponding to each viewing angle;

[0028] A second loss function corresponding to the different perspectives is determined based on the discriminant score, where the second loss function represents a minimized loss function.

[0029] According to another aspect of the present application, a vehicle re-identification method is provided, the method comprising:

[0030] Obtaining an image of a vehicle to be identified;

[0031] Inputting the vehicle image to be identified into a vehicle re-identification model to obtain a corresponding vehicle target image;

[0032] The vehicle re-identification model is obtained by using any of the above vehicle re-identification model training methods.

[0033] In the above solution, the step of inputting the to-be-identified vehicle image into the vehicle re-identification model to obtain the corresponding vehicle target image includes:

[0034] Inputting the vehicle image to be identified into a vehicle re-identification model to obtain a local feature distance matrix and a global feature distance matrix of the vehicle image to be identified;

[0035] The local feature distance matrix and the global feature distance matrix are fused to obtain a vehicle target image list corresponding to the vehicle image to be identified, wherein the vehicle target image list includes at least one vehicle target image.

[0036] According to a third aspect of the present application, a training device for a vehicle re-identification model is provided, the device comprising:

[0037] A data acquisition module is used to input randomly selected target vehicle images into the first network model and the second network model respectively to obtain the corresponding vehicle global features and vehicle local features from different perspectives;

[0038] A super-resolution enhancement module, configured to enhance the local features of the vehicle to obtain fine local features at different viewing angles;

[0039] A calculation module, configured to calculate a first loss function corresponding to the global feature of the vehicle and a second loss function corresponding to the different perspectives;

[0040] A joint optimization module is used to optimize the first network model and the second network model based on the first loss function and the second loss function to obtain a vehicle re-identification model.

[0041] According to a fourth aspect of the present application, a training device for a vehicle re-identification model is provided, the device comprising:

[0042] a memory for storing computer programs capable of running on the processor;

[0043] The processor is configured to execute any one of the above-mentioned vehicle re-identification model training methods when running the computer program.

[0044] The present application provides a vehicle re-identification model training method, device, and vehicle re-identification method, which obtains global and local features of a vehicle through global branches and local branches, and enhances the local features of the vehicle at different perspectives through super-resolution variation enhancement technology, and fuses the global features of the vehicle with the enhanced fine local features to achieve modeling of the vehicle re-identification model. In this way, it can not only save computing time and resources, but also effectively improve the model's ability to extract local features and robustness of different vehicle perspectives without the need to annotate key points of vehicle images in advance, thereby improving the recognition accuracy of vehicle re-identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 A schematic diagram of the process implementation of the vehicle re-identification model training method in this application;

[0046] Figure 2 This is a schematic diagram of the network structure for the vehicle re-identification model training in this application;

[0047] Figure 3 This is a schematic diagram of the process implementation of the vehicle re-identification method in this application;

[0048] Figure 4 Schematic diagram of the structure of the training device for the vehicle re-identification model in this application;

[0049] Figure 5 This is the model training device for vehicle re-identification in this application. DETAILED DESCRIPTION

[0050] The technical solution of the present application is further elaborated in detail below with reference to the accompanying drawings and specific embodiments.

[0051] The various specific technical features in the various embodiments described in the specific implementation methods can be combined in various ways without contradiction. For example, different implementation methods can be formed by combining different specific technical features. In order to avoid unnecessary repetition, the various possible combinations of the specific technical features in this application will not be described separately.

[0052] It should be noted that the terms "first, second, and third" used in the embodiments of the present application are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understood that the terms "first, second, and third" may be interchanged to represent a specific order or precedence, where permitted. It should be understood that the terms "first, second, and third" may be interchanged to represent objects, where appropriate, such that the embodiments of the present application described herein may be implemented in an order other than that illustrated or described herein.

[0053] Figure 1 This is a schematic diagram of the process implementation of the vehicle re-identification model training method in this application. The method can be applied to electronic devices that implement vehicle re-identification model training, such as Figure 1 As shown, the method includes:

[0054] Step 101: Input a randomly selected target vehicle image into a first network model and a second network model respectively to obtain corresponding vehicle global features and vehicle local features at different perspectives;

[0055] Here, the first network model can be a ResNet network structure, such as a ResNet50 neural network model. Non-local blocks are embedded in each stage of the first network model. When a randomly selected target vehicle image is input into the first network model, the non-local blocks can be used to capture the long-range dependencies between various features in the target vehicle image in a preset image library. Based on these long-range dependencies, the corresponding global vehicle features can be obtained.

[0056] The operation definition of the non-local block in the first network model is shown in formula (2.1):

[0057]

[0058] Here, i is the index of the target response output location, that is, the pixel coordinate of the current location, and j is the index of all possible locations, that is, the predicted pixel coordinates associated with the current location. x is the input image feature, and y is the output feature of the same size as x. Function f is used to calculate the distance between the current location i and all other locations j, that is, to traverse all points by calculating the relative position information of all pixel coordinates. Function g is used to calculate the representation of the input feature at location j, and C(x) is the normalization factor.

[0059] Here, we can also introduce a Gaussian function to calculate the similarity of the non-local block embedding space in the first network model, as shown in Formula 2.2:

[0060]

[0061] Among them, θ(x i )=Wθ x i and φ(x j )=W φ x j are two embedded features, and the embedded features refer to the one-dimensional vectors extracted by the network model, where W refers to the learned network model parameters. And set

[0062] The activation function of the network model is shown in Formula 2.3. Its function is to convert a real number vector into a probability distribution so that the output value is between 0 and 1 and the sum of all output values ​​is 1.

[0063]

[0064] Finally, the non-local block Non_local Block is defined as:

[0065] z i =W z y i +x i (2.4)

[0066] Among them, W Z is the parameter learned by the network model. i From formula (2.1), we can conclude that “+x i " represents the residual connection. Through the residual connection, the Non_local Block can be embedded into any stage of the deep learning network without changing the size of the output features.

[0067] Figure 2 This is a diagram of the network structure for training the vehicle re-identification model used in this application. Generally speaking, embedding Non-local Blocks in shallow layers of the network can capture more detailed features with long-range correlations, but doing so increases the number of parameters required during training and consumes computing resources. Therefore, Non-local Blocks can be embedded after Stage 2 and Stage 3 of the ResNet50 backbone network. The output feature dimensions of Stage 2 are 512×28×28, and those of Stage 3 are 1024×14×14.

[0068] By embedding non-local blocks in the first network model, the present application can not only make up for the inherent deficiencies of the convolution operation of the first network model, but also improve the vehicle re-identification accuracy.

[0069] In the present application, the second network model can be a U-Net network structure. The second network model can be used to first perform image segmentation on the target vehicle image to obtain multiple perspective images of the target vehicle image; then, multiple perspective features corresponding to the multiple perspective images are extracted from different data sets such as the VeRi-776 data set, the VehicleID data set, and the VRAI data set; and the multiple perspective features are determined as the vehicle local features of the target vehicle image.

[0070] Here, the multiple perspective images can be the front, back, side, top and background of the vehicle, where the front represents the front side of the vehicle, the back represents the rear side of the vehicle, the side represents the left door side or the right door side of the vehicle, the top represents the roof side of the vehicle, and the background represents the background of the vehicle.

[0071] The images in the VeRi-776 dataset were captured in real-world, unconstrained surveillance scenarios and annotated with various attributes such as bounding box, type, color, and brand. Each vehicle was captured by 2 to 18 cameras under different viewpoints, lighting, resolution, and occlusion conditions, providing rich spatiotemporal information.

[0072] The images in the VehicleID dataset are captured by multiple real surveillance cameras distributed under different lighting conditions during the day.

[0073] The images in the VRAI dataset are taken by aircraft rather than by fixed-position traffic cameras, so they contain a large amount of drastically changing perspective information. In order to improve the generalization ability of various views, this application can annotate images in the VRAI dataset that represent more perspectives, randomly select a portion of the annotated images to train the U-Net network, and then use the trained model on another portion of the annotated images for verification, so as to obtain images with different perspectives that meet the preset image segmentation.

[0074] It should be noted that what is finally obtained here are the local features of the four perspective images: front, back, side, and top. Since the image background is interference information, it is discarded in the final stage of model training.

[0075] Here, in order to embed the mask obtained from the U-Net segmentation network into the vehicle global feature, the maximum pooling method can be used to pool it into a size of 16×16 and defined as {M i |i∈{1,2,3,4}}, and then mapped to the vehicle global feature F through mask average pooling to calculate the four perspective local features of the vehicle {f l i |i∈{0,1,2,3}}, which represent the front, rear, side and top view of the vehicle respectively. li The calculation method is:

[0076]

[0077] This application can achieve accurate segmentation of different types of vehicles by extracting the view features of the target vehicle on the VeRi-776 dataset, VehicleID dataset and VRAI dataset respectively.

[0078] Step 102: Enhance the local features of the vehicle to obtain fine local features of different viewing angles.

[0079] Here, due to the complexity of actual traffic conditions, the collected vehicle images generally have problems such as low resolution, and a large amount of discriminative detail information is lost, so it is very necessary to perform super-resolution enhancement on low-resolution perspective images. The present application can use super-resolution technology to enhance all features in the local features of the vehicle to obtain fine local features of different perspectives; or, use super-resolution technology to enhance the first target feature in the local features of the vehicle that meets the preset low-resolution requirements to obtain fine local features of different perspectives; or, use super-resolution technology to enhance the first target feature in the local features of the vehicle that meets the preset low-resolution requirements, and perform feature fusion of the enhanced first target feature with the second target feature in the local features of the vehicle that meets the preset high-resolution requirements to obtain fine local features of different perspectives.

[0080] Since high-frequency details of the same image at different resolutions are similar, residual connections can be used to learn the correlation between high-resolution and low-resolution images. This involves obtaining shallow features from the vehicle's local features, extracting high-frequency features from these shallow features, fusing these shallow features with the high-frequency features to generate a fused feature, and enhancing the fused feature to output a refined local feature.

[0081] like Figure 2 As shown, when the second network model is used to obtain the local features of the target vehicle image at different perspectives, the first convolution layer is used to obtain the local features of the perspective at low resolution. LRThe shallow feature F0 is then extracted through the subsequent RIR module to extract the high-frequency features in F0. In addition, the low-frequency redundant information contained in F0 is spliced ​​together with the output of RIR through a long skip connection. Finally, the fused features are amplified using an upsampling operation, and the enhanced high-resolution perspective image, also known as fine local features, is output after the second convolutional layer of the second network model. In this way, the low-resolution problem of the original vehicle image can be effectively solved. Compared with the original image in the dataset, it can not only eliminate the interference caused by the vehicle background, but also perform image enhancement on local features. Compared with image enhancement on global features, it effectively saves computing time and resources.

[0082] Step 103, calculating a first loss function corresponding to the vehicle global feature and a second loss function corresponding to the different perspectives;

[0083] Here, the ID loss function corresponding to the vehicle global feature can be calculated based on the feature distance between the identity label of the target vehicle image and the identity label of each vehicle image; the triplet loss function corresponding to the vehicle global feature can be calculated based on the feature distance between the target vehicle image and different vehicle images; the center loss function corresponding to the vehicle global feature can be calculated based on the number of vehicle image samples selected for each iterative training of the first network model; the ID loss function, the triplet loss function and the center loss function can be used as the first loss function corresponding to the vehicle global feature.

[0084] Among them, in order to prevent overfitting of vehicle global features, label smoothing can be introduced into the ID loss function. For the input target vehicle image, the ID loss function is formally equivalent to the cross entropy loss function and can be calculated by the following formula after introducing label smoothing:

[0085]

[0086] Among them, y is the true label of the image, which refers to the real ID of the input target vehicle, which can also be understood as the correct answer. The purpose of the model is to predict the real ID of a given image through image information. i is the predicted probability of the i-th class, that is, the probability that the predicted target vehicle is the i-th ID. N is the total number of vehicles, that is, the total number of images input to the model. ε is a small constant that allows for fewer constraints on the global vehicle features during training.

[0087] The triplet loss function can make the image features of vehicles belonging to the same ID closer and the image features of vehicles belonging to different IDs farther apart. Therefore, introducing it into the global vehicle features can effectively solve the common problems of inter-class similarity and intra-class deviation in vehicle re-identification.

[0088] The triplet loss is calculated as follows:

[0089] L Tri =max(D ap -D an +γ,0) (2.7)

[0090] Among them, D ap It means<anchor,positive> The Euclidean distance between image pairs, and D an It means<anchor,negative> The Euclidean distance between image pairs. The threshold γ represents the distance between the two, that is, the distance between the positive image and the negative image. It is a variable parameter. The triplet loss function aims to optimize so that the distance between anchors of the same type and the positive image decreases, while the distance between anchors of different types and the negative image increases.

[0091] However, the triplet loss function only considers the relative distance between three samples and ignores D ap With D an In order to minimize the absolute distance between image pairs and enhance the compactness of the instance, the center loss function is introduced here:

[0092]

[0093] Where B is the number of vehicle image samples selected for each iteration of training of the first network model, c j is the feature center of the image that belongs to the same ID as the j-th image, f j is the embedding feature of the j-th image.

[0094] In the present application, when obtaining fine local features of different perspectives, the fine local features of different perspectives can be first aligned with the perspective image features in the vehicle image sample set, and the discriminant score of the fine local features corresponding to each perspective can be calculated; based on the discriminant score, the second loss function corresponding to the different perspectives is determined, and the second loss function can represent the minimization loss function.

[0095] Here we use P i Denotes the discriminative local features corresponding to the fine local features detected at each viewpoint in the query image, i∈{1,2,…,d} denotes the number of discriminative local features. For different parts P from different images, m and P n , respectively, using l m and l nRepresents the feature matrix of the two, and uses Euclidean distance as the metric function to calculate their feature distance. m and P n It represents the discriminative local features corresponding to the fine local features detected by the mth part and the nth part at each viewing angle. The discriminant score of each part is calculated as follows:

[0096]

[0097] Among them, h and w represent the height and width of the fine local features detected at different viewing angles, respectively. The fine local features P can be eliminated by minimizing the discriminant score. i Finally, the minimization loss of fine local features is as follows:

[0098]

[0099] Among them, k i Represents the fine local features P at each viewpoint in the query image i The number of correct matches is denoted by k, ignoring the number of incorrect matches. Considering the different numbers of discriminative local features from different perspectives of a specific vehicle, Dis is divided by k to normalize the result to avoid deviations caused by the different number of occurrences of each local feature.

[0100] Here, the discriminant score of each view is calculated through a local collaboration mechanism, and a new discriminant loss is designed to minimize the distance of the local feature matrix.

[0101] Step 104 : Optimize the first network model and the second network model based on the first loss function and the second loss function to obtain a vehicle re-identification model.

[0102] Here, the first loss function and the second loss function may be added together to obtain a joint loss function, and the first network model and the second network model may be optimized based on the joint loss function to obtain a vehicle re-identification model.

[0103] The joint loss function is:

[0104]

[0105] During model training, the loss function continuously decreases. When the loss function reaches a point where there is no further room for improvement, it is considered optimal. The optimization process for L involves selecting an appropriate optimizer, adjusting the learning rate, and utilizing regularization, data augmentation, and other methods to reduce the loss function and optimize the model.

[0106] This application embeds non-local blocks in the first network model to obtain global features of vehicle images, and uses the second network model to obtain local features of vehicle images. The local features are then subjected to super-resolution enhancement, perspective alignment, and optimized discriminant scores. Ultimately, the second network model is made more sensitive to fine local features, which not only improves the model's generalization performance across different datasets but also adapts to dramatic changes in vehicle perspectives in different scenarios, improving the recognition accuracy of vehicle re-identification. By continuously learning parameters, the model reduces the loss function, thereby improving the model's fit, prediction accuracy, and generalization ability.

[0107] Figure 3 This is a schematic diagram of the process implementation of the vehicle re-identification method in this application, such as Figure 3 As shown, the method includes:

[0108] Step 301: Acquire an image of a vehicle to be identified;

[0109] Here, the vehicle image to be identified is a vehicle image to be identified that is collected in real time in a real environment.

[0110] Step 302: input the to-be-identified vehicle image into a vehicle re-identification model to obtain a corresponding vehicle target image;

[0111] The vehicle re-identification model is obtained using the aforementioned vehicle re-identification model training method. By inputting the vehicle image to be identified into the vehicle re-identification model, a local feature distance matrix and a global feature distance matrix for the vehicle image to be identified are obtained. The local feature distance matrix and the global feature distance matrix are weightedly fused to obtain a list of vehicle target images corresponding to the vehicle image to be identified, wherein the list of vehicle target images includes at least one vehicle target image.

[0112] Here, the vehicle target image list is represented by the local distance matrix M Pi and the global distance matrix M g After fusion, we get:

[0113]

[0114] The vehicle re-identification method provided in this application re-identifies the target vehicle through the vehicle re-identification model trained in this application, which can not only improve the recognition efficiency but also provide recognition accuracy under different perspectives.

[0115] Figure 4 This is a schematic diagram of the structure of the training device for the vehicle re-identification model in this application, as shown in FIG. Figure 4 As shown, the device includes:

[0116] The data acquisition module 401 is used to input the randomly selected target vehicle image into the first network model and the second network model respectively to obtain the corresponding vehicle global features and vehicle local features from different perspectives;

[0117] A super-resolution enhancement module 402 is configured to enhance the local features of the vehicle to obtain fine local features at different viewing angles;

[0118] A calculation module 403 is configured to calculate a first loss function corresponding to the vehicle global feature and a second loss function corresponding to the different perspectives;

[0119] The joint optimization module 404 is configured to optimize the first network model and the second network model based on the first loss function and the second loss function to obtain a vehicle re-identification model.

[0120] Here, the data acquisition module 401 can specifically utilize the non-local blocks embedded in the first network model to capture the long-distance dependencies between the features in the target vehicle image in the preset image library; and obtain the corresponding vehicle global features based on the long-distance dependencies.

[0121] Here, the data acquisition module 401 can also specifically use the second network model to perform image segmentation on the target vehicle image to obtain multiple perspective images of the target vehicle image; extract multiple perspective features corresponding to the multiple perspective images from different data sets respectively; and determine the multiple perspective features as vehicle local features of the target vehicle image.

[0122] Here, the super-resolution enhancement module 402 can specifically use super-resolution technology to enhance the features of all the local features of the vehicle to obtain the fine local features of different perspectives; or use super-resolution technology to enhance the features of the first target features of the local features of the vehicle that meet the preset low-resolution requirements to obtain the fine local features of different perspectives; or use super-resolution technology to enhance the features of the first target features of the local features of the vehicle that meet the preset low-resolution requirements, and feature fuse the enhanced first target features with the second target features of the local features of the vehicle that meet the preset high-resolution requirements to obtain the fine local features of different perspectives.

[0123] Here, the calculation module 403 can specifically calculate the ID loss function corresponding to the vehicle global feature based on the feature distance between the target vehicle image and each vehicle image; and calculate the triplet loss function corresponding to the vehicle global feature based on the feature distance between the target vehicle image and different vehicle images; and calculate the center loss function corresponding to the vehicle global feature based on the number of vehicle image samples selected for each iterative training of the first network model; and then use the ID loss function, the triplet loss function and the center loss function as the first loss function corresponding to the vehicle global feature.

[0124] Here, the calculation module 403 can specifically perform image alignment on the fine local features of different perspectives with the perspective image features in the vehicle image sample set, and calculate the discriminant score of the fine local features corresponding to each perspective; based on the discriminant score, determine the second loss function corresponding to the different perspectives, and the second loss function represents the minimization loss function.

[0125] It should be noted that the vehicle re-identification model training device provided in the above embodiment is the same as the above Figure 1 The training method of the provided vehicle re-identification model belongs to the same concept. The specific implementation process can refer to the above method embodiment and will not be repeated here.

[0126] The training device for the vehicle re-identification model provided in the present application obtains the global features and local features of the vehicle through global branches and local branches, enhances the local features of the vehicle at different perspectives, and fuses the global features of the vehicle with the enhanced fine local features to achieve modeling of the vehicle re-identification model. In this way, it can not only save computing time and resources, but also effectively improve the model's ability to extract local features and robustness for different vehicle perspectives without the need to annotate key points of vehicle images in advance, thereby improving the recognition accuracy of vehicle re-identification.

[0127] Figure 5 This is a schematic diagram of the structure of the vehicle re-identification model training device in this application, such as Figure 5As shown, the model training device 500 includes at least one processor 501 and a memory 502 for storing a computer program that can be run on the processor 501. When the processor 501 runs the computer program, it executes the vehicle re-identification model training method suggested in the above embodiments of the present application. The model training device 500 also includes at least one network interface 504 and a user interface 503. The various components in the model training device 500 are coupled together through a bus system 505. It can be understood that the bus system 505 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 505 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, Figure 5 Various buses are labeled as bus system 505.

[0128] The user interface 503 may include a display, a keyboard, a mouse, a trackball, a click wheel, keys, buttons, a touch pad or a touch screen.

[0129] It is understood that the memory 502 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, a compact disc or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a magnetic disk or a magnetic tape. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM).The memory 502 described in the embodiments of the present application is intended to include, but is not limited to, these and any other suitable types of memory.

[0130] The memory 502 in the embodiment of the present application is used to store various types of data to support the operation of the model training device 500. Examples of these data include: any computer program for operating on the model training device 500, such as an operating system 5021 and an application 5022; messages; pictures; videos, etc. Among them, the operating system 5021 includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., which are used to implement various basic services and process hardware-based tasks. The application 5022 can include various applications, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. The program that implements the method of the embodiment of the present application can be included in the application 5022.

[0131] The processor 501 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor 501 or by instructions in the form of software. The above-mentioned processor 501 may be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 501 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the memory 502. The processor 501 reads the information in the memory 502 and completes the steps of the above method in combination with its hardware.

[0132] In an exemplary embodiment, the model training device 500 can be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to execute the aforementioned method.

[0133] In an exemplary embodiment, the present application also provides a computer-readable storage medium, such as a memory 502 including a computer program, which can be executed by a processor 501 of a model training device 500 to complete the steps of the aforementioned method. The computer-readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface storage, optical disk, or CD-ROM; or various devices including one or any combination of the above memories, such as AR devices, CR devices, VR devices, MR devices, etc.

[0134] A computer-readable storage medium stores a computer program, which, when executed by a processor, executes the vehicle re-identification model training method suggested in the above-mentioned embodiment of the present application.

[0135] In the several embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. In addition, the features disclosed in the several method or device embodiments provided herein can be combined in any manner to obtain new method or device embodiments, provided that they do not conflict.

[0136] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A method for training a vehicle re-identification model, characterized in that: The method comprises: Input the randomly selected target vehicle image into the first network model and the second network model respectively to obtain the corresponding vehicle global features and vehicle local features from different perspectives; Performing feature enhancement on the local features of the vehicle to obtain fine local features of the different perspectives; Calculating a first loss function corresponding to the global feature of the vehicle and a second loss function corresponding to the different perspectives; The first network model and the second network model are optimized based on the first loss function and the second loss function to obtain a vehicle re-identification model.

2. The method according to claim 1, characterized in that Obtaining the corresponding vehicle global features includes: Using the non-local blocks embedded in the first network model, capturing long-range dependencies between features in the target vehicle image in a preset image library; The corresponding vehicle global features are obtained based on the long-distance dependency.

3. The method according to claim 1, characterized in that The vehicle local features obtained from different perspectives include: Performing image segmentation on the target vehicle image using the second network model to obtain multiple perspective images of the target vehicle image; Extracting multiple perspective features corresponding to the multiple perspective images from different data sets respectively; The multiple viewing angle features are determined as vehicle local features of the target vehicle image.

4. The method according to claim 1, wherein The step of enhancing the local features of the vehicle to obtain the fine local features at different viewing angles includes: Using super-resolution technology to enhance all features of the vehicle local features to obtain fine local features of different perspectives; Alternatively, a super-resolution technique is used to enhance the first target feature among the local features of the vehicle that meets the preset low-resolution requirement, thereby obtaining the fine local features at different viewing angles; Alternatively, super-resolution technology is used to enhance the first target feature among the local features of the vehicle that meets the preset low-resolution requirements, and the enhanced first target feature is fused with the second target feature among the local features of the vehicle that meets the preset high-resolution requirements to obtain the fine local features of the different perspectives.

5. The method according to claim 1, characterized in that The calculating a first loss function corresponding to the vehicle global feature includes: Calculating an ID loss function corresponding to the vehicle global feature based on a feature distance between the target vehicle image and each vehicle image; Calculating a triplet loss function corresponding to the vehicle global feature based on the feature distance between the target vehicle image and different vehicle images; Calculating a central loss function corresponding to the vehicle global feature based on the number of vehicle image samples selected in each iterative training of the first network model; The ID loss function, the triplet loss function and the center loss function are used as the first loss function corresponding to the vehicle global feature.

6. The method according to claim 1, characterized in that The calculating of the second loss function corresponding to different perspectives includes: Performing image alignment on the fine local features of different viewing angles with the viewing angle image features in the vehicle image sample set, and calculating the discriminant score of the fine local features corresponding to each viewing angle; A second loss function corresponding to the different perspectives is determined based on the discriminant score, where the second loss function represents a minimized loss function.

7. A vehicle re-identification method, characterized in that: The method comprises: Obtaining an image of a vehicle to be identified; Inputting the vehicle image to be identified into a vehicle re-identification model to obtain a corresponding vehicle target image; The vehicle re-identification model is obtained using the training method according to any one of claims 1 to 6.

8. The method according to claim 7, characterized in that The step of inputting the to-be-identified vehicle image into a vehicle re-identification model to obtain a corresponding vehicle target image includes: Inputting the vehicle image to be identified into a vehicle re-identification model to obtain a local feature distance matrix and a global feature distance matrix of the vehicle image to be identified; The local feature distance matrix and the global feature distance matrix are fused to obtain a vehicle target image list corresponding to the vehicle image to be identified, wherein the vehicle target image list includes at least one vehicle target image.

9. A training device for a vehicle re-identification model, characterized in that: The device comprises: A data acquisition module is used to input randomly selected target vehicle images into the first network model and the second network model respectively to obtain the corresponding vehicle global features and vehicle local features from different perspectives; A super-resolution enhancement module, configured to enhance the local features of the vehicle to obtain fine local features at different viewing angles; A calculation module, configured to calculate a first loss function corresponding to the global feature of the vehicle and a second loss function corresponding to the different perspectives; A joint optimization module is used to optimize the first network model and the second network model based on the first loss function and the second loss function to obtain a vehicle re-identification model.

10. A training device for a vehicle re-identification model, characterized in that: The device comprises: a memory for storing computer programs capable of running on the processor; The processor is configured to execute the vehicle re-identification model training method according to any one of claims 1 to 6 when running the computer program.