Vehicle re-identification method, device, system and computer-readable storage medium
The vehicle image is decomposed into multiple components through the vehicle analysis network, and the component category information is extracted and feature vectors are generated, which solves the problem of difficult to provide fine-grained vehicle information in the prior art, and improves the accuracy and robustness of vehicle re-identification.
Patent Information
- Application Number
- CN202011053858.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-29
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2040-09-29
AI Technical Summary
Existing vehicle re-identification technology is difficult to provide fine-grained information, especially when the image is incomplete or the vehicle posture changes, it is difficult to achieve effective recognition.
Through the vehicle analysis network, the vehicles in the image are analyzed according to multiple components, and the component category information is extracted, and the feature vectors are extracted based on this information.
It provides finer-grained vehicle information, improves the accuracy and robustness of vehicle re-identification, and can effectively deal with incomplete images or changes in vehicle posture.
Smart Images

Figure CN113762000B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, and more specifically, to a vehicle re-identification method, device, system, and computer-readable storage medium. Background Art
[0002] Vehicle re-identification technology refers to finding vehicle images with the same identity among vehicle images taken by large-scale surveillance cameras given a query vehicle image. For example, in the field of traffic law enforcement, vehicle re-identification can be used to find the target vehicle from images taken by different cameras and determine the whereabouts of the target vehicle.
[0003] The existing vehicle re-identification technology mainly uses the entire vehicle image as a sample for image feature extraction and model training, thereby extracting the global appearance features of the vehicle for feature matching and similarity calculation during search. In the process of realizing the concept of the present disclosure, the inventor found that there are at least the following problems in the prior art: the existing vehicle re-identification technology is difficult to provide more fine-grained information related to the vehicle. When performing vehicle re-identification, if there is an incomplete vehicle picture in the image or the vehicle posture is seriously queried in the vehicle posture in the vehicle image, it is difficult to achieve effective recognition. Summary of the invention
[0004] In view of this, embodiments of the present disclosure provide a vehicle re-identification method, device, system, and computer-readable storage medium for performing vehicle identification based on information of various components of the vehicle.
[0005] One aspect of the disclosed embodiment provides a vehicle re-identification method. The method includes: obtaining a first image containing a first vehicle and N second images, each of which contains a vehicle, wherein N is an integer greater than or equal to 1; processing each of the first image and the N second images to be processed to obtain a feature vector corresponding to the vehicle in each of the images to be processed, wherein the feature vector corresponding to the vehicle in the first image is a first feature vector, and the feature vector corresponding to the vehicle in each of the second images is a second feature vector; and determining R vehicles corresponding to the R second feature vectors that are most similar to the first feature vector as the first vehicle, wherein R is an integer greater than or equal to 0 and less than or equal to N. Wherein, processing the first image and each of the N second images to be processed to obtain a feature vector corresponding to the vehicle in each of the images to be processed includes: using a vehicle parsing network to parse the vehicle in the image to be processed according to C components, to obtain component category information corresponding to the vehicle in the image to be processed, wherein the component category information is used to characterize the component category to which each pixel in the image belongs, wherein C is an integer greater than or equal to 2; and obtaining a feature vector corresponding to the vehicle in the image to be processed based on the component category information corresponding to the image to be processed.
[0006] According to an embodiment of the present disclosure, the component category information includes a component category matrix, and the value of each element in the component category matrix represents the component category to which the corresponding pixel position in the image belongs.
[0007] According to an embodiment of the present disclosure, the vehicle parsing network includes a first convolutional neural network. The vehicle in the image to be processed is parsed according to C parts by using the vehicle parsing network to obtain the component category information corresponding to the vehicle in the image to be processed, including: inputting the image to be processed into the first convolutional neural network; obtaining the feature tensor output by the first convolutional neural network, the feature tensor including information in three dimensions of width, height and category, wherein the length and width dimensions in the feature tensor correspond to the size information of the image to be processed, and the category dimension outputs the information of the component category to which each pixel position of the image to be processed belongs; statistically processing the elements on each category in the feature tensor to obtain the category to which each pixel position belongs; and, based on the category to which each pixel position belongs, converting the feature tensor into the category feature matrix.
[0008] According to an embodiment of the present disclosure, the method also includes: obtaining S first training images for training the vehicle parsing network, each of the first training images including a vehicle, where S is an integer greater than or equal to 1; labeling the component category information corresponding to the vehicle in each of the first training images to obtain category labeling information of each of the first training images; and training the vehicle parsing network using the S first training images and the category labeling information of each of the S first training images as training sample data.
[0009] According to an embodiment of the present disclosure, obtaining a feature vector corresponding to a vehicle in the image to be processed based on the component category information corresponding to the image to be processed includes: using a local feature extraction network to process the image to be processed and the component category matrix corresponding to the vehicle in the image to be processed to obtain a local feature vector corresponding to the vehicle in the image to be processed, the local feature vector being a feature vector obtained based on C' feature vectors corresponding one-to-one to C' components of a vehicle, wherein C' is an integer greater than or equal to 1 and less than or equal to C; and obtaining a feature vector corresponding to the vehicle in the image to be processed based on the local feature vector corresponding to the vehicle in the image to be processed.
[0010] According to an embodiment of the present disclosure, the step of obtaining a feature vector corresponding to a vehicle in the image to be processed based on the component category information corresponding to the image to be processed further includes processing the image to be processed using a global feature extraction network to obtain a global feature vector corresponding to the vehicle in the image to be processed. The step of obtaining a feature vector corresponding to the vehicle in the image to be processed based on the local feature vector corresponding to the vehicle in the image to be processed includes combining the local feature vector corresponding to the vehicle in the image to be processed with the global feature vector to obtain a feature vector corresponding to the vehicle in the image to be processed.
[0011] According to an embodiment of the present disclosure, the local feature extraction network includes a second convolutional neural network. The use of the local feature extraction network to process the image to be processed and the component category matrix corresponding to the vehicle in the image to be processed to obtain the local feature vector corresponding to the vehicle in the image to be processed includes: using the second convolutional neural network to process the image to be processed to obtain the overall feature matrix corresponding to the vehicle in the image to be processed; based on the component category matrix corresponding to the vehicle in the image to be processed, obtaining the mask matrix of each of the C components of the vehicle in the image to be processed, wherein the element value corresponding to the pixel position belonging to the component in the mask matrix of a component is the first value, and the element values corresponding to other pixel positions are all the second value; and based on the overall feature matrix corresponding to the vehicle in the image to be processed and the mask matrix of each component of the vehicle in the image to be processed, obtaining the component feature vector corresponding to each component of the vehicle in the image to be processed, wherein C components correspond to C component feature vectors; and obtaining the local feature vector based on the C component feature vectors.
[0012] According to an embodiment of the present disclosure, the local feature extraction network also includes a graph convolutional network, and obtaining the local feature vector based on the C component feature vectors also includes: combining the C component feature vectors to obtain a first component feature matrix, wherein each row of the first component feature matrix corresponds to one component feature vector; processing the first component feature matrix using the graph convolutional neural network to obtain a second component feature matrix corresponding to the first component feature matrix; and processing the second component feature matrix to obtain the local feature vector.
[0013] According to an embodiment of the present disclosure, the method further includes training the local feature extraction network. The training of the local feature extraction network includes: obtaining T second training images and the identity label of the vehicle in each second training image, where T is an integer greater than or equal to 1; and using the T second training images and the identity label of the vehicle therein to train the second convolutional neural network and the graph convolutional network. Specifically, each second training image is processed using the second convolutional neural network to obtain the overall feature matrix corresponding to the vehicle in the second training image; based on the overall feature matrix corresponding to the vehicle in the second training image and the component category matrix corresponding to the vehicle in the second training image, the input of the graph convolutional network is obtained; based on the output of the graph convolutional network, the loss function of the local feature extraction network is obtained; and based on the loss function of the local feature extraction network, the second convolutional neural network and the graph convolutional network are trained.
[0014] According to an embodiment of the present disclosure, the step of obtaining the input of the graph convolution network based on the overall feature matrix corresponding to the vehicle in the second training image and the component category matrix corresponding to the vehicle in the second training image includes: obtaining the mask matrix corresponding to each of the C components of the vehicle in the second training image based on the component category matrix corresponding to the vehicle in the second training image; obtaining the component feature vector corresponding to each component of the vehicle in the second training image based on the overall feature matrix corresponding to the vehicle in the second training image and the mask matrix of each component of the vehicle in the second training image; wherein C component feature vectors are obtained corresponding to C components; combining the C component feature vectors corresponding to the C components of the vehicle in the second training image to obtain the first component feature matrix corresponding to the vehicle in the second training image; setting the elements of some rows in the first component feature matrix corresponding to the vehicle in the second training image to 0 to obtain the first associated component feature matrix corresponding to the vehicle in the second training image; and using the first component feature matrix corresponding to the vehicle in the second training image and the first associated component feature matrix corresponding to the vehicle in the second training image as the input of the graph convolution network respectively.
[0015] According to an embodiment of the present disclosure, the loss function of the local feature extraction network obtained based on the output of the graph convolutional network includes: obtaining the second component feature matrix corresponding to the vehicle in the second training image output by the graph convolutional network when the first component feature matrix corresponding to the vehicle in the second training image is used as input; obtaining the second associated component feature matrix corresponding to the vehicle in the second training image output by the graph convolutional network when the first associated component feature matrix corresponding to the vehicle in the second training image is used as input; processing the second component feature matrix corresponding to the vehicle in the second training image to obtain the local feature vector corresponding to the vehicle in the second training image; processing the second associated component feature matrix corresponding to the vehicle in the second training image to obtain the associated local feature vector corresponding to the vehicle in the second training image; and obtaining the loss function of the local feature extraction network based on the difference between the local feature vector and the associated local feature vector.
[0016] According to an embodiment of the present disclosure, the loss function of the local feature extraction network obtained based on the difference between the local feature vector and the associated local feature vector includes: obtaining a self-supervised loss function based on the norm after subtraction operation between the local feature vector and the associated local feature vector; and obtaining the loss function based on the self-supervised loss function.
[0017] According to an embodiment of the present disclosure, the method further includes training the global feature extraction network. Specifically, Z third training images and the identity label of the vehicle in each of the third training images are obtained, where Z is an integer greater than or equal to 1; the component category matrix corresponding to the vehicle in each of the third training images is obtained; based on the Q third training images and the component category matrix corresponding to the vehicles therein, Q local occlusion images corresponding to the Q third training images are obtained, where each local occlusion image is an image in which part of the components in the corresponding third training image are occluded, where Q is an integer greater than or equal to 1 and less than or equal to Z; training sample data is obtained based on the Z third training images and the Q local occlusion images to train the global feature extraction network.
[0018] According to an embodiment of the present disclosure, the obtaining of Q locally occluded images corresponding one-to-one to the Q third training images based on the component category matrix corresponding to the Q third training images and the vehicles therein comprises: obtaining the mask matrix of each of the C components of the vehicle in the third training image based on the component category matrix corresponding to the vehicle in the third training image; and obtaining the locally occluded image corresponding to the third training image based on the third training image and the mask matrix of some components in the vehicle in the third training image.
[0019] According to an embodiment of the present disclosure, the training sample data is obtained based on Z of the third training images and Q of the local occlusion images to train the global feature extraction network, including: determining one of the third images as a benchmark sample image each time in a traversal manner, obtaining positive sample images and negative sample images from the Z third images, so as to form a triplet through the benchmark sample image, the positive sample image and the negative sample image; the positive sample image is another third training image with the same identity label as the vehicle in the benchmark sample; the negative sample image is any third training image with a different identity label from the vehicle in the benchmark sample; converting any at least one image in the triplet into the corresponding local occlusion image to obtain a conversion triplet; and in each traversal, the three images in the conversion triplet are distributed as a group of input data and input into the global feature extraction network to train the global feature extraction network.
[0020] In another aspect of the disclosed embodiment, a method for training a vehicle re-identification network is provided. The vehicle re-identification network includes a vehicle parsing network, and the vehicle parsing network is used to parse the vehicle in the image according to C component categories, where C is an integer greater than or equal to 2. The training method includes training the vehicle parsing network, specifically including: obtaining S first training images for training the vehicle parsing network, each of the first training images including a vehicle, where S is an integer greater than or equal to 1; labeling component category information corresponding to the vehicle in each of the first training images to obtain category labeling information of each of the first training images; wherein the component category information is used to characterize the component category to which each pixel in the image belongs when classified according to the C component categories; and, using the S first training images and the category labeling information of each of the S first training images as training sample data, the vehicle parsing network is trained.
[0021] According to an embodiment of the present disclosure, the component category information includes a component category matrix, and the value of each element in the component category matrix represents the component category to which the corresponding pixel position in the image belongs.
[0022] According to an embodiment of the present disclosure, the vehicle re-identification network further includes a local feature extraction network, which is used to extract a local feature vector corresponding to the vehicle in the image; the local feature extraction network includes a second convolutional neural network and a graph convolutional network. The training method further includes training the local feature extraction network, specifically including: obtaining T second training images and the identity label of the vehicle in each second training image, where T is an integer greater than or equal to 1; and using the T second training images and the identity labels of the vehicles therein to train the second convolutional neural network and the graph convolutional network. Among them, using the T second training images and the identity labels of the vehicles therein to train the second convolutional neural network and the graph convolutional network specifically includes: using the second convolutional neural network to process the second training image to obtain the overall feature matrix corresponding to the vehicle in the second training image; based on the overall feature matrix corresponding to the vehicle in the second training image and the component category matrix corresponding to the vehicle in the second training image, obtaining the input of the graph convolutional network; obtaining the output of the graph convolutional network; based on the output of the graph convolutional network, obtaining the loss function of the local feature extraction network; and based on the loss function of the local feature extraction network, training the second convolutional neural network and the graph convolutional network.
[0023] According to an embodiment of the present disclosure, the input of the graph convolution network is obtained based on the overall feature matrix corresponding to the vehicle in the second training image and the component category matrix corresponding to the vehicle in the second training image, including: based on the component category matrix corresponding to the vehicle in the second training image, a mask matrix for each of the C components of the vehicle in the second training image is obtained; wherein, in the mask matrix of a component, the element value corresponding to the pixel position belonging to the component is a first value, and the element values corresponding to other pixel positions are all second values; based on the overall feature matrix corresponding to the vehicle in the second training image and the mask matrix for each component of the vehicle in the second training image, a mask matrix for each of the C components of the vehicle in the second training image is obtained. The component feature vector corresponding to each component of the vehicle in the second training image; wherein, C component feature vectors are obtained corresponding to C components; the C component feature vectors corresponding to the C components of the vehicle in the second training image are combined to obtain the first component feature matrix corresponding to the vehicle in the second training image; the elements of some rows in the first component feature matrix corresponding to the vehicle in the second training image are set to 0 to obtain a first associated component feature matrix corresponding to the vehicle in the second training image; and the first component feature matrix corresponding to the vehicle in the second training image and the first associated component feature matrix corresponding to the vehicle in the second training image are respectively used as inputs of the graph convolutional network.
[0024] According to an embodiment of the present disclosure, the loss function of the local feature extraction network obtained based on the output of the graph convolutional network includes: obtaining the second component feature matrix corresponding to the vehicle in the second training image output by the graph convolutional network when the first component feature matrix corresponding to the vehicle in the second training image is used as input; obtaining the second associated component feature matrix corresponding to the vehicle in the second training image output by the graph convolutional network when the first associated component feature matrix corresponding to the vehicle in the second training image is used as input; processing the second component feature matrix corresponding to the vehicle in the second training image to obtain the local feature vector corresponding to the vehicle in the second training image; processing the second associated component feature matrix corresponding to the vehicle in the second training image to obtain the associated local feature vector corresponding to the vehicle in the second training image; and obtaining the loss function of the local feature extraction network based on the difference between the local feature vector and the associated local feature vector.
[0025] According to an embodiment of the present disclosure, the loss function of the local feature extraction network is obtained based on the difference between the local feature vector and the associated local feature vector, including: obtaining a self-supervised loss function based on the norm after subtraction operation between the local feature vector and the associated local feature vector; and obtaining the loss function based on the self-supervised loss function.
[0026] According to an embodiment of the present disclosure, the vehicle re-identification network also includes a global feature extraction network. The global feature extraction network is used to extract the global feature vector of the vehicle in the image. The training method also includes training the global feature extraction network, specifically including: obtaining Z third training images and the identity label of the vehicle in each of the third training images, where Z is an integer greater than or equal to 1; obtaining the component category matrix corresponding to the vehicle in each of the third training images; based on the Q third training images and the component category matrix corresponding to the vehicle therein, obtaining Q partially occluded images corresponding to the Q third training images one by one, wherein each of the partially occluded images is an image in which some components in the corresponding third training image are occluded, where Q is an integer greater than or equal to 1 and less than or equal to Z; obtaining training sample data based on the Z third training images and the Q partially occluded images to train the global feature extraction network.
[0027] According to an embodiment of the present disclosure, the obtaining of Q locally occluded images corresponding one-to-one to the Q third training images based on the component category matrix corresponding to the Q third training images and the vehicles therein comprises: obtaining a mask matrix of each of the C components of the vehicle in the third training image based on the component category matrix corresponding to the vehicle in the third training image; wherein the element value corresponding to the pixel position belonging to the component in the mask matrix of a component is a first value, and the element values corresponding to other pixel positions are all second values; and obtaining the locally occluded image corresponding to the third training image based on the third training image and the mask matrix of some components in the vehicle in the third training image.
[0028] According to an embodiment of the present disclosure, the training sample data is obtained based on Z of the third training images and Q of the local occlusion images to train the global feature extraction network, including: determining one of the third images as a benchmark sample image each time in a traversal manner, obtaining positive sample images and negative sample images from the Z third images, so as to form a triplet through the benchmark sample image, the positive sample image and the negative sample image; the positive sample image is another third training image with the same identity label as the vehicle in the benchmark sample; the negative sample image is any third training image with a different identity label from the vehicle in the benchmark sample; converting any at least one image in the triplet into the corresponding local occlusion image to obtain a conversion triplet; and in each traversal, the three images in the conversion triplet are distributed as a group of input data and input into the global feature extraction network to train the global feature extraction network.
[0029] In another aspect of the disclosed embodiment, a vehicle re-identification device is provided. The device includes an image acquisition module, an image processing module, and a determination module. The image acquisition module is used to acquire a first image containing a first vehicle and N second images, each of which contains a vehicle, where N is an integer greater than or equal to 1. The image processing module is used to process each of the first image and the N second images to be processed, and obtain a feature vector corresponding to the vehicle in each of the images to be processed, where the feature vector corresponding to the vehicle in the first image is a first feature vector, and the feature vector corresponding to the vehicle in each of the second images is a second feature vector. The determination module is used to determine the R vehicles corresponding to the R second feature vectors that are most similar to the first feature vector as the first vehicle, where R is an integer greater than or equal to 0 and less than or equal to N. The image processing module includes a vehicle parsing submodule and a feature vector acquisition submodule. The vehicle parsing submodule is used to parse the vehicle in the image to be processed according to C components using a vehicle parsing network to obtain component category information corresponding to the vehicle in the image to be processed, wherein the component category information is used to characterize the component category to which each pixel in the image belongs, wherein C is an integer greater than or equal to 2; and the feature vector acquisition submodule is used to obtain a feature vector corresponding to the vehicle in the image to be processed based on the component category information corresponding to the image to be processed.
[0030] In another aspect of the disclosed embodiment, a training device for a vehicle re-identification network is provided. The vehicle re-identification network includes the vehicle parsing network, which is used to parse the vehicle in the image according to C component categories, where C is an integer greater than or equal to 2. The training device includes a first image acquisition module, a first labeling module, and a first training module. The first image acquisition module is used to acquire S first training images for training the vehicle parsing network, each of which includes a vehicle, where S is an integer greater than or equal to 1. The first labeling module is used to label the component category information corresponding to the vehicle in each of the first training images to obtain the category labeling information of each of the first training images; wherein the component category information is used to characterize the component category to which each pixel in the image belongs when it is classified according to the C component categories. The first training module is used to train the vehicle parsing network using the S first training images and the category labeling information of each of the S first training images as training sample data.
[0031] According to an embodiment of the present disclosure, the vehicle re-identification network also includes a local feature extraction network. The local feature extraction network is used to extract local feature vectors corresponding to vehicles in an image. The local feature extraction network includes a second convolutional neural network and a graph convolutional network, and the training device also includes a second image acquisition module and a second training module. The second image acquisition module is used to acquire T second training images and the identity label of the vehicle in each of the second training images, where T is an integer greater than or equal to 1. The second training module is used to train the second convolutional neural network and the graph convolutional network using the T second training images and the identity labels of the vehicles therein. The second training module is specifically used to: use the second convolutional neural network to process the second training image to obtain the overall feature matrix corresponding to the vehicle in the second training image; obtain the input of the graph convolutional network based on the overall feature matrix corresponding to the vehicle in the second training image and the component category information corresponding to the vehicle in the second training image; the component category information is a component category matrix, and the value of each element in the component category matrix represents the component category to which the corresponding pixel position in the image belongs when classified according to the C components of the vehicle; obtain the output of the graph convolutional network; based on the output of the graph convolutional network, obtain the loss function of the local feature extraction network; and train the second convolutional neural network and the graph convolutional network based on the loss function of the local feature extraction network.
[0032] According to an embodiment of the present disclosure, the vehicle re-identification network further includes a global feature extraction network, which is used to extract a global feature vector of a vehicle in an image. The training device includes a third image acquisition module, a third category acquisition module, a third occlusion image acquisition module, and a third training module. The third image acquisition module is used to acquire Z third training images and the identity label of the vehicle in each of the third training images, where Z is an integer greater than or equal to 1. The third category acquisition module is used to acquire the component category information corresponding to the vehicle in each of the third training images, the component category information is a component category matrix, and the value of each element in the component category matrix represents the component category to which the corresponding pixel position in the image belongs when the C components of the vehicle are classified. The third occlusion image acquisition module is used to acquire Q partial occlusion images corresponding to the Q third training images one by one based on the component category matrix corresponding to the Q third training images and the vehicles therein, wherein each partial occlusion image is an image in which part of the components in the corresponding third training image are occluded, where Q is an integer greater than or equal to 1 and less than or equal to Z. The third training module is used to obtain training sample data based on the Z third training images and the Q local occlusion images to train the global feature extraction network.
[0033] Another aspect of the embodiment of the present disclosure is a vehicle re-identification system. The system includes one or more memories and one or more processors. The memory stores computer executable instructions. The processor executes the instructions to implement the vehicle re-identification method or the vehicle re-identification network training method as described above.
[0034] Another aspect of the embodiments of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the vehicle re-identification method or the vehicle re-identification network training method as described above when executed.
[0035] Another aspect of an embodiment of the present disclosure provides a computer program, which includes computer executable instructions, and when the instructions are executed, are used to implement the vehicle re-identification method or the vehicle re-identification network training method as described above.
[0036] One or more of the above embodiments have the following advantages or beneficial effects: by parsing the vehicle according to its components during vehicle re-identification, more accurate positioning and feature information of each component can be provided, thereby providing more fine-grained information and improving the accuracy of vehicle re-identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The above and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0038] Figure 1 The following schematically shows a process flow of a vehicle re-identification method according to an embodiment of the present disclosure;
[0039] Figure 2 A flow chart of a vehicle re-identification method according to an embodiment of the present disclosure is schematically shown;
[0040] Figure 3 A flowchart of obtaining component category information by using a vehicle parsing network in a vehicle re-identification method according to an embodiment of the present disclosure is schematically shown;
[0041] Figure 4 A schematic diagram of a training vehicle parsing network according to an embodiment of the present disclosure is schematically shown;
[0042] Figure 5 A flowchart of training a vehicle parsing network according to an embodiment of the present disclosure is schematically shown;
[0043] Figure 6 The schematic diagram shows a process of obtaining a local feature vector corresponding to a vehicle in an image to be processed according to an embodiment of the present disclosure;
[0044] Figure 7 A flowchart for obtaining a feature vector corresponding to a vehicle in an image to be processed according to an embodiment of the present disclosure is schematically shown;
[0045] Figure 8 The following schematically shows a process of obtaining a feature vector corresponding to a vehicle in an image to be processed according to another embodiment of the present disclosure;
[0046] Fig. 9 A flowchart for obtaining a feature vector corresponding to a vehicle in an image to be processed according to another embodiment of the present disclosure is schematically shown;
[0047] Fig.10 A flowchart of extracting a local feature vector using a local feature extraction network according to an embodiment of the present disclosure is schematically shown;
[0048] Fig.11 The schematic diagram shows the process of training a local feature extraction network according to an embodiment of the present disclosure;
[0049] Fig.12 Schematically shows a flow chart of training a local feature extraction network according to an embodiment of the present disclosure;
[0050] Fig.13 A flowchart of obtaining the input of a graph convolutional network in a method for training a local feature extraction network according to an embodiment of the present disclosure is schematically shown;
[0051] Fig.14 A flowchart schematically illustrates a loss function of a local feature extraction network obtained during training of a local feature extraction network according to an embodiment of the present disclosure;
[0052] Fig.15 The schematic diagram shows the process of training a global feature extraction network according to an embodiment of the present disclosure;
[0053] Fig.16 Schematically shows a flow chart of training a global feature extraction network according to an embodiment of the present disclosure;
[0054] Fig.17 Schematically shows a flow chart of obtaining a local occlusion image during the process of training a global feature extraction network according to an embodiment of the present disclosure;
[0055] Fig.18 A flowchart of training a global feature extraction network by data augmentation in the process of training a global feature extraction network according to an embodiment of the present disclosure is schematically shown;
[0056] Fig.19 A block diagram schematically shows a vehicle re-identification device according to an embodiment of the present disclosure; and
[0057] Fig. 20 A block diagram of a computer system suitable for implementing the vehicle identification method or training method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0058] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0059] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise", "include", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.
[0060] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.
[0061] In the case of using expressions such as "at least one of A, B, and C, etc.", it should generally be interpreted in accordance with the meaning of the expression generally understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.). In the case of using expressions such as "at least one of A, B, or C, etc.", it should generally be interpreted in accordance with the meaning of the expression generally understood by those skilled in the art (for example, "a system having at least one of A, B, or C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0062] Embodiments of the present disclosure provide a vehicle re-identification method, device, system, and computer-readable storage medium that can parse a vehicle according to predefined component categories to obtain corresponding component category information and re-identify the vehicle based on the component category information.
[0063] According to an embodiment of the present disclosure, the vehicle re-identification method may include: obtaining a first image containing a first vehicle and N second images, each of which contains a vehicle, wherein N is an integer greater than or equal to 1; processing each of the first image and the N second images to be processed to obtain a feature vector corresponding to the vehicle in each of the images to be processed, wherein the feature vector corresponding to the vehicle in the first image is a first feature vector, and the feature vector corresponding to the vehicle in each of the second images is a second feature vector; and determining R vehicles corresponding to the R second feature vectors that are most similar to the first feature vector as the first vehicle, wherein R is an integer greater than or equal to 0 and less than or equal to N. Wherein, processing the first image and each of the N second images to be processed to obtain a feature vector corresponding to the vehicle in each of the images to be processed includes: using a vehicle parsing network to parse the vehicle in the image to be processed according to C components to obtain component category information corresponding to the vehicle in the image to be processed, wherein the component category information is used to characterize the component category to which each pixel in the image belongs, wherein C is an integer greater than or equal to 2; and obtaining a feature vector corresponding to the vehicle in the image to be processed based on the component category information corresponding to the image to be processed.
[0064] The embodiments of the present disclosure can process images through a vehicle analysis network, use image segmentation technology for fine-grained analysis of various vehicle components, and then apply the fine-grained vehicle analysis to vehicle re-identification.
[0065] Although there are also identification methods for detecting vehicle parts to extract local features of vehicles in the related art, the identification method is specifically to first locate the square bounding boxes of the three parts of the headlights, windows, and logos, then extract the local features respectively, and finally connect the local features with the global features of the vehicle to jointly calculate the vehicle similarity. In comparison, the embodiments of the present disclosure can provide more detailed features of components. Moreover, compared to the use of bounding boxes, the embodiments of the present disclosure parse vehicle parts through a vehicle parsing network, which can provide more accurate positioning and feature information of each component, and the categories of components can be customized. For example, in addition to lights, windows, and logos, multiple components such as the front, rear, top cover, windows, and rearview mirrors can also be defined. Thereby, more fine-grained information can be provided to improve the accuracy of vehicle re-identification.
[0066] According to other embodiments of the present disclosure, in the vehicle re-identification method, based on the component category information corresponding to the image to be processed, a feature vector corresponding to the vehicle in the image to be processed is obtained, for example, the local feature vector of the vehicle in the image may be extracted through a local feature extraction network; or, for example, the global feature vector of the vehicle in the image may be extracted through a global feature extraction network; or, for example, based on the local feature vector and the global feature vector of the vehicle in the image extracted above, the local feature vector and the global feature vector may be combined (vector connection may be performed) and other processing may be performed.
[0067] An embodiment of the present disclosure also provides a method for training a vehicle re-identification network. The vehicle re-identification network may include a vehicle parsing network. Accordingly, the training method may include training the vehicle parsing network. In other embodiments, the vehicle re-identification network may further include a local feature extraction network and / or a global feature extraction network. Accordingly, the training method may further include training a local feature extraction network and / or training a global feature extraction network.
[0068] Figure 1 The figure schematically shows a process flow of a vehicle re-identification method according to an embodiment of the present disclosure.
[0069] like Figure 1 As shown, in the process 100 of the vehicle re-identification method, vehicle re-identification is achieved based on the processing of the first image 101 and at least one second image 102 by the vehicle parsing network 110 (multiple images are shown in the figure). Among them, the vehicle parsing network 110 can parse the vehicle in each image according to predefined component categories. The predefined component categories may include, for example, multiple components such as the front of the vehicle, the rear of the vehicle, the top cover, the window, the rearview mirror, etc. Among them, the first image 101 contains a first vehicle, and each of the second images 102 contains a vehicle, where N is an integer greater than or equal to 1 (multiple second images 102 are shown in the figure).
[0070] The vehicle identification method obtains component category information 11 corresponding to the vehicle in the first image and component category information 12 corresponding to the vehicle in the second image 102 through analysis by a vehicle analysis network 110, and on this basis obtains a first feature vector 13 corresponding to the vehicle in the first image and a second feature vector 14 corresponding to the vehicle in the second image 102. Then, based on the similarity between the first feature vector 13 and the N second feature vectors 14, a vehicle with the same identity as the first vehicle can be identified from the vehicles in the N second images 102.
[0071] Figure 2 The flowchart of a vehicle re-identification method according to an embodiment of the present disclosure is schematically shown.
[0072] Combination Figure 1 and Figure 2 According to an embodiment of the present disclosure, the vehicle re-identification method may include operations S210 to S240.
[0073] In operation S210, a first image 101 including a first vehicle and N second images 102 are obtained, each of the second images 102 including a vehicle, where N is an integer greater than or equal to 1. The second images 102 are, for example, images of vehicles cut out from a traffic monitoring video.
[0074] Then in operation S220, for each image to be processed 101 / 102 in the first image 101 and the N second images 102, the vehicle in the image to be processed is parsed according to C parts using the vehicle parsing network 110 to obtain part category information 11, 12 corresponding to the vehicle in the image to be processed, wherein the part category information 11, 12 is used to characterize the part category to which each pixel in the image belongs, where C is an integer greater than or equal to 2.
[0075] The component category information 11 and 12 may be, for example, information about which category of the C categories of vehicles each pixel position in the image belongs to. According to some embodiments of the present disclosure, the component category information 11 and 12 may be a component category matrix M, and the value of each element in the component category matrix M represents the component category to which the corresponding pixel position in the image belongs.
[0076] Next, in operation S230, feature vectors 13 and 14 corresponding to the vehicles in the images to be processed are obtained based on the component category information 11 and 12 corresponding to the images to be processed 101 and 102. The feature vector corresponding to the vehicle in the first image 101 is a first feature vector 13, and the feature vector corresponding to each vehicle in the second image 102 is a second feature vector 14.
[0077] Next, in operation S240 , R vehicles corresponding to the R second feature vectors 14 that are most similar to the first feature vector 13 are determined as the first vehicles, where R is an integer greater than or equal to 0 and less than or equal to N.
[0078] Figure 3 The flowchart schematically shows operation S220 of the vehicle re-identification method according to an embodiment of the present disclosure, in which the vehicle parsing network 110 is used to obtain component category information.
[0079] like Figure 3 As shown, according to an embodiment of the present disclosure, the vehicle parsing network 110 may include, for example, a first convolutional neural network, and operation S220 may include operations S321 to S324.
[0080] First, in operation S321, the image to be processed is input into the first convolutional neural network.
[0081] Then in operation S322, a feature tensor output by the first convolutional neural network is obtained, wherein the feature tensor includes information in three dimensions: width W, height H, and category C, wherein the length and width dimensions in the feature tensor correspond to the size information of the image to be processed, and the category dimension outputs information on the component category to which each pixel position of the image to be processed belongs.
[0082] In one embodiment, the vehicle parsing network 110 may be composed of a fully convolutional neural network, which is composed of n convolutional layers and m deconvolutional layers. x is the input, and the output Y of the previous convolutional layer is l-1 is the input and the output is the feature tensor Y l Or the segmentation result M′, the entire network can be expressed as:
[0083] M′=P(I)=f L (f L-1 (...f 1 (I x ))), where Y l =f l (Y l-1 ), Y 0 =I x
[0084] The convolution kernel size in the n convolutional layers of this fully convolutional neural network is k l ×k l ×c l , operation step length s l =k l / 2, so the length of the characteristic tensor is H l and width W lReduce layer by layer; the convolution kernel size of m deconvolution layers is k l ×k l ×c l , operation step length s l = 1, so the length of the output feature tensor is H l and width W l Gradually expand. Each layer of convolution operation is followed by a batch normalization operation and an activation function, where the activation function is a linear rectifier function, but is not limited to this activation function. The convolution kernel size of the last convolution layer of the fully convolutional neural network is k l ×k l ×C, where C is the predefined number of components, and outputs a feature tensor W×H×C with the same width and height as the original image.
[0085] Then in operation S323, statistical processing is performed on the elements of each category in the feature tensor to obtain the category to which each pixel position belongs.
[0086] Then in operation S324, based on the category to which each pixel position belongs, the feature tensor is converted into the category feature matrix.
[0087] For example, for the feature vector W×H×C, a normalized exponential function is used at each pixel position to normalize each element of the category dimension channel. For example, formula (1) can be expressed as:
[0088]
[0089] In the formula,
[0090] z c The value of each element of the category dimension channel at each pixel position in the feature vector W×H×C output by the full convolutional network;
[0091] K is the number of elements in the category dimension channel at each pixel position in the feature vector W×H×C output by the full convolutional network;
[0092] σ(z) c It is the value after normalization operation on each element of the category dimension channel in each pixel position.
[0093] After normalization processing is performed by formula (1), the element values of the category dimension channel are counted at each pixel position, such as averaging or taking the highest value, to obtain the value corresponding to each pixel position, and the component category corresponding to the value is used as the category to which the pixel position belongs. Then, a W×H component category matrix M is finally obtained, in which each pixel position corresponds to a vehicle component category in the image.
[0094] Figure 4A schematic diagram of a training vehicle parsing network according to an embodiment of the present disclosure is schematically shown.
[0095] The vehicle re-identification method according to the embodiment of the present disclosure may also include training a vehicle parsing network. Figure 4 As shown, in the embodiment 400, the vehicle parsing network is illustrated as a vehicle parsing network 410, which may be an embodiment of the vehicle parsing network 110. Specifically, when training the vehicle parsing network 410, the vehicle parsing network 410 may be trained by training sample data consisting of a plurality of first training images 401, 402, 403, and component category information 41, 42, 43 corresponding to each first training image.
[0096] Figure 5 A flowchart of training a vehicle parsing network according to an embodiment of the present disclosure is schematically shown.
[0097] Combination Figure 4 ,like Figure 5 As shown, the process of training a vehicle parsing network according to an embodiment of the present disclosure may include operations S510 to S530.
[0098] First, in operation S510 , S first training images for training the vehicle parsing network 410 are obtained, each of the first training images includes a vehicle, where S is an integer greater than or equal to 1.
[0099] Next, in operation S520, the component category information corresponding to the vehicle in each of the first training images is labeled to obtain category labeling information of each of the first training images.
[0100] Then, in operation S530 , the vehicle parsing network 410 is trained using the S first training images and the category labeling information of each of the S first training images as training sample data.
[0101] For example, the vehicle parsing network 410 uses the first training image I 1x and the sample pairs consisting of the component category matrix M 1x , M> for training, use the pixel-level cross entropy loss function to calculate the model loss value, and use the stochastic gradient descent algorithm to optimize the model.
[0102] Accordingly, the embodiment of the present disclosure also provides a training device for a vehicle parsing network. The training device can be used to implement a reference Figure 4 and 5 The training method described herein comprises a first image acquisition module, a first annotation module, and a first training module.
[0103] The first image acquisition module is used to acquire S first training images for training the vehicle parsing network 410 , each of the first training images includes a vehicle, where S is an integer greater than or equal to 1.
[0104] The first labeling module is used to label the component category information corresponding to the vehicle in each of the first training images to obtain the category labeling information of each of the first training images; wherein the component category information is used to represent the component category to which each pixel in the image belongs when classified according to the C component categories.
[0105] The first training module is used to train the vehicle parsing network 410 using the S first training images and the category labeling information of each of the S first training images as training sample data.
[0106] Figure 6 The figure schematically shows a process of obtaining a local feature vector corresponding to a vehicle in an image to be processed according to an embodiment of the present disclosure.
[0107] like Figure 6 As shown, when extracting the feature vector of the image in this embodiment 600, the image to be processed 601 can be first parsed by the vehicle parsing network 610 to obtain the component category matrix 61. Then, the image to be processed 601 and the component category matrix 61 are used as inputs of the local feature extraction network 620, and the image to be processed 601 and the component category matrix 61 are processed by the local feature extraction network 620 to obtain the local feature vector 62 corresponding to the vehicle in the image to be processed 601. Among them, the image to be processed 601 can be Figure 1 For any one of the first image 101 or the N second images 102 shown in FIG. 1 , the vehicle parsing network 610 is an embodiment of the vehicle parsing network 110 .
[0108] Figure 7 The flowchart of obtaining a feature vector corresponding to a vehicle in an image to be processed in operation S230 according to an embodiment of the present disclosure is schematically shown.
[0109] like Figure 7 As shown, combined Figure 6 According to this embodiment, operation S230 may include operation S710 and operation S720.
[0110] In operation S710, the image to be processed 601 and the component category matrix 61 corresponding to the vehicle in the image to be processed 601 are processed using a local feature extraction network 620 to obtain a local feature vector 62 corresponding to the vehicle in the image to be processed 601, wherein the local feature vector 62 is a feature vector obtained based on C' feature vectors corresponding one-to-one to C' components of a vehicle, wherein C' is an integer greater than or equal to 1 and less than or equal to C.
[0111] The local feature extraction network 620 can accurately segment vehicle components based on the component category matrix 61 and extract visual features of each component respectively, so that the obtained local feature vector 62 can provide more accurate component segmentation feature information.
[0112] According to one embodiment of the present disclosure, the local feature extraction network 620 may include a second convolutional neural network. According to another embodiment of the present disclosure, the local feature extraction network 620 may further include a graph convolutional network, wherein the output of the second convolutional network may be used as at least a part of the input of the graph convolutional network.
[0113] In operation S720 , a feature vector corresponding to the vehicle in the image to be processed 601 is obtained based on the local feature vector 62 corresponding to the vehicle in the image to be processed 601 .
[0114] According to the embodiments of the present disclosure, by extracting local detail features of a vehicle through the full local feature extraction network 620, more fine-grained information for vehicle identification can be provided, thereby improving the distinguishing power and robustness of features during vehicle identification.
[0115] Figure 8 The following schematically shows a process of obtaining a feature vector corresponding to a vehicle in an image to be processed according to another embodiment of the present disclosure.
[0116] like Figure 8 As shown, in this embodiment 800, when extracting the feature vector of the image, the image to be processed 801 can be first parsed by the vehicle parsing network 810 to obtain the component category matrix 81. Then, the image to be processed 801 and the component category matrix 81 are used as inputs of the local feature extraction network 820, and the image to be processed 801 and the component category matrix 81 are processed by the local feature extraction network 820 to obtain the local feature vector 82 corresponding to the vehicle in the image to be processed 801. At the same time, the image to be processed 801 can be processed by the global feature extraction network 820 to obtain the global feature vector 83 of the image to be processed 801. Finally, the local feature vector 82 and the global feature vector 83 can be combined together to form a feature vector 84 corresponding to the vehicle in the image to be processed 801. Among them, the image to be processed 801 can be Figure 1 For any one of the first image 101 or the N second images 102 shown in FIG. 1 , the vehicle parsing network 810 may be an embodiment of the vehicle parsing network 110 .
[0117] In this embodiment, a dual-branch neural network model is used to extract effective features of the vehicle for vehicle similarity matching, wherein one branch is a global feature extraction network for extracting global features, and the other branch is a local feature extraction network 820, which extracts local features based on the precise segmentation of vehicle parts.
[0118] Fig. 9 The flowchart of obtaining a feature vector corresponding to a vehicle in an image to be processed in operation S230 according to another embodiment of the present disclosure is schematically shown.
[0119] like Fig. 9 As shown, combined Figure 8 , operation S230 may include operations S910 to S930.
[0120] Operation S910, using the local feature extraction network 820 to process the image to be processed 801 and the component category matrix 81 corresponding to the vehicle in the image to be processed 801, to obtain a local feature vector 82 corresponding to the vehicle in the image to be processed 801, wherein the local feature vector 82 is a feature vector obtained based on C' feature vectors corresponding to C' components of a vehicle, wherein C' is an integer greater than or equal to 1 and less than or equal to C. According to one embodiment of the present disclosure, the local feature extraction network 820 may include a second convolutional neural network. According to another embodiment of the present disclosure, the local feature extraction network 820 may further include a graph convolutional network.
[0121] In operation S920 , the image to be processed 801 is processed by using a global feature extraction network 830 to obtain a global feature vector 83 corresponding to the vehicle in the image to be processed 801 .
[0122] In operation S930 , the local feature vector 82 and the global feature vector 83 corresponding to the vehicle in the image to be processed 801 are combined to obtain a feature vector 84 corresponding to the vehicle in the image to be processed 801 .
[0123] For example, in one embodiment, the trained vehicle parsing network P(·), global feature extraction network F(·) and local feature extraction network G(.) may be loaded first.
[0124] Then for the vehicle database set {I d} and a first image I containing a first vehicle q For each image to be processed in the vehicle parsing network P(·), the image to be processed can be input into the vehicle parsing network P(·) to calculate the corresponding component category matrix {M d} and M q .
[0125] Next, the data set {I d} and the corresponding component category matrix set {M d}, input F(·) and G(·) together to obtain the global feature vector set {g d} and the local eigenvector {r d Then, the global feature vector and the local feature vector corresponding to the vehicle in the same second image can be connected to form the feature vector corresponding to the vehicle in the second image, thereby obtaining the second feature vector set {v d =[g d , r d ]}.
[0126] Similarly, the first image I q and the corresponding component category matrix M q , input F(·) and G(·) together to obtain the first image I q The global feature vector g corresponding to the vehicle in q and the local eigenvector r q Then the global feature vector and the local feature vector are connected to form the first image I q The eigenvector corresponding to the vehicle in (i.e., the first eigenvector) v q =[g q , r q ].
[0127] Next, in one embodiment, for example, the first eigenvector v can be calculated according to the following formula (2): q With each second eigenvector v in the second eigenvector set d The cosine distance of :
[0128]
[0129] Finally, the first image I q With each second image {I d} from large to small cosine distance, and return the sorting result. According to the sorting result, from each second image {I d}determine the first vehicle.
[0130] According to the embodiments of the present disclosure, the global appearance features and local detail features of a vehicle are extracted respectively through a global feature extraction network and a local feature extraction network, thereby improving the distinguishing power and robustness of features during vehicle recognition and achieving accurate vehicle re-recognition results.
[0131] Fig.10 The flowchart of extracting a local feature vector by using a local feature extraction network in operation S910 according to an embodiment of the present disclosure is schematically shown.
[0132] like Fig.10 As shown, according to this embodiment, the local feature extraction network 820 may include a second convolutional neural network and a graph convolutional network, and extracting the local feature network in operation S910 may include operations S1011 to S1061.
[0133] In operation S1011, the image to be processed is processed using the second convolutional neural network to obtain an overall feature matrix corresponding to the vehicle in the image to be processed.
[0134] In operation S1021, based on the component category matrix corresponding to the vehicle in the image to be processed, a mask matrix of each of the C components of the vehicle in the image to be processed is obtained, wherein the element value corresponding to the pixel position belonging to the component in the mask matrix of a component is a first value, and the element values corresponding to other pixel positions are all second values. Of the first value and the second value, one may be 0 and the other may be 1, but is not limited to the other.
[0135] In operation S1031, based on the overall feature matrix corresponding to the vehicle in the image to be processed and the mask matrix of each component of the vehicle in the image to be processed, a component feature vector corresponding to each component of the vehicle in the image to be processed is obtained; wherein C component feature vectors are obtained corresponding to C components.
[0136] In operation S1041, the C component feature vectors are combined to obtain a first component feature matrix, wherein each row of the first component feature matrix corresponds to one component feature vector.
[0137] In operation S1051, the first component feature matrix is processed using the graph convolutional neural network to obtain a second component feature matrix corresponding to the first component feature matrix.
[0138] In operation S1061, the second component feature matrix is processed to obtain the local feature vector.
[0139] For example, in one embodiment, the second convolutional neural network is a multi-layer convolutional neural network G(·), and the graph convolutional network is a multi-layer graph convolutional network H(·). In operation S1011, the image to be processed I x Input G(·), and then perform forward propagation operation through multiple convolutional layers to obtain the overall feature matrix X′=G(I x ). Then, in operation S1021, the component category matrix M of size W×H can be converted into a matrix of corresponding size according to the output format of G(·) using the nearest neighbor interpolation method. And based on the matrix Calculate each type of component C according to the following formula (3): X The mask matrix is:
[0140]
[0141] in is an indicator function. When the input satisfies the condition x=C X The function returns 1 when , otherwise 0. The mask matrix obtained at this time , corresponding to the vehicle component C in the image X The value of the element is 1, and the others are 0.
[0142] Further, in operation S1031, component C can be calculated according to the following formula (4): X The characteristic matrix of
[0143]
[0144] Where ⊙ is the element-wise multiplication operation;
[0145] Then add component C X The characteristic matrix Convert the matrix to a vector and get component C X The eigenvector r c . And the component feature vectors of all C components are combined to form the first component feature matrix X = [r′ 1 , r′ 2 , ..., r′ c ] T .
[0146] Next, in this embodiment, the first component feature matrix X can be input into the multi-layer graph convolutional network H(·) in operation S1051. Each layer in H(·) operates on X according to the following formula: Where A is the adjacency matrix of X, D is the degree matrix of matrix A, and X (L-1) is the output of the L-1th layer, W (L) is the parameter of the Lth layer, and σ(·) is the activation function. The first layer input of the multi-layer graph convolutional network H(·) is the first feature matrix X, and the output is the second component feature matrix X after L layers of operation. L .
[0147] Then, in operation S1061, the two-component feature matrix X L Convert to a local feature vector r. For example, convert the two-component feature matrix X L The elements of each row in are averaged to obtain the local eigenvector r.
[0148] Fig.11The flowchart of training a local feature extraction network according to an embodiment of the present disclosure is schematically shown.
[0149] like Fig.11 As shown, the method according to the embodiment 1100 may further include training a local feature network. The local feature network 1120 may include a second convolutional neural network 1121 and a graph convolutional network 1122, wherein the second convolutional neural network 1121 and the graph convolutional network 1122 may be in a series relationship. The input of the graph convolutional network 1122 may be obtained based on the output of the second convolutional neural network 1121.
[0150] When training the local feature network 1120 , the training data may include second training images 1101 , 1102 , 1103 , and identity tags 111 , 112 , 113 of the vehicles in the respective second training images.
[0151] In the process of training the local feature network 1120, the loss function 114 can be obtained based on the output of the graph convolution network 1122. Then, the parameters of the second convolution neural network 1121 and the graph convolution network 1122 are adjusted and optimized according to the loss function 114.
[0152] Fig.12 The flowchart of training a local feature extraction network according to an embodiment of the present disclosure is schematically shown.
[0153] like Fig.12 As shown, combined Fig.11 According to an embodiment of the present disclosure, the process of training the local feature extraction network may include S1210 to S1250.
[0154] In operation S1210, T second training images and an identity label of a vehicle in each of the second training images are obtained, where T is an integer greater than or equal to 1.
[0155] In operation S1220, the second training image is processed using the second convolutional neural network 1121 to obtain the overall feature matrix X′ corresponding to the vehicle in the second training image.
[0156] In operation S1230, based on the overall feature matrix X′ corresponding to the vehicle in the second training image and the component category matrix M corresponding to the vehicle in the second training image, an input of the graph convolutional network 1122 is obtained.
[0157] In operation S1240, based on the output of the graph convolutional network 1122, a loss function of the local feature extraction network is obtained.
[0158] In operation S1250, based on the loss function of the local feature extraction network 1120, the second convolutional neural network 1121 and the graph convolutional network 1122 are trained.
[0159] Fig.13 The flowchart of obtaining the input of the graph convolutional network in the method of training the local feature extraction network in operation S1230 according to one embodiment of the present disclosure is schematically shown.
[0160] like Fig.13 As shown, according to an embodiment of the present disclosure, operation S1230 may include operations S1331 to S1335.
[0161] In operation S13331, based on the component category matrix M corresponding to the vehicle in the second training image, the mask matrix of each of the C components of the vehicle in the second training image is obtained. For example, the mask matrix of each component is obtained by formula (3):
[0162] Operation S1332: based on the overall feature matrix X′ corresponding to the vehicle in the second training image and the mask matrix of each component of the vehicle in the second training image, The component feature vector corresponding to each component of the vehicle in the second training image is obtained. Among them, C components correspond to C component feature vectors r c .
[0163] In operation S1333, the C component feature vectors r corresponding to the C components of the vehicle in the second training image are c The first component feature matrix X corresponding to the vehicle in the second training image is obtained by combining.
[0164] In operation S1334, elements of some rows in the first component feature matrix corresponding to the vehicle in the second training image are set to 0 to obtain a first associated component feature matrix corresponding to the vehicle in the second training image.
[0165] In operation S1335, the first component feature matrix X corresponding to the vehicle in the second training image and the first associated component feature matrix X corresponding to the vehicle in the second training image are converted into They are respectively used as the input of the graph convolutional network.
[0166] For example, in a practical application, when training a local feature extraction network, the second training image I 2x Input G(·), and perform forward propagation operation through multiple convolutional layers to obtain the overall feature matrix X′=G(I 2x). Then the second training image I can be interpolated using the nearest neighbor method. 2x The component category matrix M is converted into a matrix of corresponding size according to the output format of G(·) And based on the matrix Calculate each component C according to formula (3): X The mask matrix of component C can then be calculated according to the following formula (4): X The characteristic matrix Then, place component C X The characteristic matrix Convert the matrix to a vector and get component C X The eigenvector r c And the component feature vectors of all C components are combined to form the second training image I 2x The first component feature matrix X corresponding to the vehicle in 1 , r′ 2 , ..., r′ c ] T When training the local feature extraction network, for the first component feature matrix X = [r′ 1 , r′ 2 , ..., r′ c ] T , randomly set a row in the matrix to 0, and get the first correlation feature matrix Thus, in operation S1335, the first component feature matrix X and the first association feature matrix They are respectively used as the input of the multi-layer graph convolutional network H(·).
[0167] Fig.14 The flowchart schematically shows a loss function of a local feature extraction network obtained during training of the local feature extraction network in operation S1240 according to an embodiment of the present disclosure.
[0168] like Fig.14 As shown, according to an embodiment of the present disclosure, operation S1240 may include operations S1441 to S1445.
[0169] In operation S1441, the second component feature matrix X corresponding to the vehicle in the second training image output by the graph convolution network is obtained when the first component feature matrix X corresponding to the vehicle in the second training image is used as input. L .
[0170] In operation S1442, the graph convolutional network is used to obtain the first associated component feature matrix corresponding to the vehicle in the second training image. When is input, the second associated component feature matrix corresponding to the vehicle in the second training image is output
[0171] In operation S1443, the second component feature matrix X corresponding to the vehicle in the second training image is processed. L , obtain the local feature vector r corresponding to the vehicle in the second training image.
[0172] In operation S1444, the second associated component feature matrix corresponding to the vehicle in the second training image is processed. Get the associated local feature vector corresponding to the vehicle in the second training image
[0173] In operation S1445, a loss function of the local feature extraction network is obtained based on the difference between the local feature vector and the associated local feature vector.
[0174] According to an embodiment of the present disclosure, the graph convolution network in the local feature network is constructed by processing the first component feature matrix X and the first association feature matrix Through learning processing, we can construct an adjacency graph of vehicle components, use graph convolutional neural networks to model the relationships between components, and improve the representation ability of local features.
[0175] For example, the first component feature matrix X and the first association feature matrix They are input into the multi-layer graph convolutional network H(·) respectively. Each layer in H(·) performs the following formula on X and Calculate according to formula (5). Output the second component feature matrix X L and the second associated component feature matrix Then the two-component feature matrix X L and the second associated component feature matrix Each corresponding feature vector is converted into a local feature vector r and an associated local feature vector Thus, in operation S1445, based on the local feature vector r and the associated local feature vector Get the loss function.
[0176] In one embodiment, a self-supervised loss function can be obtained based on the norm after subtraction of the local feature vector and the associated local feature vector, and then the loss function can be obtained based on the self-supervised loss function. The self-supervised loss function can enable the graph convolution network to predict invisible parts through visible parts of the vehicle, thereby improving the characterization ability of local features under different viewing angles. Through the self-supervised loss function, the robustness of the local feature extraction network to changes in viewing angle and invisible occlusion can be improved.
[0177] In an application example, when performing local extraction network training, for a training set data consisting of T second training images and the identity labels of the vehicles therein, the second training images and their identity labels in the training set are sequentially traversed. A second training image I is selected in each traversal. 2x and its identity label y, and select it as the benchmark sample I 2x,a , randomly select an image with the same identity label y as the positive sample image I 2x,p , randomly select an image with a different identity label y as the negative sample image I 2x,n .Will 2x,a , I 2x,p , I 2x,n >Form triplets as training sample sets.
[0178] The triple 2x,a , I 2x,p , I 2x,n The images in > are input into the local feature extraction network G(H(·)) to obtain the corresponding local feature vectors <r 2x,a , r 2x,p , r 2x,n > and the associated local eigenvector Then, the triple loss values are calculated according to the following formulas (6) and (7):
[0179]
[0180]
[0181] Any local feature vector r 2x Input a fully connected layer f(·), and then obtain 1×Y through softmax operation 2x The classification probability vector f(r 2x ), where Y 2x For the number of all identity labels in the training set, the cross entropy loss function is calculated according to the following formula:
[0182]
[0183] Given a pair of local eigenvectors r and associated local eigenvectors The self-supervised loss function is calculated according to the following formula (9):
[0184]
[0185] Then the local feature extraction network is calculated to the total loss value by the following formula (10):
[0186]
[0187] The network parameters of the multi-layer convolutional neural network G(·) and the multi-layer graph convolutional network H(·) are optimized using the stochastic gradient descent algorithm and the back-propagation algorithm.
[0188] The present disclosure also provides a training device for a local feature extraction network. The training device can be used to implement the reference Figure 11 to Figure 14 The training method of the local feature extraction network described. The local feature extraction network 1120 includes a second convolutional neural network 1121 and a graph convolutional network 1122. The training device includes a second image acquisition module and a second training module.
[0189] Specifically, the second image acquisition module is used to acquire T second training images and the identity label of the vehicle in each of the second training images, where T is an integer greater than or equal to 1.
[0190] The second training module is used to train the second convolutional neural network 1121 and the graph convolutional network 1122 using T second training images and the identity labels of the vehicles therein.
[0191] The second training module is specifically used to first use the second convolutional neural network 1121 to process the second training image to obtain the overall feature matrix corresponding to the vehicle in the second training image. Then, based on the overall feature matrix corresponding to the vehicle in the second training image and the component category matrix corresponding to the vehicle in the second training image, the input of the graph convolutional network 1122 is obtained. Then, the output of the graph convolutional network 1122 is obtained. Then, based on the output of the graph convolutional network 1122, the loss function 114 of the local feature extraction network 1120 is obtained. Then, based on the loss function of the local feature extraction network 1122, the second convolutional neural network 1121 and the graph convolutional network 1122 are trained.
[0192] Fig.15 The figure schematically shows a process diagram of training a global feature extraction network according to an embodiment of the present disclosure.
[0193] like Fig.15 As shown, the method according to this embodiment 1500 may also include training a global feature network, wherein the global feature network is Fig.15 It is illustrated as a global feature network 1520.
[0194] In this embodiment 1500, the process of training the global feature network 1520 may be, first, parsing the plurality of third training images 1501 through the vehicle parsing network 1510 to obtain the component category matrix 151 corresponding to the vehicle in each third training image. Then, based on the component category matrix 151 corresponding to each third training image 1501, the corresponding third training image 1501 is processed to obtain the corresponding local occlusion image 1502. The local occlusion image 1502 may be, for example, an image formed after one or more components in the third training image 1501 are occluded. Finally, the third training image 1501 and the local occlusion image 1502 may be combined according to a certain strategy to form training data, so as to train the global feature extraction network 1530. Among them, the vehicle parsing network 1510 may be an embodiment of the vehicle parsing network 110.
[0195] Fig.16 The flowchart of training a global feature extraction network according to an embodiment of the present disclosure is schematically shown.
[0196] like Fig.16 As shown, combined Fig.15 According to an embodiment of the present disclosure, the process of training a global feature extraction network may include operations S1610 to S1640.
[0197] In operation S1610, Z third training images 1501 and an identity label of a vehicle in each of the third training images are obtained, where Z is an integer greater than or equal to 1.
[0198] In operation S1620 , the component category matrix 151 corresponding to each vehicle in the third training image 1501 is obtained.
[0199] In operation S1630, based on the Q third training images and the component category matrix 151 corresponding to the vehicles therein, Q partial occlusion images 1502 corresponding one to one to the Q third training images are obtained, wherein each of the partial occlusion images 1502 is an image in which a part of the component in the corresponding third training image is occluded, wherein Q is an integer greater than or equal to 1 and less than or equal to Z.
[0200] In operation S1640, training sample data is obtained based on the Z third training images 1501 and the Q partial occlusion images 1502 to train the global feature extraction network. For example, the Q partial occlusion images 1502 may be expanded into the training sample data.
[0201] Fig.17 The flowchart of operation S1630 of obtaining a local occlusion image in the process of training a global feature extraction network according to an embodiment of the present disclosure is schematically shown.
[0202] like Fig.17 As shown, according to an embodiment of the present disclosure, operation S1630 may include operation S1731 and operation S1732.
[0203] In operation S1731, based on the component category matrix corresponding to the vehicle in the third training image, the mask matrix of each of the C components of the vehicle in the third training image is obtained. The mask matrix M is calculated based on the component category matrix M using the following formula (11): * :
[0204]
[0205] in is an indicator function. When the input meets the condition x=c, the function return value is 1, otherwise it is 0. c is a randomly selected vehicle component category. In the mask matrix M obtained at this time, the element value corresponding to the vehicle component c in the image is 1, and the others are 0.
[0206] In operation S1732, the local occlusion image corresponding to the third training image is obtained based on the third training image and the mask matrix of the partial components in the vehicle in the third training image.
[0207] For example, for a third training image I 3x , based on the mask matrix M of a certain component c, use the following formula (12) to calculate the image after the component is erased:
[0208]
[0209] Among them, ⊙ is the element-wise multiplication operation.
[0210] At this time, the occluded image after the random component c corresponding to the input image I is erased is obtained In the occlusion image In the figure, the pixel value of the pixel corresponding to the component c area is 0, and the other areas are the original image I 3x The pixel value of .
[0211] Fig.18 The flowchart of operation S1640 of training the global feature extraction network by data augmentation in the process of training the global feature extraction network according to an embodiment of the present disclosure is schematically shown.
[0212] like Fig.18 As shown, according to an embodiment of the present disclosure, operation S1640 may include operations S1841 to S1843.
[0213] In operation S1841, one of the third images is determined as a benchmark sample image each time in a traversal manner, and a positive sample image and a negative sample image are obtained from the Z third images to form a triplet through the benchmark sample image, the positive sample image and the negative sample image; the positive sample image is another third training image having the same identity label as the vehicle in the benchmark sample; and the negative sample image is any third training image having a different identity label from the vehicle in the benchmark sample.
[0214] In operation S1842, any at least one image in the triplet is converted into the corresponding local occlusion image to obtain a converted triplet.
[0215] In operation S1843, three images in the conversion triplet are distributed as a group of input data and input to the global feature extraction network in each traversal to train the global feature extraction network.
[0216] In an application example, when training a global feature extraction network, for a training set data consisting of Z third training images and the identity labels of the vehicles therein, the training set images and labels are sequentially traversed. 3x and its identity label y 3x , select it as the benchmark sample I 3x,a , randomly select a picture with identity label y in the training set 3x The same image is used as the positive sample image I 3x,p , randomly select a picture with identity label y in the training set 3x Different images, as negative sample images I 3x,n .
[0217] Will 3x,a , I 3x,p , I 3x,n >Form a triple as a training sample set. Given probability P, for the triple 3x,a , I 3x,p , I 3x,n >Use a random number generator to generate a random number p between 0 and 1. If p>P, then 3x,a , I 3x,p , I 3x,n >Any image in the random component erasing operation is performed according to operation S1732 to obtain an occlusion image with any component erased. Then, the occlusion image is used to replace the corresponding triplet 3x,a , I 3x,p , I 3x,n >In the original graph, update the triplet.
[0218] Then the triple 3x,a , I 3x,p , I 3x,n > Input the global feature extraction network F(·) respectively to obtain the corresponding global feature vector <g 3x,a , g 3x,p , g 3x,n >. Then calculate the triple according to the following formula (13): 3x,a , I 3x,p , I 3x,n >Loss value:
[0219] L 3x,t =max(0,||g 3x,a -g 3x,n ||-||g 3x,a -g 3x,p ||-m)+α||g 3x,a -g 3x,n || (13)
[0220] Where α is the weight coefficient, ||·|| is the L of the calculation vector 2 Norm.
[0221] Any global eigenvector g 3x Input a fully connected layer f(·), and then obtain the classification probability vector f(g) of 1×Y dimension through softmax operation. 3x ), where Y is the number of all identity labels in the training set, and the cross entropy loss function is calculated according to the following formula (14):
[0222]
[0223] The total loss value of the global feature extraction network is calculated according to the following formula (15):
[0224] L g,3x =L t,3x +L ce,3x (15)
[0225] The parameters of the global feature extraction network are optimized using the stochastic gradient descent algorithm and the back propagation algorithm.
[0226] In the disclosed embodiment, when training the global feature extraction network, the occluded images erased by random components are used to expand the training data, which can improve the robustness of the global feature vector extracted by the global feature extraction network to local occlusion of the vehicle in the image.
[0227] According to some embodiments of the present disclosure, when both the local feature extraction network and the global feature extraction network are used in the process of vehicle re-identification (e.g. Figure 8 As shown in the figure, the training process of the global feature extraction network and the local feature extraction network can be trained simultaneously. After calculating the loss function according to the training process of the two networks, the loss function is added according to the following formula (16):
[0228] L=L g +λL r (16)
[0229] Among them, λ is the weight coefficient; L σ is the loss function of the global feature extraction network, L r is the loss function of the local feature extraction network. Finally, the stochastic gradient descent algorithm is used to optimize the parameters of the two networks.
[0230] The present disclosure also provides a training device for a global feature extraction network, for implementing a reference Figure 15 to Figure 18 The training method described. The global feature extraction network 1530 is used to extract the global feature vector of the vehicle in the image. The training device includes a third image acquisition module, a third category acquisition module, a third occlusion image acquisition module, and a third training module.
[0231] Specifically, the third image acquisition module is used to acquire Z third training images 1501 and the identity label of the vehicle in each of the third training images 1501, where Z is an integer greater than or equal to 1.
[0232] The third category acquisition module is used to obtain the component category matrix 151 corresponding to each vehicle in the third training image 1501; the value of each element in the component category matrix 151 represents the component category to which the corresponding pixel position in the image belongs when classified according to the C components of the vehicle, where C is an integer greater than or equal to 2.
[0233] The third occlusion image acquisition module is used to obtain Q partial occlusion images 1502 corresponding to the Q third training images 1501 one by one based on the Q third training images 1501 and the component category matrix 151 corresponding to the vehicles therein, wherein each of the partial occlusion images 1502 is an image in which part of the components in the corresponding third training image is occluded, wherein Q is an integer greater than or equal to 1 and less than or equal to Z.
[0234] The third training module is used to obtain training sample data based on the Z third training images 1501 and the Q local occlusion images 1502 to train the global feature extraction network 1530 .
[0235] Fig.19 The block diagram of a vehicle re-identification device 1900 according to an embodiment of the present disclosure is schematically shown.
[0236] like Fig.19 As shown, the vehicle re-identification device 1900 may include an image acquisition module 1910, an image processing module 1920, and a determination module 1930. The image processing module 1920 may include a vehicle analysis submodule 1921 and a feature vector acquisition submodule 1922. The vehicle re-identification device 1900 may be used to implement reference Figures 1 to 18 The method described.
[0237] The image acquisition module 1910 is used to acquire a first image containing a first vehicle and N second images, each of the second images containing a vehicle, where N is an integer greater than or equal to 1.
[0238] The image processing module 1920 is used to process the first image and each of the N second images to be processed to obtain a feature vector corresponding to the vehicle in each of the images to be processed, wherein the feature vector corresponding to the vehicle in the first image is a first feature vector, and the feature vector corresponding to the vehicle in each of the second images is a second feature vector.
[0239] Specifically, the vehicle parsing submodule 1921 is used to parse the vehicle in the image to be processed according to C components using a vehicle parsing network to obtain component category information corresponding to the vehicle in the image to be processed, wherein the component category information is used to characterize the component category to which each pixel in the image belongs, wherein C is an integer greater than or equal to 2. The feature vector acquisition submodule 1922 is used to obtain a feature vector corresponding to the vehicle in the image to be processed based on the component category information corresponding to the image to be processed.
[0240] The determination module 1930 is used to determine R vehicles corresponding to the R second feature vectors that are most similar to the first feature vector as the first vehicles, where R is an integer greater than or equal to 0 and less than or equal to N.
[0241] According to the embodiments of the present invention, any one or more of the modules, submodules, units, and subunits, or at least part of the functions of any one of them can be implemented in one module. According to the embodiments of the present invention, any one or more of the modules, submodules, units, and subunits can be split into multiple modules for implementation. According to the embodiments of the present invention, any one or more of the modules, submodules, units, and subunits can be at least partially implemented as hardware circuits, such as field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), systems on chips, systems on substrates, systems on packages, application specific integrated circuits (ASICs), or can be implemented by hardware or firmware in any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of them. Alternatively, according to the embodiments of the present invention, one or more of the modules, submodules, units, and subunits can be at least partially implemented as computer program modules, and when the computer program modules are run, the corresponding functions can be performed.
[0242] For example, any multiple of the image acquisition module 1910, the image processing module 1920, the determination module 1930, the vehicle analysis submodule 1921, the feature vector acquisition submodule 1922, the first image acquisition module, the first annotation module, the first training module, the second image acquisition module, the second training module, the third image acquisition module, the third category acquisition module, the third occlusion image acquisition module, and the third training module can be combined into one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the image acquisition module 1910, the image processing module 1920, the determination module 1930, the vehicle analysis submodule 1921, the feature vector acquisition submodule 1922, the first image acquisition module, the first labeling module, the first training module, the second image acquisition module, the second training module, the third image acquisition module, the third category acquisition module, the third occlusion image acquisition module, and the third training module can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging circuits, or can be implemented in any one of the three implementation methods of software, hardware and firmware, or in an appropriate combination of any of them. Alternatively, at least one of the image acquisition module 1910, the image processing module 1920, the determination module 1930, the vehicle analysis sub-module 1921, the feature vector acquisition sub-module 1922, the first image acquisition module, the first labeling module, the first training module, the second image acquisition module, the second training module, the third image acquisition module, the third category acquisition module, the third occlusion image acquisition module, and the third training module can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is executed.
[0243] Fig. 20 The block diagram of a computer system 2000 suitable for implementing the vehicle re-identification method or training method according to an embodiment of the present disclosure is schematically shown. Fig. 20 The computer system shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0244] like Fig. 20As shown, the computer system 2000 according to the embodiment of the present disclosure includes a processor 2001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 2002 or the program loaded from the storage part 2008 to the random access memory (RAM) 2003. The processor 2001 may include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (for example, an application-specific integrated circuit (ASIC)), etc. The processor 2001 may also include an onboard memory for caching purposes. The processor 2001 may include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of the present disclosure.
[0245] In RAM 2003, various programs and data required for the operation of system 2000 are stored. Processor 2001, ROM 2002 and RAM 2003 are connected to each other via bus 2004. Processor 2001 performs various operations of the method flow according to the embodiment of the present disclosure by executing the program in ROM 2002 and / or RAM 2003. It should be noted that the program can also be stored in one or more memories other than ROM 2002 and RAM 2003. Processor 2001 can also perform various operations of the method flow according to the embodiment of the present disclosure by executing the program stored in the one or more memories.
[0246] According to an embodiment of the present disclosure, the system 2000 may further include an input / output (I / O) interface 2005, which is also connected to the bus 2004. The system 2000 may further include one or more of the following components connected to the I / O interface 2005: an input section 2006 including a keyboard, a mouse, etc.; an output section 2007 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 2008 including a hard disk, etc.; and a communication section 2009 including a network interface card such as a LAN card, a modem, etc. The communication section 2009 performs communication processing via a network such as the Internet. A drive 2010 is also connected to the I / O interface 2005 as needed. A removable medium 2011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 2010 as needed, so that a computer program read therefrom is installed into the storage section 2008 as needed.
[0247] According to an embodiment of the present disclosure, the method flow according to an embodiment of the present disclosure can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program contains a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 2009, and / or installed from the removable medium 2011. When the computer program is executed by the processor 2001, the above-mentioned functions defined in the system of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the system, equipment, device, module, unit, etc. described above can be implemented by a computer program module.
[0248] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present disclosure is implemented.
[0249] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus or a device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 2002 and / or RAM 2003 described above and / or one or more memories other than ROM 2002 and RAM 2003.
[0250] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0251] It will be appreciated by those skilled in the art that the features described in the various embodiments and / or claims of the present disclosure may be combined and / or combined in a variety of ways, even if such combinations and / or combinations are not explicitly described in the present disclosure. In particular, the features described in the various embodiments and / or claims of the present disclosure may be combined and / or combined in a variety of ways without departing from the spirit and teachings of the present disclosure. All of these combinations and / or combinations fall within the scope of the present disclosure.
[0252] The embodiments of the present disclosure are described above. However, these embodiments are only for illustrative purposes and are not intended to limit the scope of the present disclosure. Although the embodiments are described above separately, this does not mean that the measures in the various embodiments cannot be used in combination to advantage. The scope of the present disclosure is defined by the attached claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art may make a variety of substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. A vehicle re-identification method, comprising: obtaining a first image including a first vehicle and N second images, each of the second images including a vehicle, where N is an integer greater than or equal to 1; processing each of the first image and the N second images to be processed to obtain a feature vector corresponding to the vehicle in each of the images to be processed, including: The vehicle in the image to be processed is parsed according to C components using a vehicle parsing network to obtain component category information corresponding to the vehicle in the image to be processed, wherein the component category information is used to characterize the component category to which each pixel in the image belongs, wherein C is an integer greater than or equal to 2; wherein the component category information includes a component category matrix, wherein the component category matrix is The matrix of is the width of the image to be processed, H is the height of the image to be processed, and the value of each element in the component category matrix represents the component category to which each pixel position in the image belongs; and obtaining a feature vector corresponding to the vehicle in the image to be processed based on the component category information corresponding to the image to be processed; wherein, the feature vector corresponding to the vehicle in the first image is the first feature vector, and the feature vector corresponding to the vehicle in each of the second images is the second feature vector; and determining the R vehicles corresponding to the R second feature vectors that are most similar to the first feature vector as the first vehicle, where R is an integer greater than or equal to 0 and less than or equal to N; wherein, the vehicle parsing network includes a first convolutional neural network, and using the vehicle parsing network to parse the vehicle in the image to be processed into C components to obtain the component category information corresponding to the vehicle in the image to be processed includes: inputting the image to be processed into the first convolutional neural network; obtaining a feature tensor output by the first convolutional neural network, the feature tensor including information in three dimensions of width W, height H, and category C, wherein, the size information of the image to be processed is output corresponding to the height H and width W dimensions in the feature tensor, and the information of the component category to which each pixel position in the image to be processed belongs is output in the category C dimension; performing statistical processing on the elements in each category in the feature tensor to obtain the category to which each pixel position belongs, including performing statistics on the element values of the category C dimension channels at each pixel position to obtain the value corresponding to each pixel position, and using the component category corresponding to the value as the category to which the pixel position belongs; and converting the feature tensor into the component category matrix based on the category to which each pixel position belongs.
2. The method according to claim 1, wherein, the vehicle parsing network is trained in the following manner: obtaining S first training images for training the vehicle parsing network, each of the first training images including a vehicle, where S is an integer greater than or equal to 1; annotating the component category information corresponding to the vehicle in each of the first training images to obtain the category annotation information of each of the first training images; using the S first training images and the category annotation information of each of the S first training images as training sample data to train the vehicle parsing network.
3. The method according to claim 1, wherein, the obtaining a feature vector corresponding to the vehicle in the image to be processed based on the component category information corresponding to the image to be processed includes: Processing the image to be processed and the component category matrix corresponding to the vehicle in the image to be processed by a local feature extraction network to obtain a local feature vector corresponding to the vehicle in the image to be processed, wherein the local feature vector is a feature vector obtained based on C' feature vectors corresponding to C' components of a vehicle, wherein C' is an integer greater than or equal to 1 and less than or equal to C; and Based on the local feature vector corresponding to the vehicle in the image to be processed, a feature vector corresponding to the vehicle in the image to be processed is obtained.
4. The method according to claim 3, in: The obtaining, based on the component category information corresponding to the image to be processed, a feature vector corresponding to the vehicle in the image to be processed further includes: Processing the image to be processed using a global feature extraction network to obtain a global feature vector corresponding to the vehicle in the image to be processed; The obtaining, based on the local feature vector corresponding to the vehicle in the image to be processed, a feature vector corresponding to the vehicle in the image to be processed comprises: The local feature vector corresponding to the vehicle in the image to be processed is combined with the global feature vector to obtain a feature vector corresponding to the vehicle in the image to be processed.
5. The method according to claim 3 or 4, in, The local feature extraction network includes a second convolutional neural network, and the using of the local feature extraction network to process the image to be processed and the component category matrix corresponding to the vehicle in the image to be processed to obtain the local feature vector corresponding to the vehicle in the image to be processed includes: Processing the image to be processed using the second convolutional neural network to obtain an overall feature matrix corresponding to the vehicle in the image to be processed; Based on the component category matrix corresponding to the vehicle in the image to be processed, a mask matrix of each of the C components of the vehicle in the image to be processed is obtained, wherein the element value corresponding to the pixel position belonging to the component in the mask matrix of one component is a first value, and the element values corresponding to other pixel positions are all second values; and Based on the overall feature matrix corresponding to the vehicle in the image to be processed and the mask matrix of each component of the vehicle in the image to be processed, a component feature vector corresponding to each component of the vehicle in the image to be processed is obtained; wherein C component feature vectors are obtained corresponding to C components; and The local feature vector is obtained based on the C component feature vectors.
6. The method according to claim 5, in, The local feature extraction network further includes a graph convolutional network, and obtaining the local feature vector based on the C component feature vectors further includes: Combining C of the component feature vectors to obtain a first component feature matrix, wherein each row of the first component feature matrix corresponds to one of the component feature vectors; and Processing the first component feature matrix using the graph convolutional network to obtain a second component feature matrix corresponding to the first component feature matrix; and The second component feature matrix is processed to obtain the local feature vector.
7. The method according to claim 6, in, The local feature extraction network is trained in the following way: Obtain T second training images and an identity label of a vehicle in each of the second training images, where T is an integer greater than or equal to 1; Using the T second training images and the identity labels of the vehicles therein, training the second convolutional neural network and the graph convolutional network includes: Processing the second training image using the second convolutional neural network to obtain the overall feature matrix corresponding to the vehicle in the second training image; Obtaining an input of the graph convolutional network based on the overall feature matrix corresponding to the vehicle in the second training image and the component category matrix corresponding to the vehicle in the second training image; Based on the output of the graph convolutional network, obtaining a loss function of the local feature extraction network; and Based on the loss function of the local feature extraction network, the second convolutional neural network and the graph convolutional network are trained.
8. The method according to claim 7, in, The step of obtaining the input of the graph convolution network based on the overall feature matrix corresponding to the vehicle in the second training image and the component category matrix corresponding to the vehicle in the second training image comprises: Based on the component category matrix corresponding to the vehicle in the second training image, obtaining the mask matrix for each of the C components of the vehicle in the second training image; Based on the overall feature matrix corresponding to the vehicle in the second training image and the mask matrix of each component of the vehicle in the second training image, obtaining the component feature vector corresponding to each component of the vehicle in the second training image; wherein C component feature vectors are obtained corresponding to C components; Combining C component feature vectors corresponding to C components of the vehicle in the second training image to obtain a first component feature matrix corresponding to the vehicle in the second training image; Setting the elements of some rows in the first component feature matrix corresponding to the vehicle in the second training image to 0 to obtain a first associated component feature matrix corresponding to the vehicle in the second training image; and The first component feature matrix corresponding to the vehicle in the second training image and the first associated component feature matrix corresponding to the vehicle in the second training image are respectively used as inputs of the graph convolutional network.
9. The method according to claim 8, in, The loss function of the local feature extraction network obtained based on the output of the graph convolutional network includes: obtaining the second component feature matrix corresponding to the vehicle in the second training image output by the graph convolutional network when the first component feature matrix corresponding to the vehicle in the second training image is used as input; obtaining a second associated component feature matrix corresponding to the vehicle in the second training image output by the graph convolutional network when the first associated component feature matrix corresponding to the vehicle in the second training image is used as input; Processing the second component feature matrix corresponding to the vehicle in the second training image to obtain the local feature vector corresponding to the vehicle in the second training image; processing the second associated component feature matrix corresponding to the vehicle in the second training image to obtain an associated local feature vector corresponding to the vehicle in the second training image; and Based on the difference between the local feature vector and the associated local feature vector, a loss function of the local feature extraction network is obtained.
10. The method according to claim 9, in, The step of obtaining a loss function of the local feature extraction network based on a difference between the local feature vector and the associated local feature vector includes: Obtaining a self-supervisory loss function based on a norm after subtracting the local feature vector from the associated local feature vector; and Based on the self-supervised loss function, the loss function is obtained.
11. The method according to claim 4, in, The global feature extraction network is trained in the following way: Obtain Z third training images and an identity label of a vehicle in each of the third training images, where Z is an integer greater than or equal to 1; Obtaining the component category matrix corresponding to each vehicle in the third training image; Based on the Q third training images and the component category matrix corresponding to the vehicles therein, obtaining Q partial occlusion images corresponding to the Q third training images one by one, wherein each of the partial occlusion images is an image in which a part of the component in the corresponding third training image is occluded, wherein Q is an integer greater than or equal to 1 and less than or equal to Z; Training sample data is obtained based on the Z third training images and the Q local occlusion images to train the global feature extraction network.
12. The method according to claim 11, in, The obtaining, based on the Q third training images and the component category matrix corresponding to the vehicles therein, Q partially occluded images corresponding one-to-one to the Q third training images comprises: Based on the component category matrix corresponding to the vehicle in the third training image, obtaining a mask matrix for each of the C components of the vehicle in the third training image; and Based on the third training image and the mask matrix of the partial components in the vehicle in the third training image, the local occlusion image corresponding to the third training image is obtained.
13. The method according to claim 11, in, The obtaining of training sample data based on the Z third training images and the Q local occlusion images to train the global feature extraction network comprises: Determine one third image as a reference sample image each time in a traversal manner, obtain a positive sample image and a negative sample image from the Z third images, so as to form a triplet through the reference sample image, the positive sample image and the negative sample image; the positive sample image is another third training image having the same identity label as the vehicle in the reference sample; the negative sample image is any third training image having a different identity label from the vehicle in the reference sample; Convert any at least one image in the triplet into the corresponding local occlusion image to obtain a converted triplet; and In each traversal, three images in the transformation triplet are distributed as a group of input data and input into the global feature extraction network to train the global feature extraction network.
14. A training method for a vehicle re-identification network, in, The vehicle re-identification network includes a vehicle parsing network, and the vehicle parsing network is used to parse the vehicle in the image according to C component categories, where C is an integer greater than or equal to 2; the training method includes training the vehicle parsing network, including: Acquire S first training images for training the vehicle parsing network, each of the first training images includes a vehicle, where S is an integer greater than or equal to 1; Labeling component category information corresponding to the vehicle in each of the first training images to obtain category labeling information of each of the first training images; wherein the component category information is used to represent the component category to which each pixel in the image belongs when classified according to the C component categories; wherein the component category information includes a component category matrix, and the component category matrix is The matrix of is the width of the first training image, H is the height of the first training image, and the value of each element in the component category matrix represents the component category to which each pixel position in the image belongs; Using the S first training images and the category labeling information of each of the S first training images as training sample data, training the vehicle parsing network; Wherein, the vehicle parsing network includes a first convolutional neural network, and the training of the vehicle parsing network includes: Inputting the first training image into the first convolutional neural network; Obtain a feature tensor output by the first convolutional neural network, the feature tensor including information in three dimensions: width W, height H, and category C, wherein the height H and width W dimensions in the feature tensor correspond to outputting size information of the first training image, and the category C dimension outputs information on the component category to which each pixel position of the first training image belongs; Performing statistical processing on the elements of each category in the feature tensor to obtain the category to which each pixel position belongs, including performing statistical processing on the element values of the category C dimension channel at each pixel position to obtain the value corresponding to each pixel position, and taking the component category corresponding to the value as the category to which the pixel position belongs; and Based on the category to which each pixel position belongs, converting the feature tensor into a category feature matrix to obtain a component category matrix output by the vehicle parsing network; Based on the component category matrix in the component category information annotated on the first training image and the component category matrix output by the vehicle parsing network, a pixel-level cross entropy loss function is used to calculate a model loss value, and a stochastic gradient descent algorithm is used to optimize the vehicle parsing network.
15. The training method according to claim 14, in, The vehicle re-identification network further includes a local feature extraction network, and the local feature extraction network is used to extract a local feature vector corresponding to the vehicle in the image; the local feature extraction network includes a second convolutional neural network and a graph convolutional network, and the training method further includes training the local feature extraction network, including: Obtain T second training images and an identity label of a vehicle in each of the second training images, where T is an integer greater than or equal to 1; Using the T second training images and the identity labels of the vehicles therein, training the second convolutional neural network and the graph convolutional network includes: Processing the second training image using the second convolutional neural network to obtain an overall feature matrix corresponding to the vehicle in the second training image; Obtaining an input of the graph convolutional network based on the overall feature matrix corresponding to the vehicle in the second training image and the component category matrix corresponding to the vehicle in the second training image; Obtaining an output of the graph convolutional network; Based on the output of the graph convolutional network, obtaining a loss function of the local feature extraction network; and Based on the loss function of the local feature extraction network, the second convolutional neural network and the graph convolutional network are trained.
16. The training method according to claim 15, in, The step of obtaining the input of the graph convolution network based on the overall feature matrix corresponding to the vehicle in the second training image and the component category matrix corresponding to the vehicle in the second training image comprises: Based on the component category matrix corresponding to the vehicle in the second training image, a mask matrix corresponding to each of the C components of the vehicle in the second training image is obtained; wherein the element value corresponding to the pixel position belonging to the component in the mask matrix of a component is a first value, and the element values corresponding to other pixel positions are all second values; Based on the overall feature matrix corresponding to the vehicle in the second training image and the mask matrix of each component of the vehicle in the second training image, a component feature vector corresponding to each component of the vehicle in the second training image is obtained; wherein C component feature vectors are obtained corresponding to C components; Combining C component feature vectors corresponding to C components of the vehicle in the second training image to obtain a first component feature matrix corresponding to the vehicle in the second training image; Setting the elements of some rows in the first component feature matrix corresponding to the vehicle in the second training image to 0 to obtain a first associated component feature matrix corresponding to the vehicle in the second training image; and The first component feature matrix corresponding to the vehicle in the second training image and the first associated component feature matrix corresponding to the vehicle in the second training image are respectively used as inputs of the graph convolutional network.
17. The training method according to claim 16, in, The loss function of the local feature extraction network obtained based on the output of the graph convolutional network includes: obtaining a second component feature matrix corresponding to the vehicle in the second training image output by the graph convolutional network when the first component feature matrix corresponding to the vehicle in the second training image is used as input; obtaining a second associated component feature matrix corresponding to the vehicle in the second training image output by the graph convolutional network when the first associated component feature matrix corresponding to the vehicle in the second training image is used as input; Processing the second component feature matrix corresponding to the vehicle in the second training image to obtain the local feature vector corresponding to the vehicle in the second training image; processing the second associated component feature matrix corresponding to the vehicle in the second training image to obtain an associated local feature vector corresponding to the vehicle in the second training image; and Based on the difference between the local feature vector and the associated local feature vector, the loss function of the local feature extraction network is obtained.
18. The training method according to claim 17, wherein, the obtaining the loss function of the local feature extraction network based on the difference between the local feature vector and the associated local feature vector includes: obtaining a self-supervised loss function based on the norm after performing subtraction operation on the local feature vector and the associated local feature vector; and obtaining the loss function based on the self-supervised loss function.
19. The training method according to claim 14, wherein, the vehicle re-identification network further includes a global feature extraction network, and the global feature extraction network is used to extract the global feature vector of the vehicle in the image, and wherein the training method further includes training the global feature extraction network, including: obtaining Z third training images and the identity labels of the vehicles in each of the third training images, where Z is an integer greater than or equal to 1; obtaining the part category matrix corresponding to the vehicle in each of the third training images based on Q of the third training images and the part category matrices corresponding to the vehicles therein, obtaining Q local occlusion images corresponding one-to-one to the Q third training images, where each local occlusion image is an image in which some parts of the corresponding third training image are occluded, and where Q is an integer greater than or equal to 1 and less than or equal to Z; obtaining training sample data based on the Z third training images and the Q local occlusion images to train the global feature extraction network.
20. The training method according to claim 19, wherein, the obtaining Q local occlusion images corresponding one-to-one to the Q third training images based on the Q third training images and the part category matrices corresponding to the vehicles therein includes: obtaining a mask matrix for each of the C parts of the vehicle in the third training image based on the part category matrix corresponding to the vehicle in the third training image; wherein, for a mask matrix of a part, the element value corresponding to the pixel position belonging to the part is a first value, and the element values corresponding to other pixel positions are all second values; and obtaining the local occlusion image corresponding to the third training image based on the third training image and the mask matrices of some parts of the vehicle in the third training image.
21. The training method according to claim 19, wherein, the obtaining training sample data based on the Z third training images and the Q local occlusion images to train the global feature extraction network includes: each time determining a third image as a reference sample image in a traversing manner, obtaining a positive sample image and a negative sample image from the Z third images, so as to form a triple through the reference sample image, the positive sample image and the negative sample image; the positive sample image is another third training image with the same identity label as the vehicle in the reference sample; the negative sample image is any one of the third training images with a different identity label from the vehicle in the reference sample; Convert any at least one image in the triplet into the corresponding local occlusion image to obtain a converted triplet; and In each traversal, three images in the transformation triplet are distributed as a group of input data and input into the global feature extraction network to train the global feature extraction network.
22. A vehicle re-identification device, include: An image acquisition module, configured to acquire a first image including a first vehicle and N second images, each of the second images including a vehicle, wherein N is an integer greater than or equal to 1; An image processing module, used for processing the first image and each of the N second images to be processed to obtain a feature vector corresponding to a vehicle in each of the images to be processed, comprises: a vehicle parsing submodule, configured to parse the vehicle in the image to be processed according to C components using a vehicle parsing network, and obtain component category information corresponding to the vehicle in the image to be processed, wherein the component category information is used to characterize the component category to which each pixel in the image belongs, wherein C is an integer greater than or equal to 2; wherein the component category information includes a component category matrix, wherein the component category matrix is a W×H matrix, wherein W is the width dimension of the image to be processed, and H is the height dimension of the image to be processed, and the value of each element in the component category matrix characterizes the component category to which each pixel position in the image belongs; and A feature vector acquisition submodule, used to obtain a feature vector corresponding to a vehicle in the image to be processed based on the component category information corresponding to the image to be processed; wherein the feature vector corresponding to the vehicle in the first image is a first feature vector, and the feature vector corresponding to each vehicle in the second image is a second feature vector; as well as A determination module, configured to determine R vehicles corresponding to R second feature vectors that are most similar to the first feature vector as the first vehicles, where R is an integer greater than or equal to 0 and less than or equal to N; Wherein, the vehicle parsing network includes a first convolutional neural network, and the vehicle parsing submodule is specifically used for: Inputting the image to be processed into the first convolutional neural network; Obtaining a feature tensor output by the first convolutional neural network, the feature tensor including information in three dimensions: width W, height H, and category C, wherein the height H and width W dimensions in the feature tensor correspond to outputting size information of the image to be processed, and the category C dimension outputs information on the component category to which each pixel position of the image to be processed belongs; Performing statistical processing on the elements of each category in the feature tensor to obtain the category to which each pixel position belongs, including performing statistical processing on the element values of the category C dimension channel at each pixel position to obtain the value corresponding to each pixel position, and taking the component category corresponding to the value as the category to which the pixel position belongs; and Based on the category to which each pixel position belongs, the feature tensor is converted into the component category matrix.
23. A training device for a vehicle re-identification network, in, The vehicle re-identification network includes a vehicle parsing network, and the vehicle parsing network is used to parse the vehicle in the image according to C component categories, where C is an integer greater than or equal to 2; the training device includes: A first image acquisition module, used to acquire S first training images for training the vehicle parsing network, each of the first training images includes a vehicle, where S is an integer greater than or equal to 1; The first annotation module is used to annotate the component category information corresponding to the vehicles in each of the first training images, so as to obtain the category annotation information of each of the first training images; wherein, the component category information is used to represent the component category to which each pixel in the image belongs when classified according to the C component categories; wherein, the component category information includes a component category matrix, and the component category matrix is matrix of is the width dimension of the first training image, H is the height dimension of the first training image, and the value of each element in the component category matrix represents the component category to which each pixel position in the image belongs; A first training module, configured to train the vehicle parsing network using the S first training images and the category labeling information of each of the S first training images as training sample data; The vehicle parsing network includes a first convolutional neural network, and the first training module is specifically used for: Inputting the first training image into the first convolutional neural network; Obtain a feature tensor output by the first convolutional neural network, the feature tensor including information in three dimensions: width W, height H, and category C, wherein the height H and width W dimensions in the feature tensor correspond to outputting size information of the first training image, and the category C dimension outputs information on the component category to which each pixel position of the first training image belongs; Performing statistical processing on the elements of each category in the feature tensor to obtain the category to which each pixel position belongs, including performing statistical processing on the element values of the category C dimension channel at each pixel position to obtain the value corresponding to each pixel position, and taking the component category corresponding to the value as the category to which the pixel position belongs; and Based on the category to which each pixel position belongs, converting the feature tensor into a category feature matrix to obtain a component category matrix output by the vehicle parsing network; Based on the component category matrix in the component category information annotated on the first training image and the component category matrix output by the vehicle parsing network, a pixel-level cross entropy loss function is used to calculate a model loss value, and a stochastic gradient descent algorithm is used to optimize the vehicle parsing network.
24. A vehicle re-identification system, include: one or more memories having computer-executable instructions stored thereon; One or more processors, the processors executing the instructions to implement: The vehicle re-identification method according to any one of claims 1 to 13; or A method for training a vehicle re-identification network according to any one of claims 14 to 21.
25. A computer-readable storage medium having executable instructions stored thereon, which instructions, when executed by a processor, cause the processor to perform: The vehicle re-identification method according to any one of claims 1 to 13; or A method for training a vehicle re-identification network according to any one of claims 14 to 21.
Citation Information
Patent Citations
Similar vehicle identification method and device
CN110097068A
Identification model training and vehicle re-identification method and device based on component segmentation
CN111104867A