Person re-identification method, device, computer equipment and storage medium
By training a personnel identification model containing an unsupervised classifier, multiple feature calculations are performed on video clips, and the problem of relying on manual annotation and insufficient global features in the prior art is solved, and high-precision unsupervised personnel re-identification is achieved.
Patent Information
- Application Number
- CN202111214929.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-19
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-10-19
AI Technical Summary
The personnel re-identification method in the prior art re-reliance relies on manual labeling of body parts of sample data, resulting in high labor costs, and global features are easily occluded or misaligned, and the accuracy is not high.
By training a personnel identification model containing an unsupervised classifier, multiple global features and multiple local features are calculated on video clips, unsupervised local feature extraction is realized, manual labeling costs are reduced, and personnel re-identification accuracy is improved.
Unsupervised local feature extraction is realized, the cost of manual labeling is reduced, the accuracy of personnel re-identification is improved, and the problem of insufficient relying on manual labeling and global features in the prior art is solved.
Smart Images

Figure CN113989835B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of image processing and deep learning technology, and in particular, to a person re-identification method, apparatus, computer equipment and storage medium. Background Art
[0002] Person Re-ID is a basic task in intelligent video surveillance systems. In the field of computer vision, this task refers to determining whether two persons are the same person in views (images or videos) captured by one or more cameras at different times.
[0003] The existing Re-ID method extracts global feature vectors for the entire image and classifies the image based on the global feature vectors through deep learning technology. However, the global feature vectors are easily affected by problems such as occlusion or misalignment of the human body in the image. The Re-ID method based on local features extracts body parts or regional features on the image and identifies body parts through methods such as human body analysis or posture estimation. However, this method relies on manual labeling of body parts in sample data, which has high labor costs. Summary of the invention
[0004] The embodiments of the present invention provide a person re-identification method, apparatus, computer equipment and storage medium to achieve unsupervised local feature extraction, reduce manual labeling costs and improve the accuracy of person re-identification.
[0005] In a first aspect, an embodiment of the present invention provides a method for person re-identification, the method comprising:
[0006] Obtain at least two video clips, and input the video clips into a person recognition model;
[0007] The person identification model includes an unsupervised classifier for mapping body parts of a person to calculate multi-path local features;
[0008] According to the personnel features of each video clip output by the personnel recognition model, the personnel feature distance between two video clips is calculated; the personnel features are obtained by splicing multiple global features and multiple local features;
[0009] If it is determined that the person feature distance satisfies the feature distance condition, it is determined that the person features in the two video clips that match the person feature distance correspond to the same person.
[0010] In a second aspect, an embodiment of the present invention further provides a person re-identification device, the device comprising:
[0011] A video clip input module, used for acquiring at least two video clips and inputting the video clips into a person recognition model;
[0012] The person identification model includes an unsupervised classifier for mapping body parts of a person to calculate multi-path local features;
[0013] A personnel feature distance calculation module is used to calculate the personnel feature distance between two video clips according to the personnel features of each video clip output by the personnel recognition model; the personnel features are obtained by splicing multiple global features and multiple local features;
[0014] The person re-identification module is used to determine that the person features in two video clips that match the person feature distance correspond to the same person if it is determined that the person feature distance meets the feature distance condition.
[0015] In a third aspect, an embodiment of the present invention further provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a person re-identification method as described in any one of the embodiments of the present invention is implemented.
[0016] In a fourth aspect, an embodiment of the present invention further provides a storage medium comprising computer executable instructions, which, when executed by a computer processor, are used to execute a person re-identification method as described in any one of the embodiments of the present invention.
[0017] The embodiment of the present invention trains a person recognition model including an unsupervised classifier, calculates multiple global features and multiple local features of the input video clips, thereby obtaining the person features of the video clips, and calculates the person feature distance between the person features of each video clip. When the person feature distance meets the feature distance condition, it is determined that the person features in the two video clips corresponding to the person feature distance are the same person. The problem of the person re-identification method in the prior art, which relies on the manual labeling of the body parts of the sample data and has high labor costs, is solved, and unsupervised local feature extraction is realized, the manual labeling cost is reduced, and the accuracy of person re-identification is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a flow chart of a person re-identification method in Embodiment 1 of the present invention;
[0019] Figure 2 is a flow chart of a person re-identification method in Embodiment 2 of the present invention;
[0020] Figure 3 is a schematic diagram of the structure of a person re-identification device in Embodiment 3 of the present invention;
[0021] Figure 4 It is a structural diagram of a computer device in Embodiment 4 of the present invention. DETAILED DESCRIPTION
[0022] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.
[0023] Embodiment 1
[0024] Figure 1 This is a flowchart of a person re-identification method provided in Example 1 of the present invention. This embodiment can be applied to determine whether the persons appearing in different video clips are the same person. The method can be executed by a person re-identification device, which can be implemented by software and / or hardware and is generally integrated in a computer device and used in conjunction with a shooting device.
[0025] like Figure 1 As shown, the technical solution of the embodiment of the present invention specifically includes the following steps:
[0026] S110: Obtain at least two video clips, and input the video clips into a person recognition model.
[0027] The person recognition model includes an unsupervised classifier for mapping body parts of a person to calculate multi-path local features.
[0028] In an embodiment of the present invention, an unsupervised classifier is trained in a personnel recognition model. Through the unsupervised classifier, multi-channel local feature extraction can be achieved, which not only makes up for the shortcomings of using only global features, but also does not require labeling information of personnel body parts, reduces the demand for labeling of personnel body parts, and reduces labor costs. At the same time, not relying on manually labeled data also makes the model more versatile.
[0029] S120: Calculate the personnel feature distance between any two video clips according to the personnel features of each video clip output by the personnel recognition model.
[0030] The personnel features are obtained by concatenating multiple global features and multiple local features.
[0031] In an embodiment of the present invention, a video clip is obtained by cropping the portion containing the human body after performing human body detection and target tracking on a video screen. For an input video clip, a person recognition model outputs a person feature, and the person feature is obtained by splicing multiple global features and multiple local features extracted by the person recognition model from the video clip.
[0032] For each person feature corresponding to each video clip, the person feature distance between the person features corresponding to each video clip is calculated respectively. The person feature distance may be Euclidean distance, Manhattan distance or Chebyshev distance, which is not limited in this embodiment. For example, if video clips A1 and A2 correspond to person features a1 and a2, and video clip B corresponds to person feature b, the person feature distance between a1 and b, and the person feature distance between a2 and b are calculated respectively.
[0033] S130: If it is determined that the person feature distance satisfies a feature distance condition, it is determined that the person features in two video clips that match the person feature distance correspond to the same person.
[0034] When a certain person feature distance satisfies the feature distance condition, that is, when it is less than or equal to the preset distance threshold, or when the person feature distance is the minimum, the person features in the two video clips that match the person feature distance correspond to the same person, that is, the person in the two video clips is the same person, thereby achieving person re-identification. Exemplarily, when the feature distance condition is less than or equal to the preset distance threshold, if video clips A1 and A2 correspond to person features a1 and a2, and the person feature distance between a1 and a2 is less than the preset distance threshold, then it is considered that video clips A1 and A2 correspond to the same person. Exemplarily, when the feature distance condition is that the person feature distance is the minimum, if video clips A, B...N correspond to person features a, b...n, for video clip A, the feature distances between a and b, between a and c, and up to a and n are calculated respectively. If the feature distance between a and b is the minimum, then it is considered that video clips B and video clip A correspond to the same person.
[0035] In a specific application scenario of an embodiment of the present invention, by inputting surveillance video segments in a video surveillance system into a personnel recognition model, personnel in the same camera viewing area or in different camera viewing areas are re-identified, thereby realizing the detection of abnormal behavior of personnel.
[0036] In another specific application scenario of the embodiment of the present invention, when it is necessary to re-identify a known person in a certain video clip, the video clip is input into the person recognition model to determine the person features that match the person. Multiple video clips are input into the person recognition model, and the feature distances between the person features that match the person are calculated using the person features of each video clip output by the person recognition model, thereby determining in which video clips the person appears.
[0037] The technical solution of this embodiment is to train a person recognition model including an unsupervised classifier, calculate multiple global features and multiple local features of the input video clips, thereby obtaining the person features of the video clips, and calculate the person feature distance between the person features of two video clips. When the person feature distance meets the feature distance condition, it is determined that the person features in the two video clips corresponding to the person feature distance are the same person. This solves the problem that the person re-identification method in the prior art relies on manual labeling of body parts of sample data and has high labor costs, realizes unsupervised local feature extraction, reduces the cost of manual labeling, and improves the accuracy of person re-identification.
[0038] Embodiment 2
[0039] Figure 2 It is a flowchart of a method for person re-identification provided in Embodiment 2 of the present invention. Based on the above embodiments, this embodiment of the present invention adds a specific process of training a person recognition model and a specific process of obtaining multi-channel global features and multi-channel local features through the person recognition model.
[0040] Correspondingly, such as Figure 2 As shown, the technical solution of the embodiment of the present invention specifically includes the following steps:
[0041] S210. Train a preset machine learning model using sample video clips until the total loss of the machine learning model meets the model training conditions, thereby obtaining a personnel recognition model.
[0042] The total loss includes at least one of the following: triplet loss, cross entropy loss, differentiation loss, attention loss, and human structure loss.
[0043] Exemplarily, the total loss may be the sum of triplet loss, cross entropy loss, differentiation loss, attention loss, and human structure loss.
[0044] In an embodiment of the present invention, a stochastic gradient descent method may be used to train the model until the total loss of the model meets the model training conditions. Stochastic Gradient Descent Algorithm (SGD) is an iterative method for solving the global minimum of the objective function in the negative gradient direction. In an embodiment of the present invention, the minimum value of the total loss of the model is solved by the stochastic gradient descent method. The model training condition may be that the total loss of the model is less than or equal to a preset total loss value, but this embodiment does not limit the model training method and model training conditions.
[0045] In the embodiment of the present invention, the person recognition model is trained through sample video clips, taking the connection relationship of the person's body parts or the person's skeleton structure as prior knowledge.
[0046] S220: Obtain at least two video clips, and input the video clips into a person recognition model.
[0047] Each video clip is input into the person recognition model respectively, and the person features in each video clip are calculated respectively.
[0048] S230: Obtain multi-channel global features of the video clip by calculating a person recognition model.
[0049] In an embodiment of the present invention, multi-channel global features of video clips are calculated, and the accuracy of person re-identification is improved by combining the multi-channel global features with the multi-channel local features, thereby solving the problem of insensitivity to heavy occlusion, background clutter, etc. caused by the single use of global features.
[0050] Accordingly, S230 may include:
[0051] S231, extracting coding features from the target video frame image of the video clip to obtain at least two coding feature maps.
[0052] For example, a 50-layer ResNet deep residual network can be used to extract coding features from each video frame image of the video clip to obtain 5 stages of coding feature maps. Among them, I t is the tth video frame image in the video clip, stage 0 is the input convolution stage (stage 0), Among them, stage i It is the i-th feature convolution stage. However, it should be noted that this embodiment does not limit the neural network used to extract the coding features and the number of layers of the neural network.
[0053] S232. Select a target coding feature map from each coding feature map, perform coding feature enhancement processing on the target coding feature map, and obtain a target enhanced coding feature map.
[0054] In an embodiment of the present invention, taking the example of extracting coding features for each video frame image of a video clip through a 50-layer ResNet deep residual network to obtain 5 stages of coding feature maps, the coding feature maps of the 3rd and 4th stages are selected for processing, but this embodiment does not limit which coding feature maps are specifically selected as target coding feature maps, and the number of selected target coding feature maps.
[0055] Specifically, a TSE (Temporal Saliency Erasing) method may be used to enhance the coding features of a target coding feature map using information between two adjacent video frame images, but this embodiment does not limit the method of enhancing the coding features.
[0056] by For example, in, is the original global feature pair of the t-1th video frame image and the tth video frame image, It is the global feature pair of the t-1th video frame image and the tth video frame image after the encoding feature enhancement.
[0057] S233. Obtain odd features and even features according to the target enhanced coding feature map corresponding to each video frame image, and concatenate the odd features and the even features to obtain multi-channel global features of the video clip.
[0058] After the target coding feature map of each video frame image is coded and enhanced, the parity features of the target enhanced coding feature map of each video frame image are extracted through GAP (Global Average Pooling) and TAP (Temporal Average Pooling).
[0059] Specifically, in, Target enhanced encoding feature map for even video frame images. in, Target enhanced encoding feature maps for odd video frame images.
[0060] This embodiment uses For example, if the encoding feature maps of the third and fourth stages are selected as the target encoding feature maps, then Repeat S233 to obtain and
[0061] The odd and even features of each target enhanced coding feature map are spliced to obtain the multi-channel global features of the video clip. Specifically, Among them, F g represents the global features of the video clip, [...] represents the concatenation operation, and BN(...) represents the batch normalization operation.
[0062] S240. Mapping the block features in each video frame image of the video clip to body parts of the person through an unsupervised classifier in the person recognition model to obtain multi-channel local features of the video clip.
[0063] The mapping of image block features to body parts is achieved in an unsupervised manner, which is not only highly versatile but also can reduce the cost of manual labeling of body parts.
[0064] Accordingly, S240 may include:
[0065] S241, performing block processing on the target coding feature map to obtain multiple block features, and embedding the position information of the block features into the block features.
[0066] by For example, Divided into M block features and concatenated into The Transformer structure is used to embed the location information of the blocks into the block features.
[0067] Specifically, the position information of the t-th video frame image is embedded. Among them, z t The position information embedding vector of the t-th video frame image is E is the word embedding matrix, E pos is the position embedding vector. For each video frame image, z t Fusion, get z 0 . Multi-head self-attention operation is performed through the following formula: l′ =MSA(LN(z l-1 ))+z l-1 , where LN(…) represents layer normalization operation, MSA(…) represents multi-head self-attention operation, and l represents the lth layer of the Transformer network. The fully connected mapping is performed by the following formula: l =MLP(LN(z l′ ))+z l′ , where MLP(…) represents a multi-layer perceptron operation. The output y=LN(z L ), where L is the last layer, and y is decomposed in the time domain to obtain
[0068] S242. Map the block features to body parts of the person through an unsupervised classifier to obtain body part features of the person.
[0069] Accordingly, S242 may include:
[0070] S2421. Calculate the attention matrix of the block features through an unsupervised classifier;
[0071] Specifically, the attention matrix is calculated by the following formula: in, represents the attention matrix, represents transpose, FC(…) is a fully connected layer, Represents the Softmax function operation.
[0072] S2422. Calculate the features of the body parts of the person based on the attention matrix and the block features.
[0073] Specifically, the characteristics of the body parts of the person are calculated by the following formula: N represents the number of body parts of a person.
[0074] S243. Generate a spatial relationship diagram of the person's body structure according to the characteristics of the person's body parts corresponding to each video frame image, and perform a convolution operation on the spatial relationship diagram of the person's body structure to obtain a local feature vector.
[0075] For the tth video frame image, construct the human body structure spatial relationship graph G t =(V t , E t ), where v t is a node, that is, the body part of a person in each video frame image, and Indicates that E t is the edge between nodes, corresponding to the connection relationship between the body parts of the person in each video frame image.
[0076] Calculate the sum of nodes of each video frame image V = ∪ t V t , and the sum of the edges E = U t E t ∪{e s,t},e s,t is the edge between the same person's body parts in different video frame images. Wherein, E includes both the edges of different person's body parts in different video frame images and the edges between the same person's body parts in different video frame images.
[0077] The spatial relationship graph of the body structure of each video frame image G = (V, E) is convolved to achieve the re-integration of the body part features of multiple video frame images to obtain the final local feature vector. This setting enhances the feature representation capability and improves the accuracy of person re-identification.
[0078] Specifically, the graph convolution layer is calculated in, for The corresponding degree matrix, A is the adjacency matrix of the human body structure spatial relationship diagram, and I is the identity matrix.
[0079] The input layer graph convolution is calculated by the following formula: Among them, Relu(…) is a linear rectification function. The hidden layer graph convolution is calculated by the following formula: Here k is the kth hidden layer of the graph convolutional network, and the output layer graph convolution is calculated by the following formula: Where K represents the last hidden layer, and the local feature vector is obtained
[0080] S244, aligning and concatenating the local feature vectors to obtain multi-channel local features of the video clip.
[0081] Accordingly, S244 may include:
[0082] S2441. According to the connection order of the body parts of the person, the local feature vectors are concatenated to obtain a local feature vector that matches the target coding feature map.
[0083] For each local feature vector, we splice them according to the connection relationship between the body parts of the person, and get calculate
[0084] right Repeat S240 to obtain
[0085] S2442. Concatenate the local feature vectors that match the target coding feature maps to obtain multi-channel local features of the video clip.
[0086] By splicing, we can obtain multi-channel local features after fusing the features of multiple body parts in multiple frames:
[0087] It should be noted that during the training of the person recognition model, the triplet loss, cross entropy loss, differentiation loss, attention concentration loss, and human structure loss are calculated in the following way.
[0088] Specifically, i =FC i (x i ), i = 1, 2, ..., 6, where x i Corresponding to
[0089] The triplet loss L is calculated as follows tri :calculate Where a represents the anchor sample, p represents the sample of the same category as a, n represents the sample of a different category from a, m is a constant greater than 0, and d(…) is the cosine similarity function; Calculation Among them, B is the training batch size.
[0090] The cross entropy loss L is calculated as follows xent :calculate Among them, G represents the true person identification, and MSE(…) is the mean square error loss function.
[0091] The differentiation loss L is calculated by div : Among them, ||…|| F is the Frobenius norm.
[0092] Attention loss is calculated by in, yes The elements in .
[0093] The human body structure loss L is calculated by body : Calculate the human body structure matrix: Where D is the distance matrix between image blocks. Calculate the masked human body structure matrix: Where Mask is the human body structure mask matrix, and ⊙ is the Hadamard product. Among them, ||…||2 is the 2-norm, s g is the global human body structure matrix.
[0094] S250: Calculate the personnel feature distance between any two video clips according to the personnel features of each video clip output by the personnel recognition model.
[0095] Personnel characteristics are F g and F p Splicing to get, F = [F g , F p ], for each video clip, calculate the corresponding personnel features respectively, and for the personnel features between two video clips, calculate the personnel feature distance d xy =d(F x , F y ), where F x and F y are the person features belonging to the two video clips respectively.
[0096] S260, determining whether the personnel characteristic distance meets the characteristic distance condition, if so, executing S270, otherwise executing S280.
[0097] When the person feature distance is less than or equal to the preset distance threshold, or the person feature distance is the smallest, it indicates that the person features in the two video clips matching the person feature distance correspond to the same person, thereby achieving person re-identification.
[0098] S270: Determine that the personnel features in the two video clips that match the personnel feature distances correspond to the same person.
[0099] S280, end.
[0100] Embodiment 3
[0101] Figure 3 3 is a schematic diagram of the structure of a person re-identification device provided in the third embodiment of the present invention, the device comprises: a video clip input module 310, a person feature distance calculation module 320 and a person re-identification module 330. Among them:
[0102] A video clip input module 310, for acquiring at least two video clips and inputting the video clips into a person recognition model;
[0103] The person identification model includes an unsupervised classifier for mapping body parts of a person to calculate multi-path local features;
[0104] A personnel feature distance calculation module 320 is used to calculate the personnel feature distance between two video clips according to the personnel features of each video clip output by the personnel recognition model; the personnel features are obtained by splicing multiple global features and multiple local features;
[0105] The person re-identification module 330 is configured to determine that the person features in two video clips that match the person feature distance correspond to the same person if it is determined that the person feature distance satisfies a feature distance condition.
[0106] The technical solution of this embodiment is to train a person recognition model including an unsupervised classifier, calculate multiple global features and multiple local features of the input video clips, thereby obtaining the person features of the video clips, and calculate the person feature distance between the person features of two video clips. When the person feature distance meets the feature distance condition, it is determined that the person features in the two video clips corresponding to the person feature distance are the same person. This solves the problem that the person re-identification method in the prior art relies on manual labeling of body parts of sample data and has high labor costs, realizes unsupervised local feature extraction, reduces the cost of manual labeling, and improves the accuracy of person re-identification.
[0107] Based on the above embodiment, the device further includes:
[0108] A personnel recognition model training module is used to train a preset machine learning model through sample video clips until the total loss of the machine learning model meets the model training conditions, thereby obtaining a personnel recognition model;
[0109] The total loss includes at least one of the following: triplet loss, cross entropy loss, differentiation loss, attention loss, and human structure loss.
[0110] Based on the above embodiment, the device further includes:
[0111] A multi-channel global feature calculation module, used for calculating the multi-channel global features of the video clip through a personnel recognition model;
[0112] The multi-channel local feature calculation module is used to map the block features in each video frame image of the video clip to the body parts of the person through the unsupervised classifier in the person recognition model to obtain the multi-channel local features of the video clip.
[0113] Based on the above embodiment, the multi-channel global feature calculation module includes:
[0114] A coding feature extraction unit, used to extract coding features from a target video frame image of a video clip to obtain at least two coding feature maps;
[0115] A coding feature enhancement unit is used to select a target coding feature map from each coding feature map, perform coding feature enhancement processing on the target coding feature map, and obtain a target enhanced coding feature map;
[0116] The odd-even feature extraction unit is used to obtain odd features and even features according to the target enhanced coding feature map corresponding to each video frame image, and to splice the odd features and even features to obtain the multi-channel global features of the video clip.
[0117] Based on the above embodiment, the multi-channel local feature calculation module includes:
[0118] The target coding feature map block unit is used to perform block processing on the target coding feature map to obtain multiple block features, and embed the position information of the block features into the block features;
[0119] A body part mapping unit is used to map the block features to body parts of a person through an unsupervised classifier to obtain body part features of the person;
[0120] A local feature vector acquisition unit is used to generate a spatial relationship diagram of a person's body structure according to the characteristics of the person's body parts corresponding to each video frame image, and perform a convolution operation on the spatial relationship diagram of the person's body structure to obtain a local feature vector;
[0121] The multi-channel local feature acquisition unit is used to align and splice the local feature vectors to obtain the multi-channel local features of the video clip.
[0122] Based on the above embodiment, the body part mapping unit is specifically used for:
[0123] Through the unsupervised classifier, the attention matrix of the block features is calculated;
[0124] Based on the attention matrix and block features, the features of the person’s body parts are calculated.
[0125] On the basis of the above embodiment, the multi-channel local feature acquisition unit is specifically used for:
[0126] According to the connection order of the body parts of the person, the local feature vectors are concatenated to obtain the local feature vectors matching the target encoding feature map;
[0127] The local feature vectors matching each target encoding feature map are concatenated to obtain multi-channel local features of the video clip.
[0128] The person re-identification device provided in the embodiment of the present invention can execute the person re-identification method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0129] Embodiment 4
[0130] Figure 4 A schematic diagram of the structure of a computer device provided in Embodiment 4 of the present invention is shown in FIG. Figure 4 As shown, the computer device includes a processor 70, a memory 71, an input device 72 and an output device 73; the number of processors 70 in the computer device can be one or more. Figure 4 A processor 70 is taken as an example; the processor 70, memory 71, input device 72 and output device 73 in the computer device can be connected by a bus or other means. Figure 4 The example of connecting through bus is taken in the following.
[0131] The memory 71 is a computer-readable storage medium that can be used to store software programs, computer executable programs and modules, such as the modules corresponding to the person re-identification method in the embodiment of the present invention (for example, the video clip input module 310, the person feature distance calculation module 320 and the person re-identification module 330 in the person re-identification device). The processor 70 executes various functional applications and data processing of the computer device by running the software programs, instructions and modules stored in the memory 71, that is, implements the above-mentioned person re-identification method. The method includes:
[0132] Obtain at least two video clips, and input the video clips into a person recognition model;
[0133] The person identification model includes an unsupervised classifier for mapping body parts of a person to calculate multi-path local features;
[0134] According to the personnel features of each video clip output by the personnel recognition model, the personnel feature distance between two video clips is calculated; the personnel features are obtained by splicing multiple global features and multiple local features;
[0135] If it is determined that the person feature distance satisfies the feature distance condition, it is determined that the person features in the two video clips that match the person feature distance correspond to the same person.
[0136] The memory 71 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and at least one application required for a function; the data storage area may store data created according to the use of the terminal, etc. In addition, the memory 71 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 71 may further include a memory remotely arranged relative to the processor 70, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0137] The input device 72 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the computer device. The output device 73 may include a display device such as a display screen.
[0138] Embodiment 5
[0139] Embodiment 5 of the present invention further provides a storage medium containing computer executable instructions, wherein the computer executable instructions are used to execute a person re-identification method when executed by a computer processor, the method comprising:
[0140] Obtain at least two video clips, and input the video clips into a person recognition model;
[0141] The person identification model includes an unsupervised classifier for mapping body parts of a person to calculate multi-path local features;
[0142] According to the personnel features of each video clip output by the personnel recognition model, the personnel feature distance between two video clips is calculated; the personnel features are obtained by splicing multiple global features and multiple local features;
[0143] If it is determined that the person feature distance satisfies the feature distance condition, it is determined that the person features in the two video clips that match the person feature distance correspond to the same person.
[0144] Of course, the computer executable instructions of a storage medium including computer executable instructions provided in an embodiment of the present invention are not limited to the operations of the method described above, but can also execute related operations in the person re-identification method provided in any embodiment of the present invention.
[0145] Through the above description of the implementation methods, the technicians in the relevant field can clearly understand that the present invention can be implemented by means of software and necessary general hardware, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (FLASH), a hard disk or an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0146] It is worth noting that in the embodiment of the above-mentioned person re-identification device, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.
[0147] Note that the above are only preferred embodiments of the present invention and the technical principles used. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present invention, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A person re-identification method, characterized in that: include: Obtain at least two video clips, and input the video clips into a person recognition model; The person recognition model includes an unsupervised classifier for mapping body parts of a person to calculate multi-path local features, wherein the unsupervised classifier is trained in an unsupervised manner without labeling body parts of the person; According to the personnel features of each video clip output by the personnel recognition model, the personnel feature distance between two video clips is calculated; the personnel features are obtained by splicing multiple global features and multiple local features; If it is determined that the person feature distance satisfies the feature distance condition, determining that the person features in the two video clips that match the person feature distance correspond to the same person; After the video clip is input into the person recognition model, the method further includes: Obtaining multi-channel global features of the video clip by calculating a person recognition model; By using an unsupervised classifier in a person recognition model, the block features in each video frame image of the video clip are mapped to the body parts of the person, so as to obtain multi-channel local features of the video clip; The multi-channel global features of the video clip are obtained by calculating the person recognition model, including: Extracting coding features from the target video frame image of the video clip to obtain at least two coding feature maps; Select a target coding feature map from each coding feature map, perform coding feature enhancement processing on the target coding feature map, and obtain a target enhanced coding feature map; The odd and even features of the target enhanced coding feature map of each video frame image are extracted by global average pooling and time domain average pooling to obtain odd features and even features, and the odd features and even features are spliced to obtain multi-channel global features of the video clip, wherein the odd features are the target enhanced coding feature map of the odd video frame image, and the even features are the target enhanced coding feature map of the even video frame image.
2. The method according to claim 1, characterized in that: Before acquiring at least two video clips, it also includes: Using sample video clips, a preset machine learning model is trained until the total loss of the machine learning model meets the model training conditions, thereby obtaining a person recognition model; The total loss includes at least one of the following: triplet loss, cross entropy loss, differentiation loss, attention loss, and human structure loss.
3. The method according to claim 1, characterized in that Through the unsupervised classifier in the person recognition model, the block features in each video frame image of the video clip are mapped to the body parts of the person to obtain multi-channel local features of the video clip, including: The target coding feature map is processed into blocks to obtain a plurality of block features, and the position information of the block features is embedded into the block features; Through an unsupervised classifier, the block features are mapped to the body parts of the person to obtain the body part features of the person; Generate a spatial relationship diagram of the body structure of the person according to the body part features of the person corresponding to each video frame image, and perform a convolution operation on the spatial relationship diagram of the body structure of the person to obtain a local feature vector; The local feature vectors are aligned and concatenated to obtain multi-channel local features of the video clips.
4. The method according to claim 3, characterized in that Through the unsupervised classifier, the block features are mapped to the body parts of the person to obtain the body part features of the person, including: Through the unsupervised classifier, the attention matrix of the block features is calculated; Based on the attention matrix and block features, the features of the person’s body parts are calculated.
5. The method according to claim 3, characterized in that: Align and concatenate the local feature vectors to obtain multiple local features of the video clip, including: According to the connection order of the body parts of the person, the local feature vectors are concatenated to obtain the local feature vectors matching the target encoding feature map; The local feature vectors matching each target encoding feature map are concatenated to obtain multi-channel local features of the video clip.
6. A person re-identification device, characterized in that: include: A video clip input module, used for acquiring at least two video clips and inputting the video clips into a person recognition model; The person recognition model includes an unsupervised classifier for mapping body parts of a person to calculate multi-path local features, wherein the unsupervised classifier is trained in an unsupervised manner without labeling body parts of the person; A personnel feature distance calculation module is used to calculate the personnel feature distance between two video clips according to the personnel features of each video clip output by the personnel recognition model; the personnel features are obtained by splicing multiple global features and multiple local features; A person re-identification module, configured to determine that the person features in two video clips matching the person feature distances correspond to the same person if it is determined that the person feature distance satisfies a feature distance condition; The device further comprises: A multi-channel global feature calculation module, used for calculating the multi-channel global features of the video clip through a personnel recognition model; A multi-channel local feature calculation module is used to map the block features in each video frame image of the video clip to the body parts of the person through the unsupervised classifier in the person recognition model to obtain the multi-channel local features of the video clip; The multi-path global feature calculation module comprises: A coding feature extraction unit, used to extract coding features from a target video frame image of a video clip to obtain at least two coding feature maps; A coding feature enhancement unit is used to select a target coding feature map from each coding feature map, perform coding feature enhancement processing on the target coding feature map, and obtain a target enhanced coding feature map; An odd-even feature extraction unit is used to extract odd-even features from the target enhanced coding feature map of each video frame image through global average pooling and time domain average pooling to obtain odd features and even features, and to splice the odd features and even features to obtain multi-channel global features of the video clip, wherein the odd features are the target enhanced coding feature maps of odd video frame images, and the even features are the target enhanced coding feature maps of even video frame images.
7. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the person re-identification method as described in any one of claims 1-5 is implemented.
8. A storage medium containing computer executable instructions, characterized in that: The computer executable instructions are used to execute the person re-identification method as described in any one of claims 1 to 5 when executed by a computer processor.
Citation Information
Patent Citations
Pedestrian re-identification method and system based on part attention
CN111259837A