Re-identification method, training method of target re-identification network and related devices
By extracting the static images and moving images of the moving targets, calculating the similarity of image features, the problem of insufficient accuracy of re-recognition of moving targets in the prior art is solved, and a higher accuracy of re-recognition is achieved.
Patent Information
- Application Number
- CN202110768934.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-07
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-07-07
AI Technical Summary
The existing target re-identification methods still have room for improvement in the accuracy of the moving target re-identification results.
By performing feature extraction of static images and moving images of the recognition target and candidate target, the similarity of image features is calculated using the target re-identification model to determine the re-identification recognition result.
The accuracy of the results of moving target re-identification is improved, and the representation ability of image features is enhanced by combining texture features and motion features.
Smart Images

Figure CN113705329B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision, and particularly to a method for re-identifying moving targets, a method for training a target re-identification network, an electronic device, and a computer-readable storage medium. Background Art
[0002] To meet the increasing public safety needs, intelligent monitoring technology has emerged. Intelligent monitoring is applicable to various moving targets, such as people, various animals (such as cows, pigs), etc. It plays a key role in the field of public safety. As an important branch of intelligent monitoring, target re-identification has also attracted more and more attention from researchers.
[0003] Target re-identification is a technology that uses computer vision technology to retrieve whether a specific target exists in an image or a video sequence. However, the accuracy of the re-identification results obtained by current target re-identification methods still needs to be improved. Summary of the Invention
[0004] The present application provides a method for re-identifying moving targets, a method for training a target re-identification network, an electronic device, and a computer-readable storage medium, which can improve the accuracy of the re-identification results of moving targets.
[0005] To solve the above technical problems, one technical solution adopted by the present application is: to provide a method for re-identifying moving targets. The re-identification method includes performing the following processing using a target re-identification model: respectively extracting features from the static images and moving images of the target to be identified and multiple candidate targets to obtain image features, where the moving images are used to represent the motion information of each pixel point of the static images; calculating the similarity between the image features of the target to be identified and the image features of the multiple candidate targets; and determining the re-identification result of the target to be identified from the multiple candidate targets based on the similarity.
[0006] To solve the above technical problems, one technical solution adopted by the present application is: to provide a method for training a target re-identification network. The training method includes: extracting training image features from the training static images and training moving images of the moving target using a target re-identification model, where the training moving images are used to represent the motion information of each pixel point of the training static images; classifying using the target re-identification model based on the training image features to obtain a classification result; and adjusting the parameters of the target re-identification model based on the classification result.
[0007] To solve the above technical problems, another technical solution adopted by the present application is: to provide an electronic device, which includes a processor and a memory connected to the processor. The memory stores program instructions; the processor is configured to execute the program instructions stored in the memory to implement the above method.
[0008] To solve the above technical problems, another technical solution adopted by this application is: to provide a computer-readable storage medium storing program instructions that can implement the above method when executed.
[0009] In the above manner, since this application extracts features from static images and dynamic images of the same target, compared with the method of only extracting features from static images, the extracted image features include not only texture features but also motion features, so that the representation ability of the image features is stronger, and thus the re-identification result of the target to be identified based on the image features is more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a schematic flowchart of the first embodiment of the method for re-identifying moving targets in this application;
[0011] Figure 2 is a schematic flowchart of the second embodiment of the method for re-identifying moving targets in this application;
[0012] Figure 3 is Figure 2 the specific flowchart of S21 in
[0013] Figure 4 is Figure 2 the specific flowchart of S22 in
[0014] Figure 5 is a schematic flowchart of the third embodiment of the method for re-identifying moving targets in this application;
[0015] Figure 6 is a schematic flowchart of the fourth embodiment of the method for re-identifying moving targets in this application;
[0016] Figure 7 is a schematic flowchart of the fifth embodiment of the method for re-identifying moving targets in this application;
[0017] Figure 8 is a schematic structural diagram of the target re-identification network in this application;
[0018] Figure 9 is a schematic diagram of the key point detection result in RGB;
[0019] Figure 10 is a schematic diagram of the occlusion situation of the local feature map;
[0020] Figure 11 is a schematic flowchart of the first embodiment of the training method of the target re-identification network in this application;
[0021] Figure 12 is a schematic structural diagram of an embodiment of an electronic device in this application;
[0022] Figure 13 It is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners
[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0024] The terms "first", "second", and "third" in the present application are only used for descriptive purposes, and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of the present application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
[0025] Referring to "embodiment" in this article means that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that, without conflict, the embodiments described herein may be combined with other embodiments.
[0026] Figure 1 It is a schematic flowchart of the first embodiment of the method for re-identifying moving targets in the present application. It should be noted that if there are substantially the same results, this embodiment does not limit Figure 1 to the process sequence shown. As Figure 1 shown, this embodiment may include:
[0027] S11: Extract features from the static images and moving images of the target to be identified and multiple candidate targets respectively to obtain image features.
[0028] Among them, the moving image is used to characterize the motion information of each pixel point of the static image.
[0029] The method provided by this application can be implemented using a target re-identification model. The static image is a single image in the video sequence obtained by a camera device, and the color space to which the static image belongs can be RGB, HSV, SILTP, etc. The motion information of each pixel point of the static image is the motion information between this static image and a static image with an earlier time in the video sequence. The motion image can be an optical flow image or other images obtained by processing the optical flow image.
[0030] The database includes static images of multiple candidate targets obtained historically. The image features of the static images of the candidate targets can be pre-extracted or synchronously extracted with the image features of the static images of the target to be identified.
[0031] The static image contains texture information, and the motion image contains motion information. Thus, the image features cover texture features and motion features.
[0032] S12: Calculate the similarity between the image features of the target to be identified and the image features of multiple candidate targets.
[0033] The similarity can be the cosine distance, Euclidean distance, Hamming distance, etc. between the image features, and no specific limitation is made here.
[0034] S13: Determine the re-identification result of the target to be identified from multiple candidate targets based on the similarity.
[0035] The static images of multiple candidate targets can be arranged in descending order according to the corresponding similarity. The candidate target in Top-1 is the target that best matches the target to be identified and can be used as the re-identification result. Thus, the re-identification of the moving target is completed.
[0036] Through the implementation of this embodiment, since this application extracts features from the static images and dynamic images of the same target, compared with the method of only extracting features from static images, the extracted image features not only contain texture features but also contain motion features. Therefore, the representation ability of the image features is stronger, and further, the re-identification result of the target to be identified determined based on the image features is more accurate.
[0037] In a specific embodiment, the image features involved in S11 may include an overall feature map. Thus, in this step, feature extraction can be performed on the splicing result of the static image and the motion image to obtain an overall feature map. Taking the static image as an RGB image and the motion image as an optical flow image as an example, horizontal optical flow extraction and vertical optical flow extraction can be respectively performed on the static image to obtain a first motion image and a second motion image, where the method of optical flow extraction includes but is not limited to the HS optical flow method; the splicing result of the static image, the first motion image, and the second motion image is input into a pre-trained five-channel feature extraction network to obtain an overall feature map.
[0038] As a specific example, the five-channel feature extraction network can be Resnet-50 without the global average pooling layer and the fully connected layer. The number of channels of the convolutional kernels in its first convolutional layer is 5, and the Stride of the last convolutional layer is 1 to increase the resolution of the overall feature map. When the resolution of the splicing result of the RGB image and the motion image is H×W, the resolution of the feature map obtained through the five-channel feature extraction network is H / 16×W / 16.
[0039] It can be understood that the method of using the feature extraction network to extract features from the RGB image and the motion image reduces the number of channels of the image features compared with the method of using the feature extraction network to extract features from multiple static images in different color spaces, thereby reducing the computational amount of the feature extraction network and making the feature extraction network lighter and faster.
[0040] In another specific embodiment, the image features may include key point feature vectors obtained based on the overall feature map. In this case, the above-mentioned Embodiment 1 can be further extended to obtain Embodiment 2. Specifically as follows:
[0041] Figure 2 It is a schematic flowchart of Embodiment 2 of the re-identification method for moving targets in this application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 2 the shown process sequence. This embodiment is a further extension of S11. As Figure 2 shown, this embodiment may include:
[0042] S21: Extract features from the splicing result of the static image and the motion image to obtain an overall feature map; and extract key points belonging to the target to be recognized and the candidate target based on the static image and the motion image.
[0043] The confidence of each pixel point as a key point (key point confidence) in the static image can be obtained first, and the confidence of each pixel point as a foreground point (foreground point confidence) can be obtained based on the motion image, and then the key points can be determined based on the key point confidence and / or the foreground point confidence of each pixel point. Among them, the key points can be pixel points whose corresponding key point confidence is greater than the first threshold and / or the foreground point confidence is greater than the second threshold.
[0044] It can be understood that if only the key point confidence being greater than the first threshold is used as the condition for determining key points, it is very likely to misidentify the pixel points at the occluders where the texture information is similar to the moving object in the static image as key points. And if only the foreground point confidence being greater than the second threshold is used as the condition for determining key points, it is very likely to misidentify the pixel points at the occluders with a large motion rate in the static image as key points. Therefore, using the key point confidence being greater than the first threshold and the foreground point confidence being greater than the second threshold as the determination key conditions has higher accuracy.
[0045] Refer to in combination Figure 3 , in the case where the key points are pixel points with the key point confidence being greater than the first threshold and / or the foreground point confidence being greater than the second threshold, S21 can be further expanded into the following sub-steps:
[0046] S211: Perform pose evaluation on the static image to obtain the key point confidence of each pixel point; and perform normalization processing on the motion information of the motion image to obtain the foreground point confidence of each pixel point.
[0047] The key point detection of the static image can be performed through a pose estimation network to obtain the key point confidence of each pixel point.
[0048] The motion information of the motion image is obtained from the motion information of the first motion image and the motion information of the second motion image. For example, if the motion information of the first motion image is U and the motion information of the second motion image is V, then the motion information of the motion image is
[0049] The motion information of the motion image reflects the motion rate of each pixel point. The foreground point confidence of the pixel point is the result of the normalization processing. The higher the foreground point confidence of the pixel point, the greater the motion rate, and the more likely it is to be a foreground point; conversely, the lower the foreground point confidence of the pixel point, the smaller the motion rate, and the more likely it is to be a background point.
[0050] S212: Screen out the pixel points with the key point confidence being greater than or equal to the first threshold and the foreground point confidence being greater than or equal to the second threshold as key points.
[0051] It can be understood that the pixel points with the key point confidence being less than the first threshold and / or the foreground point confidence being greater than or equal to the second threshold are the occluded pixel points in the static image.
[0052] Among them, if the key point confidence of a pixel is less than the first threshold and the foreground point confidence is less than the second threshold, it is determined that the pixel is occluded by a static and texture - dissimilar occluder, such as a fence. If the key point confidence of a pixel is greater than the first threshold and the foreground point confidence is less than the second threshold, it is determined that the pixel is occluded by a static and texture - similar occluder, such as a billboard containing a target. If the key point confidence of a pixel is less than the first threshold and the foreground point confidence is greater than the set threshold, it is determined that the pixel is occluded by a dynamic and texture - dissimilar occluder, such as a moving car.
[0053] S22: Determine the key - point feature vector based on the key points and the overall feature map.
[0054] Combined with reference to Figure 4 , S22 may include the following sub - steps:
[0055] S221: Set the response degree of the key points to be equal to the key - point confidence of the key points, and set the response degrees of other pixel points except the key points to zero to obtain a response - degree image.
[0056] The process of obtaining the response - degree image can be reflected by the following formula:
[0057]
[0058] Among them, x represents the abscissa of the pixel point, y represents the ordinate of the pixel point, C xy represents the key - point confidence, UV xy represents the foreground - point confidence, η 1 represents the first threshold, η 2 represents the second threshold.
[0059] S222: Multiply the response - degree image and the overall feature image point - by - point to obtain the key - point features.
[0060] Point - by - point multiplication means multiplying the corresponding feature values in the response - degree image and the overall feature map.
[0061] S223: Perform max - pooling on the key - point features to obtain the max - pooling result; and perform average - pooling on the overall feature map to obtain the average - pooling result.
[0062] S224: Concatenate the max - pooling result and the average - pooling result to obtain the key - point feature vector.
[0063] In addition, based on the above - mentioned Embodiment 2, S12 can be further extended to S23, and S13 can be extended to S24. Thus, the following Embodiment 3 is obtained:
[0064] Figure 5It is a schematic flowchart of the third embodiment of the method for re-identifying moving objects in this application. It should be noted that if there are substantially the same results, this embodiment does not limit to Figure 5 the process sequence shown. For example, Figure 5 as shown, this embodiment may include:
[0065] S23: Calculate the first similarity between the key-point feature vectors of the target to be identified and the key-point feature vectors of multiple candidate targets.
[0066] The features corresponding to the key-point feature vectors of the target to be identified and the candidate targets can be aligned one by one. Therefore, the calculated first similarity is beneficial to obtaining the subsequent re-identification result.
[0067] S24: Determine the re-identification result of the target to be identified based at least on the first similarity.
[0068] The re-identification result of the target to be identified can be determined only based on the first similarity, or the re-identification result of the target to be identified can be determined based on the summation result of the first similarity and the second similarity mentioned in the subsequent embodiment. The re-identification result determined by the latter is more accurate.
[0069] In yet another specific implementation, the image features may include local feature maps obtained by dividing the overall feature map. In this case, the above-mentioned Embodiment 1 can be further extended to obtain Embodiment 4. Specifically as follows:
[0070] Figure 6 It is a schematic flowchart of the fourth embodiment of the method for re-identifying moving objects in this application. It should be noted that if there are substantially the same results, this embodiment does not limit to Figure 6 the process sequence shown. For example, Figure 6 as shown, this embodiment may include:
[0071] S31: Extract features from the splicing result of the static image and the moving image to obtain an overall feature map.
[0072] S32: Divide the overall feature maps of the target to be identified and the candidate targets into multiple corresponding local feature maps along a preset direction, and extract local feature vectors.
[0073] The preset direction can be the horizontal direction or the vertical direction. For example, divide the overall feature map into 6 local feature maps along the preset direction, and extract the corresponding local feature vectors.
[0074] Based on the above-mentioned Embodiment 4, S12 can be further extended to obtain S33 - S34, and S13 can be further extended to obtain S35. Thus, the following Embodiment 5 is obtained:
[0075] Figure 7 This is a schematic flowchart of the fifth embodiment of the method for re-identifying moving objects in this application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 7 the process sequence shown. For example, Figure 7 as shown, this embodiment may include:
[0076] S33: Determine the occlusion situation of the local feature map based on the foreground point confidence and key point confidence of each pixel point of the static images of the target to be identified and the candidate targets.
[0077] In a specific implementation, if there are pixel points in the local feature map where the key point confidence is less than the first threshold and / or the foreground point confidence is less than the second threshold (there are occluded key points), it is determined that the local feature map is occluded; otherwise, it is determined that the local feature map is not occluded.
[0078] In another specific implementation, if there are pixel points in the local feature map where the key point confidence is greater than or equal to the first threshold and the foreground point confidence is greater than or equal to the second threshold (i.e., there are key points), it is determined that the local feature map is not occluded; otherwise, it is determined that the local feature map is occluded.
[0079] S34: When the corresponding local feature maps are not occluded, calculate the second similarity between the local feature vectors of the corresponding local feature maps and sum them up.
[0080] For example, if 3 out of 6 local feature maps are occluded, calculate the second similarity between the local feature vectors of the 3 unoccluded local feature maps.
[0081] S35: Determine the re-identification result of the target to be identified based at least on the sum result of the second similarity.
[0082] The re-identification result of the target to be identified can be determined only based on the sum result of the second similarity, or the re-identification result of the target to be identified can be determined based on the sum result of the second similarity and the first similarity mentioned in the previous embodiments. The re-identification result determined by the latter is more accurate.
[0083] In S24 / S35, if the re-identification result of the target to be identified is determined based on the sum result of the second similarity and the first similarity, then the sum result of the second similarity and the first similarity can be summed up twice, and the twice-summed result can be divided by the number of the second similarities to obtain the third similarity; the re-identification result of the target to be identified is determined based on the third similarity. The specific implementation process can be based on the following formula:
[0084]
[0085] Among them, Similarityall Denotes the third similarity, Similarity i Denotes the second similarity between the i-th local feature vectors; if the i-th local feature map is occluded, then I i is 0, otherwise I i is 1; Similarity keypoint Denotes the first similarity.
[0086] The re-identification method provided by the present application will be described in detail below through an example.
[0087] Example 1:
[0088] With reference to Figure 8 , Figure 8 which is a schematic structural diagram of the target re-identification network of the present application. As Figure 8 shown, the static image of the target (pedestrian) is denoted as RGB, the first motion image (horizontal optical flow image) is denoted as U, the second motion image (vertical optical flow image) is denoted as V, and the concatenation result of RGB, U, and V is denoted as UVRGB. The target re-identification model includes a feature extraction backbone network (i.e., the five-channel feature extraction network mentioned above) and a key point feature extraction network.
[0089] Input UVRGB into the feature extraction backbone network to obtain an overall feature map; divide the overall feature map horizontally into 6 local feature maps, and perform average pooling and 1×1 convolution processing on the 6 local feature maps respectively to obtain corresponding local feature vectors.
[0090] Input RGB, U, and V into the key point feature extraction network, perform pose estimation on RGB to obtain the key point confidence of each pixel point therein, and use U and V to obtain the motion image UV, and perform normalization processing on the motion information of the motion image UV to obtain the foreground point confidence of each pixel point.
[0091] Among them, the pixel points whose key point confidence is greater than or equal to the first threshold and the foreground point confidence is greater than or equal to the second threshold are used as key points. Figure 9 is a schematic diagram of the key point detection result in RGB, where both A and B represent the key points of the pedestrian, and A represents the detected key point, and B represents the key point that is occluded by an occluding object and not detected.
[0092] Set the responsiveness of the key points to be equal to the key point confidence of the key points, and set the responsiveness of other pixel points outside the key points to zero to obtain a responsiveness image. Multiply the responsiveness image and the overall feature map, perform max pooling to obtain the max pooling result, perform average pooling on the overall feature map to obtain the average pooling result, and splice the max pooling result and the average pooling result to obtain the key point feature vector. Calculate the first similarity between the key point feature vectors corresponding to the target to be recognized and the candidate target.
[0093] Determine the occluded local feature map for the local feature map whose corresponding region has pixel points with a key point confidence less than the first threshold and / or a foreground point confidence less than the second threshold; and determine other local feature maps as unoccluded local feature maps. Calculate the sum of the second similarities between the unoccluded local feature vectors corresponding to the local feature maps of the target to be recognized and the candidate target. Figure 10 It is a schematic diagram of the occlusion situation of the local feature map, where C is the static image of the target to be recognized, C' is the overall feature map of the target to be recognized, which is divided into 6 local feature maps, D is the static image of the candidate target, and D' is the overall feature map of the candidate target. The upper 3 local feature maps in C are unoccluded in the corresponding regions, and the 6 local feature maps in D are unoccluded, so the upper three local feature maps are the unoccluded local feature maps in both C and D. Calculate the sum of the second similarities between the corresponding upper three local feature maps in C' and D'.
[0094] Perform a secondary summation on the sum of the second similarities and the first similarity to obtain the third similarity.
[0095] Determine the re-identification result of the target to be recognized based on the third similarity.
[0096] In addition, before putting the above target re-identification network into use, it is necessary to train the target re-identification network.
[0097] Figure 11 It is a schematic flowchart of the first embodiment of the training method of the target re-identification network of the present application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 11 the process sequence shown. As Figure 11 shown, this embodiment may include:
[0098] S41: Use the target re-identification model to extract features from the training static image and the training motion image of the moving target to obtain the training image features.
[0099] Wherein the training motion image is used to characterize the motion information of each pixel point of the training static image.
[0100] S42: Classify based on the training image features using the object re-identification model to obtain a classification result.
[0101] The classification result can be used to represent the category to which the training image features belong, that is, the identity information of the moving object.
[0102] S43: Adjust the parameters of the object re-identification model based on the classification result.
[0103] The classification result can be used to calculate the loss of the object re-identification model (such as cross-entropy loss), and the parameters of the object re-identification model are adjusted based on the cross-entropy loss and the stochastic gradient descent method.
[0104] The process of obtaining the training image features in the training process is the same as that in the application process. For details, please refer to the previous embodiments and will not be elaborated here. In addition, a classification network (fully connected layer) is provided in the object re-identification model during the training process for classifying the training image features.
[0105] Figure 12 It is a schematic structural diagram of an embodiment of the electronic device of the present application. As Figure 12 shown, the electronic device may include a processor 51 and a memory 52 coupled to the processor 51.
[0106] Among them, the memory 52 stores program instructions for implementing the method of any of the above embodiments; the processor 51 is configured to execute the program instructions stored in the memory 52 to implement the steps of the above method embodiments. Among them, the processor 51 may also be referred to as a CPU (Central Processing Unit, central processing unit). The processor 51 may be an integrated circuit chip with signal processing capabilities. The processor 51 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor 51 may also be any conventional processor, etc.
[0107] Figure 13 It is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. As Figure 13As shown, the computer-readable storage medium 60 of the embodiment of the present application stores program instructions 61, and when the program instructions 61 are executed, the method provided in the above embodiments of the present application is implemented. Among them, the program instructions 61 can form a program file and be stored in the above computer-readable storage medium 60 in the form of a software product, so that a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor can execute all or part of the steps of the method in various embodiments of the present application. The aforementioned computer-readable storage medium 60 includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, or terminal devices such as computers, servers, mobile phones, and tablets.
[0108] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0109] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units. The above is only the embodiment of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied to other related technical fields, is equally included in the patent protection scope of the present application.
Claims
1. A re-identification method for moving objects, characterized in that, it includes using a target re-identification model to perform the following processing: Extract features from the static images and motion images of the target to be identified and multiple candidate targets respectively to obtain image features, where the motion images are used to characterize the motion information of each pixel point of the static images; Calculate the similarity between the image features of the target to be identified and the image features of the multiple candidate targets; Based on the similarity, determine the re-identification result of the target to be identified from the multiple candidate targets; Among them, the step of extracting features from the static images and motion images of the target to be identified and multiple candidate targets respectively includes: Extract features from the splicing result of the static image and the motion image to obtain an overall feature map; Evaluate the pose of the static image to obtain the key point confidence of each pixel point; Normalize the motion information of the motion image to obtain the foreground point confidence of each pixel point; Select the pixel points whose key point confidence is greater than or equal to the first threshold and whose foreground point confidence is greater than or equal to the second threshold as the key points; Set the response degree of the key points to be equal to the key point confidence of the key points, and set the response degree of other pixel points except the key points to zero to obtain a response degree image; Multiply the response degree image with the overall feature image point by point to obtain key point features; Perform max pooling on the key point features to obtain a max pooling result; Perform average pooling on the overall feature map to obtain an average pooling result; Splice the max pooling result and the average pooling result to obtain the key point feature vector; The step of calculating the similarity between the image features of the target to be identified and the image features of the multiple candidate targets includes: Calculate the first similarity between the key point feature vectors of the target to be identified and the key point feature vectors of the multiple candidate targets; The step of determining the re-identification result of the target to be identified from the multiple candidate targets based on the similarity includes: Determine the re-identification result of the target to be identified at least based on the first similarity.
2. The method according to claim 1, characterized in that, the static image is an RGB image; The step of extracting features from the splicing result of the static image and the motion image includes: Extract horizontal optical flow and vertical optical flow from the static image respectively to obtain a first motion image and a second motion image; Input the splicing result of the static image, the first motion image and the second motion image into a pre-trained five-channel feature extraction network to obtain the overall feature map.
3. The method according to claim 1, characterized in that, The step of extracting features from the static images and motion images of the target to be identified and multiple candidate targets further includes: Divide the overall feature maps of the target to be identified and the candidate targets into multiple corresponding local feature maps along a preset direction, and extract local feature vectors; The step of calculating the similarity between the image features of the target to be recognized and the image features of the multiple candidate targets includes: Based on the foreground point confidence and key point confidence of each pixel point of the static images of the target to be recognized and the candidate targets, determine the occlusion situation of the local feature map; In the case where the corresponding local feature maps are not occluded, calculate the second similarity between the local feature vectors of the corresponding local feature maps and sum them; The step of determining the re-identification result of the target to be recognized from the multiple candidate targets based on the similarity includes: Determine the re-identification result of the target to be recognized based at least on the summation result of the second similarity.
4. The method according to claim 3, wherein, The step of determining the re-identification result of the target to be recognized from the multiple candidate targets based on the similarity includes: Perform a secondary summation of the summation result of the second similarity and the first similarity, and divide the secondary summation result by the number of the second similarities to obtain a third similarity; Determine the re-identification result of the target to be recognized based on the third similarity.
5. A training method for a target re-identification model, wherein, includes: Use the target re-identification model to extract features from the training static image and training motion image of the moving target to obtain training image features, wherein the training motion image is used to represent the motion information of each pixel point of the training static image; Use the target re-identification model to classify based on the training image features to obtain a classification result; Adjust the parameters of the target re-identification model based on the classification result; Among them, the step of using the target re-identification model to extract features from the training static image and training motion image of the moving target to obtain training image features includes: Extract features from the splicing result of the training static image and the training motion image to obtain a training overall feature map; Perform pose evaluation on the training static image to obtain the training key point confidence of each pixel point; Normalize the motion information of the training motion image to obtain the training foreground point confidence of each pixel point; Select the pixel points whose training key point confidence is greater than or equal to the first training threshold and whose training foreground point confidence is greater than or equal to the second training threshold as training key points; Set the response degree of the training key points to be equal to the training key point confidence of the training key points, and set the response degree of other pixel points except the training key points to zero to obtain a training response degree image; Perform dot multiplication on the training response degree image and the training overall feature image to obtain training key point features; Perform max pooling on the training key point features to obtain a training max pooling result; Perform average pooling on the training overall feature map to obtain a training average pooling result; Splice the training max pooling result and the training average pooling result to obtain a training key point feature vector to obtain the training image features; The step of adjusting the parameters of the target re-identification model based on the classification result includes: Calculating the cross-entropy loss of the target re-identification model by using the classification result, and adjusting the parameters of the target re-identification model based on the cross-entropy loss and the stochastic gradient descent method.
6. An electronic device, characterized in that, it includes a processor and a memory connected to the processor, wherein, the memory stores program instructions; the processor is configured to execute the program instructions stored in the memory to implement the method according to any one of claims 1-5.
7. A computer-readable storage medium, characterized in that, the storage medium stores program instructions, and when the program instructions are executed, the method according to any one of claims 1-5 is implemented.
Citation Information
Patent Citations
Face recognition method, feature extraction model training method and device thereof
CN109492624A
Moving posture recognition method, moving posture recognition device, terminal equipment and medium
CN110942006A
Target re-identification method, network training method thereof and related device
CN111814857A
Method and device for extracting pedestrian features, electronic equipment and storage medium
CN112488071A