Pedestrian Matching Method, Device, Electronic Device and Computer Readable Storage Medium
By extracting and aligning the image features, local and higher-order features of key points in human body are obtained, and the problem of abnormal points in pedestrian matching is solved, and more accurate pedestrian recognition is achieved.
Patent Information
- Application Number
- CN202111453596.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-01
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-12-01
AI Technical Summary
In the prior art, the image matching method is not accurate enough due to abnormal points or outliers when pedestrian recognition.
By extracting the image features, local features and higher-order features of key points in the human body, different feature alignment strategies are used to align local features and higher-order features, and fuse feature information to reduce the impact of abnormal points or outliers.
It improves the accuracy of pedestrian matching, reduces the impact of abnormal points or outliers on the matching results, and obtains more accurate matching results.
Smart Images

Figure CN114140825B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a pedestrian matching method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] Image matching refers to a method of seeking similar image targets by analyzing the correspondence of image content, features, structures, relationships, textures, grayscale, etc. for similarity and consistency. The image matching method can be used for pedestrian matching; pedestrian matching can determine whether the pedestrians appearing in different images are the same pedestrian by comparison.
[0003] Traditional pedestrian matching methods mostly directly use the global features of the human body for matching or match according to the local features of the human body. However, in actual situations, most images usually have uneven quality. For example, the human body in the image is occluded, the human body postures are different, and the image brightness is different, that is, there are some abnormal points or outliers. When performing pedestrian matching, it is easy to have a situation where feature alignment does not match, resulting in inaccurate pedestrian matching results. Summary of the Invention
[0004] In view of this, embodiments of the present disclosure provide a pedestrian matching method, apparatus, electronic device, and computer-readable storage medium to solve the problem that the pedestrian matching results are inaccurate due to some abnormal points or outliers in the existing images.
[0005] In a first aspect of the embodiments of the present disclosure, a pedestrian matching method is provided, including:
[0006] Obtain a first image and a second image;
[0007] Extract features from the first image and the second image respectively to obtain a first local feature, a second local feature, a first high-order feature, and a second high-order feature; the first local feature and the second local feature respectively include local feature vectors corresponding to each key point of the human body;
[0008] Align the first local feature and the second local feature to obtain a first fusion feature and a second fusion feature;
[0009] Align the first high-order feature and the second high-order feature to obtain a first high-order fusion feature and a second high-order fusion feature;
[0010] Based on the first fusion feature, the second fusion feature, the first high-order fusion feature, and the second high-order fusion feature, determine whether the first image and the second image contain the same pedestrian.
[0011] In a second aspect of the embodiments of the present disclosure, a pedestrian matching apparatus is provided, including:
[0012] An acquisition module, configured to acquire a first image and a second image;
[0013] A feature extraction module, configured to respectively perform feature extraction on the first image and the second image to obtain a first local feature, a second local feature, a first high-order feature, and a second high-order feature; the first local feature and the second local feature respectively include local feature vectors corresponding to each key point of the human body;
[0014] A first feature alignment module, configured to perform feature alignment on the first local feature and the second local feature to obtain a first fusion feature and a second fusion feature;
[0015] A second feature alignment module, configured to perform feature alignment on the first high-order feature and the second high-order feature to obtain a first high-order fusion feature and a second high-order fusion feature;
[0016] A matching module, configured to determine whether the first image and the second image contain the same pedestrian based on the first fusion feature, the second fusion feature, the first high-order fusion feature, and the second high-order fusion feature.
[0017] In a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0018] In a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0019] The beneficial effects of the embodiments of the present disclosure compared with the prior art are as follows: By respectively performing feature extraction on the acquired first image and second image, local features and high-order features corresponding to the key points of the human body in the two images are obtained, which are respectively denoted as the first local feature, the second local feature, the first high-order feature, and the second high-order feature. Feature alignment is performed on the first local feature and the second local feature to obtain a first fusion feature and a second fusion feature, and feature alignment is performed on the first high-order feature and the second high-order feature to obtain a first high-order fusion feature and a second high-order fusion feature. Based on the first fusion feature, the second fusion feature, the first high-order fusion feature, and the second high-order fusion feature, it is determined whether the two images contain the same pedestrian. By extracting the local features and high-order features of the human body in the image, different feature alignment strategies are used to respectively perform feature alignment on the local features and high-order features, incorporating the relationships between the features, comprehensively using the information between the features, obtaining more effective features, reducing the influence of abnormal points or outliers in the image on pedestrian matching, and obtaining a more accurate matching result. Description of the Drawings
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0021] Figure 1 is a schematic flowchart of a pedestrian matching method provided by an embodiment of the present disclosure;
[0022] Figure 2 is a schematic flowchart of determining whether the same pedestrian is included in the first image and the second image based on the first fusion feature, the second fusion feature, the first high-order fusion feature, and the second high-order fusion feature provided by an embodiment of the present disclosure;
[0023] Figure 3 is a schematic flowchart of feature aggregation of local features to obtain high-order features provided by an embodiment of the present disclosure;
[0024] Figure 4 is a schematic diagram of the human body key point numbers provided by an embodiment of the present disclosure;
[0025] Figure 5 is a schematic flowchart of feature alignment of the first local feature and the second local feature to obtain the first fusion feature and the second fusion feature provided by an embodiment of the present disclosure;
[0026] Figure 6 is a schematic flowchart of feature alignment of the first local feature and the second local feature respectively provided by an embodiment of the present disclosure;
[0027] Figure 7 is a schematic flowchart of feature alignment of the first high-order feature and the second high-order feature to obtain the first high-order fusion feature and the second high-order fusion feature provided by an embodiment of the present disclosure;
[0028] Figure 8 is a schematic flowchart of feature alignment of the first high-order feature and the second high-order feature respectively provided by an embodiment of the present disclosure;
[0029] Figure 9 is a schematic flowchart of another pedestrian matching method provided by an embodiment of the present disclosure;
[0030] Figure 10 is a schematic structural diagram of a pedestrian matching device provided by an embodiment of the present disclosure;
[0031] Figure 11It is a schematic structural diagram of another pedestrian matching device provided by an embodiment of the present disclosure;
[0032] Figure 12 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0033] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.
[0034] Next, a pedestrian matching method and device according to an embodiment of the present disclosure will be described in detail with reference to the accompanying drawings.
[0035] Figure 1 It is a schematic flowchart of a pedestrian matching method provided by an embodiment of the present disclosure. As Figure 1 shown, the pedestrian matching method includes steps S101 to S105:
[0036] S101, obtain a first image and a second image.
[0037] Among them, the first image and the second image are the images used for pedestrian matching. In this embodiment, by comparing the first image and the second image, it is determined whether the same pedestrian is included therein.
[0038] Specifically, the first image and the second image can be obtained from a database, or can be extracted from the images collected by a monitoring device in real time. In one embodiment, when the pedestrian matching method is applied to a scenario where a person stays, both the first image and the second image can be extracted from the images collected by the monitoring device. Two images are extracted from the images collected by the monitoring device every preset time for comparison to confirm whether the same pedestrian is included, and further determine whether a person stays at the same location for a long time. Further, in one embodiment, when a person is detected in the images collected by the monitoring device, the pedestrian matching method can be triggered, and two images are extracted at an interval of a preset time. Through pedestrian matching, it is determined whether the person stays at the same location for a long time. In other embodiments, the pedestrian matching method can also be applied to other scenarios.
[0039] It should be noted that in the embodiments disclosed in the present application, the concepts such as "first" and "second" mentioned are only used to distinguish different names and do not represent any actual meaning.
[0040] S102. Extract features from the first image and the second image respectively to obtain the first local features, the second local features, the first high-order features, and the second high-order features.
[0041] To enable a computer to "understand" an image, useful data or information is usually extracted from the image to obtain a "non-image" representation or description of the image, such as numerical values, vectors, and symbols. This process is feature extraction, and these "non-image" representations or descriptions extracted are features. For an image, each image has its own features that can distinguish it from other types of images. Some are natural features that can be intuitively felt, such as brightness, edges, textures, and colors; some are those that need to be obtained through transformation or processing, such as moments, histograms, and principal components.
[0042] In some embodiments, feature extraction of an image can be implemented using a neural network. In this embodiment, the purpose is to perform pedestrian matching. Therefore, when extracting features, the features of the human body contained in the image need to be extracted. Thus, a human detection model and a human key point estimation model can be used for feature extraction.
[0043] The local features include local feature vectors corresponding to each key point of the human body. In this embodiment, for the sake of distinction, the features extracted from the first image are denoted as the first local features and the first high-order features, and the features extracted from the second image are denoted as the second local features and the second high-order features.
[0044] In some embodiments, extracting the first local features, the second local features, the first high-order features, and the second high-order features from the first image and the second image respectively includes: performing human body local feature extraction on the first image and the second image respectively to obtain the first local features and the second local features; performing feature aggregation on the first local features to obtain the first high-order features, and performing feature aggregation on the second local features to obtain the second high-order features.
[0045] In a specific embodiment, extracting the first local features from the first image includes: inputting the first image into a human detection model to perform human detection on the first image to obtain a human detection box in the first image; inputting the human detection box into a human key point estimation model to obtain the features of each human key point corresponding to the human detection box, that is, the first local features; it can be understood that the process of extracting the second local features from the second image is of the same type as the process of extracting local features from the first image, which will not be elaborated here. Among them, any human detection model and human key point estimation model can be used to extract features from the image.
[0046] Further, in a specific embodiment, after obtaining a human detection box by performing feature extraction on an image through a human detection model, the human detection box is input into a human key point estimation model. The specific process of the human key point estimation model for processing the input human detection box includes: First, each feature map is passed through a 1*1 convolution to unify the number of channels to 1024, and then through ROIAlign (calibration) to uniformly pool the length and width of the feature map to [16, 8]. At this moment, the dimension of each human feature map G1 is [1024, 16, 8]. G1 is passed through three convolution blocks (conv(3,3), BN, relu) to sequentially obtain feature maps [1024, 16, 8], [512, 16, 8], [256, 16, 8]. Finally, through conv(1,1), BN, and relu, a heat map of all human key points is obtained. When there are 14 human key points, the heat map can be expressed as [14, 16, 8]. Among them, BN (Batch Normalization) normalizes the input of each layer. relu (Rectified Linear Unit), also known as the rectified linear unit, is a commonly used activation function in artificial neural networks, usually referring to non-linear functions represented by the ramp function and its variants.
[0047] The dimension of the heat map of each key point is [1, 16, 8], and then it is stacked by repeat to make its dimension become [1024, 16, 8], which is multiplied by the human feature map G1 [1024, 16, 8] to obtain the feature map [2048, 16, 8] of this key point. Finally, through global average pooling, the feature vector
[1024] of the key point is obtained. Similarly, the feature vectors of each human key point can be obtained through the above method. Each human key point in the first image is the first local feature, and each human key point in the second image is the second local feature.
[0048] In some embodiments, after obtaining the local features, high-order features can be obtained by respectively performing feature aggregation on the local features. Similarly, in this embodiment, the high-order feature obtained by aggregating the first local features is denoted as the first high-order feature, and the high-order feature obtained by performing feature aggregation on the second local features is denoted as the second high-order feature. Specifically, the specific process of performing feature aggregation on local features to obtain high-order features will be described in detail in subsequent embodiments and will not be elaborated here.
[0049] S103, perform feature alignment on the first local feature and the second local feature to obtain a first fusion feature and a second fusion feature.
[0050] Pedestrians may have different postures in different images at different times. When matching pedestrians in different images, the image features can be aligned to find corresponding points in the two images, so as to determine whether the same pedestrian is included. In this embodiment, for the local human body features extracted from the two images respectively, feature alignment is performed, and the aligned features are denoted as the first fused feature and the second fused feature. Specifically, the specific steps for performing feature alignment on the first local feature and the second local feature to obtain the first fused feature and the second fused feature will be described in detail in the subsequent embodiments.
[0051] S104. Align the first high-order feature and the second high-order feature to obtain the first high-order fused feature and the second high-order fused feature.
[0052] In this embodiment, when performing feature extraction, low-order local human body features are extracted, and high-order features of the image are obtained through the low-order local human body features. Therefore, when performing feature alignment, low-order feature alignment is performed on the low-order local features, and a high-order feature alignment strategy is adopted for the high-order features, making full use of the acquired image features to make the extracted features better. Further, in this embodiment, the features obtained after aligning the high-order features are denoted as the first high-order fused feature and the second high-order fused feature.
[0053] S105. Based on the first fused feature, the second fused feature, the first high-order fused feature, and the second high-order fused feature, determine whether the first image and the second image contain the same pedestrian.
[0054] In one embodiment, as Figure 2 shown, determining whether the first image and the second image contain the same pedestrian based on the first fused feature, the second fused feature, the first high-order fused feature, and the second high-order fused feature includes steps S201 to S203:
[0055] S201. Concatenate the first fused feature and the first high-order fused feature to obtain the first concatenated feature.
[0056] S202. Concatenate the second fused feature and the second high-order fused feature to obtain the second concatenated feature.
[0057] In some embodiments, feature concatenation can be completed by the concat function, and the concat function is used to concatenate two or more arrays. In this embodiment, the feature obtained by concatenating the first fused feature and the first high-order fused feature is denoted as the first concatenated feature, and the feature obtained by concatenating the second fused feature and the second high-order fused feature is denoted as the second concatenated feature. In other embodiments, feature concatenation can also be implemented in other ways.
[0058] S203. Calculate the contrast loss value between the first stitching feature and the second stitching feature, and determine whether the first image and the second image contain the same pedestrian according to the contrast loss value.
[0059] In some embodiments, the contrast loss value between the first stitching feature and the second stitching feature is calculated by a contrast loss function; wherein, the contrast loss function can well express the matching degree of paired samples. In this embodiment, using the contrast loss function to calculate the contrast loss value for the first stitching feature and the second stitching feature can be used to describe the matching degree between the first image and the second image. Specifically, any contrast loss function can be adopted in this embodiment to calculate the contrast loss value between the two stitching features.
[0060] After obtaining the contrast loss value, it can be determined whether the first image and the second image contain the same pedestrian according to the contrast loss value. In one embodiment, to determine whether the first image and the second image contain the same pedestrian according to the contrast loss value, a preset condition can be set. When the contrast loss value meets the preset condition, it can be determined that the first image and the second image contain the same pedestrian; wherein, the preset condition can be set according to the actual situation. The higher the matching degree between the first image and the second image determined according to the contrast loss value, the more likely it is that the first image and the second image contain the same pedestrian.
[0061] In other embodiments, after obtaining the first stitching feature and the second stitching feature, it can also be determined whether the first image and the second image contain the same pedestrian in other ways based on the first stitching feature and the second stitching feature; for example, by calculating the similarity between the first stitching feature and the second stitching feature, and determining whether the first image and the second image contain the same pedestrian according to the similarity.
[0062] According to the technical solution provided by the embodiments of the present disclosure, by respectively performing feature extraction on the obtained first image and second image, the local features and high-order features corresponding to the human key points in the two images are obtained, which are respectively denoted as the first local feature, the second local feature, the first high-order feature, and the second high-order feature. Feature alignment is performed on the first local feature and the second local feature to obtain the first fusion feature and the second fusion feature. Feature alignment is performed on the first high-order feature and the second high-order feature to obtain the first high-order fusion feature and the second high-order fusion feature, and it is determined whether the two images contain the same pedestrian based on the first fusion feature, the second fusion feature, the first high-order fusion feature, and the second high-order fusion feature. By extracting the human local features and high-order features in the image, different feature alignment strategies are adopted to perform feature alignment on the local features and high-order features respectively, incorporating the relationships between the features, comprehensively using the information between the features, obtaining more effective features, reducing the influence of abnormal points or outliers in the image on pedestrian matching, and obtaining a more accurate matching result.
[0063] In some embodiments, there are more than two human key points; in this embodiment, as Figure 3 shown, feature aggregation of local features is performed to obtain high-order features, including steps S301 to S306, where the local features include first local features and second local features, and the high-order features include first high-order features and second high-order features.
[0064] S301. According to the local feature vectors corresponding to each human key point, learn to determine the first parameter, second parameter, and third parameter corresponding to the local features.
[0065] There are usually more than two human key points. When extracting human local features, local feature vectors corresponding to each human key point can be obtained, and the number of local feature vectors included in the local features corresponds to the number of human key points. The local features in this embodiment include first local features and second local features extracted from the first image and the second image respectively.
[0066] Specifically, learning the first parameter, second parameter, and third parameter corresponding to the local features according to the local feature vectors corresponding to each human key point can be implemented in any way. For example, it can be implemented through a neural network; in a specific embodiment, it can be implemented through a Transformer network (proposed by Ashish Vaswani). In some embodiments, the ranges of the first parameter, second parameter, and third parameter corresponding to the local vectors are all R 1024×256 .
[0067] S302. Generate queries, keys, and values corresponding to each local feature vector according to the first parameter, second parameter, and third parameter.
[0068] Among them, queries, keys, and values are concepts in the Transformer network. Generating queries query, keys key, and values value corresponding to each local feature vector according to the first parameter, second parameter, and third parameter can be expressed as:
[0069] query: Q i = f i * W q , where W q represents the first parameter;
[0070] key: K i = f i * W k , where W k represents the second parameter;
[0071] value: V i = f i * Wv , where W v represents the third parameter.
[0072] S303. Based on the query and the key, determine the weight coefficient between pairwise local feature vectors.
[0073] In some embodiments, based on the query and the key, determining the weight coefficient between pairwise local feature vectors includes: selecting two local feature vectors as the first target local feature vector and the second target local feature vector, and calculating the weight coefficient between the first target local feature vector and the second target local feature vector using the query of the first target local feature vector and the key of the second target local feature vector.
[0074] Furthermore, in a specific embodiment, calculating the weight coefficient between the first target local feature vector and the second target local feature vector using the query of the first target local feature vector and the key of the second target local feature vector can be expressed as:
[0075]
[0076] where a ij represents the weight coefficient between the local feature vectors numbered i and j; i and j respectively represent the numbers corresponding to the human key points, where the human key points can be numbered in advance; Q i represents the query corresponding to the first target local feature vector, K j represents the key of the second target local feature vector, represents the transpose of K j , 256 is the dimension of Q i , K j , softmax() represents the logistic regression algorithm, and a ij represents the weight coefficient between the first target local feature vector and the second target local feature vector. Among them, the definition of transpose is: all elements of A are mirrored along a ray at a 45-degree angle to the lower right starting from the element in the first row and first column, that is, the transpose of A is obtained.
[0077] Through the above method, the weight coefficient between the local feature vectors corresponding to each pair of human key points can be calculated.
[0078] S304. For any local feature vector, determine the adjustment coefficient according to the relationship between the local feature vector and other local feature vectors.
[0079] Among them, the relationship between the local feature vector and other local feature vectors can be determined by the numbers set for the human key points in advance. In a specific embodiment, each human key point is assigned a number, such as Figure 4Shown as follows: human head = 1; neck = 2; left shoulder = right shoulder = 3; left elbow = right elbow = 4; left hand = right hand = 5; left hip = right hip = 6; left knee = right knee = 7; left foot = right foot = 8. Since when performing pedestrian matching, the closer the positions of human key points are, the greater the influence between them, in this embodiment, an adjustment coefficient is set to aggregate high-order features according to the influence degree between key points.
[0080] In a specific embodiment, according to the relationship between the local feature vector and other local feature vectors, the adjustment coefficient can be expressed as:
[0081]
[0082] Where p represents the adjustment coefficient, r is a hyperparameter, and the value of the hyperparameter can be defined according to the actual situation. i and j respectively represent the numbers of key points.
[0083] S305. Obtain the aggregated feature corresponding to the local feature according to the weight coefficient, adjustment coefficient and value.
[0084] In some embodiments, obtaining the aggregated feature corresponding to the local feature according to the weight coefficient, adjustment coefficient and value can be expressed as:
[0085]
[0086] Where f i at represents the aggregated feature corresponding to the local feature vector numbered i, a ij represents the weight coefficient between the local feature vectors numbered i and numbered j, p represents the adjustment coefficient, V j represents the value corresponding to the local feature vector numbered j. i and j respectively represent the numbers of key points, and 14 represents the number of local feature vectors (human key points).
[0087] S306. Determine the high-order feature from the aggregated features corresponding to each local feature.
[0088] In step S305, the aggregated feature corresponding to each human key point (local feature vector) can be obtained, and by combining each aggregated feature together, the corresponding high-order feature is obtained. In this embodiment, the first high-order feature is aggregated according to the first local feature, and the second high-order feature is aggregated according to the second local feature.
[0089] Furthermore, the process of aggregating high-order features from local features described above can be implemented by a transformer model.
[0090] According to the technical solution provided by the embodiments of the present disclosure, the transformer model is used to separately learn the corresponding first parameter, second parameter, and third parameter for the input first local feature and second local feature, calculate the query, key, and value of each local feature vector according to the first parameter, second parameter, and third parameter, then calculate the weight coefficient between two local feature vectors according to the query and key corresponding to each local feature vector, and determine the adjustment coefficient according to the relationship between two local feature vectors. Finally, according to the weight coefficient, the value corresponding to the local feature vector, and the adjustment coefficient, a new local feature is aggregated, and finally a fusion feature aggregated from the local features is obtained. In this embodiment, based on the relationship between local feature vectors, that is, the prior topological structure relationship of human key points, an adjustment coefficient is designed, and the influence degree between local feature vectors is added to feature aggregation through the adjustment coefficient, so that the aggregated high-order features are better.
[0091] In some embodiments, as Figure 5 shown, feature alignment is performed on the first local feature and the second local feature to obtain a first fusion feature and a second fusion feature, including steps S501 to S504:
[0092] S501, calculate the local feature similarity between the first local feature and the second local feature.
[0093] Calculating the similarity between local features can be implemented in any way. In one embodiment, the first local feature and the second local feature are mapped to the same feature space, and the cosine distance between the two local features is calculated as the feature similarity between the two local features. Further, mapping the first local feature and the second local feature to the same feature space can be achieved by normalizing the distributions of the first local feature and the second local feature. In a specific embodiment, mapping the first local feature and the second local feature to the same feature space and calculating the cosine distance between the two local features as the feature similarity between the two local features can be expressed as:
[0094] M1 = normalize(F p ) * [normalize(F g )] T
[0095] where Fp represents the first local feature, Fg represents the second local feature, normalize() represents normalizing the local feature, and M1 represents the similarity between the first local feature Fp and the second local feature Fg.
[0096] In other embodiments, calculating the local feature similarity between the first local feature and the second local feature can also be implemented in other ways.
[0097] S502. Determine a first matching feature of the second local feature in the first local feature according to the second local feature and the local feature similarity.
[0098] In some embodiments, determining a first matching feature of the second local feature in the first local feature according to the second local feature and the local feature similarity includes: calculating a first product of the local feature similarity and the second local feature; determining the sum value of the first product and the first local feature as the first matching feature.
[0099] Among them, calculating the first product of the local feature similarity and the second local feature means that the points in the first local feature that match the second local feature are determined. According to the matching points, it is possible to determine which points in the first local feature are to be emphasized, thereby excluding outliers or occluded points.
[0100] Calculating the sum value between the first product and the first local feature includes both the original features in the first local feature and the relationship with the second local feature. In this embodiment, this sum value is denoted as the first matching feature. According to the first matching feature, it is possible to determine which points in the first image need to be matched with the second image.
[0101] In a specific embodiment, calculating the first product of the local feature similarity and the second local feature can be expressed as:
[0102] Among them, represents the first product, M1 represents the local feature similarity, and Fg represents the second local feature.
[0103] Determining the sum value of the first product and the first local feature as the first matching feature is expressed as:
[0104] Among them, represents the first matching feature, and Fp represents the first local feature.
[0105] S503. Determine a second matching feature of the first local feature in the second local feature according to the first local feature and the local feature similarity.
[0106] Similar to determining the first matching feature, a second matching feature can be determined according to the first local feature and the local feature similarity. The difference is that in this embodiment, the local feature similarity is first transposed, and the transposed local feature similarity is used to determine the second matching feature.
[0107] In some embodiments, determining the second matching feature of the first local feature in the second local feature according to the first local feature and the local feature similarity includes: transposing the local feature similarity to obtain the transposed local feature similarity; calculating the second product of the transposed local feature similarity and the first local feature; and determining the sum value of the second product and the second local feature as the second matching feature.
[0108] Calculating the second product of the transposed local feature similarity and the first local feature represents the points in the second local feature that match the first local feature. Based on the matching points, it is possible to determine which points in the second local feature are to be emphasized, thereby excluding outliers or occluded points.
[0109] Calculating the sum value between the second product and the second local feature contains both the original features in the first local feature and the relationship with the second local feature. In this embodiment, this sum value is denoted as the second matching feature. Based on the second matching feature, it is possible to determine which points in the second image need to be matched with the first image and which points do not match.
[0110] In a specific embodiment, calculating the second product of the transposed local feature similarity and the first local feature can be expressed as:
[0111] Where, represents the second product, represents the transposed local feature similarity, and Fp represents the first local feature.
[0112] Determining the sum value of the second product and the second local feature as the second matching feature is expressed as:
[0113] Where, represents the second matching feature, and Fg represents the second local feature.
[0114] S504, performing feature extraction on the first matching feature to obtain the first fusion feature, and performing feature extraction on the second matching feature to obtain the second fusion feature.
[0115] After obtaining the first matching feature and the second matching feature, feature extraction is respectively performed to obtain the new features fused after low-order feature alignment, that is, the first fusion feature and the second fusion feature.
[0116] In some embodiments, performing feature extraction on the input matching features respectively can be achieved through fc + relu. Among them, the fc layer is a fully connected layer, and each node of the fully connected layer is connected to all nodes of the previous layer, used to synthesize the features extracted previously. The Relu layer usually consists of activation functions. Among them, the matching features include the first matching feature and the second matching feature.
[0117] In a specific embodiment, as Figure 6 shown is a flowchart of steps for respectively performing feature alignment on a first local feature and a second local feature, where Fp represents the first local feature, and Fg represents the second local feature; represents a first fusion feature, represents a second fusion feature.
[0118] According to the technical solution provided by the embodiments of the present disclosure, by calculating the local feature similarity between local features, and calculating a first matching feature and a second matching feature according to the local feature similarity, and finally respectively performing feature extraction on the first matching feature and the second matching feature, new features after low-order feature alignment are obtained, that is, the first fusion feature and the second fusion feature. Subsequently, pedestrian matching is performed based on the aligned local features, and a more accurate matching result can be obtained.
[0119] In some other embodiments, as Figure 7 shown, feature alignment is performed on a first high-order feature and a second high-order feature to obtain a first high-order fusion feature and a second high-order fusion feature, including steps S701 to S704:
[0120] S701, calculate the high-order feature similarity between the first high-order feature and the second high-order feature.
[0121] Calculating the similarity between high-order features can be implemented in any way. In one embodiment, the similarity between high-order features can be determined through network learning, and can be expressed as:
[0122]
[0123] where M2 represents the high-order feature similarity, represents the first high-order feature, represents the second high-order feature, represents the transpose of and A represents a network parameter matrix. Among them, the dimension of the network parameter matrix A is [1024, 1024].
[0124] In other embodiments, calculating the similarity between the first high-order feature and the second high-order feature can also be implemented in any other way.
[0125] S702, determine a first high-order matching feature of the second high-order feature in the first high-order feature according to the second high-order feature and the high-order feature similarity.
[0126] In some embodiments, determining the first high-order matching feature of the second high-order feature in the first high-order feature according to the second high-order feature and the high-order feature similarity includes: calculating the third product of the high-order feature similarity and the second high-order feature; determining the sum value of the third product and the first high-order feature as the first high-order matching feature.
[0127] In a specific embodiment, calculating the third product of the high-order feature similarity and the second high-order feature can be expressed as:
[0128] Wherein, represents the third product, M2 represents the high-order feature similarity, represents the second high-order feature.
[0129] Determining the sum value of the third product and the first high-order feature as the first high-order matching feature is expressed as:
[0130] Wherein, represents the first high-order matching feature, represents the first high-order feature.
[0131] S703. Determine the second high-order matching feature of the first high-order feature in the second high-order feature according to the first high-order feature and the high-order feature similarity.
[0132] Similar to determining the first high-order matching feature, the second high-order matching feature can be determined according to the first high-order feature and the high-order feature similarity. The difference is that in this embodiment, the high-order feature similarity is first transposed, and the transposed high-order feature similarity is used to determine the second high-order matching feature.
[0133] In some embodiments, determining the second matching feature of the first high-order feature in the second high-order feature according to the first high-order feature and the high-order feature similarity includes: transposing the high-order feature similarity to obtain the transposed high-order feature similarity; calculating the fourth product of the transposed high-order feature similarity and the first high-order feature; determining the sum value of the fourth product and the second high-order feature as the second high-order matching feature.
[0134] In a specific embodiment, calculating the fourth product of the transposed high-order feature similarity and the first high-order feature can be expressed as:
[0135] Wherein, represents the fourth product, represents the transposed high-order feature similarity, represents the first high-order feature.
[0136] Determining the sum value of the fourth product and the second high-order feature as the second high-order matching feature is expressed as:
[0137] Among them, represents the second-highest-order matching feature, represents the second-highest-order feature.
[0138] S704. Feature extraction is performed on the first-highest-order matching feature to obtain the first-highest-order fusion feature, and feature extraction is performed on the second-highest-order matching feature to obtain the second-highest-order fusion feature.
[0139] After obtaining the first-highest-order matching feature and the second-highest-order matching feature, feature extraction is respectively performed to obtain the newly fused features after high-order feature alignment, that is, the first-highest-order fusion feature and the second-highest-order fusion feature.
[0140] In some embodiments, feature extraction is respectively performed on the input high-order matching features, which can be implemented by fc + relu. Among them, the fc layer is a fully connected layer, and each node of the fully connected layer is connected to all nodes of the previous layer, used to synthesize the previously extracted features. The Relu layer is usually composed of activation functions. Among them, the high-order matching features include the first-highest-order matching feature and the second-highest-order matching feature.
[0141] In a specific embodiment, as Figure 8 shown, it is a flowchart of the steps for respectively performing feature alignment on the first-highest-order feature and the second-highest-order feature. Among them, represents the first-highest-order feature, represents the second-highest-order feature; represents the first-highest-order fusion feature, represents the second-highest-order fusion feature.
[0142] According to the technical solution provided by the embodiments of the present disclosure, by calculating the high-order feature similarity between high-order features, and calculating the first-highest-order matching feature and the second-highest-order matching feature according to the high-order feature similarity, and finally respectively performing feature extraction on the first-highest-order matching feature and the second-highest-order matching feature to obtain the newly aligned high-order features, that is, the first-highest-order fusion feature and the second-highest-order fusion feature. Subsequently, pedestrian matching is performed based on the aligned high-order features, and a more accurate matching result can be obtained.
[0143] In a specific embodiment, as Figure 9 shown is a specific flow schematic diagram of the pedestrian matching method in a specific embodiment
[0144] (1) Input two pictures, probe and gallery:
[0145] Respectively pass through the human detection model to obtain the human detection box bbox, and send it to the human key point estimation module.
[0146] (2) Input the human detection bounding boxes obtained in step (1):
[0147] First, pass the feature map corresponding to each human detection bounding box through a 1×1 convolution to unify the number of channels to 1024, and then use ROIAlign to uniformly pool the length and width of the feature map to [16, 8]. At this moment, the dimension of each human feature map G1 is [1024, 16, 8].
[0148] Pass G1 through three convolutional blocks (conv(3,3), BN, relu) to obtain feature maps [1024, 16, 8], [512, 16, 8], and [256, 16, 8] in sequence. Finally, pass through conv(1,1), BN, relu to obtain the heat maps of all key points [14, 16, 8]. (Among them, 14 human key points are set, including head, neck, left hand, left elbow, left shoulder, left hip, left knee, left foot, right hand, right elbow, right shoulder, right hip, right knee, and right foot).
[0149] The dimension of the heat map of each key point is [1, 16, 8]. Repeat the heat maps of each key point to change its dimension to [1024, 16, 8], and multiply it with the human feature map G1 [1024, 16, 8] to obtain the feature map of this key point [2048, 16, 8]. Finally, through global average pooling, obtain the feature vector of the key point
[1024] . Similarly, the feature vectors of each key point can be obtained. That is, the first local feature F is obtained p and the second local feature F g , F p and F g both have dimensions of [14, 1024].
[0150] (3) Send F p and F g into the transformer encoding module to obtain the first high-order feature and the second high-order feature Among them, for the specific method of aggregating to obtain the high-order feature, please refer to the description of the embodiment shown in Figure 3 the embodiment shown
[0151] (4) At the same time, send F p and F g into the low-order feature alignment strategy module to obtain and both of which have dimensions of [14, 1024]. Among them, for the specific method of low-order feature alignment, please refer to the description of the embodiment shown in Figure 5 the embodiment shown
[0152] (5) According to step (3), obtain the high-order feature and Both of them have dimensions of [14, 1024]. They are sent to the high-order feature matching strategy module to obtain and Both of them have dimensions of [14, 1024]. Among them, for the specific method of high-order feature alignment, please refer to Figure 7 the description of the embodiments shown
[0153] (6) The and obtained in step (4) and the and obtained in step (5) are concatenated correspondingly to perform vector concatenation:
[0154]
[0155]
[0156] (7) The first concatenated feature and the second concatenated feature are sent to the contrast loss function to calculate the contrast loss value between and . Finally, it is determined whether the same pedestrian is included in the picture probe and gallery according to the contrast loss value
[0157] According to the pedestrian matching method in the embodiments of the present disclosure, a feature alignment module is designed. Through the low-order feature alignment strategy and the high-order feature alignment strategy, the relationship between features is incorporated, and the information between features is comprehensively utilized, so as to make the extracted features better. Among them, the low-order feature alignment strategy matches through the similarity of local features of human key points. Whether two points match not only depends on the similarity between the two points, but also considers the impact of the matching of these two points on other matches, so as to select a matching method that maximizes the overall energy. For this reason, the pedestrian matching method in the embodiments of the present disclosure also adds a high-order feature matching strategy (the high-order feature can be understood as the feature information of the edge). Specifically, the local features of human key points extracted in the early stage are aggregated into high-order feature information through the transformer encoding module, and then the network is allowed to learn how to match these high-order features, which has higher robustness. In addition, when aggregating high-order features through transformer encoding, according to the prior topological structure relationship of human key points, a regulation coefficient p is calculated to aggregate high-order semantic information according to the influence degree between key points
[0158] The present application also provides an application scenario that applies the above-mentioned pedestrian matching method. Specifically, the application of the pedestrian matching method in this application scenario is as follows:
[0159] In this embodiment, the pedestrian matching method is applied to the application scenario of detecting people's lingering in video surveillance. In this application scenario, it is necessary to detect and alarm suspicious phenomena of people lingering or wandering for a long time. It is possible to determine whether it is the same person by comparing the human features detected in consecutive multiple frames, so as to determine whether a person lingers at the same place for a long time.
[0160] In this application scenario, the pedestrian matching method can be triggered when a person appears in the monitored image. Two images are obtained from the monitoring device at every preset time, and pedestrian matching detection is performed on the two images. Further, when it is continuously detected that the two images taken at the preset time interval contain the same pedestrian, it means that the person appearing in the monitored image lingers here for a long time, and an alarm can be issued for this situation. Among them, the preset time can be set according to the actual situation. For example, it can be set to 1 minute, 3 minutes, 5 minutes, etc. It can also be set according to the actual situation to trigger an alarm when the number of times the same pedestrian is continuously detected in the monitored image reaches a preset number of times.
[0161] All the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present application, which will not be elaborated one by one here.
[0162] The following is an embodiment of the apparatus of the present disclosure, which can be used to execute the method embodiment of the present disclosure. For details not disclosed in the embodiment of the apparatus of the present disclosure, please refer to the method embodiment of the present disclosure.
[0163] Figure 10 It is a schematic diagram of a pedestrian matching device provided by an embodiment of the present disclosure. As Figure 10 shown, the pedestrian matching device includes:
[0164] An acquisition module 1001, configured to acquire a first image and a second image;
[0165] A feature extraction module 1002, configured to perform feature extraction on the first image and the second image respectively to obtain a first local feature, a second local feature, a first high-order feature, and a second high-order feature; the first local feature and the second local feature respectively include local feature vectors corresponding to each key point of the human body;
[0166] A first feature alignment module 1003, configured to perform feature alignment on the first local feature and the second local feature to obtain a first fusion feature and a second fusion feature;
[0167] A second feature alignment module 1004, configured to perform feature alignment on the first high-order feature and the second high-order feature to obtain a first high-order fusion feature and a second high-order fusion feature;
[0168] A matching module 1005, configured to determine whether the same pedestrian is included in the first image and the second image based on the first fusion feature, the second fusion feature, the first high-order fusion feature, and the second high-order fusion feature.
[0169] According to the technical solution provided by the embodiments of the present disclosure, by respectively performing feature extraction on the acquired first image and second image, local features and high-order features corresponding to human key points in the two images are obtained, which are respectively denoted as the first local feature, the second local feature, the first high-order feature, and the second high-order feature. Feature alignment is performed on the first local feature and the second local feature to obtain the first fusion feature and the second fusion feature, feature alignment is performed on the first high-order feature and the second high-order feature to obtain the first high-order fusion feature and the second high-order fusion feature, and it is determined whether the same pedestrian is included in the two images based on the first fusion feature, the second fusion feature, the first high-order fusion feature, and the second high-order fusion feature. By extracting the human local features and high-order features in the image, different feature alignment strategies are used to perform feature alignment on the local features and high-order features respectively, incorporating the relationships between the features, comprehensively using the information between the features, obtaining more effective features, reducing the influence of abnormal points or outliers in the image on pedestrian matching, and obtaining a more accurate matching result.
[0170] In some embodiments, the feature extraction module of the above pedestrian matching device includes:
[0171] A local feature extraction sub-module, configured to respectively perform human local feature extraction on the first image and the second image to obtain the first local feature and the second local feature;
[0172] An aggregation sub-module, configured to perform feature aggregation on the first local feature to obtain the first high-order feature, and perform feature aggregation on the second local feature to obtain the second high-order feature.
[0173] In some embodiments, there are more than two human key points; as Figure 11 shown, in this embodiment, the feature extraction module of the above device includes:
[0174] A parameter learning sub-module 1101, configured to learn and determine the first parameter, the second parameter, and the third parameter corresponding to the local feature according to the local feature vectors corresponding to each human key point;
[0175] A parameter determination sub-module 1102, configured to generate a query, a key, and a value corresponding to each local feature vector according to the first parameter, the second parameter, and the third parameter;
[0176] A weight coefficient determination sub-module 1103, configured to determine the weight coefficient between two local feature vectors based on the query and the key;
[0177] An adjustment coefficient determination sub-module 1104, configured to determine an adjustment coefficient for any local feature vector according to the relationship between the local feature vector and other local feature vectors;
[0178] An aggregation sub-module 1105, configured to obtain an aggregated feature corresponding to the local feature according to the weight coefficient, the adjustment coefficient, and the value;
[0179] A high-order feature determination sub-module 1106, configured to determine a high-order feature from the aggregated features corresponding to the local features; wherein, the local features include a first local feature and a second local feature, and the high-order features include a first high-order feature and a second high-order feature.
[0180] In some embodiments, please continue to refer to Figure 11 , the first feature alignment module of the above device includes:
[0181] A local feature calculation sub-module 1107, configured to calculate the local feature similarity between the first local feature and the second local feature;
[0182] A first matching feature determination sub-module 1108, configured to determine a first matching feature of the second local feature in the first local feature according to the second local feature and the local feature similarity;
[0183] A second matching feature determination sub-module 1109, configured to determine a second matching feature of the first local feature in the second local feature according to the first local feature and the local feature similarity;
[0184] A feature extraction sub-module 1110, configured to perform feature extraction on the first matching feature to obtain a first fusion feature, and perform feature extraction on the second matching feature to obtain a second fusion feature.
[0185] In some embodiments, the first matching feature determination sub-module of the above device includes: a multiplication operation unit, configured to calculate a first product of the local feature similarity and the second local feature; an addition operation unit, configured to determine the sum value of the first product and the first local feature as the first matching feature.
[0186] In some embodiments, the second matching feature determination sub-module of the above device includes: a transposition unit, configured to transpose the local feature similarity to obtain a transposed local feature similarity; a multiplication operation unit, configured to calculate a second product of the transposed local feature similarity and the first local feature; an addition operation unit, configured to determine the sum value of the second product and the second local feature as the second matching feature.
[0187] In some embodiments, please continue to refer to Figure 11 , the second feature alignment module of the above device includes:
[0188] The high-order feature similarity calculation sub-module 1111 is used to calculate the high-order feature similarity between the first high-order feature and the second high-order feature;
[0189] The first high-order matching determination sub-module 1112 is used to determine the first high-order matching feature of the second high-order feature in the first high-order feature according to the second high-order feature and the high-order feature similarity;
[0190] The second high-order matching determination sub-module 1113 is used to determine the second high-order matching feature of the first high-order feature in the second high-order feature according to the first high-order feature and the high-order feature similarity;
[0191] The feature extraction sub-module 1114 is used to extract features from the first high-order matching feature to obtain the first high-order fusion feature, and extract features from the second high-order matching feature to obtain the second high-order fusion feature.
[0192] In some embodiments, please continue to refer to Figure 11 The matching module of the above device includes:
[0193] The first splicing sub-module 1115 is used to splice the first fusion feature and the first high-order fusion feature to obtain the first splicing feature;
[0194] The second splicing sub-module 1116 is used to splice the second fusion feature and the second high-order fusion feature to obtain the second splicing feature;
[0195] The contrast loss calculation sub-module 1117 is used to calculate the contrast loss value between the first splicing feature and the second splicing feature, and determine whether the first image and the second image contain the same pedestrian according to the contrast loss value.
[0196] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure.
[0197] Figure 12 is a schematic diagram of the electronic device 12 provided by the embodiment of the present disclosure. As Figure 12 shown, the electronic device 12 of this embodiment includes: a processor 1201, a memory 1202, and a computer program 1203 stored in the memory 1202 and executable on the processor 1201. When the processor 1201 executes the computer program 1203, the steps in the above various method embodiments are implemented. Alternatively, when the processor 1201 executes the computer program 1203, the functions of each module / unit in the above device embodiments are implemented.
[0198] Exemplarily, the computer program 1203 can be divided into one or more modules / units, which are stored in the memory 1202 and executed by the processor 1201 to complete the present disclosure. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 1203 in the electronic device 12.
[0199] The electronic device 12 can be a desktop computer, a notebook, a palm computer, a cloud server, and other electronic devices. The electronic device 12 can include, but is not limited to, the processor 1201 and the memory 1202. Those skilled in the art can understand that Figure 12 merely examples of the electronic device 12, which do not constitute a limitation on the electronic device 12, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, a bus, etc.
[0200] The processor 1201 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0201] The memory 1202 can be an internal storage unit of the electronic device 12. For example, the hard disk or memory of the electronic device 12. The memory 1202 can also be an external storage device of the electronic device 12. For example, a plug-in hard disk equipped on the electronic device 12, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 1202 can also include both the internal storage unit and the external storage device of the electronic device 12. The memory 1202 is used to store computer programs and other programs and data required by the electronic device. The memory 1202 can also be used to temporarily store data that has been output or will be output.
[0202] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0203] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not described or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0204] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in the form of hardware or software depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this disclosure.
[0205] In the embodiments provided by this disclosure, it should be understood that the disclosed device / electronic device and method can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical functional division. In actual implementation, there may be other division methods. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0206] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0207] In addition, in each of the embodiments of the present disclosure, each functional unit may be integrated into a processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0208] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present disclosure, it may also be completed by instructing relevant hardware through a computer program. The computer program may be stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments may be implemented. The computer program may include computer program code, and the computer program code may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0209] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be included within the protection scope of the present disclosure.
Claims
1. A pedestrian matching method, characterized in that, Comprising: Obtaining a first image and a second image; Performing feature extraction on the first image and the second image respectively to obtain a first local feature and a second local feature; the first local feature and the second local feature respectively include local feature vectors corresponding to each human body key point; performing feature aggregation on the first local feature to obtain a first high-order feature, and performing feature aggregation on the second local feature to obtain a second high-order feature; Performing feature alignment on the first local feature and the second local feature to obtain a first fused feature and a second fused feature; Performing feature alignment on the first high-order feature and the second high-order feature to obtain a first high-order fused feature and a second high-order fused feature; Based on the first fused feature, the second fused feature, the first high-order fused feature and the second high-order fused feature, determining whether the first image and the second image contain the same pedestrian; The performing feature alignment on the first local feature and the second local feature to obtain a first fused feature and a second fused feature includes: Calculating the local feature similarity between the first local feature and the second local feature; Determining a first matching feature of the second local feature in the first local feature according to the second local feature and the local feature similarity; Determining a second matching feature of the first local feature in the second local feature according to the first local feature and the local feature similarity; Performing feature extraction on the first matching feature to obtain the first fused feature, and performing feature extraction on the second matching feature to obtain the second fused feature; The performing feature alignment on the first high-order feature and the second high-order feature to obtain a first high-order fused feature and a second high-order fused feature includes: Calculating the high-order feature similarity between the first high-order feature and the second high-order feature; Determining a first high-order matching feature of the second high-order feature in the first high-order feature according to the second high-order feature and the high-order feature similarity; Determining a second high-order matching feature of the first high-order feature in the second high-order feature according to the first high-order feature and the high-order feature similarity; Performing feature extraction on the first high-order matching feature to obtain the first high-order fused feature, and performing feature extraction on the second high-order matching feature to obtain the second high-order fused feature.
2. The method according to claim 1, wherein Performing feature extraction on the first image and the second image respectively to obtain a first local feature and a second local feature includes: Performing human body local feature extraction on the first image and the second image respectively to obtain the first local feature and the second local feature.
3. The method according to claim 2, wherein The human body key points include more than two; Performing feature aggregation on local features to obtain high-order features includes: Learning and determining a first parameter, a second parameter and a third parameter corresponding to the local features according to the local feature vectors corresponding to each of the human body key points; Generating queries, keys and values corresponding to each of the local feature vectors according to the first parameter, the second parameter and the third parameter; Determining the weight coefficients between pairwise local feature vectors based on the queries and the keys; For any of the local feature vectors, determine an adjustment coefficient according to the relationship between the local feature vector and each of the other local feature vectors; Obtain the aggregated feature corresponding to the local feature according to the weight coefficient, the adjustment coefficient, and the value; Determine the high-order feature from the aggregated features corresponding to the respective local features; Wherein, the local features include the first local feature and the second local feature, and the high-order features include the first high-order feature and the second high-order feature.
4. The method according to claim 1, characterized in that, Determine the first matching feature of the second local feature in the first local feature according to the second local feature and the local feature similarity, including: calculating a first product of the local feature similarity and the second local feature; determining the sum value of the first product and the first local feature as the first matching feature; Determine the second matching feature of the first local feature in the second local feature according to the first local feature and the local feature similarity, including: transposing the local feature similarity to obtain the transposed local feature similarity; calculating a second product of the transposed local feature similarity and the first local feature; determining the sum value of the second product and the second local feature as the second matching feature.
5. The method according to any one of claims 1 to 4, characterized in that Based on the first fusion feature, the second fusion feature, the first high-order fusion feature, and the second high-order fusion feature, determine whether the first image and the second image contain the same pedestrian, including: Concatenate the first fusion feature and the first high-order fusion feature to obtain a first concatenated feature; Concatenate the second fusion feature and the second high-order fusion feature to obtain a second concatenated feature; Calculate a contrast loss value between the first concatenated feature and the second concatenated feature, and determine whether the first image and the second image contain the same pedestrian according to the contrast loss value.
6. A pedestrian matching device, characterized in that, Including: An acquisition module for acquiring a first image and a second image; A feature extraction module for respectively performing feature extraction on the first image and the second image to obtain a first local feature and a second local feature; the first local feature and the second local feature respectively include local feature vectors corresponding to each human body key point; perform feature aggregation on the first local feature to obtain a first high-order feature, and perform feature aggregation on the second local feature to obtain a second high-order feature; A first feature alignment module for aligning the first local feature and the second local feature to obtain a first fusion feature and a second fusion feature; A second feature alignment module for aligning the first high-order feature and the second high-order feature to obtain a first high-order fusion feature and a second high-order fusion feature; A matching module for determining whether the first image and the second image contain the same pedestrian based on the first fusion feature, the second fusion feature, the first high-order fusion feature, and the second high-order fusion feature; The specific function of the first feature alignment module is as follows: calculating the local feature similarity between the first local feature and the second local feature; determining a first matching feature of the second local feature in the first local feature according to the second local feature and the local feature similarity; determining a second matching feature of the first local feature in the second local feature according to the first local feature and the local feature similarity; performing feature extraction on the first matching feature to obtain a first fused feature, and performing feature extraction on the second matching feature to obtain a second fused feature. The specific function of the second feature alignment module is as follows: calculating the high-order feature similarity between the first high-order feature and the second high-order feature; determining a first high-order matching feature of the second high-order feature in the first high-order feature according to the second high-order feature and the high-order feature similarity; determining a second high-order matching feature of the first high-order feature in the second high-order feature according to the first high-order feature and the high-order feature similarity; performing feature extraction on the first high-order matching feature to obtain a first high-order fused feature, and performing feature extraction on the second high-order matching feature to obtain a second high-order fused feature.
7. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Image retrieval method, electronic equipment and related products
CN112765381A
Global and local sensing pedestrian re-identification method fusing segmentation information
CN113657355A