A target identity recognition method, device, and storage medium

By conducting face detection and feature extraction on multiple target pedestrian pictures in video surveillance scenarios, training models are built for identity recognition, solving the problem of reduced accuracy of single-modal biometrics in video surveillance, and achieving higher recognition accuracy and reliability.

CN114495220BActive Publication Date: 2025-07-22GUILIN UNIV OF ELECTRONIC TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210060310.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-19
Publication Date
2025-07-22
Estimated Expiration
2042-01-19

AI Technical Summary

Technical Problem

In video surveillance scenarios, due to factors such as camera installation position and angle, capture distance, target activity and light changes, single-modal biometrics lead to a decrease in the accuracy of target identity recognition. The existing feature fusion methods cannot effectively utilize multimodal biometrics, and different quality characteristics have different impacts on identity decisions.

Method used

By performing face detection on multiple target pedestrian pictures, pedestrian and face features are extracted, training models are constructed for identity recognition analysis, feature extraction is performed by combining PCB+RPP model and insightface model, classification processing is used using a fully connected neural network, evidence vectors are generated and overall opinion vectors are calculated to improve recognition accuracy.

Benefits of technology

It improves the accuracy and reliability of target identity recognition, can be applied to complex video surveillance scenarios, solves the impact of different quality characteristics on identity decisions, and improves the reliability and accuracy of identification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114495220B_ABST
    Figure CN114495220B_ABST
Patent Text Reader

Abstract

The present invention provides a target identity recognition method, device and storage medium, belonging to the technical field of image recognition. The method includes: S1: Import multiple target pedestrian pictures, and respectively perform face detection on each target pedestrian picture to obtain a target face picture; S2: Respectively extract features of each target pedestrian picture and each target face picture to obtain pedestrian features and face features; S3: Construct a training model, and perform identity recognition analysis on multiple pedestrian features and multiple face features through the training model to obtain a target identity recognition result. Compared with the existing method of using only face or pedestrian for identity recognition, the target identity recognition of the present invention has a higher accuracy rate, and the target recognition result has stronger reliability, and can be well applied to the video surveillance scenario, solving the problem that different quality features have different influences on the target identity decision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the technical field of image recognition, and particularly relates to a method and device for target identity recognition and a storage medium. Background Art

[0002] In the "non-cooperative" scenario of actual monitoring, affected by real factors such as the installation position and angle of the camera, the capture distance, the activities of the target, and the light change, the target identity information in the unimodal biometric feature (such as a single frontal face) is missing and the interference information increases, resulting in a sharp decline in the accuracy of target dynamic identity recognition. Researchers have pointed out that multi-modal feature fusion can obtain richer target identity information by combining multiple biometric features, thus alleviating the challenge of low accuracy of target identity recognition using unimodal features. Previous identity recognition methods based on feature fusion, such as iris and fingerprint feature fusion, face and palmprint feature fusion, iris and fingerprint feature fusion, etc., the features used in these methods cannot be extracted in the video surveillance scenario. How to make full use of the biometric features that can be collected has become a challenge for target dynamic identity recognition in the surveillance scenario.

[0003] At the same time, due to the influence of various real factors in the surveillance scenario, the quality of the features extracted from the same modality varies greatly. For example, a completely clear frontal face photo and a side face photo with a mask, and the influence of features of different qualities on target identity decision-making is also different. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method and device for target identity recognition and a storage medium in view of the deficiencies of the prior art.

[0005] The technical solution of the present invention to solve the above technical problem is as follows: A method for target identity recognition includes the following steps:

[0006] S1: Import multiple target pedestrian pictures, and perform face detection on each of the target pedestrian pictures to obtain target face pictures corresponding to each of the target pedestrian pictures;

[0007] S2: Extract features from each of the target pedestrian pictures and each of the target face pictures to obtain pedestrian features corresponding to each of the target pedestrian pictures and face features corresponding to each of the target face pictures;

[0008] S3: Construct a training model, and perform identity recognition analysis on multiple pedestrian features and multiple face features through the training model to obtain a target identity recognition result.

[0009] Another technical solution of the present invention to solve the above technical problem is as follows: A target identity recognition device includes:

[0010] A face detection module, configured to import multiple target pedestrian pictures, perform face detection on each of the target pedestrian pictures respectively, and obtain target face pictures corresponding to each of the target pedestrian pictures respectively;

[0011] A feature extraction module, configured to perform feature extraction on each of the target pedestrian pictures and each of the target face pictures respectively, and obtain pedestrian features corresponding to each of the target pedestrian pictures and face features corresponding to each of the target face pictures;

[0012] An identity recognition result obtaining module, configured to construct a training model, perform identity recognition analysis on multiple pedestrian features and multiple face features through the training model, and obtain a target identity recognition result.

[0013] Another technical solution for the present invention to solve the above technical problems is as follows: A target identity recognition device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned target identity recognition method is implemented.

[0014] Another technical solution for the present invention to solve the above technical problems is as follows: A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned target identity recognition method is implemented.

[0015] The beneficial effects of the present invention are as follows: By performing face detection on each target pedestrian picture to obtain target face pictures, performing feature extraction on each target pedestrian picture and each target face picture to obtain pedestrian features and face features, constructing a training model, and performing identity recognition analysis on multiple pedestrian features and multiple face features through the training model to obtain a target identity recognition result. Compared with the existing method of using only faces or pedestrians for identity recognition, the target identity recognition accuracy of the present invention is higher, the target recognition result has stronger reliability, and it can be well applied to the video surveillance scenario, solving the problem that different quality features have different influences on target identity decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a schematic flowchart of a target identity recognition method provided by an embodiment of the present invention;

[0017] Figure 2 It is a module block diagram of a target identity recognition device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The principles and features of the present invention are described below in conjunction with the accompanying drawings. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention.

[0019] Figure 1 It is a schematic flow chart of a target identity recognition method provided by an embodiment of the present invention.

[0020] As Figure 1 shown, a target identity recognition method includes the following steps:

[0021] S1: Import multiple target pedestrian pictures, and perform face detection on each of the target pedestrian pictures respectively to obtain target face pictures corresponding to each of the target pedestrian pictures;

[0022] S2: Extract features from each of the target pedestrian pictures and each of the target face pictures respectively to obtain pedestrian features corresponding to each of the target pedestrian pictures and face features corresponding to each of the target face pictures;

[0023] S3: Construct a training model, and perform identity recognition analysis on multiple pedestrian features and multiple face features through the training model to obtain a target identity recognition result.

[0024] It should be understood that the target pedestrian picture is complete and contains a picture of the target human body.

[0025] It should be understood that the target pedestrian picture P1 (completely containing the target human body) (i.e., multiple target pedestrian pictures) is collected, and the target face picture P2 (i.e., the target face picture) is extracted from the pedestrian picture P1 (i.e., the target pedestrian picture) to obtain the target face (i.e., the target pedestrian picture) and pedestrian data (i.e., the target face picture).

[0026] In the above embodiment, the target face picture is obtained by performing face detection on each target pedestrian picture respectively, the pedestrian features and face features are obtained by extracting features from each target pedestrian picture and each target face picture respectively, a training model is constructed, and the target identity recognition result is obtained by performing identity recognition analysis on multiple pedestrian features and multiple face features through the training model. Compared with the existing method of using only face or pedestrian for identity recognition, the target identity recognition accuracy of the present invention is higher, and the target recognition result has stronger reliability, and it can be well applied to the video surveillance scenario, solving the problem that different quality features have different influences on the target identity decision.

[0027] Optionally, as an embodiment of the present invention, in the step S1, the process of performing face detection on each of the target pedestrian pictures respectively to obtain target face pictures corresponding to each of the target pedestrian pictures includes:

[0028] Use the MTCNN face detection algorithm to perform face detection on each of the target pedestrian pictures, and obtain target face pictures corresponding to each of the target pedestrian pictures respectively.

[0029] It should be understood that the target face and pedestrian data include the target pedestrian picture P1 (i.e., the target pedestrian picture) and its corresponding face picture P2 (i.e., the target face picture), where P2 (i.e., the target face picture) is obtained by detecting and extracting the face part in P1 (i.e., the target pedestrian picture) using the MTCNN face detection algorithm.

[0030] It should be understood that the MTCNN face detection algorithm is to balance performance and accuracy and avoid the huge performance consumption brought by traditional ideas such as sliding windows plus classifiers. It first uses a small model to generate candidate bounding boxes of target regions with a certain possibility, and then uses a more complex model for fine classification and higher-precision region bounding box regression, and recursively executes this step. With this idea, a three-layer network is formed, namely P-Net, R-Net, and O-Net, to achieve fast and efficient face detection. At the input layer, an image pyramid is used for scale transformation of the initial image, and P-Net is used to generate a large number of candidate target region bounding boxes. Then, R-Net is used to perform the first selection and bounding box regression on these target region bounding boxes, excluding most of the negative examples. Then, a more complex and higher-precision network O-Net is used to discriminate and perform region bounding box regression on the remaining target region bounding boxes.

[0031] In the above embodiment, the MTCNN face detection algorithm is used to perform face detection on each target pedestrian picture to obtain the target face picture, extract a more accurate picture, and improve the accuracy of target identity recognition.

[0032] Optionally, as an embodiment of the present invention, the process of step S2 includes:

[0033] Based on the PCB+RPP model, perform pedestrian feature extraction on each of the target pedestrian pictures to obtain pedestrian features corresponding to each of the target pedestrian pictures;

[0034] Based on the insightface model, perform face feature extraction on each of the target face pictures to obtain face features corresponding to each of the target face pictures.

[0035] It should be understood that the insightface model refers to a face recognition model proposed by Deng J et al. in 2018, which combines a 100-layer ResNet backbone network with a new margin loss.

[0036] It should be understood that the PCB+RPP model refers to a classic pedestrian re-identification model proposed by Sun Y et al. in 2017, which combines uniform partitioning (PCB) and refined regionalization (RPP).

[0037] It should be understood that for the pedestrian image P1 (i.e., the target pedestrian image) and the face image P2 (i.e., the target face image), the face recognition model and the pedestrian re-identification model based on the convolutional neural network are respectively used for feature extraction to obtain the corresponding feature tensors T1 (i.e., the pedestrian feature) and T2 (i.e., the face feature).

[0038] Specifically, through the feature extraction parts of the pedestrian re-identification framework PCB+RPP (i.e., the PCB+RPP model) and the face recognition framework insightface (i.e., the insightface model), the corresponding pedestrian features and face features in the pedestrian image P1 (i.e., the target pedestrian image) and the face image P2 (i.e., the target face image) are extracted and represented by the corresponding feature tensors T1 (i.e., the pedestrian feature) and T2 (i.e., the face feature).

[0039] In the above embodiment, the pedestrian features of each target pedestrian image are extracted based on the PCB+RPP model, and the face features of each target face image are extracted based on the insightface model, providing accurate data for subsequent processing and further improving the accuracy of target identity recognition.

[0040] Optionally, as an embodiment of the present invention, the process of step S3 includes:

[0041] Construct a fully connected neural network, and classify each of the pedestrian features and each of the face features through the fully connected neural network to obtain a pedestrian evidence vector corresponding to each target pedestrian image and a face evidence vector corresponding to each target face image;

[0042] Analyze multiple pedestrian evidence vectors and multiple face evidence vectors to obtain an overall opinion vector, and use the overall opinion vector as the target identity recognition result.

[0043] It should be understood that the evidence theory is used to combine the recognition probabilities and uncertainties of faces and pedestrians to obtain the final identity recognition result (i.e., the target identity recognition result).

[0044] It should be understood that a simple fully connected neural network for multi-classification is respectively constructed for the pedestrian feature T1 (i.e., the pedestrian feature) and the face feature T2 (i.e., the face feature).

[0045] It should be understood that the feature tensor T1 (i.e., the pedestrian feature) and T2 (i.e., the face feature) are respectively used by a fully connected network to obtain the evidence vectors E1 (i.e., the pedestrian evidence vector) and E2 (i.e., the face evidence vector) of the face and the pedestrian in terms of identity categories.

[0046] In the above embodiment, the target identity recognition result is obtained through the identity recognition and analysis of multiple pedestrian features and multiple face features by the trained model, which improves the reliability and accuracy of the target recognition result, and can be well applied to the video surveillance scenario, solving the problem that different quality features have different influences on the target identity decision.

[0047] Optionally, as an embodiment of the present invention, the fully connected neural network includes a plurality of fully connected layers and RELU activation layers corresponding to the number of the fully connected layers, and the fully connected layers and the RELU activation layers are alternately connected; the process of constructing the fully connected neural network and respectively classifying each of the pedestrian features and each of the face features through the fully connected neural network to obtain the pedestrian evidence vector corresponding to each of the target pedestrian pictures and the face evidence vector corresponding to each of the target face pictures includes:

[0048] S311: Linearly map each of the pedestrian features and each of the face features through the current fully connected layer to obtain the mapped pedestrian features corresponding to each of the target pedestrian pictures and the mapped face features corresponding to each of the target face pictures;

[0049] S312: Non-linearly map each of the mapped pedestrian features and each of the mapped face features through the current RELU activation layer, input the result of the non-linear mapping into the next fully connected layer, and execute step S311 again until all the fully connected layers and all the RELU activation layers are passed, so as to obtain the pedestrian evidence vector corresponding to each of the target pedestrian pictures and the face evidence vector corresponding to each of the target face pictures.

[0050] It should be understood that both the pedestrian feature and the face feature are first input into the fully connected layer, and then the output result of the fully connected layer is input into the RELU activation layer.

[0051] It should be understood that the fully connected layer actually performs a linear mapping by combining the local features obtained from all feature extraction parts.

[0052] It should be understood that the RELU activation layer is the part where the network has non-linear mapping through the activation function.

[0053] Specifically, starting from the first fully connected layer, linear mapping is performed on each of the pedestrian features and each of the face features to obtain the mapped pedestrian features and the mapped face features; then, through the first ReLU activation layer, non-linear mapping is performed on each of the mapped pedestrian features and each of the mapped face features, and the results after non-linear mapping are used as the inputs of the second fully connected layer and input into the second fully connected layer for linear mapping. This step is alternately performed until all the fully connected layers and all the ReLU activation layers are passed through, and the results output by the last ReLU activation layer are the pedestrian evidence vectors corresponding to each of the target pedestrian pictures and the face evidence vectors corresponding to each of the target face pictures.

[0054] Specifically, the traditional multi-classification network often connects a softmax layer at the end to present the results of multi-classification in the form of probabilities. However, due to its characteristics, the softmax layer cannot represent the certainty degree of the network for the classification results. Therefore, we use the activation layer ReLU (i.e., the ReLU activation layer) to replace the softmax layer, so that the network outputs a non-negative evidence vector E, and thus we can obtain the evidence vectors of pedestrians and faces respectively. (i.e., multiple of the pedestrian evidence vectors) and (i.e., multiple of the face evidence vectors). Evidence refers to the metrics collected from the input to support classification, and K represents the number of identity IDs to be recognized.

[0055] It should be understood why not use softmax: because the softmax calculation formula will inflate the results through exponents, and its results only have comparative significance and cannot be used to represent the certainty degree of the network for the results.

[0056] In the above embodiment, through the classification processing of each pedestrian feature and each face feature by the fully connected neural network, the pedestrian evidence vector and the face evidence vector are obtained, which can more intuitively obtain the certainty degree of the results, and improve the reliability and accuracy of the target recognition results.

[0057] Optionally, as an embodiment of the present invention, each of the pedestrian evidence vectors includes multiple pedestrian evidence values corresponding to the identity of the target to be recognized, and each of the face evidence vectors includes multiple face evidence values corresponding to the identity of the target to be recognized. The process of analyzing the overall opinion vector from the multiple pedestrian evidence vectors and the multiple face evidence vectors includes:

[0058] By the first formula, each of the pedestrian evidence values is respectively converted into a Dirichlet distribution to obtain the pedestrian Dirichlet distributions corresponding to each of the identities of the targets to be recognized. The first formula is:

[0059]

[0060] Among them, is the k-th pedestrian evidence value, where k is the identity of the k-th target to be recognized, is the k-th pedestrian Dirichlet distribution;

[0061] Each of the face evidence values is converted into a Dirichlet distribution through the second formula, and the face Dirichlet distributions corresponding to the identities of the respective targets to be recognized are obtained. The second formula is:

[0062]

[0063] Among them, is the k-th face evidence value, where k is the identity of the k-th target to be recognized, is the k-th face Dirichlet distribution;

[0064] The confidence quality of each of the face Dirichlet distributions is calculated through the third formula, and the pedestrian Dirichlet intensity and the pedestrian confidence quality corresponding to the identities of the respective targets to be recognized are obtained. The third formula is:

[0065]

[0066] Among them,

[0067] Among them, is the k-th pedestrian confidence quality, is the k-th pedestrian Dirichlet distribution, is the i-th pedestrian Dirichlet distribution, where i is the i-th target pedestrian picture, S 1 is the pedestrian Dirichlet intensity, and K is the total number of identities of the targets to be recognized;

[0068] The confidence quality of each of the face Dirichlet distributions is calculated through the fourth formula, and the face Dirichlet intensity and the face confidence quality corresponding to the identities of the respective targets to be recognized are obtained. The fourth formula is:

[0069]

[0070] Among them,

[0071] Among them, is the k-th face confidence quality, the k-th face Dirichlet distribution, is the i-th face Dirichlet distribution, where i is the i-th target face picture, S 2 is the face Dirichlet intensity, and K is the total number of identities of the targets to be recognized;

[0072] The uncertainty of the pedestrian Dirichlet intensity is calculated through the fifth formula to obtain the pedestrian uncertainty. The fifth formula is:

[0073]

[0074] where u 1 is the pedestrian uncertainty, K is the total number of target identities to be recognized, and S 1 is the pedestrian Dirichlet intensity;

[0075] The uncertainty of the face Dirichlet intensity is calculated through the sixth formula to obtain the face uncertainty. The sixth formula is:

[0076]

[0077] where u 2 is the pedestrian uncertainty or the face uncertainty, K is the total number of target identities to be recognized, and S 2 is the face Dirichlet intensity;

[0078] The pedestrian opinion vector is constructed through the seventh formula for all pedestrian confidence qualities and the pedestrian uncertainty to obtain the pedestrian opinion vector. The seventh formula is:

[0079]

[0080] where M 1 is the pedestrian opinion vector, is the confidence quality of the k-th pedestrian, and u 1 is the pedestrian uncertainty;

[0081] The face opinion vector is constructed through the eighth formula for all face confidence qualities and the face uncertainty to obtain the face opinion vector. The eighth formula is:

[0082]

[0083] where M 2 is the face opinion vector, is the confidence quality of the K-th face, and u 2 is the face uncertainty;

[0084] The overall opinion vector is calculated through the ninth formula for the pedestrian opinion vector and the face opinion vector to obtain the overall opinion vector. The ninth formula is:

[0085] M = [b1, b2, …, b K , u],

[0086] where

[0087] where

[0088] Among them, M is the overall opinion vector, i, j ∈ k, b k is the k-th overall confidence quality, u is the certainty degree, is the k-th pedestrian confidence quality, u 1 is the pedestrian uncertainty, is the k-th face confidence quality, u 2 is the face uncertainty, is the scale factor, C is the measurement sum, and K is the total number of target identities to be recognized.

[0089] It should be understood that the evidence (i.e., the pedestrian evidence vector or the face evidence vector) is converted into a Dirichlet distribution (i.e., the pedestrian Dirichlet distribution or the face Dirichlet distribution), where the +1 here is to make the prior distribution become a uniform distribution, that is, assuming no evidence is observed, the opinion α v = [1, 1,..., 1] corresponds to a uniform distribution, indicating complete uncertainty. When v is 1, it is to process the pedestrian evidence vector; when v is 2, it is to process the face evidence vector.

[0090] Specifically, the confidence quality and the uncertainty u v are calculated through the following two formulas

[0091]

[0092]

[0093] Among them, is called the Dirichlet strength, and the opinion vector is composed of the confidence quality v and the uncertainty u

[0094] For example: when the neural network output is E = [40, 1, 1], calculate the Dirichlet distribution α = [41, 2, 2], and obtain the opinion vector Finally, the opinion vectors M 1 (i.e., the pedestrian opinion vector) and M 2 (i.e., the face opinion vector) of the pedestrian and face parts are obtained respectively through the above method.

[0095] It should be understood that when generating the Dirichlet distribution opinion of a multi-classification task through a neural network, when an observation result of a sample is related to one of the K class attributes, the corresponding Dirichlet parameter will increase to update the Dirichlet distribution using the new observation result. The Dirichlet distribution parameterized by evidence represents the density of each such probability assignment, thereby simulating second-order probability and uncertainty.

[0096] Specifically, after obtaining the opinion vectors of the pedestrian part and the face part respectively (i.e., the pedestrian opinion vector) and (i.e., the face opinion vector), the overall opinion vector M = [b1, b2, …, b K , u] is obtained by combining the two opinion vectors, and the calculation method is as follows:

[0097]

[0098]

[0099] Where represents the measure sum of the conflicting parts of the two opinion vectors, is the scaling factor for normalization. The largest b in the overall opinion vector represents the final result of identity recognition, and u represents the confidence level of this result.

[0100] It should be understood that k is the kth target identity to be recognized. For example, when calculating b1, all k inside are 1, and when calculating b2, all k inside are 2.

[0101] In the above embodiment, the target identity recognition result is obtained by separately analyzing and recognizing the acceptance region, rejection region, and uncertainty region, which can adopt a suitable recognition method for the target in different scenarios. Different from the single recognition method that only uses the face or the pedestrian, the recognition accuracy of the target identity has been significantly improved and can be applied to complex real situations.

[0102] Optionally, as an embodiment of the present invention, the step S3 further includes:

[0103] Importing multiple true values corresponding to the target identity to be recognized, and calculating the total distribution of the overall opinion vector through the tenth formula to obtain the total distribution. The tenth formula is:

[0104] α = [α1, α2, …, α K ,

[0105] Where,

[0106] Where α is the total distribution, b k is the kth overall confidence quality, u is the confidence level, and K is the total number of target identities to be recognized;

[0107] The overall loss value is calculated for the pedestrian Dirichlet intensity, the face Dirichlet intensity, the overall distribution, all true values, all pedestrian Dirichlet distributions, and all face Dirichlet distributions through the eleventh formula, obtaining the overall loss value. The eleventh formula is:

[0108]

[0109] where,

[0110]

[0111] where,

[0112]

[0113] where,

[0114] where,

[0115] where L overall is the overall loss value, L(α i ) is the total opinion loss, is the pedestrian opinion loss, is the face opinion loss, i is the i-th target pedestrian picture or the i-th target face picture, N is the total number of target pedestrian pictures or target face pictures, K is the total number of target identities to be recognized, k is the k-th target identity to be recognized, λ t is the balance factor, Γ() is the gamma function, Ψ() is the digamma function, ⊙ is the Hadamard product, y k is the k-th true value, S 1 is the pedestrian Dirichlet intensity, S 2 is the face Dirichlet intensity, is the k-th pedestrian Dirichlet distribution, the k-th face Dirichlet distribution, is the multinomial view formed by the Dirichlet distribution with parameter α i , D(p|1) is the multinomial view formed by the Dirichlet distribution with parameter 1, is the multinomial view formed by the Dirichlet distribution with parameter , is the multinomial view formed by the Dirichlet distribution with parameter , is the relative entropy, and p is the class assignment probability;

[0116] Update the parameters of the training model according to the overall loss value, and return to step S2 until the number of iterations is reached, thereby obtaining the target training model.

[0117] It should be understood that since the neural network is not trained with one picture at a time but with a batch of pictures at a time, i represents the number of the picture in a batch, and N represents the total number of pictures in a batch.

[0118] It should be understood that the overall loss value includes the sum of the total opinion loss and each part of the opinion loss (i.e., the pedestrian opinion loss and the face opinion loss).

[0119] Specifically, This part is a simple modification of the cross-entropy loss. Specifically, it is the integral of the product of the cross-entropy loss and the probability density function of the Dirichlet distribution. Since L ace (α i ) only constrains the correct label to produce more evidence and does not constrain the wrong label to produce less evidence. The opinion loss is obtained by introducing the KL loss, that is constrain the wrong label to shrink to the expected value of 0.

[0120] In the above embodiment, the total distribution is calculated for the total distribution of the overall opinion vector through the tenth formula, and the overall loss value of the pedestrian Dirichlet intensity, face Dirichlet intensity, total distribution, all true values, all pedestrian Dirichlet distributions, and all face Dirichlet distributions is calculated through the eleventh formula. The parameters of the training model are updated according to the overall loss value until the number of iterations is reached to obtain the target training model, which can constrain the wrong label to shrink to the expected value, improve the accuracy of target identity recognition, and can be well applied to the video surveillance scenario, solving the problem that different quality features have different impacts on target identity decision-making.

[0121] Figure 2 It is a block diagram of a target identity recognition device provided by an embodiment of the present invention.

[0122] Optionally, as another embodiment of the present invention, as Figure 2 shown, a target identity recognition device includes:

[0123] A face detection module, configured to import multiple target pedestrian pictures, perform face detection on each of the target pedestrian pictures respectively, and obtain target face pictures corresponding to each of the target pedestrian pictures;

[0124] A feature extraction module, configured to extract features from each of the target pedestrian images and each of the target face images respectively, to obtain pedestrian features corresponding to each of the target pedestrian images and face features corresponding to each of the target face images;

[0125] An identity recognition result obtaining module, configured to construct a training model, and perform identity recognition analysis on multiple pedestrian features and multiple face features through the training model to obtain a target identity recognition result.

[0126] Optionally, another embodiment of the present invention provides a target identity recognition device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned target identity recognition method is implemented. The device may be a computer or the like.

[0127] Optionally, another embodiment of the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned target identity recognition method is implemented.

[0128] It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

[0129] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0130] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0131] The unit described as a separation component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention.

[0132] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0133] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0134] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A target identity recognition method, characterized in that, It includes the following steps: S1: Import multiple target pedestrian pictures, and perform face detection on each of the target pedestrian pictures respectively to obtain target face pictures corresponding to each of the target pedestrian pictures; S2: Extract features from each of the target pedestrian pictures and each of the target face pictures respectively to obtain pedestrian features corresponding to each of the target pedestrian pictures and face features corresponding to each of the target face pictures; S3: Construct a training model, and perform identity recognition analysis on multiple pedestrian features and multiple face features through the training model to obtain a target identity recognition result; The process of step S3 includes: Construct a fully connected neural network, and perform classification processing on each of the pedestrian features and each of the face features through the fully connected neural network to obtain a pedestrian evidence vector corresponding to each of the target pedestrian pictures and a face evidence vector corresponding to each of the target face pictures; Analyze multiple pedestrian evidence vectors and multiple face evidence vectors to obtain an overall opinion vector, and use the overall opinion vector as the target identity recognition result; The fully connected neural network includes multiple fully connected layers and RELU activation layers corresponding to the number of fully connected layers, and the fully connected layers and the RELU activation layers are alternately connected; the process of constructing the fully connected neural network and performing classification processing on each of the pedestrian features and each of the face features through the fully connected neural network to obtain a pedestrian evidence vector corresponding to each of the target pedestrian pictures and a face evidence vector corresponding to each of the target face pictures includes: S311: Perform linear mapping on each of the pedestrian features and each of the face features through the current fully connected layer to obtain mapped pedestrian features corresponding to each of the target pedestrian pictures and mapped face features corresponding to each of the target face pictures; S312: Perform non-linear mapping on each of the mapped pedestrian features and each of the mapped face features through the current RELU activation layer respectively, and input the result of the non-linear mapping into the next fully connected layer, and then execute step S311 again until all the fully connected layers and all the RELU activation layers are passed, so as to obtain a pedestrian evidence vector corresponding to each of the target pedestrian pictures and a face evidence vector corresponding to each of the target face pictures.

2. The target identity recognition method according to claim 1, wherein In step S1, the process of performing face detection on each of the target pedestrian pictures respectively to obtain target face pictures corresponding to each of the target pedestrian pictures includes: Use the MTCNN face detection algorithm to perform face detection on each of the target pedestrian pictures respectively to obtain target face pictures corresponding to each of the target pedestrian pictures.

3. The target identity recognition method according to claim 1, wherein The process of step S2 includes: Extract pedestrian features from each of the target pedestrian pictures based on the PCB+RPP model to obtain pedestrian features corresponding to each of the target pedestrian pictures; Based on the insightface model, face features are extracted from each of the target face images to obtain face features corresponding to each of the target face images.

4. The target identity recognition method according to claim 1, wherein Each of the pedestrian evidence vectors includes multiple pedestrian evidence values corresponding to the target identity to be recognized, and each of the face evidence vectors includes multiple face evidence values corresponding to the target identity to be recognized. The process of analyzing the multiple pedestrian evidence vectors and the multiple face evidence vectors to obtain the overall opinion vector includes: Converting each of the pedestrian evidence values into a Dirichlet distribution respectively through the first formula to obtain a pedestrian Dirichlet distribution corresponding to each of the target identities to be recognized. The first formula is: Among them, is the k-th pedestrian evidence value, where k is the identity of the k-th target to be recognized, is the k-th pedestrian Dirichlet distribution; Converting each of the face evidence values into a Dirichlet distribution respectively through the second formula to obtain a face Dirichlet distribution corresponding to each of the target identities to be recognized. The second formula is: Among them, is the k-th face evidence value, where k is the identity of the k-th target to be recognized, is the k-th face Dirichlet distribution; Calculating the confidence quality of each of the face Dirichlet distributions respectively through the third formula to obtain the pedestrian Dirichlet intensity and the pedestrian confidence quality corresponding to each of the target identities to be recognized. The third formula is: Among them, Among them, is the confidence quality of the k-th pedestrian, is the Dirichlet distribution of the k-th pedestrian, is the Dirichlet distribution of the i-th pedestrian, where i is the i-th target pedestrian picture, S 1 is the pedestrian Dirichlet intensity, and K is the total number of target identities to be recognized; Calculating the confidence quality of each of the face Dirichlet distributions respectively through the fourth formula to obtain the face Dirichlet intensity and the face confidence quality corresponding to each of the target identities to be recognized. The fourth formula is: Among them, Among them, is the confidence quality of the k-th face, the Dirichlet distribution of the k-th face, is the Dirichlet distribution of the i-th face, where i is the i-th target face image, and S 2 is the face Dirichlet intensity, and K is the total number of target identities to be recognized; Calculating the uncertainty of the pedestrian Dirichlet intensity through the fifth formula to obtain the pedestrian uncertainty. The fifth formula is: where u 1 is the pedestrian uncertainty, K is the total number of target identities to be recognized, and S 1 is the pedestrian Dirichlet intensity; Calculating the uncertainty of the face Dirichlet intensity respectively through the sixth formula to obtain the face uncertainty. The sixth formula is: where, u 2 is the pedestrian uncertainty or face uncertainty, K is the total number of target identities to be recognized, and S 2 is the face Dirichlet intensity; Constructing a pedestrian opinion vector for all the pedestrian confidence qualities and the pedestrian uncertainty through the seventh formula to obtain the pedestrian opinion vector. The seventh formula is: where, M 1 is the pedestrian opinion vector, is the confidence quality of the k-th pedestrian, and u 1 is the pedestrian uncertainty; Constructing a face opinion vector for all the face confidence qualities and the face uncertainty through the eighth formula to obtain the face opinion vector. The eighth formula is: Among them, M 2 is the face opinion vector, is the confidence quality of the K-th face, and u 2 is the face uncertainty; Calculating the overall opinion vector for the pedestrian opinion vector and the face opinion vector through the ninth formula to obtain the overall opinion vector. The ninth formula is: M = [b1, b2, …, b K , u], Among them, Among them, Among them, M is the overall opinion vector, i, j ∈ k, b k is the k-th overall confidence quality, u is the degree of certainty, is the k-th pedestrian confidence quality, u 1 is the pedestrian uncertainty, is the k-th face confidence quality, u 2 is the face uncertainty, is the scale factor, C is the measure sum, and K is the total number of target identities to be recognized.

5. The target identity recognition method according to claim 4, wherein The step S3 further includes: Importing multiple true values corresponding to the target identity to be recognized, and calculating the total distribution for the overall opinion vector through the tenth formula to obtain the total distribution. The tenth formula is: α=[α1,α2,…,α K ], Among them, where α is the total distribution, b k is the confidence quality of the k-th population, u is the degree of certainty, and K is the total number of target identities to be recognized; Calculating the overall loss value for the pedestrian Dirichlet intensity, the face Dirichlet intensity, the total distribution, all the true values, all the pedestrian Dirichlet distributions, and all the face Dirichlet distributions through the eleventh formula to obtain the overall loss value. The eleventh formula is: Among them, Among them, Among them, Among them, Among them, L overall is the overall loss value, L(α i ) is the total opinion loss, is the pedestrian opinion loss, is the face opinion loss, i is the i-th target pedestrian image or the i-th target face image, N is the total number of target pedestrian images or target face images, K is the total number of target identities to be recognized, k is the k-th target identity to be recognized, λ t is the balance factor, Γ() is the gamma function, Ψ() is the digamma function, ⊙ is the Hadamard product, y k is the k-th true value, S 1 is the pedestrian Dirichlet intensity, S 2 is the face Dirichlet intensity, is the k-th pedestrian Dirichlet distribution, the k-th face Dirichlet distribution, is the multinomial view formed by the Dirichlet distribution with parameter α i , D(p|1) is the multinomial view formed by the Dirichlet distribution with parameter 1, is the multinomial view formed by the Dirichlet distribution with parameter , is the multinomial view formed by the Dirichlet distribution with parameter , is the relative entropy, p is the class assignment probability; Updating the parameters of the training model according to the overall loss value, and returning to step S2 until the iteration times are reached, so as to obtain the target training model.

6. A target identity recognition device, characterized in that, Including: A face detection module, configured to import multiple target pedestrian images, and perform face detection on each of the target pedestrian images respectively to obtain target face images corresponding to each of the target pedestrian images; A feature extraction module, configured to perform feature extraction on each of the target pedestrian pictures and each of the target face pictures respectively, to obtain pedestrian features corresponding to each of the target pedestrian pictures and face features corresponding to each of the target face pictures; An identity recognition result obtaining module, configured to construct a training model, and perform identity recognition analysis on a plurality of the pedestrian features and a plurality of the face features through the training model, to obtain a target identity recognition result; Specifically, the identity recognition result obtaining module is configured to: Construct a fully-connected neural network, and perform classification processing on each of the pedestrian features and each of the face features through the fully-connected neural network, to obtain a pedestrian evidence vector corresponding to each of the target pedestrian pictures and a face evidence vector corresponding to each of the target face pictures; Analyze a plurality of the pedestrian evidence vectors and a plurality of the face evidence vectors to obtain an overall opinion vector, and use the overall opinion vector as the target identity recognition result; The fully-connected neural network includes a plurality of fully-connected layers and RELU activation layers corresponding to the number of the fully-connected layers, and the fully-connected layers and the RELU activation layers are alternately connected; the process of constructing the fully-connected neural network and performing classification processing on each of the pedestrian features and each of the face features through the fully-connected neural network to obtain a pedestrian evidence vector corresponding to each of the target pedestrian pictures and a face evidence vector corresponding to each of the target face pictures includes: S311: Perform linear mapping on each of the pedestrian features and each of the face features through a current fully-connected layer, to obtain mapped pedestrian features corresponding to each of the target pedestrian pictures and mapped face features corresponding to each of the target face pictures; S312: Perform non-linear mapping on each of the mapped pedestrian features and each of the mapped face features through a current RELU activation layer, input the result after non-linear mapping into the next fully-connected layer, and perform step S311 again until all the fully-connected layers and all the RELU activation layers are passed through, so as to obtain a pedestrian evidence vector corresponding to each of the target pedestrian pictures and a face evidence vector corresponding to each of the target face pictures.

7. A target identity recognition system, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the target identity recognition method according to any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the target identity recognition method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Target tracking and identity recognition method based on dual-task learning

    CN110796072A

  • Pedestrian recognition method and system

    CN112560720A