Face recognition method and system based on skeleton calibration algorithm
Through the bone calibration algorithm and adaptive distillation model, combined with the coupling degree calculation and attention mechanism, the accuracy and stability of face recognition in complex occlusion scenarios are solved, and efficient and accurate no-manual intervention recognition is achieved.
Patent Information
- Application Number
- CN202510587295.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-15
AI Technical Summary
The prior art has poor accuracy and stability in complex occlusion scenarios, insufficient generalization ability, and relies on a large amount of labeled data and manual intervention, resulting in low recognition efficiency.
Using a bone calibration algorithm, the three-dimensional bone coordinate sequence and pixel point feature information is extracted from the video stream through preset adaptive distillation model and attention mechanism, and the three-dimensional bone coordinate sequence and pixel point feature information are calculated, and the coupling degree value and multimodal feature aggregation amount are calculated to achieve completion and accurate identification of the occluded face.
It improves the recognition accuracy in complex occlusion scenarios, enhances generalization capabilities, reduces labor costs, and achieves efficient face recognition without manual intervention.
Smart Images

Figure CN120496145A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of face recognition technology, and in particular to a face recognition method and system based on a skeleton calibration algorithm. Background Art
[0002] With the development of artificial intelligence technology, facial recognition has been widely used in security monitoring, access control, financial payment and other fields. Traditional facial recognition methods mainly rely on two-dimensional image features, extracting facial texture, contour and other information for identity recognition.
[0003] However, in actual application scenarios, face recognition often faces many challenges, such as complex lighting conditions, posture changes, occlusion (such as wearing masks, sunglasses, hats, etc.), etc. These factors can cause facial features to be incomplete or change, seriously affecting the accuracy and stability of recognition.
[0004] To address the occlusion problem, some existing technologies attempt to directly extract and complete features of occluded areas using deep learning models. However, these methods often rely on large amounts of labeled data and lack generalization capabilities in complex occlusion scenarios. Furthermore, in practice, manual intervention may be required to complete the recognition process, which is labor-intensive.
[0005] Therefore, there is an urgent need for a technical solution that can be used for more flexible, accurate and efficient face recognition in complex occlusion scenes. Summary of the Invention
[0006] To solve the above problems, the embodiments of the present application provide a face recognition method and system based on a skeleton calibration algorithm, which is used to achieve more flexible, accurate and efficient face recognition in complex occlusion scenes.
[0007] On the one hand, an embodiment of the present application provides a face recognition method based on a skeleton calibration algorithm, the method comprising:
[0008] Determine, based on a face recognition video stream from a preset face recognition device group, each first three-dimensional bone coordinate sequence and each first pixel feature information of an object to be recognized within a preset time;
[0009] Determine a corresponding second three-dimensional skeleton coordinate sequence based on the first three-dimensional skeleton coordinate sequence, the first pixel feature information, and a preset adaptive distillation model; wherein the second three-dimensional skeleton coordinate sequence is a three-dimensional skeleton coordinate sequence after completing the occluded face of the object to be identified;
[0010] Calculating a coupling value corresponding to each selected pixel point in a preset area based on a preset coupling calculation formula, each of the first three-dimensional bone coordinate sequences, and each of the first pixel feature information, so as to determine a corresponding multimodal feature aggregation amount based on a weighted calculation result of each of the coupling values, the first pixel feature information, and the second three-dimensional bone coordinate sequence; the coupling value is used to characterize the degree of dynamic association between the bone and the pixel point;
[0011] Based on the aggregation amount of each multimodal feature, the second three-dimensional bone coordinate sequence and the preset attention mechanism, the bone pixel feature vector is determined, and the face recognition result of the object to be identified is determined according to the bone pixel feature vector and the preset face recognition database.
[0012] In one implementation of the present application, based on a face recognition video stream from a preset face recognition device group, determining each first three-dimensional skeletal coordinate sequence and each first pixel feature information of an object to be recognized within a preset time period specifically includes:
[0013] Inputting the face recognition video stream from the preset facial recognition device group into a pre-trained skeleton calibration model to determine the coordinates of a plurality of skeleton approximate points corresponding to the object to be recognized in the face recognition video stream according to the model output result;
[0014] Mapping the coordinates of each of the skeleton approximate points to a preset depth map coordinate system, and sequentially adding the skeleton mapping coordinate points to each of the first three-dimensional skeleton coordinate sequences in chronological order;
[0015] Determining RGB images of each face in the face recognition video stream; wherein each frame of the face recognition video stream includes an RGB image and a depth image;
[0016] Dividing each of the facial RGB images into a grid according to a preset grid to determine a feature vector corresponding to each of the facial RGB images as second pixel feature information;
[0017] According to each of the first three-dimensional bone coordinate sequences and each of the second pixel feature information, the corresponding bone pixel correlation coefficient matrix is determined, so as to determine the first pixel feature information according to the bone pixel correlation coefficient matrix, the preset correlation coefficient threshold and the second pixel feature information.
[0018] In one implementation of the present application, determining a corresponding bone pixel correlation coefficient matrix according to each of the first three-dimensional bone coordinate sequences and each of the second pixel feature information specifically includes:
[0019] Determine a dynamic bone pixel correlation coefficient based on the first three-dimensional bone coordinate sequence and the second pixel feature information corresponding to the same frame, and construct a dynamic bone pixel correlation coefficient set corresponding to each video frame; the dynamic bone pixel correlation coefficient is calculated based on the distance between the bone mapping coordinate point and the corresponding grid center;
[0020] The dynamic skeleton pixel association coefficient set is matched with a preset standard association coefficient library to determine the skeleton pixel association coefficient matrix according to the matching result; wherein the preset standard association coefficient set is constructed based on an unobstructed face recognition video stream of a pre-recorded object.
[0021] In one implementation of the present application, determining the first pixel feature information according to the bone pixel correlation coefficient matrix, the preset correlation coefficient threshold, and the second pixel feature information specifically includes:
[0022] Traversing the bone pixel correlation coefficient matrix, and screening correlation coefficients that are greater than the preset correlation coefficient threshold as screening correlation coefficients;
[0023] The pixel corresponding to the filtered correlation coefficient in the second pixel feature information is determined as the first pixel, so as to determine the first pixel feature information according to the first pixel obtained by filtering.
[0024] In one implementation of the present application, determining a corresponding second three-dimensional bone coordinate sequence based on the first three-dimensional bone coordinate sequence, the first pixel feature information, and a preset adaptive distillation model specifically includes:
[0025] Inputting the first three-dimensional bone coordinate sequence and the first pixel feature information into the preset adaptive distillation model to match the initial complete bone coordinate sequence in the standard face template library through a pre-trained teacher model;
[0026] Performing convolution processing on the preprocessed first pixel feature information and the first three-dimensional bone coordinate sequence using a student model to determine a pending three-dimensional bone coordinate sequence; the student model is trained based on a number of three-dimensional bone coordinate sequence samples of occluded faces;
[0027] A knowledge distillation loss function value is calculated based on the initial complete bone coordinate sequence and the undetermined three-dimensional bone coordinate sequence, so that when the knowledge distillation loss function value is less than a preset threshold, the undetermined three-dimensional bone coordinate sequence is used as the second three-dimensional bone coordinate sequence.
[0028] In one implementation of the present application, the coupling value corresponding to each selected pixel point in the preset area is calculated according to a preset coupling calculation formula, each of the first three-dimensional bone coordinate sequences and each of the first pixel feature information, specifically including:
[0029] Calculating the displacement variance of each bone point coordinate according to each of the first three-dimensional bone coordinate sequences, using the inverse of the displacement variance of the bone point coordinate as the time series stability value of the bone point, and calculating the pixel point texture entropy according to each of the first pixel point feature information;
[0030] Determining, based on each of the first three-dimensional bone coordinate sequences and the first pixel feature information, a Euclidean distance value between each bone point and each of the selected pixel points; wherein a correlation coefficient of the selected pixel points is greater than a preset correlation coefficient threshold;
[0031] The temporal stability value of the skeleton point, the texture entropy of the pixel point and the Euclidean distance value are input into the preset coupling degree calculation formula to calculate the coupling degree value of the selected pixel point.
[0032] In one implementation of the present application, determining a corresponding multimodal feature aggregation amount based on a weighted calculation result of each of the coupling values, the first pixel feature information, and the second three-dimensional bone coordinate sequence specifically includes:
[0033] Determining a facial occlusion ratio value of the object to be identified based on the feature information of each of the first pixels and a preset occlusion ratio prediction model, and determining a pixel feature weight and a bone feature weight based on the facial occlusion ratio value; the sum of the pixel feature weight and the bone feature weight is 1;
[0034] Performing vector encoding on the feature information of each selected pixel point in the first pixel feature information to obtain a first feature vector, and performing vector encoding on the feature attributes corresponding to each bone point in the second three-dimensional bone coordinate sequence to determine a second feature vector; wherein the feature attributes corresponding to the bone point include at least position code, motion speed, and posture;
[0035] Calculate a weighted sum of the first eigenvector and the second eigenvector of the k-th skeleton point according to the pixel feature weight and the skeleton feature weight, and use the weighted sum as the k-th first component; k is the total number of the selected pixels, and k is a natural number;
[0036] The product value of the kth first component and the coupling value of the kth skeleton point is calculated as the kth second component, and the sum of the k second components is calculated to obtain the multimodal feature aggregation amount of the corresponding selected pixel point.
[0037] In one implementation of the present application, determining a bone pixel feature vector based on each of the multimodal feature aggregation amounts, the second three-dimensional bone coordinate sequence, and a preset attention mechanism specifically includes:
[0038] Performing feature splicing on each of the multimodal feature aggregation amounts and the second three-dimensional bone coordinate sequence to obtain a splicing vector;
[0039] Input the concatenated vector into the preset attention mechanism to map the query vector, key vector, and value vector, and calculate the attention mechanism output value corresponding to each attention head;
[0040] The output values of each attention mechanism are concatenated and linearly transformed to obtain the bone pixel feature vector.
[0041] In one implementation of the present application, determining the face recognition result of the object to be identified based on the skeleton pixel feature vector and a preset face recognition database specifically includes:
[0042] Calculating the posterior probability between the bone pixel feature vector and each identity entry information in the preset face recognition database according to a preset Bayesian decision rule;
[0043] The maximum posterior probability value among the posterior probabilities is determined, and the identity identifier corresponding to the maximum posterior probability value is used as the face recognition result of the object to be identified.
[0044] On the other hand, an embodiment of the present application further provides a face recognition system based on a skeleton calibration algorithm, the system comprising:
[0045] A first determination module is configured to determine, based on a face recognition video stream from a preset face recognition device group, each first three-dimensional bone coordinate sequence and each first pixel feature information of an object to be recognized within a preset time;
[0046] A second determination module is configured to determine a corresponding second three-dimensional skeleton coordinate sequence based on the first three-dimensional skeleton coordinate sequence, the first pixel feature information, and a preset adaptive distillation model; wherein the second three-dimensional skeleton coordinate sequence is a three-dimensional skeleton coordinate sequence after completing the occluded face of the object to be identified;
[0047] a calculation module, configured to calculate a coupling value corresponding to each selected pixel point in a preset area according to a preset coupling calculation formula, each of the first three-dimensional bone coordinate sequences, and each of the first pixel feature information, so as to determine a corresponding multimodal feature aggregation amount based on a weighted calculation result of each of the coupling values, the first pixel feature information, and the second three-dimensional bone coordinate sequence; the coupling value is used to represent the degree of dynamic association between the bone and the pixel point;
[0048] The third determination module is used to determine the bone pixel feature vector based on each of the multimodal feature aggregation amounts, the second three-dimensional bone coordinate sequence and the preset attention mechanism, and determine the face recognition result of the object to be identified based on the bone pixel feature vector and the preset face recognition database.
[0049] Compared with the prior art, this application has the following significant effects:
[0050] The above scheme uses a pre-set adaptive distillation model and the collaboration of the teacher model and student model to complete the second three-dimensional bone coordinate sequence of the occluded face based on the first three-dimensional bone coordinate sequence and the first pixel feature information. This provides complete bone feature information for accurate recognition, effectively overcomes the interference of occlusion on recognition, improves recognition accuracy in complex occlusion scenes, and can cope with different occlusion levels and scenes, enhancing the generalization ability of face recognition. At the same time, the coupling value is calculated to characterize the dynamic relationship between bones and pixels, fully exploring the intrinsic connection between multimodal data and calculating the multimodal feature aggregation amount, making the feature expression more comprehensive and enhancing the accuracy of recognition. It does not require human intervention and reduces labor costs. This solves the problems of low face recognition accuracy, poor stability, weak generalization ability and high labor costs in complex occlusion scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0052] Figure 1 A schematic diagram of a flow chart of a face recognition method based on a skeleton calibration algorithm in an embodiment of the present application;
[0053] Figure 2 This is a structural diagram of a face recognition system based on a skeleton calibration algorithm in an embodiment of the present application. DETAILED DESCRIPTION
[0054] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0055] The embodiments of the present application provide a face recognition method and system based on a skeleton calibration algorithm, which is used to solve the current technical problems of insufficient generalization ability of face recognition in complex occlusion scenarios, making it difficult to perform recognition flexibly, accurately and efficiently, and consuming manpower costs.
[0056] The following describes in detail various embodiments of the present application with reference to the accompanying drawings.
[0057] The embodiment of the present application provides a face recognition method based on a skeleton calibration algorithm, such as Figure 1 As shown, the method may include steps S101-S104:
[0058] S101, the server determines each first three-dimensional skeleton coordinate sequence and each first pixel feature information of an object to be recognized within a preset time based on a face recognition video stream from a preset face recognition device group.
[0059] It should be noted that the server, as the executor of the face recognition method based on the skeleton calibration algorithm, is only an example. The executor is not limited to the server, and this application does not make any specific restrictions on this.
[0060] The preset facial recognition device group includes at least an RGB camera and a depth camera. The face recognition video stream can be understood as a video shot of the face of the object to be recognized within a preset time. The present application can be applied to access control recognition scenarios such as hospitals and laboratories. The bone points corresponding to the first three-dimensional bone coordinate sequence include at least one or more of the following: frontal bone, eye socket, nasal bone, zygomatic bone, mandible, and dental occlusion line. The present application can pre-train a convolutional neural network model, which is obtained by training a number of RGB images and depth images that mark bone points. Thus, the bone points are identified through the face recognition video stream and the convolutional neural network model.
[0061] In an embodiment of the present application, based on a face recognition video stream from a preset face recognition device group, determining each first three-dimensional skeletal coordinate sequence and each first pixel feature information of an object to be recognized within a preset time specifically includes:
[0062] The face recognition video stream from the preset face recognition device group is input into the pre-trained bone calibration model to determine the coordinates of several bone approximate points corresponding to the object to be recognized in the face recognition video stream based on the model output results. The coordinates of each bone approximate point are mapped to a preset depth map coordinate system, and the bone mapping coordinate points are added to each first three-dimensional bone coordinate sequence in chronological order. The RGB images of each face in the face recognition video stream are determined. Each frame of the face recognition video stream contains an RGB image and a depth image. According to the preset grid, each facial RGB image is grid-divided to determine the feature vector corresponding to each facial RGB image as the second pixel feature information. According to each first three-dimensional bone coordinate sequence and each second pixel feature information, the corresponding bone pixel correlation coefficient matrix is determined, so as to determine the first pixel feature information based on the bone pixel correlation coefficient matrix, the preset correlation coefficient threshold and the second pixel feature information.
[0063] That is, the present application pre-trains a skeleton calibration model, a machine learning model trained on a large amount of data, which is used to analyze and detect the coordinates of several skeletal approximate points corresponding to the object to be identified from the face recognition video stream. For example, the model can determine the coordinates of these skeletal approximate points by identifying key facial features, such as the corners of the eyes, the tip of the nose, and the corners of the mouth. These coordinates reflect the approximate position of the facial bones in the video stream. Subsequently, the obtained coordinates of each skeletal approximate point are mapped to a preset depth map coordinate system. The depth map coordinate system is a coordinate system used to represent the depth information of an object in three-dimensional space. Through this mapping, the position of the skeletal points can be combined with the depth information. Then, the skeletal mapping coordinate points are sequentially added to each first three-dimensional skeletal coordinate sequence in chronological order. Because the face recognition video stream is composed of a series of continuous frames, the position of the skeletal points will change over time. Therefore, recording these coordinate points in chronological order forms a first three-dimensional skeletal coordinate sequence that can reflect the movement and position changes of the bones over a period of time.
[0064] Subsequently, the server can also extract the RGB image, and extract the facial RGB image containing the color and texture information of the face. Then, each facial RGB image is grid-divided according to the preset grid. The preset grid is a predefined division method that divides the facial RGB image into multiple small grid areas. For each grid area, a feature vector is determined by a certain algorithm (such as extracting the color, texture and other features of the pixels in the area), and these feature vectors constitute the second pixel feature information. In this way, the features of the facial RGB image can be represented in a structured manner, which is convenient for subsequent analysis and processing. Then, the present application can also establish a connection between bone information and pixel feature information, construct a bone pixel correlation coefficient matrix, and use the matrix, the preset correlation coefficient threshold and the second pixel feature information to determine the first pixel feature information. The preset correlation coefficient threshold is a threshold set by the user according to the actual usage scenario, and the present application does not make specific restrictions on this.
[0065] Furthermore, the above-mentioned determining the corresponding bone pixel correlation coefficient matrix based on each first three-dimensional bone coordinate sequence and each second pixel feature information specifically includes:
[0066] Based on the first three-dimensional skeleton coordinate sequence and the second pixel feature information corresponding to the same frame, the dynamic skeleton pixel correlation coefficient is determined, and a dynamic skeleton pixel correlation coefficient set corresponding to each video frame is constructed. The dynamic skeleton pixel correlation coefficient is calculated based on the distance between the skeleton mapping coordinate point and the corresponding grid center. The dynamic skeleton pixel correlation coefficient set is matched with a preset standard correlation coefficient library to determine the skeleton pixel correlation coefficient matrix based on the matching results. The preset standard correlation coefficient set is constructed based on the unobstructed face recognition video stream of the pre-recorded object.
[0067] In other words, the present application calculates the distance between the skeleton mapping coordinate point and the grid center for the first three-dimensional skeleton coordinate sequence and the second pixel feature information corresponding to the same frame in the video stream. The distance can be the Euclidean distance. For example, the Euclidean distance between the skeleton mapping coordinate point P (x1, y1, z1) in the first three-dimensional skeleton coordinate sequence and a grid center Q (x2, y2, z2) in the first pixel feature information is This Euclidean distance is used as the dynamic bone pixel correlation coefficient of P(x1, y1, z1) and Q(x2, y2, z2). By calculating the dynamic bone pixel correlation coefficients between all bone points and all grid centers, a dynamic bone pixel correlation coefficient set is constructed. Then, the present application matches the constructed dynamic bone pixel correlation coefficient set with a preset standard correlation coefficient library. The preset standard correlation coefficient library is constructed based on the unobstructed face recognition video stream of the pre-entered object. When constructing this library, the correlation coefficients between the bones and pixels in the video stream under the unobstructed condition are also calculated. These correlation coefficients reflect the association pattern of facial bones and pixels under the normal (unobstructed) state. This library can be regarded as a reference standard for measuring the association between bones and pixels in the current video stream to be identified. When the two are matched, they can be compared for similarity, such as calculating the similarity metric (such as cosine similarity, etc.) between the two coefficient sets to find the closest matching result. Through this matching, the difference between the association between bones and pixels in the current video frame and the standard situation can be determined. If, during the matching process, the correlation coefficient between a particular bone point and a pixel grid area is found to be very close to the corresponding correlation coefficient in the standard library, such as the difference is less than a first set value, then the corresponding position in the matrix is assigned an appropriate value (e.g., a larger value indicates a strong correlation), which can be set according to the actual use scenario; conversely, if the difference is large, a smaller value is assigned, which can be set according to the actual use scenario. In this way, by matching and processing all video frames, a bone pixel correlation coefficient matrix is constructed that can fully reflect the dynamic correlation between bones and pixels in the entire video stream.
[0068] Furthermore, the above-mentioned determination of the first pixel feature information based on the bone pixel correlation coefficient matrix, the preset correlation coefficient threshold and the second pixel feature information specifically includes:
[0069] The skeleton pixel correlation coefficient matrix is traversed to select correlation coefficients greater than a preset correlation coefficient threshold as the screening correlation coefficients. The pixel corresponding to the screening correlation coefficient in the second pixel feature information is determined as the first pixel, and the first pixel feature information is determined based on the first pixel obtained by screening.
[0070] In other words, the present application filters the correlation coefficients in the bone pixel correlation coefficient matrix by presetting the correlation coefficient threshold value to find the correlation coefficients with a high degree of correlation between bones and pixels. According to the correspondence between the elements in the bone pixel correlation coefficient matrix and the pixels in the second pixel feature information, the first pixel corresponding to the filtered correlation coefficient is determined, and then the feature vector corresponding to the first pixel can be extracted from the second pixel feature information based on these filtered first pixel points, thereby determining the first pixel feature information.
[0071] Through the above scheme, the features of the face can be reflected more accurately, which helps to improve the accuracy and effect of face recognition. In particular, the first pixel feature information can play an important role when processing facial feature analysis related to bones and dealing with complex scenes (such as occlusion).
[0072] S102: The server determines a corresponding second three-dimensional bone coordinate sequence based on the first three-dimensional bone coordinate sequence, the first pixel feature information, and a preset adaptive distillation model.
[0073] The second three-dimensional skeleton coordinate sequence is a three-dimensional skeleton coordinate sequence after completing the occluded face of the object to be identified.
[0074] In the embodiment of the present application, determining the corresponding second three-dimensional bone coordinate sequence based on the first three-dimensional bone coordinate sequence, the first pixel feature information, and the preset adaptive distillation model specifically includes:
[0075] The first three-dimensional skeletal coordinate sequence and the first pixel feature information are input into a preset adaptive distillation model to match the initial complete skeletal coordinate sequence in the standard face template library through a pre-trained teacher model. The pre-processed first pixel feature information and the first three-dimensional skeletal coordinate sequence are convolved by the student model to determine the pending three-dimensional skeletal coordinate sequence. The student model is trained based on a number of three-dimensional skeletal coordinate sequence samples of occluded faces. The knowledge distillation loss function value is calculated based on the initial complete skeletal coordinate sequence and the pending three-dimensional skeletal coordinate sequence, so that when the knowledge distillation loss function value is less than a preset threshold, the pending three-dimensional skeletal coordinate sequence is used as the second three-dimensional skeletal coordinate sequence.
[0076] That is to say, the present application uses a preset adaptive distillation model including a teacher model and a student model. First, a pre-trained teacher model, such as a convolutional neural network model (CNN), uses the input data to match the initial complete bone coordinate sequence in the standard face template library. Among them, the standard face template library stores a variety of standard and complete facial bone coordinate sequence information, which can be used as a reference standard. The teacher model finds the most matching initial complete bone coordinate sequence by comparing the input bone coordinate sequence with the sequence in the template library. The purpose of this step is to provide a more accurate reference framework for the subsequent bone coordinate sequence generation, and to use standard bone information to guide the learning and generation process of the model.
[0077] The present application can also preprocess the input first pixel feature information and the first three-dimensional bone coordinate sequence (e.g., normalization, feature enhancement, etc.), and then input the preprocessed data into the student model. The student model can perform convolution processing on the input data based on technologies such as convolutional neural networks (CNN). Convolution processing can automatically extract feature information from the data. Through the operation of multiple convolutional layers, the complex relationship between facial bones and pixel features can be learned.
[0078] After processing by the student model, it outputs a sequence of undetermined 3D skeletal coordinates. This sequence is the possible 3D skeletal coordinate sequence predicted by the student model based on the input data, but it is not necessarily the final accurate result at this point. Because the student model is trained on a number of 3D skeletal coordinate sequence samples of occluded faces, it has a certain ability to handle skeletal information in occluded situations and can attempt to recover the skeletal coordinates of the occluded parts.
[0079] The present application also calculates the value of the knowledge distillation loss function based on the initial complete bone coordinate sequence obtained by matching the teacher model and the pending three-dimensional bone coordinate sequence output by the student model. The knowledge distillation loss function is used to measure the degree of difference between the pending sequence output by the student model and the reference sequence provided by the teacher model. The knowledge distillation loss function can use the mean square error to calculate the average value of the sum of squares of the differences between the corresponding coordinate points of the two sequences to quantify the degree of difference. At the same time, the present application sets a preset threshold value, which is compared with the knowledge distillation loss function value to determine whether the accuracy of the pending three-dimensional bone coordinate sequence meets the requirements. When the calculated knowledge distillation loss function value is less than the preset threshold value, it means that the pending sequence output by the student model is close enough to the reference sequence of the teacher model. At this time, the pending three-dimensional bone coordinate sequence can be used as the second three-dimensional bone coordinate sequence. The preset threshold value can be set by the user according to the actual usage scenario, and the present application does not make specific restrictions on this.
[0080] The aforementioned second 3D skeletal coordinate sequence further refines and supplements the initial first 3D skeletal coordinate sequence. In particular, for occluded facial portions, the model's learning and processing yields more complete skeletal coordinate information, providing more accurate skeletal feature data for subsequent face recognition and other tasks. Utilizing a preset adaptive distillation model, combined with the collaboration of the teacher and student models, starting from the first 3D skeletal coordinate sequence and the first pixel feature information, the corresponding second 3D skeletal coordinate sequence is ultimately determined through operations such as matching, convolution processing, loss calculation, and judgment, improving the accuracy and completeness of the skeletal coordinate information.
[0081] S103, the server calculates the coupling value corresponding to each selected pixel point in the preset area according to the preset coupling calculation formula, each first three-dimensional bone coordinate sequence and each first pixel point feature information, so as to determine the corresponding multimodal feature aggregation amount based on the weighted calculation results of each coupling value, the first pixel point feature information and the second three-dimensional bone coordinate sequence.
[0082] The above coupling value is used to characterize the degree of dynamic correlation between bones and pixels.
[0083] In the embodiment of the present application, the above calculation of the coupling value corresponding to each selected pixel point in the preset area according to the preset coupling calculation formula, each first three-dimensional bone coordinate sequence and each first pixel feature information specifically includes:
[0084] The coordinate displacement variance of each skeletal point is calculated based on each first three-dimensional skeletal coordinate sequence, and the inverse of the coordinate displacement variance of the skeletal point is used as the temporal stability value of the skeletal point. The texture entropy of the pixel point is calculated based on the feature information of each first pixel point. The Euclidean distance value between each skeletal point and each selected pixel point is determined based on each first three-dimensional skeletal coordinate sequence and the feature information of the first pixel point. The correlation coefficient of the selected pixel point is greater than a preset correlation coefficient threshold. The temporal stability value of the skeletal point, the texture entropy of the pixel point, and the Euclidean distance value are input into a preset coupling degree calculation formula to calculate the coupling degree value of the selected pixel point.
[0085] That is, the present application uses the first three-dimensional bone coordinate sequence and the first pixel feature information, as well as a preset coupling degree calculation formula, to calculate the coupling degree value that characterizes the degree of dynamic association between the bone and the pixel. The preset coupling degree calculation formula is as follows:
[0086]
[0087] in, Represents the pixel point p in the i-th row and j-th column of the preset area at time t ij (i.e. selected pixel point) and the kth bone point b k The coupling value between them; Represents the pixel point p at time t ij The corresponding pixel texture entropy; Represents the bone point b at time t k The corresponding skeletal point temporal stability value; Represents the pixel point p at time t ij , that is, the coordinates of the selected pixel point and the bone point b k The Euclidean distance value of the coordinates on the two-dimensional plane, Represents pixel p ij The coordinates projected onto the two-dimensional plane, Represents the bone point b kThe coordinates projected onto the two-dimensional plane; ∈ is a preset positive value to avoid the denominator being zero and ensure the stability of the formula; is the time decay term, t is the current moment; t ref is a preset reference moment, which can be set by the user according to the actual usage scenario and is not specifically limited here; τ is a preset time attenuation coefficient, which can be set by the user according to the actual usage scenario and is not specifically limited here, and is used to control the attenuation rate of the coupling value over time.
[0088] The above skeleton point temporal stability value The calculation formula is as follows:
[0089] First calculate the bone point b k The variance of the bone point coordinate displacement at n consecutive moments: Then pass Calculate the temporal stability value of the skeleton point. (σ x ,σ y ,σ z ) is the bone point b k The coordinates (x k ,y k ,z k ) The variance of the bone point coordinate displacement at n consecutive moments, is the mean of the coordinates at n consecutive moments.
[0090] The above pixel texture entropy The calculation formula is as follows:
[0091]
[0092] in, is the grayscale value of pixel (i, j) at time t The occurrence probability value, g, represents all possible grayscale values traversed, which can be set according to the actual usage scenario and is not specifically limited here.
[0093] In one embodiment of the present application, the weighted calculation result based on each coupling value, the first pixel feature information, and the second three-dimensional bone coordinate sequence is used to determine the corresponding multimodal feature aggregation amount, specifically including:
[0094] Based on the feature information of each first pixel point and a preset occlusion ratio prediction model, the facial occlusion ratio value of the object to be identified is determined, and the pixel feature weight and the bone feature weight are determined based on the facial occlusion ratio value. The sum of the pixel feature weight and the bone feature weight is 1. The feature information of each selected pixel point in the first pixel feature information is vector-encoded to obtain a first feature vector, and the feature attributes corresponding to each bone point in the second three-dimensional bone coordinate sequence are vector-encoded to determine a second feature vector. The feature attributes corresponding to the bone point include at least position code, motion speed, and posture. Based on the pixel feature weight and the bone feature weight, a weighted sum of the first feature vector and the second feature vector of the kth bone point is calculated, and the weighted sum is used as the kth first component. k is the total number of selected pixels, and k is a natural number. The product of the kth first component and the coupling value of the kth bone point is calculated as the kth second component, and the sum of the k second components is calculated to obtain the multimodal feature aggregation amount of the corresponding selected pixel point.
[0095] In other words, for each frame of RGB image, the present application can identify the first pixel feature information and the facial occlusion ratio value of the object to be identified through a pre-trained lightweight preset occlusion ratio prediction model, which can be a neural network model, a machine learning model, etc. At the same time, the server pre-stores the correspondence between different facial occlusion ratio values and different feature weight groups. Different feature weight groups contain different pixel feature weights and bone feature weights. The correspondence can be set according to the actual usage scenario and is not specifically limited here. Among them, the bone feature weight and the facial occlusion ratio value are inversely proportional, that is, the more the face is occluded, the greater the bone feature weight. For example, if the facial occlusion ratio value is 0.6, the bone feature weight may be set to 0.6 and the pixel feature weight may be set to 0.4.
[0096] Subsequently, the present application performs vector encoding on the feature information of each selected pixel point in the first pixel feature information, that is, mapping the features of the pixel point into a high-dimensional vector space so that each vector can accurately represent the feature information of the pixel point; at the same time, for each bone point in the second three-dimensional bone coordinate sequence, its corresponding feature attributes (including at least position coding, movement speed and posture) are vector encoded. Then, the first feature vector and the second feature vector are obtained, and based on the pixel feature weight and bone feature weight obtained above, the weighted sum value c of the first feature vector and the second feature vector of the kth bone point is calculated. k , and then the second component c is obtained according to the above coupling value k ×C k , C k is the coupling value of the kth skeleton point, C k is with c k At the same frame and at the same moment, all C ij,kThen, the server calculates the sum of the second components of the k skeleton points, and then the multimodal feature aggregation of the k-th selected pixel point (i, j) can be obtained. The calculation formula of the multimodal feature aggregation is as follows:
[0097] Through the above scheme, pixel features, bone features and the coupling relationship between them are comprehensively considered, multimodal information is effectively aggregated, and more representative feature quantities are obtained, providing richer and more accurate information for subsequent tasks such as face recognition.
[0098] S104, the server determines the bone pixel feature vector based on the aggregation amount of each multimodal feature, the second three-dimensional bone coordinate sequence and the preset attention mechanism, and determines the face recognition result of the object to be identified according to the bone pixel feature vector and the preset face recognition database.
[0099] In the embodiment of the present application, the above-mentioned determination of the bone pixel feature vector based on the aggregation amount of each multimodal feature, the second three-dimensional bone coordinate sequence and the preset attention mechanism specifically includes:
[0100] The multimodal feature aggregations are concatenated with the second 3D bone coordinate sequence to produce a concatenated vector. This concatenated vector is then fed into a pre-defined attention mechanism to map the query vector, key vector, and value vector. The attention mechanism output corresponding to each attention head is then calculated. The outputs of each attention mechanism are concatenated and linearly transformed to produce a bone pixel feature vector.
[0101] That is to say, the present application can also splice each multimodal feature aggregation amount with the second three-dimensional bone coordinate sequence, that is, connect them in a certain order to form a new vector. If the multimodal feature aggregation amount is an m-dimensional vector and the second three-dimensional bone coordinate sequence is represented as an n-dimensional vector, then the splicing vector is an (m+n)-dimensional vector. Then, the input preset attention mechanism is linearly transformed and mapped to the query vector Q, key vector K and value vector V. Multi-head attention is performed in parallel through the preset attention mechanism to calculate the output value of each attention mechanism. The preset attention mechanism can be pre-trained by several splicing vectors. Then, the server splices the output value to obtain the output value splicing vector and further linearly transforms it to obtain the bone pixel feature vector. Further linear transformation can be achieved by multiplying the preset weight matrix with the output value splicing vector and adding the bias term. The preset weight matrix and bias term can be set according to the actual use scenario and are not specifically limited here.
[0102] Furthermore, based on the skeleton pixel feature vector and the preset face recognition database, the face recognition result of the object to be recognized is determined, specifically including:
[0103] Based on a preset Bayesian decision rule, the posterior probabilities between the skeleton pixel feature vectors and the identity information entered in the preset face recognition database are calculated. The maximum posterior probability value among these posterior probabilities is determined, and the identity corresponding to the maximum posterior probability value is used as the face recognition result for the object to be identified.
[0104] In other words, the server can be pre-set with a preset face recognition database, which can store the identity entry information of the pre-entered objects. This application uses the Bayesian decision rule to take the bone pixel feature vector as the observed feature vector and the identity entry information as the sample category. It is assumed that the feature vector obeys a preset probability distribution such as a Gaussian distribution. The probability distribution can be selected by the user according to the actual scenario and is not specifically limited here. Then, the class conditional probability density is calculated based on the feature vector data of the identity entry information; then, the prior probability is obtained based on the ratio of the number of identity entry information in the preset face recognition database; then, the posterior probability corresponding to the bone pixel feature vector and different identity entry information is calculated according to the Bayesian formula. By finding the maximum value from each posterior probability, the identity identifier corresponding to the corresponding identity entry information can be determined, and it is used as the face recognition result of the object to be identified.
[0105] The above-mentioned method adopts a preset Bayesian decision rule, which provides a scientific basis for recognition decision-making from a probability perspective by calculating the posterior probability of the bone pixel feature vector and the information entered for each identity in the preset face recognition database. The identity identification corresponding to the maximum posterior probability value is used as the recognition result, which enhances the reliability and scientific nature of the recognition result.
[0106] The above scheme uses a pre-set adaptive distillation model and the collaboration of the teacher model and student model to complete the second three-dimensional bone coordinate sequence of the occluded face based on the first three-dimensional bone coordinate sequence and the first pixel feature information. This provides complete bone feature information for accurate recognition, effectively overcomes the interference of occlusion on recognition, improves recognition accuracy in complex occlusion scenes, and can cope with different occlusion levels and scenes, enhancing the generalization ability of face recognition. At the same time, the coupling value is calculated to characterize the dynamic relationship between bones and pixels, fully exploring the intrinsic connection between multimodal data and calculating the multimodal feature aggregation amount, making the feature expression more comprehensive and enhancing the accuracy of recognition. It does not require human intervention and reduces labor costs. This solves the problems of low face recognition accuracy, poor stability, weak generalization ability and high labor costs in complex occlusion scenes.
[0107] Figure 2 A structural diagram of a face recognition system based on a skeleton calibration algorithm provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, the face recognition system 200 based on the skeleton calibration algorithm includes:
[0108] The first determination module 201 is configured to determine, based on a face recognition video stream from a preset facial recognition device group, each first three-dimensional skeletal coordinate sequence and each first pixel feature information of the object to be identified within a preset time period. The second determination module 202 is configured to determine a corresponding second three-dimensional skeletal coordinate sequence based on the first three-dimensional skeletal coordinate sequence, the first pixel feature information, and a preset adaptive distillation model. The second three-dimensional skeletal coordinate sequence is a three-dimensional skeletal coordinate sequence after completing the occluded face of the object to be identified. The calculation module 203 is configured to calculate the coupling value corresponding to each selected pixel within a preset area based on a preset coupling calculation formula, each first three-dimensional skeletal coordinate sequence, and each first pixel feature information. The calculation module 203 then determines the corresponding multimodal feature aggregation value based on a weighted calculation result of each coupling value, the first pixel feature information, and the second three-dimensional skeletal coordinate sequence. The coupling value is used to characterize the degree of dynamic correlation between the skeleton and the pixel. The third determination module 204 is configured to determine a skeletal pixel feature vector based on each multimodal feature aggregation value, the second three-dimensional skeletal coordinate sequence, and a preset attention mechanism. The skeletal pixel feature vector and a preset face recognition database are then used to determine the face recognition result for the object to be identified.
[0109] The various embodiments in this application are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences from other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment.
[0110] The system and method provided in the embodiments of the present application correspond one to one. Therefore, the system also has similar beneficial technical effects to its corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the system will not be repeated here.
[0111] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0112] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A face recognition method based on a skeleton calibration algorithm, characterized in that: The method comprises: Determine, based on a face recognition video stream from a preset face recognition device group, each first three-dimensional bone coordinate sequence and each first pixel feature information of an object to be recognized within a preset time; Determine a corresponding second three-dimensional skeleton coordinate sequence based on the first three-dimensional skeleton coordinate sequence, the first pixel feature information, and a preset adaptive distillation model; wherein the second three-dimensional skeleton coordinate sequence is a three-dimensional skeleton coordinate sequence after completing the occluded face of the object to be identified; Calculating a coupling value corresponding to each selected pixel point in a preset area based on a preset coupling calculation formula, each of the first three-dimensional bone coordinate sequences, and each of the first pixel feature information, so as to determine a corresponding multimodal feature aggregation amount based on a weighted calculation result of each of the coupling values, the first pixel feature information, and the second three-dimensional bone coordinate sequence; the coupling value is used to characterize the degree of dynamic association between the bone and the pixel point; Based on the aggregation amount of each multimodal feature, the second three-dimensional bone coordinate sequence and the preset attention mechanism, the bone pixel feature vector is determined, and the face recognition result of the object to be identified is determined according to the bone pixel feature vector and the preset face recognition database.
2. A face recognition method based on skeleton calibration algorithm according to claim 1, characterized in that: Determining, based on a face recognition video stream from a preset face recognition device group, each first three-dimensional bone coordinate sequence and each first pixel feature information of an object to be recognized within a preset time period, specifically includes: Inputting the face recognition video stream from the preset facial recognition device group into a pre-trained skeleton calibration model to determine the coordinates of a plurality of skeleton approximate points corresponding to the object to be recognized in the face recognition video stream according to the model output result; Mapping the coordinates of each of the skeleton approximate points to a preset depth map coordinate system, and sequentially adding the skeleton mapping coordinate points to each of the first three-dimensional skeleton coordinate sequences in chronological order; Determining RGB images of each face in the face recognition video stream; wherein each frame of the face recognition video stream includes an RGB image and a depth image; Dividing each of the facial RGB images into a grid according to a preset grid to determine a feature vector corresponding to each of the facial RGB images as second pixel feature information; According to each of the first three-dimensional bone coordinate sequences and each of the second pixel feature information, the corresponding bone pixel correlation coefficient matrix is determined, so as to determine the first pixel feature information according to the bone pixel correlation coefficient matrix, the preset correlation coefficient threshold and the second pixel feature information.
3. A face recognition method based on skeleton calibration algorithm according to claim 2, characterized in that: Determining a corresponding bone pixel correlation coefficient matrix according to each of the first three-dimensional bone coordinate sequences and each of the second pixel feature information specifically includes: Determine a dynamic bone pixel correlation coefficient based on the first three-dimensional bone coordinate sequence and the second pixel feature information corresponding to the same frame, and construct a dynamic bone pixel correlation coefficient set corresponding to each video frame; the dynamic bone pixel correlation coefficient is calculated based on the distance between the bone mapping coordinate point and the corresponding grid center; The dynamic skeleton pixel association coefficient set is matched with a preset standard association coefficient library to determine the skeleton pixel association coefficient matrix according to the matching result; wherein the preset standard association coefficient set is constructed based on an unobstructed face recognition video stream of a pre-recorded object.
4. A face recognition method based on skeleton calibration algorithm according to claim 2, characterized in that: Determining the first pixel feature information according to the bone pixel correlation coefficient matrix, the preset correlation coefficient threshold, and the second pixel feature information specifically includes: Traversing the bone pixel correlation coefficient matrix, and screening correlation coefficients that are greater than the preset correlation coefficient threshold as screening correlation coefficients; The pixel corresponding to the filtered correlation coefficient in the second pixel feature information is determined as the first pixel, so as to determine the first pixel feature information according to the first pixel obtained by filtering.
5. A face recognition method based on skeleton calibration algorithm according to claim 1, characterized in that: Determining a corresponding second three-dimensional bone coordinate sequence according to the first three-dimensional bone coordinate sequence, the first pixel feature information, and a preset adaptive distillation model specifically includes: Inputting the first three-dimensional bone coordinate sequence and the first pixel feature information into the preset adaptive distillation model to match the initial complete bone coordinate sequence in the standard face template library through a pre-trained teacher model; Performing convolution processing on the preprocessed first pixel feature information and the first three-dimensional bone coordinate sequence using a student model to determine a pending three-dimensional bone coordinate sequence; the student model is trained based on a number of three-dimensional bone coordinate sequence samples of occluded faces; A knowledge distillation loss function value is calculated based on the initial complete bone coordinate sequence and the undetermined three-dimensional bone coordinate sequence, so that when the knowledge distillation loss function value is less than a preset threshold, the undetermined three-dimensional bone coordinate sequence is used as the second three-dimensional bone coordinate sequence.
6. A face recognition method based on skeleton calibration algorithm according to claim 1, characterized in that: Calculating the coupling value corresponding to each selected pixel point in the preset area according to a preset coupling calculation formula, each of the first three-dimensional bone coordinate sequences, and each of the first pixel feature information, specifically including: Calculating the displacement variance of each bone point coordinate according to each of the first three-dimensional bone coordinate sequences, using the inverse of the displacement variance of the bone point coordinate as the time series stability value of the bone point, and calculating the pixel point texture entropy according to each of the first pixel point feature information; Determining, based on each of the first three-dimensional bone coordinate sequences and the first pixel feature information, a Euclidean distance value between each bone point and each of the selected pixel points; wherein a correlation coefficient of the selected pixel points is greater than a preset correlation coefficient threshold; The temporal stability value of the skeleton point, the texture entropy of the pixel point and the Euclidean distance value are input into the preset coupling degree calculation formula to calculate the coupling degree value of the selected pixel point.
7. A face recognition method based on skeleton calibration algorithm according to claim 1, characterized in that: Determining a corresponding multimodal feature aggregation amount based on a weighted calculation result of each of the coupling values, the first pixel feature information, and the second three-dimensional bone coordinate sequence specifically includes: Determining a facial occlusion ratio value of the object to be identified based on the feature information of each of the first pixels and a preset occlusion ratio prediction model, and determining a pixel feature weight and a bone feature weight based on the facial occlusion ratio value; the sum of the pixel feature weight and the bone feature weight is 1; Performing vector encoding on the feature information of each selected pixel point in the first pixel feature information to obtain a first feature vector, and performing vector encoding on the feature attributes corresponding to each bone point in the second three-dimensional bone coordinate sequence to determine a second feature vector; wherein the feature attributes corresponding to the bone point include at least position code, motion speed, and posture; Calculate a weighted sum of the first eigenvector and the second eigenvector of the k-th skeleton point according to the pixel feature weight and the skeleton feature weight, and use the weighted sum as the k-th first component; k is the total number of the selected pixels, and k is a natural number; The product value of the kth first component and the coupling value of the kth skeleton point is calculated as the kth second component, and the sum of the k second components is calculated to obtain the multimodal feature aggregation amount of the corresponding selected pixel point.
8. A face recognition method based on skeleton calibration algorithm according to claim 1, characterized in that: Determining a bone pixel feature vector based on each of the multimodal feature aggregation amounts, the second three-dimensional bone coordinate sequence, and a preset attention mechanism specifically includes: Performing feature splicing on each of the multimodal feature aggregation amounts and the second three-dimensional bone coordinate sequence to obtain a splicing vector; Input the concatenated vector into the preset attention mechanism to map the query vector, key vector, and value vector, and calculate the attention mechanism output value corresponding to each attention head; The output values of each attention mechanism are concatenated and linearly transformed to obtain the bone pixel feature vector.
9. A face recognition method based on skeleton calibration algorithm according to claim 1, characterized in that: Determining a face recognition result of the object to be recognized based on the skeleton pixel feature vector and a preset face recognition database, specifically including: Calculating the posterior probability between the bone pixel feature vector and each identity entry information in the preset face recognition database according to a preset Bayesian decision rule; The maximum posterior probability value among the posterior probabilities is determined, and the identity identifier corresponding to the maximum posterior probability value is used as the face recognition result of the object to be identified.
10. A face recognition system based on a skeleton calibration algorithm, characterized in that: The system comprises: A first determination module is configured to determine, based on a face recognition video stream from a preset face recognition device group, each first three-dimensional bone coordinate sequence and each first pixel feature information of an object to be recognized within a preset time; A second determination module is configured to determine a corresponding second three-dimensional skeleton coordinate sequence based on the first three-dimensional skeleton coordinate sequence, the first pixel feature information, and a preset adaptive distillation model; wherein the second three-dimensional skeleton coordinate sequence is a three-dimensional skeleton coordinate sequence after completing the occluded face of the object to be identified; a calculation module, configured to calculate a coupling value corresponding to each selected pixel point in a preset area according to a preset coupling calculation formula, each of the first three-dimensional bone coordinate sequences, and each of the first pixel feature information, so as to determine a corresponding multimodal feature aggregation amount based on a weighted calculation result of each of the coupling values, the first pixel feature information, and the second three-dimensional bone coordinate sequence; the coupling value is used to represent the degree of dynamic association between the bone and the pixel point; The third determination module is used to determine the bone pixel feature vector based on each of the multimodal feature aggregation amounts, the second three-dimensional bone coordinate sequence and the preset attention mechanism, and determine the face recognition result of the object to be identified based on the bone pixel feature vector and the preset face recognition database.
Citation Information
Patent Citations
Action recognition method based on depth image and skeleton information
CN110263720A
Motion recognition method and device based on video image, equipment and storage medium
CN114511931A
Face blind restoration method based on three-dimensional decomposition
CN114862697A
Three-dimensional key point prediction method, training method and related equipment
CN115994944A
Visual target tracking
US20100197391A1