Face recognition detection method and device, computer equipment and program product
By obtaining occlusion measurement from 3D face point clouds and performing occlusion adaptive graph convolution operations, combined with face geometric prior models and feature manifolds for feature completion, the problem of low face recognition accuracy under occlusion is solved, achieving high-quality feature generation and improved recognition accuracy.
Patent Information
- Application Number
- CN202511646265.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-11-11
AI Technical Summary
Existing 3D face recognition technology lacks the ability to perceive occlusion during feature extraction, resulting in feature completion results that deviate from the actual physiological structure and low recognition accuracy.
By acquiring the occlusion metric of the 3D face point cloud, performing occlusion-adaptive graph convolution operations, and combining the face geometric prior model and feature manifold for feature completion, a high-quality complete feature map is generated.
It significantly improves the accuracy of face recognition in occluded scenarios, can handle various symmetrical and asymmetrical occlusions, and has stronger robustness and wide applicability to real-world scenarios.
Smart Images

Figure CN121096006A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of computer vision and three-dimensional data processing, and particularly relates to a face recognition detection method and device, a computer device and a program product. BACKGROUND
[0002] Three-dimensional face recognition technology is an important research direction in the field of computer vision. In actual application scenarios, the collected three-dimensional face point cloud data often becomes incomplete due to the occlusion of objects such as masks, hats, hair, glasses and the like, resulting in the loss of key geometric and texture features of the face, which poses a great challenge to existing three-dimensional face recognition methods.
[0003] On the one hand, traditional feature extraction methods, such as fixed graph convolutional networks, have difficulty in extracting effective information in the feature missing area when processing such incomplete data, resulting in low quality of the generated feature expression.
[0004] On the other hand, some existing technologies attempt to complete the missing data before recognition, such as using the natural symmetry of the face to repair the occluded area, i.e., using the unoccluded side of the face data to fill in the occluded side through mirroring operation. However, this method only has a certain effect on ideal symmetric occlusion, and is powerless for asymmetric occlusion, such as wearing single-sided glasses, head tilt or side face posture, etc. Even more noise may be introduced due to the wrong filling, which interferes with the recognition result. SUMMARY
[0005] The embodiments of the present application provide a face recognition detection method, device, computer device and program product, which can solve the technical problem of low face recognition accuracy in the prior art when processing occluded three-dimensional face point cloud, due to the lack of occlusion perception ability in the feature extraction process and the deviation of the feature completion result from the true physiological structure.
[0006] In a first aspect, the embodiments of the present application provide a face recognition detection method, comprising: obtaining a three-dimensional face point cloud each first data point corresponding occlusion metric ; based on the corresponding occlusion metric of each first data point , performing an occlusion-adaptive graph convolution operation on the three-dimensional face point cloud to generate an initial feature map ; wherein the initial feature map includes occluded points and non-occluded points; for the initial feature map The occluded points in the image are used to perform feature completion based on a preset face geometry prior model and a feature manifold composed of unoccluded points, generating a complete feature map. ; Based on the complete feature map Facial recognition detection is performed to determine the identity of the person.
[0007] Optionally, the acquisition of three-dimensional face point clouds First data points in each of the middle Corresponding occlusion metric ,include: The three-dimensional face point cloud Perform normalization processing; Normalized 3D face point cloud The input is fed into a preset occlusion prediction network, and the output is a normalized 3D face point cloud. First data points in each of the middle The corresponding occlusion probability, and the first data points The occlusion probability is used as the first data point Occlusion metric .
[0008] Optionally, the first data point Corresponding occlusion metric The three-dimensional face point cloud Perform occlusion-adaptive graph convolution operations to generate initial feature maps. ,include: For 3D face point cloud Each first data point in Based on this first data point Occlusion metric Dynamically adjust the first data point The neighborhood range used for information aggregation during graph convolution operations is used to obtain the first data point. dynamic neighborhood ;in, For dynamic neighborhood The second data point in; For each first data point Based on this first data point Occlusion metric The first data point Its dynamic neighborhood The second data point within The correlation metric parameters are used to calculate the first data point. With the second data point Information transfer weight between ; based on a dynamic neighborhood of each first data point and an information passing weight , a graph convolution operation is performed on each first data point , the features of the second data points in the dynamic neighborhood of the first data point are aggregated, and updated features of each first data point are obtained; updated features of all first data points are integrated to generate an initial feature map .
[0009] Optionally, for each first data point in the three-dimensional face point cloud , a neighborhood range of the first data point for information aggregation in the graph convolution operation process is dynamically adjusted based on an occlusion metric of the first data point , to obtain a dynamic neighborhood of the first data point , including: for each first data point , a neighborhood radius of the first data point in the graph convolution operation is determined based on an occlusion metric of the first data point through a preset neighborhood radius calculation formula; wherein the neighborhood radius calculation formula is , is a preset basic neighborhood radius, is used to limit the maximum contribution of the occlusion metric to the neighborhood radius , so as to avoid excessive increase of the neighborhood radius due to the occlusion metric being too large; an initial spatial neighborhood is determined in the three-dimensional face point cloud with the first data point as the center and the neighborhood radius as the radius; based on the geometric features of the first data point and each second data point in the initial spatial neighborhood, geometric consistency filtering is performed on the initial spatial neighborhood, including: curvature filtering: the curvature of the first data point and the curvature of the second data point are calculated , retain satisfaction The second data point ; Normal filtering: Calculate the first data point unit normal Second data point unit normal , retain satisfaction The second data point with an angle greater than the preset value ; If the first data point Occlusion metric If the occlusion metric is greater than the first preset occlusion metric, then in the initial spatial neighborhood after geometric consistency filtering, the occlusion metric will be... The second data point is less than or equal to the second preset occlusion metric. As the first data point dynamic neighborhood ; If the first data point Occlusion metric If the occlusion metric is less than or equal to the first preset occlusion metric, then the initial spatial neighborhood after geometric consistency filtering is directly used as the first data point. dynamic neighborhood .
[0010] Optionally, the association metric parameter includes geometric distance. Feature similarity; For each first data point Based on this first data point Occlusion metric The first data point Its dynamic neighborhood The second data point within The correlation metric parameters are used to calculate the first data point. With the second data point Information transfer weight between ,include: Based on the geometric distance and the feature similarity, combined with the first data point Second data point The occlusion relationship is determined by the occlusion relationship factor. The first data point was calculated. With the second data point Information transfer weight between .
[0011] Optionally, the training process of the face geometric prior model includes: Collecting a plurality of unoccluded three-dimensional face data to form a training sample set Obtaining individual facial difference features and expression deformation features from the training sample set Performing statistical analysis on the training sample set to calculate an average face model Extracting the individual facial difference features by principal component analysis to obtain identity basis vectors Extracting the expression deformation features by principal component analysis to obtain expression basis vectors Based on the above average face model , identity basis vectors and expression basis vectors , a function representing the shape of the face is constructed: ; wherein, is the shape of the face, is the identity parameter, is the expression parameter; The L2 norm constraint is applied to the expression parameter ; The coordinate error between the face shape and the corresponding real face shape in the training sample set is used as the optimization objective to construct a loss function; Minimizing the loss function obtains the face geometry prior model.
[0012] Optionally, for the occluded points in the initial feature map , the feature of the occluded points is completed in combination with a preset face geometry prior model and a feature manifold composed of non-occluded points to generate a complete feature map , including: Obtaining the features of the non-occluded points in the initial feature map to form a non-occluded point feature set; wherein the non-occluded points are first data points satisfying a preset non-occluded condition; Based on the non-occluded point feature set, a low-dimensional feature manifold is constructed by Isomap algorithm; Based on the feature manifold and the real face coordinates of the non-occluded points, a multilayer perceptron is trained as a mapping function; Based on the corresponding occlusion metric of each first data point , a label of whether each first data point is an occluded point is generated to generate an initial occlusion mask; The initial occlusion mask is optimized by removing occlusion points at the edge of the initial occlusion mask that have similar features to the non-occlusion points in the local spatial neighborhood of the occlusion point at that edge, thus obtaining the optimized occlusion mask. For each occlusion point in the optimized occlusion mask, select the non-occlusion points in the local spatial neighborhood of the occlusion point. Through the trained mapping function, map the features of the non-occlusion points in the local spatial neighborhood to the parameters of the face geometry prior model. Then, take the average value of the parameters based on the number of non-occlusion points in the local spatial neighborhood to obtain the parameter estimate of the face geometry prior model. The parameter estimates are input into the face geometry prior model to generate physical constraint features for each occluded point in the optimized occlusion mask. Simultaneously, based on the constructed feature manifold, the geodesic distance between the occluded point and the non-occluded points in the local spatial neighborhood of the occluded point is calculated on the manifold. The attention weight is calculated based on the geodesic distance. The features of the non-occluded points in the local spatial neighborhood of the occluded point are weighted and aggregated through the attention weight to generate the context features of the occluded point. Based on the physical constraint features and context features of each occlusion point, the completion features of the occlusion point are obtained through weighted fusion; The completed features of all occluded points are merged with the features of unoccluded points to generate a complete feature map. .
[0013] Secondly, embodiments of this application provide a face recognition detection device, comprising: The acquisition module is used to acquire 3D face point clouds. First data points in each of the middle Corresponding occlusion metric ; The graph convolution module is used to perform graph convolution based on each first data point. Corresponding occlusion metric The three-dimensional face point cloud Perform occlusion-adaptive graph convolution operations to generate initial feature maps. ; wherein, the initial feature map Including occluded points and unoccluded points; The feature completion module is used to complete the initial feature map. The occluded points in the image are used to perform feature completion based on a preset face geometry prior model and a feature manifold composed of unoccluded points, generating a complete feature map. ; The identification and detection module is used to identify and detect based on the complete feature map. Facial recognition detection is performed to determine the identity of the person.
[0014] In a third aspect, an embodiment of the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and the processor implements the face recognition detection method of any one of the first aspect when executing the computer program.
[0015] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program implements the face recognition detection method of any one of the first aspect when executed by a processor.
[0016] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when running on a computer device, causes the computer device to execute the face recognition detection method of any one of the first aspect.
[0017] Compared with the prior art, the embodiment of the present application has the following beneficial effects: By combining the preset face geometric prior model and the feature manifold composed of non-occluded points, the feature of the occluded point is completed, compared with the prior art which only relies on simple symmetry, the completed feature generated by the present application not only conforms to the physiological law of the face, but also is highly consistent with the features of the visible area of the individual, thereby generating high-quality complete features, and significantly improving the face recognition accuracy in the occlusion scene.
[0018] Further, in the feature extraction stage, a processing mechanism for occlusion perception is introduced, and by dynamically adjusting the neighborhood and information transmission weight, the feature extraction process can actively adapt to the occlusion, aggregate more information from reliable visible areas for high-occluded points, and improve the robustness of the initial features, thereby laying a solid foundation for subsequent high-quality completion.
[0019] Further, the completion mechanism of the present application does not rely on the symmetry assumption, and therefore can effectively handle various symmetric and asymmetric occlusions, such as mask, single-sided hair occlusion, large-angle pose, etc., and has stronger robustness and wider real scene applicability.
[0020] It can be understood that the beneficial effects of the above-mentioned second aspect to fifth aspect can be referred to the related description in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0022] Figure 1 A flow chart of a face recognition detection method provided by an embodiment of the present application; Figure 2 A structural schematic diagram of a face recognition detection system provided by an embodiment of the present application; Figure 3 A schematic diagram of an occlusion adaptive graph convolution operation provided by an embodiment of the present application; Figure 4 A schematic diagram of a feature completion working principle provided by an embodiment of the present application; Figure 5 A schematic diagram of data interaction between modules in a system provided by an embodiment of the present application; Figure 6 A structural schematic diagram of a computer device provided by an embodiment of the present application.
[0023] Legend: 10 - acquisition module; 20 - graph convolution module; 30 - feature completion module; 40 - recognition module; S100 - occlusion metric acquisition step; S200 - occlusion-aware feature extraction step; S300 - feature completion step; S400 - face recognition detection step; 301 - center point; 303 - dynamic neighborhood; 304 - occlusion point; 305 - non-occlusion point; 306 - information transmission weight; 401 - occlusion point; 402 - non-occlusion point; 403 - face geometry prior model; 404 - feature manifold; 405 - physical constraint feature; 406 - context feature; 407 - completed feature; 408 - fusion; 501 - three-dimensional face point cloud; 502 - occlusion metric; 503 - initial feature map; 504 - complete feature map; 505 - face identity. DETAILED DESCRIPTION
[0024] In the following description, specific details are set forth, such as particular system configurations, techniques, etc., in order to provide a thorough understanding of the present application. However, persons skilled in the art will understand that the present application can be practiced without these specific details. In other instances, well-known structures, devices, circuits, and methods have not been described in detail in order to avoid obscuring the present application.
[0025] It should be understood that, when used in the specification and the appended claims, the term "comprises" indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0026] It should also be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items, and that the term “at least one of’ denotes one, or a combination of two or more items.
[0027] As used in the description of the application and the appended claims, the term “if’ can be interpreted to mean “when” or “upon” or “in response to determining” or “in response to detecting” depending on the context. Similarly, the phrase “if it is determined” or “if [a described condition or event] is detected” can be interpreted to mean “upon determining” or “in response to determining” or “upon [the described condition or event] being detected” or “in response to [the described condition or event] being detected,” depending on the context.
[0028] In addition, the terms “first,” “second,” “third,” etc. as used in the description of embodiments herein and throughout the claims, are used for differentiation only and are not meant to signify relative importance or order of magnitude.
[0029] Reference to “one embodiment” or “some embodiments” or “one implementation” or “some implementations” etc. in the present description means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrases “in one embodiment” or “in some embodiments” or “in other embodiments” or “in still other embodiments” or the like in various places in the specification are not necessarily all referring to the same embodiment, although they can. The terms “comprising,” “including,” “having” and the like are meant to be interpreted open-ended, unless otherwise specifically noted. They include the case when the specified feature, structure, or characteristic is included but not when it is not.
[0030] Embodiment 1 This embodiment elaborates a complete flow and system implementation of a face recognition detection method. Referring to Figure 1 , the method mainly includes an occlusion metric acquisition step S100, an occlusion-aware feature extraction step S200, a feature completion step S300, and a face recognition detection step S400. Correspondingly, referring to Figure 2 and Figure 5 , a system implementing the method can include an acquisition module 10, a graph convolution module 20, a feature completion module 30, and a recognition module 40. The modules work together to process a complete flow from an input three-dimensional face point cloud 501 to an output final face identity 505.
[0031] In the step of obtaining the occlusion metric S100, the system receives a set of first data points with three-dimensional coordinates, i.e., a three-dimensional face point cloud 501, which represents the geometric shape of the face. In order to eliminate the scale difference caused by the acquisition device, distance and other factors, the obtaining module 10 first performs normalization processing on the input three-dimensional face point cloud 501. As a specific implementation manner, the three-dimensional coordinates of each point can be scaled based on the maximum coordinate value of the point cloud in the vertical direction (usually the Y axis). For example, for any one first data point in the point cloud , the normalized coordinate can be calculated as , wherein is the maximum value of the Y coordinates of all points in the entire point cloud. It can be understood that this processing ensures the scale consistency of subsequent network processing.
[0032] After normalization, the obtaining module 10 feeds the processed three-dimensional face point cloud into a pre-trained occlusion prediction network. The network can be a lightweight convolutional neural network, for example, stacked by a plurality of 1x1 convolutional layers and max pooling layers, and the design goal is to efficiently predict the occlusion probability of each first data point in the point cloud. The network outputs a scalar value between 0 and 1 for each first data point , which is the occlusion probability of the point. In an embodiment of the present application, the occlusion probability is directly used as the occlusion metric 502 (also referred to as the occlusion metric ) of the first data point . A higher occlusion metric (for example, close to 1) indicates that the point is likely to be located in the area occluded by a mask, hair or other objects, while a lower occlusion metric (for example, close to 0) indicates that the point is located in a clearly visible facial area. By calculating a continuous and fine-grained occlusion metric 502 for each point, the system can accurately quantify the information reliability of each position in the point cloud, thereby providing key guiding information for subsequent differentiated processing.
[0033] , wherein the specific structure of the occlusion prediction network can be: input layer (receiving normalized three-dimensional point cloud coordinates, dimension Nx3, N is the total number of first data points)→1x1 convolutional layer 1 (output dimension Nx64, activation function ReLU)→max pooling layer (pooling kernel size 2)→1x1 convolutional layer 2 (output dimension Nx12, activation function ReLU)→batch normalization layer→1x1 convolutional layer 3 (output dimension Nx1, activation function Sigmoid)→output layer (output the occlusion probability of each point). The network training adopts a cross-entropy loss function.
[0034] , wherein the occlusion metric is the quantized three-dimensional face point cloud first data point A scalar value representing the degree of occlusion by external objects (such as masks, hair, glasses, etc.), with a value range of [0, 1]. Among them, represents that the point is completely unoccluded, and the facial information is 100% reliable; represents that the point is completely occluded, and the facial information is completely unreliable; represents that the point is in a semi-occluded state, and the larger the value, the lower the information reliability. This measure provides a core basis for subsequent dynamic neighborhood adjustment, information transmission weight calculation, and occlusion mask generation.
[0035] Subsequently, the process enters the occlusion-aware feature extraction step S200 performed by the graph convolution module 20. This step aims to generate an initial feature map 503 based on the original three-dimensional face point cloud 501 and the occlusion measure 502 obtained in the previous step. Unlike traditional graph convolution networks that use fixed neighborhoods and the same weight calculation logic for all points, the graph convolution module 20 in this embodiment performs an occlusion-adaptive graph convolution operation, the core idea of which is to dynamically adjust the information aggregation method according to the occlusion measure of each point. Specifically, this operation includes two aspects: dynamically adjusting the neighborhood range and dynamically calculating the information transmission weight.
[0036] Please refer to Figure 3 , which illustrates the principle of occlusion-adaptive graph convolution operation. For each first data point in the three-dimensional face point cloud as the information aggregation center (i.e., the center point 301), the graph convolution module 20 first determines a dynamic neighborhood 303 (also referred to as a dynamic neighborhood ) for it. Specifically, for the center point 301, its neighborhood search radius is not a fixed value, but is dynamically calculated by a preset neighborhood radius calculation formula, which is: In this formula, is the dynamic neighborhood radius of the center point 301, is a preset basic neighborhood radius (for example, it can be set to 0.05 meters), which defines the basic receptive field size of non-occluded points, is the occlusion measure of the center point 301. It should be noted that part plays a role in limiting the maximum contribution of the occlusion measure to the neighborhood radius, so as to avoid excessive increase of the neighborhood radius due to abnormally large occlusion measure of individual points, thereby ensuring the stability of the calculation. The intuitive effect of this formula is that when the center point 301 is an occluded point 304 (i.e., the value is large), its neighborhood radius will increase accordingly, so that it can pass through the adjacent occluded area to connect and aggregate more distant non-occluded points 305 with more reliable information.
[0037] In the formula, After the initial spatial neighborhood is determined for the radius, to ensure that the points in the neighborhood belong to the same local surface and avoid aggregating irrelevant information from different areas of the face (such as the nose and the cheek), the graph convolution module 20 also performs geometric consistency screening on the initial spatial neighborhood. The screening includes curvature screening and normal screening, wherein the curvature screening requires that the difference between the curvature of the second data point in the neighborhood and the curvature of the center point 301 is less than a preset threshold; and the normal screening requires that the dot product of the unit normal vector of the second data point in the neighborhood and the unit normal vector of the center point 301 is greater than a preset angle value.
[0038] The preset angle value corresponds to the threshold of the dot product of the unit normal vectors, which is preferably 0.85. According to the characteristics of the three-dimensional face local surface, when the dot product of the unit normal vectors of two points is greater than or equal to 0.85, the normal angle between them is less than or equal to 31.79°, indicating that the two points belong to the same continuous surface (such as the same cheek or the same forehead region).
[0039] The threshold of the curvature screening is The curvature is The average curvature is calculated (the formula is , , The threshold ensures that the bending degree of the neighborhood point and the center point is consistent, and avoids including the bridge of the nose (high curvature) into the neighborhood of the cheek point (low curvature).
[0040] After the geometric consistency screening, the system performs the last step of neighborhood point selection according to the occlusion measure of the center point 301. If the occlusion measure of the center point 301 is greater than a first preset occlusion measure (for example, 0.6), indicating that it is a high-occlusion point, then the system selects only those second data points whose occlusion measure is less than or equal to a second preset occlusion measure (for example, 0.3) from the geometrically consistent neighborhood of the center point 301 as its final dynamic neighborhood 303. This can force the high-occlusion point to aggregate information only from reliable non-occluded points 305. Conversely, if the occlusion measure of the center point 301 is less than or equal to the first preset occlusion measure, indicating that it is a reliable non-occluded point, then all the neighborhood points that have passed the geometric consistency screening are directly taken as its dynamic neighborhood 303.
[0041] The first preset occlusion measure is 0.6, and the second preset occlusion measure is 0.3. These values are obtained by statistical analysis of multiple groups (such as 1000 groups) of three-dimensional face samples containing mask, hair, and glasses occlusions: when , the probability that the data point is actually occluded is more than 92%, and it is determined to be a high-occlusion point; and when , the probability that the data point is not occluded is more than 95%, and it is determined to be a reliable non-occluded point. The threshold setting can balance the accuracy and robustness of the occlusion judgment.
[0042] Among them, dynamic neighborhood 303 is for the first data point. A dynamically generated set of local points used for graph convolution information aggregation, distinct from the fixed-radius neighborhood of traditional graph convolution—its neighborhood range (neighborhood radius) (From the first data point) Occlusion metric The decision, and the second data point in the neighborhood. It needs to pass geometric consistency screening (curvature and normal matching) and occlusion reliability screening (non-occlusion or low occlusion). The core function of dynamic neighborhood is to allow high-occlusion points to aggregate more information from distant reliable non-occlusion points, and to allow non-occlusion points to aggregate information from locally geometrically consistent points, thus avoiding interference from invalid occlusion information.
[0043] This application employs a two-stage filtering method for dynamic neighborhood construction (first geometric consistency filtering, then occlusion reliability filtering), which offers significant advantages over existing technologies (which filter neighborhoods solely based on distance). Existing technologies tend to include occluded points on different surfaces (e.g., a forehead point obscured by hair and an unoccluded cheek point) in the same neighborhood, leading to information aggregation noise. In contrast, this application first excludes points on different surfaces through geometric consistency filtering, and then excludes highly occluded points through occlusion reliability filtering, ensuring that only reliable points on the same surface with low occlusion are retained within the neighborhood. This improves the robustness of the initial features extracted by graph convolution.
[0044] After determining the dynamic neighborhood 303, the graph convolution module 20 calculates the information transfer weight 306 for each second data point within the center point 301 and its dynamic neighborhood 303. This information transfer weight... Also modulated by the occlusion metric, its calculation is based on association metric parameters (including geometric distance and feature similarity) and the occlusion relationship between two points. The specific calculation process is as follows: Step a1, for each first data point Calculate the first data point Its dynamic neighborhood Each second data point within geometric distance ; Step a2, based on the geometric distance and preset distance attenuation coefficient The geometric distance weight is calculated using the distance weight formula. The distance weight formula is: ; Step a3, for each first data point Calculate the first data point Its dynamic neighborhood Each second data point within Feature similarity ; wherein is a feature vector of the first data point , is a feature vector of the second data point ; Step a4, mapping the feature similarity to the interval [0, 1] by a similarity weight formula to obtain a feature similarity weight ; wherein the similarity weight formula is: ; Step a5, determining the occlusion relationship of the two based on the occlusion metric of the first data point and the occlusion metric of the second data point ; Step a6, obtaining an occlusion relationship factor of the occlusion relationship based on a preset factor rule; Step a7, obtaining the information transmission weight between the first data point and the second data point based on the geometric distance weight , the feature similarity weight and the occlusion relationship factor obtained above by a weight product formula; wherein the weight product formula is: .
[0045] In the embodiments of the present application, first, the geometric distance weight is calculated. For the center point 301 and the second data point in its neighborhood, the Euclidean geometric distance between the two is calculated, and the geometric distance weight is calculated by a distance weight formula , wherein is a preset distance attenuation coefficient, and the formula makes the points with closer distance have higher weight. Second, the feature similarity weight is calculated. The cosine similarity between the feature vector of the center point 301 and the feature vector of the second data point (the feature vector here comes from the output of the last layer graph convolution or the initial input) is calculated, and the similarity value is mapped to the interval [0, 1] by a similarity weight formula to obtain the feature similarity weight, which makes the points with more similar features contribute more in information transmission. Then, based on the occlusion metric of the center point 301 and the occlusion metric of the second data point , the occlusion relationship factor is determined according to a preset factor rule For example, when a non-occluded point passes information to an occluded point (A , ), can be set to a value greater than 1 (such as 1.5) to enhance the transmission of reliable information, corresponding to the weight enhancement in Figure 3 ; when information is transmitted between two occluded points (B , ), can be set to a value less than 1 (such as 0.5) to suppress the spread of noise, corresponding to the weight suppression in Figure 3 ; in other cases, can be set to 1. Finally, the above three weight components are multiplied by the weight product formula to obtain the final information transmission weight 306.
[0046] Based on the dynamic neighborhood 303 and the corresponding information transmission weight 306 calculated for each point, the graph convolution module 20 performs a graph convolution operation on each first data point, updates its own features by weighted aggregation of the features of the second data points in its dynamic neighborhood. This process can be represented as , where and are the learnable weight matrix and bias term in the graph convolution layer, is the dynamic neighborhood of the first data point , ReLU is a nonlinear activation function, and the addition represents a residual connection. By stacking multiple layers of such occlusion-adaptive graph convolution layers, information is effectively propagated in the point cloud, and finally the updated features of all first data points are integrated to form the initial feature map 503. In this initial feature map 503, the features of non-occluded points are relatively accurate, while the features of occluded points, although robustly aggregated, may still have the problem of incomplete information.
[0047] Then, the process enters the core feature completion step S300, which is performed by the feature completion module 30, and its working principle is shown in Figure 4 . The goal of this step is to complete the features of the occluded points 401 in the initial feature map 503 using a dual constraint mechanism, and finally generate a high-quality complete feature map 504. The two constraints are the global physical constraint from the pre-set face geometry prior model 403, and the local contextual constraint from the feature manifold 404 composed of non-occluded points.
[0048] First, the feature completion module 30 needs a pre-set face geometry prior model 403. The training process of the pre-set face geometry prior model 403 includes: Step b1, collect several unoccluded three-dimensional face data to form a training sample set; Step b2, obtain individual facial difference features and expression deformation features from the training sample set; Step b3, perform statistical analysis on the training sample set to calculate an average face model ; Step b4, perform dimensionality reduction extraction on the individual facial difference features by principal component analysis to obtain an identity base vector ; Step b5, perform dimensionality reduction extraction on the expression deformation features by principal component analysis to obtain an expression base vector ; Step b6, construct a function representing a face shape based on the average face model , the identity base vector , and the expression base vector : ; wherein, is a face shape, is an identity parameter, is an expression parameter; Step b7, impose an L2 norm constraint on the expression parameter ; Step b8, take the coordinate error between the face shape and the corresponding real face shape in the training sample set as an optimization objective to construct a loss function: ; wherein, is a face shape generated for the mth training sample, is a real face shape of the mth training sample, is an identity parameter corresponding to the mth training sample, is an expression parameter corresponding to the mth training sample, is the sum of squared Euclidean distances between corresponding points of the generated face shape and the real face shape; Step b9, minimize the loss function to obtain the face geometry prior model.
[0049] In this embodiment, the model is a parameterized three-dimensional face model that can generate a three-dimensional face shape conforming to human facial anatomical structure and expression change rules through a set of low-dimensional parameters. The model itself can be represented as a function: . Wherein, is a generated face shape; is a pre-computed average face model; and Identity basis vector and expression basis vector, which are usually extracted by dimension reduction methods such as principal component analysis on large-scale unoccluded three-dimensional face database, respectively encode the identity difference between individuals (such as face shape, facial feature ratio) and the expression change within an individual (such as joy, anger, sorrow, and happiness); is an identity parameter, is an expression parameter, which is a low-dimensional parameter for controlling the identity and expression of a face. The model provides global and structurally reasonable constraints for feature completion.
[0050] The specific process of feature completion is as follows: first, the feature completion module 30 extracts the features of all non-occluded points 402 (i.e., points with an occlusion metric less than or equal to a preset threshold, such as 0.3) from the initial feature map 503 to form a non-occluded point feature set. Since these features all correspond to real and valid facial information, their distribution in the high-dimensional feature space is not random, but contains an inherent low-dimensional structure. The system can use a manifold learning algorithm such as Isomap to construct a low-dimensional feature manifold 404 based on the non-occluded point feature set. This manifold can be regarded as a map of the individual's valid facial features, reflecting the non-linear structural relationship between the features.
[0051] wherein the feature manifold is a non-linear structural mapping of high-dimensional non-occluded point features in a low-dimensional space - due to the correlation of local facial structures (e.g., strong correlation between eye corner points, eyelid points, and brow points), the distribution of high-dimensional features of non-occluded points (such as geometric coordinates, texture features, etc.) will form a continuous low-dimensional surface, i.e., a feature manifold. This manifold can accurately reflect the internal correlation between non-occluded point features (e.g., points with similar distances have small geodesic distances on the manifold and high feature similarity), providing local structural constraints for context feature generation on occluded points.
[0052] wherein the specific parameter settings of the Isomap algorithm are: number of neighbors (selecting 10 nearest neighbors of each non-occluded point to construct a neighborhood graph, which is consistent with the correlation of local facial features), and the dimension of the low-dimensional space after dimension reduction is 20 - determined by reconstruction error analysis: when the dimension is ≥ 20, the reconstruction error of the manifold for high-dimensional non-occluded point features is ≤ 5%, which can both preserve the core structure of the features and reduce the subsequent computational complexity; the geodesic distance calculation uses the Dijkstra algorithm (for connected components in the neighborhood graph) to ensure that the distance calculation conforms to the non-linear structure of the manifold.
[0053] Secondly, the system uses a preset mapping function (for example, a trained multilayer perceptron), which can map the feature vector of a non-occluded point to the parameters of the face geometry prior model 403 that can best reconstruct the geometry of the point and its surrounding area.
[0054] The structure of the multi-layer perception is: input layer (receiving high-dimensional features of non-occluded points, dimension D, D is consistent with the feature dimension of the initial feature map) → hidden layer 1 (number of neurons 256, activation function ReLU) → hidden layer 2 (number of neurons 128, activation function ReLU) → batch normalization layer → output layer (output dimension is 51, corresponding to the parameters of the face geometry prior model: identity parameter (50 dimensions) + expression parameter (1 dimension)). The network training adopts a mean square error loss function (the loss is the coordinate error of the face shape reconstructed by the model parameters generated by the mapping and the real face shape), and the training data set is a plurality of non-occluded point features of unoccluded three-dimensional faces and corresponding real model parameters.
[0055] In addition, in order to make the boundary of the completed area more smooth and natural, the system will also optimize the initial binary occlusion mask generated based on the occlusion degree of each point. Specifically, the system will check the occlusion points located at the edge of the mask, and if the initial features of an edge occlusion point are highly similar to the features of the non-occluded points in its local spatial neighborhood, it is considered that the point is likely to be incorrectly labeled as an occlusion point, and it is excluded from the set of occlusion points, thereby obtaining an optimized occlusion mask.
[0056] Among them, the occlusion mask of the prior art only generates a binary result (0 / 1) based on the occlusion degree, which is prone to edge misjudgment (for example, non-occluded points at the junction of the cheek edge and the hair are misjudged as occluded points). And the application optimizes the feature similarity, compares the feature similarity (such as cosine similarity ≥0.85) of the occlusion points at the edge of the mask with the non-occluded points in the neighborhood, and removes the misjudged points. This optimization can reduce the misjudgment rate of the occlusion mask, avoid the incorrect completion of non-occluded points by subsequent feature completion, and reduce invalid calculations.
[0057] Then, for each occlusion point 401 in the optimized occlusion mask, double-constraint completion is performed. On the one hand, physical constraint features 405 are generated. The system selects non-occluded points in the local spatial neighborhood of the occlusion point 401, and uses the above-mentioned preset mapping function to map the features of these neighborhood non-occluded points into the parameters of the face geometry prior model, respectively, and then takes the average of all obtained parameters to obtain a robust parameter estimate . The estimate represents the identity and expression of the whole face inferred from the visible part. This parameter estimate is input into the face geometry prior model 403, which generates a complete 3D face that is biologically plausible. From this generated face, the feature corresponding to the current occluded point 401 is extracted, i.e., the physically constrained feature 405. This feature ensures that the completion result is a natural face globally. On the other hand, the contextual feature 406 is generated. The system computes the geodesic distance on the feature manifold 404 between the current occluded point 401 (using its feature in the initial feature map) and the non-occluded points in its local spatial neighborhood, based on the constructed feature manifold 404. It can be understood that the geodesic distance better reflects the intrinsic similarity between features than the Euclidean distance. The attention weights are calculated based on the geodesic distance (e.g., the closer the distance, the higher the weight), and then the features of the non-occluded points in the neighborhood are aggregated by weighting with these attention weights, and the aggregated result is the contextual feature 406. This feature ensures that the completion result is highly consistent with the local feature style and details of the current visible part of the face.
[0058] where the physically constrained feature 405 is the occluded point feature generated based on the face geometry prior model, which is consistent with the physiological structure of the human face. The core constraint is global reasonableness, i.e., the model parameters (identity parameters , expression parameters ) inferred from the neighborhood non-occluded points ensure that the generated occluded point feature is consistent with the geometric structure of the whole face (e.g., the ratio of nose height to nose width, the correlation of lip thickness to chin contour), avoiding unnatural results such as disjointed nose features and cheek features.
[0059] The contextual feature is the occluded point feature generated based on the feature manifold, which is consistent with the local neighborhood feature style of the occluded point. The core constraint is local consistency, i.e., the attention weights are calculated based on the geodesic distance (rather than the Euclidean distance) on the manifold, ensuring that the generated occluded point feature is highly matched with the surrounding non-occluded point features (e.g., the texture style and curvature features of the surrounding non-occluded points), avoiding problems such as sudden changes in local features (e.g., smooth cheeks with rough textures).
[0060] Finally, the physically constrained feature 405 and the contextual feature 406 are fused by a fusion step 408 to obtain the final completion feature 407 of the occluded point. The fusion formula can be wherein is a hyperparameter for balancing the importance of global constraints and local constraints, is the physically constrained feature, For contextual features, the initial features of all occluded points are replaced with their corresponding complete features 407, and then merged with the original features of all non-occluded points 402 to generate the final, high-quality complete feature map 504.
[0061] The final step of this application's method is the face recognition detection step S400, executed by the recognition module 40. The recognition module 40 receives the complete feature map 504 output by the feature completion module 30. Since the feature map is now complete and high-fidelity, the recognition module 40 can operate on near-ideal data. Specifically, it first applies a global feature aggregation operation (such as average pooling or max pooling) to the complete feature map 504, compressing it into a fixed-dimensional global face feature vector. Then, it compares this feature vector with the feature vectors of known identities pre-stored in the database, calculating the distance or similarity between them (such as Euclidean distance or cosine similarity). Finally, the identity with the smallest distance and less than a preset judgment threshold is taken as the recognition result, i.e., the face identity 505, and output to the user or subsequent applications.
[0062] Example 2 As an optional implementation, the occlusion-aware feature extraction step S200 in Embodiment 1 can also be implemented using other technical solutions. In Embodiment 1, the graph convolution module 20 uses a specially designed occlusion-adaptive graph convolutional network. In this embodiment, the occlusion-aware feature extraction function can be implemented through other advanced point cloud processing network architectures, such as attention-based networks.
[0063] Specifically, the graph convolution module 20 can be replaced with a point cloud Transformer-based architecture. In a standard point cloud Transformer, the calculation of attention scores mainly depends on the relationship between queries, keys, and values. To incorporate occlusion awareness into the variant of this embodiment, when calculating the attention score between any two points (the first data point and the second data point), in addition to the relationship between their features, a bias term determined by their occlusion metric is introduced.
[0064] For example, when calculating the center point For neighboring points When focusing on a given amount of attention, the attention score can be adjusted to... ,in This refers to the occlusion bias term. From the center point The query vector is obtained by linear transformation of the features. For neighborhood points The transpose of the key vector obtained by linear transformation of the features. Key vector The dimension, for is scaled (divided by ) by the dot product of , is the scaled dot product attention score.
[0065] where the bias term is calculated in the same way as the occlusion relationship factor in embodiment 1. The idea is similar: if is an occluded point and is a reliable non-occluded point, then is a positive value, thus increasing the attention weight of on to encourage occluded points to learn from reliable points; if and are both occluded points, then is a negative value to reduce the attention between them, thus suppressing noise propagation. In this way, the attention mechanism can also dynamically and adaptively adjust the weight of information transmission according to the occlusion of the points, achieving the same technical idea as embodiment 1, i.e. enabling the feature extraction process to actively adapt to occlusion, thus providing a more robust initial feature map 503 for the subsequent feature completion step. This embodiment shows that the core idea of realizing occlusion-aware feature extraction is universal and not limited to a specific network structure.
[0066] Embodiment 3 This embodiment provides a variant selection of the core component used in the feature completion step S300 in embodiment 1, i.e. the face geometry prior model 403. Embodiment 1 describes a general parametric face model constructed based on statistical analysis. However, the technical solution of the present application is not limited to this specific model.
[0067] As a preferred implementation scheme, the feature completion module 30 can use a more advanced and detailed parametric face model in the industry, such as the FLAME model. The FLAME model not only represents identity and expression, but also explicitly models head pose, mandibular joint movement, etc. to generate more realistic and dynamic face geometry.
[0068] In this embodiment, the overall procedure of the feature completion step S300 remains the same, but all the operations related to the face geometry prior model 403 will be based on the FLAME model. Specifically, a pre-set mapping function will be trained to map the features of the non-occluded points to the parameters of the FLAME model (including shape, expression, pose, etc. parameters). When generating the physically constrained features 405 for the occluded points 401, the system will estimate the parameters of the FLAME model through the non-occluded points within its neighborhood, and input these parameters into the FLAME model to generate a complete, pose-correct, and expression-natural three-dimensional face. Then the features of the corresponding positions are extracted from this generated face as the physically constrained features 405. The subsequent fusion step with the contextual features 406 is also the same as in embodiment 1.
[0069] This embodiment demonstrates the flexibility and scalability of the dual-constrained feature completion framework proposed in this application, whose core idea does not rely on a certain specific face prior model. Any prior model that can provide a parameterized representation and generate reasonable face geometry can be integrated into this framework to provide global physical constraints.
[0070] Embodiment 4 This embodiment provides a variant implementation of the two specific technical points in the feature completion step S300 in embodiment 1, which respectively involve the construction algorithm of the feature manifold 404 and the fusion manner of the features 408.
[0071] On the one hand, regarding the construction of the feature manifold 404, the Isomap algorithm is used in embodiment 1 to construct the low-dimensional manifold of the non-occluded point features. As an alternative implementation, other manifold learning algorithms can also be used, such as the local linear embedding algorithm. The core idea of the local linear embedding algorithm is to assume that the data points can be linearly reconstructed by the points within their neighborhood, and it aims to maintain this local linear reconstruction relationship in the low-dimensional space. In the feature completion module 30, the local linear embedding algorithm can be used to process the non-occluded point feature set, and an effective low-dimensional feature manifold 404 can also be obtained. The subsequent steps of calculating the geodesic distance and the contextual features 406 based on this manifold remain unchanged. This shows that the technical choice of constructing the feature manifold is diverse.
[0072] On the other hand, regarding the fusion manner of the physically constrained features 405 and the contextual features 406, a fixed weighted fusion manner is used in embodiment 1, i.e. where the weight is a hyperparameter that needs to be manually adjusted. As a more advanced and adaptive implementation, this embodiment adopts a learnable gating fusion mechanism.
[0073] Specifically, a small neural network can be designed as a gating unit. The input to this gating unit is the physical constraint feature 405 and the context feature 406 to be fused. Through several fully connected layers and a sigmoid activation function, the gating unit outputs a dynamic fusion weight. Its value is between 0 and 1. The final completion feature 407 is calculated using the following formula: This gating unit can be trained end-to-end with the entire recognition network. During training, the network automatically learns, based on different input samples and occlusion conditions, when to trust physical constraint features representing the global structure more and when to prioritize contextual features that maintain local consistency. For example, when the visible area information is very rich, the network may learn to give higher weights to contextual features; while when the visible area information is sparse, it may rely more on the physical constraints provided by the geometric prior model. This adaptive fusion mechanism makes the feature completion process more intelligent, further improving completion quality and final recognition performance.
[0074] The gating fusion mechanism in this application can dynamically adjust the weights according to the degree of occlusion: when the number of non-occluded neighboring points of an occluded point is ≥15 (rich local information), the dynamic weights output by the gating unit will be adjusted accordingly. The value ranges from 0.3 to 0.4, emphasizing contextual features (ensuring local consistency); when the number of non-occluded neighboring points is less than 5 (sparse local information), The value is set to 0.6~0.7, placing greater emphasis on physical constraint features (ensuring global rationality). This adaptive adjustment improves the matching accuracy of the completed features compared to fixed weights. The accuracy improved by 8.7%, especially in scenarios where a large area of hair on one side was obscured, where the Euclidean distance between the completed features and the real features was significantly reduced.
[0075] Example 5 The method in this application is compatible with cross-modal inputs of geometric and texture features: if the 3D face point cloud also contains texture information (such as RGB color values), the texture features and geometric features can be concatenated and input into the initial feature extraction process. During dynamic neighborhood adjustment, texture similarity (calculating the cosine similarity of RGB values) is added to the association metric parameter, and the texture feature dimension is incorporated when constructing the feature manifold, with the input dimension of the mapping function being expanded synchronously. After adding texture features, the face recognition accuracy is further improved in scenarios with double occlusion such as masks and glasses.
[0076] In this embodiment, the input is a textured 3D face point cloud, which includes each first data point. geometric coordinates ( In addition to geometric features, it also includes each Corresponding texture information (i.e., RGB color values) , value range [0, 255], i.e. texture feature). The texture information is synchronously collected by the depth camera: while the depth camera outputs the three-dimensional geometric coordinates, the RGB value of the corresponding pixel is obtained through the color imaging module, and then each point cloud data point is bound to the RGB value through pixel-point cloud coordinate matching.
[0077] After the three-dimensional face point cloud normalization processing in the occlusion measurement step S100 of embodiment 1, a texture feature normalization step is added to ensure that the numerical scales of the geometric features and the texture features are consistent: the RGB value of each first data point is normalized: , , After normalization, the texture feature has a value range of [0, 1]; the geometric feature and the texture feature are spliced to form a cross-modal initial feature: , where is the normalized geometric coordinate in embodiment 1 , The dimension is 6 (3D geometry + 3D texture).
[0078] On the basis of the correlation measurement parameters (geometric distance, feature similarity) for calculating the information transmission weight in embodiment 1, a texture similarity is added, and the specific calculation and fusion logic is as follows: 1) Texture similarity calculation For the first data point and the second data point in its dynamic neighborhood, the cosine similarity of the normalized texture features of the two is calculated, and the formula is:
[0079] If , it means that and are completely consistent in texture (such as the same skin color area); if , it means that the texture is completely irrelevant (such as the skin texture of the face and the hair texture).
[0080] 2) Information transmission weight fusion texture similarity The texture similarity is converted into a texture weight , and is integrated into the original weight formula. The modified information transmission weight formula is: .
[0081] , where is the geometric distance weight: same as embodiment 1, the formula is: , (based on 6D cross-modal feature statistical optimization).
[0082] wherein, is the cross-modal feature similarity weight: replace the original geometric feature similarity with the cosine similarity of the cross-modal initial feature , the formula is the same as embodiment 1, and the value after mapping is [0, 1].
[0083] wherein, is the texture weight): perform sigmoid mapping on , the formula is , to ensure that the texture similarity is greater than or equal to 0.5 0.5 (enhance the contribution of the texture consistent point).
[0084] wherein, is the occlusion relationship factor: the same as embodiment 1 (non-occlusion to occlusion transmission takes 1.5, and occlusion to occlusion transmission takes 0.5).
[0085] After that, the input of the graph convolution layer in embodiment 1 is changed from 3-dimensional geometric features to 6-dimensional cross-modal initial features , and the learnable parameters of the graph convolution layer are adjusted accordingly: the dimension of the original weight matrix W is changed from 3xC (C is the output feature channel number) to 6xC; the dimension of the bias term b remains unchanged as 1xC; the output feature channel number C is still set to 64.
[0086] In embodiment 1, the low-dimensional feature manifold is constructed based on the non-occluded point feature set. In this embodiment, the non-occluded point feature set is replaced by cross-modal features, and the specific adjustments are as follows: 1) Non-occluded point cross-modal feature screening Screen non-occluded points that satisfy , extract 64-dimensional cross-modal features (instead of original geometric features) output by the graph convolution, and form a non-occluded point cross-modal feature set .
[0087] 2) Isomap algorithm parameter adjustment Because the feature dimension is expanded from 3-dimensional geometric features to 64-dimensional cross-modal features, the Isomap algorithm parameters are adjusted to ensure the accuracy of manifold reconstruction: the number of neighbors : from 10 to 15 (more neighbors ensure the preservation of local structure of high-dimensional features); the dimension of the low-dimensional space after dimension reduction: from 20 to 30; geodesic distance calculation: still uses Dijkstra algorithm, but the neighborhood graph is constructed based on the Euclidean distance of cross-modal features (instead of the original geometric distance).
[0088] In addition, the input of the mapping function in embodiment 1 is changed from non-occluded point geometric features to non-occluded point cross-modal features, and the specific modifications are as follows:
[0089] Training data set: On the basis of the original unoccluded three-dimensional face geometry data, corresponding texture data is supplemented to form a geometry + texture paired training sample (such as 1200 groups, 200 more texture sample groups than the original data set); Loss function: The coordinate error of the face shape generated by the model and the real face shape is still used (since the texture feature is only used to assist parameter mapping, the final physical constraint feature is generated by the geometry prior model, which is irrelevant to the texture); Training iteration number: Adjusted from 500 rounds to 600 rounds to ensure network convergence under high-dimensional input.
[0090] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. The present application can have various changes and modifications for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0091] It should be understood that the size of the serial number of each step in the above embodiment does not mean the order of execution. The execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0092] Corresponding to the face recognition detection method described in the above embodiment, the face recognition detection device provided in the embodiments of the present application is not described again.
[0093] It should be noted that the information interaction, execution process, etc. between the modules, since based on the same concept as the method embodiments of the present application, the specific functions and the technical effects brought by them can be referred to the method embodiments part, and will not be repeated here.
[0094] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or software. In addition, the specific name of each functional unit and module is only for easy distinction, and does not limit the protection scope of the present application. The specific working process of the unit and module in the above system can be referred to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0095] The embodiment of the present application further provides a computer device, comprising at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement the steps in any of the above method embodiments.
[0096] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the steps in any of the above method embodiments.
[0097] The embodiment of the present application provides a computer program product, which, when executed on a computer device, enables the computer device to implement the steps in any of the above method embodiments.
[0098] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the embodiment of the present application can implement all or part of the above-mentioned method processes through a computer program to instruct related hardware to complete, and the computer program can be stored in a computer readable storage medium. The computer program is executed by a processor to implement the steps in each of the above method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable file or some intermediate form. The computer readable medium at least includes any entity or device capable of carrying the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium. For example, U disk, mobile hard disk, magnetic disk or optical disk, etc. In some jurisdictions, according to legislation and patent practice, the computer readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0099] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0100] Those skilled in the art can appreciate that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized in electronic hardware or in a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solutions. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0101] In the embodiments provided in the present application, it should be understood that the disclosed apparatus / computer device and method can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or in other forms.
[0102] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., they can be located in one place, or distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.
[0103] The above-described embodiments are only used to illustrate the technical solutions of the present application, but not limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
[0104] Figure 6 The structural schematic diagram of the computer device provided by an embodiment of the present application is shown in FIG. 1. As shown in the figure, the computer device of the embodiment includes at least one processor 60 (only one is shown in the figure), a memory 61, and a computer program 62 stored in the memory 61 and executable on the at least one processor 60, wherein the processor 60 implements the steps in any of the face recognition detection method embodiments described above when executing the computer program 62. Figure 6 Figure 6
[0105] The computer device can include, but is not limited to, a processor 60, a memory 61. Those skilled in the art can understand that Figure 6 The computer device is only an example and does not constitute a limitation on the computer device, and can include more or fewer components than shown, or combine certain components, or different components, for example, can also include an input / output device, a network access device, etc.
[0106] The processor 60 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor or can also be any conventional processor.
[0107] The memory 61 can be an internal storage unit of the computer device in some embodiments, such as a hard disk or a memory of the computer device. The memory 61 can also be an external storage device of the computer device in other embodiments, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 61 can include both the internal storage unit and the external storage device of the computer device. The memory 61 is used to store an operating system, application programs, a boot loader, data and other programs, such as program codes of the computer program, etc. The memory 61 can also be used to temporarily store data that has been output or will be output.
[0108] The related user personal information that can be involved in the embodiments of the present application is strictly in accordance with the requirements of laws and regulations, and follows the principles of legality, legitimacy and necessity, and processes the personal information of the user based on the reasonable purpose of the business scene, which is provided by the user in the process of using the product / service or generated due to the use of the product / service, and authorized by the user.
[0109] The user's personal information handled by the applicant will vary depending on the specific product / service scenario and will be subject to the specific scenario in which the user uses the product / service and may involve the user's account information, device information, driving information, vehicle information, or other related information. The applicant will treat the user's personal information and its processing with a high degree of diligence.
[0110] The applicant attaches great importance to the security of user personal information and has taken security protection measures in accordance with industry standards that are reasonable and feasible to protect user information and prevent unauthorized access, public disclosure, use, modification, damage, or loss of personal information.
Claims
1. A face recognition detection method, characterized in that, include: Obtaining 3D face point clouds First data points in each of the middle Corresponding occlusion metric ; Based on each first data point Corresponding occlusion metric The three-dimensional face point cloud Perform occlusion-adaptive graph convolution operations to generate initial feature maps. ; wherein, the initial feature map Including occluded points and unoccluded points; For the initial feature map The occluded points in the image are used to perform feature completion based on a preset face geometry prior model and a feature manifold composed of unoccluded points, generating a complete feature map. ; Based on the complete feature map Facial recognition detection is performed to determine the identity of the person.
2. The face recognition detection method as described in claim 1, characterized in that, The acquisition of 3D face point clouds First data points in each of the middle Corresponding occlusion metric ,include: The three-dimensional face point cloud Perform normalization processing; Normalized 3D face point cloud The input is fed into a preset occlusion prediction network, and the output is a normalized 3D face point cloud. First data points in each of the middle The corresponding occlusion probability, and the first data points The occlusion probability is used as the first data point Occlusion metric .
3. The face recognition detection method as described in claim 1, characterized in that, The first data points Corresponding occlusion metric The three-dimensional face point cloud Perform occlusion-adaptive graph convolution operations to generate initial feature maps. ,include: For 3D face point cloud Each first data point in Based on this first data point Occlusion metric Dynamically adjust the first data point The neighborhood range used for information aggregation during graph convolution operations is used to obtain the first data point. dynamic neighborhood ;in, For dynamic neighborhood The second data point in; For each first data point Based on this first data point Occlusion metric The first data point Its dynamic neighborhood The second data point within The correlation metric parameters are used to calculate the first data point. With the second data point Information transfer weight between ; Based on each first data point dynamic neighborhood And information transmission weight For each first data point Perform graph convolution operations to aggregate their dynamic neighborhoods. Second data point The features are used to obtain each first data point. Update features; All first data points The updated features are integrated to generate an initial feature map. .
4. The face recognition detection method as described in claim 3, characterized in that, The above refers to 3D face point clouds Each first data point in Based on this first data point Occlusion metric Dynamically adjust the first data point The neighborhood range used for information aggregation during graph convolution operations is used to obtain the first data point. dynamic neighborhood ,include: For each first data point Based on the first data point Occlusion metric The first data point is determined by a preset neighborhood radius calculation formula. Neighborhood radius during graph convolution operation The formula for calculating the neighborhood radius is as follows: , The preset basic neighborhood radius, Used to limit occlusion measurement for neighborhood radius Maximum contribution, avoiding neighborhood radius Measurement due to occlusion Too large and excessively enlarged; With the first data point Center, neighborhood radius With radius, in the 3D face point cloud Determine the initial spatial neighborhood; Based on the first data point and each second data point in the initial spatial neighborhood The geometric features are used to perform geometric consistency screening on the initial spatial neighborhood, including: Curvature filtering: Calculate the first data point curvature Second data point curvature , retain satisfaction The second data point ; Normal filtering: Calculate the first data point unit normal Second data point unit normal , retain satisfaction The second data point with an angle greater than the preset value ; If the first data point Occlusion metric If the occlusion metric is greater than the first preset occlusion metric, then in the initial spatial neighborhood after geometric consistency filtering, the occlusion metric will be... The second data point is less than or equal to the second preset occlusion metric. As the first data point dynamic neighborhood ; If the first data point Occlusion metric If the occlusion metric is less than or equal to the first preset occlusion metric, then the initial spatial neighborhood after geometric consistency filtering is directly used as the first data point. dynamic neighborhood .
5. The face recognition detection method as described in claim 4, characterized in that, The correlation metric parameters include geometric distance. Feature similarity; For each first data point Based on this first data point Occlusion metric The first data point Its dynamic neighborhood The second data point within The correlation metric parameters are used to calculate the first data point. With the second data point Information transfer weight between ,include: Based on the geometric distance and the feature similarity, combined with the first data point Second data point Occlusion relation factors determined by the occlusion relation The first data point was calculated. With the second data point Information transfer weight between .
6. The face recognition detection method as described in claim 1, characterized in that, The training process of the facial geometric prior model includes: Collect a number of unobstructed 3D face data to form a training sample set; Individual facial difference features and expression change features are obtained from the training sample set; Statistical analysis was performed on the training sample set to calculate the average face model. ; Principal component analysis was used to perform dimensionality reduction extraction of the individual facial difference features to obtain the identity basis vector. ; Principal component analysis was used to extract the dimensionality-reducing features of the facial expressions, resulting in expression basis vectors. ; Based on the above average face model Identity basis vectors and expression basis vectors Construct a function to represent the shape of a human face: ;in, The shape of a human face, For identity parameters, For facial expression parameters; For facial expression parameters Apply L2 norm constraints; In the shape of a human face The coordinate error between the actual face shape and the coordinates in the training sample set is used as the optimization objective to construct a loss function; Minimize the loss function to obtain the face geometric prior model.
7. The face recognition detection method as described in claim 6, characterized in that, The initial feature map The occluded points in the image are used to perform feature completion based on a preset face geometry prior model and a feature manifold composed of unoccluded points, generating a complete feature map. ,include: Obtain the initial feature map The features of the unoccluded points in the dataset are used to form a set of unoccluded point features; wherein, the unoccluded point is a first data point that satisfies a preset unoccluded condition. ; Based on the set of unoccluded point features, a low-dimensional feature manifold is constructed using the Isomap algorithm; Based on the real face coordinates of the feature manifold and unoccluded points, a multilayer perceptron is trained as a mapping function. Based on each first data point Corresponding occlusion metric For each first data point Mark whether a point is an occlusion point and generate an initial occlusion mask; The initial occlusion mask is optimized by removing occlusion points at the edge of the initial occlusion mask that have similar features to the non-occlusion points in the local spatial neighborhood of the occlusion point at that edge, thus obtaining the optimized occlusion mask. For each occlusion point in the optimized occlusion mask, select the non-occlusion points in the local spatial neighborhood of the occlusion point. Through the trained mapping function, map the features of the non-occlusion points in the local spatial neighborhood to the parameters of the face geometry prior model. Then, take the average value of the parameters based on the number of non-occlusion points in the local spatial neighborhood to obtain the parameter estimate of the face geometry prior model. The parameter estimates are input into the face geometry prior model to generate physical constraint features for each occluded point in the optimized occlusion mask. Simultaneously, based on the constructed feature manifold, the geodesic distance between the occluded point and the non-occluded points in the local spatial neighborhood of the occluded point is calculated on the manifold. The attention weight is calculated based on the geodesic distance. The features of the non-occluded points in the local spatial neighborhood of the occluded point are weighted and aggregated through the attention weight to generate the context features of the occluded point. Based on the physical constraint features and context features of each occlusion point, the completion features of the occlusion point are obtained through weighted fusion; The completed features of all occluded points are merged with the features of unoccluded points to generate a complete feature map. .
8. A face recognition detection device, characterized in that, include: The acquisition module is used to acquire 3D face point clouds. First data points in each of the middle Corresponding occlusion metric ; The graph convolution module is used to perform graph convolution based on each first data point. Corresponding occlusion metric The three-dimensional face point cloud Perform occlusion-adaptive graph convolution operations to generate initial feature maps. ; wherein, the initial feature map Including occluded points and unoccluded points; The feature completion module is used to complete the initial feature map. The occluded points in the image are used to perform feature completion based on a preset face geometry prior model and a feature manifold composed of unoccluded points, generating a complete feature map. ; The identification and detection module is used to identify and detect based on the complete feature map. Facial recognition detection is performed to determine the identity of the person.
9. A computer device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as claimed in any one of claims 1 to 7.
10. A computer program product, characterized in that, When the computer program product is run on a computer device, it causes the computer device to perform the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Shielded face recognition method and device, storage medium and terminal
CN112487886A
Facial expression recognition method, system and device and readable storage medium
CN112528764A
Mask-shielded face recovery method based on an adaptive context attention mechanism
CN113378980A
Shielded face restoration method based on geometric perception priori guidance
CN114764754A
Personnel identification and positioning method based on RGB-D image under face shielding condition
CN117133032A