Face recognition method and device based on face depth features, medium and product

By decomposing images based on the principle of single-scale retinal cortex and attention-guided residual network model, combining frequency-time feature extraction and Diloney triangulation method, the accuracy and stability of face recognition under complex lighting conditions are solved, and the adaptability and recognition effect of face recognition are improved.

CN120544249APending Publication Date: 2025-08-26AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510634323.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The prior art has low accuracy, poor stability and insufficient adaptability under complex lighting conditions, mainly due to poor image feature extraction effect and high noise sensitivity.

Method used

The face image is decomposed into illuminated images and reflective images based on the single-scale retinal cortex principle, and the attention-guided residual network model is used to denoise, and the image details are enhanced by the pre-trained frequency-time feature extraction model, and the Dilony triangle network is constructed through the Dilony triangulation method for key point annotation and three-dimensional spatial distance calculation for identification.

Benefits of technology

Under complex lighting conditions, the accuracy, stability and adaptability of face recognition are effectively improved. By effectively separating light and reflective components, noise is suppressed, image texture details are enhanced, and the accuracy of recognition and system robustness are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544249A_ABST
    Figure CN120544249A_ABST
Patent Text Reader

Abstract

The invention discloses a face recognition method and device based on face depth features, a medium and a product. The method comprises the following steps: acquiring a face image, decomposing the face image into an illumination image and a reflection image, and inputting the illumination image and the reflection image into an attention-guided residual network to obtain a de-noised image; inputting the denoised image into a pre-trained frequency time feature extraction model to obtain an enhanced image; constructing triangulation networks for the enhanced image and the reference image based on a Delauni triangulation method, and marking key points; and calculating the distance between each key point of the enhanced image and the key point corresponding to the reference image, and performing face recognition. Illumination and reflection components of an image are separated, a face recognition model is constructed by using a reflection image, noise of the reflection image is suppressed by combining an illumination image, the texture quality of the image is enhanced, and finally an enhanced image is obtained. The three-dimensional space distance of the key points is calculated through the Delauni triangulation method, face recognition is achieved, and the accuracy, stability and adaptability of face recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and face recognition technology, and in particular to a face recognition method, device, medium and product based on face depth features. Background Art

[0002] Facial recognition is a technical method that uses computer vision and feature engineering. It can accurately identify and verify facial identity based on facial information, and plays a key role in scenarios such as security control, identity confirmation, and automated processing.

[0003] Image quality significantly impacts facial recognition. High-quality images facilitate feature extraction and precise matching, while poor quality compromises recognition accuracy. In real-world applications, captured images often suffer from poor quality due to poor user environments or limited device capture capabilities. Existing algorithms struggle to effectively extract image features when modeling captured images, particularly those acquired in extreme light conditions such as strong or low light. This problem is further exacerbated by the instability of traditional neural network model training and its sensitivity to noise, reducing facial recognition accuracy and adaptability. Summary of the Invention

[0004] The present invention provides a face recognition method, device, medium and product based on face depth features to solve the problems of poor image feature extraction under complex lighting conditions, as well as low accuracy, low stability and poor adaptability of face recognition.

[0005] According to one aspect of an embodiment of the present invention, a face recognition method based on face depth features is provided, comprising:

[0006] Obtain the face image to be recognized collected by the terminal, and decompose the face image into an illumination image and a reflection image based on the single-scale retinal cortex principle;

[0007] The illumination image and the reflection image are input into the attention-guided residual network model to obtain the denoised image;

[0008] The denoised image is input into the pre-trained frequency-time feature extraction model to obtain the enhanced image;

[0009] Constructing Delaunay triangulation methods for the enhanced image and at least one standard user's reference image stored in the database, respectively, and marking key points in the enhanced image and the reference image according to the constructed Delaunay triangulation methods;

[0010] The three-dimensional spatial distance between each key point position of the enhanced image and the corresponding key point position of the reference image is calculated, and face recognition is performed based on the calculation of each spatial distance.

[0011] According to another aspect of an embodiment of the present invention, a face recognition device based on face depth features is provided, comprising:

[0012] A decomposition module is used to obtain the face image to be recognized collected by the terminal and decompose the face image into an illumination image and a reflection image based on the single-scale retinal cortex principle;

[0013] A denoising module, which is used to input the illumination image and the reflection image into the attention-guided residual network model to obtain the denoised image;

[0014] An enhancement module is used to input the denoised image into a pre-trained frequency-time feature extraction model to obtain an enhanced image;

[0015] a labeling module for constructing a Delaunay triangulation method for the enhanced image and a reference image of at least one standard user stored in the database, respectively, and labeling key points in the enhanced image and the reference image according to the constructed Delaunay triangulation method;

[0016] The recognition module is used to calculate the three-dimensional spatial distance between the key point positions of the enhanced image and the corresponding key point positions of the reference image, and perform face recognition based on the calculation of each spatial distance.

[0017] According to another aspect of an embodiment of the present invention, an electronic device is provided, the electronic device comprising:

[0018] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the face recognition method based on face depth features as described in any embodiment of the present invention.

[0019] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the face recognition method based on face depth features described in any embodiment of the present invention when executed.

[0020] According to another aspect of an embodiment of the present invention, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the steps of the method according to any embodiment of the present invention are implemented.

[0021] The technical solution of the embodiment of the present invention obtains a facial image to be recognized collected by a terminal, and decomposes the facial image into an illumination image and a reflection image based on the principle of single-scale retinal cortex, which can effectively separate the illumination and reflection components in strong or weak light environments; inputs the illumination image and the reflection image into an attention-guided residual network model to obtain a denoised image, uses the reflection image to build a face recognition model, and combines the illumination image to suppress the noise of the reflection image, thereby improving the overall image quality. The denoised image is input into a pre-trained frequency-time feature extraction model to obtain an enhanced image, thereby enhancing the image texture details; based on the Delaunay triangulation method, a Delaunay triangulation network is constructed for the enhanced image and a reference image of at least one standard user stored in a database, and key points are respectively marked in the enhanced image and the reference image based on the constructed Delaunay triangulation network; the three-dimensional spatial distance between the position of each key point of the enhanced image and the corresponding key point position of the reference image is calculated, and face recognition is performed based on the calculation of each spatial distance, thereby solving the problem of poor image feature extraction effect under complex lighting conditions and improving the accuracy, stability and adaptability of face recognition.

[0022] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0024] Figure 1 This is a flowchart of a face recognition method based on face depth features provided according to the first embodiment of the present invention;

[0025] Figure 2 This is a flowchart of another face recognition method based on face depth features provided by Embodiment 2 of the present invention;

[0026] Figure 3 is a schematic diagram of face recognition based on face depth features applicable to an embodiment of the present invention;

[0027] Figure 4 2 is a schematic diagram of the structure of a face recognition device based on face depth features according to a third embodiment of the present invention;

[0028] Figure 5The present invention is a schematic diagram of the structure of an electronic device for implementing the face recognition method based on face depth features according to an embodiment of the present invention. DETAILED DESCRIPTION

[0029] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0030] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0031] Example 1

[0032] Figure 1 This is a flowchart of a face recognition method based on face depth features provided in the first embodiment of the present invention. This embodiment is applicable to the case of performing face recognition based on collected face images. The method can be performed by a face recognition device based on face depth features. The face recognition device based on face depth features can be implemented in the form of hardware and / or software and can generally be configured in an electronic device. Figure 1 As shown, the method includes:

[0033] S110: Obtain a face image to be recognized collected by the terminal, and decompose the face image into an illumination image and a reflection image based on a single-scale retinal cortex principle.

[0034] In embodiments of the present invention, the retinal cortex principle can be specifically understood as an image processing theory that simulates the response characteristics of retinal cells to visual features at different scales. By leveraging color constancy, it achieves a balance between dynamic range compression, edge enhancement, and color constancy, enabling adaptive enhancement of various image types, rather than solely enhancing a single type of feature. The single-scale retinal cortex principle assumes that an image consists of an illumination image and a reflection image. The illumination image can be specifically understood as the incident component information of an object, primarily containing low-frequency information about the image and reflecting the overall illumination distribution of the image. The illumination component is typically obtained by filtering the red, green, and blue channels of the image using a Gaussian surround function. The reflection image can be specifically understood as the reflected portion of an object, primarily containing high-frequency information about the image and reflecting inherent surface properties such as color and texture. The reflection component is typically obtained by subtracting the original image from the illumination component in the logarithmic domain. In addition, a single-scale retinal cortex model can be pre-trained (such as a convolutional neural network model, in which the convolution kernel size can be set to a fixed value to capture single-scale image features) to perform feature extraction and image decomposition on the preprocessed face image to obtain the corresponding illumination image and reflection image.

[0035] Specifically, before facial recognition, users need to check documents such as the "Facial Information Collection Authorization" and the "Personal Information Authorization". Only after the customer agrees to the authorization can the user's facial image be collected.

[0036] The facial image to be identified is obtained through a terminal device (such as a terminal device or a mobile phone camera), and based on the single-scale (such as spatial scale, frequency scale or time scale can be selected, depending on the application scenario and requirements) retinal cortex principle, the facial image is decomposed into an illumination image and a reflection image.

[0037] S120. Input the illumination image and the reflection image into the attention-guided residual network model to obtain a denoised image.

[0038] In an embodiment of the present invention, the attention-guided residual network model can be specifically understood as: a model based on an improved U-type network model that combines an attention mechanism and a residual network structure. Based on the U-type network model, this model inserts an attention module (such as a convolutional block attention module or a spatial attention module) after each or specific convolution layer (such as inserting an attention mechanism in the early layers of the encoder to help capture low-level features of the image (such as edges and textures)), weights the feature map in the channel and spatial dimensions, highlights important features and suppresses unimportant features.

[0039] Residual connections are built around convolutional layers and an attention mechanism. After the input feature map is processed by the convolutional layers and the attention module, it is added to the original input feature map to form a residual connection. Residual connections allow the network to directly pass the output of previous layers to subsequent layers, thereby facilitating the flow of gradients during backpropagation and alleviating the vanishing gradient problem in deep networks. The attention mechanism enables the network to focus on key areas of the input data, suppressing irrelevant or minor information, and helping the network extract image features.

[0040] Specifically, the illumination image and the reflection image are input into an attention-guided residual network model, where features are extracted from each image. The illumination image's feature map contains relatively smooth, low-frequency illumination variations. The reflection image's feature map contains high-frequency details and noise. The attention mechanism analyzes the channel and spatial characteristics of the feature map to identify and highlight regions containing valid information while suppressing regions containing noise. Through a residual connection, the reflection image's feature map, processed by the convolution and attention modules, is added to the original reflection image's feature map, preserving key details of the original image while incorporating new features. Then, at a specific level of the attention-guided residual network model (e.g., an intermediate level in the encoder or the initial level in the decoder), the illumination image's feature map and the reflection image's feature map are fused (e.g., by weighted summation). Using the smoothness of the illumination image as a constraint, the reflection image's feature map is corrected to remove high-frequency noise.

[0041] As mentioned above, noise is removed and image details are enhanced at each iteration of the corresponding multiple layers of the network model. The final output reflection image is a high-quality denoised image after multiple steps of feature extraction, attention guidance, and residual connections.

[0042] S130: Input the denoised image into a pre-trained frequency-time feature extraction model to obtain an enhanced image.

[0043] In an embodiment of the present invention, the pre-trained frequency-time feature extraction model can be specifically understood as: a model pre-trained based on a stable diffusion model, for extracting features of frequency and time dimensions from an image. Among them, the stable diffusion model can be specifically understood as: a generative model that can generate high-quality images by learning the feature distribution of the image, that is, the stable diffusion model compresses the image into a low-dimensional latent space to generate the image through a diffusion process. This process can be understood as gradually adding Gaussian noise to the compressed image latent representation during the forward diffusion process, gradually interfering with the image, causing the image to gradually lose details and eventually become noise. By gradually adding noise, the model can learn the distribution characteristics of the data. During the reverse diffusion process, the model starts from pure noise and gradually removes the noise, so that the image gradually recovers details from the noise to restore a matching clear image.

[0044] It is understandable that high-quality images can provide richer and clearer detail information, so that the pre-trained frequency-time feature extraction model can more accurately capture and retain image features during processing, thereby improving the quality of the generated image.

[0045] Specifically, the obtained high-quality denoised image is input into a pre-trained frequency-time feature extraction model. The model processes the image through an optimization algorithm, enhances the texture quality, makes the image texture clearer and more regular, and finally obtains an enhanced image.

[0046] S140 , constructing Delaunay triangulation networks for the enhanced image and a reference image of at least one standard user stored in the database based on a Delaunay triangulation method, and marking key points in the enhanced image and the reference image according to the constructed Delaunay triangulation networks.

[0047] In the embodiments of the present invention, Delaunay triangulation can be specifically understood as a method of dividing a set of points on a plane into a triangular mesh, ensuring that no point lies within the circumcircle of any triangle. Key points are connected by triangles, forming a Delaunay triangulation network that reflects the shape and structure of a face, effectively embodying its geometric features and topological structure. The resulting triangulation network minimizes the appearance of narrow, long triangles, effectively improving approximation accuracy and maintaining optimal overall network quality.

[0048] A standard user can be understood as a legitimate user pre-registered in the facial recognition system and recognized by the system. A benchmark image can be understood as a facial image of a standard user captured under favorable lighting and viewing angles and stored in the database, serving as a reference for subsequent facial recognition. Key points can be understood as representative and discriminative feature points in a facial image, such as pixels at the eyes, nose tip, and mouth corners, which accurately locate the face and reflect its shape and structure. Typically, 68 key points for facial recognition can be selected, including: eyebrows (5 key points for each eyebrow on the left and right sides, sampled evenly from the left to the right boundary of the eyebrow, used to describe the shape and position of the eyebrows), eyes (6 key points for each eye, located at the left and right boundaries and the upper and lower eyelids, sampled evenly, used to locate the outline and shape of the eyes), lips (a total of 20 key points, including 2 points at the corners of the mouth, 6 points evenly sampled at the outer boundaries of the upper and lower lips, and 3 points evenly sampled at the inner boundaries of the upper and lower lips, used to describe the shape and movement of the lips), nose (including 4 key points on the bridge of the nose and 5 points evenly sampled at the tip of the nose, used to describe the shape and position of the nose) and facial contour (17 key points are evenly sampled to describe the overall contour of the face, including some boundary points of the chin, cheeks and forehead).

[0049] Specifically, the enhanced image is compared with a baseline image of at least one standard user stored in a database, and a facial key point detection algorithm (such as a convolutional neural network) is used to identify key points in the enhanced image and the baseline image. The Delaunay triangulation method is used to construct triangulated networks for the enhanced image and the baseline image of at least one standard user stored in the database with the key points in the image as vertices. Based on the constructed Delaunay triangulation network, the key points are marked in the enhanced image and the baseline image respectively.

[0050] Optionally, based on the above embodiments, a Delaunay triangulation network (DTM) can be pre-constructed for the reference images of standard users stored in the database and stored in the database for rapid use in face recognition and matching, thereby improving system efficiency. Furthermore, to address changes in a person's appearance, the system sets a periodicity based on age-related changes to collect and update the DTM to maintain recognition accuracy. That is, the DTM is typically collected during user registration or system initialization. To account for changes in appearance due to factors such as age, the system can periodically re-collect DTM images to update the DTM and their DTM to ensure accurate face recognition.

[0051] S150 , calculating the three-dimensional spatial distances between the key point positions of the enhanced image and the corresponding key point positions of the reference image, and performing face recognition based on the calculation of the spatial distances.

[0052] Specifically, the three-dimensional distances between the key points of the enhanced image and the corresponding key points of each reference image are calculated. If the three-dimensional distances for all key points are less than a set threshold, the user corresponding to the enhanced image is considered to be the user corresponding to the reference image, and a facial recognition success notification can be issued. Otherwise, a facial recognition failure notification (such as a re-acquisition prompt) is issued.

[0053] The technical solution of the embodiment of the present invention obtains a facial image to be recognized collected by a terminal, and decomposes the facial image into an illumination image and a reflection image based on the principle of single-scale retinal cortex, which can effectively separate the illumination and reflection components in strong or weak light environments; inputs the illumination image and the reflection image into an attention-guided residual network model to obtain a denoised image, uses the reflection image to build a face recognition model, and combines the illumination image to suppress the noise of the reflection image, thereby improving the overall image quality. The denoised image is input into a pre-trained frequency-time feature extraction model to obtain an enhanced image, thereby enhancing the image texture details; based on the Delaunay triangulation method, a Delaunay triangulation network is constructed for the enhanced image and a reference image of at least one standard user stored in a database, and key points are respectively marked in the enhanced image and the reference image based on the constructed Delaunay triangulation network; the three-dimensional spatial distance between the position of each key point of the enhanced image and the corresponding key point position of the reference image is calculated, and face recognition is performed based on the calculation of each spatial distance, thereby solving the problem of poor image feature extraction effect under complex lighting conditions and improving the accuracy, stability and adaptability of face recognition.

[0054] Optionally, based on the above embodiments, decomposing a face image into an illumination image and a reflection image based on the single-scale retinal cortex principle may include:

[0055] According to the Gaussian surround function, the three channels of the face image are filtered respectively to obtain the illumination components as the illumination image;

[0056] The face image and the illumination component are subtracted in the logarithmic domain to obtain the reflection component as the reflection image.

[0057] Specifically, construct a Gaussian surround function Among them, x and y are the horizontal and vertical coordinates of the pixel in the image respectively. σ is the scale parameter of Gaussian surround. A smaller σ means that the scale of the Gaussian template is small, which can better preserve the edge details and the dynamic range becomes larger, but the color cannot be preserved; on the contrary, the color restoration is better, but the dynamic range becomes smaller and the details are poorly preserved. When there are obvious illumination changes in the image (such as strong shadows or highlights), a smaller σ value can be used to preserve the details of these areas. In areas with relatively uniform illumination, the σ value can be appropriately increased to better extract the overall illumination information. The illumination components obtained by filtering the red, green and blue channels of the image respectively using the Gaussian surround function are used as the illumination image, and then the reflection component obtained by subtracting the original image and the illumination component in the logarithmic domain is used as the reflection image, that is, Among them, l i (x, y) is the original face image of the i-th color channel, R i (x, y) is the reflection image of the i-th color channel, L i (x, y) is the illumination image of the i-th color channel, which can be obtained by l i The result is (x, y)*G(x, y), which is the convolution of the original image of the i-th color channel with the Gaussian surround function (using the Gaussian surround function to filter the red, green, and blue channels of the image separately). By filtering the three channels of the facial image using the surround function, the illumination component can be effectively extracted, resulting in an illumination image. The facial image and the illumination component are then subtracted in the logarithmic domain to separate the reflection component, resulting in a reflection image. This decomposition of the facial image is achieved, providing high-quality image features for subsequent face recognition tasks.

[0058] Optionally, based on the above embodiments, the attention-guided residual network model may include: an input module, an encoder module, a decoder module, and an output module connected in sequence, and each module includes at least one convolutional layer;

[0059] The attention-guided residual network model may also include a jump connection module for establishing a jump connection between each convolutional layer and at least one subsequent convolutional layer, and each module of the attention-guided residual network model has a residual connection with at least one subsequent module.

[0060] Specifically, the attention-guided residual network model achieves image denoising and overall image quality enhancement through the collaborative work of multiple modules. The input module receives and preprocesses the image, the encoder module gradually extracts features, the decoder module restores spatial resolution, and the output module generates the final result. Each module contains at least one convolutional layer, which is responsible for feature extraction and transformation. The skip connection module establishes a skip connection between each convolutional layer and at least one subsequent convolutional layer, directly transferring early features to subsequent layers to enhance feature fusion. At the same time, each module has a residual connection with at least one subsequent module, that is, D(x) = x + f(x), where x is the input parameter, f(x) is the output feature of the original module, and D(x) is the residual connection output. Adding the module input directly to the output helps gradient propagation, alleviates the gradient vanishing problem of deep networks, and enables the model to more effectively transfer information and gradients, improving training efficiency and model performance. Through the combination of convolutional layers, skip connections and residual connections, the attention-guided residual network model can learn richer feature representations, enhance the overall quality and detail expression of the image, solve the problem of poor image feature extraction under complex lighting conditions, and improve the accuracy, stability and adaptability of face recognition.

[0061] Optionally, based on the above embodiments, a convolution block attention module attention mechanism is used between the first two convolution layers of the encoder module, and asymmetric convolution layers are used in the skip connection module, the decoder module and the output module.

[0062] In embodiments of the present invention, the convolutional block attention module can be specifically understood as an attention mechanism module used to enhance the network's ability to focus on key features. Attention weights are calculated in both the channel and spatial dimensions, and then applied to the feature map to highlight important features and suppress unimportant ones. An asymmetric convolutional layer can be specifically understood as a layer that uses convolution kernels of different sizes to perform convolution operations in different directions. Typically, two convolution kernels of different sizes can be used, such as 1×k and k×1, corresponding to horizontal convolution operations to capture horizontal features and vertical convolution operations to capture vertical features, where k is the size of the convolution kernel. Asymmetric convolution can expand the receptive field by using convolution kernels of different sizes in different directions. The receptive field can be specifically understood as the area of ​​the input image that a neuron in the network can perceive. A larger receptive field allows the network to capture more contextual information, and asymmetric convolution can decompose a large convolution kernel into two smaller kernels, reducing computational effort.

[0063] Specifically, since faces are concentrated in the center of the image in face recognition tasks, excessive attention to other areas of the image will reduce the model's ability to recognize faces. Therefore, in the encoder module, a convolutional block attention module attention mechanism is integrated between the first two convolutional layers to highlight key features in the initial feature extraction stage. Through the channel attention mechanism, the input feature map is subjected to global maximum pooling and global average pooling to obtain two different feature vectors. These are then fed into a multi-layer perceptron to obtain corresponding channel attention weights. The obtained channel attention weights are summed and the final channel attention weight map is generated through an activation function (highlighting important feature channels and suppressing unimportant feature channels). Through the spatial attention mechanism, the feature map is subjected to maximum pooling and average pooling along the channel dimension to obtain two two-dimensional feature maps. These two feature maps are spliced ​​in the channel dimension, processed using a convolutional layer, and the spatial attention weight map is generated through an activation function (focusing on key areas and suppressing irrelevant areas).

[0064] The asymmetric convolutional layers in the skip connection module are used to fuse features at different levels. When linking low-level and high-level features, convolution kernels of different sizes extract multi-scale features, enhancing expressiveness. The asymmetric convolutional layers in the decoder module reduce computational effort and improve efficiency when restoring spatial resolution. Asymmetric convolutions in the output module refine the output, enhance texture features, and improve output accuracy.

[0065] By using the convolution block attention module between the first two convolutional layers of the encoder module, the attention mechanism can enhance the model's attention to key features. The asymmetric convolutional layers in the skip connection module, decoder module and output module help reduce the amount of computation, expand the receptive field and improve the feature extraction capability, enabling the model to more effectively process and generate high-quality image features, improve the texture quality of the image, solve the problem of poor image feature extraction under complex lighting conditions, and improve the accuracy, stability and adaptability of face recognition.

[0066] Example 2

[0067] Figure 2 A flowchart of another face recognition method based on face depth features provided in Example 2 of the present invention. This embodiment is a refinement of the face recognition method based on face depth features in the above embodiment. Specifically, it can be: before inputting the denoised image into the pre-trained frequency-time feature extraction model to obtain an enhanced image, it can also include: obtaining historical face images collected by the terminal at at least one historical moment, calculating the clarity of each historical face image, and screening out images greater than the target clarity threshold as images to be trained; inputting the images to be trained into the stable diffusion model for pre-training to obtain a pre-trained frequency-time feature extraction model.

[0068] Correspondingly, such as Figure 2 As shown, the method includes:

[0069] S210: Obtain a face image to be recognized collected by the terminal, and decompose the face image into an illumination image and a reflection image based on a single-scale retinal cortex principle.

[0070] S220. Input the illumination image and the reflection image into the attention-guided residual network model to obtain a denoised image.

[0071] S230: Obtain historical facial images collected by the terminal at at least one historical moment, calculate the clarity of each historical facial image, and select images with a clarity greater than a target clarity threshold as images to be trained.

[0072] S240 , inputting the image to be trained into a stable diffusion model for pre-training to obtain a pre-trained frequency-time feature extraction model.

[0073] Specifically, the system obtains historical facial images collected by the terminal at at least one historical moment, calculates the clarity of each historical facial image (e.g., by calculating the clarity through the image gradient), and selects images with a clarity threshold greater than the target as training images. Blurry or low-quality images are excluded to ensure the quality of the training data and the performance of subsequent pre-training models. The screened training images are input into the stable diffusion model for pre-training, resulting in a pre-trained frequency-time feature extraction model to support subsequent tasks such as face recognition.

[0074] S250: Input the denoised image into a pre-trained frequency-time feature extraction model to obtain an enhanced image.

[0075] S260: constructing Delaunay triangulation methods for the enhanced image and at least one standard user's reference image stored in the database, respectively, and marking key points in the enhanced image and the reference image according to the constructed Delaunay triangulation methods.

[0076] S270: Calculate the three-dimensional spatial distances between the key point positions of the enhanced image and the key point positions corresponding to the reference image, and perform face recognition based on the calculation of the spatial distances.

[0077] The technical solution of the embodiment of the present invention is to obtain the facial image to be recognized collected by the terminal, and decompose the facial image into an illumination image and a reflection image based on the principle of single-scale retinal cortex; input the illumination image and the reflection image into the attention-guided residual network model to obtain a denoised image, use the reflection image to build a facial recognition model, and combine the illumination image to suppress the noise of the reflection image to improve the overall image quality. Historical facial images collected by the terminal at historical moments are obtained, the clarity of each image is calculated, and images greater than the target clarity threshold are screened out as images to be trained; the images to be trained are input into a stable diffusion model for pre-training to obtain a pre-trained frequency-time feature extraction model, and the pre-training enhances the model's ability and reliability in extracting image features in actual application scenarios under different lighting and angle conditions. The denoised image is input into the pre-trained frequency-time feature extraction model to obtain an enhanced image, thereby enhancing image texture details. Based on the Delaunay triangulation method, a Delaunay triangulation network is constructed for the enhanced image and the reference image of at least one standard user stored in the database. According to the constructed Delaunay triangulation network, key points are marked in the enhanced image and the reference image respectively; the three-dimensional spatial distance between the key point position of the enhanced image and the corresponding key point position of the reference image is calculated, and face recognition is performed based on the calculation of each spatial distance, which solves the problem of poor image feature extraction effect under complex lighting conditions and improves the accuracy, stability and adaptability of face recognition.

[0078] Optionally, based on the above embodiments, constructing Delaunay triangulation networks for the enhanced image and the reference image of at least one standard user stored in the database based on the Delaunay triangulation method may include:

[0079] In the scenario of target user identity authentication, Delaunay triangulation method is used to construct Delaunay triangulation networks for the enhanced image and the reference image of the target user stored in the database.

[0080] Face recognition based on spatial distance calculations can include:

[0081] Count the number of key points whose spatial distance is greater than or equal to the preset distance threshold;

[0082] If the number of key points is greater than or equal to the preset number threshold, it is determined that the identity authentication of the target user has failed; otherwise, it is determined that the identity authentication of the target user has passed.

[0083] Specifically, in the identity authentication scenario of the target user, the system already knows the identity of the user (which can be determined through the personal information authorized by the user), and the purpose is to verify whether the current user is the real person. The saved benchmark image of the target user is retrieved from the database. Based on the Delaunay triangulation method, a Delaunay triangulation network is constructed for the enhanced image and the benchmark image of the target user saved in the database. For each corresponding key point in the enhanced image and the benchmark image to be verified, the three-dimensional space distance between them is calculated. The number of key points whose three-dimensional space distance is greater than or equal to the preset distance threshold is counted and compared with the preset number threshold. If the number of key points is greater than or equal to the preset number threshold, it is considered that the difference between the two images is too large, the identity authentication of the target user fails, and an authentication failure prompt is issued; otherwise, it is determined that the identity authentication of the target user is passed, and an authentication pass prompt is issued.

[0084] By setting distance and number thresholds for key points, the similarity between the captured image of the target user and the corresponding reference image can be determined, improving the security of identity authentication. Based on specific application scenarios and security requirements, the preset distance and number thresholds can be flexibly adjusted to meet different recognition accuracy requirements, improving the accuracy, stability, and adaptability of facial recognition.

[0085] Furthermore, based on the above embodiments, before constructing Delaunay triangulation networks for the enhanced image and the reference image of at least one standard user stored in the database based on the Delaunay triangulation method, the following steps may be further included:

[0086] Measuring the distance between the enhanced image and the human eyes in the reference image of each standard user stored in the database;

[0087] According to the measured distances, the enhanced image is scaled at least once, so that the distance between human eyes in the enhanced image after the scaling process is consistent with the distance between human eyes in the reference image of the matched standard user.

[0088] Specifically, in face recognition, due to differences in facial features, the distance between the eyes of different people is different, or due to differences in shooting angles, the actual sizes of different face images collected will be different. The coordinates of the key points of the eyes in the enhanced image and each reference image stored in the database are respectively located, and the distance between the eyes (such as the distance from the inner corner of the left eye to the inner corner of the right eye) is calculated. Based on the eye distance of the enhanced image and the eye distance of the reference image, the scaling factor is calculated, and then the enhanced image is interpolated (such as bilinear interpolation) and other operations are performed to achieve scaling. The eye distance of the scaled enhanced image is consistent with the eye distance in the matching reference image, reducing the impact caused by individual differences or different image acquisition conditions, and improving the accuracy, stability and adaptability of face recognition.

[0089] For ease of understanding, the specific application scenarios to which the above-mentioned embodiments of the invention are applicable are described. Face recognition is a process of identifying and verifying faces through computer vision and feature engineering, which can be used in a variety of scenarios such as security control, identity authentication, and automatic face tagging. Among them, image quality directly affects the performance of the face recognition model. Currently, face recognition methods have a high recognition rate under conditions of suitable lighting and angles and clear images. However, in actual application scenarios, there are often situations that affect image quality, such as undesirable user usage environments and poor terminal device image capture capabilities, which lead to the failure of existing face recognition algorithms and reduce the accuracy and adaptability of face recognition. In order to solve the above problems, an embodiment of the present invention proposes a face recognition method based on face depth features, Figure 3 This is a schematic diagram of face recognition based on face depth features applicable to an embodiment of the present invention, which may specifically include:

[0090] 1. Image acquisition

[0091] The facial image is collected through a terminal device or mobile phone, and the collected original image is defined as l(x,y), l(x,y)=L(x,y)+R(x,y), where L(x,y) represents the illumination component of the surrounding light intensity information, R(x,y) represents the reflection component of the inherent properties of the object itself, and x and y are the horizontal and vertical coordinates of the pixel points in the image, respectively.

[0092] 2. Image Decomposition

[0093] Based on the principle of single-scale retinal cortex (Retinex), the image decomposition module is used to decompose the collected image into illumination image and reflection image. The image decomposition module first constructs the Gaussian surround function Among them, σ is the scale parameter of Gaussian surround. A smaller σ means that the scale of the Gaussian template is small, which can better preserve the edge details and increase the dynamic range, but the color cannot be preserved; conversely, the color restoration is better, but the dynamic range becomes smaller and the detail preservation is poor.

[0094] Then, the red, green and blue channels of the image are filtered using the Gaussian surround function to obtain the illumination component as the illumination image, and then the original image and the illumination component are subtracted in the logarithmic domain to obtain the reflection component as the reflection image, that is, Among them, l i (x, y) is the original image of the i-th color channel, R i (x, y) is the reflection image of the i-th color channel, L i (x, y) is the illumination image of the i-th color channel, which can be obtained by l i (x, y)*G(x, y) is calculated, that is, the original image of the i-th color channel is convolved with the Gaussian surround function.

[0095] 3. Image denoising

[0096] The decomposed illumination image and reflection image are used as the input of the denoising module. The illumination image is used as a constraint for the reflection image to suppress the noise in the reflection image and obtain a high-quality image.

[0097] The denoising module is an improved Attention Guided Residual Network (ARes-Net) model based on the U-net model. The ARes-Net model includes an input module, an encoder module, a decoder module, and an output module connected in sequence, and each module contains at least one convolutional layer. The ARes-Net model also includes a skip connection module for establishing a skip connection between each convolutional layer and at least one subsequent convolutional layer. Each module of the ARes-Net model has a residual connection with at least one subsequent module, i.e., D(x) = x + f(x), where x is the input parameter, f(x) is the output feature of the original convolutional module, and D(x) is the residual connection output. By adding residual connections, gradient vanishing can be avoided, image features can be better preserved, and network convergence can be accelerated.

[0098] Because faces in face recognition tasks are concentrated in the center of the image, excessive attention to other areas of the image can reduce the model's ability to recognize faces. Therefore, a convolutional block attention mechanism is used between the first two convolutional layers of the encoder module to enhance the model's focus on facial features in the central area. Furthermore, asymmetric convolutional layers are used in the skip connection module, decoder module, and output module to weight the center of the receptive field, reduce the model's computational load, and improve inference speed.

[0099] 4. Texture enhancement

[0100] The denoised image is used as input, and the pre-trained frequency-time feature extraction model is used to fine-tune the high-definition face image to make it more suitable for face image texture enhancement, and the enhanced image is output for the final comparison.

[0101] Among them, the pre-trained frequency-time feature extraction model is based on the Stable-Diff (stable diffusion model), which obtains historical facial images collected by the terminal at at least one historical moment, calculates the clarity of each historical facial image, and screens out images with a clarity threshold greater than the target, that is, high-quality facial images as images to be trained, and inputs the images to be trained into the model for pre-training.

[0102] 5. 3D Modeling and Face Matching

[0103] A Delaunay triangulation is used to construct a Delaunay triangulation network between the enhanced image and the user images stored in the database, and key monitoring points are annotated. Typically, 68 facial key points are selected. When evaluating deviation, due to the differences in the actual size of different facial images, to facilitate comparison of algorithm performance at the same scale, a normalization strategy for face size is adopted using interocular distance. Specifically, the distance between the eyes of the enhanced image and the baseline images of each standard user stored in the database is measured. Based on the measured distances, the enhanced image is scaled at least once to ensure that the distance between the eyes in the scaled enhanced image is consistent with the distance between the eyes in the baseline images of the matched standard users.

[0104] The normalized data is used to calculate the three-dimensional spatial distance between the key point position of the enhanced image and the real key point position (the key point position corresponding to the reference image) to identify the user or verify whether the user is the real user.

[0105] The face recognition method based on facial depth features proposed in an embodiment of the present invention decomposes the image into an illumination map and a reflection map based on the Retinex principle, and uses the reflection image as the modeling basis of the face recognition algorithm. The proposed ARes-UNet model enhances the model's capture of facial features in the image by introducing asymmetric convolution, and uses a residual structure to retain image features. During the diffusion process, the illumination image is used to suppress noise in the reflection image, thereby improving image quality. By stacking an attention-guided residual network model and a frequency-time feature extraction model, the overall image quality and texture are enhanced respectively. The diffusion model provides a more stable and controllable generation process that is easy to understand and use, has stronger anti-interference capabilities, and is not easily affected by input noise. The enhanced facial image is then three-dimensionally modeled using Delaunay three-piece decomposition, and the relationship between feature points in three-dimensional space is calculated, which has a higher recognition accuracy than calculations only in two-dimensional space.

[0106] Example 3

[0107] Figure 4 This is a structural diagram of a face recognition device based on face depth features provided by the third embodiment of the present invention. Figure 4 As shown, the apparatus includes: a decomposition module 410, a denoising module 420, an enhancement module 430, a labeling module 440 and a recognition module 450, wherein:

[0108] Decomposition module 410, configured to obtain a face image to be recognized collected by a terminal, and decompose the face image into an illumination image and a reflection image based on the single-scale retinal cortex principle;

[0109] a denoising module 420 for inputting the illumination image and the reflection image into an attention-guided residual network model to obtain a denoised image;

[0110] An enhancement module 430 is configured to input the denoised image into a pre-trained frequency-time feature extraction model to obtain an enhanced image;

[0111] a labeling module 440 for constructing a Delaunay triangulation method for the enhanced image and a reference image of at least one standard user stored in the database, respectively, and labeling key points in the enhanced image and the reference image according to the constructed Delaunay triangulation method;

[0112] The recognition module 450 is used to calculate the three-dimensional spatial distance between each key point position of the enhanced image and the corresponding key point position of the reference image, and perform face recognition based on the calculation of each spatial distance.

[0113] The technical solution of the embodiment of the present invention obtains a facial image to be recognized collected by a terminal, and decomposes the facial image into an illumination image and a reflection image based on the principle of single-scale retinal cortex, which can effectively separate the illumination and reflection components in strong or weak light environments; inputs the illumination image and the reflection image into an attention-guided residual network model to obtain a denoised image, uses the reflection image to build a face recognition model, and combines the illumination image to suppress the noise of the reflection image, thereby improving the overall image quality. The denoised image is input into a pre-trained frequency-time feature extraction model to obtain an enhanced image, thereby enhancing the image texture details; based on the Delaunay triangulation method, a Delaunay triangulation network is constructed for the enhanced image and a reference image of at least one standard user stored in a database, and key points are respectively marked in the enhanced image and the reference image based on the constructed Delaunay triangulation network; the three-dimensional spatial distance between the position of each key point of the enhanced image and the corresponding key point position of the reference image is calculated, and face recognition is performed based on the calculation of each spatial distance, thereby solving the problem of poor image feature extraction effect under complex lighting conditions and improving the accuracy, stability and adaptability of face recognition.

[0114] Based on the above embodiments, the decomposition module 410 is specifically configured to:

[0115] According to the Gaussian surround function, the three channels of the face image are filtered respectively to obtain the illumination components as the illumination image;

[0116] The face image and the illumination component are subtracted in the logarithmic domain to obtain the reflection component as the reflection image.

[0117] Based on the above embodiments, the attention-guided residual network model may include: an input module, an encoder module, a decoder module, and an output module connected in sequence, and each module includes at least one convolutional layer;

[0118] The attention-guided residual network model may also include a jump connection module for establishing a jump connection between each convolutional layer and at least one subsequent convolutional layer, and each module of the attention-guided residual network model has a residual connection with at least one subsequent module.

[0119] Based on the above embodiments, a convolutional block attention module attention mechanism is used between the first two convolutional layers of the encoder module, and asymmetric convolutional layers are used in the skip connection module, the decoder module and the output module.

[0120] Furthermore, based on the above embodiments, the face recognition device based on face depth features may further include: a training image module and a pre-training module, wherein:

[0121] A training image module is configured to obtain historical facial images collected by the terminal at at least one historical moment, calculate the clarity of each historical facial image, and select images with clarity greater than a target clarity threshold as training images before inputting the denoised image into a pre-trained frequency-time feature extraction model to obtain an enhanced image;

[0122] The pre-training module is used to input the image to be trained into the stable diffusion model for pre-training to obtain a pre-trained frequency-time feature extraction model.

[0123] Based on the above embodiments, the marking module 440 is specifically configured to:

[0124] In the scenario of target user identity authentication, Delaunay triangulation method is used to construct Delaunay triangulation networks for the enhanced image and the reference image of the target user stored in the database.

[0125] Based on the above embodiments, the identification module 450 is specifically configured to:

[0126] Count the number of key points whose spatial distance is greater than or equal to the preset distance threshold;

[0127] If the number of key points is greater than or equal to the preset number threshold, it is determined that the identity authentication of the target user has failed; otherwise, it is determined that the identity authentication of the target user has passed.

[0128] Furthermore, based on the above embodiments, the face recognition device based on face depth features may further include: a measurement module and a scaling module, wherein:

[0129] a measurement module for measuring the distance between the enhanced image and the human eyes in the reference image of each standard user stored in the database before constructing Delaunay triangulation networks for the enhanced image and the reference image of at least one standard user stored in the database based on the Delaunay triangulation method;

[0130] The scaling module is used to perform at least one scaling process on the enhanced image according to the measured distances, so that the distance between the human eyes in the enhanced image after the scaling process is consistent with the distance between the human eyes in the baseline image of the matched standard user.

[0131] The face recognition device based on face depth features provided in the embodiment of the present invention can execute the face recognition method based on face depth features provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0132] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0133] Example 4

[0134] Figure 5 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0135] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0136] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0137] The processor 11 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the face recognition method based on facial depth features, that is,

[0138] Obtain the face image to be recognized collected by the terminal, and decompose the face image into an illumination image and a reflection image based on the single-scale retinal cortex principle;

[0139] The illumination image and the reflection image are input into the attention-guided residual network model to obtain the denoised image;

[0140] The denoised image is input into the pre-trained frequency-time feature extraction model to obtain the enhanced image;

[0141] Constructing Delaunay triangulation methods for the enhanced image and at least one standard user's reference image stored in the database, respectively, and marking key points in the enhanced image and the reference image according to the constructed Delaunay triangulation methods;

[0142] The three-dimensional spatial distance between each key point position of the enhanced image and the corresponding key point position of the reference image is calculated, and face recognition is performed based on the calculation of each spatial distance.

[0143] In some embodiments, the face recognition method based on facial depth features may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the face recognition method based on facial depth features described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the face recognition method based on facial depth features in any other appropriate manner (e.g., by means of firmware).

[0144] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0145] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0146] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0147] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0148] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0149] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within a cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0150] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0151] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A face recognition method based on face depth features, characterized in that: include: Obtain the face image to be recognized collected by the terminal, and decompose the face image into an illumination image and a reflection image based on the single-scale retinal cortex principle; The illumination image and the reflection image are input into the attention-guided residual network model to obtain the denoised image; The denoised image is input into the pre-trained frequency-time feature extraction model to obtain the enhanced image; Constructing Delaunay triangulation methods for the enhanced image and at least one standard user's reference image stored in the database, respectively, and marking key points in the enhanced image and the reference image according to the constructed Delaunay triangulation methods; The three-dimensional spatial distance between each key point position of the enhanced image and the corresponding key point position of the reference image is calculated, and face recognition is performed based on the calculation of each spatial distance.

2. The method according to claim 1, characterized in that Based on the principle of single-scale retinal cortex, the face image is decomposed into illumination image and reflection image, including: According to the Gaussian surround function, the three channels of the face image are filtered respectively to obtain the illumination components as the illumination image; The face image and the illumination component are subtracted in the logarithmic domain to obtain the reflection component as the reflection image.

3. The method according to claim 1, characterized in that The attention-guided residual network model includes an input module, an encoder module, a decoder module, and an output module connected in sequence, and each module contains at least one convolutional layer; The attention-guided residual network model also includes: a jump connection module for establishing a jump connection between each convolutional layer and at least one subsequent convolutional layer, and each module of the attention-guided residual network model has a residual connection with at least one subsequent module.

4. The method according to claim 3, characterized in that The convolutional block attention module attention mechanism is used between the first two convolutional layers of the encoder module, and asymmetric convolutional layers are used in the skip connection module, decoder module and output module.

5. The method according to claim 1, characterized in that Before the denoised image is input into the pre-trained frequency-time feature extraction model to obtain the enhanced image, it also includes: Obtain historical facial images collected by the terminal at at least one historical moment, calculate the clarity of each historical facial image, and select images with a clarity greater than a target threshold as images to be trained; The image to be trained is input into the stable diffusion model for pre-training to obtain a pre-trained frequency-time feature extraction model.

6. The method according to claims 1-5, characterized in that Constructing Delaunay triangulation networks for the enhanced image and a reference image of at least one standard user stored in a database based on the Delaunay triangulation method, including: In the scenario of target user identity authentication, Delaunay triangulation method is used to construct Delaunay triangulation networks for the enhanced image and the reference image of the target user stored in the database. Face recognition is performed based on spatial distance calculations, including: Count the number of key points whose spatial distance is greater than or equal to the preset distance threshold; If the number of key points is greater than or equal to the preset number threshold, it is determined that the identity authentication of the target user has failed; otherwise, it is determined that the identity authentication of the target user has passed.

7. The method according to any one of claims 1 to 5, characterized in that Before constructing Delaunay triangulation networks for the enhanced image and the reference image of at least one standard user stored in the database based on the Delaunay triangulation method, the method further includes: Measuring the distance between the enhanced image and the human eyes in the reference image of each standard user stored in the database; According to the measured distances, the enhanced image is scaled at least once, so that the distance between human eyes in the enhanced image after the scaling process is consistent with the distance between human eyes in the reference image of the matched standard user.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the face recognition method based on face depth features according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the face recognition method based on face depth features according to any one of claims 1 to 7 when executed.

10. A computer program product, characterized in that The computer program product includes a computer program, which, when executed by a processor, implements the face recognition method based on face depth features according to any one of claims 1 to 7.

Citation Information

Cited By

  • Dynamic face recognition method in ultra-low illumination environment based on event camera

    CN121640551A