Coding and decoding three-dimensional face recognition method and system based on frequency characteristic assistance
The frequency feature-assisted encoding and decoding 3D face recognition method uses an encoder and decoder combined with an attention mechanism to dynamically select frequency features and local features, which solves the problem of insufficient preservation of texture details and edge structures in depth map face denoising, and improves the processing efficiency and recognition accuracy of depth maps.
Patent Information
- Application Number
- CN202510936677.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-11-28
AI Technical Summary
Existing depth map face denoising methods fail to preserve sufficient texture details and edge structures, resulting in excessive smoothing of subtle texture features in the face region and blurring or distortion of important edge structures, which affects the accuracy of subsequent applications.
A frequency feature-assisted encoding and decoding 3D face recognition method is adopted. By combining the encoder and decoder with an attention mechanism, frequency features and local features are dynamically selected to reduce redundant information and improve the feature refinement effect.
It effectively preserves key geometric information of skin texture and facial contours, reduces computational redundancy, and improves the processing efficiency and recognition accuracy of face depth maps.
Smart Images

Figure CN121033906A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of face recognition, and particularly relates to a coding and decoding three-dimensional face recognition method and system based on frequency feature assistance. BACKGROUND
[0002] Face depth map denoising technology shows broad application prospects in the fields of mobile terminal face unlocking, medical image analysis and digital content creation. Current face depth map denoising methods can be mainly divided into two categories: fusion-based methods and learning-based methods.
[0003] Fusion-based methods mainly generate high-quality depth maps by integrating multiple consecutive low-quality depth frames. This kind of method usually adopts the technical route of spatial-temporal filtering or multi-frame registration fusion, which can effectively reduce random noise and fill data gaps. However, this kind of method has three main defects: first, the algorithm implementation complexity is high, which requires accurate inter-frame registration and complex optimization calculation; second, the input data requirements are strict, which must rely on consecutive depth image sequences, which is often difficult to guarantee in actual application; finally, the processing delay is significant, since multiple frame accumulation and iterative optimization are required, it is difficult to meet the real-time requirement of high application scenarios. In the fusion-based method, since multiple consecutive frames of data are required for spatial-temporal optimization, the high similarity between adjacent frames leads to a large amount of repeated calculation, especially in static or low dynamic scenes, redundant information not only increases the storage burden, but also reduces the processing efficiency. In the fusion process of high-level features and low-level features of deep neural networks, there is also information redundancy. In addition, high parameter quantity and complex architecture (such as deep CNN or Transformer) will introduce huge calculation overhead, when processing high-resolution depth maps, pixel-by-pixel convolution or self-attention operation leads to a sharp increase in memory occupation and inference delay.
[0004] The problem of information redundancy is particularly prominent in fusion-based methods. These methods often rely on continuous multi-frame data for joint optimization in space and time. However, due to the high temporal correlation between adjacent video frames, especially in static backgrounds or low dynamic scenes, there is a lot of repeated feature extraction and calculation. This redundancy not only significantly increases the memory storage burden, but also causes about 30-50% of invalid calculation, which seriously reduces the system processing efficiency. At the feature fusion level, the cross-layer fusion process of high-level semantic features and low-level detail features in deep neural networks often leads to a large number of similar or repeated feature representations between different levels due to the lack of effective feature selection mechanism. This feature space redundancy further reduces the representation efficiency of the network. In terms of model architecture, although the current mainstream deep CNN or Transformer structure has strong representation ability, its high parameter quantity (usually reaching millions or even hundreds of millions) will introduce huge computational overhead. Especially when processing high-resolution depth maps, pixel-by-pixel convolution operations or global self-attention mechanisms will produce O(N^2) computational complexity, which not only leads to exponential growth of memory usage, but also causes significant inference delay.
[0005] Learning-based methods refer to establishing an end-to-end mapping relationship between low-quality depth maps and high-quality depth maps by designing deep neural networks, and realizing image quality optimization in a data-driven manner. These methods usually use convolutional neural networks (CNN) or Transformer architecture to learn the noise distribution and geometric features in depth images through a large amount of training data. However, there are still several key technical bottlenecks in this method. First, existing models still lack sufficient modeling of noise characteristics in three-dimensional face depth images. The noise produced by depth sensing devices has significant spatial correlation and distance dependence characteristics. For example, the multi-path interference noise of time-of-flight (ToF) cameras presents a complex spatial distribution pattern. The current network architecture has limitations in the mathematical representation of noise formation mechanisms, resulting in two typical technical defects: one is the "over-smoothing phenomenon", where the model misjudges subtle facial geometric features (such as nasal labial groove morphology, eye corner wrinkles, etc.) as noise components and over-inhibits them; the second is the "under-de-noising problem", where the removal effect of certain noise patterns (such as stripe noise, speckle noise, etc.) is not good.
[0006] In summary, existing denoising algorithms for depth maps generally face the challenge of insufficient preservation of texture details and edge structures. On the one hand, subtle texture features of the facial region (such as skin pores and wrinkles) are easily over-smoothed during filtering, causing the denoised depth map to lose the unique microscopic geometric features of a real face. On the other hand, important edge structures (such as facial contours and hairstyle boundaries) are prone to blurring or distortion during noise suppression, resulting in edge blunting or step-like artifacts. This loss of texture and edge information severely affects the accuracy of subsequent applications. Summary of the Invention
[0007] The purpose of this invention is to provide a frequency feature-assisted encoding and decoding three-dimensional face recognition method and system to solve at least one of the technical problems existing in the background art.
[0008] To achieve the above objectives, the present invention adopts the following technical solution:
[0009] In a first aspect, the present invention provides a frequency feature-assisted encoding and decoding method for three-dimensional face recognition, comprising:
[0010] Acquire the face image to be identified;
[0011] A pre-trained recognition model is used to process the acquired face image to obtain a 3D face depth map. The recognition model includes a feature extraction module, an encoder, a decoder, and an image generation module. The feature extraction module extracts 3D face depth map features and normal map features from the face image. The encoder obtains frequency features and local features based on the depth map and normal map features, establishes the relationship between frequency features and local features using an attention mechanism, and dynamically selects frequency features with high information content to extract frequency information from the 3D face data, thus obtaining effective face features. The decoder integrates the frequency features and local features to reduce redundant information from similar features during the fusion process and effectively integrates contextual information. The image generation module obtains the final refined 3D face depth map based on the denoised and fused features.
[0012] As a further limitation of the first aspect of the present invention, the encoder includes four stacked frequency feature-assisted denoising units, each including a main branch and an auxiliary branch; the main branch is used to extract local features, and the auxiliary branch is used to extract frequency features; wherein, the local features are obtained by downsampling using three cascaded convolutional blocks, each convolutional block containing a convolutional layer, a batch normalization layer, and a ReLU activation function; the convolutional layer in the second convolutional block is a depthwise separable convolution; the auxiliary branch employs learnable convolutional operations equivalent to high-pass and low-pass filters to generate high-frequency feature maps and low-frequency feature maps.
[0013] As a further limitation of the first aspect of the present invention, high-frequency features and low-frequency features are respectively used as the key K and value V of the attention mechanism, and the local features extracted from the main branch are used as the query Q of the attention mechanism to dynamically select effective frequency features, thereby obtaining high-frequency features and low-frequency features after cross-attention processing.
[0014] As a further limitation of the first aspect of the present invention, a convolutional layer and a batch normalization layer are used to fuse the high-frequency features and low-frequency features after cross-attention processing, and the fused features are added to the local features as an auxiliary denoising branch.
[0015] As a further limitation of the first aspect of the present invention, the decoder includes four decoding feature enhancement units, each of which contains two branches. One branch receives the feature output from the previous decoding feature enhancement unit and then processes the feature using 3×3 convolution and reshape operations to obtain upsampled local features. The other branch receives the features output from the frequency feature auxiliary denoising unit in the corresponding order in the encoder, and fuses the features obtained through depthwise separable convolution and batch normalization to obtain the final features.
[0016] As a further limitation of the first aspect of the present invention, the calculation process of the second decoding feature enhancement unit is as follows:
[0017]
[0018] Wherein, Conv(·) represents the convolution operation, Reshape(·) represents the feature map size transformation operation, BN(·) represents the batch normalization operation, EFD(·) represents the encoding and decoding feature fusion module, and Relu(·) represents the Relu activation function.
[0019] Secondly, the present invention provides a frequency feature-assisted encoding and decoding three-dimensional face recognition system, comprising:
[0020] The acquisition module is used to acquire the face image to be identified;
[0021] The recognition module processes the acquired face image to be recognized using a pre-trained recognition model to obtain a 3D face depth map. The recognition model includes a feature extraction module, an encoder, a decoder, and an image generation module. The feature extraction module extracts 3D face depth map features and normal map features from the face image. The encoder obtains frequency features and local features based on the depth map and normal map features, establishes the relationship between frequency features and local features using an attention mechanism, and dynamically selects frequency features with high information content to extract frequency information from the 3D face data, thus obtaining effective face features. The decoder integrates the frequency features and local features to reduce redundant information from similar features during the fusion process and effectively integrates contextual information. The image generation module obtains the final refined 3D face depth map based on the noise-reduced and fused features.
[0022] Thirdly, the present invention provides a non-transitory computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the frequency feature-assisted encoding and decoding three-dimensional face recognition method as described in the first aspect.
[0023] Fourthly, the present invention provides a computer device including a memory and a processor, wherein the processor and the memory communicate with each other, the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the frequency feature-assisted encoding and decoding three-dimensional face recognition method as described in the first aspect.
[0024] Fifthly, the present invention provides an electronic device, comprising: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to cause the electronic device to execute instructions for implementing the frequency feature-assisted encoding and decoding three-dimensional face recognition method as described in the first aspect.
[0025] Face Depth Map Denoising: This technique is specifically designed to remove noise and enhance details in facial regions of 3D depth images. Since data acquired by depth sensors (such as Time-of-Flight and structured light) often contains noise, holes, and blurred edges, this technique uses deep learning (such as 3D CNNs and GANs) or traditional filtering methods to remove noise while preserving key facial geometric features (such as facial contours and surface curvature). Its core challenge lies in balancing denoising effectiveness with geometric fidelity, and it is primarily applied in 3D face reconstruction, biometrics (such as Face ID), and AR / VR.
[0026] Encoder-Decoder Network: An encoder-decoder network consists of two parts: an encoder and a decoder. The encoder compresses the input data step by step through operations such as convolution and pooling to extract high-level low-dimensional feature representations; the decoder reconstructs or generates the target output from these features through upsampling, transposed convolution, and other methods.
[0027] Frequency features: Frequency features refer to the frequency domain representation extracted from data such as images and speech through methods such as Fourier transform (FFT) and wavelet transform. They can effectively capture the periodicity, texture and structural information of signals.
[0028] Feature fusion refers to the effective integration of features from different levels, sources, or modalities to improve the model's expressive power and task performance. Its core objective is to enhance the discriminative power and robustness of features through the combination of complementary information.
[0029] The beneficial effects of this invention are as follows: Utilizing an encoder-decoder network and employing a multi-scale encoder structure, hierarchical features of the face depth map are extracted during progressive downsampling, effectively preserving key geometric information including skin texture and facial contours; in the decoding stage, the network selectively fuses features from different levels through skip connections and attention mechanisms, significantly reducing redundancy and improving the efficiency of the face depth map; the attention mechanism dynamically selects frequency features, choosing those with higher information content to enhance denoising capabilities; and the effective fusion of encoded and decoded features improves feature refinement.
[0030] The advantages of additional aspects of the invention will be set forth more clearly in the following description or will be learned by practice of the invention. Attached Figure Description
[0031] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 This is a flowchart illustrating the encoding and decoding process framework for 3D face recognition based on frequency feature assistance, as described in an embodiment of the present invention.
[0033] Figure 2 This is a flowchart illustrating the frequency feature extraction process described in an embodiment of the present invention.
[0034] Figure 3This is a functional flowchart of the frequency feature and local feature relationship establishment module and the cross-attention module described in the embodiments of the present invention.
[0035] Figure 4 This is a flowchart illustrating the encoding and decoding feature fusion process according to an embodiment of the present invention.
[0036] Figure 5 This is a comparison chart showing the noise reduction effects. Detailed Implementation
[0037] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0038] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0039] It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as here.
[0040] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, and / or groups thereof.
[0041] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.
[0042] To facilitate understanding of the present invention, the present invention will be further explained and described below with reference to the accompanying drawings and specific embodiments. However, the specific embodiments do not constitute a limitation on the embodiments of the present invention.
[0043] Those skilled in the art should understand that the accompanying drawings are merely schematic diagrams of embodiments, and the components in the drawings are not necessarily essential for implementing the present invention.
[0044] Example 1
[0045] In this embodiment 1, a frequency feature-assisted encoding and decoding 3D face recognition system is first provided, including: an acquisition module for acquiring a face image to be recognized; and a recognition module for processing the acquired face image to be recognized using a pre-trained recognition model to obtain a 3D face depth map. The recognition model includes: a feature extraction module, an encoder, a decoder, and an image generation module. The feature extraction module is used to extract the 3D face depth map features and normal map features of the face image, respectively. The encoder is used to acquire frequency features and local features based on the depth map features and normal map features, establish the relationship between frequency features and local features using an attention mechanism, and dynamically select frequency features with high information content to extract the frequency information of the 3D face data, thus obtaining effective face features. The decoder is used to integrate the frequency features and local features to reduce redundant information of similar features during the fusion process and effectively integrate contextual information. The image generation module is used to obtain the final refined 3D face depth map based on the noise-reduced and fused features.
[0046] In this embodiment, the above-described system is used to implement a frequency feature-assisted encoding and decoding three-dimensional face recognition method, including: acquiring a face image to be recognized; and processing the acquired face image to be recognized using a pre-trained recognition model to obtain a three-dimensional face depth map.
[0047] The encoder includes four stacked frequency feature-assisted denoising units, each comprising a main branch and auxiliary branches. The main branch extracts local features, while the auxiliary branches extract frequency features. The local features are obtained by downsampling using three cascaded convolutional blocks, each containing a convolutional layer, a batch normalization layer, and a ReLU activation function. The convolutional layer in the second convolutional block is a depthwise separable convolution. The auxiliary branches employ learnable convolutional operations equivalent to high-pass and low-pass filters to generate high-frequency and low-frequency feature maps.
[0048] Specifically, the frequency feature-assisted enhancement unit (FRAME) introduces a frequency feature branch to assist the local feature branch, thereby achieving more effective facial feature extraction. Taking the first FRAME unit in the encoder as an example, the features are first input into the main branch and the auxiliary branch, respectively. The main branch is used to extract local features, while the auxiliary branch is used to extract frequency features, ultimately yielding the facial features. and in, It was obtained by downsampling using three cascaded convolutional blocks, each containing a convolutional layer, a batch normalization layer, and a ReLU activation function. The convolutional layer in the second convolutional block is a depthwise separable convolution (DSConv). The calculation process is as follows:
[0049]
[0050] Here, Convblock(·) and DSConvblock(·) represent ordinary convolutional blocks and depthwise separable convolutional blocks, respectively, both of which contain convolution operations, batch normalization layers, and ReLU activation functions.
[0051] The auxiliary branch is used to extract the frequency information of the input features. This branch employs a learnable convolution operation equivalent to a high-pass filter and a low-pass filter to generate a high-frequency feature map F. h and low-frequency feature map F l Specifically, suppose the given feature map is Where H×W represents the spatial dimension and C represents the number of channels, the principle of the low-frequency filter is as follows:
[0052] X = Conv(BN(F) in ))
[0053] Low filter =Softmax(BN(Conv(GAP(X))))
[0054] Here, Conv(·), BN(·), and GAP(·) represent convolution, batch normalization, and global average pooling, respectively. Softmax(·) refers to applying the softmax function to the filter.
[0055] The input feature map X is expanded and multiplied element-wise by the low-frequency filter. The purpose is to extract low-frequency information through the filter. The principles of low-frequency feature maps and high-frequency feature maps are as follows:
[0056]
[0057] F h =XF l
[0058] Where p and q ∈ {-1, 0, 1}, c is the index of the feature map channel, and h and w represent spatial coordinates.
[0059] The obtained frequency characteristics and local features After inputting the frequency feature and local feature relationship establishment module (FLRC Module), the output of the FADE module is obtained. The entire calculation process is shown in the following formula:
[0060]
[0061] Where FADE(·) represents the module for establishing the relationship between frequency features and local features, and Relu(·) represents the Relu activation function.
[0062] In this embodiment, high-frequency features and low-frequency features are used as the key K and value V of the attention mechanism, respectively. Local features extracted from the main branch are used as the query Q of the attention mechanism to dynamically select effective frequency features, resulting in high-frequency and low-frequency features after cross-attention processing. Convolutional layers and batch normalization layers are used to fuse the high-frequency and low-frequency features after cross-attention processing, and the fused features are added to the local features as an auxiliary denoising branch.
[0063] Specifically, the relationship between frequency features and local features established in this embodiment allows for the extraction of more detailed information and optimization of local features within the visible area of the face. An attention mechanism is used to dynamically select frequency features with high information content, thereby enhancing the expressive power of local features. The high-frequency features F extracted by the frequency feature-assisted enhancement unit are then used... h and low-frequency characteristics F l Input the cross-attention module separately, using it as the key K and value V of the attention mechanism. Extract local features F from the main branch. e_l The query Q serves as the attention mechanism to dynamically select effective frequency features. The computation process of the attention mechanism is shown below:
[0064] Q = W Q F local
[0065] K = W K F local
[0066] V = W V F local
[0067]
[0068] Among them, W Q W K and W V d represents the weight matrix for query, key, and value, respectively. k This represents the scaling factor. After processing by the cross-attention module, the high-frequency feature F is obtained. att_h and low-frequency characteristics Fatt_l Subsequently, convolutional layers and batch normalization layers are used to fuse the two features, and the fused features are added to the local features as an auxiliary denoising branch. The entire computation process of the FLRC module can be represented as follows:
[0069]
[0070] Where Conv(·) represents the convolution operation, BN(·) represents the batch normalization operation, and Relu(·) represents the Relu activation function.
[0071] The decoder includes four decoding feature enhancement units. Each decoding feature enhancement unit contains two branches. One branch receives the feature output from the previous decoding feature enhancement unit and then processes the feature using 3×3 convolution and reshape operations to obtain the upsampled local feature. The other branch receives the feature output from the frequency feature auxiliary denoising unit in the corresponding order in the encoder. The feature obtained after depthwise separable convolution and batch normalization is fused to obtain the final feature.
[0072] Specifically, the goal is to recover as much original information as possible from the features compressed by the encoder, while further reducing the impact of noise. The decoder section proposed in this embodiment consists of four decoding feature enhancement units. Taking the second decoding feature enhancement unit in the encoder as an example, it includes two branches, one of which receives features output from the previous DFE module. Then, a 3×3 convolution and a reshape operation are used to process the feature to obtain the upsampled local features. The other branch receives the features output from the frequency feature-assisted denoising encoding module corresponding to the frequency features in the encoder. Features Features obtained through depthwise separable convolution and batch normalization With features The final features are obtained by fusion. Encoding-decoding feature fusion can enhance feature representations while reducing information redundancy during context information fusion. The entire computation process of the second decoding feature enhancement unit can be represented as follows:
[0073]
[0074]
[0075] Wherein, Conv(·) represents the convolution operation, Reshape(·) represents the feature map size transformation operation, BN(·) represents the batch normalization operation, EFD(·) represents the encoding and decoding feature fusion module, and Relu(·) represents the Relu activation function.
[0076] Example 2
[0077] In this embodiment 2, a frequency-feature-assisted encoder-decoder network for 3D face denoising (FAEDNet) is provided. This network employs an encoder-decoder architecture, consisting of an encoder and a decoder. The encoder portion stacks four frequency-feature-assisted denoising encoding modules (FADEModules). These modules utilize a dual-branch architecture, combining frequency features obtained by the frequency feature extraction module (FFEModule) with local features obtained by convolutional blocks. A novel frequency-local feature relationship construction module (FLRCModule) is also embedded. This module dynamically selects information-rich frequency features through an attention mechanism, thereby assisting in local feature denoising. The decoder section also stacks four Decode Feature Enhancement Modules (DFE Modules). These modules integrate encoded and decoded features through an introduced Encoder-Decoder Feature Fusion Module (EDF Module) to reduce redundant information from similar features during fusion and effectively integrate contextual information. Experiments on the Bosphorus dataset demonstrate that the proposed network can effectively denoise and refine low-quality 3D face depth maps, generating high-quality 3D face depth maps, thereby improving the performance of subsequent 3D face recognition. The encoder section contains four frequency feature-assisted denoising encoding modules. Each module enhances feature representation by fusing frequency features as auxiliary branches with local features. The frequency feature and local feature relationship establishment module uses an attention mechanism to dynamically select frequency features, choosing those with higher information content to enhance denoising capabilities. The decoder feature enhancement module, through the introduction of the EDF Module, effectively fuses encoded and decoded features, thereby reducing feature redundancy and improving feature refinement.
[0078] The frequency feature-assisted encoding and decoding 3D face denoising model provided in this embodiment can be applied to subsequent 3D face recognition tasks, such as recognizing the current face in video surveillance; and recognizing the face in face payment to confirm identity information.
[0079] In this embodiment, the module structure diagram of the frequency feature-assisted encoding and decoding 3D face denoising network (FAEDNet) is as follows: Figure 1 As shown, FAEDNet consists of a feature extraction module, an encoder, a decoder, and an image generation module. First, the feature extraction module extracts 3D face depth map features F through two convolutional blocks. n and normal map feature F d The two features are then added together and input into an encoder, which stacks four frequency feature-assisted denoising modules (FADE Modules). The FADE Modules obtain the frequency information of the 3D face data through the Frequency Feature Extraction Module (FFE Module), and input it along with local features into the Frequency Feature and Local Feature Relationship Establishment Module (FLRC Module) to obtain the features. Then, this information is input into four Decoding Feature Enhancement (DFE) modules, each receiving the output of the previous DFE module as another input. This allows us to obtain the denoised features. Finally, the features The input image generation module obtains the final refined 3D face depth map.
[0080] The structure of the Frequency Feature Assisted Enhancement (FADE) module proposed in this embodiment is as follows: Figure 2 As shown. To compensate for the lack of spatial domain information in local features, a frequency feature-assisted enhancement module was designed. This module introduces a frequency feature branch to assist the local feature branch, thereby achieving more effective facial feature extraction. Taking the first FADE module in the encoder as an example, the features are first input into the main branch and the auxiliary branch, respectively. The main branch is used to extract local features, while the auxiliary branch is used to extract frequency features, ultimately obtaining the facial features. and in, It was obtained by downsampling using three cascaded convolutional blocks, each containing a convolutional layer, a batch normalization layer, and a ReLU activation function. The convolutional layer in the second convolutional block is a depthwise separable convolution (DSConv). The calculation process is as follows:
[0081]
[0082] Here, Convblock(·) and DSConvblock(·) represent ordinary convolutional blocks and depthwise separable convolutional blocks, respectively, both of which contain convolution operations, batch normalization layers, and ReLU activation functions.
[0083] The auxiliary branch is used to extract the frequency information of the input features. The module structure is as follows: Figure 3 As shown, this module employs a learnable convolution operation equivalent to a high-pass filter and a low-pass filter to generate a high-frequency feature map F. h and low-frequency feature map F l Specifically, suppose the given feature map is Where H×W represents the spatial dimension and C represents the number of channels, the principle of the low-frequency filter is as follows:
[0084] X = Conv(BN(F) in ))
[0085] Low filter =Softmax(BN(Conv(GAP(X))))
[0086] Here, Conv(·), BN(·), and GAP(·) represent convolution, batch normalization, and global average pooling, respectively. Softmax(·) refers to applying the softmax function to the filter.
[0087] The input feature map X is expanded and multiplied element-wise by the low-frequency filter. The purpose is to extract low-frequency information through the filter. The principles of the low-frequency feature map and the high-frequency feature map are shown in the following formulas:
[0088]
[0089] F h =XF l
[0090] Where p and q ∈ {-1, 0, 1}, c is the index of the feature map channel, and h and w represent spatial coordinates.
[0091] The obtained frequency characteristics and local features After inputting the frequency feature and local feature relationship establishment module (FLRC Module), the output of the FADE module is obtained. The entire calculation process is shown in the following formula:
[0092]
[0093] Where FADE(·) represents the module for establishing the relationship between frequency features and local features, and Relu(·) represents the Relu activation function.
[0094] This embodiment proposes a Frequency Feature and Local Feature Relationship Establishment (FLRC) module, which can extract more detailed information and optimize local features in the visible area of the face. This module uses an attention mechanism to dynamically select frequency features with high information content, thereby enhancing the expressive power of local features. Figure 4 The specific structure of the internal modules in the FLRC module is presented. This module extracts the high-frequency features F from the FADE module. h and low-frequency characteristics F l Input the cross-attention module separately, using it as the key K and value V of the attention mechanism. Extract local features F from the main branch. e_l The query Q serves as the attention mechanism to dynamically select effective frequency features. The computation process of the attention mechanism is shown below:
[0095] Q = W Q F local
[0096] K = W K F local
[0097] V = W V F local
[0098]
[0099] Among them, W Q W K and W V d represents the weight matrix for query, key, and value, respectively. k This represents the scaling factor. After processing by the cross-attention module, the high-frequency feature F is obtained. att_h and low-frequency characteristics F att_l Subsequently, convolutional layers and batch normalization layers are used to fuse the two features, and the fused features are added to the local features as an auxiliary denoising branch. The entire computation process of the FLRC module can be represented as follows:
[0100]
[0101] Where Conv(·) represents the convolution operation, BN(·) represents the batch normalization operation, and Relu(·) represents the Relu activation function.
[0102] This embodiment proposes a Decoding Feature Enhancement (DFE) module, which can recover as much original information as possible from the features compressed by the encoder, while further reducing the impact of noise. Specifically, the decoder part of the proposed FAEDNet model consists of four DFE modules. Taking the second DFE module in the encoder as an example, this module contains two branches, one of which receives features output from the previous DFE module. Then, a 3×3 convolution and a reshape operation are used to process the feature to obtain the upsampled local features. The other branch receives the features output from the frequency feature-assisted denoising encoding module corresponding to the frequency features in the encoder. Features Features obtained through depthwise separable convolution and batch normalization With features The final features are obtained by fusion. The Encoder-Decoder Feature Fusion Module (EDF Module) can enhance feature representations while reducing information redundancy during context information fusion. The entire computation process of the second DFE module can be represented as follows:
[0103]
[0104] Wherein, Conv(·) represents the convolution operation, Reshape(·) represents the feature map size transformation operation, BN(·) represents the batch normalization operation, EFD(·) represents the encoding and decoding feature fusion module, and Relu(·) represents the Relu activation function.
[0105] In summary, this embodiment utilizes an encoder-decoder network with a multi-scale encoder structure to extract hierarchical features from the face depth map during progressive downsampling, effectively preserving key geometric information including skin texture and facial contours. During the decoding stage, the network selectively fuses features from different levels through skip connections and an attention mechanism, significantly reducing information redundancy issues in traditional multi-scale fusion methods and improving the efficiency of face depth map extraction. The proposed frequency feature and local feature relationship establishment module uses an attention mechanism to dynamically select frequency features, choosing those with higher information content to enhance denoising capabilities. The decoder section includes four decoding feature enhancement modules. By introducing an encoder-decoder feature fusion module, encoded and decoded features are effectively fused, improving feature refinement.
[0106] like Figure 5The image shows a comparison of the denoising performance of FAEDNet on the Bosphorus dataset. The first and third rows show the unprocessed raw normal maps and depth maps, respectively, which suffer from significant detail loss due to noise interference. In contrast, the second and fourth rows show the normal maps and depth maps processed by FAEDNet, which significantly improve image quality.
[0107] Example 3
[0108] This embodiment 3 provides a non-transitory computer-readable storage medium for storing computer instructions. When executed by a processor, the computer instructions implement the frequency feature-assisted encoding and decoding three-dimensional face recognition method described above. The method includes:
[0109] Acquire the face image to be identified;
[0110] A pre-trained recognition model is used to process the acquired face image to obtain a 3D face depth map. The recognition model includes a feature extraction module, an encoder, a decoder, and an image generation module. The feature extraction module extracts 3D face depth map features and normal map features from the face image. The encoder obtains frequency features and local features based on the depth map and normal map features, establishes the relationship between frequency features and local features using an attention mechanism, and dynamically selects frequency features with high information content to extract frequency information from the 3D face data, thus obtaining effective face features. The decoder integrates the frequency features and local features to reduce redundant information from similar features during the fusion process and effectively integrates contextual information. The image generation module obtains the final refined 3D face depth map based on the denoised and fused features.
[0111] Example 4
[0112] This embodiment 4 provides a computer device, including a memory and a processor, wherein the processor and the memory communicate with each other, and the memory stores program instructions executable by the processor. The processor calls the program instructions to execute the frequency feature-assisted encoding and decoding three-dimensional face recognition method described above, the method including:
[0113] Acquire the face image to be identified;
[0114] A pre-trained recognition model is used to process the acquired face image to obtain a 3D face depth map. The recognition model includes a feature extraction module, an encoder, a decoder, and an image generation module. The feature extraction module extracts 3D face depth map features and normal map features from the face image. The encoder obtains frequency features and local features based on the depth map and normal map features, establishes the relationship between frequency features and local features using an attention mechanism, and dynamically selects frequency features with high information content to extract frequency information from the 3D face data, thus obtaining effective face features. The decoder integrates the frequency features and local features to reduce redundant information from similar features during the fusion process and effectively integrates contextual information. The image generation module obtains the final refined 3D face depth map based on the denoised and fused features.
[0115] Example 5
[0116] This embodiment 5 provides an electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device is running, the processor executes the computer program stored in the memory to cause the electronic device to execute instructions for implementing the frequency feature-assisted encoding and decoding three-dimensional face recognition method described above. The method includes:
[0117] Acquire the face image to be identified;
[0118] A pre-trained recognition model is used to process the acquired face image to obtain a 3D face depth map. The recognition model includes a feature extraction module, an encoder, a decoder, and an image generation module. The feature extraction module extracts 3D face depth map features and normal map features from the face image. The encoder obtains frequency features and local features based on the depth map and normal map features, establishes the relationship between frequency features and local features using an attention mechanism, and dynamically selects frequency features with high information content to extract frequency information from the 3D face data, thus obtaining effective face features. The decoder integrates the frequency features and local features to reduce redundant information from similar features during the fusion process and effectively integrates contextual information. The image generation module obtains the final refined 3D face depth map based on the denoised and fused features.
[0119] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0120] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0121] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0122] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment, whereby a series of operational steps are performed to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0123] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that, based on the technical solutions disclosed in the present invention, various modifications or variations that can be made by those skilled in the art without creative effort should be included within the scope of protection of the present invention.
Claims
1. A frequency-feature-assisted encoding and decoding method for 3D face recognition, characterized in that, include: Acquire the face image to be identified; A pre-trained recognition model is used to process the acquired face image to obtain a 3D face depth map. The recognition model includes a feature extraction module, an encoder, a decoder, and an image generation module. The feature extraction module extracts 3D face depth map features and normal map features from the face image. The encoder obtains frequency features and local features based on the depth map and normal map features, establishes the relationship between frequency features and local features using an attention mechanism, and dynamically selects frequency features with high information content to extract frequency information from the 3D face data, thus obtaining effective face features. The decoder integrates the frequency features and local features to reduce redundant information from similar features during the fusion process and effectively integrates contextual information. The image generation module obtains the final refined 3D face depth map based on the denoised and fused features.
2. The frequency feature-assisted encoding and decoding three-dimensional face recognition method according to claim 1, characterized in that, The encoder includes four stacked frequency feature-assisted denoising units, each comprising a main branch and auxiliary branches. The main branch extracts local features, while the auxiliary branches extract frequency features. The local features are obtained by downsampling using three cascaded convolutional blocks, each containing a convolutional layer, a batch normalization layer, and a ReLU activation function. The convolutional layer in the second convolutional block is a depthwise separable convolution. The auxiliary branches employ learnable convolutional operations equivalent to high-pass and low-pass filters to generate high-frequency and low-frequency feature maps.
3. The frequency feature-assisted encoding and decoding method for 3D face recognition according to claim 2, characterized in that, High-frequency features and low-frequency features are used as the key K and value V of the attention mechanism, respectively. Local features extracted from the main branch are used as the query Q of the attention mechanism to dynamically select effective frequency features, thus obtaining high-frequency features and low-frequency features after cross-attention processing.
4. The frequency feature-assisted encoding and decoding three-dimensional face recognition method according to claim 3, characterized in that, Convolutional layers and batch normalization layers are used to fuse high-frequency and low-frequency features after cross-attention processing, and the fused features are added to the local features as an auxiliary denoising branch.
5. The frequency feature-assisted encoding and decoding three-dimensional face recognition method according to claim 4, characterized in that, The decoder includes four decoding feature enhancement units. Each decoding feature enhancement unit contains two branches. One branch receives the feature output from the previous decoding feature enhancement unit and then processes the feature using 3×3 convolution and reshape operations to obtain the upsampled local feature. The other branch receives the feature output from the frequency feature auxiliary denoising unit in the corresponding order in the encoder. The feature obtained after depthwise separable convolution and batch normalization is fused to obtain the final feature.
6. The frequency feature-assisted encoding and decoding method for 3D face recognition according to claim 5, characterized in that, The calculation process of the second decoding feature enhancement unit is as follows: Wherein, Conv(·) represents the convolution operation, Reshape(·) represents the feature map size transformation operation, BN(·) represents the batch normalization operation, EFD(·) represents the encoding and decoding feature fusion module, and Relu(·) represents the Relu activation function.
7. A frequency-feature-assisted encoding and decoding 3D face recognition system, characterized in that, include: The acquisition module is used to acquire the face image to be identified; The recognition module processes the acquired face image to be recognized using a pre-trained recognition model to obtain a 3D face depth map. The recognition model includes a feature extraction module, an encoder, a decoder, and an image generation module. The feature extraction module extracts 3D face depth map features and normal map features from the face image. The encoder obtains frequency features and local features based on the depth map and normal map features, establishes the relationship between frequency features and local features using an attention mechanism, and dynamically selects frequency features with high information content to extract frequency information from the 3D face data, thus obtaining effective face features. The decoder integrates the frequency features and local features to reduce redundant information from similar features during the fusion process and effectively integrates contextual information. The image generation module obtains the final refined 3D face depth map based on the noise-reduced and fused features.
8. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, which, when executed by a processor, implement the frequency feature-assisted encoding and decoding three-dimensional face recognition method as described in any one of claims 1-6.
9. A computer device, characterized in that, The method includes a memory and a processor, the processor and the memory communicating with each other, the memory storing program instructions executable by the processor, and the processor calling the program instructions to execute the frequency feature-assisted encoding and decoding three-dimensional face recognition method as described in any one of claims 1-6.
10. An electronic device, characterized in that, include: The device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to cause the electronic device to execute instructions for implementing the frequency feature-assisted encoding and decoding three-dimensional face recognition method as described in any one of claims 1-6.
Citation Information
Cited By
Fast moving face recognition method and system based on multi-frame image enhancement
CN121214529A
A fast moving face recognition method and system based on multi-frame image enhancement
CN121214529B