Face image recognition method

By using key point detection and occlusion recognition processing, the visible and occluded areas of the face are automatically located, and deep feature extraction and scoring fusion are performed. This solves the problem of low recognition accuracy caused by occlusion and achieves high-accuracy identity recognition.

CN121330735APending Publication Date: 2026-01-13成都浩景云格科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511400084.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

In existing technologies, facial recognition accuracy is low due to facial occlusion, especially when covered by masks, helmets, or other obstructions. Traditional methods struggle to extract complete facial features, leading to a decrease in recognition accuracy.

Method used

By using key point detection and occlusion recognition, the system automatically locates the visible and occluded areas of the face, extracts deep features, supplements missing features based on a preset facial supplementation pattern, and obtains overall features through scoring fusion processing, thereby improving the ability to identify individuals.

Benefits of technology

In cases of occlusion, the accuracy of facial recognition is improved by supplementing missing features, solving the problem of low recognition accuracy caused by occlusion and achieving high-accuracy identity recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121330735A_ABST
    Figure CN121330735A_ABST
Patent Text Reader

Abstract

The invention provides a face image recognition method, and relates to the technical field of image recognition. The method comprises the following steps: acquiring a face image of a target object; carrying out key point detection processing and shielding recognition processing on the face image to obtain a face visible area and a face shielding area; and performing feature extraction on the face visual area to obtain visual features of the face visual area. On the basis of a preset face supplement mode, performing supplement prediction processing on the face shielding area, and performing supplement to obtain missing features of the face shielding area; carrying out scoring fusion processing on the visual features and the missing features to obtain overall features of the face; wherein the face supplement mode is used for performing score fusion processing on the missing features and the visual features according to preset score information. Performing face detection processing on the overall features to obtain detection result information; wherein the detection result information represents the identity label of the target object. Therefore, the technical problem that the accuracy of face recognition is low due to the fact that the face is shielded is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition technology, and in particular to a method for facial image recognition. Background Technology

[0002] Currently, with the development of artificial intelligence and computer vision, facial recognition technology has been widely applied in scenarios such as public security, access control and attendance, financial payments, airports, and subways. In traditional applications, facial recognition systems typically rely on complete facial feature information for feature extraction and comparison to achieve high-accuracy identity recognition.

[0003] However, numerous obstructions exist in real-world applications. Face masks are routinely worn in public places, while safety helmets are essential protective equipment in industrial production and construction sites. These obstructions can cover key areas of the lower and upper parts of the face, making it difficult for traditional full-face recognition methods to extract complete facial features, resulting in a significant drop in recognition accuracy.

[0004] Therefore, existing technologies have several shortcomings in recognition. For example, key features are lost due to facial occlusion, resulting in low accuracy of facial recognition. Different wearing methods, different mask or helmet models, tightness of wearing, and changes in head posture can all cause different local occlusion positions, increasing the difficulty of feature extraction and thus reducing the accuracy of facial recognition. Summary of the Invention

[0005] To address the aforementioned problems in the prior art, this invention provides a face image recognition method that solves the technical problem of low accuracy in face recognition due to facial occlusion.

[0006] In a first aspect, embodiments of this application provide a face image recognition method, including: Acquire a facial image of the target object; perform key point detection and occlusion recognition processing on the facial image to obtain the visible facial region and the occluded facial region; and extract features from the visible facial region to obtain the visible features of the visible facial region. Based on a preset facial enhancement mode, the facial occlusion area is subjected to enhancement prediction processing to obtain the missing features of the facial occlusion area; and the visible features and the missing features are fused together to obtain the overall facial features; wherein, the facial enhancement mode is used to perform scoring fusion processing on the missing features and the visible features according to preset scoring information. The overall features are subjected to facial detection processing to obtain detection result information; wherein, the detection result information represents the identity of the target object.

[0007] Secondly, embodiments of this application provide a face image recognition device, comprising: Acquisition device, used to acquire facial images of a target object; The recognition device is used to perform key point detection processing and occlusion recognition processing on the face image to obtain the visible area of ​​the face and the occluded area of ​​the face. An extraction device is used to extract features from the visible area of ​​the face to obtain the visible features of the visible area of ​​the face. A prediction device is used to perform supplementary prediction processing on the occluded facial area based on a preset facial supplementation pattern, and supplement the missing features of the occluded facial area. A fusion device is used to perform scoring fusion processing on the visible features and the missing features to obtain the overall features of the face; wherein, the facial supplementation mode is used to perform scoring fusion processing on the missing features and the visible features according to preset scoring information; A detection device is used to perform facial detection processing on the overall features to obtain detection result information; wherein, the detection result information represents the identity of the target object.

[0008] The beneficial effects of this invention are reflected in the following aspects: Key point detection and occlusion recognition are performed on the input facial image to automatically locate the visible and occluded areas of the face. The visible area includes usable local regions (such as the periorbital region, forehead region, and temple region). Subsequently, depth features are extracted from each local region within the visible area to obtain visual features. Then, based on a preset facial supplementation mode, supplementation prediction processing is performed on the occluded areas to obtain missing features. Based on the scoring information in the facial supplementation mode, the visual features and missing features are fused to obtain the overall facial features, ensuring that the output overall features maintain a high level of identity recognition capability even under occlusion. Finally, facial detection processing is performed on the overall features to generate the final detection result, i.e., the identity recognition result. Therefore, for occluded facial images, the missing features of the occluded areas can be supplemented using the facial supplementation mode, and all facial features can be obtained based on these missing features. Facial recognition is then performed on all features, greatly improving the accuracy of facial recognition and solving the technical problem of low accuracy in facial recognition due to occlusion. Attached Figure Description

[0009] Figure 1 A flowchart illustrating a face image recognition method provided in an embodiment of the present invention. Figure 1 ; Figure 2 A flowchart illustrating a face image recognition method provided in an embodiment of the present invention. Figure 2 ; Figure 3This is a schematic diagram of the structure of a face image recognition device provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of another face image recognition device provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0011] Currently, with the development of artificial intelligence and computer vision, facial recognition technology has been widely applied in scenarios such as public security, access control and attendance, financial payments, airports, and subways. In traditional applications, facial recognition systems typically rely on complete facial feature information for feature extraction and comparison to achieve high-accuracy identity recognition.

[0012] However, numerous obstructions exist in real-world applications. Face masks are routinely worn in public places, while safety helmets are essential protective equipment in industrial production and construction sites. These obstructions can cover key areas of the lower and upper parts of the face, making it difficult for traditional full-face recognition methods to extract complete facial features, resulting in a significant drop in recognition accuracy.

[0013] Therefore, there are several shortcomings in the existing technology, such as: (1) the key features are missing due to facial occlusion, which leads to low accuracy of face recognition; among them, the missing key features include: the mask covers the lower half of the face, and the features of the nose, mouth and jawline are missing; the helmet covers the forehead or hairline area, resulting in incomplete forehead and brow features. (2) Different wearing methods, different mask or helmet models, tightness of wearing and changes in head posture will cause different local occlusion positions, increasing the difficulty of feature extraction, which leads to low accuracy of face recognition. (3) Insufficient robustness of recognition: traditional whole face feature matching relies on global information and is very sensitive to occlusion or missing local information. Once the occlusion area is too large, recognition failure or misjudgment is likely to occur. (4) Real-time and multi-environment adaptability issues: in security access control and other scenarios, the system needs to respond quickly, but the efficiency and stability of feature extraction under lighting, dynamic posture and partial occlusion are difficult to guarantee.

[0014] Example 1: Figure 1 A flowchart illustrating a face image recognition method provided in an embodiment of the present invention. Figure 1 , refer to Figure 1 The method includes: S101. Obtain the face image of the target object; perform key point detection and occlusion recognition processing on the face image to obtain the visible area of ​​the face and the occluded area of ​​the face; and extract features from the visible area of ​​the face to obtain the visible features of the visible area of ​​the face.

[0015] For example, the executing entity of this embodiment can be an electronic device, a terminal device, a face image recognition method apparatus or device, or other apparatus or device capable of executing this embodiment, and there is no limitation thereto. In this embodiment, the executing entity is described as an electronic device.

[0016] First, a facial image refers to an image containing a human face. This image can be a complete face image or a face image with occlusion. An image containing a face can be a headshot with a face, an upper-body image with a face, or a full-body image with a face, etc., without limitation. The visible facial area refers to the unoccluded exposed area of ​​the face, which includes multiple localized areas. For example, multiple localized areas in the visible facial area include the corners of the eyes, brow peaks, brow ridges, forehead, etc. The areas within the visible facial area change depending on the occluded area, without limitation. The occluded facial area refers to the obscured area of ​​the face, which includes multiple localized areas. For example, multiple localized areas in the occluded facial area include the mouth, nose, cheeks, chin, etc. The areas within the occluded facial area change depending on the occluded area, without limitation.

[0017] In this step, a facial image of the target object is acquired. Taking an occluded facial image as an example, keypoint detection processing is performed on the facial image to obtain facial keypoint information. Occlusion recognition processing is then performed on the facial image, and the visible facial region and the occluded facial region are obtained based on the facial keypoint information. Then, depth feature extraction is performed on the visible facial region to obtain its visual features. These visual features may include features of the corners of the eyes, brow peaks, brow ridges, forehead features, etc. The visual features vary depending on the occluded area and are not limited thereto.

[0018] S102. Based on a preset facial supplementation mode, perform supplementation prediction processing on the occluded facial area to supplement the missing features of the occluded facial area; and perform scoring fusion processing on the visible features and the missing features to obtain the overall facial features; wherein, the facial supplementation mode is used to perform scoring fusion processing on the missing features and the visible features according to the preset scoring information.

[0019] For example, the facial enhancement mode is used to perform score fusion processing on missing features and visible features based on preset scoring information to obtain a fused overall feature. The electronic device performs enhancement prediction processing on the occluded facial area based on the preset facial enhancement mode to obtain the missing features of the occluded facial area, and determines a first uncertainty score value for each region in the occluded facial area and a second uncertainty score value for each region in the visible facial area based on the scoring information in the facial enhancement mode. Then, based on the first uncertainty score value and the second uncertainty score, the visible features and missing features are fused to obtain the overall facial feature.

[0020] S103. Perform facial detection processing on the overall features to obtain detection result information; wherein, the detection result information represents the identity of the target object.

[0021] For example, the electronic device performs facial detection processing on the overall features according to a preset facial detection method, obtains detection result information, and then determines the identity of the target object based on the detection result information. The preset facial detection method can be a liveness detection method, etc., and is not limited thereto.

[0022] The facial image recognition method provided in this application acquires a facial image of a target object; performs key point detection and occlusion recognition processing on the facial image to obtain the visible facial region and the occluded facial region; and extracts features from the visible facial region to obtain the visible features of the visible facial region. Based on a preset facial supplementation mode, the occluded facial region is supplemented with supplementation prediction processing to obtain the missing features of the occluded facial region; and the visible features and missing features are fused with scores to obtain the overall facial features; wherein, the facial supplementation mode is used to perform score fusion processing on the missing features and visible features according to preset scoring information. Facial detection processing is performed on the overall features to obtain detection result information; wherein, the detection result information represents the identity of the target object. Therefore, key point detection and occlusion recognition processing are performed on the input facial image to automatically locate the visible facial region and the occluded facial region. The visible facial region includes available local areas (such as the periorbital region, forehead region, and temple region). Subsequently, depth feature extraction is performed on each local area in the visible facial region to obtain the visible features. Then, based on a preset facial augmentation pattern, the occluded facial areas are augmented with predictive processing to obtain missing features. Using the scoring information from the facial augmentation pattern, the visible features and missing features are fused to obtain the overall facial features, ensuring that the output features maintain high identity recognition capability even under occlusion. Finally, facial detection processing is performed on the overall features to generate the final detection result, i.e., the identity recognition result. Therefore, for occluded face images, the missing features of the occluded facial areas can be obtained through the facial augmentation pattern, and all facial features can be obtained based on these missing features. This allows for facial recognition based on all features, significantly improving the accuracy of facial recognition and solving the technical problem of low accuracy due to facial occlusion.

[0023] Figure 2 A flowchart illustrating a face image recognition method provided in an embodiment of the present invention. Figure 2 ,like Figure 2 As shown, in this embodiment... Figure 1 Based on the embodiments, the method is described in detail below, and the method includes: S201. Obtain the face image of the target object.

[0024] For example, the electronic device acquires a facial image of the target object, wherein the description of the facial image is given in step S101 and will not be repeated here.

[0025] S202. Perform key point detection and occlusion recognition processing on the face image to obtain the visible area of ​​the face and the occluded area of ​​the face.

[0026] In one example, S202 includes: performing key point detection and occlusion recognition processing on a face image according to a preset face key point detection network to extract face key point information; performing face image segmentation processing according to the face key point information and a preset semantic segmentation network to obtain the face occlusion region; and generating a heat map of the face occlusion region; and performing block processing on the face image according to the face key point information and the heat map to generate the face visible region.

[0027] For example, the electronic device extracts facial landmark information from the captured facial image using a pre-set lightweight facial landmark detection network. Since most feature points of the jaw and hairline may be missing when the user wears a mask and helmet, the electronic device can construct a "periocular coordinate system" based on the facial landmark information, using the corner of the eye, pupil center, and brow peak as anchor points. This weakens the registration error caused by the missing chin / hairline, ensuring that the forehead can still be accurately located under the helmet, providing basic geometric positioning for subsequent region cropping and occlusion judgment.

[0028] Based on the acquired facial landmark information, a pre-defined lightweight semantic segmentation network is used to segment the input face image, yielding segmentation results, which are probability distribution maps for categories such as mask, helmet, skin, and hair. These segmentation results clearly identify facial occlusion areas, i.e., which areas of the face belong to masks and which belong to helmets, thus generating a heatmap H of the occluded facial areas.

[0029] Combining key point information with H, the face region is segmented, and high-value regions (ROIs) that are not occluded are dynamically cropped out, i.e., the visible facial areas. The visible facial areas include: the periorbital area; the forehead-brow bone area; and the temple-upper cheekbone area, etc. Adaptive cropping is performed through deformable convolution to ensure that effective visible local areas can still be extracted under different wearing methods (mask moved up / down, helmet tilted).

[0030] Therefore, in cases of facial occlusion, the system can automatically locate usable local facial areas (around the eyes, brow bone to forehead, temples / upper cheekbone) even with the double coverage of a mask and helmet, facilitating subsequent facial recognition and greatly improving the accuracy of subsequent facial recognition.

[0031] S203. Determine the quality weight of each region in the visible facial area based on the preset quality weight information.

[0032] For example, after generating the visible area of ​​interest (ROI) of the face, the electronic device performs a quality assessment on each region within the ROI according to preset quality weight information. The quality is categorized as follows: , Where qi represents the overall quality score of the i-th region in the visible region of face (ROI), and the higher the overall quality score, the more suitable the region is for feature extraction; This represents the weighting coefficient of each preset indicator, used to adjust the contribution of different factors to the overall quality score; vis i The visibility rate metric represents the proportion of unoccluded pixels in the i-th region, reflecting the integrity of that region; sharp i Sharpness index, obtained through Laplacian variance or other sharpness measurements, reflects the degree of sharpness of an image within a region; i The nir index represents the uniformity of illumination, used to measure whether the illumination distribution within a region is uniform, reducing the impact of uneven brightness on feature extraction; i It represents modal consistency and detects the degree of alignment between a face image and visible light.

[0033] Finally, through softmax normalization, the quality weight w of each region in the visible facial region is obtained. i The quality weight w i This will determine the importance of different regions during subsequent local feature extraction.

[0034] S204. Based on the quality weights and heatmap, generate an attention mask for each region in the visible facial area.

[0035] For example, the electronic device uses the quality weights and heatmap from the previous step to generate an attention mask for each region of the visible facial area. Specifically, the unoccluded pixel regions in the heatmap are element-wise multiplied with the attention map of the intermediate layer features of the convolutional network, and then activated by a sigmoid function to obtain a normalized attention mask M. This attention mask M is subsequently used to suppress occluded region features and increase the contribution of local regions to feature extraction.

[0036] Therefore, even with the double cover of a mask and a helmet, it can automatically locate available local facial areas (around the eyes, brow bone to forehead, temples / upper cheekbone) and generate learnable attention masks and quality weights based on the automatically located local facial areas.

[0037] S205. Extract features from the visible area of ​​the face to obtain the visible features of the visible area of ​​the face.

[0038] In one example, S205 includes: during the feature depth extraction process, each region in the visible area of ​​the face is input into multiple preset feature extraction branch networks for feature depth extraction to obtain the visible features of each region; based on the mapping relationship between the region and the attention mask, and the preset region-aware pooling mechanism, the visible features of the region are weighted and pooled with the attention mask corresponding to the region to obtain the weighted visible features.

[0039] For example, after the visible area of ​​the face is extracted by the electronic device, each region within that visible area needs to undergo preprocessing. This preprocessing includes size normalization, scaling the face to a preset size (e.g., 112×112, this is just an example and not a limitation), illumination equalization, and attention-based masking to ensure that each region within the visible area of ​​the face has a uniform scale and quality before entering the feature extraction network. Therefore, through preprocessing, the interference of factors such as illumination and blurring on recognition performance can be greatly reduced in subsequent face recognition.

[0040] Subsequently, during the feature depth extraction process, each region within the visible facial area is input into a pre-defined feature extraction branch network for deep feature extraction, yielding the visual features of each region. This feature extraction branch network is divided into three local branches: P, F, and T. These branches are composed of lightweight convolutional neural networks used to extract deep representation features of the regions. Then, based on the mapping relationship between regions and attention masks, the attention mask M corresponding to the visual features of a region is determined. According to the pre-defined region-aware pooling mechanism introduced during convolution, the visual features of the region and the corresponding attention mask M are weighted and pooled to obtain the weighted visual features. Therefore, through weighted pooling of the attention mask and visual features, visual features are more prominent, thereby highlighting the effective, unoccluded areas while suppressing occlusion noise.

[0041] Optionally, based on the feature depth extraction process, a pre-defined hybrid attention mechanism is introduced, comprising two parts: channel attention and spatial attention. Channel attention automatically determines which feature channels contain more useful identity information, such as fine lines around the eyes or the iris edge; spatial attention further enhances edge contours and textured regions within the two-dimensional space of the feature map. Therefore, by optimizing visual features through the hybrid attention mechanism, more accurate and effective visual features are obtained, facilitating improvements in the accuracy of subsequent face recognition.

[0042] S206. Based on the preset facial supplementation pattern, perform supplementation prediction processing on the facial occlusion area to supplement the missing features of the facial occlusion area.

[0043] In one example, step S206 includes three implementation methods: The first implementation of step S206 is as follows: Based on the local consistency distillation loss information in the preset facial supplementation mode and the preset facial global features, the preset facial global features are mapped to the facial occlusion area and the facial visible area respectively, to obtain multiple areas in the facial occlusion area, multiple areas in the facial visible area, and the missing features of each area in the facial occlusion area; wherein, the preset facial global features include the features of each area in the multiple areas.

[0044] The second implementation of step S206 is as follows: Acquire the sound data of the target object; perform timbre recognition processing on the sound data to determine the gender of the target object; determine the first global facial feature corresponding to the target object based on the preset mapping relationship between gender and global facial features; based on the local consistency distillation loss information in the preset facial supplementation mode, map the first global facial feature to the occluded facial area and the visible facial area respectively, obtaining multiple regions in the occluded facial area, multiple regions in the visible facial area, and the missing features of each region in the occluded facial area; wherein, the first global facial feature includes the features of each region in the multiple regions.

[0045] The third implementation of step S206 is as follows: Collect upper body data of the target object; perform gender feature recognition processing on the upper body data, and determine the gender of the target object based on the recognized gender features; determine the second global facial feature corresponding to the target object based on the preset mapping relationship between gender and global facial features; based on the local consistency distillation loss information in the preset facial supplementation mode, map the second global facial feature to the occluded facial area and the visible facial area respectively, obtaining multiple areas in the occluded facial area, multiple areas in the visible facial area, and the missing features of each area in the occluded facial area; wherein, the second global facial feature includes the features of each area in the multiple areas.

[0046] For example, to avoid severe feature shifts under occlusion, electronic devices introduce consistency constraints during the training phase. When an occluded face image is available, global face features are extracted first as preset global facial features; however, when the face is occluded, local features are forced to maintain a certain consistency with the global face features in the embedding space. That is, in the feature learning phase, instead of relying on a single classification supervision, a combination of multiple discriminative losses is used to improve robustness. Therefore, face completion modes include ArcFace (additive angular distance loss information), local consistency distillation loss information, etc.

[0047] For example, in the first implementation of step S206, the preset global facial features include features of each region in multiple regions. Based on the local consistency distillation loss information in the preset facial supplementation mode and the preset global facial features, the features of each region included in the preset global facial features are mapped to the occluded facial region and the visible facial region, respectively, to obtain multiple regions in the occluded facial region, multiple regions in the visible facial region, and the missing features of each region in the occluded facial region.

[0048] In the second implementation of step S206, the first global facial feature includes features of each region in multiple regions. First, the target object's voice data is collected, and the timbre of the voice data is processed to determine the target object's gender (male or female). Based on a preset mapping relationship between gender and global facial features, the first global facial feature corresponding to the target object is determined. Then, based on the local consistency distillation loss information in a preset facial supplementation mode, the features of each region included in the first global facial feature are mapped to the occluded facial region and the visible facial region, respectively, resulting in multiple regions in the occluded facial region, multiple regions in the visible facial region, and the missing features of each region in the occluded facial region. Therefore, by determining the first global facial feature based on the target object's gender and supplementing the missing features of each region in the occluded facial region based on the first global facial feature, facial supplementation can be performed according to gender, greatly improving the accuracy of facial supplementation.

[0049] In the third implementation of step S206, the second global facial feature includes features of each region in multiple regions. First, upper body data of the target object is collected; this upper body data can be data collected by an infrared sensor or an upper body image, without limitation; the upper body data can include data from the head to the neck, or data from the head to the waist, etc., without limitation; it can also be full-body data, etc., without limitation. Gender features of the upper body data are identified, and the gender of the target object is determined based on the identified gender features, such as a prominent Adam's apple or a prominent chest, without limitation. Then, according to a preset mapping relationship between gender and global facial features, the second global facial feature corresponding to the target object is determined. Based on the local consistency distillation loss information in the preset facial supplementation mode, the second global facial feature is mapped to the occluded facial region and the visible facial region, respectively, resulting in multiple regions in the occluded facial region, multiple regions in the visible facial region, and the missing features of each region in the occluded facial region.

[0050] Therefore, the gender of the target object is determined based on the upper body data, and the second global facial features are determined based on the gender. Based on the second global facial features, the missing features of each area in the occluded facial region are obtained, and the face can be supplemented according to gender, which greatly improves the accuracy of face supplementation.

[0051] S207. The visible features and missing features are fused together to obtain the overall features of the face; wherein, the face supplementation mode is used to perform scoring fusion processing on the missing features and visible features according to the preset scoring information.

[0052] In one example, S207 includes: determining a first uncertainty score value for each region in the visible area of ​​the face and a second uncertainty score value for each region in the occluded area of ​​the face based on preset scoring information; and performing a scoring fusion process on the face image based on the visible features, missing features, the first uncertainty score value, and the second uncertainty score value to obtain the overall features of the face.

[0053] In one example, “the facial image is scored and fused based on the visible features, missing features, first uncertainty score, and second uncertainty score to obtain the overall facial features” includes: if there are multiple facial images, the local features in adjacent frames are aggregated and enhanced according to a preset temporal model to obtain the enhanced overall features.

[0054] For example, facial enhancement methods include ArcFace (additive angular distance loss information) and local consistency distillation loss information. ArcFace (additive angular distance loss information) is used for identity classification and can significantly increase the inter-class angular distance and reduce the variance of intra-class feature distribution, thereby achieving more stable differentiation between highly similar faces. Local consistency distillation loss information uses a preset global facial feature f* of an unoccluded face as the target. , Among them, f i The feature vector extracted from the i-th local branch; : Region projection function, which maps a preset global facial feature f* to the corresponding local region; The obtained loss value; the local area is the area in the visible area of ​​the face and the area in the occluded area of ​​the face.

[0055] Different local regions contribute differently to recognition. For example, periorbital features are most stable when obscured by a mask, while forehead features may be unavailable when a helmet is pulled down. Therefore, based on the scoring information in the facial supplementation pattern, an uncertainty score is assigned to the feature vector of each local region. This uncertainty score includes a first uncertainty score for each region within the occluded facial region and a second uncertainty score for each region within the visible facial region. , Among them, f i,j d represents the component of the feature vector in the j-th dimension within the i-th local region; d represents the dimension of the feature vector. This represents the mean of the feature vectors within the i-th region; This represents the uncertainty score of the features in the i-th region, reflecting the stability of the features in that region. The smaller the value, the more stable and reliable the features in that region are.

[0056] In the final fusion stage, the features of each local region are weighted and averaged based on the first uncertainty score and the second uncertainty score to obtain the overall facial features. : , Where N represents the number of local regions. This indicates that the uncertainty score is expressed as a negative value; f i w is the feature vector extracted from the i-th local branch; i Let be the quality weight of the i-th region.

[0057] Optionally, in video or continuous capture scenarios, a single frame image may have incomplete features due to momentary occlusion or blinking. A pre-defined temporal model can be introduced to aggregate local features from adjacent frames. Let the fused feature at time t be F. t Then, the enhanced representation is obtained through ConvLSTM: , Among them, H t This is a temporally enhanced hidden representation that integrates information from multiple past frames; H t t is the hidden representation at time t-1; ConvLSTM represents a long short-term memory network with convolutional units, used to capture spatiotemporal dependencies.

[0058] Therefore, after score fusion and temporal enhancement, a stable global representation is obtained, which is the enhanced overall feature. This is then mapped to a low-dimensional space through an embedding layer to form the final identity vector, which is the feature of the complete face.

[0059] S208. Perform facial detection processing on the overall features to obtain detection result information; wherein, the detection result information represents the identity of the target object.

[0060] In one example, before performing facial detection processing on the overall features to obtain the detection result information, the method further includes: acquiring multimodal data; wherein the multimodal data includes multiple local features extracted by infrared sensors and / or depth sensors; determining the region and quality weight corresponding to each local feature; wherein the quality weight is determined according to preset quality weight information; the region is a region in the visible area of ​​the face, or a region in the occluded area of ​​the face; for the visible area of ​​the face, multiplying the local features and visible features in each region by the corresponding quality weight to obtain multiple first values; for the occluded area of ​​the face, multiplying the local features and missing features in each region by the corresponding quality weight to obtain multiple second values; and performing normalized weighted summation on the first values ​​and second values ​​to obtain a comprehensive feature vector.

[0061] In one example, S208 includes: updating the overall features to a comprehensive feature vector, and performing liveness detection processing on the comprehensive feature vector to obtain a liveness detection result; if the detection result information indicates that the liveness detection is passed, comparing the comprehensive feature vector with the registered templates in the preset database to obtain comparison result information; wherein, the comparison result information includes the matching score between the comprehensive feature vector and each registered template; generating a confidence score based on the maximum matching score, the weight of each region, the first uncertainty score, the second uncertainty score, and the liveness detection result; and determining the detection result information based on the confidence score and the preset confidence score threshold.

[0062] For example, the electronic device acquires enhanced feature vectors and corresponding uncertainty scores for multiple local regions. Multiple local features are extracted by an infrared sensor or depth sensor using multimodal data reception. After collecting all local features, they are first unified into a vector space of the same dimension, and the source region and corresponding quality weight of each local feature are labeled to prepare for subsequent fusion. The quality weight is determined based on preset quality weight information; the source region is a region within the visible area of ​​the face, or a region within an occluded area of ​​the face. Because different regions have varying importance in the current scenario (e.g., the periorbital region is more stable than the jawline when obscured by a mask), electronic devices can adaptively fuse features based on uncertainty scores and region quality weights. Specifically, for the visible facial region, both the local features (i.e., local feature vectors) and visible features in each region are multiplied by their corresponding quality weights to obtain multiple first values. For the occluded facial region, both the local features and missing features in each region are multiplied by their corresponding quality weights to obtain multiple second values. Finally, the first and second values ​​are normalized and weighted to obtain a comprehensive feature vector. Therefore, this weighted fusion effectively suppresses the influence of occluded or blurred regions, making the final embedding more robust.

[0063] To ensure the security of the recognition results, liveness detection is performed after fusing the comprehensive feature vectors. Liveness detection can be divided into passive and active methods. Liveness detection is modeled using temporal signals or depth features; the rPPG signal can be derived from image sequences I... t extract: , Where R represents the skin region. This represents the grayscale value of the pixel (x, y) at time, where x represents the pixel's horizontal position and y represents the pixel's vertical position. This is the result of a liveness detection. The system uses Fourier transform to detect heart rate periodicity to determine if it's a real face. If it is a face, the liveness detection is passed; otherwise, the identity result is rejected.

[0064] After successful liveness detection, the fused feature vector is compared with registered templates in a pre-defined database to obtain comparison results, including the matching score between the fused feature vector and each registered template. Specifically, cosine similarity is used for comparison and matching to obtain a matching score for each registered template. To enhance stability, the pre-defined matching threshold can be adjusted according to environmental conditions to achieve dynamic calibration.

[0065] After calculating the matching score, the maximum matching score is determined. This maximum matching score is then combined with the weights of each local feature (i.e., the weights of each region: the weights of features within each region of the visible facial area and the weights of features within each region of the occluded facial area), the maximum matching score, the first uncertainty score, the second uncertainty score, and the liveness detection result to generate the final confidence score. Finally, based on the confidence score and a preset confidence score threshold, the detection result information C is determined.

[0066] , Among them, S maxThe highest similarity score, i.e., the maximum matching score, is represented by a ∈ [0,1], which is the adjustment coefficient; N represents the number of local regions. This represents the uncertainty score. The electronic device can output explanatory information, such as "70% contribution weight around the eyes, liveness detection passed." After completing the above steps, the final identity determination result is output, including the identified identity, matching score, confidence level, and liveness verification status. Therefore, this result can be directly transmitted to access control, attendance, or security management systems, thereby enabling automated and highly reliable facial recognition applications. It should be noted that registered templates can be updated online according to conditions to ensure long-term stability and robustness.

[0067] The facial image recognition method provided in this application acquires a facial image of a target object. Keypoint detection and occlusion recognition are performed on the facial image to obtain the visible facial region and the occluded facial region. A quality weight is determined for each region within the visible facial region based on preset quality weight information. An attention mask is generated for each region within the visible facial region based on the quality weight and a heatmap. Feature extraction is performed on the visible facial region to obtain its visual features. Based on a preset facial supplementation mode, supplementation prediction processing is performed on the occluded facial region to obtain missing features. The visual features and missing features are then fused through scoring to obtain the overall facial features; wherein, the facial supplementation mode is used to perform scoring fusion processing on the missing features and visual features based on preset scoring information. Facial detection processing is performed on the overall features to obtain detection result information; wherein, the detection result information represents the identity of the target object. Therefore, this application solves the technical problem of low accuracy in facial recognition due to facial occlusion.

[0068] Figure 3 This is a schematic diagram of the structure of a face image recognition device provided in an embodiment of the present invention, as shown below. Figure 3 As shown, the face image recognition device 30 provided in this embodiment includes: Acquisition device 31 is used to acquire a facial image of a target object; The recognition device 32 is used to perform key point detection processing and occlusion recognition processing on the face image to obtain the visible area of ​​the face and the occluded area of ​​the face. Extraction device 33 is used to extract features from the visible area of ​​the face to obtain the visible features of the visible area of ​​the face; The prediction device 34 is used to perform supplementary prediction processing on the occluded area of ​​the face based on a preset face supplementation pattern, and supplement the missing features of the occluded area of ​​the face. The fusion device 35 is used to perform scoring fusion processing on visible features and missing features to obtain the overall features of the face; wherein, the face supplementation mode is used to perform scoring fusion processing on missing features and visible features according to preset scoring information. The detection device 36 is used to perform facial detection processing on the overall features to obtain detection result information; wherein, the detection result information represents the identity of the target object.

[0069] Figure 4 This is a schematic diagram of another face image recognition device provided in an embodiment of the present invention. Figure 3 Based on the illustrated embodiments, as Figure 4 As shown, the prediction device 34 includes: Based on the local consistency distillation loss information in the preset facial supplementation pattern and the preset facial global features, the preset facial global features are mapped to the facial occlusion area and the facial visible area respectively, to obtain multiple regions in the facial occlusion area, multiple regions in the facial visible area, and the missing features of each region in the facial occlusion area; wherein, the preset facial global features include the features of each region in the multiple regions.

[0070] In one possible implementation, the prediction device 34 includes: Collect the voice data of the target object; and perform timbre recognition processing on the voice data to determine the gender of the target object; Based on the preset mapping relationship between gender and global facial features, the first global facial feature corresponding to the target object is determined; Based on the local consistency distillation loss information in the preset facial supplementation mode, the first global facial features are mapped to the facial occlusion area and the facial visible area respectively, to obtain multiple regions in the facial occlusion area, multiple regions in the facial visible area, and the missing features of each region in the facial occlusion area; wherein, the first global facial features include the features of each region in the multiple regions.

[0071] In one possible implementation, the prediction device 34 includes: Collect upper body data of the target object; identify the gender characteristics of the upper body data; and determine the gender of the target object based on the identified gender characteristics. Based on the preset mapping relationship between gender and facial global features, the second facial global feature corresponding to the target object is determined; Based on the local consistency distillation loss information in the preset facial supplementation mode, the second global facial features are mapped to the facial occlusion area and the facial visible area respectively, to obtain multiple regions in the facial occlusion area, multiple regions in the facial visible area, and the missing features of each region in the facial occlusion area; wherein, the second global facial features include the features of each region in the multiple regions.

[0072] In one possible implementation, the fusion device 35 includes: The first determining unit 351 is used to determine a first uncertainty score value for each region in the visible area of ​​the face and a second uncertainty score value for each region in the occluded area of ​​the face based on preset scoring information. The second determining unit 352 is used to perform scoring fusion processing on the face image based on visible features, missing features, a first uncertainty score, and a second uncertainty score to obtain the overall features of the face.

[0073] In one possible implementation, the second determining unit 352 includes: If there are multiple face images, then according to the preset temporal model, the local features in the adjacent multiple frames are aggregated and enhanced to obtain the enhanced overall features.

[0074] In one possible implementation, the identification device 32 includes: Based on the preset facial key point detection network, key point detection and occlusion recognition are performed on the facial image to extract facial key point information; Based on facial key point information and a pre-set semantic segmentation network, the face image is segmented to obtain the occluded facial region; and a heat map of the occluded facial region is generated. Based on facial key point information and heat map, the facial image is segmented to generate a visible facial region.

[0075] In one possible implementation, the device is also specifically used for: After performing key point detection and occlusion recognition on the face image to obtain multiple visible facial regions and occluded facial regions, the quality weight of each region in the visible facial region is determined according to the preset quality weight information. Based on quality weights and heatmaps, an attention mask is generated for each region of the visible facial area.

[0076] In one possible implementation, the extraction device 33 includes: During the feature depth extraction process, each region of the visible facial area is... Each region is input into multiple preset feature extraction branch networks for feature depth extraction processing to obtain the visible features of each region. Based on the mapping relationship between regions and attention masks, and the preset region-aware pooling mechanism, the visual features of a region are weighted and pooled with the attention mask corresponding to the region to obtain the weighted visual features.

[0077] In one possible implementation, the device is also specifically used for: Before performing facial detection processing on the overall features to obtain the detection results, multimodal data is acquired; the multimodal data includes multiple local features extracted by infrared sensors and / or depth sensors. Determine the region and quality weight corresponding to each local feature; where the quality weight is determined based on preset quality weight information; the region is either a region within the visible facial area or a region within the occluded facial area. For the visible facial region, the local features and visible features in each region are multiplied by the corresponding quality weight to obtain multiple first values; for the occluded facial region, the local features and missing features in each region are multiplied by the corresponding quality weight to obtain multiple second values. The first and second values ​​are normalized and weighted, and then summed to obtain the comprehensive feature vector.

[0078] In one possible implementation, the detection device 36 includes: Update the overall features to a comprehensive feature vector, and then process the comprehensive feature vector. Perform liveness detection processing to obtain liveness detection results; If the detection result indicates that the liveness detection is passed, the comprehensive feature vector is compared with the registered templates in the preset database to obtain the comparison result information; wherein, the comparison result information includes the matching score between the comprehensive feature vector and each registered template; A confidence score is generated based on the maximum matching score in the matching score, the weight of each region, the first uncertainty score, the second uncertainty score, and the liveness detection result. The detection result information is determined based on the confidence score and the preset confidence score threshold.

[0079] The apparatus provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.

[0080] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 5As shown, the electronic device 50 provided in this embodiment includes at least one processor 501 and a memory 502. Optionally, the device 50 further includes a communication component 503. The processor 501, memory 502, and communication component 503 are connected via a bus 504.

[0081] In a specific implementation, at least one processor 501 executes computer execution instructions stored in memory 502, causing at least one processor 501 to perform the above-described method.

[0082] The specific implementation process of processor 501 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0083] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0084] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0085] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0086] In the description of the embodiments of the present invention, it should be understood that the terms "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "center," "top," "bottom," "top," "bottom," "inner," "outer," "inner side," and "outer side," etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the accompanying drawings and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. "Inner side" refers to the interior or enclosed area or space. "Outer perimeter" refers to the area surrounding a specific component or specific area.

[0087] In the description of embodiments of the present invention, the terms "first," "second," "third," and "fourth" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first," "second," "third," or "fourth" may explicitly or implicitly include one or more of that feature. In the description of the present invention, unless otherwise stated, "a plurality of" means two or more.

[0088] In the description of the embodiments of the present invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," "joining," and "assembly" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.

[0089] In the description of embodiments of the present invention, specific features, structures, materials or characteristics may be combined in any suitable manner in one or more embodiments or examples.

[0090] In the description of the embodiments of the present invention, it should be understood that "-" and "~" represent a range of two numerical values, and this range includes the endpoints. For example, "AB" represents a range greater than or equal to A and less than or equal to B. "A~B" represents a range greater than or equal to A and less than or equal to B.

[0091] In the description of embodiments of the present invention, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.

[0092] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A face image recognition method, characterized in that, include: Acquire a facial image of the target object; and perform key point detection and occlusion recognition processing on the facial image to obtain the visible facial area and the occluded facial area; The visible facial region is then subjected to feature extraction to obtain the visible features of the visible facial region. Based on a preset facial enhancement mode, the facial occlusion area is subjected to enhancement prediction processing to obtain the missing features of the facial occlusion area; and the visible features and the missing features are fused together to obtain the overall facial features; wherein, the facial enhancement mode is used to perform scoring fusion processing on the missing features and the visible features according to preset scoring information. The overall features are subjected to facial detection processing to obtain detection result information; wherein, the detection result information represents the identity of the target object.

2. The method according to claim 1, characterized in that, The surface based on the preset The partial supplementation mode performs supplementation prediction processing on the facial occlusion area to supplement the missing features of the facial occlusion area, including: Based on the local consistency distillation loss information in the preset facial supplementation mode and the preset facial global features, the preset facial global features are mapped to the facial occlusion area and the facial visible area respectively, to obtain multiple areas in the facial occlusion area, multiple areas in the facial visible area, and the missing features of each area in the facial occlusion area; wherein, the preset facial global features include the features of each area in the multiple areas.

3. The method according to claim 1, characterized in that, The surface based on the preset The partial supplementation mode performs supplementation prediction processing on the facial occlusion area to supplement the missing features of the facial occlusion area, including: Collect the voice data of the target object; and perform timbre recognition processing on the voice data to determine the gender of the target object; Based on the preset mapping relationship between gender and global facial features, the first global facial feature corresponding to the target object is determined; Based on the local consistency distillation loss information in the preset facial supplementation mode, the first facial global features are mapped to the facial occlusion area and the facial visible area respectively, to obtain multiple areas in the facial occlusion area, multiple areas in the facial visible area, and the missing features of each area in the facial occlusion area; wherein, the first facial global features include the features of each area in the multiple areas.

4. The method according to claim 1, characterized in that, The surface based on the preset The partial supplementation mode performs supplementation prediction processing on the facial occlusion area to supplement the missing features of the facial occlusion area, including: Collect upper body data of the target object; identify the gender characteristics of the upper body data; and determine the gender of the target object based on the identified gender characteristics. Based on the preset mapping relationship between gender and facial global features, the second facial global feature corresponding to the target object is determined; Based on the local consistency distillation loss information in the preset facial supplementation mode, the second facial global features are mapped to the facial occlusion area and the facial visible area respectively, to obtain multiple regions in the facial occlusion area, multiple regions in the facial visible area, and the missing features of each region in the facial occlusion area; wherein, the second facial global features include the features of each region in the multiple regions.

5. The method according to claim 1, characterized in that, Combine the visual features with The missing features are subjected to scoring fusion processing to obtain the overall facial features, including: Based on preset scoring information, determine a first uncertainty score value for each region in the visible facial area and a second uncertainty score value for each region in the occluded facial area; The facial image is scored and fused based on the visible features, the missing features, the first uncertainty score, and the second uncertainty score to obtain the overall features of the face.

6. The method according to claim 1, characterized in that, Process the face image Keypoint detection and occlusion recognition are performed to obtain the visible facial area and the occluded facial area, including: Based on a preset facial landmark detection network, the facial image is subjected to landmark detection. The facial key point information is extracted through measurement and occlusion recognition processing. Based on the facial key point information and a preset semantic segmentation network, the facial structure is analyzed. The face image is segmented to obtain the occluded facial region; and a heat map of the occluded facial region is generated. Based on the facial key point information and the heatmap, the facial image is segmented. Block processing generates the visible facial area.

7. The method according to claim 6, characterized in that, In the face image After performing key point detection and occlusion recognition to obtain multiple visible facial regions and occluded facial regions, the method further includes: The quality weight of each region in the visible facial area is determined based on the preset quality weight information. An attention mask is generated for each region of the visible facial area based on the quality weights and the heatmap.

8. The method according to claim 7, characterized in that, For the facial visible area Feature extraction is performed on the domain to obtain the visual features of the visible facial region, including: During the feature depth extraction process, each region of the visible facial area is... Each region is input into multiple preset feature extraction branch networks for feature depth extraction processing to obtain the visual features of each region. Based on the mapping relationship between the region and the attention mask, and the preset region-aware pooling mechanism, the visual features of the region are weighted and pooled with the attention mask corresponding to the region to obtain the weighted visual features.

9. The method according to claim 1, characterized in that, In terms of the overall features Before performing facial detection processing and obtaining the detection result information, the method further includes: Acquire multimodal data; wherein the multimodal data includes infrared sensor and / or depth sensor. Multiple local features extracted by the degree sensor; Determine the region and quality weight corresponding to each of the local features; wherein, the quality... The weight is determined based on preset quality weight information; the region is a region within the visible area of ​​the face, or a region within the occluded area of ​​the face. For the visible facial region, the local features and visible features in each region are multiplied by the corresponding quality weight to obtain multiple first values; for the occluded facial region, the local features and missing features in each region are multiplied by the corresponding quality weight to obtain multiple second values. The first and second values ​​are normalized and weighted, and then summed to obtain a comprehensive feature vector.

10. The method according to claim 9, characterized in that, Regarding the overall features Facial detection processing is performed to obtain detection result information, including: The overall features are updated to a comprehensive feature vector, and the comprehensive feature vector is then processed. Perform liveness detection processing to obtain liveness detection results; If the detection result information indicates that the liveness detection is passed, the comprehensive feature vector is compared with the registered templates in the preset database to obtain the comparison result information; wherein, the comparison result information includes the matching score of the comprehensive feature vector with each registered template; A confidence score is generated based on the maximum matching score in the matching scores, the weight of each region, the first uncertainty score, the second uncertainty score, and the liveness detection result. The detection result information is determined based on the confidence score and the preset confidence score threshold.