Facial recognition method and apparatus, and electronic device and storage medium

By extracting and fusing facial and ambient lighting features and using an autoencoder model for face recognition, the problem of misidentification caused by irregular camera lighting is solved, and efficient, low-cost lighting-interference-resistant face recognition is achieved.

WO2025194760A1PCT designated stage Publication Date: 2025-09-25CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD

Patent Information

Application Number
PCT/CN2024/125809
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-22
Filing Date
2024-10-18
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

In the field of face recognition, due to the large differences in camera placement and ambient lighting, the lighting distribution is irregular, resulting in an increased misrecognition rate. Existing technologies are difficult to effectively solve the impact of lighting changes on face recognition, especially when computing power is insufficient. Lightweight models cannot be applied to scenes with changing lighting, and the cost of training customized models is high.

Method used

By extracting facial lighting features and ambient lighting features from facial images, the fusion vector is obtained and input into the autoencoder model for reconstruction processing, which reduces the impact of lighting and improves recognition accuracy. The autoencoder model is used for unsupervised adaptive updates to reduce the workload of large-scale sample data collection.

Benefits of technology

It improves the accuracy and efficiency of face recognition, reduces the impact of light on recognition, reduces manpower and hardware costs, and is suitable for terminal devices with limited computing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024125809_25092025_PF_FP_ABST
    Figure CN2024125809_25092025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to a facial recognition method and apparatus, and an electronic device and a storage medium. The method comprises: acquiring real-time video data, and performing face detection processing on the real-time video data, in order to obtain a facial image and an environmental image; performing facial illumination extraction processing and facial recognition processing on the facial image, in order to obtain a facial illumination feature distribution vector and a facial feature vector; performing environmental illumination extraction processing on the environmental image, in order to obtain an environmental illumination feature distribution vector; performing fusion processing on the facial illumination feature distribution vector, the facial feature vector and the environmental illumination feature distribution vector, in order to obtain a fused vector; inputting the fused vector into an auto-encoder model for reconstruction processing, in order to obtain a multi-modal feature vector; and performing retrieval and matching processing on the multi-modal feature vector, in order to obtain a facial recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Face recognition method, device, electronic device and storage medium

[0001] Related applications

[0002] This application claims priority to Chinese patent application number 202410338451X, filed on March 22, 2024, entitled “A face recognition method, device, electronic device and storage medium,” the entire text of which is hereby incorporated by reference. Technical Field

[0003] The present application relates to the field of face recognition technology, and in particular to a face recognition method, device, electronic device and storage medium. Background Art

[0004] In the field of face recognition, facial images are obtained through cameras for recognition. However, due to the large differences in the location of each camera and the ambient lighting, the corresponding lighting distribution and field of view are irregular. The irregular lighting distribution and field of view seriously reduce the accuracy of the face recognition system. The main problem caused is the increase in the false recognition rate, resulting in a large number of false recognitions and rejections.

[0005] In summary, the technical problems existing in the relevant technologies need to be improved.

[0006] Summary of the Invention

[0007] The embodiments of the present application provide a face recognition method, device, electronic device, and storage medium.

[0008] An embodiment of the present application provides a face recognition method, comprising:

[0009] Acquire real-time video data, and perform face detection processing on the real-time video data to obtain a face image and an environment image;

[0010] Performing facial illumination extraction processing and facial recognition processing on the facial image to obtain a facial illumination feature distribution vector and a facial feature vector;

[0011] Performing ambient lighting extraction processing on the ambient image to obtain an ambient lighting feature distribution vector;

[0012] Fusing the facial illumination feature distribution vector, the facial feature vector, and the ambient illumination feature distribution vector to obtain a fusion vector;

[0013] Inputting the fusion vector into the autoencoder model for reconstruction to obtain a multimodal feature vector;

[0014] Performing retrieval and matching processing on the multimodal feature vector to obtain a face recognition result.

[0015] In some embodiments, performing face detection processing on the real-time video data to obtain a face image and an environment image includes:

[0016] Performing frame processing on the real-time video data to obtain an image set;

[0017] The face detection model is used to perform face detection processing on the image set to obtain a face image and an environment image.

[0018] In some embodiments, the face detection model is constructed based on a cross-stage local network.

[0019] In some embodiments, performing facial illumination extraction processing and facial recognition processing on the facial image to obtain a facial illumination feature distribution vector and a facial feature vector includes:

[0020] Performing key point detection processing on the face image to obtain a set of face key points;

[0021] Performing facial illumination extraction processing on the facial image according to the facial key point set to obtain a facial illumination feature distribution vector;

[0022] Performing face recognition processing on the face image to obtain a face feature vector.

[0023] In some embodiments, performing facial illumination extraction processing on the facial image according to the facial key point set to obtain a facial illumination feature distribution vector includes:

[0024] Dividing the face image according to the face key point set to obtain a face block image;

[0025] The illumination mean value of each block area in the face block image is calculated and processed to obtain a face illumination feature distribution vector.

[0026] In some embodiments, the facial key point set includes external contour points and internal points; and the segmenting process of the facial image according to the facial key point set to obtain facial segmented images includes:

[0027] The facial image is subjected to contour segmentation processing based on the external contour points, and the facial image is subjected to horizontal and vertical segmentation processing based on the internal points to obtain the facial block image.

[0028] In some embodiments, the set of facial key points includes points of eyebrows, eyes, nose, mouth, and facial contour areas.

[0029] In some embodiments, performing ambient lighting extraction processing on the ambient image to obtain an ambient lighting feature distribution vector includes:

[0030] determining a recognition area according to the environmental image;

[0031] performing regional illumination intensity extraction processing on the environment image according to the identified area to obtain an illumination intensity histogram of the identified area;

[0032] Perform illumination mean calculation processing on the illumination intensity histogram of the recognition area to obtain an ambient illumination feature distribution vector.

[0033] In some embodiments, performing regional illumination intensity extraction processing on the environment image according to the identified area to obtain an illumination intensity histogram of the identified area includes:

[0034] Converting the environment image into a grayscale image according to the identified area;

[0035] A histogram of light intensity in the identification area is obtained according to the grayscale image.

[0036] In some embodiments, performing retrieval and matching processing on the multimodal feature vector to obtain a face recognition result includes:

[0037] Obtaining a facial feature database;

[0038] Feature matching processing is performed on the facial feature database according to the multimodal feature vector to obtain a facial recognition result.

[0039] In some embodiments, before inputting the fused vector into the autoencoder model for reconstruction, the method further includes:

[0040] Pre-training an initial autoencoder model to obtain the trained autoencoder model.

[0041] In some embodiments, the pre-training of the initial autoencoder model to obtain the trained autoencoder model includes:

[0042] Obtain environmental video data and erroneous face sample data at different time periods;

[0043] Performing ambient lighting extraction processing on the ambient video data to obtain an ambient lighting training vector;

[0044] Performing face illumination extraction processing and face recognition processing on the erroneous face sample data to obtain a face illumination training vector and a face feature training vector;

[0045] Fusing the environmental illumination training vector, the facial illumination training vector, and the facial feature training vector to obtain a training fusion vector;

[0046] The training fusion vector is input into the initial autoencoder model, and the parameters of the initial autoencoder model are updated based on the back propagation algorithm and the gradient descent method to obtain the autoencoder model.

[0047] Another aspect of the present application provides a face recognition device, comprising:

[0048] The first module is used to acquire real-time video data and perform face detection processing on the real-time video data to obtain a face image and an environment image;

[0049] The second module is used to perform facial illumination extraction and facial recognition processing on the facial image to obtain a facial illumination feature distribution vector and a facial feature vector;

[0050] The third module is used to perform ambient lighting extraction processing on the ambient image to obtain an ambient lighting feature distribution vector;

[0051] A fourth module is configured to fuse the facial illumination feature distribution vector, the facial feature vector, and the ambient illumination feature distribution vector to obtain a fusion vector;

[0052] A fifth module is configured to input the fused vector into an autoencoder model for reconstruction to obtain a multimodal feature vector;

[0053] The sixth module is used to perform retrieval and matching processing on the multimodal feature vector to obtain a face recognition result.

[0054] Another aspect of an embodiment of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor implements the aforementioned method when executing the computer program.

[0055] Another aspect of the embodiments of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.

[0056] The details of one or more embodiments of the present application are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the present application will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without any creative work.

[0058] FIG1 is a flow chart of a face recognition method provided in an embodiment of the present application;

[0059] FIG2 is a flow chart of a multi-terminal interactive face recognition method provided by an embodiment of the present application;

[0060] FIG3 is a flow chart of step S101 in FIG1 ;

[0061] FIG4 is a flow chart of step S102 in FIG1 ;

[0062] FIG5 is a flow chart of step S402 in FIG4 ;

[0063] FIG6 is a schematic diagram of a facial illumination extraction method provided in an embodiment of the present application;

[0064] FIG7 is a flow chart of step S103 in FIG1 ;

[0065] FIG8 is a schematic diagram of the operation of a face recognition method provided in an embodiment of the present application;

[0066] FIG9 is a schematic structural diagram of a face recognition device provided in an embodiment of the present application;

[0067] FIG10 is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0068] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.

[0069] It will be understood that the terms "first," "second," and the like used herein may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are merely used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words "if," "if" as used herein, may be interpreted as "at the time of," "when," or "in response to determining."

[0070] The terms "at least one", "plurality", "each", "any", etc. used in this application include "at least one", "two" or more, "plurality" or "each", "any" or "any one", "each" or "any one" as used herein.

[0071] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0072] Before explaining the embodiments of the present application in detail, some of the nouns and terms involved in the embodiments of the present application are first explained. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0073] False Positive Rate (FPR): One of the performance metrics associated with binary classification problems, it measures the model's ability to incorrectly classify negative samples as positive. FPR is an important component when calculating the ROC curve and is often used to evaluate model performance in ROC curve (Receiver Operating Characteristic Curve) analysis. The ROC curve shows the trade-off between the model's true positive rate (TPR) and FPR at different thresholds. Typically, the X-axis of the ROC curve represents FPR, the Y-axis represents TPR, and the area under the curve (AUC) is a model evaluation metric used to measure model performance. The closer the AUC is to 1, the better the performance.

[0074] Gamma Correction: A common technique used in image processing to adjust the brightness and contrast of an image. It is a nonlinear transformation typically used to correct the brightness and color of an image. The core idea of ​​gamma correction is to apply a nonlinear function to change the image's brightness distribution to better match the human eye's perception.

[0075] Region of Interest (ROI): A common term in image processing and computer vision, it refers to selecting a specific area within an image, typically an area encompassing a specific feature or object. ROIs are often used to extract or analyze a specific region within an image for further processing or analysis.

[0076] In related technologies, in facial recognition scenarios for regional personnel management, the location and environment of each camera vary significantly, resulting in irregular lighting distribution and field of view. These irregular lighting distributions and field of view significantly reduce the accuracy of facial recognition systems, primarily leading to increased FPRs and a large number of false positives and false negatives. Currently, to optimize the model, a large amount of data must be collected from each camera. Because lighting and imaging vary in each time period, this long data collection process increases the workload and slows down the optimization algorithm. Furthermore, in the field of on-device facial recognition, especially when computing power is insufficient, heavyweight facial recognition models cannot be used. Lightweight facial recognition models are not fully applicable to scenarios with varying lighting conditions. Therefore, customized model training is often required for these scenarios, increasing costs.

[0077] For example, one approach is to improve the generalization capability of the face recognition system by collecting data from each camera to obtain a variety of facial data. However, this approach has drawbacks. Due to the huge number of cameras, the difficulty of filtering out data with ID tags increases exponentially. Moreover, learning hundreds of millions of data requires a larger model, which increases the requirements for equipment and hardware, making it more expensive. However, this approach is not effective in application scenarios with limited computing power. A lightweight recognition model cannot handle such complex situations. Another approach is to analyze the illumination intensity of the face and use different gamma correction coefficients to correct it, making the face to be recognized clearer. However, this approach also has drawbacks. Currently, the analysis of facial illumination intensity can only be done through statistical methods to determine the threshold value of each camera. The facial illumination intensity varies with the time of day. Since the imaging of each camera at each time is different, the robustness of the threshold setting is not good. Moreover, the gamma coefficient correction algorithm cannot solve the problems caused by uneven facial illumination (such as yin-yang face, too dark, overexposed, etc.), resulting in low face recognition accuracy.

[0078] In view of this, embodiments of the present application provide a face recognition method, device, electronic device, and storage medium. This solution obtains a face illumination feature distribution vector and a face feature vector by performing face illumination extraction and face recognition processing on a face image; and obtains an environment illumination feature distribution vector by performing environment illumination extraction processing on an environment image. The face illumination feature distribution vector, the face feature vector, and the environment illumination feature distribution vector are then fused to obtain a fused vector, which is then input into an autoencoder model for reconstruction to obtain a multimodal feature vector. This fusion of multimodal information allows for face recognition, reduces the impact of illumination on face recognition, and improves the accuracy of face recognition. Furthermore, this solution utilizes an autoencoder model, which enables unsupervised adaptive updates, reduces the workload of large-scale sample data collection, and improves the efficiency of face recognition.

[0079] The embodiment of the present application provides a face recognition method, which relates to the field of face recognition technology. The face recognition method provided in the embodiment of the present application can be applied to a terminal, can also be applied to a server, and can also be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements a face recognition method, etc., but is not limited to the above forms.

[0080] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0081] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.

[0082] FIG1 is an optional flowchart of a face recognition method provided in an embodiment of the present application. The method in FIG1 may include but is not limited to steps S101 to S106.

[0083] Step S101, acquiring real-time video data, and performing face detection processing on the real-time video data to obtain a face image and an environment image;

[0084] Step S102, performing facial illumination extraction processing and facial recognition processing on the facial image to obtain a facial illumination feature distribution vector and a facial feature vector;

[0085] Step S103, performing ambient lighting extraction processing on the ambient image to obtain an ambient lighting feature distribution vector;

[0086] Step S104, fusing the facial illumination feature distribution vector, the facial feature vector, and the ambient illumination feature distribution vector to obtain a fusion vector;

[0087] Step S105: inputting the fusion vector into the autoencoder model for reconstruction to obtain a multimodal feature vector;

[0088] Step S106: performing search and matching processing on the multimodal feature vector to obtain a face recognition result.

[0089] In the steps S101 to S106 shown in the embodiment of the present application, with reference to FIG2 , real-time video data is obtained by arranging a camera or other shooting device at the front end or mobile end, and then face detection is performed on the real-time video data using a face detection model to obtain a face image and an environment image, wherein the face image is an image in which a face is detected, and the environment image is an image in which no face is detected. The face image is then input into the face key point detection model and the face recognition feature extraction model respectively to obtain a face illumination distribution vector and a face sample vector to be tested, that is, a face illumination feature distribution vector and a face feature vector are extracted from the face image. The face key point detection model and the face recognition feature extraction model can be obtained by training a pre-built model, or an existing related model can be used. The ROI region illumination distribution statistics of the environment image are performed to obtain a histogram of the illumination intensity of the ROI region in continuous time, thereby calculating the environment variable distribution vector, that is, extracting the environment illumination feature distribution vector from the environment image. The ambient lighting feature distribution vector, facial lighting feature distribution vector, and facial feature vector are then fused with multi-modal information. The model is fine-tuned using an autoencoder, and the fused vector is input into the autoencoder model for reconstruction. This allows the final multi-modal feature vector to be used for facial feature matching, identifying the face and determining whether it has passed facial recognition. Feature vectors such as misidentification data are then saved and input into the cloud or server to update the autoencoder model, reducing the workload of collecting sample data and enabling local adaptation to different scenarios. This solution can be applied to every camera, eliminating the need for manual inspection of large-scale recognition data, reducing labor costs. It can also customize the recognition model for each camera, eliminating the need to retrain the facial recognition model, which would otherwise increase product launch time. Furthermore, this solution can still be used on older camera devices, effectively improving camera utilization.

[0090] 3 , in step S101 of some embodiments, face detection processing is performed on the real-time video data to obtain a face image and an environment image, including:

[0091] S301, performing frame processing on the real-time video data to obtain an image set;

[0092] S302, performing face detection processing on the image set through a face detection model to obtain a face image and an environment image; the face detection model is constructed based on a cross-stage local network.

[0093] In an embodiment of the present application, real-time video data can be obtained through a shooting device such as a camera or a camera, and then the obtained real-time video is frame-processed to obtain a set of each frame image, and then each frame image in the image set is subjected to face detection processing by a face detection model to obtain a face image and an environmental image, where the face image is an image in which a face is detected, and the environmental image is an image in which a face is not detected. Among them, the face detection model of this scheme is constructed using a cross-stage local network (CSPNet), wherein the cross-stage local network is a lightweight network designed for edge computing systems, which improves the learning ability of the convolutional network by reducing the duplication of gradient information and memory costs. The embodiment of the present application performs face detection processing on the image set through a face detection model to obtain a face image and an environmental image, thereby providing a data basis for subsequent face illumination extraction processing and face recognition processing.

[0094] Referring to FIG. 4 , in step S102 of some embodiments, performing facial illumination extraction processing and facial recognition processing on the facial image to obtain a facial illumination feature distribution vector and a facial feature vector includes:

[0095] S401, performing key point detection processing on the face image to obtain a set of face key points;

[0096] S402, performing facial illumination extraction processing on the facial image according to the facial key point set to obtain a facial illumination feature distribution vector;

[0097] S403: Perform face recognition processing on the face image to obtain a face feature vector.

[0098] In an embodiment of the present application, facial illumination extraction processing and facial recognition processing are performed on the facial image respectively. First, key point detection processing is performed on the facial image to obtain a set of facial key points. Facial key point detection refers to locating the key points of the facial face, including eyebrows, eyes, nose, mouth, and points in the facial contour area, given a facial image. Then, facial illumination extraction processing is performed on the facial image according to the facial key point set to obtain a facial illumination feature distribution vector. Finally, facial recognition processing is performed on the facial image through a facial recognition model to obtain a facial feature vector, wherein the facial recognition model can be constructed using a yolo model or the like. In an embodiment of the present application, facial illumination feature distribution vectors, facial feature vectors, and ambient illumination feature distribution vectors are fused by performing facial illumination extraction processing and facial recognition processing on the facial image, thereby improving the robustness of facial recognition and reducing the influence of illumination on facial recognition.

[0099] 5 , in step S302 of some embodiments, performing facial illumination extraction processing on the facial image according to the facial key point set to obtain a facial illumination feature distribution vector includes:

[0100] S501, dividing the face image according to the face key point set to obtain face block images;

[0101] S502: Calculate the illumination mean of each block area in the face block image to obtain a face illumination feature distribution vector.

[0102] In an embodiment of the present application, the face key point set includes 98 points in total, including external contour points and internal points. As shown in FIG6 , the external contour points are used to perform contour segmentation processing on the face image, and the internal points are used to perform horizontal and vertical segmentation processing on the face image. The face key points provide a positional basis for the subsequent extraction of the face illumination distribution. After the face key point set is obtained using the face key point model, the face illumination distribution extraction model is input to perform the face illumination extraction processing. The face key point model and the face illumination distribution extraction model can be obtained by training a pre-built model, or an existing related model can be used. Among them, in order to accurately capture the distribution of facial illumination, the face illumination distribution extraction model divides the face image into blocks. Referring to FIG5 , the face is divided into 25 blocks according to the face key points, and the illumination mean of each block is calculated as the illumination value of the current block, and finally a face illumination feature distribution vector with a length of 25 dimensions is obtained. In related technologies, facial illumination information extraction can only be performed by using statistical methods to determine the threshold of each camera. However, this threshold is often a fixed value and cannot describe the distribution of facial illumination intensity in a fixed manner. Therefore, the embodiment of the present application uses a facial illumination analysis algorithm to perform a block operation on the face using key points, thereby obtaining the horizontal, vertical and overall illumination distribution of the face, and obtaining a facial illumination distribution vector with relatively high robustness, and then fusing the multimodal information of the facial illumination distribution vector, the ambient illumination distribution vector and the facial recognition feature vector, thereby improving the accuracy of facial recognition.

[0103] 7 , in step S103 of some embodiments, performing ambient lighting extraction processing on the ambient image to obtain an ambient lighting feature distribution vector includes:

[0104] S701, determining a recognition area according to the environment image;

[0105] S702, performing regional illumination intensity extraction processing on the environment image according to the identified area to obtain an illumination intensity histogram of the identified area;

[0106] S703: Perform illumination mean calculation processing on the illumination intensity histogram of the recognition area to obtain an ambient illumination feature distribution vector.

[0107] In an embodiment of the present application, a camera or other shooting device is first set to identify an ROI area within the field of view, thereby determining the identification area in the environmental image based on the ROI area. Then, the environmental image is subjected to regional light intensity extraction processing based on the identification area. The environmental image can be converted into a grayscale image through tools such as OpenCV, thereby obtaining a light intensity histogram of the identification area based on the grayscale image, and then the illumination mean of the light intensity histogram of the identification area is calculated to obtain an environmental light feature distribution vector. Here, the identification area of ​​the environmental image can also be divided, and then the illumination mean of the divided area is calculated to obtain a multi-dimensional environmental light feature distribution vector. Since the location and environment of each camera are quite different, the corresponding illumination distribution and field of view are irregular. Therefore, the embodiment of the present application obtains the environmental light feature distribution vector of the corresponding camera by performing environmental light extraction on the environmental image, which is used for subsequent fusion of multimodal information, thereby improving the accuracy of face recognition.

[0108] In step S104 of some embodiments, the facial illumination feature distribution vector, the facial feature vector, and the ambient illumination feature distribution vector are fused to obtain a fused vector;

[0109] In an embodiment of the present application, the facial illumination feature distribution vector, the facial feature vector and the ambient illumination feature distribution vector can be normalized into vectors of the same dimension, and then the facial illumination feature distribution vector, the facial feature vector and the ambient illumination feature distribution vector are weighted fused. Different weights are given to different feature vectors according to their importance, and then the feature vectors are weightedly summed to obtain the fused feature vector. The specific weights can be set according to actual conditions.

[0110] In step S105 of some embodiments, the fused vector is input into an autoencoder model for reconstruction to obtain a multimodal feature vector;

[0111] In an embodiment of the present application, the autoencoding model includes an encoder and a decoder, which encodes the high-dimensional input fusion vector into a low-dimensional latent variable, thereby forcing the neural network to learn the most informative features; then the latent variables of the hidden layer are restored to the initial dimension through the decoder, so that the output of the decoder can perfectly or approximately restore the original input. This allows the neural network to better learn the illumination features and facial features in the fusion vector and obtain a multimodal feature vector. The embodiment of the present application can perform local adaptation according to different application scenarios by using an autoencoding learning model, and can improve the illumination interference resistance of lightweight face recognition in application scenarios with limited computing power, thereby improving the accuracy of the face recognition system.

[0112] In step S106 of some embodiments, performing retrieval and matching processing on the multimodal feature vector to obtain a face recognition result includes:

[0113] Obtaining a facial feature database;

[0114] Feature matching processing is performed on the facial feature database according to the multimodal feature vector to obtain a facial recognition result.

[0115] In an embodiment of the present application, a facial feature database is first established, which is used to store facial features that can be identified. By performing feature matching between the multimodal feature vector and the facial features in the facial feature database, the distance between the feature vectors can be calculated using Euclidean distance or cosine distance, and then the distance relationship between the two image feature points is obtained to determine whether the multimodal feature vector matches the facial features. When the matching features match, it means that the facial recognition has passed, and when they do not match, it means that the facial recognition has failed. In the subsequent process, it is possible to find out whether there is a case of misidentification by manual methods, and record and save the relevant multimodal feature vectors of the misidentification, and input the saved feature vectors into the cloud or server to update the autoencoder model, so as to perform local adaptation according to different application scenarios. The autoencoder used in the embodiment of the present application is more efficient than the optimized facial recognition model, and can be more conveniently deployed on a mobile terminal with limited computing power, thereby reducing costs.

[0116] In the steps of some embodiments, before inputting the fused vector into the autoencoder model for reconstruction, the method further includes pre-training the initial autoencoder model, specifically obtaining the trained autoencoder model.

[0117] In some embodiments, the pre-training of the initial autoencoder model to obtain the trained autoencoder model includes:

[0118] Obtain environmental video data and erroneous face sample data at different time periods;

[0119] Performing ambient lighting extraction processing on the ambient video data to obtain an ambient lighting training vector;

[0120] Performing face illumination extraction processing and face recognition processing on the erroneous face sample data to obtain a face illumination training vector and a face feature training vector;

[0121] Fusing the environmental illumination training vector, the facial illumination training vector, and the facial feature training vector to obtain a training fusion vector;

[0122] The training fusion vector is input into the initial autoencoder model, and the parameters of the initial autoencoder model are updated based on the back propagation algorithm and the gradient descent method to obtain the autoencoder model.

[0123] In an embodiment of the present application, environmental video data and erroneous face sample data are first obtained at different time periods. The environmental video data can be collected by using a camera device installed on non-working days and during periods when no one is passing by, in the morning, noon, and evening, to obtain the corresponding environmental video data. The erroneous face sample data is facial feature data that is not in the facial feature database. A large amount of erroneous sample data can be provided by the experience recognition system, thereby obtaining an erroneous sample vector, which can be used to increase the threshold for face recognition. Similarly, by performing facial illumination extraction and facial recognition on the erroneous face sample data, a facial illumination training vector and a facial feature training vector are obtained. Then, by performing ambient illumination extraction on the environmental video data, an ambient illumination training vector is obtained. The ambient illumination training vector, the facial illumination training vector, and the facial feature training vector are fused to obtain a training fusion vector and input into an initial autoencoder model. The initial autoencoder model is fitted with a vector relationship with a true registration template, thereby updating the parameters of the initial autoencoder model based on a backpropagation algorithm and a gradient descent method to obtain a pre-trained autoencoder model. The true registration template is the true registered face feature in the facial feature database. The embodiment of the present application can also select corresponding environmental distribution vectors according to different time periods, fuse the environmental distribution variables with the facial illumination feature distribution vector and facial feature vector in the real-time video data, and input them into the autoencoder model to obtain the final multimodal feature vector, thereby performing retrieval and matching processing to obtain face recognition results.

[0124] The following is a detailed description of the embodiments of the present application with reference to specific application examples:

[0125] Referring to Figure 8, the embodiment of the present application can be applied to scenarios such as security or regional personnel management in factory work areas. For example, it can be applied to building access security systems. Video data is acquired through a camera, and the acquired video data is input into a face detection model for face detection. When a face is detected, facial key point detection and facial feature extraction are performed on the detected face to obtain a facial illumination distribution vector and a facial feature vector. Furthermore, an environmental image is extracted from the video data, which can be the previous frame of the facial image. The environmental image is extracted to obtain an environmental illumination distribution vector. The extracted environmental illumination distribution vector, facial illumination distribution vector, and facial feature vector are then fused and input into an autoencoder to obtain a multimodal feature vector. Finally, a database matching is performed, that is, the multimodal feature vector is matched with the facial features in the facial feature database. If the match is successful, the user identification information is returned and the user is released. Otherwise, the user is considered a stranger and is not released. The embodiments of this application use an autoencoding model, eliminating the need for manual review of large-scale recognition data and reducing labor costs. Furthermore, the autoencoding model can be locally adapted to different scenarios, making it more efficient than optimizing facial recognition models and more conveniently deployable on mobile devices with limited computing power, thus reducing costs. The embodiments of this application can be applied to scenarios with limited computing power, improving the light interference resistance of lightweight facial recognition and thereby increasing the accuracy of the facial recognition system.

[0126] Referring to FIG. 9 , an embodiment of the present application further provides a face recognition device that can implement the above-mentioned face recognition method. The device includes:

[0127] The first module 901 is used to obtain real-time video data and perform face detection processing on the real-time video data to obtain a face image and an environment image;

[0128] The second module 902 is configured to perform facial illumination extraction and facial recognition processing on the facial image to obtain a facial illumination feature distribution vector and a facial feature vector;

[0129] The third module 903 is configured to perform ambient lighting extraction processing on the ambient image to obtain an ambient lighting feature distribution vector;

[0130] The fourth module 904 is configured to fuse the facial illumination feature distribution vector, the facial feature vector, and the ambient illumination feature distribution vector to obtain a fusion vector.

[0131] The fifth module 905 is configured to input the fused vector into an autoencoder model for reconstruction to obtain a multimodal feature vector;

[0132] The sixth module 906 is used to perform search and matching processing on the multimodal feature vector to obtain a face recognition result.

[0133] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0134] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned face recognition method when executing the computer program. The electronic device can be any smart terminal including a tablet computer, an in-vehicle computer, or the like.

[0135] It can be understood that the contents of the above method embodiments are applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0136] Please refer to FIG10 , which illustrates a hardware structure of an electronic device according to another embodiment. The electronic device includes:

[0137] The processor 1001 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.

[0138] The memory 1002 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1002 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called by the processor 1001 to execute the face recognition method of the embodiments of this application.

[0139] Input / output interface 1003, used to implement information input and output;

[0140] Communication interface 1004, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0141] Bus 1005 , which transmits information between various components of the device (e.g., processor 1001 , memory 1002 , input / output interface 1003 , and communication interface 1004 );

[0142] The processor 1001 , the memory 1002 , the input / output interface 1003 and the communication interface 1004 are connected to each other in communication within the device via the bus 1005 .

[0143] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned face recognition method is implemented.

[0144] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0145] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0146] Embodiments of the present application provide a face recognition method, device, electronic device, and storage medium. This solution performs facial illumination extraction and facial recognition processing on a facial image to obtain a facial illumination feature distribution vector and a facial feature vector; and then performs ambient illumination extraction processing on an ambient image to obtain an ambient illumination feature distribution vector. The facial illumination feature distribution vector, the facial feature vector, and the ambient illumination feature distribution vector are then fused to obtain a fused vector, which is then input into an autoencoder model for reconstruction to obtain a multimodal feature vector. This fusion of multimodal information allows for face recognition, reduces the impact of illumination on face recognition, and improves face recognition accuracy. Furthermore, this solution utilizes an autoencoder model, enabling unsupervised adaptive updates, reducing the workload of large-scale sample data collection and improving face recognition efficiency.

[0147] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0148] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0149] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0150] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0151] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0152] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0153] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0154] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0155] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0156] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the traditional technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store programs.

[0157] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0158] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A face recognition method, comprising: Acquire real-time video data, and perform face detection processing on the real-time video data to obtain a face image and an environment image; Performing facial illumination extraction processing and facial recognition processing on the facial image to obtain a facial illumination feature distribution vector and a facial feature vector; Performing ambient lighting extraction processing on the ambient image to obtain an ambient lighting feature distribution vector; Fusing the facial illumination feature distribution vector, the facial feature vector, and the ambient illumination feature distribution vector to obtain a fusion vector; Inputting the fusion vector into the autoencoder model for reconstruction to obtain a multimodal feature vector; Performing retrieval and matching processing on the multimodal feature vector to obtain a face recognition result.

2. The method according to claim 1, wherein The performing face detection processing on the real-time video data to obtain a face image and an environment image includes: Performing frame processing on the real-time video data to obtain an image set; The face detection model is used to perform face detection processing on the image set to obtain a face image and an environment image.

3. The method according to claim 1, wherein The face detection model is constructed based on a cross-stage local network.

4. The method according to claim 1, wherein The performing of facial illumination extraction processing and facial recognition processing on the facial image to obtain a facial illumination feature distribution vector and a facial feature vector includes: Performing key point detection processing on the face image to obtain a set of face key points; Performing facial illumination extraction processing on the facial image according to the facial key point set to obtain a facial illumination feature distribution vector; Performing face recognition processing on the face image to obtain a face feature vector.

5. The method according to claim 4, wherein The step of performing facial illumination extraction processing on the facial image according to the facial key point set to obtain a facial illumination feature distribution vector includes: Dividing the face image according to the face key point set to obtain a face block image; The illumination mean value of each block area in the face block image is calculated and processed to obtain a face illumination feature distribution vector.

6. The method according to claim 5, wherein: The facial key point set includes external contour points and internal points; and the facial image is divided according to the facial key point set to obtain a facial block image, including: The facial image is subjected to contour segmentation processing based on the external contour points, and the facial image is subjected to horizontal and vertical segmentation processing based on the internal points to obtain the facial block image.

7. The method according to claim 5 or 6, wherein: The facial key point set includes points in the eyebrows, eyes, nose, mouth, and facial contour areas.

8. The method according to claim 1, wherein The performing ambient light extraction processing on the ambient image to obtain an ambient light feature distribution vector includes: determining a recognition area according to the environmental image; performing regional illumination intensity extraction processing on the environment image according to the identified area to obtain an illumination intensity histogram of the identified area; Perform illumination mean calculation processing on the illumination intensity histogram of the recognition area to obtain an ambient illumination feature distribution vector.

9. The method according to claim 8, wherein The performing regional illumination intensity extraction processing on the environment image according to the identified area to obtain an illumination intensity histogram of the identified area includes: Converting the environment image into a grayscale image according to the identified area; A histogram of light intensity in the identification area is obtained according to the grayscale image.

10. The method according to claim 1, wherein The performing retrieval and matching processing on the multimodal feature vector to obtain a face recognition result includes: Obtaining a facial feature database; Feature matching processing is performed on the facial feature database according to the multimodal feature vector to obtain a facial recognition result.

11. The method according to any one of claims 1 to 10, wherein: Before inputting the fused vector into the autoencoder model for reconstruction, the method further includes: Pre-training an initial autoencoder model to obtain the trained autoencoder model.

12. The method according to claim 11, wherein The pre-training of the initial autoencoder model to obtain the trained autoencoder model includes: Obtain environmental video data and erroneous face sample data at different time periods; Performing ambient lighting extraction processing on the ambient video data to obtain an ambient lighting training vector; Performing face illumination extraction processing and face recognition processing on the erroneous face sample data to obtain a face illumination training vector and a face feature training vector; Fusing the environmental illumination training vector, the facial illumination training vector, and the facial feature training vector to obtain a training fusion vector; The training fusion vector is input into the initial autoencoder model, and the parameters of the initial autoencoder model are updated based on the back propagation algorithm and the gradient descent method to obtain the autoencoder model.

13. A face recognition device, comprising: The first module is used to acquire real-time video data and perform face detection processing on the real-time video data to obtain a face image and an environment image; The second module is used to perform facial illumination extraction and facial recognition processing on the facial image to obtain a facial illumination feature distribution vector and a facial feature vector; The third module is used to perform ambient lighting extraction processing on the ambient image to obtain an ambient lighting feature distribution vector; A fourth module is configured to fuse the facial illumination feature distribution vector, the facial feature vector, and the ambient illumination feature distribution vector to obtain a fusion vector; A fifth module is configured to input the fused vector into an autoencoder model for reconstruction to obtain a multimodal feature vector; The sixth module is used to perform retrieval and matching processing on the multimodal feature vector to obtain a face recognition result.

14. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 12 when executing the computer program. 15 . A computer-readable storage medium storing a computer program, wherein the computer program implements the method according to claim 1 when executed by a processor.

Citation Information

Patent Citations

  • Near infrared human face identification method and device

    CN106250877A

  • Method and device for face three-dimensional model texture blending

    CN107945267A

  • Face recognition method and device and electronic equipment

    CN111414803A

  • Face recognition method and device, electronic equipment and storage medium

    CN118230385A

  • Object recognition apparatus and method based on environment matching

    US20230015295A1

Cited By

  • Electric meter reading method, device and equipment based on image recognition and medium

    CN120894771A