Face recognition method and related equipment

The ambient light sensor of the terminal device detects the dark light environment and uses an image enhancement model to process dark light images. Combined with face feature extraction and template image degradation strategies, the problem of low facial recognition accuracy in dark light scenes is solved, improving user experience and reducing hardware costs.

CN120340077AActive Publication Date: 2025-07-18HONOR DEVICE CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202410044054.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-18
Estimated Expiration
2044-01-10

AI Technical Summary

Technical Problem

In specific environments (such as dark light scenes), the image quality collected by conventional shooting devices is poor, resulting in a reduced accuracy of face recognition and a poor user experience, and adding auxiliary fill-up devices will increase hardware costs.

Method used

The ambient light sensor of the terminal device detects the illuminance, determine whether it is a dark-light environment, and use the pre-trained image enhancement model to denoise and brightness to the dark-light image, combine it with the face feature extraction model to match features. If the matching fails, the face template image will be degraded for secondary matching.

Benefits of technology

It improves the accuracy and user experience of facial recognition in dark light environments, reduces the error rejection rate, and avoids increasing hardware costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340077A_ABST
    Figure CN120340077A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a face recognition method and related equipment, and the method comprises the steps: determining whether a current environment is a dark light environment or not according to the illumination collected by an ambient light sensor during face recognition; under the condition that the current environment is determined to be the dark light environment, performing image enhancement on a to-be-recognized image collected by a shooting device; performing similarity verification on the face feature of the enhanced image and the face feature of a preset face template image, and if the similarity verification is passed, determining that the user identity verification is passed; if the verification is not passed, performing degradation processing on the face template image; performing similarity verification on the face features of the to-be-recognized image and the face features of the degraded image, and if the similarity verification is passed, determining that the user identity verification is passed; and if the two times of verification are not passed, determining that the user identity verification is not passed. According to the embodiment of the invention, the accuracy of face recognition in a dark light environment can be improved through a dark light image enhancement strategy and a template image degradation strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of terminals, and particularly relates to a face recognition method and related devices. Background Art

[0002] In the field of computer vision, face recognition technology is a technology for authenticating the identity of a user in an image to be recognized based on the face features in the image to be recognized. Face recognition technology has been applied to various scenarios. For example, by using a shooting device (such as the camera of a mobile phone, tablet, security access control, etc.) of a terminal device (such as a mobile phone, tablet) to collect real-time images of a user, and then performing face recognition on the collected images to authenticate the identity of the user, it can facilitate the user to use the face to unlock the terminal device or enter a specific function interface of the terminal device, improving the user experience.

[0003] However, due to the poor quality of the images collected by a conventional shooting device in a specific environment (such as a low-light scene), problems such as a decrease in accuracy rate and inability to unlock the device due to incorrect recognition may occur when using related technologies for face recognition, resulting in a poor user experience. Although the terminal device can be equipped with an additional auxiliary lighting device for face recognition, it will lead to an increase in hardware costs. Summary of the Invention

[0004] In view of the above, it is necessary to provide a face recognition method and related devices, which can solve problems such as a decrease in face recognition accuracy and user experience caused by the poor quality of the images collected by a conventional shooting device in a specific environment.

[0005] In a first aspect, this application provides a face recognition method, which is applied to a terminal device. The terminal device includes a shooting device and an ambient light sensor. The method includes: in response to a user operation that triggers face recognition, using the shooting device to collect an image to be recognized, and using the ambient light sensor to obtain the illuminance of the current environment; in the case where it is determined that the current environment is a low-light environment according to the illuminance, inputting the image to be recognized into a pre-trained image enhancement model to obtain an enhanced image of the image to be recognized; using a preset face feature extraction model to extract first face features from the enhanced image; if the first similarity between the first face features and second face features in a preset face template image is greater than or equal to a preset similarity threshold, determining that the identity verification of the user passes.

[0006] Through the above technical solution, when performing face recognition on a terminal device with a non-supplementary light camera, it is possible to determine whether the current is a low-light environment according to the illuminance collected by the ambient light sensor of the terminal device, so as to determine whether the image to be recognized captured by the imaging device is a low-light image; perform image enhancement on the low-light image captured in the low-light environment, so as to obtain an enhanced image with significantly reduced noise, appropriately restored brightness, and undamaged face features; by performing feature matching on the face features in the enhanced image and the face features in the preset face template image, the accuracy of face recognition in the low-light environment can be improved.

[0007] In a possible implementation manner, the method further includes: if the first similarity is less than the similarity threshold, performing image degradation processing on the face template image to obtain a degraded image of the face template image; using the face feature extraction model to extract a third face feature from the image to be recognized, and using the face feature extraction model to extract a fourth face feature from the degraded image; if the second similarity between the third face feature and the fourth face feature is greater than or equal to the similarity threshold, determining that the identity verification of the user passes; or, if the second similarity is less than the similarity threshold, determining that the identity verification of the user fails.

[0008] Through the above technical solution, if the result of face feature matching using the enhanced image fails, the noise of the low-light image can be further added to the face template image to degrade the template image, and the degraded image is used to perform a second matching attempt with the low-light image, that is, comparing the face features of the low-light image before enhancement and the degraded template image. By introducing the face template image degradation strategy, the false rejection rate (FRR) of face recognition in the low-light scenario can be further reduced, and the accuracy of face recognition and the user experience can be improved.

[0009] In a possible implementation manner, the method further includes: using the ambient light sensor to collect multiple consecutive illuminances of the current environment; based on the sliding average algorithm, determining the average value of the multiple consecutive illuminances; if the average value is less than a preset illuminance threshold, determining that the current environment is a low-light environment.

[0010] Through the above technical solution, the average value of multiple consecutive illuminances can be determined as the illuminance of the current environment by the sliding average algorithm, avoiding misjudgment of whether the current environment is a low-light environment caused by instantaneous illuminance increase or decrease, fluctuations of the ambient light sensor, etc., and improving the judgment accuracy of the low-light environment.

[0011] In a possible implementation, the extracting of the first facial feature from the enhanced image by using a preset facial feature extraction model includes: based on a preset face detection model, determining whether there is a face to be recognized in the enhanced image; if there is a face to be recognized in the enhanced image, intercepting an image of the face region from the enhanced image; and extracting the first facial feature from the image of the face region by using the facial feature extraction model.

[0012] Through the above technical solution, when performing facial feature recognition on the image to be enhanced, it is possible to first determine whether there is a face in the enhanced image and the image of the region where the face exists, so as to specifically recognize the facial features in the image of the face region, enabling the facial feature extraction model to not need to consider the features of the regions that do not contain faces, reducing the model computing power, and improving the efficiency and accuracy of facial feature extraction.

[0013] In a possible implementation, the extracting of the first facial feature from the enhanced image by using a preset facial feature extraction model further includes: if there is no face to be recognized in the enhanced image, determining that the identity verification of the user fails.

[0014] Through the above technical solution, when it is determined that there is no face to be recognized in the enhanced image, it is not necessary to perform other unnecessary steps, such as calling the facial feature extraction model to extract facial features, etc., thereby improving the efficiency of face recognition.

[0015] In a possible implementation, the method further includes: when it is determined that the current environment is not a low-light environment according to the illuminance, based on a preset face detection model, determining whether there is a face to be recognized in the image to be recognized; if there is a face to be recognized in the image to be recognized, intercepting an image of the face region from the image to be recognized; extracting a fifth facial feature from the image of the face region by using the facial feature extraction model; if it is determined that the third similarity between the fifth facial feature and the second facial feature is greater than or equal to the similarity threshold, determining that the identity verification of the user passes; or, if it is determined that the third similarity is less than the similarity threshold, determining that the identity verification of the user fails.

[0016] Through the above technical solution, when it is determined that the current environment is not a low-light environment, it can be considered that the image to be recognized captured is not a low-light image, and thus there is no need to perform image enhancement and other processing; the face recognition can be directly performed on the image to be recognized, and the face region in the image to be recognized can be intercepted; then the facial features of the intercepted face region image are extracted and feature-matched with the facial features in the face template image, so as to determine the face recognition result, and efficient and accurate face recognition can be achieved in a non-low-light environment.

[0017] In a possible implementation, the image enhancement model includes a first model and a second model. Inputting the image to be recognized into the pre-trained image enhancement model to obtain the enhanced image of the image to be recognized includes: denoising the image to be recognized by using the first model and outputting the denoised image to be recognized; enhancing the brightness of the denoised image to be recognized by using the second model and using the output brightness-enhanced image as the enhanced image.

[0018] Through the above technical solution, the noise in the low-light image can be removed by using the first model, and then the brightness of the denoised low-light image can be enhanced by using the second model, so as to obtain an enhanced image with significantly reduced noise, appropriately restored brightness, and the facial features not being significantly damaged.

[0019] In a possible implementation, the training method of the image enhancement model includes: obtaining multiple sample images, where the multiple sample images include multiple first images and multiple second images, and the image noise of any first image is more than that of any second image; constructing a training set and a test set according to the multiple sample images, and performing at least one iterative update on a preset initial model by using the training set until the image enhancement model that meets the preset requirements is obtained.

[0020] Through the above technical solution, the objective condition limitation of difficult collection of real paired data sets can be addressed, and a training strategy of non-paired training can be adopted, that is, there is no need for high-quality images corresponding one by one to real low-light and low-quality images as training true value labels. Among them, the first image represents a low-light and low-quality image with more image noise, and the second image represents a high-quality image, and there is no corresponding relationship between the first image and the second image. Therefore, the difficulty of obtaining the sample images used for training the image enhancement model can be reduced.

[0021] In a possible implementation, the preset requirements include a combination of one or more of the following requirements: the loss function corresponding to the initial model reaches the convergence condition; the test result obtained by testing the initial model by using the test set indicates that the model performance of the initial model reaches the expectation; the number of times of the at least one iterative update reaches a preset iteration number threshold.

[0022] Through the above technical solution, various methods can be used to determine whether the trained model meets the preset requirements, improving the usability of the model.

[0023] In a possible implementation, the initial model includes a first initial model, and any one of the at least one iterative update includes: inputting any first image in the training set into the first initial model, denoising the any first image by using the first initial model to obtain a first denoised image corresponding to the any first image, and a first difference image between the any first image and the first denoised image; determining a first loss value corresponding to the first difference image by using a preset first loss function, where the first loss value is used to indicate the noise degree in the first difference image; if the first loss value is less than or equal to a preset first threshold, using another first image to perform the next update on the first initial model.

[0024] In a possible implementation, any one of the at least one iterative update further includes: if the first loss value is greater than the first threshold, performing image fusion on the first difference image and any second image in the training set to obtain a fused image corresponding to the any second image; inputting the fused image into the first initial model, denoising the fused image by using the first initial model to obtain a second denoised image corresponding to the fused image; determining a second loss value between the second denoised image and the corresponding any second image by using a preset second loss function; if the second loss value is greater than a preset second threshold, using the next first image and the next second image to perform the next update on the first initial model.

[0025] Through the above technical solution, after any low-quality image A (such as the first image) passes through the denoising neural network (such as the first model) of the image enhancement model, it is expected to obtain a denoised image A and a difference map between the denoised image A and the low-quality image A (such as the first difference image, which is a pure noise map); by superimposing the pure noise map corresponding to the low-quality image A on any real high-quality image B (such as the second image), a low-quality image B (such as the fused image corresponding to the second image) can be synthesized; the real high-quality image B can thus be used as the enhancement target for synthesizing the low-quality image B to train the denoising neural network, that is, after denoising the synthesized low-quality image B using the denoising neural network, the obtained denoised image B should be as close as possible to the high-quality image B. Among them, during the training process of the above first model, the denoising neural network was used for denoising twice, and each denoising corresponded to a loss function. Specifically, during the first denoising, the low-quality image A was denoised to obtain a pure noise map, so the first loss function can be used to constrain the noise in this pure noise map, that is, to measure the degree of noise existence in this pure noise map; during the second denoising, the synthesized low-quality image B was denoised to obtain a denoised image B, and the second loss function can be used to constrain this denoised image B so that the denoised image B can be as close as possible to the high-quality image B. Through the above training method, it is possible to reduce the training difficulty of the first model and improve the performance of the model based on the unpaired training strategy.

[0026] In a possible implementation manner, the initial model further includes a second initial model, and any update in the at least one iterative update further includes: if the second loss value is less than or equal to the second threshold, input the second denoised image corresponding to any second image into the second initial model, use the second initial model to enhance the brightness of the second denoised image, and obtain a brightness-enhanced image corresponding to the second denoised image; use a preset third loss function to determine a third loss value between the brightness-enhanced image and the corresponding any second image; if the third loss value is greater than a preset third threshold, use the second denoised image corresponding to the next second image to perform the next update on the second initial model; or, if the third loss value is less than or equal to the third threshold, determine that the initial model meets the preset requirements, and use the initial model as the image enhancement model.

[0027] Through the above technical solution, the third loss function can be used to constrain the brightness-enhancing neural network (such as the second model) of the image enhancement model, so that the denoised image B (such as the second denoised image) before brightness enhancement is as close as possible to the original high-quality image B after passing through the brightness-enhancing neural network, improving the performance of the brightness-enhancing neural network.

[0028] In addition, through the iterative update of the first model and the second model by the above technical solution, the performance of the model can be continuously optimized to obtain an image enhancement model that meets the usage requirements.

[0029] In a possible implementation, the image degradation process for the face template image to obtain the degraded image of the face template image includes: obtaining a preset third image; inputting the third image into the image enhancement model, using the first model of the image enhancement model to denoise the third image to obtain the third denoised image corresponding to the third image, and the second difference image between the third image and the third denoised image; reducing the brightness of the face template image to obtain an image with reduced brightness; and performing image fusion on the second difference image and the image with reduced brightness to obtain the degraded image of the face template image.

[0030] Through the above technical solution, it is not necessary to separately train a degradation neural network for image degradation. Instead, the image enhancement model can be directly used to obtain the pure noise map corresponding to the low-quality image D (such as the third image), and this pure noise map is used as the noise image for degrading the face template (such as the second difference image). When the denoising model (such as the first model) in the aforementioned image enhancement model is properly and sufficiently trained, the degraded image here will have very realistic noise interference and an approximate noise distribution to the real low-quality image D. At this time, face feature extraction is performed on the degraded image, and a matching algorithm is used to compare the face features in the face template image with the face features in the degraded features, which can weaken the influence of noise to a certain extent and focus on the face feature differences, thereby improving the accuracy of face recognition.

[0031] In a second aspect, the present application provides a terminal device, which includes a memory and a processor: wherein, the memory is used to store program instructions; the processor is used to read and execute the program instructions stored in the memory, and when the program instructions are executed by the processor, the terminal device executes the above-mentioned face recognition method.

[0032] In a third aspect, the present application provides a chip coupled to the memory in the terminal device, and the chip is used to control the processor of the terminal device to execute the above-mentioned face recognition method.

[0033] In a fourth aspect, the present application provides a computer storage medium, which stores program instructions, and when the program instructions run on the terminal device, the processor of the terminal device is caused to execute the above-mentioned face recognition method.

[0034] In addition, the technical effects brought by the second aspect to the fourth aspect can be referred to the descriptions related to the methods of each design in the above method part, and will not be elaborated here. Description of the Drawings

[0035] Figure 1Schematic diagram of a low-light image and schematic diagram of related solutions provided by an embodiment of the present application.

[0036] Figure 2 Software architecture diagram of a terminal device provided by an embodiment of the present application.

[0037] Figure 3 Flowchart of a face recognition method provided by an embodiment of the present application.

[0038] Figure 4 Flowchart of a method for determining whether the current environment is a low-light environment provided by an embodiment of the present application.

[0039] Figure 5 Flowchart of a training method for an image enhancement model provided by an embodiment of the present application.

[0040] Figure 6 Schematic diagram of the application inference stage of an image enhancement model provided by an embodiment of the present application.

[0041] Figure 7 Flowchart of a method for performing image degradation processing on a face template image provided by an embodiment of the present application.

[0042] Figure 8 Schematic diagram of a template image degradation strategy provided by an embodiment of the present application.

[0043] Figure 9 Schematic diagram of a training strategy for unpaired training provided by an embodiment of the present application.

[0044] Figure 10 Flowchart of a method for any update to the first initial model provided by an embodiment of the present application.

[0045] Figure 11 Flowchart of a method for any update to the second initial model provided by an embodiment of the present application.

[0046] Figure 12 Flowchart of a method for performing face recognition on a high-quality image provided by an embodiment of the present application.

[0047] Figure 13 Schematic diagram of the overall architecture diagram of a face recognition method provided by an embodiment of the present application.

[0048] Figure 14 Flowchart of a face recognition method provided by an embodiment of the present application.

[0049] Figure 15 Hardware architecture diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners

[0050] In an embodiment of the present application, the terms "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, words such as "exemplary" or "for example" are used to mean for example, illustration, or explanation. Any embodiment or design described as "exemplary" or "for example" in an embodiment of the present application should not be construed as being more preferred or having more advantages than other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific manner.

[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. It should be understood that unless otherwise stated in this application, " / " means "or". For example, A / B may mean A or B. The "and / or" in this application is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A alone, A and B existing simultaneously, and B alone, these three situations. "At least one" means one or more. "Multiple" means two or more than two. For example, at least one of a, b, or c may mean: a, b, c, a and b, a and c, b and c, a, b, and c, these seven situations. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0052] In the field of computer vision, face recognition technology is a technology for authenticating the identity of a user in an image to be recognized based on the face features in the image to be recognized. Face recognition technology has been applied to a variety of scenarios. For example, by using the shooting device (such as the camera of a mobile phone) of a terminal device (such as a mobile phone, a tablet, a security access control, etc.) to collect real-time images of a user, and then performing face recognition on the collected images to authenticate the user's identity, it can facilitate the user to use the face to unlock the terminal device or enter a specific function interface of the terminal device, improving the user experience.

[0053] However, due to the poor quality of the images collected by a conventional shooting device in a specific environment (such as a low-light scene), problems such as a decrease in the accuracy rate when using related technologies for face recognition and the inability to unlock the device due to incorrect recognition, resulting in a poor user experience, may occur. Although the terminal device can add additional auxiliary devices for face recognition, it will lead to an increase in hardware costs.

[0054] Refer toFigure 1 As shown in the figure, it is a schematic diagram of a low-light image provided by an embodiment of the present application and a schematic diagram of a related solution. For example, if the illumination in the current environment is insufficient, the terminal device only has a conventional front camera and no other auxiliary devices (such as a Time of flight (TOF) sensor or other fill lights, etc.), the low-light image captured by the camera is often darker and full of obvious noise, as Figure 1 shown in the low-light image in Figure (a) of the figure. Only a very blurred outline of the human body can be seen. Among them, although the face has been mosaicked, referring to the blurred outline of the human body, the fine outlines in the face will be even more unidentifiable. For this situation, if only Figure 1 shown in Figure (b) of the figure, simply increasing the brightness of the image, the resulting image with increased brightness will contain a large amount of noise, destroying the face features in the image and making it difficult to achieve accurate face recognition; in the related art, as Figure 1 shown in Figure (c) of the figure, using the method of filling light with the maximum screen brightness, not only the filling light is limited and the filling light effect is poor, but also the enhanced screen brightness in the low-light scene is extremely dazzling, resulting in a poor user experience; in addition, as Figure 1 shown in Figure (d) of the figure, if additional auxiliary devices are added to the terminal device, such as adding an infrared fill light, it will increase a lot of hardware costs, resulting in cost increase.

[0055] To solve the above problems, an embodiment of the present application provides a face recognition method. When performing face recognition on a terminal device using a non-fill light camera, it is possible to determine whether the current is a low-light environment according to the illuminance collected by the ambient light sensor of the terminal device, and perform image enhancement on the low-light image captured in the low-light environment, so as to obtain an enhanced image with significantly reduced noise, appropriately restored brightness, and face features not significantly damaged; by performing feature matching between the face features in the enhanced image and the face features in the preset face template image, the accuracy of face recognition in the low-light environment can be improved. The specific process of the face recognition method will be described in detail below in combination with the Figure 3 process shown in the figure.

[0056] Referring to Figure 2 As shown in the figure, it is a software architecture diagram of a terminal device provided by an embodiment of the present application. The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. For example, the Android system from top to bottom is the application layer 101, the framework layer 102, the Android runtime and system libraries 103, the hardware abstraction layer 104, the kernel layer 105, and the hardware layer 106.

[0057] The application layer 101 may include a series of application packages. For example, the application packages may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, device control service, etc.

[0058] The framework layer 102 provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions. For example, the application framework layer may include window manager, content provider, view system, phone manager, resource manager, notification manager, etc.

[0059] Among them, the window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc. The content provider is used to store and obtain data, and make this data accessible to applications. The data may include videos, images, audio, dialed and answered calls, browsing history and bookmarks, phone book, etc. The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a short message notification icon may include a view for displaying text and a view for displaying pictures. The phone manager is used to provide the communication functions of the terminal device. For example, the management of call status (including connected, hung up, etc.). The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, etc. The notification manager enables applications to display notification information in the status bar, can be used to convey notification-type messages, and can automatically disappear after a short stay without user interaction. For example, the notification manager is used to inform that the download is completed, message reminder, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as the notification of a background-running application, and can also be a notification that appears on the screen in the form of a dialog window. For example, prompt text information in the status bar, emit a prompt tone, the terminal device vibrates, the indicator light flashes, etc.

[0060] Android Runtime includes core libraries and a virtual machine. Android runtime is responsible for the scheduling and management of the Android system. The core libraries contain two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core libraries of Android.

[0061] The application layer 101 and the framework layer 102 run in a virtual machine. The virtual machine executes the Java files of the application layer and the framework layer as binary files. The virtual machine is used to perform functions such as management of object life cycles, stack management, thread management, security and exception management, and garbage collection.

[0062] The system library 103 can include multiple functional modules. For example, a surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0063] Among them, the surface manager is used to manage the display subsystem and provide the fusion of 2D and 3D layers for multiple applications. The media libraries support the playback and recording of various common audio and video formats, as well as static image files, etc. The media libraries can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc. The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc. The 2D graphics engine is a drawing engine for 2D drawing.

[0064] The hardware abstraction layer 104 runs in the user space, encapsulates the kernel layer drivers, and provides call interfaces to the upper layer.

[0065] The kernel layer 105 is the layer between hardware and software. The kernel layer 105 at least includes a display driver, a touch driver, an audio driver, and a sensor driver.

[0066] The kernel layer 105 is the core of the operating system of the terminal device, is the first layer of software expansion based on the hardware, provides the most basic functions of the operating system, is the basis for the operation of the operating system, and is responsible for managing the system's processes, memory, device drivers, files, and network systems, and determines the performance and stability of the system. For example, the kernel layer can determine the operation time of an application program for a certain part of the hardware.

[0067] The kernel layer 105 includes programs closely related to the hardware, such as interrupt handlers, device drivers, etc., and also includes basic, common, and frequently running modules, such as a clock management module, a process scheduling module, etc., and also includes key data structures. The kernel layer can be set in the processor or solidified in the internal memory.

[0068] The hardware layer 106 includes the hardware of the terminal device, such as a display screen, buttons, a camera, an ambient light sensor, etc. The hardware layer 106 may not include auxiliary fill light devices, such as a time-of-flight sensor or other fill lights, etc.

[0069] Refer to Figure 3As shown in the figure, it is a flowchart of a face recognition method provided by an embodiment of the present application. The face recognition method is applied to a terminal device and includes the following processes.

[0070] S101, in response to a user operation that triggers face recognition, use a photographing device to collect an image to be recognized, and use an ambient light sensor to obtain the illuminance of the current environment.

[0071] In an embodiment of the present application, the user operation for triggering face recognition may vary due to differences in the application scenarios of the face recognition method. For example, when the face recognition method is applied to environments such as security access control, the user operation may include face scanning in front of an access control system or a security camera to verify identity; when the face recognition method is applied to scenarios such as unlocking a terminal device, the user operation may include using face recognition to unlock devices such as a mobile phone, a tablet computer, or a laptop computer, or using face recognition to unlock an application (APP) of the terminal device; when the face recognition method is applied to environments such as payment and financial services, the user operation may include using face recognition for payment verification or identity verification to access a bank account, etc. For example, when an application needs to verify the user's identity, it will prompt the user on an application interface to perform face recognition. The user clicks on a relevant icon or control on the application interface to trigger the operation of recognizing the user's face. The user operation can also be a preset gesture operation. For example, when the terminal device is a mobile phone, when the user raises the mobile phone facing the screen of the terminal device, the photographing device of the terminal device can start scanning and comparing the user's face. The above are only examples, and the actual application is not limited thereto.

[0072] In an embodiment of the present application, the image to be recognized can be a multi-channel image. For example, the multi-channel image can be an RGB image in which each pixel point includes three color channels: red (R), green (G), and blue (B). The image to be recognized can also be a grayscale image or an image in other formats, and the present application does not make specific limitations thereto.

[0073] In an embodiment of the present application, without the aid of an auxiliary light supplementing device, the environment when the shooting device of the terminal device captures the image to be recognized may be a low-light environment. Since the low-light image captured in the low-light environment has low image quality, it cannot be directly used for face recognition. Therefore, the illuminance of the current environment can be obtained by using an ambient light sensor, and then it can be determined whether the current environment is a low-light environment according to the illuminance. The ambient light sensor is a sensor used to sense the light intensity (hereinafter referred to as illuminance) of the surrounding environment. An ambient light sensor is usually installed in a terminal device (such as a mobile phone, a notebook, a tablet computer, etc.). The ambient light sensor is generally used to inform the processing chip of the terminal device of the illuminance of the environment, so as to automatically adjust the display brightness of the display of the terminal device, thereby reducing the power consumption of the display.

[0074] S102. When it is determined that the current environment is a low-light environment according to the illuminance, input the image to be recognized into a pre-trained image enhancement model to obtain an enhanced image of the image to be recognized.

[0075] In an embodiment of the present application, a low-light environment refers to an environment with insufficient illuminance. For example, an environment with an illuminance lower than a preset illuminance threshold. For example, the illuminance threshold can be 0.005 Lux. When determining whether the current environment is a low-light environment according to the illuminance, due to reasons such as the illuminance in the environment may change at any time and the fluctuation of the ambient light sensor, the average value of multiple consecutive illuminances can be used as the illuminance of the current environment, so as to avoid misjudging whether the current environment is a low-light environment due to instantaneous illuminance increase or decrease, and improve the judgment accuracy of the low-light environment. The method for determining whether the current environment is a low-light environment can also refer to the description of the Figure 4 illustrated embodiment below.

[0076] In an embodiment of the present application, the image to be recognized captured in a low-light environment can be called a low-light image, or can also be called a low-illuminance image, a low-light image, a low-quality image, etc. The image quality of a low-light image is generally low. For example, there may be a lack of image details, low contrast, a large amount of noise, etc., so that the image includes a lot of useless information and interference information, which will seriously affect the accuracy of algorithms such as face detection and face feature extraction in the face recognition process. The accuracy and reliability of face recognition can be improved by using a pre-trained image enhancement model and invoking the low-light image enhancement algorithm provided in the embodiment of the present application to enhance the low-light image to eliminate the noise in the image, and perform the first face feature matching on the image to be recognized.

[0077] In an embodiment of the present application, the image enhancement model includes a first model and a second model. The first model can also be referred to as a denoising neural network, which is used to remove image noise in low-light images; the second model can also be referred to as a brightening neural network, which is used to adaptively increase the image brightness.

[0078] In the related art, when training an image enhancement model, it is usually necessary to use real paired training samples to train the image enhancement model. Since each pair of training samples needs to include a low-quality image m and a high-quality image M corresponding to the low-quality image m, it is difficult to collect a large number of paired training samples to construct a sample set.

[0079] To solve the above problems, an embodiment of the present application provides a training strategy for unpaired training, that is, there is no need to use high-quality images corresponding one by one to real low-light low-quality images as training ground truth labels, and any low-quality image and high-quality image can be used to form a sample set to train the image enhancement model. Among them, the specific training method for the image enhancement model can refer to the description of the Figure 5 illustrated embodiment below.

[0080] In an embodiment of the present application, in the application inference stage of the image enhancement model, there is no need to adopt a complex strategy design. Refer to Figure 6 shown, which is a schematic diagram of the image enhancement model provided by the embodiment of the present application in the application inference stage. As Figure 6 , any real low-quality image C can be directly passed through the denoising neural network (such as the first model) and the brightening neural network (such as the second model) in sequence to obtain the final enhanced image C. Specifically, inputting the image to be recognized into the pre-trained image enhancement model to obtain the enhanced image of the image to be recognized includes: using the first model to denoise the image to be recognized and output the denoised image to be recognized; using the second model to increase the brightness of the denoised image to be recognized, and taking the output brightness-increased image as the enhanced image.

[0081] In the model application process of the above embodiment, the first model is used first, and then the output of the first model is used as the input of the second model; in other embodiments, the second model can also be used first, and then the output of the second model is used as the input of the first model. The present application does not limit this.

[0082] S103, extracting first face features from the enhanced image by using a preset face feature extraction model.

[0083] In an embodiment of the present application, the to-be-recognized image collected by the imaging device may not include a human face. Therefore, before extracting the first human face feature, it is also possible to first determine whether there is a to-be-recognized human face in the enhanced image; if there is a to-be-recognized human face in the enhanced image, extract the first human face feature of the to-be-recognized human face; if there is no to-be-recognized human face in the enhanced image, there is no need to execute the subsequent process, and it can be directly determined that the user's identity verification fails.

[0084] Specifically, extracting the first human face feature from the enhanced image by using a preset human face feature extraction model includes: based on a preset human face detection model, determining whether there is a to-be-recognized human face in the enhanced image; if there is a to-be-recognized human face in the enhanced image, intercepting the human face region image from the enhanced image; and extracting the first human face feature from the human face region image by using the human face feature extraction model.

[0085] Among them, the human face detection model is a model used to identify whether there is a human face and the position of the human face in an image or video frame. For example, the human face detection model can be a model based on a convolutional neural network (CNN) that learns human face features by training a large number of image samples containing human faces and non-human faces. The CNN-based human face detection model usually consists of multiple convolutional layers, pooling layers, and fully connected layers to extract the human face features in the image and generate a predicted bounding box for the human face region. The image within the predicted bounding box is the human face region image. The human face detection model can adopt existing detection models in related technologies. For example, the human face detection model can include but is not limited to the following models: Haar feature classifier, R-CNN (Region-based Convolutional Neural Networks) model, Fast R-CNN model, Faster R-CNN model, Single Shot MultiBox Detector (SSD) model, You Only Look Once (YOLO) model, etc. The embodiments of the present application do not limit this.

[0086] A face feature extraction model is a model used to extract face feature vectors from images. The face feature vectors can be used to indicate key information of a face, such as features like facial contours, eyes, nose, and mouth, as well as more abstract attributes such as age, gender, expression, etc. The face feature extraction model can adopt existing feature extraction models in related technologies. For example, the face feature extraction model can include but is not limited to the following models: VGGFace model, FaceNet model, DeepFace model, ArcFace model, etc. The embodiments of this application do not limit this. The face feature vectors extracted by the face feature extraction model usually have a fixed dimension and can be used for subsequent tasks such as face verification. For example, when the face feature extraction model is the VGGFace model, the first face feature extracted can be a first face feature vector with a dimension of 4096.

[0087] Through the above embodiments, when performing face feature recognition on the image to be enhanced, it is possible to first determine whether there is a face in the enhanced image and the regional image where the face exists, so as to specifically identify the face features in the face regional image. This enables the face feature extraction model to not need to consider the features of areas that do not contain faces, reducing the model's computing power and improving the efficiency and accuracy of face feature extraction.

[0088] S104, determine whether the first similarity between the first face feature and the second face feature in the preset face template image is greater than or equal to the preset similarity threshold. If the first similarity is greater than or equal to the similarity threshold, execute S105.

[0089] In an embodiment of this application, the face template image represents the face images of pre-acquired legitimate users. Among them, the number of legitimate users can be not unique, and the face template image includes the face images of each legitimate user. If the face recognition method determines that the face to be recognized in the image to be recognized matches the face in any face template image, it can be determined that the face in the image to be recognized belongs to a certain legitimate user, and then the user's identity verification passes.

[0090] In an embodiment of the present application, a face feature extraction model can be used to extract the second face feature in the face template image. For example, when the face feature extraction model is the VGGFace model, the extracted second face feature can be a second face feature vector with a dimension of 4096. The first face feature vector has the same dimension as the second face feature vector. A preset distance formula can be used to calculate the distance between the first face feature vector and the second face feature vector, and the first similarity between the first face feature and the second face feature can be determined according to this distance. Among them, the distance formula can include, but is not limited to, a combination of one or more of the following distance formulas: cosine distance formula, Euclidean distance formula, Chebyshev distance formula, etc. For example, the first similarity can be made inversely proportional to the cosine distance between the first face feature vector and the second face feature vector, and the formula used can include: First similarity = 0.5×(1 - cosine distance between the first face feature vector and the second face feature vector).

[0091] In an embodiment of the present application, the similarity threshold can be set according to actual needs. For example, the similarity threshold can be 0.9, and the present application does not make specific restrictions on this.

[0092] S105. Determine that the identity verification of the user is passed.

[0093] In an embodiment of the present application, if the first similarity is greater than or equal to the similarity threshold, it indicates that the face in the enhanced image and the face in the face template image belong to the same user. Therefore, it can be determined that the identity verification of the user is passed.

[0094] Through the above embodiments, when performing face recognition on a terminal device without the aid of a fill light, it is possible to determine whether the current is a low-light environment according to the illuminance collected by the ambient light sensor of the terminal device, so as to determine whether the to-be-recognized image captured by the imaging device is a low-light image; perform image enhancement on the low-light image captured in the low-light environment, so as to obtain an enhanced image with significantly reduced noise, appropriately restored brightness, and the face features not being significantly damaged; by performing feature matching on the face features in the enhanced image and the face features in the preset face template image, the accuracy of face recognition in the low-light environment can be improved.

[0095] In an embodiment of the present application, if the first similarity is less than the similarity threshold, it indicates that the above first face feature matching for the to-be-recognized image fails, and the second face feature matching can be further performed. The second face feature matching can match the face features of the to-be-recognized image before enhancement with the face features of the degraded face template image, and return the final result of the face feature matching.

[0096] Specifically, if the first similarity in S104 is less than the similarity threshold, the process proceeds to S106.

[0097] S106, perform image degradation processing on the face template image to obtain a degraded image of the face template image.

[0098] In an embodiment of the present application, the face template image can be subjected to image degradation processing by adding noise to the face template image, reducing the brightness of the face template image, etc. Specifically, the method for performing image degradation processing on the face template image can refer to the description of the Figure 7 illustrated embodiment. As Figure 7 shown, it is a flowchart of a method for performing image degradation processing on a face template image provided by an embodiment of the present application. The method is applied to a terminal device, and the method for performing image degradation processing on the face template image includes:

[0099] S401, obtain a preset third image.

[0100] In an embodiment of the present application, the third image represents a low-quality image, such as the first image that has not been used to train the image enhancement model, or the third image can also be an image to be recognized.

[0101] S402, input the third image into the image enhancement model, and use the first model of the image enhancement model to denoise the third image to obtain a third denoised image corresponding to the third image, and a second difference image between the third image and the third denoised image.

[0102] In an embodiment of the present application, the method in Figure 10 shown in S501 can be used to obtain the second difference image. For example Figure 8 shown, the image enhancement model can be used to denoise a low-quality image to be verified D (such as the third image) to obtain a denoised image D (such as the third denoised image), and a pure noise image D (such as the second difference image) can be obtained according to the low-quality image to be verified D and the denoised image D.

[0103] S403, reduce the brightness of the face template image to obtain an image with reduced brightness.

[0104] In an embodiment of the present application, the methods for reducing the brightness of a face template image include, but are not limited to, one or a combination of the following methods: (1) Linear adjustment: reducing the brightness value of each pixel point in the face template image by a fixed offset. For example, when the face template image is an RGB image, the RGB channel values of each pixel point in the face template image can be subtracted by a preset constant (such as 5); (2) Gamma correction: obtaining an updated brightness value by taking the power of the brightness value of each pixel point in the face template image. For example, let the updated brightness value = original brightness ^ gamma value; (3) Histogram equalization: increasing the contrast and reducing the brightness by redistributing the brightness values of the pixel points in the face template image. For example, stretching the grayscale histogram of the face template image to achieve the redistribution of the brightness values of the pixel points in the face template image. For example Figure 8 As shown, after reducing the brightness of the template image E (such as the face template image), a darkened template image E (such as the image after brightness reduction) is obtained.

[0105] S404, performing image fusion on the second difference image and the image after brightness reduction to obtain a degraded image of the face template image.

[0106] In an embodiment of the present application, the image after brightness reduction can be converted into a grayscale image, and then pixel-level fusion is performed on the second difference image and the grayscale image after brightness reduction to obtain a degraded image of the face template image. For example Figure 8 As shown, after fusing the pure noise image D (such as the second difference image) with the darkened template image E (such as the image after brightness reduction), a degraded template image E (such as the degraded image) is obtained.

[0107] Through the above embodiments, there is no need to separately train a degradation neural network for image degradation. Instead, the image enhancement model can be directly used to obtain the pure noise map corresponding to the low-quality image D (such as the third image), and this pure noise map is used as the noise image (such as the second difference image) for degrading the face template. When the denoising model (such as the first model) in the aforementioned image enhancement model is properly and sufficiently trained, the degraded image here will have very realistic noise interference and an approximate noise distribution to the real low-quality image D. At this time, face features are extracted from the degraded image, and a matching algorithm is used to compare the face features in the face template image with the face features in the degraded features, which can, to a certain extent, weaken the influence of noise and focus on the face feature differences, thereby improving the accuracy of face recognition.

[0108] S107, extracting the third face features from the image to be recognized using the face feature extraction model, and extracting the fourth face features from the degraded image using the face feature extraction model.

[0109] In an embodiment of the present application, the third face feature may be the face feature vector of the image to be recognized, and the fourth face feature may be the face feature vector of the degraded image.

[0110] S108. Determine whether the second similarity between the third face feature and the fourth face feature is greater than or equal to the similarity threshold. If the second similarity is greater than or equal to the similarity threshold, execute S105; or, if the second similarity is less than the similarity threshold, execute S109.

[0111] In an embodiment of the present application, a distance formula such as cosine similarity or Euclidean distance can be used to calculate the second similarity between the third face feature and the fourth face feature.

[0112] In an embodiment of the present application, if the second similarity is greater than or equal to the similarity threshold, it indicates that the face to be recognized in the image to be recognized and the face in the degraded image belong to the same user. Therefore, it can be determined that the user's identity verification is passed.

[0113] S109. Determine that the user's identity verification fails.

[0114] In an embodiment of the present application, if both of the above two face feature matching processes fail, it can be determined that the user's identity verification fails.

[0115] Through the above embodiments, if the result of face feature matching using the enhanced image fails, the face template image can be further degraded by adding the noise of the low-light image to the face template image, and a second matching attempt can be made using the degraded image and the low-light image. For example, the face features of the low-light image before enhancement and the degraded template image are compared. By introducing the face template image degradation strategy, the false rejection rate (FRR) of face recognition in low-light scenarios can be further reduced, and the accuracy of face recognition and the user experience can be improved.

[0116] Refer to Figure 4 As shown in the figure, it is a flowchart of a method for determining whether the current environment is a low-light environment provided by an embodiment of the present application. The method is applied to a terminal device. The method for determining whether the current environment is a low-light environment includes:

[0117] S201. Use an ambient light sensor to collect multiple consecutive illuminances of the current environment.

[0118] In an embodiment of the present application, multiple consecutive illuminances represent multiple illuminances with a consecutive chronological order. The number of multiple consecutive illuminances can be set according to the sampling frequency of the ambient light sensor. For example, this number can be proportional to the sampling frequency of the ambient light sensor, and the present application does not limit this. For example, if the sampling frequency of the ambient light sensor is 10 times per second, then this number can be 5.

[0119] S202. Based on the moving average algorithm, determine the average value of multiple consecutive illuminances.

[0120] In an embodiment of the present application, the specific process of determining the average value of multiple consecutive illuminances based on the moving average algorithm includes: determining the size of the window, that is, the number of illuminances that the window can accommodate. The size of the window can be less than or equal to the total number of consecutive illuminances. For example, if the number of multiple consecutive illuminances is 5, the window size can be 3; initializing the window position, making the first data in the window be the first illuminance among the multiple consecutive illuminances, and calculating the average value p1 of the illuminances in the window; moving the window position one data position backward, making the first data in the window be the second illuminance among the multiple consecutive illuminances, and calculating the average value p2 of the illuminances in the window; and so on, continuously moving the sliding window backward until the last data in the window is the last illuminance among the multiple consecutive illuminances, and taking the position where the window is located at this time as the last placement position of the window; determining the number m of the window placement positions (also called moving positions) in the above process and the average value pi of the illuminances in the window at each position, and determining the average value of multiple consecutive illuminances according to ∑pi / m, where i represents a natural number.

[0121] S203. Determine whether the above average value is less than a preset illuminance threshold. If the average value is less than the illuminance threshold, execute S204. Or, if the average value is greater than or equal to the illuminance threshold, execute S205.

[0122] In an embodiment of the present application, the illuminance threshold can be set according to the sensitivity of the shooting device. For example, the illuminance threshold can be proportional to the sensitivity of the shooting device, and the present application does not limit this. For example, the illuminance threshold can be set to 0.005 Lux.

[0123] S204. Determine that the current environment is a low-light environment.

[0124] In an embodiment of the present application, in the case of determining that the current environment is a low-light environment according to the illuminance, the description of the embodiment shown in S102 can be referred to.

[0125] S205. Determine that the current environment is not a low-light environment.

[0126] Through the above embodiments, the average value of multiple consecutive illuminances can be determined by a moving average algorithm as the illuminance of the current environment, avoiding misjudgment of whether the current environment is a low-light environment due to instantaneous increases or decreases in illuminance, fluctuations of the ambient light sensor, etc., and improving the judgment accuracy of the low-light environment.

[0127] In an embodiment of the present application, when it is determined according to the illuminance that the current environment is not a low-light environment, it can be considered that the image to be recognized is a high-quality image that can be directly used for face recognition and is not in low light. At this time, the method for face recognition of the high-quality image is a traditional face recognition method, which can refer to the description of the Figure 12 embodiment shown below.

[0128] The application process of the image enhancement model is described in the above embodiments. Next, the process of training the image enhancement model based on the training strategy of unpaired training will be described.

[0129] First, the training strategy of unpaired training used in the present application is introduced. Refer to the Figure 9 figure shown below, which is a schematic diagram of the training strategy of unpaired training provided by the embodiment of the present application. Specifically, the training strategy of unpaired training includes: after any low-quality image A (hereinafter referred to as the first image) passes through the denoising neural network (such as the first model) of the image enhancement model, it is expected to obtain a denoised image A and a difference map between the denoised image A and the low-quality image A (hereinafter referred to as the first difference image, which is a pure noise map); the pure noise map corresponding to the low-quality image A is superimposed on any real high-quality image B (hereinafter referred to as the second image), and the low-quality image B (such as the fused image corresponding to the second image) can be synthesized; the real high-quality image B can thus be used as the enhancement target for synthesizing the low-quality image B to train the denoising neural network, that is, after the synthesized low-quality image B is denoised using the denoising neural network, the obtained denoised image B should be as close as possible to the high-quality image B. Further, the real high-quality image B can also be used as the enhancement target for synthesizing the low-quality image B to train the brightening neural network (such as the second model), so that the denoised image B before brightening (such as the second denoised image) is as close as possible to the original high-quality image B after passing through the brightening neural network.

[0130] Among them, during the training process of the first model, a denoising neural network was used for denoising twice, and each denoising corresponded to a loss function. For example, during the first denoising, the low-quality image A was denoised to obtain a pure noise image, so the first loss function can be used to constrain the noise in the pure noise image, that is, to measure the degree of noise existence in the pure noise image; during the second denoising, the synthesized low-quality image B was denoised to obtain the denoised image B, and the second loss function can be used to constrain the denoised image B so that the denoised image B can be as close as possible to the high-quality image B. In addition, a third loss function can be used to constrain the brightening neural network (such as the second model) of the image enhancement model, so that the denoised image B before brightening (such as the second denoised image) is as close as possible to the original high-quality image B after passing through the brightening neural network. By adding loss functions to form constraints in the above three links, the model training can be prevented from falling into local optima and the usage performance of the model can be improved.

[0131] After that, the image enhancement model is trained according to the training strategy of the above non-paired training. Refer to Figure 5 As shown, it is a flowchart of a training method for an image enhancement model provided by an embodiment of the present application. The method is applied to a terminal device, and the training method for the image enhancement model includes:

[0132] S301, obtain multiple sample images.

[0133] In an embodiment of the present application, the multiple sample images include multiple first images and multiple second images. Based on the training strategy of non-paired training provided by the embodiment of the present application, multiple (for example, 10,000) arbitrary low-quality images (for example, images with more noise and lower brightness) and multiple (for example, 10,000) arbitrary high-quality images (for example, images with less noise and higher brightness) can be obtained from an open-source database, so as to obtain multiple (for example, 20,000) sample images. Among them, the first image represents a low-quality image, the second image represents a high-quality image, and the image noise of any first image is more than that of any second image.

[0134] In an embodiment of the present application, after obtaining multiple sample images, the sample images can also be preprocessed so that the sample images can be applicable to the subsequent model training process. The preprocessing can include, but is not limited to, a combination of one or more of the following processing methods: image cropping, image correction, contrast adjustment, and grayscale processing.

[0135] S302, construct a training set and a test set according to the multiple sample images, and use the training set to perform at least one iterative update on a preset initial model until an image enhancement model that meets the preset requirements is obtained.

[0136] In one embodiment of the present application, the multiple sample images can be divided into a training set and a test set according to a preset ratio (e.g., 8:2). For example, 80% of the images are selected from multiple first images to construct the training set, and 80% of the images are also selected from multiple second images to construct the training set, and the remaining unselected images are used as the images in the test set. In another embodiment, the multiple sample images can also be divided into a training set, a test set, and a validation set. For example, the multiple sample images are divided into a training set, a test set, and a validation set according to a ratio of 7:2:1.

[0137] In one embodiment of the present application, the initial model can be a model obtained by initializing a preset neural network (e.g., a graph-to-graph conversion neural network). The initial model can include a first initial model and a second initial model. Among them, the first initial model can be a denoising neural network for removing image noise, and the second initial model can be a brightening neural network for enhancing the image brightness. Specifically, reference can be made to the introduction in the subsequent embodiments.

[0138] In one embodiment of the present application, the process of initializing a preset neural network to obtain an initial model may include, but is not limited to: defining the network structure of the model. For example, determining the type, number, and connection method of the layers included in the model's network; initializing the parameters of the model. For example, initializing the weights and biases of each network layer of the model to random values or specific preset values. Common methods include zero initialization, random normal distribution initialization (such as Xavier initialization), etc.; defining the loss function of the model. The loss function can be used to measure the difference between the model output and the target value. Based on the non-paired training strategy used in the present application, the loss function of the model can include a first loss function, a second loss function, and a third loss function. Specifically, reference can be made to the description in the subsequent embodiments; determining the optimization algorithm of the model. For example, optimization algorithms such as stochastic gradient descent and Adam algorithm. The optimization algorithm can be used to iteratively update the model parameters to minimize the loss function; determining the sample batch used when training the model. The sample batch is used to determine the number of samples used in each iterative training. Usually, it is trained in the form of mini-batches, which can improve the training efficiency and stability. For example, one first image is used each time.

[0139] In one embodiment of the present application, when performing at least one iterative update on the initial model, the input data (such as the data in the training set) can be propagated forward through the network to obtain an output result; calculate the error between the output result and the ideal result according to the loss function, calculate the gradient through the backpropagation algorithm, and update the parameters in the network; repeat the above process of forward propagation and backpropagation, and gradually optimize the model parameters through multiple iterations so that the model can better fit the training data until an image enhancement model that meets the preset requirements is obtained. After that, the data in the validation set can also be used to verify the performance of the image enhancement model. When the verification result does not meet the expectation, the training set is updated and the updated training set is used to tune the model, so as to avoid overfitting of the model. Further, the test set can also be used to further test the model to further improve the generalization ability of the model.

[0140] In one embodiment of the present application, since the initial model includes a first initial model and a second initial model, at least one iterative update can be performed on the first initial model and the second initial model respectively. The preset requirements include one or more combinations of the following requirements: the loss function corresponding to the initial model reaches the convergence condition; the test result obtained by testing the initial model using the test set indicates that the model performance (such as accuracy, recall rate, F1 score, etc.) of the initial model reaches the expectation. For example, the accuracy of the model reaches the preset accuracy threshold (such as 0.9); the number of times of the at least one iterative update reaches the preset iterative number threshold (such as 10,000 times). Among them, the loss function can include a first loss function corresponding to the first initial model and a second loss function, and the loss function can also include a third loss function corresponding to the second initial model. The loss function reaching the convergence condition can include but is not limited to: the sum value of the first loss value corresponding to the first loss function and the second loss value corresponding to the second loss function reaches a preset threshold (hereinafter referred to as the fourth threshold); the third loss value corresponding to the third loss function is less than or equal to a preset threshold (hereinafter referred to as the third threshold).

[0141] The process of any update in the at least one iterative update can also refer to the description of the embodiment shown below for Figures 10 to 11 the illustrated embodiment.

[0142] In one embodiment of the present application, at least one iterative update can be first performed on the first initial model. Refer to Figure 10 As shown, it is a flowchart of a method for any update of the first initial model provided by an embodiment of the present application. The method is applied to a terminal device, and the process of the method for any update of the first initial model includes:

[0143] S501. Input any first image in the training set into the first initial model, and use the first initial model to denoise any first image to obtain the first denoised image corresponding to any first image and the first difference image between any first image and the first denoised image.

[0144] In an embodiment of the present application, the first initial model can be a general graph-to-graph conversion neural network, which can be selected according to actual needs, such as screening according to requirements such as specific scenario performance, parameter quantity limitation of the landing scenario hardware, and speed limitation.

[0145] For example, the first initial model can include, but is not limited to, a combination of one or more of the following models: Autoencoder model: It includes an encoder and a decoder. The encoder compresses the input image (such as the first image) into a low-dimensional representation and then the decoder decodes it back to an image representation (such as the first denoised image), which can effectively remove the noise in the image; U-Net model: It is a convolutional neural network with a symmetric U-shaped structure, which can capture features of different scales of the input image and use a skip connection mechanism to connect different network layers in the decoding stage, thereby improving the performance of image denoising.

[0146] In an embodiment of the present application, the method for obtaining the first difference image can include: obtaining the grayscale image corresponding to the first image (hereinafter referred to as the first grayscale image) and the grayscale image corresponding to the first denoised image (hereinafter referred to as the second grayscale image); calculating the difference between the grayscale values of each pixel point in the first grayscale image and the pixel point at the corresponding position in the second grayscale image, and using the obtained difference as the pixel value of the corresponding pixel point in the first difference image to obtain the first difference image.

[0147] In an embodiment of the present application, the first difference image is a pure noise image that only contains noise, and the preliminary denoising performance of the first initial model can be determined by analyzing the noise in the first difference image.

[0148] In an embodiment of the present application, at the next update after each update, another first image is used to perform the next update on the first initial model. Wherein, another first image refers to a first image that has not been used in the previous update.

[0149] S502. Use a preset first loss function to determine the first loss value corresponding to the first difference image.

[0150] In an embodiment of the present application, the first loss value calculated by the first loss function is used to indicate the noise level (or the degree of noise presence) in the first difference image. The larger the first loss value, the more noise there is in the first difference image.

[0151] Specifically, the first loss function can be a statistical analysis function. By calculating the mean, variance, standard deviation, etc. of the gray values of all pixel points in the first difference image, a statistical analysis function can be constructed based on these values as the first loss function. For example, the first loss function can be a weighted sum function of the mean, variance, and standard deviation. The first loss function can also be a spectral analysis function. By using methods such as the fast Fourier transform, the first difference image can be transformed into the frequency domain, so as to perform spectral analysis on the first difference image, determine the distribution of the signal energy of the first difference image in the high-frequency range corresponding to the noise, and then construct a spectral analysis function based on this distribution. For example, this spectral analysis function can be the distribution probability of the signal energy of the first difference image in the high-frequency range corresponding to the noise.

[0152] S503, determine whether the first loss value is less than or equal to a preset first threshold. If the first loss value is less than or equal to the first threshold, execute S501; or, if the first loss value is greater than the first threshold, execute S504.

[0153] In an embodiment of the present application, the first threshold can be set according to actual needs. For example, the first threshold can be proportional to the size of the first image. The present application does not make specific limitations on this.

[0154] In an embodiment of the present application, the first loss value being less than or equal to the first threshold indicates that the degree of noise existence in the first difference image is relatively low. Therefore, the denoising performance of the first initial model in this iteration is relatively low, and the next batch of training samples (such as another first image) need to be used for the next update.

[0155] In another embodiment, it is also possible not to judge the first loss value, that is, the process can be transferred to S504 after S502, and then combined with the second loss value in the subsequent process S507 to determine whether to perform the next update on the first initial model.

[0156] S504, fuse the first difference image with any second image in the training set to obtain a fused image corresponding to the second image.

[0157] In an embodiment of the present application, the first loss value being greater than the first threshold indicates that the first initial model can be used for preliminary denoising in this iteration, but the performance of whether the first initial model can denoise and restore a low-quality image to the corresponding high-quality image still needs to be further verified.

[0158] Based on the training strategy of non-paired training in this application, there are no paired low-quality images (such as images with more noise) and high-quality images (such as images with less noise) corresponding to each other in the training samples used in this application. Therefore, a pure noise image (such as the first difference image) of the intermediate result generated during the training process of the above first initial model can be used to reduce the quality of a high-quality image (such as the second image) to obtain a low-quality image corresponding to the second image (such as the fused image corresponding to the second image), so as to realize the acquisition of paired low-quality images and high-quality images.

[0159] In an embodiment of this application, when fusing the first difference image with any second image in the training set, a grayscale image corresponding to the second image (hereinafter referred to as the third grayscale image) can be obtained, and the grayscale value of each pixel point in the first difference image is added to the grayscale value of the corresponding pixel point in the third grayscale image using the per-pixel fusion method, thereby generating a fused image corresponding to the second image.

[0160] S505, input the fused image into the first initial model, and use the first initial model to denoise the fused image to obtain a second denoised image corresponding to the fused image.

[0161] In an embodiment of this application, the first initial model is expected to denoise the input fused image so that the obtained second denoised image is infinitely close to the original second image.

[0162] S506, use a preset second loss function to determine the second loss value between the second denoised image and any corresponding second image.

[0163] In an embodiment of this application, the second loss value calculated by the second loss function is used to measure the difference between the second denoised image and the corresponding any second image. The larger the second loss value, the greater the difference between the second denoised image and the corresponding any second image. The second loss function can include, but is not limited to, a combination of one or more of the following functions: L1 norm, perceptual loss function.

[0164] S507, determine whether the second loss value is less than or equal to a preset second threshold. If the second loss value is greater than the second threshold, execute S501; or, if the second loss value is less than or equal to the second threshold, end the iterative update of the first initial model.

[0165] In an embodiment of this application, the second threshold can be set according to actual needs. For example, the second threshold can be 0.1, and this application does not make specific limitations on this.

[0166] In an embodiment of the present application, that the second loss value is greater than the second threshold indicates that the difference between the second denoised image and any one of the corresponding second images is large. Therefore, the denoising performance of the first initial model in this iteration is low, and the next batch of training samples need to be used for the next update. Specifically, the first initial model can be updated next using the next first image and the next second image. In another embodiment, if the second loss value is greater than the second threshold, the first initial model can also be updated next using the next first difference image and the next second image, that is, the process proceeds to S504.

[0167] In an embodiment of the present application, if the second loss value is less than or equal to the second threshold, it can be considered that the denoising performance of the first initial model obtained in this iteration meets the expectation. Therefore, the iterative update of the first initial model can be stopped, and at least one iterative update of the second initial model can be performed.

[0168] In another embodiment, if the above S503 is not executed, that is, the process directly proceeds to S504 after S502, it is also possible to determine whether to perform the next update of the first initial model using the sum value of the first loss value and the second loss value in S507. For example, if the sum value of the first loss value and the second loss value is less than or equal to the fourth threshold, the first initial model is updated next using the next first image and the next second image; or, if the sum value of the first loss value and the second loss value is greater than the fourth threshold, the iterative update of the first initial model is stopped, and at least one iterative update of the second initial model is performed.

[0169] Refer to Figure 11 As shown, it is a flowchart of a method for any update of the second initial model provided by an embodiment of the present application. The method is applied to a terminal device, and the process of the method for any update of the second initial model includes:

[0170] S601, input the second denoised image corresponding to any second image into the second initial model, and use the second initial model to enhance the brightness of the second denoised image to obtain a brightness-enhanced image corresponding to the second denoised image.

[0171] In an embodiment of the present application, based on the training strategy of non-paired training of the present application, there are no paired low-quality images (such as images with lower brightness) and high-quality images (such as images with higher brightness) corresponding to each other in the training samples used in the present application. Therefore, the second denoised image corresponding to any second image generated as an intermediate result during the training of the first initial model can be used as the low-quality image, and the any second image corresponding to the second denoised image can be used as the high-quality image, so as to obtain paired low-quality images and high-quality images to train the second initial model.

[0172] In an embodiment of the present application, the second initial model may be a general graph-to-graph conversion neural network, which can be selected according to actual needs, for example, screened according to requirements such as specific scenario performance, parameter quantity limitations of the landing scenario hardware, and speed limitations.

[0173] The second initial model may include, but is not limited to, a combination of one or more of the following models: Contrast enhancement network: including a convolutional layer, a batch normalization layer, and an activation function, which can perform contrast enhancement operations on the input image (such as the second denoised image) to make the image more distinct and clear, thereby improving the image brightness; Gamma correction network: can adjust the brightness and contrast of the image through non-linear operations, thereby improving the display effect of the image.

[0174] In an embodiment of the present application, at the next update after each update, the second initial model is updated next time using the second denoised image corresponding to the next second image. Among them, the second denoised image corresponding to the next second image refers to the second denoised image corresponding to the second image that has not been used in the previous update.

[0175] S602, determining a third loss value between the brightness-enhanced image and any one of the corresponding second images using a preset third loss function.

[0176] In an embodiment of the present application, the third loss value calculated by the third loss function is used to indicate the difference between the brightness-enhanced image and any one of the corresponding second images. The larger the third loss value, the greater the difference between the brightness-enhanced image and any one of the corresponding second images. The third loss function may include, but is not limited to, a combination of one or more of the following functions: L1 norm, perceptual loss function.

[0177] S603, determining whether the third loss value is less than or equal to a preset third threshold. If the third loss value is greater than the third threshold, execute S601; or, if the third loss value is less than or equal to the third threshold, execute S604.

[0178] In an embodiment of the present application, the third threshold can be set according to actual needs. For example, the third threshold can be 0.1, and the present application does not make specific limitations on this.

[0179] In an embodiment of the present application, that the third loss value is greater than the third threshold indicates that the difference between the brightness-enhanced image and any one of the corresponding second images is relatively large. Therefore, the brightness enhancement performance of the second initial model in this iteration is relatively low, and the next update needs to use the training samples of the next batch. Specifically, the second initial model can be updated next time using the second denoised image corresponding to the next second image.

[0180] Execute S604, determine that the initial model meets the preset requirements, and use the initial model as the image enhancement model.

[0181] In one embodiment of the present application, if the third loss value is less than or equal to the third threshold, it can be considered that the brightness improvement of the second initial model obtained in this iteration meets the expectation; since the training of the first initial model has also been completed before, the iterative update of the initial model can be stopped, it is determined that the initial model meets the preset requirements, and the initial model is used as the image enhancement model.

[0182] Through the above embodiments, compared with the conventional solutions including preprocessing operations such as noise removal or brightness enhancement, the denoising neural network provided in the present application is trained with unpaired real data, without assuming the noise distribution, without restricting the usage scenario, and without manually adjusting the degradation parameters, and has better generalization performance. In addition, the above embodiments connect the denoising neural network and the brightening neural network in series, improving the usage performance and interpretability of the image enhancement model.

[0183] Refer to Figure 12 As shown, it is a flowchart of a method for face recognition of high-quality images provided by an embodiment of the present application. The method is applied to a terminal device, and the method for face recognition of high-quality images includes:

[0184] S701, based on a preset face detection model, determine whether there is a face to be recognized in the image to be recognized; if there is a face to be recognized in the image to be recognized, execute S702; if there is no face to be recognized in the image to be recognized, execute S706.

[0185] S702, intercept the face region image from the image to be recognized.

[0186] In one embodiment of the present application, referring to the relevant description in S103, according to the result of face recognition, using image processing techniques such as cropping and scaling, the corresponding face region image is intercepted from the image to be recognized. The intercepted face region image may include the entire face for subsequent processing and recognition.

[0187] S703, use the face feature extraction model to extract the fifth face feature from the face region image.

[0188] In one embodiment of the present application, using the face feature extraction model to extract the fifth face feature from the face region image can refer to the relevant description in S103.

[0189] S704, determine whether the third similarity between the fifth face feature and the second face feature is greater than or equal to the similarity threshold. If it is determined that the third similarity is greater than or equal to the similarity threshold, execute S705; or, if it is determined that the third similarity is less than the similarity threshold, execute S706.

[0190] In an embodiment of the present application, the extraction and comparison of face features may refer to the methods adopted in related technologies, and the embodiments of the present application do not limit this. Specifically, reference may also be made to the relevant description in S103.

[0191] S705. Determine that the user's identity verification is passed.

[0192] S706. Determine that the user's identity verification fails.

[0193] Through the above technical solution, when it is determined that the current environment is not a low-light environment, it can be considered that the to-be-recognized image captured is not a low-light image, and thus there is no need to perform image enhancement and other processing; the to-be-recognized image can be directly used for face recognition, and the face area in the to-be-recognized image can be intercepted; then, face features are extracted from the intercepted face area image, and feature matching is performed with the face features in the face template image, so as to determine the face recognition result, and efficient and accurate face recognition can be achieved in a non-low-light environment.

[0194] Refer to Figure 13 shown in the figure, which is a schematic diagram of the overall architecture of the face recognition method provided by an embodiment of the present application. It can be seen from Figure 13 that compared with the module architecture of a conventional face recognition system, the face recognition method of the present application adds a low-light image enhancement algorithm and a template image degradation strategy.

[0195] Specifically, a face recognition system generally includes two parts: a face detection algorithm and a feature extraction and matching algorithm. The former determines whether a face can be found and detected in an image, and the latter is used to determine whether the detected face matches a preset template face. However, low-brightness and high-noise images under low light have a serious impact on both detection and feature extraction and matching.

[0196] The low-light image enhancement algorithm provided by the embodiments of the present application is applied before the face detection algorithm. As a graph-to-graph conversion neural network model algorithm, the low-light image enhancement algorithm can take the original low-light image as input and output an enhanced image with significantly reduced noise, appropriately restored brightness, and no obvious damage to face features. This enhanced image can be used in face detection, face feature extraction, and face feature matching modules. In another embodiment, the low-light image enhancement algorithm can also be applied after the face detection algorithm to perform image enhancement on the face area image determined by the face detection algorithm.

[0197] The template image degradation strategy provided by the embodiments of this application does not require the introduction of a new neural network model. Instead, based on the derivative of the aforementioned low-light image enhancement algorithm, i.e., the input image noise, a secondary matching mechanism is introduced. Specifically, the embodiments of this application will introduce a double matching process for low-light images. The first match is an attempt to match the enhanced image with the template image. If it is unsuccessful, the noise of the low-light image is added to the template image to degrade the template image, and a second match attempt is made, that is, the features of the low-light image before enhancement are compared with the features of the degraded template image. By introducing this template image degradation strategy, the false rejection rate in low-light scenarios can be further reduced, improving the user experience.

[0198] Refer to Figure 14 As shown, it is a flowchart of a face recognition method provided by another embodiment of this application. Figure 14 The flowchart shown can be briefly described as follows: When the face recognition system is activated, the camera captures the original image (such as the image to be recognized), and after basic preprocessing of the original image (such as image correction, size adjustment, etc.), the ambient light sliding average value collected by the ambient light sensor is obtained. If the ambient light sliding average value is greater than the preset illuminance threshold p, the current environment is not a low-light environment, and at this time, it directly enters the conventional face detection and recognition process, and directly returns the recognition result. If the ambient light sliding average value is less than the illuminance threshold p, the current environment is a low-light environment, and it enters the first matching process for the low-light environment.

[0199] In the first matching process for the low-light environment, the low-light image enhancement algorithm is used to enhance the original image, and the enhanced image is used for face detection and face recognition algorithms, and an attempt is made to match the features of the enhanced face image with the features of the template face image (or called the face template image). If the result is a match, the result is directly returned; if the result is still a mismatch, the second matching process for the low-light environment is executed. In the second matching process for the low-light environment, based on the intermediate result in the first matching process, the low-light image noise extracted by the aforementioned image enhancement algorithm is used to actively degrade the template face image, and then the features of the low-light face image before enhancement are matched with the features of the degraded template face image, and the final result of match or mismatch is returned.

[0200] Through the above embodiments, when performing face recognition on a terminal device using a non-supplementary light camera, it is possible to determine whether the current environment is a low-light environment based on the illuminance collected by the ambient light sensor of the terminal device, so as to determine whether the image to be recognized captured by the imaging device is a low-light image; perform image enhancement on the low-light image captured in the low-light environment, so as to obtain an enhanced image with significantly reduced noise, appropriately restored brightness, and the facial features not being significantly damaged; by performing feature matching between the facial features in the enhanced image and the facial features in the preset facial template image, the accuracy of face recognition in the low-light environment can be improved. If the result of face feature matching using the enhanced image fails, the template image can be further degraded by adding the noise of the low-light image to the facial template image, and the degraded image is used to perform a second matching attempt with the low-light image, that is, comparing the facial features of the low-light image before enhancement with the degraded template image. By introducing the facial template image degradation strategy, the false rejection rate (FRR) of face recognition in the low-light scenario can be further reduced, and the accuracy and user experience of face recognition can be improved.

[0201] An embodiment of the present application further provides a terminal device 100. Refer to Figure 15 As shown, the terminal device 100 may be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, as well as a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) device, a virtual reality (VR) device, an artificial intelligence (AI) device, a wearable device, a vehicle-mounted device, a smart home device, and / or a smart city device. The specific type of the terminal device 100 is not particularly limited in the embodiment of the present application.

[0202] The terminal device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a Universal Serial Bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a Subscriber Identification Module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0203] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the terminal device 100. In other embodiments of the present application, the terminal device 100 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0204] The processor 110 may include one or more processing units. For example, the processor 110 may include an Application Processor (AP), a modem processor, a Graphics Processing Unit (GPU), an Image Signal Processor (ISP), a controller, a video codec, a Digital Signal Processor (DSP), a baseband processor, and / or a Neural-network Processing Unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0205] The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.

[0206] A memory may also be provided in the processor 110 for storing instructions and data. In an embodiment of the present application, the memory in the processor 110 is a cache memory. The memory may store instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instructions or data again, it can directly call them from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0207] In an embodiment of the present application, the processor 110 may include one or more interfaces. The interfaces may include an Inter-integrated Circuit (I2C) interface, an Inter-integrated Circuit Sound (I2S) interface, a Pulse Code Modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a Mobile Industry Processor Interface (MIPI), a General-Purpose Input / Output (GPIO) interface, a Subscriber Identity Module (SIM) interface, and / or a Universal Serial Bus (USB) interface, etc.

[0208] The I2C interface is a bidirectional synchronous serial bus that includes a Serial Data Line (SDA) and a Serial Clock Line (SCL). In an embodiment of the present application, the processor 110 may include multiple groups of I2C buses. The processor 110 may be respectively coupled to the touch sensor 180K, the charger, the flashlight, the camera 193, etc. through different I2C bus interfaces. For example: the processor 110 may be coupled to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface to implement the touch function of the terminal device 100.

[0209] The I2S interface can be used for audio communication. In an embodiment of the present application, the processor 110 may include multiple groups of I2S buses. The processor 110 can be coupled to the audio module 170 through the I2S bus to achieve communication between the processor 110 and the audio module 170. In an embodiment of the present application, the audio module 170 can transmit an audio signal to the wireless communication module 160 through the I2S interface to implement the function of answering a call through a Bluetooth headset.

[0210] The PCM interface can also be used for audio communication to sample, quantize, and encode analog signals. In an embodiment of the present application, the audio module 170 and the wireless communication module 160 can be coupled through the PCM bus interface. In an embodiment of the present application, the audio module 170 can also transmit an audio signal to the wireless communication module 160 through the PCM interface to implement the function of answering a call through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.

[0211] The UART interface is a general-purpose serial data bus for asynchronous communication. The bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In an embodiment of the present application, the UART interface is generally used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 through the UART interface to implement the Bluetooth function. In an embodiment of the present application, the audio module 170 can transmit an audio signal to the wireless communication module 160 through the UART interface to implement the function of playing music through a Bluetooth headset.

[0212] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a Camera Serial Interface (CSI), a Display Serial Interface (DSI), etc. In an embodiment of the present application, the processor 110 and the camera 193 communicate through the CSI interface to implement the shooting function of the terminal device 100. The processor 110 and the display screen 194 communicate through the DSI interface to implement the display function of the terminal device 100.

[0213] The GPIO interface can be configured through software. The GPIO interface can be configured as a control signal or a data signal. In an embodiment of the present application, the GPIO interface can be used to connect the processor 110 to the camera 193, the display screen 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0214] The USB interface 130 is an interface that complies with the USB standard specification. Specifically, it can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the terminal device 100, and can also be used to transfer data between the terminal device 100 and peripheral devices. It can also be used to connect headphones to play audio. The interface can also be used to connect other terminal devices 100, such as AR devices, etc.

[0215] It can be understood that the interface connection relationship between the modules illustrated in the embodiments of the present invention is only for illustrative purposes and does not constitute a structural limitation on the terminal device 100. In other embodiments of the present application, the terminal device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0216] The charging management module 140 is used to receive a charging input from a charger. Among them, the charger can be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 140 can receive the charging input of the wired charger through the USB interface 130. In some embodiments of wireless charging, the charging management module 140 can receive the wireless charging input through the wireless charging coil of the terminal device 100. While charging the battery 142, the charging management module 140 can also supply power to the terminal device 100 through the power management module 141.

[0217] The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110. The power management module 141 receives the inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the display screen 194, the camera 193, the wireless communication module 160, etc. The power management module 141 can also be used to monitor parameters such as the battery capacity, the number of battery cycles, and the battery health status (leakage, impedance). In some other embodiments, the power management module 141 can also be provided in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be provided in the same device.

[0218] The wireless communication function of the terminal device 100 can be implemented by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.

[0219] Antenna 1 and Antenna 2 are used for transmitting and receiving electromagnetic wave signals. Each antenna in the terminal device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, Antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0220] The mobile communication module 150 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the terminal device 100. The mobile communication module 150 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves through Antenna 1, and perform processing such as filtering and amplifying on the received electromagnetic waves, and then transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves through Antenna 1 and radiate it out. In an embodiment of the present application, at least some functional modules of the mobile communication module 150 can be disposed in the processor 110. In an embodiment of the present application, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 can be disposed in the same device.

[0221] The modulation and demodulation processor can include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, receiver 170B, etc.), or displays an image or video through the display screen 194. In an embodiment of the present application, the modulation and demodulation processor can be an independent device. In some other embodiments, the modulation and demodulation processor can be independent of the processor 110 and be disposed in the same device as the mobile communication module 150 or other functional modules.

[0222] The wireless communication module 160 may provide solutions for wireless communications applied to the terminal device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The wireless communication module 160 may be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 may also receive the signals to be sent from the processor 110, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through the antenna 2 for radiation.

[0223] In an embodiment of the present application, antenna 1 of the terminal device 100 is coupled to the mobile communication module 150, and antenna 2 is coupled to the wireless communication module 160, so that the terminal device 100 can communicate with the network and other devices through wireless communication technologies. The wireless communication technologies may include Global System For Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), Beidou Navigation Satellite System (BDS), Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS).

[0224] The terminal device 100 realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.

[0225] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a Microled, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In an embodiment of the present application, the terminal device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1.

[0226] The terminal device 100 can implement the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.

[0227] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and the light passes through the lens and is transmitted to the camera photosensitive element. The optical signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize the noise, brightness, and skin color of the image through algorithms. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In an embodiment of the present application, the ISP can be set in the camera 193.

[0228] The camera 193 is used to capture static images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard RGB, YUV, etc. formats. In an embodiment of the present application, the terminal device 100 may include 1 or N cameras 193, where N is a positive integer greater than 1.

[0229] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the terminal device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0230] The video codec is used to compress or decompress digital videos. The terminal device 100 can support one or more video codecs. In this way, the terminal device 100 can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0231] The NPU is a Neural-Network (NN) computing processor. By learning from the biological neural network structure, such as learning from the transmission mode between human brain neurons, it can quickly process the input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the terminal device 100 can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.

[0232] The internal memory 121 may include one or more Random Access Memories (RAM) and one or more Non-Volatile Memories (NVM).

[0233] The random access memory may include Static Random-Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM, for example, the fifth-generation DDR SDRAM is generally called DDR5 SDRAM), etc.;

[0234] The non-volatile memory may include disk storage devices, flash memory.

[0235] Flash memories can be classified according to their operating principles, including NOR Flash, NAND Flash, 3D NAND Flash, etc. According to the number of potential levels of storage cells, they can include single-level cells (SLC), multi-level cells (MLC), triple-level cells (TLC), quad-level cells (QLC), etc. According to storage specifications, they can include Universal Flash Storage (UFS), embedded Multi Media Card (eMMC), etc.

[0236] The random access memory can be directly read and written by the processor 110. It can be used to store the operating system or executable programs (such as machine instructions) of other running programs, and can also be used to store data of users and application programs, etc.

[0237] The non-volatile memory can also store executable programs and data of users and application programs, etc. It can be pre-loaded into the random access memory for direct reading and writing by the processor 110.

[0238] The external memory interface 120 can be used to connect to an external non-volatile memory to expand the storage capacity of the terminal device 100. The external non-volatile memory communicates with the processor 110 through the external memory interface 120 to implement the data storage function. For example, files such as music and videos are saved in the external non-volatile memory.

[0239] The internal memory 121 or the external memory interface 120 is used to store one or more computer programs. One or more computer programs are configured to be executed by the processor 110. One or more computer programs include multiple instructions. When the multiple instructions are executed by the processor 110, the screen display detection method executed on the terminal device 100 in the above embodiments can be implemented to realize the screen display detection function of the terminal device 100.

[0240] The terminal device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc. For example, music playback, recording, etc.

[0241] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used for encoding and decoding audio signals. In an embodiment of the present application, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.

[0242] The speaker 170A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The terminal device 100 can listen to music or hands-free calls through the speaker 170A.

[0243] The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the terminal device 100 answers a call or a voice message, the user can listen to the voice by bringing the receiver 170B close to the ear.

[0244] The microphone 170C, also known as the "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak by bringing the mouth close to the microphone 170C to input the sound signal into the microphone 170C. The terminal device 100 can be provided with at least one microphone 170C. In some other embodiments, the terminal device 100 can be provided with two microphones 170C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the terminal device 100 can also be provided with three, four or more microphones 170C to implement functions such as collecting sound signals, noise reduction, identifying the sound source, and implementing a directional recording function.

[0245] The headphone jack 170D is used to connect a wired headphone. The headphone jack 170D can be a USB interface 130, or a 3.5 mm Open Mobile Terminal Platform (OMTP) standard interface, or a Cellular Telecommunications Industry Association of the USA (CTIA) standard interface.

[0246] The keys 190 include a power-on key, volume keys, etc. The keys 190 can be mechanical keys or touch keys. The terminal device 100 can receive key inputs to generate key signal inputs related to the user settings and function controls of the terminal device 100.

[0247] The motor 191 can generate vibration prompts. The motor 191 can be used for incoming call vibration prompts and also for touch vibration feedback. For example, touch operations for different applications (such as taking pictures, playing audio, etc.) can correspond to different vibration feedback effects. For touch operations on different areas of the display screen 194, the motor 191 can also correspond to different vibration feedback effects. Different application scenarios (such as time reminder, receiving messages, alarm clock, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.

[0248] The indicator 192 can be an indicator light and can be used to indicate the charging state, power change, and can also be used to indicate messages, missed calls, notifications, etc.

[0249] The SIM card interface 195 is used to connect the SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation from the terminal device 100. The terminal device 100 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The terminal device 100 interacts with the network through the SIM card to implement functions such as calls and data communication. In an embodiment of the present application, the terminal device 100 uses an eSIM, that is, an embedded SIM card. The eSIM card can be embedded in the terminal device 100 and cannot be separated from the terminal device 100. The embodiments of the present application also provide a computer storage medium. Computer instructions are stored in the computer storage medium. When the computer instructions run on the terminal device 100, the terminal device 100 is caused to execute the above-mentioned related method steps to implement the face recognition method in the above embodiment.

[0250] The embodiments of the present application also provide a computer program product. When the computer program product runs on a computer, the computer is caused to execute the above-mentioned related steps to implement the face recognition method in the above embodiment.

[0251] In addition, the embodiments of the present application also provide a device. This device can specifically be a chip, component or module. The device can include a processor and a memory connected to each other. Among them, the memory is used to store computer execution instructions. When the device runs, the processor can execute the computer execution instructions stored in the memory to cause the chip to execute the face recognition method in each of the above method embodiments.

[0252] Among them, the terminal device, computer storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be elaborated here.

[0253] From the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0254] In several embodiments provided in this application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point, the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0255] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0256] In addition, each functional unit in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0257] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0258] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them. Although the present application has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. A face recognition method, applied to a terminal device, characterized in that The terminal device includes a photographing device and an ambient light sensor, and the method includes: In response to a user operation that triggers face recognition, using the photographing device to collect an image to be recognized, and using the ambient light sensor to obtain the illuminance of the current environment; When it is determined that the current environment is a low-light environment according to the illuminance, inputting the image to be recognized into a pre-trained image enhancement model to obtain an enhanced image of the image to be recognized; Using a preset face feature extraction model to extract first face features from the enhanced image; If the first similarity between the first face features and second face features in a preset face template image is greater than or equal to a preset similarity threshold, it is determined that the user's identity verification is passed.

2. The face recognition method according to claim 1, wherein The method further includes: If the first similarity is less than the similarity threshold, performing image degradation processing on the face template image to obtain a degraded image of the face template image; Using the face feature extraction model to extract third face features from the image to be recognized, and using the face feature extraction model to extract fourth face features from the degraded image; If the second similarity between the third face features and the fourth face features is greater than or equal to the similarity threshold, it is determined that the user's identity verification is passed; or, If the second similarity is less than the similarity threshold, it is determined that the user's identity verification fails.

3. The face recognition method according to claim 1, wherein The method further includes: Using the ambient light sensor to collect multiple consecutive illuminances of the current environment; Based on the sliding average algorithm, determining the average value of the multiple consecutive illuminances; If the average value is less than a preset illuminance threshold, it is determined that the current environment is a low-light environment.

4. The face recognition method according to claim 1, wherein The step of using a preset face feature extraction model to extract first face features from the enhanced image includes: Based on a preset face detection model, determining whether there is a face to be recognized in the enhanced image; If there is a face to be recognized in the enhanced image, intercepting a face region image from the enhanced image; Using the face feature extraction model to extract the first face features from the face region image.

5. The face recognition method according to claim 4, wherein The method further includes: If there is no face to be recognized in the enhanced image, it is determined that the user's identity verification fails.

6. The face recognition method according to claim 1, wherein The method further includes: When it is determined that the current environment is not a low-light environment according to the illuminance, based on a preset face detection model, determining whether there is a face to be recognized in the image to be recognized; If there is a face to be recognized in the image to be recognized, intercepting a face region image from the image to be recognized; Using the face feature extraction model to extract fifth face features from the face region image; If it is determined that the third similarity between the fifth face features and the second face features is greater than or equal to the similarity threshold, it is determined that the user's identity verification is passed; or, If it is determined that the third similarity is less than the similarity threshold, it is determined that the user's identity verification fails.

7. The face recognition method according to claim 1, wherein The image enhancement model includes a first model and a second model, and the step of inputting the image to be recognized into a pre-trained image enhancement model to obtain an enhanced image of the image to be recognized includes: Denoise the image to be recognized using the first model, and output the denoised image to be recognized; Enhance the brightness of the denoised image to be recognized using the second model, and use the output brightness-enhanced image as the enhanced image.

8. The face recognition method according to claim 1, wherein The training method of the image enhancement model includes: Obtain multiple sample images, where the multiple sample images include multiple first images and multiple second images, and among them, the image noise of any first image is more than that of any second image; Construct a training set and a test set according to the multiple sample images, and use the training set to perform at least one iteration update on a preset initial model until the image enhancement model that meets the preset requirements is obtained.

9. The face recognition method according to claim 8, wherein The preset requirements include one or more combinations of the following requirements: the loss function corresponding to the initial model reaches the convergence condition; the test result obtained by testing the initial model using the test set indicates that the model performance of the initial model reaches the expectation; the number of times of the at least one iteration update reaches a preset iteration number threshold.

10. The face recognition method according to claim 8, wherein The initial model includes a first initial model, and any update in the at least one iteration update includes: Input any first image in the training set into the first initial model, and use the first initial model to denoise the any first image to obtain the first denoised image corresponding to the any first image, and the first difference image between the any first image and the first denoised image; Determine the first loss value corresponding to the first difference image using a preset first loss function, and the first loss value is used to indicate the noise degree in the first difference image; If the first loss value is less than or equal to a preset first threshold, use another first image to perform the next update on the first initial model.

11. The face recognition method according to claim 10, characterized in that, Any update in the at least one iteration update further includes: If the first loss value is greater than the first threshold, fuse the first difference image with any second image in the training set to obtain the fused image corresponding to the any second image; Input the fused image into the first initial model, and use the first initial model to denoise the fused image to obtain the second denoised image corresponding to the fused image; Determine the second loss value between the second denoised image and the corresponding any second image using a preset second loss function; If the second loss value is greater than a preset second threshold, use the next first image and the next second image to perform the next update on the first initial model.

12. The face recognition method according to claim 11, wherein The initial model further includes a second initial model, and any update in the at least one iteration update further includes: If the second loss value is less than or equal to the second threshold, input the second denoised image corresponding to any second image into the second initial model, and use the second initial model to enhance the brightness of the second denoised image to obtain the brightness-enhanced image corresponding to the second denoised image; Determine the third loss value between the brightness-enhanced image and the corresponding any second image using a preset third loss function; If the third loss value is greater than a preset third threshold, use the second denoised image corresponding to the next second image to perform the next update on the second initial model; or, If the third loss value is less than or equal to the third threshold, determine that the initial model meets the preset requirements, and use the initial model as the image enhancement model.

13. The face recognition method according to claim 2, wherein The performing image degradation processing on the face template image to obtain a degraded image of the face template image includes: Obtain a preset third image; Input the third image into the image enhancement model, and use the first model of the image enhancement model to denoise the third image to obtain a third denoised image corresponding to the third image and a second difference image between the third image and the third denoised image; Reduce the brightness of the face template image to obtain an image with reduced brightness; Perform image fusion on the second difference image and the image with reduced brightness to obtain the degraded image of the face template image.

14. A terminal device, characterized in that, The terminal device includes a memory and a processor: The memory is used for storing program instructions; The processor is used for reading and executing the program instructions stored in the memory. When the program instructions are executed by the processor, the terminal device executes the face recognition method according to any one of claims 1 to 13.

15. A computer storage medium, characterized in that, The computer storage medium stores program instructions. When the program instructions run on the terminal device, the processor of the terminal device executes the face recognition method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Method for enhancing low-light video image based on space-time accumulation and image degradation model

    CN106327450A

  • Unlocking method and related product

    CN110287668A

  • Dark light image enhancement method

    CN111064904A

  • Image processing method and device, electronic equipment and readable storage medium

    CN113674159A

  • Training method of image enhancement model, image enhancement method and electronic equipment

    CN114331918A