Face feature detection method and electronic device
Through the dual camera solution, color and infrared images are switched to detect face features under different lighting conditions, the problems of poor recognition effect and high power consumption when the light conditions are poor, and efficient and low-consumption user experience is improved.
Patent Information
- Application Number
- CN202311371268.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-20
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2043-10-20
AI Technical Summary
The existing contactless face feature detection technology has poor recognition effect and high power consumption when the light conditions are poor, resulting in poor user experience.
The dual camera scheme is adopted, and the first camera is used to shoot color images when the light is good to detect the face features, and the second camera is used to shoot infrared images when the light is bad to detect the face features. By analyzing the image quality score, the camera switching is determined, reducing power consumption and improving the success rate.
The success rate of face feature detection under various lighting conditions is improved, power consumption is reduced, and user experience is improved.
Smart Images

Figure CN118447545B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of electronic devices, and particularly relates to a face feature detection method and an electronic device. Background Art
[0002] With the rapid development of terminal technology and the maturity of communication technology, people have begun to explore new human-computer interaction methods that are divorced from touch operations or keyboard and mouse inputs, that is, non-contact interaction methods. For example, speech recognition, gesture recognition, face recognition, eye tracking, etc., which can provide users with more convenient and diverse interaction methods and enhance the user experience.
[0003] However, at present, the technologies related to non-contact interaction methods are not yet mature, and there are still defects such as poor recognition effects and high power consumption. For example, when implementing a non-contact interaction method by detecting face features, the detection of face features fails in many scenarios, resulting in the inability to achieve human-computer interaction and a deterioration of the user experience. Summary of the Invention
[0004] The embodiments of this application provide a face feature detection method and an electronic device, which can improve the success rate of detecting face features and reduce power consumption.
[0005] In a first aspect, a face feature detection method is provided, which is applied to an electronic device having a first camera and a second camera. The method may specifically include: taking a first image by using the first camera and detecting face features according to the first image, where the first image includes a first face image and the first image is a color image. When the detection of face features fails, obtain the image quality score of the first image. When the image quality score of the first image is less than or equal to a first threshold, take a second image by using the second camera and re-detect face features according to the second image, where the second image includes a second face image and the second image is an infrared image.
[0006] Therefore, in the above solution, the color image (i.e., the first image) is first collected by the first camera to detect the face features. Since the power consumption caused by collecting the color image is relatively low, and most electronic devices are used in well-lit scenarios, the imaging quality of the color image meets the requirements in most scenarios. Therefore, by calling the first camera to detect the face features, both low power consumption can be maintained and a good detection success rate can be ensured. If the face feature detection fails through the first image, there may be many reasons for the detection failure. For example, it may be caused by the current shooting light environment, or it may be caused by the detected target itself, or it may also be caused by some other accidental errors, and so on. At this time, the image quality score of the first image can be obtained. If the image quality score is less than or equal to the first threshold, it means that the image quality of the first image is very poor. Therefore, the poor image quality is very likely to be the main reason for the failure of face feature detection. For example, in low light, backlight, side light or reflective scenarios, the face features in the first image may not be clear enough. Even if the first image includes the first face image, the electronic device may not be able to detect the face features from the first image. In this case, the electronic device can continue to call the second camera to collect the infrared image (i.e., the second image) to detect the face features. Since the imaging quality of the infrared image is hardly affected by the light conditions, even if the light conditions are relatively poor, a high image quality can be ensured, thereby improving the success rate of face feature detection.
[0007] In summary, the solution provided by this application involves two cameras (i.e., the first camera and the second camera). Among them, the first camera can capture color images, and the second camera can capture infrared images. In scenarios with relatively good light conditions, the first camera is used to capture color images to detect face features, which can not only maintain low power consumption but also ensure a good detection success rate. In scenarios with poor light conditions, using the second camera to capture infrared images can ensure a high detection success rate under various light conditions. However, since the light conditions are not easy to quantify, if the camera is selected only based on the ambient light intensity, the face feature detection will still fail in many scenarios. For example, in backlight, side light or glasses reflection scenarios, the ambient light intensity may be relatively high, but the quality of the captured color image may be relatively poor. The solution provided by the embodiments of this application can call the second camera to capture infrared images only after the face feature detection fails using the first image and through analysis, it is found that the reason for the failure is very likely due to the relatively poor image quality of the first image. Because capturing infrared images will cause relatively high power consumption, and through the above solution, the number of infrared image captures can be reduced, thereby reducing the power consumption, and at the same time, a high detection success rate can be ensured, and at least the situation of face feature detection failure caused by image quality can be reduced.
[0008] Optionally, detecting a face feature according to the first image includes: converting the first image into a fourth image, and then detecting the face feature according to the fourth image. The format of the fourth image is different from that of the first image. For example, the fourth image is a grayscale image obtained by converting the first image. The fourth image includes a face image corresponding to the first face image.
[0009] That is to say, in the embodiments of the present application, the face feature can be directly detected according to the first image, or the face feature can be detected according to the image obtained after the conversion of the first image. In this way, the solution of the present application can be applied to different face feature detection models to detect face features, because different models require different formats of input images.
[0010] Optionally, capturing the first image using the first camera includes: obtaining the current ambient light intensity; when the ambient light intensity is greater than or equal to a second threshold, capturing the first image using the first camera.
[0011] Optionally, when the ambient light intensity is less than the second threshold, capturing a third image using a second camera. The third image includes a third face image, and the third image is an infrared image; detecting the face feature using the third image.
[0012] Therefore, in the above solution, the first camera is used to detect the face feature only when the light intensity is relatively high (the ambient light intensity is greater than or equal to the second threshold). In other words, when the light intensity is relatively low (the ambient light intensity is less than the second threshold), the second camera can be directly used to detect the face feature. Because the ambient light intensity is relatively low, it means that it is a low-light scene at this time. If a color image is captured at this time, the imaging quality will probably be relatively poor. Directly calling the second camera to detect the face feature can save the processes of collecting the first image, detecting the face feature according to the first image, obtaining the quality score, and judging the quality score, thereby saving resources, reducing latency, and improving the user experience.
[0013] Optionally, before capturing the first image using the first camera, the method further includes: determining to detect the face feature. That is to say, when the electronic device determines to detect the face feature, the first camera is enabled to capture an image. Among them, it can be that the user controls the electronic device to perform face feature detection, or the electronic device automatically performs face feature detection based on a preset rule. The present application does not make a limitation.
[0014] Optionally, the first camera and the second camera can be one camera or two different cameras.
[0015] Optionally, the first camera and the second camera can be the front cameras of the electronic device.
[0016] Optionally, the first camera is used to acquire a two-dimensional image, and the second camera is used to acquire an image containing depth information. Optionally, the first camera is a red green blue (RGB) camera, and the second camera is a time of flight (Tof) camera.
[0017] Optionally, the first camera is an always-on (AO) camera. Among them, the AO camera is a low-power RGB camera. By capturing the first image with the AO camera, the power consumption can be further reduced.
[0018] Optionally, obtaining the image quality score of the first image includes: determining a first feature image according to the first image, where the first feature image includes one or more of the first face image, the first left eye image, and the first right eye image, and the first left eye image and the first right eye image are included in the first face image; calculating the image quality score of the first image according to the first feature image.
[0019] Since face feature detection mainly detects face features (especially eye features), in the above solution, calculating the image quality score based on one or more of the first face image, the first left eye image, and the first right eye image can improve the accuracy of the image quality score, that is, the image quality score can more accurately characterize the quality of the face features in the first image.
[0020] Optionally, calculating the image quality score of the first image according to the first feature image includes: calculating a quality evaluation coefficient according to the first feature image, where the quality evaluation coefficient includes one or more of the following: brightness evaluation coefficient, contrast evaluation coefficient, sharpness evaluation coefficient; calculating the image quality score of the first image according to the quality evaluation coefficient.
[0021] The image quality score is a parameter used to measure the image quality, and the image quality may be affected by various factors. Or rather, there are various parameters that may be factors affecting the image quality, such as the brightness, contrast, sharpness, etc. of the image. Embodiments of the present application can comprehensively consider one or more of the above factors so that the calculated image quality score can more accurately characterize the image characteristics. The higher the accuracy of the image quality score, the fewer times the second camera is misinvoked, thus saving resources. Specifically, if the failure to detect the human face features from the first image may be caused by other reasons than the image quality, and if the accuracy of the image quality score is relatively low, the calculated image quality score may still be less than or equal to the second threshold. At this time, the electronic device will invoke the second camera to capture an infrared image. However, since the reason for the detection failure is not due to poor image quality, even capturing an infrared image to detect the human face features may be of no avail and will waste resources in vain. Embodiments of the present application can reduce the probability of misinvoking the second camera by improving the quality score of the first image, thereby saving resources and reducing power consumption.
[0022] Optionally, obtaining the image quality score of the first image includes: calculating the average pixel intensity corresponding to the first feature image as the image quality score of the first image.
[0023] In the above solution, the quality score of the first image can be obtained by calculating the average pixel intensity. The calculation method is more concise, which can improve the calculation efficiency and can also more accurately indicate whether the first image meets the quality standard.
[0024] Optionally, detecting the human face features according to the second image includes: detecting whether the user's fixation point is within the display screen of the electronic device according to the first image.
[0025] Optionally, detecting the human face features according to the second image includes: detecting whether the user's fixation point is within the display screen of the electronic device according to the second image; the method further includes: controlling the on / off state of the display screen of the electronic device according to the detection result.
[0026] The solution provided by the embodiments of the present application can be applied to the scenarios of "gaze without screen off" or "gaze with screen on". That is to say, the electronic device can detect whether the user's gaze point is within the display screen of the electronic device according to the second image (or the first image in the above solution), that is, detect whether the user is currently gazing at the display screen, and control the electronic device to switch between the screen-off state and the screen-on state. For example, in the screen-on state, if the user is not currently gazing at the display screen, the electronic device can be automatically switched from the screen-on state to the screen-off state, thereby saving power; if the user is currently gazing at the display screen, the electronic device can be kept in the screen-on state, or switched from the screen-off state to the screen-on state, thereby improving the user experience.
[0027] Optionally, controlling the on / off state of the display screen of the electronic device according to the detection result includes: if the user's gaze point is within the display screen of the electronic device, maintaining the screen-on state of the display screen of the electronic device and resetting the screen-off countdown.
[0028] When the solution provided by the embodiments of the present application is applied to the scenario of "gaze without screen off", if it is detected that the user is currently gazing at the display screen, the screen-on state of the electronic device can be maintained, that is, the electronic device does not automatically enter the screen-off state, so as to reduce the situation where the user is using the electronic device but the electronic device automatically enters the screen-off state. Moreover, the electronic device can reset the screen-off countdown, that is, the screen-off countdown starts counting from the beginning again, so that after the user no longer uses the electronic device subsequently, the electronic device can still automatically turn off the screen.
[0029] Optionally, capturing the first image by using the first camera includes: when the screen-off countdown is less than or equal to a third threshold, capturing the first image by using the first camera based on the current ambient light intensity.
[0030] When the solution provided by the embodiments of the present application is applied to the scenario of "gaze without screen off", the current ambient light intensity can be obtained when the screen-off countdown is about to end (that is, when the screen-off countdown is less than or equal to the third threshold). That is to say, when the electronic device is about to automatically turn off the screen, the current ambient light intensity can be obtained so as to call the camera based on the ambient light intensity to detect whether the user's gaze point falls within the display screen of the electronic device, and decide whether to automatically turn off the screen based on the detection result, thereby being able to decide whether to automatically turn off the screen according to the user's actual usage situation and improving the user experience.
[0031] Optionally, detecting whether the user's fixation point is within the display screen of the electronic device based on the first image includes: determining a first feature image according to the first image, where the first feature image includes a first face image, a first left-eye image, and a first right-eye image, the first face image is included in the first image, and the first left-eye image and the first right-eye image are included in the first face image. Then, determining whether the user's fixation point is within the display screen of the electronic device according to the first feature image.
[0032] Optionally, detecting whether the user's fixation point is within the display screen of the electronic device based on the second image includes: determining a second feature image according to the second image, where the second feature image includes a second face image, a second left-eye image, and a second right-eye image, the second face image is included in the second image, and the left-eye image and the right-eye image are included in the second face image. Then, determining whether the user's fixation point is within the display screen of the electronic device according to the second feature image.
[0033] In the above solution, face features can be detected based on the face image and eye images (including the left-eye image and the right-eye image). In the fixation point recognition scenario, separating the eye images and using them as important parameters for face feature detection can focus more on the eye features during detection, thus enabling more accurate identification of the eye gaze direction.
[0034] Optionally, determining whether the user's fixation point is within the display screen of the electronic device according to the feature image includes: respectively extracting a second left-eye feature and a second right-eye feature from the second left-eye image and the second right-eye image using an eye feature extraction model; extracting a second face feature from the second face image using a face feature extraction model; fusing the second left-eye feature, the second right-eye feature, and the second face feature to obtain fusion data, where the fusion data is multi-modal fusion feature data; processing the fusion data using a fixation recognition model and outputting a fixation result, where the fixation result is used to indicate whether the user's fixation point is within the display screen of the electronic device.
[0035] Optionally, the above eye feature extraction model, face feature extraction model, and fixation recognition model are established based on a convolutional neural network.
[0036] In the above solution, extracting the features of the eye image and the face image through different models respectively can improve the fineness of feature extraction. Because the characteristics of different images are different, if the same model is used to extract the features of different types of images, it may affect the accuracy of subsequent fixation recognition.
[0037] Optionally, before performing face feature detection based on the first image, the method further includes: detecting the first image based on a face detection model and outputting a first detection result, where the first detection result indicates that the first image includes the first face image; before performing face feature detection based on the second image, the method further includes: detecting the second image based on a face detection model and outputting a second detection result, where the second detection result indicates that the second image includes the second face image.
[0038] In the above solution, after the electronic device captures the first image using the first camera, it can detect whether the first image includes the first face image. If it does, the subsequent process can be continued. If not, the subsequent process may not be executed, or the first camera can be called again to capture an image, or it can be detected whether other frame images captured by the first camera include a face image (the first image and the other frame images here can be multiple frame images captured by the first camera at one time), avoiding the situation where face feature detection fails due to the first image not including a face image, omitting the detection process in this case, saving resources and reducing latency.
[0039] Optionally, before capturing the first image using the first camera, the method further includes: displaying a first interface on the screen, where the first interface is any one of the following multiple interfaces: the interface of a first application installed in the electronic device, the system desktop, the negative first screen, and the lock screen interface.
[0040] In a second aspect, a face feature detection method is provided, which is applied to an electronic device having a first camera and a second camera. The method may specifically include: obtaining the current ambient light intensity; if the current ambient light intensity is greater than or equal to a second threshold, capturing a first image using the first camera and detecting face features based on the first image. Wherein, the first image includes first face features and the first image is a color image. If the current ambient light intensity is less than the second threshold, capturing a second image using the second camera and detecting face features based on the second image. Wherein, the second image includes a second face image and the second image is an infrared image.
[0041] Therefore, in the above solution, in the case of relatively high light intensity, face features are recognized by capturing a color image (i.e., the first image), reducing power consumption while ensuring a good face feature recognition effect; in the case of relatively low light intensity, face features are recognized by capturing an infrared image (i.e., the second image) to improve the face feature recognition effect in low-light scenarios.
[0042] It is understandable that the solution provided by the first aspect can be implemented based on the solution provided by the second aspect. Therefore, the solution provided by the first aspect can be implemented as a further optional solution in the second aspect. The specific solution will not be elaborated here.
[0043] In a third aspect, a face feature detection method is provided, which is applied to an electronic device having a first camera and a second camera. The method may specifically include: when the display screen of the electronic device is currently in the lit state, using the first camera to capture a first image, and detecting whether the user's gaze point is within the display screen of the electronic device based on the first image, where the first image includes a first face image and the first image is a color image. In the case of detection failure, obtaining an image quality score of the first image. When the image quality score of the first image is less than or equal to a first threshold, using the second camera to capture a second image, and re-detecting whether the user's gaze point is within the display screen of the electronic device based on the second image, where the second image includes a second face image and the second image is an infrared image.
[0044] Therefore, the solution provided by the embodiments of the present application can be applied to the "gaze without screen off" scenario. It is understandable that the solution provided by the third aspect can be a further solution based on the method provided by the first aspect. That is to say, the method provided by the third aspect can be combined with the method provided by the first aspect, and the optional solutions in the first aspect can also be used as optional solutions in the third aspect. For the sake of brevity, the above solutions will not be repeated here.
[0045] In a fourth aspect, an electronic device is provided. The electronic device has a first camera and a second camera. The electronic device specifically includes: a shooting module for using the first camera to capture a first image; a detection module for detecting face features based on the first image, where the first image includes a first face image and the first image is a color image. An acquisition module for obtaining an image quality score of the first image when face feature detection fails; the shooting module is further configured to use the second camera to capture a second image when the image quality score of the first image is less than or equal to a first threshold; the detection module is further configured to re-detect face features based on the second image, where the second image includes a second face image and the second image is an infrared image.
[0046] In a fifth aspect, an electronic device is provided, including a memory and a processor. A computer program is stored on the memory and can run on the processor. When the processor executes the computer program, the electronic device implements the steps of the face feature detection method as described in any one of the first aspect or the second aspect above.
[0047] In a sixth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the face feature detection method described in any one of the above first aspect or second aspect are implemented.
[0048] In a seventh aspect, a computer program product is provided. When the computer program product runs on an electronic device, the electronic device is caused to execute the face feature detection method described in any one of the above first aspect or second aspect.
[0049] In an eighth aspect, a chip system is provided. The chip system includes a processor, the processor is coupled to a memory, and the processor executes a computer program stored in the memory to implement the face feature detection method described in any one of the above first aspect or second aspect.
[0050] Wherein, the chip system may be a single chip or a chip module composed of multiple chips.
[0051] It can be understood that the beneficial effects of the above second aspect to eighth aspect can be referred to the relevant descriptions in the above first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 FIG. shows a schematic diagram of an application example of an automatic screen-off rule provided by an embodiment of the present application;
[0053] Figure 2 FIG. shows a schematic diagram of an application scenario provided by an embodiment of the present application;
[0054] Figure 3 FIG. shows another schematic diagram of an application scenario provided by an embodiment of the present application;
[0055] Figure 4 FIG. shows yet another schematic diagram of an application scenario provided by an embodiment of the present application;
[0056] Figure 5 FIG. shows images obtained by using an RGB camera under different ambient light intensities;
[0057] Figure 6 FIG. shows images obtained by using a Tof camera under different ambient light intensities;
[0058] Figure 7 FIG. shows a schematic diagram of an electronic device provided by an embodiment of the present application;
[0059] Figure 8 FIG. shows a schematic block diagram of a method for detecting face features provided by an embodiment of the present application;
[0060] Figure 9Shows another schematic diagram of an application scenario provided by an embodiment of the present application;
[0061] Figure 10 Shows a schematic block diagram of another method for detecting human face features provided by an embodiment of the present application;
[0062] Figure 11 Shows an exemplary image obtained by an electronic device through a camera;
[0063] Figure 12 Shows an exemplary process for identifying and cropping a human face image;
[0064] Figure 13 Shows an exemplary process for determining a human face image and a human eye image;
[0065] Figure 14 Shows an exemplary framework and an exemplary process of an electronic device for gaze recognition;
[0066] Figure 15 Shows another exemplary process for gaze recognition;
[0067] Figure 16 Shows an exemplary process for determining whether an image meets a quality standard;
[0068] Figure 17 Shows the structure of an electronic device and the implementation process of a solution provided by an embodiment of the present application;
[0069] Figure 18 Shows an exemplary flowchart of a method for detecting human face features provided by an embodiment of the present application;
[0070] Figure 19 Shows a block diagram of the software and hardware system of an electronic device provided by an embodiment of the present application;
[0071] Figure 20 Shows a block diagram of the hardware system of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0072] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. To describe the embodiments of the present application more clearly, some terms or technologies related to the embodiments of the present application will be briefly introduced first.
[0073] 1. Brightness and off state of the display screen: In the embodiments of the present application, the brightness and off state of the display screen is used to indicate whether the display screen of the electronic device is currently in the bright screen state or the off screen state.
[0074] Among them, the bright screen state refers to the state in which the display screen of the electronic device is lit. At this time, the display screen may display the lock screen interface (i.e., the unlocked interface), or it may display the system desktop or the negative one screen, or it may display the application interface of an application installed in the electronic device, etc. The bright screen state can also be called the open screen state, etc.
[0075] The screen-off state refers to the state in which the display screen of an electronic device is not lit, that is, the display screen is in a dark state. The screen-off state can also be called the screen-off state or the black screen state.
[0076] It is understandable that the display screen being in the screen-off state does not necessarily mean that the display screen does not display any content at all. For example, the display screen can still display time, date, battery level, application notifications, wallpapers and other content in part of the display screen when the display screen is in the screen-off state (that is, the screen-off display function of the corresponding electronic device).
[0077] It can also be understood that if the electronic device includes multiple display screens, when the main screen is in the screen-on state and the secondary screen is in the screen-off state (or the main screen is in the screen-off state and the secondary screen is in the screen-on state), it can be determined based on the actual strategy whether the electronic device is in the screen-on state or the screen-off state at this time. This application does not make specific limitations on this situation.
[0078] An electronic device can switch between a screen-on state and a screen-off state based on a user's instructions. In one example, the electronic device is configured with a power button, and the user can press the power button to switch the display screen on and off. That is, pressing the power button in the screen-on state switches the device to the screen-off state, and pressing the power button in the screen-off state switches the device to the screen-on state.
[0079] 2. Screen-off countdown: The screen-off countdown described in the embodiments of the present application is used for the electronic device to automatically complete the screen-off operation.
[0080] For example, when an electronic device in the screen-on state receives a user operation instruction, the electronic device automatically starts a screen-off countdown through a timer. If the electronic device receives another user instruction before the screen-off countdown ends, the electronic device resets the screen-off countdown. The instruction here can be a user touch operation anywhere on the display screen. If the electronic device does not receive any user instruction before the screen-off countdown ends, the electronic device switches the display screen to the screen-off state.
[0081] 3. Gaze Point: The gaze point in the embodiments of this application refers to a point on a target object where the user's line of sight is directed during visual perception. If the user's gaze point is within the display area of the electronic device's display (hereinafter referred to as the user's gaze point being within the display), the user is considered to be looking at the display; otherwise, the user is considered not to be looking at the display.
[0082] With the rapid development of terminal technology and the maturity of communication technology, non-contact interaction methods have begun to be applied to various electronic devices. The electronic device described in the embodiments of this application may be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The embodiments of this application do not impose any restrictions on the specific type of the electronic device. For convenience, the following takes a mobile phone as an example of the electronic device for illustration.
[0083] As an example, in the non-contact interaction method, the mobile phone can use the camera to detect the face features to perform corresponding operations. For example, it can perform operations such as unlocking or payment by detecting the face features, and can also achieve eye tracking by detecting the face features, and can also control the brightness or the on / off state of the display screen by detecting the face features. The following will be described in conjunction with examples.
[0084] In many cases, after the user finishes using the mobile phone, they do not actively press the power button to switch the display screen of the mobile phone to the screen-off state. In this case, if the mobile phone always maintains the screen-on state, it will cause a waste of power. To solve this problem, an automatic screen-off rule can be configured for the mobile phone.
[0085] Exemplarily, Figure 1 shows an application example of an automatic screen-off rule: The application interface of a shopping application (App) is currently displayed on the mobile phone (such as (a) in Figure 1 ). If the user does not perform any operation on the mobile phone within 30 seconds (s), that is, the mobile phone does not receive any instruction sent by the user within 30 s, the mobile phone automatically switches from the screen-on state to the screen-off state (such as (b) in Figure 1 ). It can be understood that the 30 s here can be the screen-off countdown duration set by the user.
[0086] Optionally, in one implementation, the mobile phone can also first reduce the brightness and then switch to the screen-off state. For example, if the user does not perform any operation on the mobile phone within 25 s, the mobile phone automatically reduces the brightness of the display screen (as shown in (c) in Figure 1 ). After reducing the brightness of the display screen, if the user performs any operation on the mobile phone within 5 s, the mobile phone restores the brightness of the display screen to Figure 1The state shown in (a); if the user does not perform any operation on the mobile phone within 5s, the mobile phone automatically switches to the screen-off state (as shown in (b) of Figure 1 ).
[0087] Through the above solution, the mobile phone can automatically enter the screen-off state when it does not receive any touch operation from the user for a period of time, thus saving resources.
[0088] However, in some scenarios, although the user does not perform any operation, the user may still be using the mobile phone. For example, after the user opens a shopping App on the mobile phone, the user enters a flash sale interface as shown in (a) of Figure 1 . As can be seen from (a) of Figure 1 , at this time, there are still 3 minutes and 6 seconds until the flash sale of the product. In order to avoid missing the flash sale of the product, the user may be staring at the current display page without performing any operation. If the mobile phone automatically enters the screen-off state at this time, it will affect the user experience.
[0089] Based on this, in the solution involved in the present application, the mobile phone can detect the user's gaze before entering the screen-off state and control the on / off state of the mobile phone based on the user's gaze. Please refer to Figure 2 . If the user does not operate within 30s, the mobile phone can use the front camera 210 to detect whether the user is staring at the display screen before entering the screen-off state. If the user is staring at the display screen at this time (that is, the user's fixation point is within the display screen), it can be considered that the user is using the mobile phone at this time, then the mobile phone does not automatically enter the screen-off state (that is, maintains the screen-on state of the mobile phone), and resets the screen-off timer (that is, the screen-off timer starts counting from the beginning again), so as to reduce the situation that the mobile phone automatically turns off the screen when the user is using the mobile phone and improve the user experience. Please refer to Figure 3 . If the user does not operate within 30s, the mobile phone can use the front camera 310 to detect whether the user is staring at the display screen before entering the screen-off state. If the user is not staring at the display screen at this time (that is, the user's fixation point is outside the display screen), it can be considered that the user is not using the mobile phone at this time, then the mobile phone automatically enters the screen-off state according to the original logic, thus saving power and reducing power consumption.
[0090] However, when using the front camera to detect the user's gaze, the detection effect may be affected due to problems such as light. Please refer to Figure 4, the electronic device is equipped with a front camera 410. The front camera 410 is a two-dimensional camera, that is, the front camera 410 is a camera for shooting two-dimensional color images, such as an RGB camera. When the user uses the mobile phone at night without turning on the light, if the front camera is called to take a face image at this time, the display effect of the captured face image may be relatively poor, so it may not be possible to successfully detect the user's gaze based on this face image. Please refer to Figure 5 , the face image captured by the user using the front camera 410 in sufficient light is relatively clear (as shown in (a) of Figure 5 ), while the face image captured by the user using the front camera 410 in insufficient light is relatively blurred (as shown in (b) of Figure 5 ). It should be noted that if the front camera is an RGB camera, the image shown in Figure 5 should be a color image. For the convenience of display, Figure 5 a black-and-white image is used as an example for illustration.
[0091] Further, if the image shown in (b) of Figure 5 is used to detect the user's gaze, it is very likely to detect failure. In this case, even if the user is currently looking at the display screen, the mobile phone may automatically turn off the screen, thus affecting the user experience.
[0092] In view of this, the embodiment of the present application provides an electronic device with a camera that can shoot infrared images. This camera can be a camera for shooting images containing depth information or a three-dimensional camera, such as a Tof camera. Since the infrared image is obtained by the infrared light emitter emitting modulated infrared light pulses, continuously hitting the object surface, being reflected and received by the receiver, calculating the time difference through the change of phase, and then combining the speed of light to calculate the object depth information, the general photo-taking effect is not interfered by ambient light. Therefore, even in dark or overly bright scenes, there can be a relatively clear user's face image in the infrared image captured by the electronic device. Please refer to Figure 6 , the three-dimensional camera can shoot infrared images. The facial features in the infrared image captured by the user using the Tof camera in sufficient light are relatively clear (as shown in (a) of Figure 6 ), and the facial features in the infrared image captured by the user using the Tof camera in insufficient light are also relatively clear (as shown in (b) of Figure 6 ). Based on this solution, when face feature detection is required, the electronic device always calls the above-mentioned Tof camera to take a face image and detects face features based on this face image. Since the Tof camera can be applied to various scenarios, relatively accurate recognition effects can be obtained in various scenarios.
[0093] However, in the above solution, calling a 3D camera to detect facial features will result in relatively high power consumption. In view of this, embodiments of the present application further provide an electronic device with two cameras, where one camera is used to obtain a color image, such as an RGB camera, and the other camera is used to obtain an infrared image, such as a Tof camera. For example, Figure 7 the mobile phone in Figure 8 has a front camera module including two cameras. Among them, camera 710 is an RGB camera, and camera 720 is a Tof camera. These two cameras can detect facial features under different ambient light conditions respectively. The following will make an exemplary description of this solution in combination with
[0094] S110, obtain the current ambient light intensity.
[0095] Exemplarily, in the case of needing to detect facial features, the electronic device obtains the current ambient light intensity.
[0096] It can be understood that detecting facial features as described in embodiments of the present application refers to the process of collecting a facial image through a camera and performing feature detection.
[0097] Embodiments of the present application can be applied to various scenarios of detecting facial features.
[0098] For example, in Figure 2 the application scenario shown, if the user is looking at the display screen, the screen remains on. Therefore, in this scenario, detecting facial features refers to collecting a facial image through a camera and determining whether the fixation point of the human eye is within the display screen based on the facial image.
[0099] Also, for example, Figure 9 shows another application scenario. As shown in Figure 9 (a) in Figure 9 , the electronic device enters the lock screen interface (or an unlocked interface). At this time, the identifier 910 is used to indicate to the user that the electronic device is currently in the locked state. At this time, the electronic device performs face recognition through the camera module 920. If the recognition is successful, the electronic device will be unlocked, and the interface after unlocking is as shown in
[0100] In (b) of Figure 9In the scenario shown, before performing face recognition, it is also possible to first determine whether the user is looking at the display screen. If so, then perform face recognition; if not, then no longer perform face recognition, thereby reducing the probability of accidental unlocking. Therefore, in Figure 9 the scenario shown, detecting facial features may also include determining whether the fixation point of the human eye is within the display screen.
[0101] It can be understood that the above several application scenarios are only examples, and this application is not limited thereto. Other scenarios that require detecting facial features are also within the protection scope of this application. For example, in the scenario of waking up the screen when looking, that is, when the electronic device detects that the user has been looking at the display screen for a period of time, it automatically wakes up the screen, or after the electronic device reduces the brightness to reduce power consumption (such as Figure 1 in (c) of ), if it detects that the user is looking at the display screen, it automatically restores the brightness of the display screen; another example is the eye tracking control scenario, that is, the electronic device detects the position of the user's fixation point on the display screen to open the application or notification in the fixation area, or to achieve page turning, page magnification, etc., or to track the user's dynamic fixation point to achieve contactless control. For example, interactive electronic games can be realized through eye tracking; another example is the electronic payment scenario, that is, after the user opens the payment interface, face recognition is used to complete password verification, etc. For the sake of convenience, the following takes the Figure 2 "never turn off the screen when looking" scenario shown as an example for illustration.
[0102] Optionally, before S110, the electronic device determines to detect facial features. For example, in Figure 2 the corresponding scenario, when the screen-off countdown is less than or equal to a preset threshold (such as 4 s), the electronic device determines to detect facial features.
[0103] S120, in the case where the current ambient light intensity is greater than or equal to the second threshold, use the first camera to take a first image; detect facial features according to the first image.
[0104] Exemplarily, after obtaining the ambient light intensity, it is determined whether the ambient light intensity is greater than or equal to the second threshold. For example, using the ambient light smoothed value to represent the ambient light intensity, if the ambient light smoothed value is greater than or equal to 15 lux (corresponding to the second threshold), it means that the lighting condition is relatively good at this time, and the first camera can be used to take a first image. Among them, the first image is a color image, or in other words, the first image is a two-dimensional image. Therefore, the first camera can be used to take color images, such as the first camera is an RGB camera.
[0105] Furthermore, detect facial features according to the first image. For example, corresponding to Figure 2In the scene shown, after the electronic device obtains the first image, it detects whether the first image includes a first face image. If so, it further determines whether the fixation point of the human eye is within the display screen based on the first face image.
[0106] Since the lighting condition is relatively good at this time, the electronic device uses a color image to detect face features, which can not only ensure a good recognition effect but also reduce power consumption.
[0107] S130, when the current ambient light intensity is less than the second threshold, use the second camera to take a second image; detect face features according to the second image.
[0108] Exemplarily, if the current ambient light intensity is less than the second threshold, it means that the lighting condition is relatively poor at this time. The second camera can be used to take a second image, where the second image is an infrared (IR) image, or in other words, the second image is an image with depth information. Therefore, the second camera can be used to take infrared images, such as when the second camera is a Tof camera.
[0109] Since the lighting condition is relatively poor at this time, the electronic device uses an infrared image to detect face features, which can ensure accurate identification of face features even when the light condition is not good.
[0110] In summary, in the above method 100, when the light intensity is relatively high, face features are recognized by collecting color images, which reduces power consumption while ensuring a good face feature recognition effect; when the light intensity is relatively low, face features are recognized by collecting infrared images to improve the face feature recognition effect in low-light scenes.
[0111] However, based on the above solutions, the recognition effect in some scenarios may still not be very good. For example, even in a scenario with high ambient light intensity, such as a backlight scenario, the quality of the captured color image may still not be high, and face feature detection may still fail in this case.
[0112] In view of this, the embodiments of the present application further provide a face feature detection method. In this method, when using a color image to detect face features, if the detection fails and the image quality score of the color image is lower than a preset threshold, an infrared image is re-taken to detect face features, which reduces the situation of face feature detection failure caused by poor image quality while minimizing power consumption. The following describes this method in detail with Figure 10 Method 200 in it. It can be understood that method 200 can be regarded as a further solution to method 100. Therefore, the descriptions of some terms or technical features in method 100 also apply to method 200.
[0113] S210, capture a first image using the first camera.
[0114] Exemplarily, the electronic device has a first camera, which is, for example, the front camera of the electronic device, such as Figure 7 the camera 410 in. It can be understood that the first camera can also be an external camera of the electronic device, and this application does not limit this.
[0115] The electronic device can capture a first image using the first camera. Among them, the first image can be one of multiple frames of images captured by the first camera. It can be understood that after the electronic device captures the first image using the first camera, the first image can be stored in an internal data buffer without displaying the first image on the display screen. That is to say, capturing the first image is an internal operation of the electronic device, and the user usually does not perceive this process.
[0116] The first image is a color image. That is to say, the first camera can be used to capture a color image, or in other words, the first camera can be used to obtain a two-dimensional image. For example, the first camera is an RGB camera. Specifically, the first camera can be an AO camera in the RGB camera. Since the AO camera is a low-power RGB camera, the power consumption can be further reduced by using the AO camera.
[0117] The first image includes a first face image. The face image described in the embodiments of this application refers to an image that contains part or all of a face. For example, assume Figure 11 the image 1110 in corresponds to the first image (for convenience, the image 1110 is shown in black and white), then the local image 1120 therein is the first face image.
[0118] The first image is used to detect face features. For example, the first image is used to detect whether the user's gaze point is within the display screen or for face recognition.
[0119] Optionally, before S210, the electronic device may further determine that the current ambient light intensity is greater than or equal to a second threshold. That is, before enabling the camera to capture an image, the electronic device may first obtain the current ambient light intensity, and then determine whether the current ambient light intensity is greater than or equal to the second threshold. If the ambient light intensity is greater than or equal to the second threshold, it indicates that the lighting condition may be good. In this case, the first camera is used to capture a color image first. If the ambient light intensity is less than the second threshold, it means that the current lighting condition is poor. If the first camera is still used to capture a color image, the probability of failure in detecting face features will be very high. Therefore, the second camera can be directly called to capture an infrared image. The specific process can refer to the description in the subsequent S240 part and will not be elaborated here. Optionally, when determining to perform face feature detection, the electronic device may obtain the current ambient light intensity. For example, corresponding to Figure 2 the application scenario shown, when the screen-off countdown of the electronic device is less than or equal to a third threshold (such as 4 s), the current ambient light intensity can be obtained.
[0120] Optionally, after S210 and before S220, the method further includes: detecting the first image based on a face detection model and outputting a first detection result, where the first detection result indicates that the first image includes the first face image; or in other words, the electronic device determines that the first image includes the first face image.
[0121] That is, after the electronic device captures the first image using the first camera, it detects whether the first image includes a face image. If it includes, the subsequent operation of S220 is continued; if it does not include, S210 can be re-executed, or it can continue to detect whether other images captured by the first camera include a face image, or the face feature detection process can be stopped. In this way, face feature detection can be performed only on the first image containing a face image, thereby saving resources.
[0122] It can be understood that this application does not limit the specific method for detecting whether a face image is included in the image. The following combines and to give an exemplary illustration of a possible implementation manner.
[0123] An exemplary process for identifying and cropping a face image is given. Optionally, the original image (i.e., the above-mentioned first image, here Taking the image 1110 in it as an example for illustration) perform scaling processing, and then input the processed image (or the original image) into a face detection model, which is, for example, a neural network model. The face detection model processes the input image and outputs a detection result. If the input image includes a face image, the output detection result includes the face image, and optionally, also includes eye images (including a left eye image and a right eye image). If the input image does not include a face image, the output detection result indicates that there is no face image in the image.
[0124] Please refer to , this face detection model can use face key point detection algorithms (such as DeepBlue Face (Dbface), Practical Facial Landmark Detector (PFLD), Face-Landmark-Factory, Multi-Task Convolutional Neural Network (MTCNN), CenterFace, etc. The algorithms are not limited here) to determine the face key points in the input image: left eye A, right eye B, nose C, left lip corner D, and right lip corner E, and determine the coordinate positions of each key point (as shown in (a) of ). Optionally, if the face is in a tilted state, such as key points A and B not being on the same horizontal line, the face image can also be corrected through a face correction algorithm, and the specific process is not limited in this application. It can be understood that if the face detection model does not detect face key points (or does not detect all five key points) from the input image, it is considered that there is no face image in the input image.
[0125] Furthermore, the face detection model outputs a detection result. The detection result includes the face image. For example, the face detection model can determine a rectangle centered on the nose C, and this rectangle contains the other several key points. The image intercepted by this rectangle is the face image (corresponding to the above first face image), as shown in (b) of . Optionally, the detection result also includes the left eye image. For example, the face detection model determines a rectangle with a fixed size centered on the left eye A, and the image intercepted by this rectangle is the left eye image, as shown in (c) of ; Optionally, the detection result also includes the right eye image. For example, the face detection model determines a rectangle with a fixed size centered on the right eye B, and the image intercepted by this rectangle is the right eye image, as shown in (d) of .
[0126] S220, detect face features according to the first image.
[0127] Exemplarily, after obtaining the first image by capturing with the first camera, face features are detected according to the first image. For example, corresponding to the application scenario shown, detecting face features according to the first image means: detecting whether the user's fixation point is within the display screen of the electronic device according to the first image. The present application does not limit the specific detection method. As an example, the electronic device can determine a first feature image according to the first image, where the first feature image includes one or more of a first face image, a first left eye image, and a first right eye image. The first left eye image and the first right eye image here include the first human eye image. The specific acquisition method can refer to the corresponding solution, which will not be elaborated here. Then, it is determined whether the user's fixation point is within the display screen of the electronic device according to the first feature image. The following takes the first feature image including three items: a first face image, a first left eye image, and a first right eye image as an example for illustration: Feature data is extracted from the first feature image, that is, a first facial feature, a first left eye feature, and a first right eye feature are respectively extracted from the first face image, the first left eye image, and the first right eye image. It can be understood that the feature data can be used to characterize the features of the feature image. Then, the extracted feature data is fused, and the fused data is processed using a gaze recognition model to obtain a gaze result, and the gaze result is used to indicate whether the user's fixation point is within the display screen of the electronic device. The following combines and to make an exemplary illustration of the specific implementation process.
[0128] shows an architecture diagram of a gaze recognition module. The gaze recognition module includes a multi-modal input module, a multi-scale feature fusion module, and a gaze classification module. Specifically, the multi-modal input module is used to input the first face image into the multi-scale feature same and module. The multi-scale feature fusion module is used to extract a first facial feature from the first face image using a face feature extraction model. The specific process is not limited in the present application. The multi-modal input module is also used to input the first human eye image and the first right eye image into the multi-scale feature fusion module. The multi-scale feature fusion module is also used to extract a first left eye feature and a first right eye feature from the first left eye image and the first right eye image respectively using an eye feature extraction model. The specific process is not limited in the present application.
[0129] Furthermore, the multi-scale feature fusion module is also used to perform multi-scale fusion on the first left eye feature, the first right eye feature, and the first facial feature, and then input the fused data into the gaze classification module. The gaze classification module is used to perform gaze recognition according to the fused data.
[0130] Another gaze recognition process is shown. In this process, the first feature image can be processed by a convolutional network, and then the gaze point position can be output. As shown, the convolutional network includes Convolution Group 1 (Conv1), Convolution Group 2 (Conv2), and Convolution Group 3 (Conv3). In one example, a convolution group includes: a convolution kernel (Convolution), a parametric rectified linear unit (PRelu) activation function, a pooling kernel (Pooling), and a local response normalization (LRN) layer. Among them, the convolution kernel of Conv1 is a 7×7 matrix, and the pooling kernel is a 3×3 matrix; the convolution kernel of Conv2 is a 5×5 matrix, and the pooling kernel is a 3×3 matrix; the convolution kernel of Conv3 is a 3×3 matrix, and the pooling kernel is a 2×2 matrix. After different images are processed by Conv1, Conv2, Conv3, and average pooling (AvgPool), they are input into different fully connected (FC) layers for full connection. For example, the first left-eye image is input into FC1 for full connection after being processed, the first right-eye image is input into FC2 for full connection after being processed, and the first face image is input into FC3 for full connection (not shown in the figure). The specific process is not limited here. It can be understood that the connection layers with different structures are constructed for different types of images, which can better obtain the features of various images, thereby improving the accuracy of the model and enabling the electronic device to more accurately identify the punctual position. Then, FC4 and FC5 can further process the data after the above processing, and finally output sigmod to indicate the human eye gaze point position.
[0131] S230. When the detection of the face feature fails, obtain the image quality score of the first image;
[0132] Exemplarily, the failure of detecting the face feature may mean that the first face image cannot be detected from the first image, or that it is impossible to identify whether the user's gaze point is within the display screen through the first face image, or that after face detection is performed on the first face image, the credibility of the detection result is lower than a preset value.
[0133] After the detection of the face feature fails, the electronic device can obtain the image quality score of the first image. Here, the image quality score is used to measure the image quality of the first image, and the image quality here may include characteristics such as the clarity, brightness, and contrast of the image.
[0134] The following gives an exemplary description of the specific implementation manner for the electronic device to obtain the image quality score of the first image.
[0135] Exemplarily, after determining the first feature image according to the first image, calculate the image quality score based on the first feature image. Taking the first feature image including the first face image, the first left eye image, and the first right eye image as an example, calculating the image quality score of the first image means calculating the image quality scores of the first face image, the first left eye image, and the first right eye image respectively. That is to say, the image quality score of the first image may include multiple quality scores.
[0136] The specific calculation method is exemplified as follows: Calculate the quality evaluation coefficient based on the first feature image, and the quality evaluation coefficient includes one or more of the following: brightness evaluation coefficient, contrast evaluation coefficient, sharpness evaluation coefficient, and then calculate the image quality score of the first image according to the quality evaluation coefficient.
[0137] Next, taking the calculation of the image quality score of the first face image as an example, where it is assumed that the image sequence of the first face image is: I i,j (i = 1, 2…M, j = 1, 2…N):
[0138] 1. The brightness evaluation coefficient can be calculated in the following way:
[0139] First, calculate the image gray value
[0140]
[0141] Then, based on the image gray value calculate the brightness evaluation coefficient λ1:
[0142]
[0143] 2. The contrast evaluation coefficient can be calculated in the following way:
[0144] First, calculate the normalized histogram p(r k ):
[0145]
[0146] Then, calculate the average value of the normalized histogram
[0147]
[0148] Next, calculate the variance δ of the normalized histogram 2 :
[0149]
[0150] Calculate the contrast evaluation coefficient λ2 based on the square root of the variance:
[0151] λ2 = 1 - δ.
[0152] 3. The clarity evaluation coefficient can be calculated in the following way:
[0153]
[0154] 4. Further, based on the brightness evaluation coefficient, the contrast evaluation coefficient, and the clarity evaluation coefficient, calculate the image quality score score of the first face image:
[0155]
[0156] where ω1 + ω2 + ω3 = 100,
[0157] It can be understood that the image quality scores of the first left eye image and the first right eye image can be calculated in a similar way. For the sake of brevity, it will not be elaborated here.
[0158] It can also be understood that the above-mentioned image quality score can be calculated in advance by a face detection model or other models in the above way. In this case, after the electronic device fails to detect the face features, it can directly obtain the pre-generated image quality score of the first image, thereby reducing the latency; or, after the electronic device fails to detect the face features, it can also calculate the image quality score of the first image in the above way. This application does not make any limitations in this regard.
[0159] It can be understood that the embodiments of this application do not limit the specific implementation manner of calculating the image quality score of the first image. The above calculation method is only an example, and other parameters that can be used to characterize the image quality can also be used as the image quality score. For example, the average pixel intensity corresponding to the first feature image can also be calculated as the image quality score of the first image. The specific process is not limited in this application.
[0160] S240. When the image quality score of the first image is less than or equal to the first threshold, use the second camera to take a second image. The second image includes a second face image, and the second image is an infrared image;
[0161] Exemplarily, after obtaining the image quality score of the first image, the electronic device can determine whether the first image meets the quality standard based on the quality score of the first image.
[0162] Among them, the first image meeting the quality standard means that when using the first image to detect facial features, the probability of detection failure due to the quality problem of the first image is less than the preset value. That is to say, if the first image meets the quality standard, it means that the image quality of the first image is relatively good. Therefore, it is highly unlikely that the facial feature detection fails due to quality problems. In other words, if the first image meets the quality standard but the facial feature detection fails using the first image, the reason for the failure is generally not the poor quality of the first image.
[0163] Correspondingly, the first image not meeting the image quality standard means that when using the first image to detect facial features, the probability of detection failure due to the quality problem of the first image is greater than the preset value. That is to say, if the first image does not meet the quality standard, it means that the image quality of the first image is relatively poor. Therefore, it is highly likely that the facial feature detection fails due to quality problems. In other words, if the first image does not meet the quality standard and the facial feature detection fails using the first image, the poor image quality is likely to be one of the reasons for the detection failure.
[0164] In a possible implementation manner, this application considers the situation where the quality score of the first image is less than or equal to the first threshold as the first image not meeting the quality standard, where the first threshold can be a value preconfigured based on business needs. Among them, if the quality score of the first image includes at least two of the quality score of the first face image, the quality score of the first left eye image, and the quality score of the first right eye image, then the quality score of the first image being less than or equal to the first threshold means that at least one of all the quality scores corresponding to the first image is less than or equal to the first threshold. That is to say, if the quality score of the first image includes multiple scores, as long as one score is less than or equal to the first threshold, it is considered that the first image does not meet the quality standard. On the contrary, if the quality score of the first image includes multiple scores, only when all scores are greater than the first threshold is it considered that the first image meets the quality standard. The following combines , and gives an exemplary introduction to the specific process of determining whether the first image meets the quality standard:
[0165] A1. Calculate the quality score face_score of the first face image.
[0166] A2. Determine whether face_score is greater than or equal to thr_score1. Exemplarily, the thr_score1 represents the above-mentioned first threshold.
[0167] A3. If face_score < thr_score1, it means that the first image does not meet the quality standard, that is, the quality score of the first image is less than or equal to the first threshold.
[0168] A4. If face_score ≥ thr_score1, optionally, the quality score left_eye_score of the first left-eye image and / or the quality score right_eye_score of the first right-eye image can be calculated.
[0169] A5. Determine whether left_eye_score is less than thr_score1 and whether right_eye_score is less than thr_score1.
[0170] A3. If left_eye_score < thr_score1 or right_eye_score < thr_score1, it means that the first image does not meet the quality standard.
[0171] A6. If left_eye_score ≥ thr_score1 and right_eye_score ≥ thr_score1, it means that the first image meets the quality standard, that is, the quality score of the first image is greater than the first threshold.
[0172] Through the above solution, it can be determined whether the image quality score of the first image is less than or equal to the first threshold. If so, the second camera is used to capture a second image. Wherein, the second image can be one of the multiple frames of images captured by the second camera.
[0173] The second image is an infrared image. That is to say, the second camera can be used to capture an infrared image, or the second camera can be used to obtain an image containing depth information. For example, the second camera is a Tof camera.
[0174] The second image includes a second face image, and the second image is used to detect face features. For example, the second image is used to detect whether the user's gaze point is within the display screen or for face recognition.
[0175] Optionally, after S240 and before S250, the method further includes: detecting the second image based on a face detection model and outputting a second detection result, where the second detection result indicates that the second image includes the second face image; or, the electronic device determines that the second image includes the second face image.
[0176] That is to say, after the electronic device captures a second image using the second camera, it detects whether the second image includes a face image. If it does, the subsequent operations of S250 are continued; if not, S240 can be re-executed, or it can continue to detect whether other images captured by the second camera include a face image, or the face feature detection process can be stopped. In this way, face feature detection can be performed only on the second image containing a face image, thereby saving resources. It can be understood that the specific implementation of detecting whether a face image is included in an image can refer to the descriptions in and this part and will not be elaborated here.
[0177] S250. Detect face features based on the second image.
[0178] Exemplarily, after capturing a second image using the second camera, face features are detected based on the second image. For example, corresponding to the application scenario shown, detecting face features based on the second image means: detecting whether the user's gaze point is within the display screen of the electronic device based on the second image.
[0179] For example, the electronic device determines a second feature image based on the second image. The second feature image includes one or more of a second face image, a second left eye image, and a second right eye image. Among them, the second face image is included in the second image, and the left eye image and the right eye image are included in the second face image. Then, it is determined whether the user's gaze point is within the display screen of the electronic device based on the second feature image. Specifically, feature data can be extracted from the second feature image, and then the extracted feature data is subjected to feature fusion to obtain fusion data, and the fusion data is processed using a gaze recognition model, and a gaze result is output. The gaze result is used to indicate whether the user's gaze point is within the display screen of the electronic device. The specific implementation is similar to the solution for detecting face features based on the first image introduced in the S220 part. For the sake of brevity, the specific solution will not be elaborated here.
[0180] Since the second image is an infrared image, the image quality of the second image is not affected by light. Therefore, detecting face features based on the second image can improve the detection success rate.
[0181] In summary, in the face feature detection method provided by the embodiments of the present application, if the face feature detection fails based on the first image captured by the first camera, and the reason for the failure may be that the image quality of the first image does not meet the quality standard, then the second camera is used to capture the second image, and the face features are detected based on the second image. That is to say, in the solution of the present application, after the face feature detection fails through the first image, it is not directly determined that the detection fails. Instead, it is necessary to first determine whether the image quality score of the first image is less than or equal to the first threshold. If so, it means that the image quality of the first image does not meet the quality standard. Therefore, it is very likely that the detection failure is caused by the image quality problem of the first image. For example, in a backlight scenario, the imaging quality of a color image is relatively poor. At this time, switching to the second camera to capture an infrared image can improve the detection success rate because the imaging quality of the infrared image is hardly affected by the light problem. Therefore, in the above solution, the first camera can be preferentially used to collect images and perform face feature detection, thereby reducing power consumption. If the detection fails and the reason for the failure may be an image quality problem, then the second camera is used to collect images to perform face feature detection, thereby improving the detection success rate.
[0182] The embodiments of the present application can be applied to various scenarios for detecting face features. After the face features are detected based on the second image, subsequent operations can be performed according to the detection results.
[0183] For example, corresponding to the scenario shown, detecting face features refers to the process of face recognition. If the detection is successful, the electronic device will automatically switch from the unlocked state (as shown in (a) of ) to the unlocked state (as shown in (b) of ).
[0184] Another example, corresponding to the scenario shown, detecting face features refers to detecting whether the user's gaze point is within the display screen of the electronic device. If the detection is successful, the electronic device controls the on / off state of the display screen of the electronic device according to the detection result. Specifically, for example, if the user's gaze point is within the display screen of the electronic device, the on-screen state of the display screen of the electronic device is maintained, and the screen-off countdown is reset; if the user's gaze point is outside the display screen of the electronic device, the screen-off countdown continues, and after the countdown expires, it automatically switches from the on-screen state to the screen-off state. To more clearly describe the application method of the solution provided by the embodiments of the present application in an actual scenario, an electronic device is introduced below in combination with this scenario. The electronic device includes modules corresponding to the above respective method embodiments for execution. The module can be software, hardware, or a combination of software and hardware. As As shown in the figure, the electronic device includes an image acquisition module, a face image acquisition module, a gaze recognition module, and a face image quality judgment module. Among them, the image acquisition module is used to judge the ambient light intensity when there are 4 seconds left in the screen-off countdown, and enable the RGB camera or the Tof camera to capture an image, and output the captured image to the face image acquisition module; the face image acquisition module is used to intercept the face image and the human eye image from the original image according to the face detection and key point detection algorithms; the gaze recognition module is used to judge whether the user is gazing at the display screen according to the gaze detection model; the face image quality judgment module is used to judge the image quality.
[0185] The following will combine the method flowchart shown in the figure to make an exemplary description of the execution processes of the above-mentioned modules.
[0186] B1. Judge whether the screen-off countdown is less than 4 seconds.
[0187] B2. Obtain the ambient light intensity.
[0188] Exemplarily, when the screen-off countdown is less than (or equal to) 4 seconds, the image acquisition module obtains the ambient light intensity (here, the ambient light smooth value is used to represent the ambient light intensity).
[0189] B3. Judge whether the ambient light smooth value is less than 15 lux.
[0190] B4. Enable the RGB camera.
[0191] Exemplarily, if the ambient light intensity is greater than or equal to 15 lux, the image acquisition module starts the RGB camera to capture an RGB image.
[0192] B5. Obtain the RGB image.
[0193] Exemplarily, the image acquisition module outputs the captured RGB image to the face image acquisition module.
[0194] B6. Judge whether a face is detected.
[0195] B7. Intercept the face image + the binocular images.
[0196] After the face image acquisition module obtains the RGB image from the image acquisition module, it judges whether a face is detected. If a face is detected, the face detection and key point detection algorithms are used to process the obtained RGB image to obtain the face image and the human eye (including the left eye and the right eye) images. If no face is detected, the RGB image is obtained again.
[0197] B8. Process the intercepted image using the gaze recognition model.
[0198] Exemplarily, the face image acquisition module outputs the acquired image to the gaze recognition module. The gaze recognition module processes the face image and the eye image according to the gaze recognition model.
[0199] B9. Determine whether the human eye is gazing at the screen.
[0200] Exemplarily, the gaze recognition module processes the face image and the eye image according to the gaze recognition model to determine whether the human eye is gazing at the screen.
[0201] B10. Perform image quality judgment.
[0202] B11. Determine whether the first image meets the quality standard.
[0203] Exemplarily, if the detection fails, the face image quality judgment module judges the quality of the RGB image. If the image quality of the RGB image does not meet the quality standard, the Tof camera is enabled to retake the IR image, and then the above detection process is executed according to the IR image, and the specific process will not be elaborated here.
[0204] B12. Keep the screen on, reset the screen-off countdown, and turn off the RGB camera.
[0205] Exemplarily, if the human eye is gazing at the screen, keep the screen on, reset the screen-off timer, and turn off the RGB camera.
[0206] B13. Enable the Tof camera.
[0207] Exemplarily, if the ambient light intensity is less than 15 lux, the image acquisition module directly activates the Tof camera to capture the IR image, and then executes the above detection process according to the IR image. The specific processes of B14 - B19 will not be elaborated here.
[0208] It should be noted that for the information interaction, execution process, etc. between the above modules / units, since they are based on the same concept as the method embodiments of this application, their specific functions and the technical effects brought can be specifically referred to in the method and system embodiment parts, and will not be elaborated here.
[0209] Corresponding to the methods given in the above method embodiments, the embodiments of this application also provide a corresponding system, which is deployed in an electronic device and used to execute the methods provided by the above method embodiments. Therefore, the content not described in detail can be referred to the above method embodiments, and for the sake of brevity, it will not be elaborated here.
[0210] Shows a system corresponding to an electronic device provided by an embodiment of the present application. The system may adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture, etc. In the embodiments of the present application, taking the Android system with a layered architecture as an example, the software system of the electronic device is exemplarily described.
[0211] See , the layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided from top to bottom into an application layer (also called the application program layer), an application framework layer (also called the application framework layer), a hardware abstraction layer (HAL), and a driver layer. Outside of this system, there is also a hardware layer.
[0212] The application layer may include multiple application programs, such as a dialing application, a gallery application, etc. In the embodiments of the present application, the application layer also includes a face feature detection software development kit (SDK), such as a gaze recognition SDK. The system of the electronic device and the third application program installed on the electronic device can detect face features by invoking the face feature detection SDK, such as identifying the position of the user's gaze point.
[0213] The framework layer provides application programming interfaces (APIs) and programming frameworks for the application programs in the application layer. The framework layer includes some predefined functions. In the embodiments of the present application, the framework layer may include a camera service interface, a face feature detection service interface, and an image quality analysis service interface. The camera service interface is used to provide an application programming interface and a programming framework for using the camera. The face feature detection service interface provides an application programming interface and a programming framework for using a face feature detection model (such as a gaze recognition model). The image quality analysis service interface provides an application programming interface and a programming framework for an algorithm for calculating an image quality score. It can be understood that the face feature detection service interface and the image quality analysis service interface can either be two interfaces or be combined into one interface, and the present application does not make a limitation. For the convenience of description, the embodiments of the present application take the face feature detection service interface and the image quality analysis service interface as two different interfaces as an example for illustration.
[0214] The hardware abstraction layer is an interface layer located between the framework layer and the driver layer, providing a virtual hardware platform for the operating system. In the embodiments of the present application, the hardware abstraction layer may include a camera hardware abstraction layer and a face feature detection process. The camera hardware abstraction layer may provide virtual hardware for camera device 1 (RGB camera), camera device 2 (Tof camera), or more camera devices. The calculation process of detecting face features through the face feature detection model is executed in this face feature detection process. For example, the calculation process of identifying the position of the user's gaze point through the gaze recognition module is executed in this face feature detection model.
[0215] The driver layer is the layer between hardware and software. The driver layer includes drivers for various hardware. The driver layer may include a camera device driver. The camera device driver is used to drive the sensor of the camera to collect images and drive the image signal processor to preprocess the images.
[0216] The hardware layer includes sensors, and optionally also includes a data buffer. Among them, the sensors include a first camera (such as an RGB camera) and a second camera (such as a Tof camera). For the functions of the first camera and the second camera, reference can be made to the description in the above method embodiments, which will not be elaborated here. Optionally, the sensors may further include a light sensor, which is used to detect the ambient light intensity.
[0217] Optionally, the data collected by the camera can be stored in the data buffer. When an upper-layer process or reference obtains the image data collected by the camera, it can obtain it from the data buffer.
[0218] Next, in combination with the above system structure, a brief introduction to the interaction process of the face feature detection method in the embodiments of the present application will be given. For the content not described in detail, reference can be made to the above method embodiments.
[0219] When the electronic device determines to detect face features (such as in the shown scenario, when the screen-off countdown is less than a preset value, it determines to detect face features), it calls the face feature detection service through the face feature detection SDK.
[0220] On the one hand, the face feature detection service can call the camera service in the framework layer to collect and obtain an image frame containing the user's facial image through the camera service. Specifically, the camera service can send an instruction to start the first camera by calling camera device 1 (the first camera) in the camera hardware abstraction layer. The camera hardware abstraction layer sends this instruction to the camera device driver in the driver layer. The camera device driver can start the first camera according to the above instruction. The instruction sent from camera device 1 to the camera device driver can be used to start the first camera. After the first camera is turned on, it collects optical signals and generates an electrical signal color image through the image signal processor.
[0221] On the other hand, the face feature detection service can create a face feature detection process and initialize a face feature detection model (such as a face image acquisition model, a gaze recognition model).
[0222] The color image generated by the image signal processor can be stored in the data buffer. After the face feature detection process is created and initialized, the color image data stored in the data buffer can be input into the face feature detection process. In the face feature detection process, the face feature detection model (such as the gaze recognition model) can be used to process the color image to determine whether the user's gaze point is within the display screen. Then the recognition result can be returned to the application layer face feature detection SDK via the camera service and the face feature detection service.
[0223] If the above face feature detection fails, the image quality analysis service can obtain the color image from the data buffer and analyze whether the image quality score of the image is less than or equal to the first threshold. If so, the image quality analysis service calls the camera service in the framework layer through the face feature detection service to re-collect and obtain an image frame containing the user's facial image through the camera service. Specifically, the camera service can send an instruction to start the second camera by calling the camera device 2 (the second camera) in the camera hardware abstraction layer. The camera hardware abstraction layer sends this instruction to the camera device driver in the driver layer. The camera device driver can start the second camera according to the above instruction. The instruction sent by the camera device 2 to the camera device driver can be used to start the second camera. After the second camera is turned on, it collects optical signals and generates an infrared image of an electrical signal through the image signal processor.
[0224] The infrared image generated by the image signal processor can be stored in the data buffer. After the face feature detection process is created and initialized, the infrared image data stored in the data buffer can be input into the face feature detection process. In the face feature detection process, the face feature detection model (such as the gaze recognition model) can be used to process the infrared image to determine whether the user's gaze point is within the display screen. Then the recognition result can be returned to the application layer face feature detection SDK via the camera service and the face feature detection service.
[0225] It should be noted that only the Android system is taken as an example in the embodiments of this application. In other operating systems (such as the IOS system, etc.), as long as the functions implemented by each functional module are similar to those of the embodiments of this application, the solution of this application can also be implemented.
[0226] Corresponding to the methods given in the above method embodiments, the embodiments of this application also provide a hardware architecture of a corresponding electronic device.
[0227] Exemplarily, Shows a detailed architecture diagram of an electronic device 1000 to which the present application is applicable.
[0228] As shown, the electronic device 1000 may include a processor 1010, an internal memory 1021, two or more cameras 1093 (the multiple displays can be represented by 2 to N), one or more displays 1094 (the multiple displays can be represented by 1 to N, where N is a positive integer greater than 1), and a sensor module 1080, where the sensor module includes an ambient light sensor 1080L.
[0229] Optionally, the electronic device 1000 may further include an external memory interface 1020, a universal serial bus (USB) interface 1030, a charging management module 1040, a power management module 1041, a battery 1042, an audio module 1070, a speaker 1070A, a receiver 1070B, a microphone 1070C, a headphone interface 1070D, a sensor module 1080, keys 1090, a motor 1091, an indicator 1092, and a subscriber identification module (SIM) card interface 1095, etc. In addition to the ambient light sensor, the sensor module 1080 may further include a pressure sensor 1080A, a gyroscope sensor 1080B, a barometric pressure sensor 1080C, a magnetic sensor 1080D, an acceleration sensor 1080E, a distance sensor 1080F, a proximity light sensor 1080G, a fingerprint sensor 1080H, a temperature sensor 1080J, a touch sensor 1080K, a bone sensor 1080M, etc.
[0230] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 1000. In other embodiments of the present application, the electronic device 1000 may include more or fewer components than shown, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0231] The processor 1010 may include one or more processing units. For example, the processor 1010 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc.
[0232] A memory may also be provided in the processor 1010 for storing instructions and data. In some embodiments, the memory in the processor 1010 is a cache memory. This memory can save the instructions or data that the processor 1010 has just used or recycled. If the processor 1010 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 1010, and thus improves the efficiency of the system.
[0233] The internal memory 1021 may be used to store computer-executable program code, and the executable program code includes instructions. The processor 1010 executes various functional applications and data processing of the electronic device 1000 by running the instructions stored in the internal memory 1021. In one example, when the computer program stored in the internal memory 1021 is called and run by the processor, the electronic device 1000 can implement the steps in any method embodiment in the above various embodiments.
[0234] The electronic device 1000 realizes the display function through the GPU, the display screen 1094, and the application processor, etc.
[0235] The camera 1093 is used to capture static images or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in standard formats such as RGB and YUV. In the embodiments of this application, the electronic device 1000 includes at least two cameras. One of the cameras can be used to capture color images, such as an RGB camera, and the other camera can be used to capture infrared images, such as a Tof camera. It can be understood that the electronic device 1000 can also include more than two cameras.
[0236] The display screen 1094 is used to display images, videos, etc. The display screen 1094 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLed, a MicroLed, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 1000 can include one or N display screens 1094, where N is a positive integer greater than 1.
[0237] The ambient light sensor 1080L is used to sense the ambient light brightness. The electronic device 1000 can adaptively adjust the brightness of the display screen 1094 according to the sensed ambient light brightness. The ambient light sensor 1080L can also be used to automatically adjust the white balance during photography. The ambient light sensor 1080L can also cooperate with the proximity light sensor 1080G to detect whether the electronic device 1000 is in a pocket to prevent accidental touch.
[0238] In the embodiments of the present application, the electronic device being in the screen-on state means that the display panel of the display screen 1094 is in the lit state, and the electronic device being in the screen-off state means that the display panel of the display screen 1094 is in the unlit state. When the electronic device is in the screen-on state, the display screen 1094 can display any interface such as the desktop, the negative first screen, the application interface of a certain application, the lock screen interface, etc.
[0239] Optionally, in one implementation, the electronic device 1000 can implement the shooting function through the ISP, the camera 1093, the video codec, the GPU, the application processor, etc., to obtain an image including a human face, and detect human face features based on the image.
[0240] Taking the application scenario shown as an example, in the embodiments of the present application, when the screen-off countdown is less than or equal to 4 s, the processor 1010 calls the ambient light sensor 1080L to obtain the current ambient light intensity. When the ambient light intensity is greater than or equal to the second threshold, it calls the RGB camera in the camera 1093 to capture a color image including a human face, and detects human face features based on the color image, such as detecting whether the user's fixation point is within the display screen 1094. If the detection fails, the processor 1010 obtains the image quality score of the image. If the image quality score is less than or equal to the first threshold, it calls the Tof camera in the camera 1093 to reshoot an infrared image including a human face, and detects human face features based on the infrared image.
[0241] It can be understood that if the units integrated in the above device embodiments are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / electronic device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0242] An embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device 1000, it enables the mobile terminal to implement the steps in the above-mentioned method embodiments when executed.
[0243] It should be noted that in the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0244] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0245] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
[0246] The terms "first", "second", "third", "fourth" and other various term labels (if any) in the specification, claims and drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or quantity. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than that illustrated or described herein.
[0247] The terms "including" and "having" and any variations thereof mean "including but not limited to", unless otherwise specifically emphasized. For example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0248] In the embodiments of the present application, words such as "exemplarily" or "for example" are used to give examples, illustrations or explanations. Any embodiment or design solution described as "exemplarily" or "for example" in the embodiments of the present application should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Rather, the use of words such as "exemplarily" or "for example" is intended to present relevant concepts in a specific manner.
[0249] In the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B; herein, "and / or" is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. In the description of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0250] In the various embodiments of the present application, if there is no special explanation and logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced to each other. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships. The specific operation methods in the method embodiments of the present application can also be applied to the device embodiments or system embodiments.
[0251] In the present application, "pre-configuration" may include predefined, for example, protocol definition. Among them, "predefined" can be implemented by pre-saving corresponding codes, tables or other means that can be used to indicate relevant information in a device (for example, including each network element). The present application does not limit its specific implementation manner.
[0252] In the schematic diagrams in the attached drawings of the specification of the present application, the dotted lines, arrows or boxes indicate optional steps or optional modules, and the dash-dotted lines or boxes indicate the content of the annotations.
Claims
1. A face feature detection method applied to an electronic device, the electronic device having a first camera and a second camera, characterized in that, The method includes: Obtaining the current ambient light intensity; When the ambient light intensity is greater than or equal to a second threshold, using the first camera to capture a first image, the first image includes a first face image, and the first image is a color image; Detecting face features based on the first image; When the face feature detection fails, calculating a first image quality score of the first face image; When the first image quality score of the first face image is greater than or equal to a first threshold, calculating a second image quality score of a first left-eye image and / or a third image quality score of a first right-eye image, wherein the first left-eye image and the first right-eye image are included in the first face image; When at least one of the second image quality score and the third image quality score is less than or equal to the first threshold, using the second camera to capture a second image, the second image includes a second face image, and the second image is an infrared image; Detecting face features based on the second image.
2. The method according to claim 1, wherein The first camera is used to obtain a two-dimensional image, and the second camera is used to obtain an image including depth information.
3. The method according to claim 2, wherein The first camera is a red, green, and blue (RGB) camera, and the second camera is a time-of-flight (ToF) camera.
4. The method according to claim 1, wherein The calculating the first image quality score of the first face image includes: Calculating a quality evaluation coefficient corresponding to the first face image, the quality evaluation coefficient including one or more of the following: brightness evaluation coefficient, contrast evaluation coefficient, sharpness evaluation coefficient; Calculating the first image quality score of the first face image according to the quality evaluation coefficient.
5. The method according to claim 1, characterized in that The calculating the first image quality score of the first face image includes: Calculating an average pixel intensity corresponding to the first face image as the image quality score of the first face image.
6. The method according to any one of claims 1 to 4, wherein The detecting face features based on the second image includes: Detecting whether a fixation point of the user is within a display screen of the electronic device according to the second image; The method further includes: Controlling a brightness on / off state of the display screen of the electronic device according to a detection result.
7. The method according to claim 6, wherein The display screen of the electronic device is currently in a lit state, and the controlling the brightness on / off state of the display screen of the electronic device according to the detection result includes: If the fixation point of the user is within the display screen of the electronic device, maintaining the lit state of the display screen of the electronic device and resetting a screen-off countdown.
8. The method according to claim 7, wherein The using the first camera to capture a first image includes: When the screen-off countdown is less than or equal to a third threshold, using the first camera to capture a first image based on the current ambient light intensity.
9. The method according to claim 6, characterized in that, The detecting whether a fixation point of the user is within the display screen of the electronic device according to the second image includes: Determine a second feature image according to the second image, where the second feature image includes a second face image, a second left eye image, and a second right eye image. Among them, the second face image is included in the second image, and the second left eye image and the second right eye image are included in the second face image; Determine whether the gaze point of the user is within the display screen of the electronic device according to the second feature image.
10. The method according to claim 9, characterized in that, The determining whether the gaze point of the user is within the display screen of the electronic device according to the feature image includes: Use an eye feature extraction model to extract a second left eye feature and a second right eye feature from the second left eye image and the second right eye image respectively; Use a face feature extraction model to extract a second face feature from the second face image; Fuse the second left eye feature, the second right eye feature, and the second face feature to obtain fused data; Use a gaze recognition model to process the fused data and output a gaze result, where the gaze result is used to indicate whether the gaze point of the user is within the display screen of the electronic device.
11. The method according to any one of claims 1 to 3, wherein Before performing face feature detection according to the first image, the method further includes: Detect the first image based on a face detection model and output a first detection result, where the first detection result indicates that the first image includes the first face image; Before performing face feature detection according to the second image, the method further includes: Detect the second image based on a face detection model and output a second detection result, where the second detection result indicates that the second image includes the second face image.
12. An electronic device, characterized in that, The structure of the electronic device includes a processor and a memory; The memory is used to store a program that supports the electronic device to execute the method provided in any one of claims 1 to 11, and to store data involved in implementing the method described in any one of claims 1 to 11; The processor is configured to execute the program stored in the memory.
13. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium, and when it runs on a computer, the computer is caused to execute the method described in any one of claims 1 to 11.
Citation Information
Patent Citations
Method and device for recognizing human face
CN108388878A
Data acquisition method, system, device and equipment and storage medium
CN116433880A