Face gaze unlocking method and electronic device
By combining image and sensor data to optimize facial and eye feature information, and using gaze models and neural network training samples, the accuracy and security issues of facial gaze unlocking in complex scenarios are solved, improving the accuracy and security of unlocking.
Patent Information
- Application Number
- CN202010452422.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-26
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2040-05-26
AI Technical Summary
Existing technologies are insufficient in terms of accuracy and security for facial gaze unlocking in complex scenarios. They are easily affected by factors such as the posture of electronic devices, ambient lighting, and shooting distance, leading to unlocking failure or reduced security.
Scene information is determined by collecting images and sensor data. An image optimizer is used to optimize facial and eye feature information. An unlocking judgment is made in combination with a gaze model. A dynamic or preset gaze score threshold is set. A neural network model is used to train samples to improve recognition accuracy.
It improves the accuracy of face gaze unlocking in complex scenarios, reduces the non-gaze unlocking rate, and enhances the user's unlocking experience and security.
Smart Images

Figure CN113723144B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of terminals, and in particular to a face gaze unlocking method and an electronic device. BACKGROUND
[0002] With the development of terminal technology, face unlocking and payment functions on electronic devices (such as mobile phones, tablet computers, etc.) have become popular. In order to improve the security of face unlocking and payment, most face unlocking functions have added a face gaze function to avoid others using their own mobile phones to unlock the mobile phone in a non-voluntary situation of the user, such as when the user is asleep or when others forcibly use the mobile phone to scan the face, thereby peeping or stealing user information.
[0003] In order to increase unlocking security, the prior art uses a face gaze unlocking scheme to unlock, but in actual use, various factors such as the posture of the electronic device, the ambient light, the shooting distance, etc. can affect face gaze recognition, which may misrecognize a face gaze picture as a non-gaze picture, thereby causing unlocking failure, affecting the user's unlocking experience, or may misrecognize a non-gaze picture as a gaze picture, thereby causing unlocking success and reducing the security of unlocking. The prior art does not have a method that can accurately implement face gaze unlocking in complex scenarios (such as strong light, long distance, etc.). SUMMARY
[0004] Embodiments of the present application provide a face gaze unlocking method and an electronic device to improve the accuracy of face gaze unlocking in complex scenarios.
[0005] In a first aspect, embodiments of the present application provide a face gaze unlocking method, which can be executed by an electronic device. The method includes: the electronic device collects a first image and acquires sensor data in the electronic device when the first image is collected, determines scene information according to the first image and the sensor data, wherein the scene information includes at least one of the following: current light intensity, shooting distance, posture of the electronic device, camera temperature. The electronic device extracts face, eye feature information and face posture information from the first image, locates a first behavior scene corresponding to the first image according to the scene information and the face posture information, when the first behavior scene is a non-normal behavior scene, uses an image optimizer corresponding to the first behavior scene to optimize the face and eye feature information to obtain optimized face and eye feature information, and then inputs the optimized face and eye feature information into a gaze model to obtain a gaze score, and when it is determined that the gaze score is greater than a gaze score threshold, unlocking is performed.
[0006] Based on the scheme, by combining the collected first image and the sensor data when the first image is collected, the current behavior scene is located, and then the corresponding image optimizer is matched according to the behavior scene to optimize the face and eye feature information, which can reduce the noise of the input data of the gaze model and enhance the effectiveness of the data, thereby improving the accuracy of face gaze unlocking in complex scenes and reducing the non-gaze unlocking rate.
[0007] In a possible design, before extracting the face, eye feature information and face pose information from the first image, the electronic device can also determine the device pose of the electronic device according to the sensor data, determine the deflection angle between the direction of the long axis of the screen when the electronic device is in the portrait screen state and the direction of the long axis of the screen when the electronic device is in the device pose, and rotate the first image by the deflection angle. Through this design, the first image collected in a non-portrait screen state can be rotated in combination with the sensor data, which facilitates accurate extraction of face and eye feature information and can well solve the problem of unlocking when the electronic device is in a reverse screen state.
[0008] In a possible design, before determining that the gaze score is greater than the gaze score threshold, the electronic device can also set different gaze score thresholds for different abnormal behavior scenes before unlocking. Two possible implementation manners are provided below: Implementation manner one, the electronic device determines a preset score threshold corresponding to the first behavior scene according to the first behavior scene and the first correspondence relationship, and takes the preset score threshold corresponding to the first behavior scene information as the gaze score threshold, where the first correspondence relationship includes the correspondence relationship between the preset behavior scene and the preset score threshold.
[0009] Implementation manner two, the electronic device determines a dynamic threshold range corresponding to the first behavior scene, determines a dynamic threshold from the dynamic threshold range corresponding to the first behavior scene according to the scene information, and determines the gaze score threshold according to the static threshold corresponding to the normal behavior scene and the determined dynamic threshold.
[0010] Through the above two implementation manners, the gaze threshold is set according to the behavior scene, which not only can obtain a face gaze detection result closer to the real situation and improve the user's unlocking experience, but also can enhance the robustness and generalization of the gaze model in certain specific scenes.
[0011] In a possible design, the method further includes: obtaining a gaze image set and a non-gaze image set, the gaze image set including gaze images collected under each preset behavior scenario, and the non-gaze image set including non-gaze images collected under each preset behavior scenario; for each gaze image in the gaze image set, performing optimization on facial and eye feature information in the gaze image by using an image optimizer corresponding to a behavior scenario at which the gaze image is collected, to obtain first optimized facial and eye feature information; for each non-gaze image in the non-gaze image set, performing optimization on facial and eye feature information in the non-gaze image by using an image optimizer corresponding to a behavior scenario at which the non-gaze image is collected, to obtain second optimized facial and eye feature information; setting a behavior scenario label for each first optimized facial and eye feature information and each second optimized facial and eye feature information; and inputting the first optimized facial and eye feature information with the set behavior scenario label and the second optimized facial and eye feature information with the set behavior scenario label into a neural network model for model training, to obtain a gaze model.
[0012] By this design, the neural network model is trained by using training samples with set behavior scenario labels, and a gaze model that can accurately recognize the category of a to-be-detected image can be obtained, thereby improving the accuracy of facial gaze unlocking.
[0013] In a possible design, the sensor data includes at least one of the following: data collected by a posture sensor in the electronic device, data collected by a distance sensor, data collected by an ambient light sensor, and data collected by a temperature sensor.
[0014] In a second aspect, an embodiment of the present application further provides an electronic device. The electronic device includes a processor and a memory; the memory is configured to store an image and one or more computer programs; when the one or more computer programs stored in the memory are executed by the processor, the electronic device is enabled to implement the technical solutions of the above-described first aspect and any possible design of the first aspect.
[0015] In a third aspect, an embodiment of the present application further provides an electronic device, which includes modules / units for executing the method of the above-described first aspect or any possible design of the first aspect; these modules / units can be implemented by hardware, or by hardware executing corresponding software.
[0016] In a fourth aspect, a chip in an embodiment of the present application is coupled with a memory in an electronic device, and executes the technical solutions of the first aspect and any possible design of the first aspect of the embodiment of the present application; in the embodiment of the present application, "coupled" means that two components are directly or indirectly combined with each other.
[0017] In a fifth aspect, a computer readable storage medium of an embodiment of the present application includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the technical solutions of the first aspect of the embodiment of the present application and any possible design of the first aspect.
[0018] In a sixth aspect, a program product of an embodiment of the present application includes instructions that, when executed on an electronic device, cause the electronic device to perform the technical solutions of the first aspect of the embodiment of the present application and any possible design of the first aspect.
[0019] These aspects and other aspects of the present application will be more apparent in the following description of the embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 1.
[0021] Figure 2 A software structural block diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 2.
[0022] Figure 3 A training process schematic diagram of a face gaze detection provided by an embodiment of the present application is shown in FIG. 3.
[0023] Figure 4 A flowchart of a face gaze unlocking method provided by an embodiment of the present application is shown in FIG. 4.
[0024] Figure 5 A schematic diagram of a user face provided by an embodiment of the present application is shown in FIG. 5.
[0025] Figure 6 A schematic diagram of a landscape state and a portrait state of an electronic device of an embodiment of the present application is shown in FIG. 6.
[0026] Figure 7 A schematic diagram of a user face provided by an embodiment of the present application is shown in FIG. 7.
[0027] Figure 8 A flowchart of another face gaze unlocking method provided by an embodiment of the present application is shown in FIG. 8.
[0028] Figure 9 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 9. DETAILED DESCRIPTION
[0029] In the following, some terms in the embodiments of the present application are explained and described to facilitate understanding by those skilled in the art.
[0030] The plurality referred to in the embodiments of the present application means greater than or equal to two.
[0031] It should be noted that the term "and / or" in this document is merely used to describe associated objects, and can exist in three forms, for example, A and / or B can mean that there are three cases, A exists alone, A and B exist together, and B exists alone. In addition, the character " / " in this document, unless otherwise specified, generally represents an "or" relationship between the front and rear associated objects. And in the description of the embodiments of the present application, "first", "second", etc. are used only for the purpose of distinguishing the description, and cannot be understood as indicating or implying relative importance, indicating or implying order, or implying that the number of indicated technical features.
[0032] Various embodiments disclosed in the present application can be applied to electronic devices capable of realizing face gaze unlocking function. In some embodiments of the present application, the electronic device can be a portable terminal, such as a mobile phone, a tablet computer, a wearable device (such as a smart watch) with wireless communication function, a camera, a notebook computer, etc. The portable terminal contains a device (such as a processor) capable of collecting images and extracting features from the collected images. Exemplary embodiments of the portable terminal include, but are not limited to, a portable terminal running an Android operating system or other operating systems. The above-mentioned portable terminal can also be other portable terminals as long as it can collect images and perform image processing (such as feature information or posture information extraction, optimization, obtaining gaze score, etc.) on the collected images. It should also be understood that in some other embodiments of the present application, the above-mentioned electronic device can not be a portable terminal, but a desktop computer capable of collecting images and performing image processing (such as feature information or posture information extraction, optimization, obtaining gaze score, etc.) on the collected images. Or other operating systems. The above-mentioned portable terminal can also be other portable terminals as long as it can collect images and perform image processing (such as feature information or posture information extraction, optimization, obtaining gaze score, etc.) on the collected images. It should also be understood that in some other embodiments of the present application, the above-mentioned electronic device can not be a portable terminal, but a desktop computer capable of collecting images and performing image processing (such as feature information or posture information extraction, optimization, obtaining gaze score, etc.) on the collected images.
[0033] In some other embodiments of the present application, the electronic device can also not have the function of image processing (such as feature information or posture information extraction, optimization, obtaining gaze score, etc.), but have communication function. For example, after the electronic device collects images, it can send the images to other devices such as servers, and other devices use the face gaze unlocking method provided by the embodiments of the present application to perform image processing (such as feature information or posture information extraction, optimization, obtaining gaze score, etc.) on the images, and then send the image processing results to the electronic device, and the electronic device determines whether to unlock according to the image processing results.
[0034] Figure 1 The structure diagram of the electronic device 100 is shown.
[0035] It should be understood that the electronic device 100 is only one example and that the electronic device 100 can have more or fewer components than shown in the figure, can combine two or more components, or can have a different configuration of components. The various components shown in the figure can be implemented in hardware, software, or a combination of both hardware and software including one or more signal processing and / or application specific integrated circuits.
[0036] As shown in Figure 1 The electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headset interface 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0037] The various components of the electronic device 100 will be described in detail below: Figure 1
[0038] The processor 110 can include one or more processing units, for example, the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices or integrated in one or more processors. The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching and executing instructions.
[0039] The processor 110 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can save instructions or data that the processor 110 has just used or repeatedly uses. If the processor 110 needs to use the instructions or data again, it can directly call from the memory, thereby avoiding repeated access and reducing the waiting time of the processor 110, thus improving the efficiency of the system.
[0040] The processor 110 can run the software code of the face gaze unlocking method provided by the embodiments of the present application. When the processor 110 integrates different devices, such as integrating CPU and GPU, the CPU and GPU can cooperate to execute the face gaze unlocking method provided by the embodiments of the present application, such as part of the algorithm in the face gaze unlocking method is executed by the CPU and another part of the algorithm is executed by the GPU, to obtain faster processing efficiency.
[0041] In some embodiments, the processor 110 can include one or more interfaces. For example, the interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc. The electronic device 100 can use different interface connection manners in the above embodiments, or a combination of multiple interface connection manners.
[0042] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement a data storage function. For example, files such as music, videos, and captured images are stored in the external memory card.
[0043] The internal memory 121 can be used to store computer executable program codes including instructions. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application required for a function (such as a sound playing function, an image playing function, etc.), etc. The data storage area can store data created during use of the electronic device 100 (such as sensor data, captured images, etc.), etc. In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various function applications and data processing of the electronic device 100 by running instructions stored in the internal memory 121 and / or instructions stored in a memory disposed in the processor.
[0044] The internal memory 121 can also store software code of the face gaze unlocking method provided by the embodiments of the present application. When the processor 110 runs the code, the face gaze unlocking process below is executed, and the face gaze unlocking function is realized.
[0045] The internal memory 121 can also store other contents, for example, the first correspondence relationship between the preset behavior scene and the preset score threshold is stored in the internal memory 121. After the first behavior scene corresponding to the first image is located, the processor 110 can determine the preset score threshold corresponding to the first behavior scene as the gaze score threshold according to the first correspondence relationship between the first behavior scene and the first correspondence relationship in the internal memory 121. In this way, the processor 110 can determine whether to perform unlocking according to the determined gaze score and gaze score threshold.
[0046] The internal memory 121 can also store the dynamic threshold range corresponding to the preset behavior scene. For example, after the first behavior scene corresponding to the first image is located, the electronic device 100 can determine the dynamic threshold range corresponding to the first behavior scene according to the dynamic threshold range corresponding to the preset behavior scene stored in the internal memory 121, then determine a dynamic threshold from the dynamic threshold range corresponding to the first behavior scene according to the scene information, and then determine the gaze score threshold according to the default scene corresponding static threshold and the determined dynamic threshold. Further, the processor 110 can determine whether to perform unlocking according to the determined gaze score and gaze score threshold.
[0047] It should be understood that the internal memory 121 can also store other contents mentioned below, such as gaze models, etc.
[0048] The wireless communication function of the electronic device 100 can be realized by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, etc.
[0049] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 1 can be multiplexed as a diversity antenna of a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0050] The mobile communication module 150 can provide a solution for wireless communication including 2G / 3G / 4G / 5G, etc. applied to the electronic device 100. The mobile communication module 150 can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves by the antenna 1, and perform filtering, amplification, etc. on the received electromagnetic waves, and transfer the same to the modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor, and radiate the same as electromagnetic waves through the antenna 1. In some embodiments, at least part of the functional modules of the mobile communication module 150 can be disposed in the processor 110. In some embodiments, at least part of the functional modules of the mobile communication module 150 can be disposed in the same device as at least part of the modules of the processor 110.
[0051] The modem processor can include a modulator and a demodulator. The modulator is configured to modulate a low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is configured to demodulate a received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. The low-frequency baseband signal processed by the baseband processor is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the microphone 170B, etc.), or displays an image or a video through the display screen 194. In some embodiments, the modem processor can be a separate device. In other embodiments, the modem processor can be independent of the processor 110, and disposed in the same device as the mobile communication module 150 or other functional modules.
[0052] The wireless communication module 160 can provide a solution for wireless communication including wireless local area networks (WLAN) (e.g., wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc. applied on the electronic device 100. The wireless communication module 160 can be one or more devices that integrate at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency-modulates and filters the electromagnetic wave signals, and transmits the processed signals to the processor 110. The wireless communication module 160 can also receive signals to be transmitted from the processor 110, frequency-modulate them, amplify them, and radiate them as electromagnetic waves via the antenna 2.
[0053] In some embodiments of the present application, the antenna 1 and the mobile communication module 150 of the electronic device 100 are coupled, and the antenna 2 and the wireless communication module 160 are coupled, so that the electronic device 100 can communicate with a network and other devices through wireless communication technology, e.g., transmit captured images to other devices, process them to determine whether the electronic device 100 can be unlocked, and then receive the results transmitted by the other devices.
[0054] The functions of several sensors included in the sensor module 180 are described below.
[0055] The pressure sensor 180A is used to sense a pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be disposed on the display screen 194. When a force is applied to the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device 100 determines the intensity of the pressure according to the change in capacitance. When a touch operation is applied to the display screen 194, the electronic device 100 detects the intensity of the touch operation according to the pressure sensor 180A. The electronic device 100 can also calculate the position of the touch according to the detection signal of the pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation instructions.
[0056] The gyroscope sensor 180B can be used to determine the motion posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., x, y, and z axes) can be determined by the gyroscope sensor 180B. The gyroscope sensor 180B can be used for anti-shake shooting. The gyroscope sensor 180B can also be used for navigation, motion sensing game scenarios.
[0057] The acceleration sensor 180E can detect the magnitude of acceleration of the electronic device 100 in various directions (generally three axes). The magnitude and direction of gravity can be detected when the electronic device 100 is stationary. It can also be used to identify the electronic device posture, applied to landscape / portrait screen switching, pedometer, etc.
[0058] The distance sensor 180F is used to measure distance. The electronic device 100 can measure distance by infrared or laser. In some embodiments, the shooting scene, the electronic device 100 can use the distance sensor 180F to measure distance to achieve fast focusing.
[0059] The touch sensor 180K, also known as a "touch panel". The touch sensor 180K can be disposed on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, also known as a "touch screen". The touch sensor 180K is used to detect touch operations acting on or near it. The touch sensor can pass the detected touch operation to the application processor to determine the touch event type. The display screen 194 can provide visual output related to the touch operation. In other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, which is different from the position of the display screen 194.
[0060] The ambient light sensor 180L is used to sense the ambient light brightness. The electronic device 100 can adaptively adjust the display screen 194 brightness according to the sensed ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance when shooting. The ambient light sensor 180L can also cooperate with the proximity light sensor 180G to detect whether the electronic device 100 is in the pocket to prevent false touch.
[0061] The temperature sensor 180J is used to detect temperature. In some embodiments, the electronic device 100 uses the temperature detected by the temperature sensor 180J to perform temperature processing strategies. For example, when the temperature reported by the temperature sensor 180J exceeds a threshold value, the electronic device 100 reduces the performance of the processor located near the temperature sensor 180J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is lower than another threshold value, the electronic device 100 heats the battery 142 to avoid abnormal shutdown of the electronic device 100 caused by low temperature. In other embodiments, when the temperature is lower than yet another threshold value, the electronic device 100 performs voltage boosting on the output voltage of the battery 142 to avoid abnormal shutdown caused by low temperature.
[0062] The keys 190 include a power key, a volume key, and the like. The keys 190 can be mechanical keys or touch keys. The electronic device 100 can receive a key input and generate a key signal input related to user settings and function control of the electronic device 100.
[0063] The electronic device 100 can implement a photographing function through an ISP, a camera 193, a video codec, a GPU, a display 194, and an application processor, and the like.
[0064] The ISP is used to process data fed back by the camera 193. For example, when taking a photo, the shutter is opened, light is transmitted to the camera photosensitive element through the lens, and the light signal is converted into an electrical signal. The camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize algorithms for image noise, brightness, and skin color. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be disposed in the camera 193.
[0065] The camera 193 is used to capture still images or videos. Objects generate optical images through lenses and project them onto photosensitive elements. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into a standard RGB, YUV, or the like image signal. In some embodiments, the electronic device 100 can include one or N cameras 193, where N is a positive integer greater than 1.
[0066] The electronic device 100 implements a display function through a GPU, a display 194, and an application processor, and the like. The GPU is a microprocessor for image processing, connected to the display 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs that execute program instructions to generate or change display information.
[0067] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.
[0068] In this embodiment, the display screen 194 can be a single flexible display screen, or it can be a spliced display screen composed of two rigid screens and a flexible screen located between the two rigid screens.
[0069] although Figure 1 As not shown in the diagram, the electronic device 100 may also include a Bluetooth device, a positioning device, a flash, a miniature projection device, a near field communication (NFC) device, etc., which will not be described in detail here.
[0070] The following embodiments can all be implemented in an electronic device 100 (such as a mobile phone, tablet computer, etc.) having the above-described hardware structure.
[0071] Figure 2 A software structure block diagram of an electronic device provided in an embodiment of this application is shown. For example... Figure 2 As shown, the software architecture of an electronic device can be a layered architecture. For example, the software can be divided into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer (framework, FWK), the Android runtime and system libraries, and the kernel layer.
[0072] The application layer can include a series of application packages. For example... Figure 2As shown, the application layer can include camera, settings, skin modules, user interface (UI), and third-party applications. Third-party applications can include WeChat, QQ, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS.
[0073] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer can include some predefined functions. For example... Figure 2 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.
[0074] The window manager is used to manage windowed applications. It can obtain the screen size, determine if a status bar is present, lock the screen, and capture screenshots. The content provider stores and retrieves data, making this data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.
[0075] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0076] A phone manager is used to provide communication functions for electronic devices. For example, it manages call status (including connection and disconnection).
[0077] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0078] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.
[0079] The Android runtime consists of core libraries and a virtual machine. The Android runtime is responsible for scheduling and managing the Android system.
[0080] The core library consists of two parts: one part contains the functionalities that the Java language needs to call, and the other part is the Android core library. The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0081] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0082] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.
[0083] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0084] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0085] A 2D graphics engine is a graphics engine for 2D drawing.
[0086] In addition, the system library may also include a face detection module, a pose detection module, a data fusion module, a threshold adjustment module, and a gaze processing module. The face detection module performs face detection on the acquired first image. If a face is detected, it extracts face and eye feature information and sends the first image to the pose detection module for face pose detection; if no face is detected, it proceeds to anomaly handling and stops the subsequent process. The pose detection module extracts face pose information from the first image. The data fusion module fuses sensor data and face pose information to locate the current behavior scene. The threshold adjustment module outputs a dynamic threshold for a specific scene based on the input behavior scene information. The gaze processing module optimizes the face and eye feature information using an image optimizer corresponding to the first behavior scene, inputs the optimized face and eye feature information into a gaze model to obtain an output gaze score, and determines whether to unlock based on the output gaze score.
[0087] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.
[0088] The hardware layer may include a display screen and various sensors, such as the attitude sensor (e.g., accelerometer, gyroscope), ambient light sensor, distance sensor, temperature sensor, etc. involved in the embodiments of this application.
[0089] The following describes the software and hardware workflow of an electronic device using a face gaze unlocking method according to an embodiment of this application. As an example, the system library inputs a first image into a face detection module. The face detection module detects that the first image contains a face, extracts face and eye feature information from the first image, and sends the first image to a face pose detection module and the face and eye feature information to a data fusion module. The pose detection module extracts face pose information from the first image and sends it to the data fusion module. The data fusion module can locate the current behavioral scene based on the received information. The gaze processing module uses an image optimizer corresponding to the first behavioral scene to optimize the face and eye feature information, and inputs the optimized face and eye feature information into a gaze model to obtain an output gaze score and output a result indicating whether unlocking is successful. The result output by the gaze processing module is displayed on a screen. For example, if the result is successful unlocking, the screen displays the unlocked interface; if the result is unsuccessful unlocking, the screen displays a failure message.
[0090] The specific implementation process of the face gaze unlocking method in the embodiments of this application is described below with reference to the accompanying drawings.
[0091] In this embodiment of the application, the process of an electronic device performing face gaze detection on a face image using a deep learning algorithm can include two parts: a training process and a testing process.
[0092] The training process is as follows: Figure 3As shown, the overall approach is as follows: Face images collected under various preset behavioral scenarios are acquired. The gazed images are classified as positive examples, resulting in a gazed image set. The non-gaze images (i.e., images with non-gaze, closed eyes, or fake eyes) are classified as negative examples, resulting in a non-gaze image set. Face detection and feature point detection are performed on each gazed image in the gazed image set, and optimized to obtain first face and eye feature information. Similarly, face detection and feature point detection are performed on each non-gaze image in the non-gaze image set, and optimized to obtain second face and eye feature information. Behavioral scenario labels are then assigned to each optimized first and second face and eye feature information. These optimized first and second face and eye feature information with behavioral scenario labels are used as a training set and input into a neural network model for training, resulting in a gaze model. Then, a test set with known classification results is selected to validate the fixation model. This involves using the test set as input to the fixation model, obtaining test results, and comparing these results with the known classification results. If the similarity is high, training is complete; if the similarity is low, retraining continues until the identified test results are similar to the known classification results. The optimal model is selected through test set validation and will be used to classify fixated and non-fixated images.
[0093] The overall idea of the actual test process is as follows: the trained gaze model is used to identify the image to be detected in order to determine whether the image to be detected is a gazed image.
[0094] In this embodiment, the electronic device can train the gaze model before leaving the factory; that is, the machine learning model in the electronic device after leaving the factory has already been trained and can be directly used for the testing process. Alternatively, the electronic device can also train the gaze model after leaving the factory and then use the trained gaze model for the testing process. The testing process is described in detail below.
[0095] Please see Figure 4 This is a flowchart illustrating a face gaze unlocking method provided in an embodiment of this application. Figure 4 As shown, the process of this method may include:
[0096] Step 401: The electronic device acquires a first image and obtains sensor data from the electronic device during the acquisition of the first image.
[0097] For example, the electronic device can capture a first image via camera 193. The first image can be an infrared image, and it can include the user's face, such as the entire face or a partial face. A partial face could be due to the user's entire face being partially obscured by a mask, sunglasses, or other occupants, or only part of the user's face entering the camera's field of view, resulting in only a partial face being captured. Alternatively, the first image may not include the user's face. In this case, no facial or eye feature information can be extracted from the first image, and the electronic device will process the error and will not proceed with subsequent unlocking steps.
[0098] The sensor data includes at least one of the following: data collected by an attitude sensor in an electronic device, data collected by a distance sensor, data collected by an ambient light sensor, and data collected by a temperature sensor.
[0099] Step 402: The electronic device determines scene information based on the first image and sensor data.
[0100] The scene information may include at least one of the following: current light intensity, shooting distance, the orientation of the electronic device, and camera temperature.
[0101] Step 403: The electronic device extracts face, eye feature information and face pose information from the first image.
[0102] Before step 403, face detection can be performed on the first image. If a face is detected, face pose detection and eye feature detection can be performed. If no face is detected, anomaly processing is performed and the face gaze unlock result is not output.
[0103] For face pose detection, if facial pose features are detected, such as tilting the head back or turning the head to the side, face pose information is extracted. This facial pose information can include types such as frontal face, side view, tilting the head back, and tilting the head down, as well as the angle of each type relative to a frontal face. If no facial pose is detected, anomaly handling is performed, and no face gaze unlock result is output.
[0104] For eye feature detection, if eye features are detected, face and eye feature information are extracted. If eye features are not detected, anomaly handling is performed, and no face gaze unlock result is output.
[0105] For example, from such Figure 5 Face and eye feature information is extracted from the first image shown in (A). This face and eye feature information may include the facial contour and feature information of various parts of the face, such as... Figure 5 The features of the eyes, eyebrows, nose, mouth, ears, etc. shown in (B) are shown.
[0106] Step 404: The electronic device locates the first behavior scene corresponding to the first image based on the scene information and the face pose information.
[0107] In one example, an electronic device can use a classifier to classify scene information and facial pose information to locate the behavioral scene, thus obtaining the first behavioral scene corresponding to the first image. For example, the classifier can be a support vector machine (SVM), or other types of classifiers can be used, such as perceptron methods, neural network methods, radial basis function (RBF) methods, etc.
[0108] In one possible implementation, if the first behavioral scenario is a normal behavioral scenario, such as normal light, normal distance, normal camera temperature, and a frontal face, the electronic device skips steps 405 and 406 after executing step 404. That is, it does not optimize the face and eye feature information corresponding to the normal behavioral scenario. The electronic device directly inputs the face and eye feature information into the gaze model to obtain the gaze score under the normal scenario. Then, when it is determined that the gaze score is greater than the gaze score threshold, it unlocks the device.
[0109] If the first behavior scenario is an abnormal behavior scenario, such as strong light, normal distance, normal camera temperature, and a frontal face, the background light is very strong when the first image is captured in such a scenario, and the captured face is relatively dark, resulting in blurred eye features in the first image. The image optimizer corresponding to the first behavior scenario in step 405 can be used to optimize the face and eye feature information in the first image, such as enhancing the face and eye feature information. Then, the enhanced face and eye feature information is input into the gaze model to obtain the gaze score under the first behavior scenario.
[0110] Step 405: When the first behavior scenario is an abnormal behavior scenario, the electronic device uses the image optimizer corresponding to the first behavior scenario to optimize the face and eye feature information to obtain the optimized face and eye feature information.
[0111] In some embodiments, the electronic device is provided with image optimizers corresponding to various preset behavioral scenarios. After locating the first behavioral scenario corresponding to the first image, the electronic device determines the image optimizer corresponding to the first behavioral scenario from the image optimizers corresponding to various preset behavioral scenarios, and then uses the image optimizer corresponding to the first behavioral scenario to optimize the facial and eye feature information.
[0112] In some other embodiments, the electronic device may be configured with optimization parameters corresponding to various preset behavior scenarios. After locating the first behavior scenario corresponding to the first image, the electronic device determines the optimization parameters corresponding to the first behavior scenario from the optimization parameters corresponding to various preset behavior scenarios, and then uses the optimization parameters corresponding to the first behavior scenario to optimize the facial and eye feature information.
[0113] Since the main sources of noise in images differ under different behavioral scenarios, by locating the behavioral scenario, we can predict what type of noise is the main source of noise in the current behavioral scenario. This allows us to set specific image optimizers or optimization parameters to optimize the image in a targeted manner, achieving better optimization results.
[0114] Step 406: The electronic device inputs the optimized facial and eye feature information into the gaze model to obtain a gaze score.
[0115] Step 407: When the electronic device determines that the gaze score is greater than the gaze score threshold, it unlocks.
[0116] Correspondingly, the electronic device will not unlock if it determines that the fixation score is less than or equal to the fixation score threshold.
[0117] Here, the same fixation score threshold can be used for all behavioral scenarios, meaning the fixation score threshold is the same for both normal and abnormal behavioral scenarios. Alternatively, different fixation score thresholds can be used for different behavioral scenarios.
[0118] For example, the gaze score threshold for normal behavior scenarios is 50. Electronic devices will unlock when the gaze score is determined to be greater than 50, and will not unlock when the gaze score is determined to be less than or equal to 50.
[0119] In this embodiment, by combining the acquired first image and the sensor data at the time of acquiring the first image, the current behavioral scene is located. Then, according to the behavioral scene, the corresponding image optimizer is matched to optimize the facial and eye feature information. This can reduce the noise of the input data of the gaze model in a targeted manner, enhance the effectiveness of the data, and thus improve the accuracy of facial gaze unlocking in complex scenes and reduce the non-gaze unlocking rate.
[0120] In the above embodiments, although the facial and eye feature information corresponding to abnormal behavior scenarios is enhanced, it is not possible to enhance the facial and eye feature information of the image to the optimal effect. Therefore, different gaze score thresholds can be set for different abnormal behavior scenarios, so as to obtain a more realistic facial gaze detection result and improve the user's unlocking experience.
[0121] In one optional implementation, the electronic device includes a preset dynamic threshold range corresponding to a behavioral scenario. This dynamic threshold range can be obtained by training a threshold model using test sets corresponding to different scenarios. After step 404 and before step 407, the face gaze unlocking method may further include the following steps:
[0122] Step 408: Determine the dynamic threshold range corresponding to the first behavior scenario.
[0123] For example, the first scenario is a strong light scenario, and its corresponding dynamic threshold range is (-5, 5).
[0124] Step 409: Based on the scene information, determine a dynamic threshold from the dynamic threshold range corresponding to the scene in the first row.
[0125] In strong light scenarios, compared to normal behavior scenarios, the main factor affecting the dynamic threshold is light intensity. The dynamic threshold corresponding to the first behavior scenario can be determined based on the relationship between the preset light intensity and the dynamic threshold range. The relationship can be linear or non-linear, and no specific limitation is made here.
[0126] Step 410: Determine the fixation score threshold based on the static threshold corresponding to the normal behavior scenario and the determined dynamic threshold.
[0127] For example, if the static threshold for a normal behavior scenario is 50 and the dynamic threshold for the first behavior scenario is -5, then the gaze score threshold for the first behavior scenario is 45. When the gaze score is greater than 45, the image is judged as gazed and unlocked; when the gaze score is less than or equal to 45, the image is judged as non-gazed and unlocked.
[0128] In another optional implementation, the electronic device includes a first correspondence between a preset behavioral scenario and a preset score threshold. After step 404 and before step 407, the electronic device may further determine the preset score threshold corresponding to the first behavioral scenario based on the first behavioral scenario and the first correspondence, and then use the preset score threshold corresponding to the first behavioral scenario information as the gaze score threshold.
[0129] In this embodiment of the application, by setting the gaze threshold according to the behavioral scenario, the robustness and generalization of the gaze model in certain specific scenarios can be enhanced in a targeted manner.
[0130] When users capture faces using electronic devices, the device's orientation varies, including landscape, portrait, and other orientations such as upside down. Capturing facial images when the device is not in portrait mode can make it difficult to extract facial and eye features, affecting the unlocking result.
[0131] To address this issue, prior to step 403, the electronic device can determine its orientation based on sensor data, determine whether the first image needs to be rotated, and then perform facial feature detection. The electronic device can determine the deflection angle between the direction of the screen's long axis when the device is in its orientation and the direction of the screen's long axis when the device is in portrait mode, and then rotate the first image by that deflection angle.
[0132] The following section describes the device posture of electronic devices.
[0133] In landscape mode, the screen of an electronic device is basically a horizontal bar. In portrait mode, the screen is basically a vertical bar. Specifically, the aspect ratio of the screen differs between landscape and portrait modes. The aspect ratio, also known as the screen's vertical-to-horizontal ratio, is the ratio of the screen's height to its width. In landscape mode, the screen's height is the length of its shorter side, and its width is the length of its longer side. In portrait mode, the screen's height is the length of its longer side, and its width is the length of its shorter side. The longer side of the screen is defined as the two longer, parallel, and equal-length sides of the screen, while the shorter side is defined as the two shorter, parallel, and equal-length sides of the screen.
[0134] For example, electronic devices in such Figure 6 In the landscape mode shown in Figure (A), the height of the display screen is Y, and the width of the display screen is X. Therefore, the aspect ratio of the display screen is Y / X, where Y / X < 1. It should be noted that when the electronic device is in a horizontal position... Figure 6 In landscape mode as shown in (A), tilting or rotating the device by a small angle (e.g., an angle not greater than a first angle threshold, such as 20°, 15°, 5°, etc.) will still maintain the device in landscape mode. For example, in... Figure 6 In landscape mode (A), the clockwise rotation angle is α, causing the electronic device to be in a certain position. Figure 6 In the state shown in (B), when α is not greater than the first angle threshold, the electronic device will... Figure 6 The state shown in (B) is considered the landscape mode. For example, when an electronic device is in... Figure 6In the portrait mode shown in (C), the height of the display screen is X, and the width of the display screen is Y. Therefore, the aspect ratio of the display screen is X / Y, where X / Y > 1. It should also be noted that when the electronic device is in a portrait mode... Figure 6 In the portrait mode shown in (C), tilting or rotating by a small angle (e.g., the angle is no greater than the second angle threshold, such as 20°, 15°, 5°, etc.) will still treat the electronic device as being in portrait mode. For example, in... Figure 6 In the portrait mode shown in (C), the counterclockwise rotation angle is β, which makes the electronic device... Figure 6 In the state shown in (D), when β is not greater than the second angle threshold, the electronic device will... Figure 6 The state shown in (D) is considered the portrait mode. It is understood that the first angle threshold and the second angle threshold can be the same or different, and can be set according to actual needs; there are no restrictions on this.
[0135] For example, when the electronic device is in portrait mode, the captured image of the user's face is as follows: Figure 5 As shown in (A), when the electronic device is in landscape mode, the captured image of the user's face is as follows: Figure 7 As shown in (A), when the electronic device is in the state of Figure 6 In the inverted screen state shown in Figure (E), the captured user face image is as follows: Figure 7 As shown in (B). Figure 7 (A) and Figure 7 The user's face image shown in (B) is difficult to extract facial features from, and it takes a long time.
[0136] For example, electronic devices in such Figure 6 In the portrait mode shown in (C), tilting or rotating by a large angle (e.g., the angle is greater than or equal to the second angle threshold) will result in a landscape mode if the angle is 90 degrees or 270 degrees, and an inverted mode if the angle is 108 degrees.
[0137] In this embodiment of the application, if the electronic device is in a non-portrait mode when capturing the first image, taking the electronic device in a landscape mode as an example, the captured image is as follows: Figure 7 The face image shown in (A) is rotated 90 degrees according to the device's orientation and the screen orientation. The resulting image is then rotated 90 degrees in the same direction as the device's orientation. Figure 5 The face image shown in (A) facilitates the accurate extraction of facial and eye feature information.
[0138] By combining sensor data to rotate the first image captured in a non-portrait state, the problem of electronic devices being unable to unlock when the screen is upside down can be effectively solved.
[0139] The following is a specific example to illustrate the implementation process of the face gaze unlocking method provided in the embodiments of this application.
[0140] like Figure 8 As shown, the face gaze unlocking method includes the following steps:
[0141] Step 801: Acquire the first image.
[0142] Step 802: Obtain sensor data from the electronic device during the acquisition of the first image.
[0143] Step 803: Determine whether the electronic device is upside down based on sensor data (such as attitude sensor). If yes, proceed to step 804; otherwise, keep the original image and proceed to step 805.
[0144] Step 804: Rotate the first image according to the inverted screen orientation of the electronic device.
[0145] Step 805: Perform face detection on the first image and determine whether a face is detected. If yes, proceed to step 806; otherwise, proceed to step 810.
[0146] Step 806: Perform face pose detection on the first image and determine whether face pose is detected. If yes, proceed to steps 807 and 808; otherwise, proceed to step 810.
[0147] Step 807: Extract facial pose features from the first image to obtain facial pose information.
[0148] Step 808: Perform eye detection on the first image and determine whether eyes are detected. If yes, proceed to step 809; otherwise, proceed to step 810.
[0149] Step 809: Extract face and eye features from the first image to obtain face and eye feature information. Then, proceed to step 811.
[0150] Step 810: Enter exception handling; process ends.
[0151] Step 811: Determine scene information based on the first image and the sensor data.
[0152] Scene information includes light intensity, shooting distance, the orientation of the electronic device, and camera temperature.
[0153] Step 812: Based on the scene information and the face pose information, locate the behavior scene corresponding to the first image, and match the image optimizer or optimization parameters corresponding to the behavior scene.
[0154] Step 813: Optimize the face and eye feature information according to the image optimizer or optimization parameters corresponding to the behavior scene to obtain the optimized face and eye feature information.
[0155] Step 814: Input the optimized facial and eye feature information into the gaze model.
[0156] Step 815: Output gaze score.
[0157] Step 816: Determine whether the fixation score is greater than the fixation score threshold. If yes, proceed to step 817; otherwise, proceed to step 818.
[0158] Step 817: Unlock.
[0159] Step 818: Do not unlock.
[0160] It should be understood that the face gaze unlocking method provided in this application embodiment can be applied to a variety of scenarios. For example, it can be used in scenarios where unlocking is required when an electronic device is locked, or in scenarios where an application (such as Alipay or WeChat) or a page of an application with a face gaze unlocking function needs to be unlocked. In short, the face gaze unlocking method provided in this application embodiment can be applied to any scenario that requires face gaze unlocking, and these will not be listed one by one in this document.
[0161] Furthermore, the method of locating behavioral scenes based on sensor data and optimizing input images by matching optimizers or optimization parameters in this embodiment is not limited to optimizing input images for face gaze scene detection; it can also be used for optimizing input data for other target detection (such as vehicles, animals, etc.). The method of setting gaze thresholds based on behavioral scenes in this solution is not limited to gaze models or face gaze classifiers; it can also be used for setting thresholds for most target classifiers (face detection, target recognition).
[0162] The various embodiments of this application can be combined arbitrarily to achieve different technical effects.
[0163] The methods provided in the embodiments of this application above are described from the perspective of an electronic device as the executing entity. To implement the functions of the methods provided in the embodiments of this application above, the electronic device may include hardware structures and / or software modules, implementing the above functions in the form of hardware structures, software modules, or a combination of hardware structures and software modules. Whether a particular function is executed in the form of hardware structures, software modules, or a combination of hardware structures and software modules depends on the specific application and design constraints of the technical solution.
[0164] When implemented in hardware, the hardware implementation of this electronic device can be found in [reference needed]. Figure 9 And its related descriptions.
[0165] See Figure 9 The electronic device 100 includes: a touchscreen 901, wherein the touchscreen 901 includes a touch panel 907 and a display screen 908; one or more processors 902; a memory 903; one or more application programs (not shown); and one or more computer programs 904; a sensor 905; and the aforementioned devices can be connected via one or more communication buses 906. The one or more computer programs 904 are stored in the memory 903 and configured to be executed by the one or more processors 902. The one or more computer programs 904 include instructions that can be used to perform the methods in any of the above embodiments.
[0166] This application also provides a computer-readable storage medium, which may include a memory that stores a program. When the program is executed, it causes an electronic device to perform the aforementioned actions. Figure 3 , Figure 4 , Figure 8 All or part of the steps described in the method embodiments shown.
[0167] This application also provides a computer program product that, when run on an electronic device, causes the electronic device to perform the aforementioned operations. Figure 3 , Figure 4 , Figure 8 All or part of the steps described in the method embodiments shown.
[0168] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component, or module. The apparatus may include a connected processor and a memory; wherein the memory is used to store computer execution instructions, and when the apparatus is running, the processor may execute the computer execution instructions stored in the memory to cause the chip to execute the methods in the above-described method embodiments.
[0169] In this application, the electronic devices, computer storage media, computer program products or chips provided in the embodiments are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0170] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0171] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0172] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0173] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0174] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0175] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A face gaze unlocking method, characterized in that, The method is applied to an electronic device, and the method comprises: collecting a first image and obtaining sensor data in the electronic device when the first image is collected; determining scene information according to the first image and the sensor data; the scene information comprises the following: current light intensity, shooting distance, posture of the electronic device, camera temperature; extracting face, eye feature information and face posture information from the first image; the face posture information comprises at least one of the following face posture types: front face, profile, looking up, looking down, and angle information of each face posture type relative to the front face; locating a first behavior scene corresponding to the first image according to the scene information and the face posture information; when the first behavior scene is an abnormal behavior scene, determining an image optimizer corresponding to the first behavior scene from preset image optimizers corresponding to behavior scenes, and optimizing face and eye feature information by using the image optimizer corresponding to the first behavior scene to obtain optimized face and eye feature information; wherein the abnormal behavior scene is a scene other than a normal behavior scene, and the normal behavior scene is a scene corresponding to normal light, normal distance, normal camera temperature and front face; inputting the optimized face and eye feature information into a gaze model to obtain a gaze score; when it is determined that the gaze score is greater than a gaze score threshold, performing unlocking; before the unlocking is performed when it is determined that the gaze score is greater than the gaze score threshold, the method further comprises: determining a dynamic threshold range corresponding to the first behavior scene from preset dynamic threshold ranges corresponding to behavior scenes, wherein the preset dynamic threshold range corresponding to the behavior scene is obtained by training a threshold model through a test set corresponding to different behavior scenes; determining a dynamic threshold from the dynamic threshold range corresponding to the first behavior scene according to the scene information; determining the gaze score threshold according to a static threshold corresponding to a normal behavior scene and the determined dynamic threshold.
2. The method of claim 1, wherein, Before the face, eye feature information and face posture information are extracted from the first image, the method further comprises: determining a device posture of the electronic device according to the sensor data; determining a deflection angle between a direction of a long axis of a screen when the electronic device is in the device posture and a direction of the long axis of the screen when the electronic device is in a vertical screen state; rotating the first image by the deflection angle.
3. The method according to any one of claims 1 to 2, wherein, The method further comprises: obtaining a gaze image set and a non-gaze image set; the gaze image set comprises gaze images collected under each preset behavior scene, and the non-gaze image set comprises non-gaze images collected under each preset behavior scene; for each gaze image in the gaze image set, optimizing face and eye feature information in the gaze image by using an image optimizer corresponding to a behavior scene when the gaze image is collected to obtain first optimized face and eye feature information; and For each non-gaze image in the set of non-gaze images, the face and eye feature information in the non-gaze image is optimized by an image optimizer corresponding to the behavior scene when the gaze image is collected, to obtain optimized second face and eye feature information; Each of the optimized first face and eye feature information and each of the optimized second face and eye feature information is provided with a behavior scene label; The optimized first face and eye feature information provided with the behavior scene label and the optimized second face and eye feature information provided with the behavior scene label are input into a neural network model for model training to obtain a gaze model.
4. The method according to any one of claims 1 to 2, wherein, The sensor data includes at least one of the following: data collected by a posture sensor in the electronic device, data collected by a distance sensor, data collected by an ambient light sensor, and data collected by a temperature sensor.
5. An electronic device, comprising: Comprise: One or more processors; Memory; And one or more computer programs, wherein the one or more computer programs are stored in the memory, the one or more computer programs include instructions, when the instructions are executed by the electronic device, the electronic device executes the following steps: Collect a first image and obtain sensor data in the electronic device when collecting the first image; According to the first image and the sensor data, determine the scene information; The scene information includes the following: current light intensity, shooting distance, posture of the electronic device, camera temperature; Extract face, eye feature information and face posture information from the first image, the face posture information includes at least one of the following face posture types: front face, side face, head up, head down, and angle information of each face posture type relative to the front face; According to the scene information and the face posture information, locate the first behavior scene corresponding to the first image; When the first behavior scene is an abnormal behavior scene, determine the image optimizer corresponding to the first behavior scene from the preset behavior scene corresponding image optimizer, and optimize the face and eye feature information by the image optimizer corresponding to the first behavior scene to obtain the optimized face and eye feature information; wherein the abnormal behavior scene is a scene other than the normal behavior scene, and the normal behavior scene is a scene corresponding to normal light, normal distance, normal camera temperature and front face; Input the optimized face and eye feature information into the gaze model to obtain a gaze score; When it is determined that the gaze score is greater than a gaze score threshold, perform unlocking; When the instructions are executed by the electronic device, the electronic device performs the following steps before determining that the gaze score is greater than the gaze score threshold and performing unlocking: From the preset behavior scene corresponding dynamic threshold range, determine the dynamic threshold range corresponding to the first behavior scene, and the preset behavior scene corresponding dynamic threshold range is obtained by training a threshold model through different behavior scene corresponding test sets; determine a dynamic threshold from the dynamic threshold range corresponding to the first behavior scene according to the scene information; determine the gaze score threshold according to the static threshold corresponding to the normal behavior scene and the determined dynamic threshold.
6. The electronic device of claim 5, wherein, When the instructions are executed by the electronic device, the electronic device is caused to perform the following steps before extracting the face, eye feature information and face pose information from the first image: determine a device pose of the electronic device according to the sensor data; determine a deflection angle between a direction of a long axis of a screen when the electronic device is in the device pose and a direction of the long axis of the screen when the electronic device is in a portrait screen state; rotate the first image by the deflection angle.
7. The electronic device of any of claims 5-6, wherein, When the instructions are executed by the electronic device, the electronic device is caused to further perform the following steps: obtain a gaze image set and a non-gaze image set, the gaze image set including gaze images collected under each preset behavior scene, and the non-gaze image set including non-gaze images collected under each preset behavior scene; for each gaze image in the gaze image set, use an image optimizer corresponding to a behavior scene when the gaze image is collected to optimize the gaze image, to obtain an optimized gaze image; for each non-gaze image in the non-gaze image set, use an image optimizer corresponding to a behavior scene when the non-gaze image is collected to optimize the non-gaze image, to obtain an optimized non-gaze image; set a behavior scene label for each of the optimized gaze images and each of the optimized non-gaze images; input the optimized gaze images with the set behavior scene labels and the optimized non-gaze images with the set behavior scene labels into a neural network model for model training, to obtain a gaze model.
8. The electronic device of any of claims 5-6, wherein, The sensor data includes at least one of the following: data collected by a posture sensor in the electronic device, data collected by a distance sensor, data collected by an ambient light sensor, and data collected by a temperature sensor.
9. A computer-readable storage medium, characterized in that, The computer instructions, when executed on an electronic device, cause the electronic device to perform the method of any one of claims 1-4.
Citation Information
Patent Citations
Face recognition method and device based on dynamic threshold value
CN103136533A
Image processing method and related products
CN107862265A
Fixation point judgment method and device, electronic device, and compute storage medium
CN109389069A
Face recognition method and device
CN110472504A