Operation control method and electronic equipment

By detecting the user's hand and eye movement cursor in an electronic device, replacing the eye movement cursor with a finger mouse, and performing operations according to gestures, the problem of insufficient eye movement tracking in the prior art is solved, and the device execution accuracy and user experience are improved.

CN120010708APending Publication Date: 2025-05-16HONOR DEVICE CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202311493804.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-09
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing eye tracking technology is difficult to accurately analyze the position of the user's eye gaze point, causing electronic devices to perform incorrect operations and affect the user experience.

Method used

By acquiring image frames, if the hand is detected and there is an eye-moving cursor on the display interface, the eye-moving cursor is replaced with a finger mouse, and the corresponding operation is performed according to the gestures of the user's hand.

Benefits of technology

It realizes hand-eye linkage scenarios, improves the execution accuracy of electronic devices, reduces the situation of wrong operations, simplifies the user operation process, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010708A_ABST
    Figure CN120010708A_ABST
Patent Text Reader

Abstract

The embodiment of the invention is applied to the technical field of terminals, and provides an operation control method and electronic equipment. The electronic device acquires a first image frame. Afterwards, under the condition that the electronic equipment detects that the first image frame comprises a hand, if an eye movement cursor exists in the display interface, the electronic equipment replaces a display finger mouse at a first position displayed by the eye movement cursor, and the first position is used for representing the fixation point position of the eyes of the user on the display interface. Afterwards, the electronic device can control a finger mouse to execute a target operation corresponding to the first preset gesture. According to the application, a hand-eye linkage scene can be realized, and the execution precision of the electronic equipment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of terminal technology, and in particular to an operation control method and an electronic device. Background Art

[0002] With the continuous development of electronic devices, electronic devices are equipped with more and more functions (such as air gestures, eye tracking, etc.) for users to use. At present, eye tracking, as a non-contact human-computer interaction function, can realize human-computer interaction by analyzing the position of the user's eye gaze point.

[0003] However, related eye tracking technology cannot accurately analyze the position of the user's eye gaze point when using an electronic device. Position tracking errors may cause the electronic device to perform incorrect operations, affecting the user's experience. Summary of the invention

[0004] The embodiments of the present application provide an operation control method and an electronic device for realizing a hand-eye linkage scenario and improving the execution accuracy of the electronic device.

[0005] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:

[0006] In a first aspect, an operation control method is provided, in which an electronic device acquires a first image frame, wherein the first image frame is an image frame currently collected by the electronic device. Afterwards, the electronic device may perform hand detection on the first image frame to obtain a hand detection result, wherein the hand detection result is used to indicate whether the first image frame includes a hand. Afterwards, when the hand detection result indicates that the first image frame includes a hand, the electronic device determines whether there is an eye cursor on the display interface. If there is an eye cursor on the display interface, the electronic device may replace the display of a finger mouse at the first position where the eye cursor is displayed, wherein the first position is used to represent the position of the user's eye gaze point on the display interface. Afterwards, the electronic device may control the finger mouse to perform a target operation corresponding to a first preset gesture, wherein the first preset gesture is obtained by performing gesture recognition based on the hand included in the target image frame, and the target image frame includes the first image frame and / or the second image frame, and the second image frame is an image frame collected after the electronic device displays the finger mouse.

[0007] In the present application, after the first image frame captured by the electronic device includes the user's hand, if an eye cursor is displayed on the display interface of the electronic device, the electronic device replaces the eye cursor with a finger mouse, that is, the electronic device switches the eye control operation to a gesture control operation, that is, the electronic device can perform corresponding operations according to the gesture corresponding to the user's hand. In this way, a hand-eye linkage scenario can be realized, which can not only improve the execution accuracy of the electronic device and reduce the occurrence of erroneous operations of the electronic device due to small changes in the user's eyes, but also simplify the user's operation process. As long as the electronic device captures an image frame containing the user's hand, the user does not need to make special gestures, which reduces the user's operation time, improves the execution efficiency of the electronic device, and thus improves the user's experience.

[0008] In a possible implementation manner of the first aspect, the method further includes: when the hand detection result indicates that the first image frame does not include a hand, the electronic device continues to acquire the first image frame.

[0009] In the present application, if the hand detection result indicates that the first image frame does not include a hand, that is, the hand detection result does not include a hand detection box, it means that the user currently does not have the possibility of triggering the finger mouse service. Therefore, the electronic device can continue to capture the first image frame to further detect whether the next image frame contains the user's hand. In this way, real-time detection of the user's hand can be achieved, reducing the occurrence of the finger mouse service not being triggered in time due to the electronic device missing image frames, thereby improving the execution efficiency of the electronic device and thereby improving the user experience.

[0010] In a possible implementation of the first aspect, the method further includes: if the display interface does not have an eye cursor, the electronic device may perform gesture recognition on the hand included in the first image frame to obtain a second target gesture. Then, when the second target gesture is a second preset gesture, the electronic device displays a finger mouse on the display interface according to a preset position. Then, the electronic device controls the finger mouse to perform a target operation corresponding to the second preset gesture.

[0011] In the present application, if there is no eye cursor on the display interface, it means that the electronic device has not received the notification message or the user has not turned on the eye tracking function. Therefore, the electronic device can add a finger mouse at a preset position on the display interface. In this way, the user can be clearly prompted that the finger mouse service has been turned on, thereby avoiding the mobile phone performing erroneous operations due to the user's incorrect triggering of the finger mouse service, thereby improving the accuracy of the operation execution.

[0012] In a possible implementation manner of the first aspect, the method further includes: when the second target gesture is not a second preset gesture, the electronic device continues to acquire the first image frame.

[0013] In the present application, if the second target gesture is not the second preset gesture, it means that the user does not want to perform the gesture control operation. Therefore, the electronic device can continue to capture the first image frame to further detect whether the next image frame contains the user's hand. In this way, real-time detection of the user's hand can be achieved, reducing the occurrence of the inability to trigger the finger mouse service in time due to the electronic device missing the image frame, thereby improving the execution efficiency of the electronic device and thus improving the user experience.

[0014] In a possible implementation of the first aspect, the method further includes: when the electronic device receives the notification message, acquiring a third image frame. Then, the electronic device performs eye movement recognition on the third image frame to obtain eye movement recognition data; wherein the eye movement recognition data includes the coordinates of the gaze point of the user's eyes. Then, the electronic device displays an eye movement cursor on the display interface according to the gaze point coordinates.

[0015] In the present application, after receiving a notification message, the electronic device can display an eye cursor on the display interface according to the gaze point coordinates corresponding to the gaze point of the user's eyes in the third image frame. That is to say, the display interface will display the eye cursor while displaying the message notification to turn on the eye tracking function. In this way, the user can interact with the electronic device without contact, thereby improving the convenience of human-computer interaction.

[0016] In a possible implementation of the first aspect, the process of performing eye movement recognition on the third image frame may specifically include: the electronic device performs face detection on the third image frame according to a face detection model to obtain a face detection result, wherein the face detection result is used to indicate whether the third image frame includes a face. Thereafter, when the face detection result indicates that the third image frame includes a face, the electronic device may perform eye movement recognition on the third image frame according to an eye movement recognition algorithm to obtain eye movement recognition data.

[0017] In the present application, since the detection accuracy of the eye movement recognition algorithm is higher than that of the face detection model, that is, the eye movement recognition algorithm requires more computing power than the face detection model, therefore, in order to reduce unnecessary waste of resources, the electronic device may perform face detection on the third image frame before performing eye movement recognition on the third image frame, thereby avoiding waste of computing resources.

[0018] In a possible implementation manner of the first aspect, the method further includes: when the face detection result indicates that the third image frame does not include a face, the electronic device does not perform eye movement recognition on the third image frame.

[0019] In the present application, if the face detection result indicates that the third image frame does not include a face, it means that the user is not watching the electronic device and the possibility of the user looking at the notification message is small. Therefore, the electronic device does not need to further perform eye movement recognition on the third image frame. In this way, unnecessary waste of resources can be reduced and the utilization rate of computing resources can be improved.

[0020] In a possible implementation of the first aspect, the process of controlling the finger mouse to perform the target operation may specifically include: the electronic device performs gesture recognition on the hand included in the first image frame to obtain a third target gesture. Thereafter, when the third target gesture is the first preset gesture, the electronic device controls the finger mouse to perform the target operation corresponding to the first preset gesture.

[0021] In this application, after replacing the eye cursor with a finger cursor, the electronic device can directly perform gesture recognition on the hand in the first image frame to obtain a third target gesture. If the third target gesture is the first preset gesture, it means that the user wants to perform the target operation corresponding to the third target gesture. Therefore, the electronic device can directly control the finger mouse to perform the target operation corresponding to the first preset gesture. In this way, the execution efficiency of the mobile phone can be improved.

[0022] In a possible implementation of the first aspect, the process of controlling the finger mouse to perform the target operation may specifically include: the electronic device acquires a second image frame. Then, when the electronic device detects that the second image frame includes a hand, the electronic device performs gesture recognition on the hand included in the second image frame to obtain a first target gesture. Then, when the first target gesture is a first preset gesture, the electronic device controls the finger mouse to perform the target operation corresponding to the first preset gesture.

[0023] In the present application, if the gesture included in the second image frame is a first preset gesture, the electronic device can control the finger mouse to perform the target operation corresponding to the first preset gesture. In this way, the operation of eye cursor movement can be realized through eye gaze, and the corresponding operation can be performed through the gesture made by the user's hand, which provides convenience for the user and improves the user's experience. In addition, after determining that the first target gesture is the first preset gesture, the electronic device promptly executes the control operation corresponding to the first target gesture to improve the execution accuracy of the mobile phone.

[0024] In a possible implementation manner of the first aspect, the method further includes: when the second image frame does not include a hand, or when the first target gesture is not a first preset gesture, the electronic device does not perform a gesture control operation.

[0025] In the present application, if the second image frame does not include a hand, or the gesture included in the second image frame is not the first preset gesture, the electronic device may not perform the gesture control operation. In this way, the occurrence of erroneous operations performed by the electronic device can be reduced, the execution accuracy of the electronic device can be improved, and the user experience can be improved.

[0026] In a possible implementation of the first aspect, the method further includes: the electronic device continuously collects a preset number of first image frames. Afterwards, the electronic device may perform hand detection on the first image frames respectively to determine whether each first image frame includes the user's hand. Afterwards, if the preset number of first image frames all include the user's hand, the electronic device determines whether there is an eye cursor on the display interface. Afterwards, if there is an eye cursor on the display interface, the eye cursor is replaced with a finger mouse.

[0027] In the present application, if a preset number of first image frames all include the user's hand, it means that the user wants to trigger the finger mouse service. Therefore, the mobile phone can further determine whether there is an eye cursor on the display interface. In this way, the situation where the mobile phone incorrectly displays the finger mouse due to the hand mistakenly appearing in the shooting range of the front camera can be reduced, thereby improving the display accuracy of the finger mouse.

[0028] In a second aspect, the present application provides an electronic device, comprising a camera, a memory and one or more processors; the camera, the memory and the processor are coupled; the camera is used to capture images, the memory is used to store computer program codes, and the computer program codes include computer instructions; when the processor executes the computer instructions, the electronic device executes the method described above.

[0029] In a third aspect, the present application provides a computer-readable storage medium, comprising computer instructions, which, when executed on an electronic device, enable the electronic device to execute the method described above.

[0030] In a fourth aspect, the present application provides a computer program product, which, when executed on an electronic device, enables the electronic device to execute the method described above.

[0031] In a fifth aspect, a chip is provided, comprising: an input interface, an output interface, a processor and a memory, wherein the input interface, the output interface, the processor and the memory are connected via an internal connection path, and the processor is used to execute the code in the memory. When the code is executed, the processor is used to execute the method as described above.

[0032] Among them, the beneficial effects that can be achieved by the electronic device described in the second aspect, the computer-readable storage medium described in the third aspect, the computer program product described in the fourth aspect, and the chip described in the fifth aspect provided above can refer to the beneficial effects in the first aspect and any possible design method thereof, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 A schematic diagram of an interface showing a pop-up window click process provided in an embodiment of the present application;

[0034] Figure 2 A schematic diagram of a scenario in which a user's eyes are looking at a screen provided in an embodiment of the present application;

[0035] Figure 3 A schematic diagram of a cursor movement interface provided in an embodiment of the present application;

[0036] Figure 4 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;

[0037] Figure 5 A schematic diagram showing a front camera and a camera of a mobile phone provided in an embodiment of the present application;

[0038] Figure 6 A schematic diagram of the software structure of an electronic device provided in an embodiment of the present application;

[0039] Figure 7 A flowchart of an operation control method provided in an embodiment of the present application;

[0040] Figure 8 A schematic diagram of setting a finger mouse function provided in an embodiment of the present application;

[0041] Fig. 9 A schematic diagram of a mobile phone performing hand detection provided in an embodiment of the present application;

[0042] Fig.10 A flowchart of another operation control method provided by an embodiment of the present application;

[0043] Fig.11 A schematic diagram of an interface for a mobile phone to receive a notification message provided in an embodiment of the present application;

[0044] Fig.12 A schematic diagram of setting an eye tracking function provided in an embodiment of the present application;

[0045] Fig.13 A schematic diagram of eye movement recognition performed by a mobile phone provided in an embodiment of the present application;

[0046] Fig.14A schematic diagram of a mobile phone performing face detection provided in an embodiment of the present application;

[0047] Fig.15 A schematic diagram of an interface for displaying an eye-movement cursor on a mobile phone provided in an embodiment of the present application;

[0048] Fig.16 A schematic diagram of an interface for deepening text on a control provided in an embodiment of the present application;

[0049] Fig.17 A schematic diagram of an interface in which an eye cursor is replaced with a finger mouse provided in an embodiment of the present application;

[0050] Fig.18 A schematic diagram of a scenario for triggering a click operation provided in an embodiment of the present application;

[0051] Fig.19 A schematic diagram of a mobile phone entering a reply interface provided in an embodiment of the present application;

[0052] Fig. 20 A flowchart of an operation control process provided by an embodiment of the present application;

[0053] Fig.21 A schematic diagram of an interface of a moving finger mouse provided in an embodiment of the present application. DETAILED DESCRIPTION

[0054] The technical scheme in the embodiment of the present application will be described below in conjunction with the accompanying drawings in the embodiment of the present application. Wherein, in the description of the present application, unless otherwise specified, the "and / or" in the present application is only a kind of association relationship describing the associated object, indicating that there can be three kinds of relationships, for example, A and / or B, which can be represented by: A exists alone, A and B exist at the same time, and B exists alone, wherein A, B can be singular or plural. And, in the description of the present application, unless otherwise specified, "multiple" refers to two or more than two. "At least one of the following (individuals)" or its similar expressions refers to any combination of these items, including any combination of single items (individuals) or plural items (individuals). For example, at least one of a, b, or c (individuals) can be represented by: a, b, c, ab, ac, bc, or abc, wherein a, b, c can be single or multiple. In addition, in order to facilitate the clear description of the technical scheme of the embodiment of the present application, in the embodiment of the present application, the words "first", "second" and the like are used to distinguish the same items or similar items with substantially the same functions and effects. Those skilled in the art will appreciate that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit the difference. At the same time, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way for easy understanding.

[0055] In some embodiments, when a user is using an electronic device, if the electronic device receives a notification message (such as a text message, etc.), the electronic device will usually display the corresponding notification message in the form of a pop-up window at the top of the interface to remind the user to check. If the user wants to view or reply to the notification message, the user can manually click on the notification pop-up window to enter the detailed interface. However, the operation process is relatively complicated, which wastes the user's operation time and makes the user's experience low. In addition, since the display area of ​​the notification pop-up window is limited, if the displayed notification message content is long, the complete message content may not be displayed due to the limited display area, resulting in the user needing to click to view the specific message content, and the operation process is relatively cumbersome.

[0056] For example, Figure 1 As shown, when the mobile phone screen is displaying short video content, if the mobile phone receives a text message notification, the mobile phone can display the notification message in the form of a pop-up window at the top of the screen. It can be understood that Figure 1The T area in the (a) interface is the pop-up window area, which displays the SMS content, contacts, SMS receiving time, reply controls, and mark read controls. Afterwards, if the user clicks any position in the pop-up window area T except the reply control and the mark read control, it means that the user wants to view the notification message. Therefore, the mobile phone can display the detailed interface of the notification message, that is, Figure 1 The (b) interface in the figure is displayed for the user to view the detailed content. If the reply control is clicked by the user, it means that the user wants to reply to the notification message. Therefore, the mobile phone can display the reply interface of the notification message to facilitate the user to reply to the SMS message. If the mark read control is clicked by the user, it means that the user does not want to view the notification message. Therefore, the mobile phone can cancel the display of the pop-up window T so that the mobile phone screen continues to display the short video content.

[0057] In one implementation, in order to simplify the user's operation steps and improve the user's experience, the electronic device can use eye tracking technology to analyze the changes in the user's eyes, so that the user can control the electronic device without touching the screen. It can be understood that when the user's eyes look at different areas or a certain position on the screen of the electronic device, the user's eyes will have subtle changes. Therefore, the electronic device can extract eye features and track the changes in the user's eyes in real time to predict the user's browsing needs. Afterwards, the electronic device can perform corresponding operations based on the user's browsing needs, thereby achieving the purpose of controlling the electronic device through the user's eyes.

[0058] For example, Figure 2 As shown, user Y is browsing the contents of an e-book on the screen of mobile phone S, and the gaze point of user Y is located in the area of ​​position Z. If the user wants to browse the contents of the e-book at the top of the screen, the user can adjust the user's eye sight (such as moving the sight upward) so that the eye gaze point is located in the area of ​​position Z1. Afterwards, the mobile phone can track the changes in the user's eyes to predict the user's browsing needs.

[0059] In some embodiments, when an electronic device receives a notification message, the electronic device can capture the current image frame through the front camera and identify the current image frame to determine the user's eye gaze point. Afterwards, if the electronic device identifies that the user's eye gaze point is within the pop-up window area, it means that the user wants to view the notification message. Therefore, the electronic device can automatically expand the notification message for the user to view, reducing the user's operation process and improving the user's experience.

[0060] For example, Figure 3As shown, when a mobile phone receives a text message, the mobile phone can display the text message content in the form of a pop-up window, that is, the text message content (When are you free tomorrow? Do you want to go shopping...) is displayed in the pop-up window T of the (a) interface, and a cursor is displayed at position Q, where the cursor is used to indicate the current position of the user's eyes. If the cursor moves from position Q to position Q1, it means that the user wants to view the content of the text message. Therefore, the mobile phone can further display the detailed content of the text message, that is, display "When are you free tomorrow? Do you want to go shopping and eat!" in the pop-up window T1 of the (b) interface, so that the user can read the complete text message content, thereby improving the operation efficiency of the mobile phone.

[0061] However, when a user browses the display interface of an electronic device, the changes in the user's eyes are relatively subtle, that is, it is difficult for the electronic device to accurately capture the changes in the user's eyes, that is, the execution accuracy of the electronic device's eye tracking operation is low. If the electronic device cannot accurately detect the changes in the user's eyes, it cannot timely predict the user's operation needs, thereby affecting the user's experience.

[0062] Therefore, in order to improve the execution accuracy of the electronic device, an embodiment of the present application provides an operation control method. In the method, the electronic device acquires a first image frame, wherein the first image frame is the image frame currently captured by the electronic device. Afterwards, when the electronic device detects that the first image frame includes the user's hand, if there is an eye cursor on the display interface, the electronic device replaces the eye cursor with a finger mouse. Afterwards, the electronic device controls the finger mouse to perform the target operation corresponding to the user's hand.

[0063] In an embodiment of the present application, after the image frame captured by the electronic device includes the user's hand, if an eye cursor is displayed on the display interface of the electronic device, the electronic device replaces the eye cursor with a finger mouse, that is, the electronic device switches the eye control operation to a gesture control operation, that is, the electronic device can perform corresponding operations according to the gesture corresponding to the user's hand. In this way, a hand-eye linkage scenario can be realized, which can not only improve the execution accuracy of the electronic device, reduce the occurrence of erroneous operations of the electronic device due to small changes in the user's eyes, but also simplify the user's operation process. As long as the electronic device captures an image frame containing the user's hand, the user does not need to make special gestures, which reduces the user's operation time, improves the execution efficiency of the electronic device, and thus improves the user's experience.

[0064] In some examples, the electronic device in the embodiments of the present application may be a mobile phone, a tablet computer, a smart watch, a desktop, a laptop, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, as well as a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) device, a virtual reality (VR) device, and other devices containing a camera. The embodiments of the present application do not impose any special restrictions on the specific form of the electronic device.

[0065] For example, Figure 4 FIG. 2 shows a schematic diagram of the structure of the electronic device 200. Figure 4 As shown, the electronic device 200 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 211, a power management module 212, a battery 213, an antenna 1, an antenna 2, a mobile communication module 240, a wireless communication module 250, an audio module 270, a sensor module 280, a button 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.

[0066] It is to be understood that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0067] The processor 210 may include one or more processing units, for example, the processor 210 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0068] The controller may be the nerve center and command center of the electronic device 200. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0069] The processor 210 may also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. The memory may store instructions or data that the processor 210 has just used or cyclically used. If the processor 210 needs to use the instruction or data again, it may be directly called from the memory. This avoids repeated access, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.

[0070] In some embodiments, the processor 210 may include one or more interfaces. The interface may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0071] It is understandable that the interface connection relationship between the modules illustrated in the embodiment of the present invention is only a schematic illustration and does not constitute a structural limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0072] The charging management module 211 is used to receive charging input from a charger. While the charging management module 211 is charging the battery 213 , it can also power the electronic device through the power management module 212 .

[0073] The wireless communication function of the electronic device 200 can be implemented through the antenna 1, the antenna 2, the mobile communication module 240, the wireless communication module 250, the modem processor and the baseband processor.

[0074] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 200 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve the utilization of the antennas. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0075] The mobile communication module 240 may provide solutions for wireless communications including 2G / 3G / 4G / 5G etc. applied to the electronic device 200. The modem processor may include a modulator and a demodulator.

[0076] The wireless communication module 250 can provide wireless communication solutions for application in the electronic device 200, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication technology (NFC), infrared technology (IR), etc.

[0077] The electronic device 200 implements the display function through a GPU, a display screen 294, and an application processor. The GPU is a microprocessor for image processing, which connects the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or change display information.

[0078] The display screen (or screen) 294 is used to display images, videos, etc. The display screen 294 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Miniled, MicroLed, Micro-oLed, a quantum dot light emitting diode (QLED), etc. In some embodiments, the electronic device 200 may include 1 or N display screens 294, where N is a positive integer greater than 1.

[0079] The electronic device 200 can realize the shooting function through ISP, camera 293, video codec, GPU, display screen 294 and application processor.

[0080] ISP is used to process the data fed back by camera 293. For example, when an electronic device takes a photo, the shutter is opened, and light is transmitted to the camera photosensitive element (or image sensor) through the lens. The light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to ISP for processing and converts it into an image visible to the naked eye. ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. ISP can also optimize the exposure, color temperature and other parameters of the shooting scene. In some embodiments, ISP can be set in camera 293. In some embodiments, camera 293 includes a shutter. The shutter is a device in the camera used to control the time that light irradiates the photosensitive element.

[0081] The camera 293 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then passes the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 200 may include 1 or N cameras 293, where N is a positive integer greater than 1.

[0082] In some embodiments, the camera 293 may include a lens, which is an optical component for generating an image.

[0083] Exemplarily, the N cameras 293 may include: one or more front cameras and one or more rear cameras. Figure 5 , taking the above-mentioned electronic device 200 as a mobile phone as an example. Figure 5 The (a) interface shows a front camera, such as front camera 20. Figure 5 The interface (b) in FIG. 1 shows three rear cameras, such as rear cameras 21, 22 and 23. Of course, the number of cameras in the mobile phone includes but is not limited to the number described in the above embodiment.

[0084] Among them, the above-mentioned N cameras 293 may include one or more of the following cameras: a main camera, a telephoto camera, a wide-angle camera, an ultra-wide-angle camera, a macro camera, a fisheye camera, an infrared camera, a depth camera and a black and white camera.

[0085] In this embodiment, the front camera is a camera with an always-on camera (AON) function. Specifically, when the electronic device uses the AON function, the front camera of the electronic device is in an always-on state, and can collect images in real time. The electronic device performs gesture recognition through image analysis, and can respond to user gestures to control the screen, so that the screen of the electronic device can be controlled without the user touching the electronic device.

[0086] The digital signal processor is used to process digital signals, and can process not only digital image signals but also other digital signals. For example, when the electronic device 200 is selecting a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.

[0087] The external memory interface 220 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 200.

[0088] The internal memory 221 can be used to store computer executable program codes, which include instructions. The processor 210 executes various functional applications and data processing of the electronic device 200 by running the instructions stored in the internal memory 221. The internal memory 221 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the electronic device 200 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 221 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0089] The electronic device 200 can implement audio functions through the audio module 270 and the application processor, such as music playing and recording, etc. The audio module 270 may include a speaker, a receiver, a microphone, and an earphone interface, etc.

[0090] The buttons 290 include a power button, a volume button, etc. The indicator 292 may be an indicator light.

[0091] The sensor module 280 may include a pressure sensor, a gyro sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, and the like.

[0092] The gyro sensor may be used to determine the motion posture of the electronic device 200. In some embodiments, the angular velocity of the electronic device 200 around three axes (ie, x, y, and z axes) may be determined by the gyro sensor.

[0093] The acceleration sensor can detect the magnitude of the acceleration of the electronic device 200 in all directions (generally three axes). When the electronic device 200 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers.

[0094] The software system of the electronic device 200 may adopt a layered architecture, an event-driven architecture, a micro-core architecture, a micro-service architecture, or a cloud architecture. The present application embodiment takes the Android system of the layered architecture as an example to exemplify the software structure of the electronic device 200.

[0095] Figure 6It is a software structure block diagram of the electronic device 200 of the embodiment of the present application. The embodiments of the present application will be discussed based on the following technical architecture. In order to facilitate the explanation of logic, only the business logic relationship is illustrated by a schematic block diagram, and the specific location of the technical architecture where each business is located is not strictly expressed. In addition, the naming of each module in the software architecture diagram is an exemplary example. The embodiment of the present application does not limit the naming of each module in the software architecture diagram. In actual implementation, the specific naming of the module can be determined according to actual needs.

[0096] The layered architecture divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system includes four layers, from top to bottom: application layer (applications), application framework layer (application framework), hardware abstraction layer (HAL), and kernel layer (kernel).

[0097] Among them, the application layer may include a series of application packages. For example, the application layer may include applications such as cameras and smart perception (applications may be referred to as applications for short), and the embodiments of the present application do not impose any restrictions on this. For another example, the application layer may also include a system user interface (UI), wherein the system user interface may provide a basic display interface for the electronic device, such as a status bar at the top of the screen, a navigation bar at the bottom of the screen, a quick setting bar of the drop-down interface, a notification bar, a lock screen interface, a volume adjustment dialog box, and a screenshot display interface.

[0098] The smart perception application provided in the embodiment of the present application supports various services. For example, the services supported by the smart perception application may include finger mouse services and eye tracking services, and may also include smart code scanning services, staring without turning off the screen services, smart screen off display services, etc. These services can be collectively referred to as smart perception services.

[0099] Optionally, the implementation of these services supported by the smart perception application depends on the AON camera (front camera) of the electronic device being in a normally open state, collecting images in real time, and obtaining data related to the smart perception service. When the smart perception application monitors the service-related data, the smart perception application sends the service-related data to the smart perception algorithm platform, which analyzes the collected images and determines whether to execute the relevant services of the smart perception application or which specific service to execute based on the analysis results.

[0100] Among them, the above-mentioned air gesture service refers to a service in which an electronic device recognizes user gestures and responds according to the preset strategy corresponding to the gesture to achieve human-computer interaction. It can be understood that the user gesture is an action made when the user's hand is within a preset distance of the electronic device. For example, the preset distance can be 20 centimeters. When the electronic device is in the screen-on state, the electronic device can support the recognition of various preset air gestures, and different gestures can correspond to different preset strategies, for example: air OK gesture → finger mouse, air up / down gesture → page turning.

[0101] In one example, taking the air OK gesture as an example, the user changes from an extended palm state to an OK gesture, or from a clenched fist state to an OK gesture, and the control method corresponding to the gesture is preset to display a finger mouse on the display interface of the electronic device. When the electronic device is in the bright screen state, the electronic device will automatically display the finger mouse when detecting the air OK gesture through the AON camera (front camera), and perform corresponding operations based on the user's subsequent gestures. In this way, human-computer interaction can be completed without the user touching the electronic device. Among them, the air OK gesture can also be called the air finger mouse gesture. It can be understood that in actual implementation, the air finger mouse gesture can also be other gestures, such as a V gesture, a pinch gesture, and the like.

[0102] In another example, taking the air up / down gesture as an example, the fingers change from a close-together and spread-out state to a downwardly bent state (upward sliding gesture), and the control mode corresponding to the gesture is preset to turn the page up or slide the screen upward, and the fingers change from a close-together and bent state to an upwardly spread-out state (downward sliding gesture), and the control mode corresponding to the gesture is preset to turn the page down or slide the screen downward. When the electronic device is in the screen-on state, the electronic device will automatically perform the operation of turning the page up when detecting an upward sliding gesture through the AON camera (front camera), and will automatically perform the operation of turning the page down when detecting a downward sliding gesture, and the human-computer interaction can be completed without the user touching the electronic device. Among them, the air up / down gesture can also be referred to as an air sliding screen gesture. Taking the upward or downward sliding gesture as an example for explanation, it can be understood that in actual implementation, the air sliding screen gesture can also be a left or right sliding gesture.

[0103] In addition, in an embodiment of the present application, the air gesture service also supports user-defined operations, for example, the user cancels the "air press gesture → answer the call" and resets it to "air press gesture → jump to the payment code interface".

[0104] Exemplarily, when the electronic device is in the screen-on state and the electronic device displays the main desktop, the user can trigger a predefined quick service through an air press gesture, such as quickly jumping from the main desktop to the payment code interface. In this scenario, the payment interface can be quickly called up through an air press gesture, so this scenario can be called a smart payment scenario. Exemplarily, the user can also reset it to "air press gesture → jump to the scan interface" or "air press gesture → jump to the ride code interface" and so on.

[0105] Among them, the above-mentioned eye tracking service refers to the service that an electronic device identifies changes in the user's eyes, determines the user's gaze point according to the eye changes, and responds according to a preset strategy based on the position of the gaze point and the gaze duration to achieve human-computer interaction. Exemplarily, when an electronic device is in a bright screen state, if a notification message is received, the electronic device can display an eye cursor on the current interface based on the user's gaze point. Afterwards, if the eye cursor is located in the area where the notification message is located and the gaze duration reaches a preset duration, the electronic device can display the complete content of the notification message.

[0106] The application framework layer provides application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions. Figure 6 As shown, the application framework layer may include AON service, camera service, intelligent perception service, input framework, etc.

[0107] Among them, the AON service is used to collect images in real time. The camera service is used to collect images. Specifically, when the camera application in the electronic device is turned on, the camera of the electronic device can perform corresponding operations according to the user's touch operation. For example, if the shooting control in the camera application is clicked by the user, the camera of the electronic device can perform a shooting operation to collect images.

[0108] The Smart Perception Service is used to process the eye movement recognition and tracking requests sent by the application, and return the processing results to the application to implement the eye movement recognition and tracking functions. At the same time, the Smart Perception Service is also used to process the gesture recognition sent by the application, and return the processing results to the application to implement the gesture recognition function.

[0109] The input framework is a GUI toolkit designed for Java, which can include graphical user interface components such as text boxes, buttons, split panes and tables. It can be responsible for registering the smart perception service fence, calling the smart perception service when there is a notification message, and performing corresponding processing when the registration result is returned.

[0110] The hardware abstraction layer is an encapsulation of the Linux kernel driver, providing an interface to the upper layer, hiding the hardware interface details of a specific platform and providing a virtual hardware platform for the operating system. In the embodiment of the present application, the hardware abstraction layer includes the camera HAL, the intelligent perception HAL, and the Android interface definition language (AIDL).

[0111] Among them, the camera HAL is the core software framework of the Camera, which can include sensor nodes and image front end (IFE) nodes. The sensor node and IFE node are components (nodes) in the image data and control instruction transmission path (also called transmission pipeline) created by the camera HAL.

[0112] Smart Perception HAL is the core software application of Smart Perception. Among them, the Smart Perception Control Module includes the Smart Perception Client Application (CA), and the Smart Perception CA runs in the REE environment.

[0113] The kernel layer is the layer between hardware and software. The kernel layer at least includes Qualcomm communication interface (QMI), air mouse detection module, eye tracking module, display driver, camera driver and sensor driver.

[0114] The Qualcomm communication interface is a multi-processor inter-process communication functional interface provided by Qualcomm, which is used to call functions, read data, etc. The camera driver is the driver layer of the Camera device, which is mainly responsible for the interaction with the camera hardware.

[0115] Understandably, Figure 6 The layers in the illustrated structure and the components contained in each layer do not constitute a specific limitation on the electronic device 200. In other embodiments of the present application, the structure may include more or fewer layers than shown, and each layer may include more or fewer components, which is not limited in the present application.

[0116] Based on the electronic device described above, an embodiment of the present application provides an operation control method. The method can be applied to intelligent perception scenarios (such as hand-eye linkage scenarios). The following takes the electronic device as a mobile phone and the application scenario as a hand-eye linkage scenario as an example to illustrate the method of the embodiment of the present application. Specifically, Figure 7 As shown, the operation control method may include S701 to S710.

[0117] S701, the mobile phone captures a first image frame.

[0118] The first image frame is the image frame currently captured by the mobile phone, that is, the image frame within the field of view that can be captured by the front camera of the mobile phone. The front camera is a camera with AON function, that is, the front camera can capture the first image frame in real time.

[0119] Exemplarily, the front camera of the mobile phone can collect the first image frame frame by frame, or can collect images according to a preset interval frame number, for example, the mobile phone can collect a first image frame every 2 frames, or can collect images according to a preset time interval, for example, the mobile phone can collect a first image frame every 2ms, etc. It can be understood that the preset interval frame number and time interval can be set according to actual needs, and are not specifically limited.

[0120] In some embodiments, the first image frame can be used to identify the user's hand so that the mobile phone can perform gesture control operations. In order to protect the user's privacy, it is necessary to determine whether the user has turned on the finger mouse function (or finger mouse service) of the smart perception application before the mobile phone captures the first image frame. When the finger mouse function in the smart perception application is turned on, that is, when the mobile phone receives the user's operation to turn on the finger mouse function in the smart perception application, the mobile phone can capture the first image frame to further perform hand detection on the first image frame.

[0121] For example, Figure 8 As shown, when the setting control in the initial interface is clicked by the user, the mobile phone can display the setting interface, wherein the setting interface includes setting items such as WLAN, Bluetooth, mobile network, desktop and wallpaper, smart assistant and smart perception. Afterwards, when the smart perception setting item in the setting interface is clicked by the user, the mobile phone can display the smart perception setting interface. Among them, the smart perception setting interface includes a smart gaze setting bar, an air gesture setting bar and other setting bars. The smart gaze setting bar includes staring at the screen without turning off the screen, staring at the screen to reduce the volume and eye tracking. The air gesture setting bar includes air sliding the screen, air screenshot, air pressing and finger mouse. Other setting bars include smart screen off display and smart horizontal and vertical screen.

[0122] After that, if any perception function in the smart perception setting interface is clicked, the phone can control the clicked perception function to be turned on or off. For example, see Figure 8 , if the "Finger Mouse" function in the air gesture setting bar is clicked by the user, it means that the user wants to turn on the finger mouse function, so the mobile phone can turn on the finger mouse function. For another example, please refer to Figure 8If the "Air Screenshot" function in the air gesture setting bar is clicked by the user, it means that the user wants to turn off the air screenshot function. Therefore, the mobile phone can turn off the air screenshot function.

[0123] It is understandable that the above Figure 8 The interface for setting the perception function of the smart perception application is only an example, and the user can also control the perception function to be turned on or off in other ways. For example, the mobile phone can directly display the air gesture option so that the user can agree to control the turning on or off of the air gesture function. That is to say, if the air gesture option is clicked by the user, all functions included in the air gesture option are turned on or off.

[0124] In some embodiments, the personal information used in the technical solution of the present application is limited to information for which the individual's separate consent has been obtained, including but not limited to notifying and reminding the user to read the relevant user agreement (notification) and sign the agreement (authorization) including authorization of relevant user information before the user uses the function.

[0125] S702: When the mobile phone detects that the first image frame does not include the user's hand, the mobile phone continues to capture the first image frame.

[0126] Specifically, after acquiring the first image frame, the mobile phone can perform hand detection on the first image frame to obtain a hand detection result. The hand detection result is used to indicate whether the first image frame contains the user's hand. It can be understood that when the hand detection result includes a hand detection frame, it means that the first image frame includes a hand, that is, the user wants to trigger the finger mouse service. Therefore, the mobile phone can further determine whether there is an eye cursor on the display interface. When the hand detection result does not include a hand detection frame, it means that the first image frame does not include a hand, that is, the user currently does not have the possibility of triggering the finger mouse service. Therefore, the mobile phone can continue to acquire the first image frame to further detect whether the first image frame contains the user's hand.

[0127] In some embodiments, the mobile phone can capture multiple consecutive first image frames. Afterwards, the mobile phone can perform hand detection on the first image frames respectively to determine whether each first image frame includes the user's hand. Afterwards, if there is at least one first image frame among the multiple first image frames that does not include the user's hand, or the acquisition time corresponding to the multiple first image frames is less than the first preset time, the mobile phone can continue to capture the first image frames. In this way, the situation where the mobile phone incorrectly displays the finger mouse due to the hand mistakenly appearing in the shooting field of view can be reduced, thereby improving the display accuracy of the finger mouse.

[0128] The hand detection result may also include the confidence of the hand detection frame, that is, if the first image frame includes a hand, the hand detection result may include the hand detection frame and the confidence of the hand detection frame. The confidence of the hand detection frame is used to indicate the confidence level of the hand detection frame. The higher the confidence of the hand detection frame, the more accurate the hand detection result is, that is, the more precise the hand detection frame is.

[0129] Specifically, the mobile phone can perform hand detection on the first image frame according to the hand detection model to obtain the hand detection result, that is, the mobile phone can input the first image frame into the hand detection model to obtain the hand detection result. The hand detection model can take the first image sample as the input of the hand detection model to be trained, output the detection frame of the hand in the first image sample through learning and prediction of the hand detection model to be trained, and adjust the parameters of the hand detection model to be trained based on the output hand detection frame and the real annotation frame of the hand in the first image sample until a trained hand detection model is obtained.

[0130] For example, Fig. 9 As shown, after the front camera 32 of the mobile phone captures the image, the mobile phone can input the image into the trained hand detection model for hand detection to obtain the hand detection result. It can be seen that the image includes the user's hand (such as the gesture of the palm being spread out). Therefore, after the mobile phone inputs the image with the user's hand into the hand detection model, the hand detection result obtained includes the hand detection frame T and the confidence level 0.87 corresponding to the hand detection frame T. Among them, the hand detection frame is used to indicate the position information of the user's hand. For example, the mobile phone can mark the user's hand in the target image frame through the detection frame, and the position information of the user's hand is represented according to the coordinates corresponding to the four vertex corners of the detection frame.

[0131] It is understandable that the hand detection model in the embodiment of the present application is not a hand detection model for a specific user, that is, the hand detection model cannot reflect the personal information of a specific user.

[0132] In some embodiments, when the hand detection result indicates that the first image frame does not include a hand, it means that there is no possibility for the user to trigger the finger mouse service. Therefore, the mobile phone can continue to capture the first image frame to further detect whether the first image frame contains the user's hand.

[0133] S703: When the mobile phone detects that the first image frame includes the user's hand, the mobile phone determines whether there is an eye movement cursor on the display interface.

[0134] The above display interface is the interface displayed by the mobile phone when the mobile phone collects the first image frame. The eye cursor is used to indicate the gaze point position (or gaze point coordinates) of the user's eyes on the display interface.

[0135] In some embodiments, when the hand detection result indicates that the first image frame includes a hand, it means that the user has triggered the air mouse function. Therefore, the mobile phone can further determine whether there is an eye cursor on the display interface. If there is an eye cursor on the display interface, it means that the mobile phone is performing eye tracking service. Therefore, the mobile phone can execute S704. If there is no eye cursor on the display interface, it means that the mobile phone is not currently performing eye tracking service. Therefore, the mobile phone can execute S709 to further perform gesture recognition on the user's hand included in the first image frame. Among them, the air mouse function is an air mouse function, which refers to the function of controlling the movement of icons (such as mouse, cursor, etc.) on the display interface through the air by moving the user's hand or control device in three-dimensional space.

[0136] It can be understood that the user's hand included in the above-mentioned first image frame is the front of the hand, that is, the mobile phone performs gesture recognition based on the feature points of the front of the hand.

[0137] In some embodiments, the control device may be a mouse connected to a mobile phone (such as a wireless mouse), or a device that can instruct a finger mouse to perform control operations (such as a cursor pen), etc., without specific limitation. For example, if the control object of the finger mouse is a wireless mouse, the user needs to connect the wireless mouse to the mobile phone via Bluetooth before triggering the finger mouse function. Afterwards, when the wireless mouse is successfully connected to the mobile phone, the user can wake up the wireless mouse to work by moving the wireless mouse. Afterwards, the mobile phone can perform the target operation corresponding to the operation instruction according to the operation instruction of the wireless mouse.

[0138] In one implementation, Fig.10 As shown, the process of the mobile phone displaying the eye movement cursor may specifically include S711 to S714:

[0139] S711, the mobile phone receives a notification message.

[0140] The notification message mentioned above refers to a notification message sent by any application in the mobile phone. For example, the notification message may be a text message notification or a video recommendation notification of a video application.

[0141] It can be understood that after receiving the notification message, the mobile phone can display the notification message in the message notification bar based on the display interface. For example, the mobile phone can display the notification message in the form of a pop-up window at the top of the current interface. Among them, the display interface is the interface displayed before the mobile phone receives the notification message. The message notification bar can also display at least one of the application name, application icon, reception time of the notification message and operation controls corresponding to the notification message. Exemplarily, the operation control can be a "reply" control, or a "mark read" control, etc., without specific limitation.

[0142] For example, Fig.11 As shown, the mobile phone displays the main interface, which includes icons of multiple applications. If the mobile phone receives a notification message, the mobile phone can display the notification message in the form of a pop-up window at the top of the main interface, that is, the SMS notification is displayed in pop-up window A. It can be understood that pop-up window A is equivalent to a message notification bar, which displays the application name (SMS), receiving time (just now), the sender of the SMS notification (sister), the SMS content (When are you free tomorrow? Do you want to go shopping...), "Reply" control and "Mark Read" control.

[0143] S712, the mobile phone captures a third image frame, wherein the third image frame is an image frame captured when the mobile phone receives a notification message.

[0144] Specifically, after receiving the notification message, the mobile phone can trigger the front camera to capture the third image frame. The third image frame is the image frame captured when the mobile phone receives the notification message, that is, the image frame within the field of view that the front camera of the mobile phone can capture.

[0145] In some embodiments, the third image frame can be used to identify the user's face so that the mobile phone can perform eye tracking operations. In order to protect the user's privacy, it is necessary to determine whether the user has turned on the eye tracking function (or eye tracking service) of the smart perception application before the mobile phone captures the third image frame. When the eye tracking function in the smart perception application is turned on, that is, when the mobile phone receives the user's operation to turn on the eye tracking function in the smart perception application, the mobile phone can capture the third image frame to further perform face detection on the third image frame.

[0146] For example, Fig.12 As shown in the figure, if any perception function in the smart perception setting interface is clicked, the mobile phone can control the clicked perception function to be turned on or off. For example, if the eye tracking function in the smart gaze setting bar is clicked by the user, it means that the user wants to turn on the eye tracking function. Therefore, the mobile phone can display the details interface of the eye tracking function so that the user can further select the sub-perception function that he wants to turn on.

[0147] Among them, the detail interface of the above eye tracking function includes sub-perception functions (gaze to expand message notifications, gaze to enter details), function description (stand 20-50 cm away from the screen, with your eyes facing the screen, gaze to expand the message notification; pause for a while to enter the details) and function prompts (make sure your eyes and face are not blocked. When it recognizes that your line of sight is focused on the message notification, the corresponding operation will be performed for you. This function is not currently supported in the horizontal screen.).

[0148] Afterwards, if any sub-sensing function in the detailed interface of the above eye tracking function is clicked, the mobile phone can control the clicked sub-sensing function to be turned on or off. For example, see Fig.12 If the user clicks the "Look at the message notification to enter details" function, it means that the user wants to enable the "Look at the message notification to enter details" function. Therefore, the mobile phone can enable the "Look at the message notification to enter details" function. After that, when the mobile phone receives a message notification, the mobile phone can determine whether it is possible to enter the details interface based on the user's gaze point. For another example, please refer to Fig.12 If the gaze-expand message notification is clicked by the user, it means that the user wants to turn off the gaze-expand message notification function, so the mobile phone can turn off the gaze-expand message notification function.

[0149] It is understandable that the above Fig.12 The interface for setting the perception function of the smart perception application is only an example, and the user can also control the perception function to be turned on or off in other ways. For example, the mobile phone can directly display the eye tracking option so that the user can agree to control the eye tracking function to be turned on or off. That is to say, if the eye tracking option is clicked by the user, all sub-perception functions included in the eye tracking option will be turned on or off.

[0150] S713, the mobile phone performs eye movement recognition on the third image frame to obtain eye movement recognition data, wherein the eye movement recognition data includes the gaze point coordinates of the user's eyes.

[0151] In some embodiments, after the mobile phone captures the third image frame, it can perform eye movement recognition on the third image frame to obtain the gaze point coordinates of the user's eyes. The gaze point coordinates are used to indicate the coordinate position of the gaze point of the user's eyes on the display interface. In this embodiment, the mobile phone uses the upper left corner vertex of the display interface as the coordinate origin to calculate the gaze point coordinates. In other embodiments, the mobile phone can also use other vertices of the display interface (such as the upper right corner vertex) or the center point of the display interface as the coordinate origin to calculate the gaze point coordinates.

[0152] Specifically, the mobile phone can perform eye movement recognition on the third image frame according to the eye movement recognition algorithm to obtain eye movement recognition data, that is, the mobile phone can input the third image frame into the eye movement recognition algorithm to obtain eye movement recognition data. The eye movement recognition algorithm is a biometric recognition technology that performs eye movement recognition based on the user's eye feature information.

[0153] For example, Fig.13 As shown, after the front camera 32 of the mobile phone captures the third image frame, the mobile phone can input the third image frame into the eye movement recognition algorithm for eye movement recognition to obtain eye movement recognition data. It can be seen that the third image frame includes a face, that is, the third image frame includes the user's eyes. Therefore, after the mobile phone inputs the third image frame into the eye movement recognition algorithm, the obtained eye movement recognition data includes the gaze point coordinates (x1, y1) of the user's eyes.

[0154] The eye movement recognition data may also include the user's eye gaze duration, which refers to the length of time the user's eyes stay on the gaze point.

[0155] In some embodiments, the mobile phone can continuously capture multiple third image frames. Afterwards, the mobile phone can perform eye movement recognition on the third image frames respectively to obtain the coordinates of the gaze point of the user's eyes in each third image frame. Afterwards, when it is determined that the coordinates of the gaze point of the user's eyes in each third image frame are the same, and the acquisition time corresponding to the multiple third image frames (or called the gaze time of the user's eyes) is greater than the second preset time, the mobile phone can display the eye movement cursor on the display interface according to the coordinates of the gaze point of the user's eyes. In this way, the occurrence of the mobile phone incorrectly displaying the eye movement cursor due to the user's mistaken gaze on the screen can be reduced, thereby improving the display accuracy of the eye movement cursor.

[0156] In one implementation, after the mobile phone captures the third image frame, it can perform face detection on the third image frame to obtain a face detection result. The face detection result is used to indicate whether the third image frame includes a face. It can be understood that when the face detection result includes a face detection frame, it means that the third image frame includes a face, that is, there is a possibility that the user is looking at the notification message, so the mobile phone can perform eye movement recognition on the third image frame to obtain eye movement recognition data. When the face detection result does not include a face detection frame, it means that the third image frame does not include a face, that is, the user is not watching the electronic device, and the possibility of the user looking at the notification message is small. Therefore, the mobile phone does not need to further perform eye movement recognition on the third image frame, so that unnecessary waste of resources can be reduced and the utilization rate of computing resources can be improved.

[0157] In some embodiments, the above-mentioned face detection result may also include the confidence of the face detection frame, that is, if the above-mentioned third image frame includes a face, the face detection result may include the face detection frame and the confidence of the face detection frame. Among them, the confidence of the face detection frame refers to a confidence score of the output result. For example, if the third image frame captured by the mobile phone is not clear, it will cause the detection and output results of the mobile phone to be inaccurate. At this time, the confidence score generated when outputting the face detection frame will be low. In other words, the confidence of the face detection frame is used to indicate the confidence level of the face detection frame. The higher the confidence of the face detection frame, the more accurate the face detection result, that is, the more precise the face detection frame. Exemplarily, the mobile phone can output the confidence in the form of a numerical value, for example, the confidence is 0.88.

[0158] Specifically, the mobile phone can perform face detection on the third image frame according to the face detection model to obtain the face detection result, that is, the mobile phone can input the third image frame into the face detection model to obtain the face detection result. The face detection model can be a model that uses the second image sample as the input of the face detection model to be trained, outputs the detection frame of the face in the second image sample through learning and prediction of the face detection model to be trained, and adjusts the parameters of the face detection model to be trained based on the output face detection frame and the real annotation frame of the face in the second image sample until a trained face detection model is obtained.

[0159] The second image sample may be the same as or different from the first image sample. If the second image sample is the same as the first image sample, the second image sample and the first image sample may contain both a face and a hand.

[0160] For example, Fig.14 As shown, after the front camera 42 of the mobile phone captures the image 43, the mobile phone can input the image 43 into the trained face detection model for face detection to obtain the face detection result. It can be seen that the image 43 includes a face. Therefore, after the mobile phone inputs the image 43 with the face into the face detection model, the face detection result obtained includes a face detection frame Y and a confidence level of 0.97 corresponding to the face detection frame Y. Among them, the face detection frame is used to indicate the location information of the user's face. For example, the mobile phone can mark the face in the third image frame through the face detection frame, and the coordinates corresponding to the four top corners of the face detection frame represent the location information of the face.

[0161] It is understandable that the face detection model in the embodiment of the present application is not a face detection model for a specific user, that is, the face detection model cannot reflect the personal information of a specific user.

[0162] S714, the mobile phone displays an eye movement cursor on the display interface according to the gaze point coordinates.

[0163] Specifically, after obtaining the gaze point coordinates of the user's eyes, the mobile phone can display an eye movement cursor at the gaze point coordinates on the display interface.

[0164] For example, Fig.15 As shown, after the mobile phone receives the SMS notification message, it can display the message on the display interface (such as Fig.15 On the home page shown in FIG. 1 , the text message notification is displayed in a pop-up window A. Afterwards, the mobile phone can display an eye movement cursor T according to the gaze point coordinates of the user's eyes on the interface displaying the text message notification.

[0165] It is understandable that the above Fig.15 The shape and color of the eye movement cursor shown are only examples, and the mobile phone can also set an eye movement cursor of other shapes and colors to be displayed on the display interface. For example, the mobile phone can display a triangular eye movement cursor on the display interface. For another example, the mobile phone can display a light gray eye movement cursor on the display interface.

[0166] In some embodiments, after determining the above-mentioned gaze point coordinates, the mobile phone may not display the eye cursor, that is, the mobile phone does not need to replace the eye cursor with a finger mouse, that is, after detecting that the above-mentioned first image frame includes the user's hand and determining the above-mentioned gaze point coordinates, the mobile phone can directly execute S705 to further perform gesture recognition on the user's hand in the second image frame. In this way, the situation where the user's visual experience is affected by the movement of the eye cursor can be reduced, thereby improving the user's usage experience.

[0167] For example, if the above-mentioned gaze point coordinates are located at any control in the message notification bar, it means that the user has the possibility of clicking on the control. Therefore, the mobile phone can perform text deepening processing on the control to facilitate the user to determine whether the control is the control that the user wants to click. For example, see Fig.16 After the mobile phone receives the SMS notification message, the mobile phone can display the SMS notification in the message notification bar B on the main interface. The message notification bar displays the application name (SMS), the receiving time (just now), the sender of the SMS notification (sister), the SMS content (when are you free tomorrow? Do you want to go shopping...), the "Reply" control, and the "Mark Read" control. Afterwards, if the gaze point coordinates are located at the "Mark Read" control in the message notification bar, it means that the user has the possibility of clicking the control, so the mobile phone can bold the "Mark Read" text.

[0168] In some embodiments, the execution order of the above-mentioned process of displaying a notification message and the process of displaying an eye movement cursor is not limited. For example, the mobile phone can first execute the above-mentioned process of displaying a notification message, and then execute the process of displaying an eye movement cursor, or the mobile phone can simultaneously execute the above-mentioned process of displaying a notification message and the process of displaying an eye movement cursor, without specific limitation.

[0169] It can be understood that the above process of displaying the eye movement cursor, that is, the above process S711 to S714 can be performed before step S701. Specifically, when the mobile phone receives the notification message, it displays the eye movement cursor on the display interface based on the coordinates of the user's eye gaze point. Afterwards, the mobile phone can collect the first image frame to further determine whether the eye movement cursor needs to be replaced based on whether the first image frame contains the user's hand.

[0170] S704, the mobile phone replaces the eye-movement cursor with a finger mouse.

[0171] Specifically, after determining that there is an eye cursor on the above display interface, the mobile phone can directly replace the eye cursor with a finger mouse, that is, the mobile phone replaces the finger mouse at the first position where the eye cursor is displayed, that is, the current position of the eye cursor is the initial position of the finger mouse. Among them, the first position is used to characterize the position of the user's eye gaze point on the display interface. In this way, not only can the eye cursor and the finger mouse be seamlessly connected, without the user having to control the finger mouse to move to the message notification bar, which provides convenience for the user, but also the user can be clearly prompted that the finger mouse service is turned on, reducing the occurrence of erroneous operations caused by the mobile phone accidentally turning on the finger mouse service, thereby improving the user's experience.

[0172] The display effects of the finger mouse and the eye cursor may be the same or different, and there is no specific limitation. For example, the eye cursor may be round, and the finger mouse may be in the shape of a small hand. For another example, the eye cursor may be dark gray, and the finger mouse may be light gray.

[0173] For example, Fig.17 As shown, after receiving the SMS notification message, the mobile phone can display the eye cursor M according to the coordinates of the user's eye gaze point. Afterwards, if the mobile phone captures an image frame with the user's hand, it means that the user wants to trigger the finger mouse service, so the mobile phone can replace the eye cursor M with the finger mouse N. Afterwards, the user can control the finger mouse by changing the gesture to make the mobile phone perform the corresponding operation.

[0174] In some embodiments, the mobile phone needs to determine that the first image frame includes a hand and the gesture corresponding to the hand is a preset gesture before it can display the finger mouse on the display interface according to the preset position, that is, the mobile phone turns on the air mouse function. Afterwards, the mobile phone can control the finger mouse to perform corresponding operations according to the gestures made by the user's hand, which makes the execution efficiency of the mobile phone low. However, in this embodiment, as long as the first image frame includes a hand and there is an eye cursor on the display interface, the mobile phone can replace the eye cursor with a finger mouse, that is, the mobile phone turns on the air mouse function, that is, the mobile phone does not need to perform gesture recognition on the hand included in the first image frame, and does not need to determine whether the recognized gesture is a specific gesture. In this way, the direct use of the air mouse function can be realized, thereby improving the execution efficiency of the mobile phone.

[0175] In one implementation, if the finger mouse and the eye cursor are the same, after the mobile phone determines that the first image frame includes a hand, it can move the eye cursor directly by detecting the user's hand, without replacing the eye cursor with a finger mouse, thereby reducing unnecessary waste of resources.

[0176] In another implementation, if the finger mouse and the eye cursor are different, the mobile phone can directly convert the eye cursor to the finger mouse after determining that the first image frame includes a hand. Alternatively, the mobile phone can delete the eye cursor first. After that, the mobile phone can display the finger mouse according to the gaze point coordinates.

[0177] In another implementation, if the mobile phone does not display the eye cursor, then after determining that the first image frame includes the hand, the mobile phone does not display the finger mouse. That is to say, after determining that the first image frame includes the hand, the mobile phone can directly move the finger mouse by detecting the user's hand. In this way, the situation where the user's visual experience is affected by the movement of the eye cursor can be reduced, thereby improving the user's experience.

[0178] Specifically, if the coordinates corresponding to the above-mentioned finger mouse are located at any control in the message notification bar, it means that there is a possibility that the user will click on the control. Therefore, the mobile phone can deepen the text of the control to facilitate the user to determine whether the control is the control he wants to click, thereby reducing the occurrence of incorrect operations performed by the mobile phone.

[0179] In some embodiments, if the first image frame is an image frame captured by the mobile phone when the user touches his eyes or hair, it means that the user does not want to trigger the finger mouse service. Therefore, even if the mobile phone detects an image frame with the user's hand, there is no need to replace the eye cursor with a finger mouse.

[0180] In one implementation, after replacing the eye cursor with the finger cursor, the mobile phone can directly perform gesture recognition on the hand in the first image frame to obtain the third target gesture. Afterwards, when the third target gesture is the first preset gesture, the mobile phone can directly execute S707, so that the execution efficiency of the mobile phone can be improved. For example, take the first preset gesture as the OK gesture, and the target operation corresponding to the OK gesture as the click operation. If the gesture included in the first image frame is the OK gesture, and the position corresponding to the finger mouse is the "Mark Read" control, it means that the user wants to click on the "Mark Read" control. Therefore, the mobile phone can click on the "Mark Read" control to delete the display of the message notification bar, that is, the mobile phone can not display the message notification bar as shown in the figure. Fig.15 Message notification bar A is shown.

[0181] When the third target gesture is not the first preset gesture, the mobile phone may execute S705 to further determine whether the mobile phone can perform a gesture control operation.

[0182] S705, the mobile phone captures a second image frame, wherein the second image frame is an image frame captured after the mobile phone displays the finger mouse.

[0183] Specifically, after the mobile phone displays the finger mouse, it can trigger the front camera of the mobile phone to continuously collect the second image frame. The second image frame is the image frame collected after the mobile phone displays the finger mouse.

[0184] Exemplarily, the front camera of the mobile phone can capture the second image frame frame by frame, or can capture images according to a preset frame interval, for example, the mobile phone can capture a second image frame every 2 frames, or can capture images according to a preset time interval, for example, the mobile phone can capture a second image frame every 2ms, etc. The preset frame interval and time interval can be set according to actual needs and are not specifically limited.

[0185] S706: The mobile phone determines whether the gesture included in the second image frame is a first preset gesture.

[0186] Specifically, after the mobile phone captures the second image frame, the mobile phone may first perform hand detection on the second image frame to determine whether the second image frame includes the user's hand. Afterwards, if the second image frame includes the user's hand, the mobile phone continues to perform gesture recognition on the hand in the second image frame to obtain the first target gesture. The first target gesture is the gesture made by the user's hand in the second image frame.

[0187] In some embodiments, the mobile phone can capture multiple consecutive third image frames. Afterwards, the mobile phone can perform gesture recognition on the hands in the third image frames respectively to determine whether the gesture included in each third image frame is the first preset gesture. Afterwards, when the gestures included in the multiple third image frames are all the first preset gestures, and the acquisition time corresponding to the multiple third image frames is greater than the third preset time, the mobile phone can control the finger mouse to perform the target operation corresponding to the preset gesture. In this way, the situation where the mobile phone performs an erroneous operation can be avoided, and the execution accuracy of the mobile phone can be improved.

[0188] Specifically, the mobile phone can perform gesture recognition on the hand in the second image frame according to a gesture recognition model to obtain the first target gesture. The gesture recognition model is a model for performing gesture analysis based on the shape and motion trajectory of the hand in the image.

[0189] Afterwards, the mobile phone can determine whether the above-mentioned first target gesture is the first preset gesture. If the first target gesture is the first preset gesture, it means that the user wants the mobile phone to perform the operation corresponding to the preset gesture. Therefore, the mobile phone can execute S707. If the first target gesture is not the first preset gesture, it means that the user does not want to perform the gesture control operation. Therefore, the mobile phone can execute S708. It can be understood that the first preset gesture can be a default gesture pre-configured by the mobile phone, or it can be a gesture pre-configured by the user according to his own habits, etc., and there is no specific limitation. For example, the operation corresponding to the preset gesture of closing the palm is pressing the finger mouse, and the operation corresponding to the preset gesture of spreading the palm is releasing the finger mouse. For another example, the operation corresponding to the preset gesture of raising the index finger is moving the finger mouse.

[0190] For example, Fig.18 As shown, Fig.18 The interface (a) in the figure is a scene corresponding to an image frame of a user's closed palm captured by the front camera 52 of the mobile phone. Fig.18The (b) interface in the figure is the scene corresponding to the image frame of the user's unfolded image captured by the front camera 52 of the mobile phone. Specifically, after the mobile phone captures the image frame of the user's closed palm, it can perform gesture recognition on the hand in the image frame of the user's closed palm based on the gesture recognition model to obtain the target gesture 1, wherein the target gesture corresponds to the action of the user pressing the finger mouse. Afterwards, after the mobile phone captures the image frame of the user's unfolded palm, it can perform gesture recognition on the hand in the image frame of the user's unfolded palm based on the gesture recognition model to obtain the target gesture 2, wherein the target gesture corresponds to the action of the user releasing the finger mouse. Afterwards, the mobile phone can determine whether there is a gesture identical to the target gesture 1 and the target gesture 2 in the first preset gesture. If there is a gesture identical to the target gesture 1 and the target gesture 2 in the first preset gesture, the mobile phone can perform the operation corresponding to the target gesture 1 and the target gesture 2, that is, the mobile phone can perform a click operation on the "reply" control.

[0191] In one implementation, the number of the first preset gestures may be multiple, that is, the mobile phone may determine whether the first target gesture is any of the first preset gestures, and if the target gesture is any of the first preset gestures, the mobile phone may execute S707. If the first target gesture is not any of the first preset gestures, the mobile phone may execute S708.

[0192] S707, the mobile phone controls the finger mouse to perform the target operation corresponding to the first preset gesture.

[0193] Specifically, after determining that the gesture included in the second image frame is the first preset gesture, the mobile phone can control the finger mouse to perform the target operation corresponding to the first preset gesture. In this way, the operation of eye cursor movement can be realized through eye gaze, and the corresponding operation can be performed through the gesture made by the user's hand, which provides convenience for the user and improves the user experience. In addition, after determining that the first target gesture is the first preset gesture, the mobile phone promptly executes the control operation corresponding to the first target gesture to improve the execution accuracy of the mobile phone.

[0194] For example, Fig.19 As shown, after determining that the image frame captured by the front camera 52 includes a preset gesture of the user's unfolded palm, the mobile phone can determine that the user has triggered a click operation, so the finger mouse in the mobile phone can perform a click operation on the "reply" control. Afterwards, in response to the click operation, the mobile phone can display a reply interface. Among them, the reply interface displays the receiving time (08:01 this morning), the sender of the SMS notification (sister), and the detailed SMS content (When are you free tomorrow? Do you want to go shopping and eat!).

[0195] S708: The mobile phone does not perform the gesture control operation.

[0196] Specifically, after determining that the gesture included in the second image frame is not a preset gesture, the mobile phone may not perform the gesture control operation. In this way, the occurrence of incorrect operations performed by the mobile phone can be reduced, the execution accuracy of the mobile phone can be improved, and the user experience can be improved.

[0197] S709: The mobile phone determines whether the gesture included in the first image frame is a second preset gesture.

[0198] In some embodiments, after determining that there is no eye cursor on the display interface, the mobile phone can perform gesture recognition on the user's hand included in the first image frame according to the gesture recognition model to obtain a second target gesture, wherein the second target gesture is a gesture made by the user's hand in the first image frame.

[0199] Afterwards, the mobile phone can determine whether the second target gesture is the second preset gesture. If the second target gesture is the second preset gesture, it means that the user wants the mobile phone to perform the operation corresponding to the second preset gesture. Therefore, the mobile phone can execute S710. If the second target gesture is not the second preset gesture, it means that the user does not want to perform the gesture control operation. Therefore, the mobile phone can return to execute S701 to continue to collect the first image frame. It can be understood that the second preset gesture can be a default gesture pre-configured by the mobile phone, or a gesture pre-configured by the user according to his own habits, etc., without specific limitation. The second preset gesture can be the same as or different from the first preset gesture.

[0200] S710, the phone displays the finger mouse according to the preset position.

[0201] The preset position may be a default position pre-configured by the mobile phone, or a position pre-configured by the user according to his / her own habits, etc., and is not specifically limited. For example, the preset position is the center position of the display interface, wherein the center position is the intersection of two diagonal lines of the display interface. For another example, the preset position is the position corresponding to the vertex of the lower left corner of the display interface.

[0202] In some embodiments, after determining that the gesture included in the first image frame is the second preset gesture, the mobile phone can display a finger mouse at a preset position on the display interface. Afterwards, the mobile phone can directly control the finger mouse to perform the target operation corresponding to the second preset gesture. In this way, the execution efficiency of the gesture control operation can be improved.

[0203] In other embodiments, after the finger mouse is displayed, the user can control the finger mouse by changing the gesture mode to make the mobile phone perform the corresponding operation, that is, after the mobile phone displays the finger mouse, the mobile phone can return to execute the above S705. In this way, the corresponding operation can be performed by the gesture made by the user's hand, which provides convenience for the user and improves the user experience.

[0204] In one implementation, when the first image frame is captured, the mobile phone can determine whether the mobile phone can perform a gesture control operation based on whether the first image frame includes a hand and whether an eye cursor exists on the display interface. Figure 4 The structure shown and Fig. 20 The operation control flow shown in the figure details how the mobile phone executes the target operation.

[0205] S201. A sensor hub in a mobile phone acquires a first image frame.

[0206] The sensor hub is an independent subsystem used to connect and process data from various sensor devices. For example, the sensor hub can receive image frames sent by a camera sensor.

[0207] In some embodiments, after the front camera of the mobile phone captures the first image frame, the first image frame can be sent to the sensor hub. It can be understood that the front camera is a sensor among the above-mentioned camera sensors.

[0208] S202: The sensor center performs hand detection on the first image frame to obtain a hand detection result.

[0209] The hand detection result is used to indicate whether the first image frame includes the user's hand.

[0210] In some embodiments, after acquiring the first image frame, the sensor hub can perform hand detection on the first image frame according to the hand detection model to obtain a hand detection result. If the hand detection result indicates that the first image frame includes the user's hand, that is, the hand detection result includes a hand detection frame, it means that the user has the possibility of triggering the finger mouse service, so the sensor hub in the mobile phone can execute S203. If the hand detection result indicates that the first image frame does not include the user's hand, that is, the hand detection result does not include a hand detection frame, it means that the user does not have the possibility of triggering the finger mouse service, so the sensor hub in the mobile phone can continue to acquire the first image frame to wait for the user to trigger the finger mouse service.

[0211] S203: When the hand detection result indicates that the first image frame includes the user's hand, the sensor hub sends a display icon instruction to the hardware abstraction layer.

[0212] Specifically, after determining that the first image frame includes the user's hand, the sensor hub in the mobile phone can send a display icon indication to the hardware abstraction layer, wherein the display icon indication indicates that the system user interface can display finger mouse operations.

[0213] In this embodiment, as long as the sensor hub determines that the first image frame includes the user's hand, the sensor hub can send a display icon indication to the hardware abstraction layer without performing gesture recognition on the user's hand, and without sending a display icon indication to the hardware abstraction layer when it is determined that the recognized gesture is a preset gesture. In this way, while ensuring the accuracy of the mobile phone's execution of operations, the user's operation process is simplified, the execution efficiency of the mobile phone is improved, convenience is provided to the user, and the user's experience is improved.

[0214] S204. When receiving the display icon indication, the hardware abstraction layer sends a start icon indication to the smart perception service.

[0215] It can be understood that since there are too many algorithms in the hand detection model, that is, the hand detection model is relatively complex, and the delay requirement when starting the air mouse function is not high, the mobile phone needs to send the icon indication carrying the first image frame to the smart perception service, so that the smart perception service can further determine whether the first image frame contains the user's hand, thereby improving the accuracy of hand detection.

[0216] S205. When receiving the display icon indication, the intelligent perception service sends a start icon indication to the system user interface.

[0217] S206: The hardware abstraction layer determines whether there is an eye movement cursor on the display interface.

[0218] In some embodiments, after the system user interface receives the display icon indication, the hardware abstraction layer can determine whether there is an eye cursor on the display interface. If there is an eye cursor on the display interface, it means that the mobile phone is currently performing eye tracking service, so the mobile phone can execute S207 to replace the eye cursor with a finger mouse. If there is no eye cursor on the display interface, it means that the mobile phone is not currently performing eye tracking service, so the mobile phone can execute S210 to further display the finger mouse.

[0219] S207: When there is an eye-movement cursor on the display interface, the hardware abstraction layer sends the gaze point coordinates of the user's eyes to the driver layer or the kernel layer.

[0220] Specifically, after determining that there is an eye movement cursor on the display interface, the hardware abstraction layer can send the user's eye gaze point coordinates to the driver layer or the kernel layer, wherein the gaze point coordinates are used to indicate the coordinate position of the user's eye gaze point on the display interface.

[0221] In some embodiments, the above-mentioned gaze point coordinates can be obtained by the sensor hub through eye movement recognition of the third image frame according to the eye movement recognition algorithm, wherein the third image frame is an image frame collected when the mobile phone receives the notification message. After obtaining the gaze point coordinates, the sensor hub can send the gaze point coordinates to the hardware abstraction layer. Afterwards, the hardware abstraction layer can send the received gaze point coordinates to the driver layer or the kernel layer.

[0222] In one example, the hardware abstraction layer may first send the gaze point coordinates to the driver layer. Afterwards, the driver layer sends the gaze point coordinates to the kernel layer upon receiving the gaze point coordinates. In another example, the hardware abstraction layer may directly send the gaze point coordinates to the kernel layer.

[0223] S208. When receiving the gaze point coordinates of the user's eyes, the driver layer or the kernel layer sends the gaze point coordinates of the user's eyes to the input framework layer.

[0224] S209. When the input framework layer receives the gaze point coordinates of the user's eyes, the input framework layer adds a finger mouse according to the gaze point coordinates of the user's eyes.

[0225] Specifically, after receiving the above-mentioned gaze point coordinates, the input framework layer can add a finger mouse according to the gaze point coordinates, so that the finger mouse can be displayed at the gaze point coordinates of the display interface. It can be understood that if there is an eye cursor on the display interface, it means that the eye cursor has been displayed at the gaze point coordinates of the display interface. Therefore, in order to avoid disturbing the user's vision due to the presence of multiple icons (eye cursor and finger mouse) on the display interface, if the input framework layer is to display the finger mouse, it is necessary to delete the eye cursor on the display interface first to improve the user's visual experience.

[0226] S210: When there is no eye cursor on the display interface, the hardware abstraction layer sends an empty mouse indication to the driver layer or the kernel layer.

[0227] Specifically, after determining that there is no eye cursor on the display interface, the hardware abstraction layer can send an empty mouse indication to the driver layer or the kernel layer, wherein the empty mouse indication indicates that the input framework layer can add a finger mouse operation on the display interface.

[0228] S211. When receiving an empty mouse indication, the driver layer or the kernel layer sends an empty mouse indication to the input framework layer.

[0229] S212: When receiving the air mouse indication, the input framework layer adds a finger mouse according to the preset position.

[0230] Specifically, after receiving the above-mentioned air mouse indication, the input framework layer can add a finger mouse according to the preset position, so that the finger mouse can be displayed at the preset position of the display interface. It can be understood that if there is no eye cursor on the display interface, it means that the mobile phone has not received the notification message or the user has not turned on the eye tracking function. Therefore, the input framework layer adds a finger mouse at the preset position of the display interface, so that the user can be clearly prompted that the finger mouse service has been turned on, avoiding the mobile phone from performing wrong operations due to the user's mistaken triggering of the finger mouse service, thereby improving the accuracy of operation execution.

[0231] It can be understood that the above S207~S209 and S210~S212 are parallel schemes, that is, if there is an eye movement cursor on the display interface, S207~S209 are executed, and there is no need to execute steps S210~S212; if there is no eye movement cursor on the display interface, S210~S212 are executed, and there is no need to execute steps S207~S209.

[0232] S213: The sensor hub acquires a second image frame.

[0233] Among them, the above-mentioned second image frame is an image frame captured by the front camera of the mobile phone after the system user interface displays a finger mouse.

[0234] Specifically, after the mobile phone displays the finger mouse, it can trigger the front camera of the mobile phone to continuously collect the second image frame. The front camera is a camera with AON function, that is, the front camera can collect the second image frame in real time.

[0235] In some embodiments, after the front camera of the mobile phone captures the second image frame, the second image frame can be sent to the sensor hub.

[0236] S214. The sensor hub determines the target coordinates of the finger mouse based on the position of the user's hand in at least two second image frames and the corresponding position coordinates when the finger mouse is added.

[0237] Specifically, after receiving at least two second image frames, the sensor hub can respectively identify the position of the user's hand in the second image frames to obtain the movement data of the user's hand, wherein the movement data of the user's hand may include the movement direction and movement length of the user's hand.

[0238] Afterwards, the sensor hub can determine the target coordinates of the finger mouse based on the above-mentioned user hand movement data and the position coordinates corresponding to the above-mentioned addition of the finger mouse. The target coordinates are the position coordinates of the finger mouse on the display interface after the user's hand moves.

[0239] For example, Fig.21 As shown in the figure, if the user's hand makes a gesture of moving upward in the air, the sensor center in the mobile phone can determine the movement data of the user's hand according to the movement position of the user's hand. After that, the sensor center can determine the target coordinate N2 of the finger mouse according to the movement data and the position coordinate N1 of the finger mouse. In other words, the finger mouse on the mobile phone display interface will move from the position coordinate N1 to the target coordinate N2.

[0240] S215. The hardware abstraction layer receives the target coordinates sent by the sensor hub.

[0241] Specifically, after determining the target coordinates of the finger mouse, the sensor hub can send the target coordinates to the hardware abstraction layer to provide a basis for the subsequent input framework layer to move the finger mouse.

[0242] In some embodiments, the sensor hub may send the target coordinates of the finger mouse to the hardware abstraction layer via the QMI interface.

[0243] S216. When receiving the target coordinates, the hardware abstraction layer sends the target coordinates to the driver layer or the kernel layer.

[0244] S217: When receiving the target coordinates, the driver layer or the kernel layer sends the target coordinates to the input framework layer.

[0245] S218. When the input framework layer receives the target coordinates, it moves the finger mouse according to the target coordinates.

[0246] Specifically, after receiving the target coordinates, the input framework layer may move the finger mouse according to the target coordinates, so that the finger mouse may be displayed at the target coordinates of the display interface.

[0247] S219: The sensor center performs gesture recognition on the user's hand in the second image frame to obtain a target gesture.

[0248] Specifically, after receiving the second image frame, the sensor hub may perform gesture recognition on the user's hand in the second image frame according to a gesture recognition model to obtain a target gesture (or referred to as a first target gesture).

[0249] S220: The hardware abstraction layer receives a target gesture sent by the sensor hub.

[0250] Specifically, after determining the target gesture, the sensor hub may send the target gesture to the hardware abstraction layer to provide a basis for the subsequent input framework layer to display an operation interface.

[0251] S221: The hardware abstraction layer determines whether the target gesture is a preset gesture.

[0252] In some embodiments, after receiving the above-mentioned target gesture, the hardware abstraction layer can determine whether the target gesture is a preset gesture (or called the first preset gesture). If the target gesture is a preset gesture, it means that the user wants the mobile phone to perform an operation corresponding to the preset gesture. Therefore, the mobile phone can execute S222. In this way, the mobile phone can promptly perform the control operation corresponding to the target gesture after determining that the target gesture is a preset gesture, thereby improving the execution accuracy of the mobile phone. If the target gesture is not a preset gesture, it means that the user does not want to perform a gesture control operation. Therefore, the mobile phone may not perform a gesture control operation. In this way, the occurrence of incorrect operations performed by the mobile phone can be reduced, the execution accuracy of the mobile phone can be improved, and the user experience can be improved.

[0253] S222: When the target gesture is a preset gesture, the hardware abstraction layer sends an operation instruction to the driver layer or the kernel layer.

[0254] Specifically, after determining that the target gesture is a preset gesture, the hardware abstraction layer may send an operation instruction to the driver layer or the kernel layer, wherein the operation instruction indicates that the input framework layer may draw an operation interface after executing the target operation.

[0255] It can be understood that due to the high latency requirement when the mobile phone performs gesture control operations, the mobile phone does not need to send the operation instructions carrying the second image frame to the smart perception service, which can improve the execution efficiency of the mobile phone.

[0256] S223. When receiving the operation instruction, the driver layer or the kernel layer sends the operation instruction to the input framework layer.

[0257] S224: When receiving the operation instruction, the input framework layer draws the operation interface according to the operation instruction.

[0258] Specifically, after receiving the above operation instruction, the input framework layer can draw the operation interface according to the operation instruction, so that the mobile phone can display the operation interface.

[0259] In some embodiments, when the mobile phone displays the operation interface, if the multiple consecutive image frames collected do not include the user's hand, or the mobile phone receives a click operation of the user on the operation interface, it means that the user has exited the finger mouse service, so the mobile phone may not perform the gesture control operation. Afterwards, the mobile phone may wait for the user to trigger the finger mouse service next time.

[0260] In this embodiment, the electronic device is taken as a mobile phone for example, but in some embodiments, the electronic device may also be a VR device. If the electronic device is a VR device, after receiving the notification message, the VR device may directly display a finger mouse according to whether the captured image frame includes the user's hand. That is to say, if the image frame captured by the VR device does not include the user's hand, it means that the user does not have the possibility of viewing the notification message, so the VR device may not display the finger mouse; if the image frame captured by the VR device includes the user's hand, it means that the user does not have the possibility of viewing the notification message, so the VR device may display the finger mouse. Afterwards, the VR device may perform corresponding control operations according to the gestures corresponding to the user's hand in the subsequent image frames.

[0261] It can be understood that the user's hand included in the above image frame is the back of the hand, that is, the VR device performs gesture recognition based on the feature points on the back of the hand.

[0262] In some embodiments, the present application provides a computer storage medium including computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the method for adjusting usage parameters as described above.

[0263] In some embodiments, the present application provides a computer program product, which, when executed on an electronic device, enables the electronic device to execute the method for adjusting usage parameters as described above.

[0264] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0265] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0266] The units described as separate components may or may not be physically separated, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple different places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0267] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0268] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.

[0269] The above contents are only specific implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application shall be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. An operation control method, characterized in that: include: The electronic device acquires a first image frame; wherein the first image frame is an image frame currently acquired by the electronic device; In the case where the electronic device detects that the first image frame includes a hand, if there is an eye cursor on the display interface, the electronic device replaces the display of the finger mouse at the first position where the eye cursor is displayed; wherein the first position is used to represent the gaze point position of the user's eyes on the display interface; The electronic device controls the finger mouse to perform a target operation corresponding to a first preset gesture; wherein the first preset gesture is obtained based on gesture recognition of a hand included in a target image frame, the target image frame includes the first image frame and / or the second image frame, and the second image frame is an image frame captured after the electronic device displays the finger mouse.

2. The method according to claim 1, characterized in that The electronic device controls the finger mouse to perform a target operation corresponding to a first preset gesture, including: When a hand is included in the second image frame and a first target gesture corresponding to the hand is the first preset gesture, the electronic device controls the finger mouse to perform a target operation corresponding to the first preset gesture.

3. The method according to claim 2, characterized in that The method further comprises: When the second image frame does not include a hand, or when the first target gesture is not the first preset gesture, the electronic device does not perform a gesture control operation.

4. The method according to claim 1, characterized in that: The electronic device controls the finger mouse to perform a target operation corresponding to a first preset gesture, including: In a case where the hand gesture corresponding to the hand included in the first image frame is the first preset gesture, the electronic device controls the finger mouse to perform a target operation corresponding to the first preset gesture.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: The electronic device acquires a third image frame when receiving the notification message; The electronic device performs eye movement recognition on the third image frame to obtain eye movement recognition data; wherein the eye movement recognition data includes the gaze point coordinates of the user's eyes; The electronic device displays the eye movement cursor on the display interface according to the gaze point coordinates.

6. The method according to claim 5, characterized in that The electronic device performs eye movement recognition on the third image frame to obtain eye movement recognition data, including: The electronic device performs face detection on the third image frame according to the face detection model to obtain a face detection result; wherein the face detection result is used to indicate whether the third image frame includes a face; When the face detection result indicates that the third image frame includes a face, the electronic device performs eye movement recognition on the third image frame according to an eye movement recognition algorithm to obtain the eye movement recognition data.

7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: If the eye cursor does not exist on the display interface, the electronic device performs gesture recognition on the hand included in the first image frame to obtain a second target gesture; In the case where the second target gesture is a second preset gesture, the electronic device displays the finger mouse on the display interface according to a preset position; The electronic device controls the finger mouse to perform a target operation corresponding to the second preset gesture.

8. The method according to claim 7, characterized in that The method further comprises: When the second target gesture is not the second preset gesture, the electronic device continues to acquire the first image frame.

9. An electronic device, characterized in that: The electronic device includes a camera, a memory and one or more processors; the camera, the memory and the processor are coupled; the camera is used to capture images, the memory is used to store computer program codes, and the computer program codes include computer instructions; when the processor executes the computer instructions, the electronic device executes the method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: The method comprises computer instructions, which, when executed on an electronic device, cause the electronic device to execute the method as claimed in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Interface display method and electronic equipment

    CN114510174A

  • Method for interacting with electronic equipment and electronic equipment

    CN116107419A

  • Operation execution method and device, electronic equipment and readable storage medium

    CN116841397A

  • Display control device, display control program and display-control-program product

    US10078416B2

  • Interactive operating system and method for carrying out an operational action in an interactive operating system

    WO2017054894A1