Body language recognition method and electronic equipment

By dynamically adjusting the image resolution on the camera of the electronic device and adjusting the resolution of the acquired image according to the distance between the user and the device, the problem of inaccurate long-distance recognition and high power consumption by electronic devices when recognizing the user's body language is solved, and efficient and accurate recognition effect is achieved.

CN119964190APending Publication Date: 2025-05-09HONOR DEVICE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311435622.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-10-30
Publication Date
2025-05-09

Smart Images

  • Figure CN119964190A_ABST
    Figure CN119964190A_ABST
Patent Text Reader

Abstract

The invention provides a body language recognition method and electronic equipment, relates to the technical field of terminals, and can solve the problems of inaccurate remote recognition and high recognition power consumption when the electronic equipment recognizes a body language of a user. The method comprises the following steps: when the electronic equipment starts an AO function, acquiring an Nth frame of image at a first resolution; if the electronic device recognizes that the Nth frame of image comprises the first target object, the electronic device determines a second resolution based on the first target object and collects an (N + 1) th frame of image at the second resolution, the (N + 1) th frame of image comprises a second target object, and the second target object is the same as the first target object; the resolution when the electronic device collects the image is related to the distance between the electronic device and the collection object. And then, the electronic equipment identifies body languages corresponding to the first target object and the second target object based on the Nth frame image and the (N + 1) th frame image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of terminal technology, and more particularly to a body language recognition method and electronic device. Background Art

[0002] At present, electronic devices (such as mobile phones, tablets, etc.) can realize functions such as air gestures and gaze-on screen through the always on (AO) function. Among them, AO can also be called always online, always on, always online, always open, etc.

[0003] Specifically, AO means that the camera of the electronic device is in a normally open state, which can collect images in real time and identify user operations. It can control the electronic device to perform corresponding functions without the user touching the electronic device, thereby improving the efficiency of human-computer interaction and enhancing the user experience. For example, it can identify air gestures, and control the display interface of the electronic device to slide up and down, take screenshots, and other functions through air gestures; for example, in the screen-on state, it can identify whether the user is looking at the screen. If the user is looking at the screen, the electronic device will not turn off the screen; if the user is not looking at the screen, the electronic device will go black.

[0004] However, in some scenarios, after the electronic device starts AO, there may be problems of inaccurate recognition of user operations and high power consumption. Summary of the invention

[0005] The present application provides a body language recognition method and electronic device, which are used to solve the problems of inaccurate long-distance recognition and high recognition power consumption when the electronic device recognizes the user's body language.

[0006] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:

[0007] In the first aspect, a method for recognizing body language is provided, which can be executed by an electronic device equipped with a camera and supporting AO function, such as a mobile phone, a tablet computer, a laptop computer, etc., and can also be a chip, a chip system or a processor that can implement the method for recognizing body language provided by the present application, or a logic module or software that can implement all or part of the functions of the electronic device. The following is an example of an electronic device as the execution subject.

[0008] The electronic device captures the Nth frame of image at a first resolution (N≥1); if the electronic device recognizes that the Nth frame of image includes a first target object, the electronic device determines a second resolution based on the first target object. The resolution of the electronic device when capturing an image is related to the distance between the electronic device and the captured object.

[0009] The electronic device captures the N+1th frame image at a second resolution, the N+1th frame image includes the second target object, and the first target object is the same as the second target object; then, the electronic device identifies the body language corresponding to the first target object and the second target object based on the Nth frame image and the N+1th frame image.

[0010] Based on the first aspect, when the electronic device starts the AO function, the electronic device dynamically adjusts the resolution of the electronic device when capturing images based on the distance between the captured object and the electronic device, thereby solving the problems of inaccurate long-distance recognition and high recognition power consumption when the electronic device recognizes the user's body language.

[0011] Optionally, the first target object is a palm of the user; or, the first target object is a face of the user; or, the first target object is a torso of the user (such as the waist, legs, arms, etc. of the user).

[0012] Optionally, the first target object is used to indicate the user's body language. Optionally, when the first target object is the user's palm, it indicates the user's air gesture; or, when the first target object is the user's face, it indicates the user's facial expression (such as looking at the screen); or, when the first target object is the user's torso, it indicates the user's posture.

[0013] Optionally, the farther the distance between the first target object and the electronic device is, the higher the second resolution is. Optionally, the closer the distance between the first target object and the electronic device is, the lower the second resolution is.

[0014] In a possible implementation of the first aspect, the electronic device determines the second resolution based on the first target object, including: the electronic device determines the pixel height of the first target object, and determines the second resolution based on the pixel height; the pixel height is used to indicate the number of pixels of the first target object in the vertical direction; the smaller the pixel height is, the greater the distance between the first target object and the electronic device.

[0015] And / or, the electronic device determines an area ratio of the first target object in the Nth frame image, and determines a second resolution based on the area ratio; the smaller the area ratio, the greater the distance between the first target object and the electronic device.

[0016] Alternatively, the electronic device determines a first distance between the first target object and the electronic device, and determines the second resolution based on the first distance.

[0017] In a possible implementation manner of the first aspect, the electronic device determines the second resolution based on the pixel height, including: if the pixel height is less than a first height threshold, the electronic device determines that the second resolution is greater than the first resolution.

[0018] Alternatively, if the pixel height is greater than the second height threshold, the electronic device determines that the second resolution is less than the first resolution.

[0019] Alternatively, if the pixel height is greater than or equal to the first height threshold and less than or equal to the second height threshold, the electronic device determines that the second resolution is equal to the first resolution.

[0020] In this implementation, if the pixel height is less than the first height threshold, it means that the distance between the first target object and the electronic device is far, and in this case, increasing the resolution can solve the problem of inaccurate recognition results caused by the long distance. If the pixel height is greater than the second height threshold, it means that the distance between the first target object and the electronic device is close, and in this case, reducing the resolution can achieve the purpose of reducing power consumption without affecting the recognition result.

[0021] In a possible implementation manner of the first aspect, the electronic device determines the second resolution based on the area ratio, including: if the area ratio is less than a first preset value, the electronic device determines that the second resolution is greater than the first resolution.

[0022] Alternatively, if the area ratio is greater than a second preset value, the electronic device determines that the second resolution is smaller than the first resolution.

[0023] Alternatively, if the area ratio is greater than or equal to the first preset value and less than or equal to the second preset value, the electronic device determines that the second resolution is equal to the first resolution.

[0024] In this implementation, if the area ratio is less than the first preset value, it means that the distance between the first target object and the electronic device is far, and improving the resolution at this time can solve the problem of inaccurate recognition results caused by the long distance. If the area ratio is greater than the second preset value, it means that the distance between the first target object and the electronic device is close, and reducing the resolution at this time can achieve the purpose of reducing power consumption without affecting the recognition result.

[0025] In a possible implementation manner of the first aspect, the electronic device determines the second resolution based on the first distance, including: if the first distance is less than a first distance threshold, the electronic device determines that the second resolution is less than the first resolution.

[0026] Alternatively, if the first distance is greater than the second distance threshold, the electronic device determines that the second resolution is greater than the first resolution.

[0027] Alternatively, if the first distance is greater than or equal to the first distance threshold and less than or equal to the second distance threshold, the electronic device determines that the second resolution is equal to the first resolution.

[0028] In this implementation, if the first distance is less than the first distance threshold, it means that the distance between the first target object and the electronic device is relatively close. In this case, lowering the resolution can achieve the purpose of reducing power consumption without affecting the recognition result. If the first distance is greater than the second distance threshold, it means that the distance between the first target object and the electronic device is relatively far. In this case, increasing the resolution can solve the problem of inaccurate recognition results caused by long distance.

[0029] In a possible implementation manner of the first aspect, the method further includes: if the electronic device does not recognize that the Nth frame image includes the first target object, the electronic device captures the N+1th frame image at a preset resolution; wherein the preset resolution is related to QVGA.

[0030] In this implementation, since the preset resolution is related to QVGA and the resolution corresponding to QVGA is relatively low, if the electronic device does not recognize that the Nth frame image includes the first target object, it captures the N+1th frame image at the preset resolution, which can further reduce power consumption.

[0031] In a second aspect, an electronic device is provided, which has the function of implementing any one of the functions in the first aspect, and the function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.

[0032] According to a third aspect, an electronic device is provided, comprising a memory and one or more processors; wherein the memory is used to store computer instructions, and when the electronic device is running, the processor executes the computer instructions of the memory so that the electronic device executes any one of the methods described in the first aspect.

[0033] In a fourth aspect, a chip system is provided, comprising: at least one processor and an interface, the interface being used to receive instructions and transmit them to at least one processor; at least one processor executes the instructions so that the electronic device executes any one of the methods described in the first aspect.

[0034] In a fifth aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, and when the computer-readable storage medium is run on a computer, the computer can execute any of the methods described in the first aspect.

[0035] In a sixth aspect, a computer program product comprising instructions is provided, which, when executed on a computer, enables the computer to execute any of the methods described in the first aspect.

[0036] Among them, the technical effects brought about by any implementation method in the above-mentioned second to sixth aspects can refer to the technical effects brought about by different implementation methods in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0038] Figure 2 A schematic diagram of a camera collecting images provided in an embodiment of the present application;

[0039] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0040] Figure 4 A schematic diagram of an air gesture provided in an embodiment of the present application;

[0041] Figure 5 A schematic diagram of another air gesture provided in an embodiment of the present application;

[0042] Figure 6 A schematic diagram of a method of staring at the screen without turning it off provided in an embodiment of the present application;

[0043] Figure 7 A schematic diagram of a ring tone automatically becoming smaller when looking at an embodiment of the present application;

[0044] Figure 8 A schematic diagram of a flow chart of a method for recognizing body language provided in an embodiment of the present application;

[0045] Fig. 9 A schematic diagram of a process for determining a second resolution provided in an embodiment of the present application;

[0046] Fig.10 A schematic diagram of another process for determining a second resolution provided in an embodiment of the present application;

[0047] Fig.11 A schematic diagram of the principle of the distance between an electronic device and a collection object provided in an embodiment of the present application;

[0048] Fig.12 A schematic diagram of another process for determining a second resolution provided in an embodiment of the present application;

[0049] Fig.13 A schematic diagram of image cropping provided in an embodiment of the present application;

[0050] Fig.14 A schematic diagram of a flow chart of another method for recognizing body language provided in an embodiment of the present application;

[0051] Fig.15 A schematic diagram of the structure of a chip system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0052] The body language recognition method provided in the embodiment of the present application can be applied to electronic devices such as mobile phones, tablet computers, laptops, ultra-mobile personal computers (UMPCs), handheld computers, netbooks, personal digital assistants (PDAs), wearable electronic devices, vehicle-mounted devices, virtual reality devices, etc., and the embodiment of the present application does not impose any restrictions on this.

[0053] For example, Figure 1 A schematic structural diagram of an electronic device 100 is shown.

[0054] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a camera 193 and a display screen 194, etc.

[0055] It is to be understood that the structure illustrated in the embodiment of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device may include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0056] The processor 110 may include one or more processing units, for example, the processor 110 may include a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0057] Among them, the controller can be the nerve center and command center of the electronic device 100.

[0058] The processor 110 may also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory may store instructions or data that the processor 110 has just used or cyclically used. If the processor 110 needs to use the instruction or data again, it may be directly called from the memory. This avoids repeated access, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0059] In some embodiments, the processor 110 may include one or more interfaces. The interface may include an I2C interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface 130, etc.

[0060] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL).

[0061] The I2S interface can be used for audio communication.

[0062] The PCM interface can also be used for audio communication to sample, quantize and encode analog signals.

[0063] The UART interface is a universal serial data bus used for asynchronous communication. The bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication.

[0064] The MIPI interface can be used to connect the processor 110 with peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), and the like.

[0065] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or a data signal.

[0066] The USB interface 130 is an interface that complies with USB standard specifications, and may specifically be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc.

[0067] It is understandable that the interface connection relationship between the modules illustrated in the embodiment of the present application is only a schematic illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.

[0068] The charging management module 140 is used to receive charging input from a charger, where the charger can be a wireless charger or a wired charger.

[0069] The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110. The power management module 141 can receive input from the battery 142 and / or the charging management module 140 to power the processor 110, the internal memory 121, the display screen 194, the camera 193, and the wireless communication module 160.

[0070] The power management module 141 may be used to monitor performance parameters such as battery capacity, battery cycle times, battery charging voltage, battery discharging voltage, and battery health status (eg, leakage, impedance).

[0071] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.

[0072] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals.

[0073] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc., applied to the electronic device 100. The mobile communication module 150 may include one or more filters, switches, power amplifiers, low noise amplifiers (LAN), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be placed in the processor 110. In some embodiments, at least some of the functions of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.

[0074] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the electronic device 100. The wireless communication module 160 can be one or more devices integrating one or more communication processing modules. The wireless communication module 160 receives electromagnetic waves via the antenna 2, modulates the frequency of the electromagnetic wave signal and performs filtering, and sends the processed signal to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, modulate the frequency of it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0075] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, so that electronic device 100 can communicate with the network and other devices through wireless communication technology.

[0076] The electronic device 100 implements the display function through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, which connects the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs that execute program instructions to generate or change display information.

[0077] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Mini-LED, Micro-LED, Micro-OLED, quantum dot light-emitting diodes (QLED), etc. In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1.

[0078] The electronic device 100 can realize the shooting function through ISP, camera 193, video codec, GPU, display screen 194 and application processor.

[0079] The camera 193 is used to capture static images or videos. In some embodiments, the mobile phone 100 may include 1 or N cameras, where N is a positive integer greater than 1. The camera 193 may be a front camera or a rear camera. The ISP is used to process the data fed back by the camera 193.

[0080] like Figure 2 As shown, the camera 193 generally includes a lens and a photosensitive element (sensor), and the photosensitive element can be any photosensitive device such as a charge-coupled device (CCD) or a complementary metal oxide semiconductor (CMOS).

[0081] Still Figure 2 As shown, in the process of taking photos or videos, the reflected light of the photographed object can generate a light signal after passing through the lens, and the light signal is projected onto the photosensitive element, which converts the received light signal into an electrical signal. Then, the camera 193 sends the obtained electrical signal to the ISP for processing, and finally obtains each frame of the image. Optionally, the ISP can also perform algorithm optimization on the noise, brightness, color, etc. of the image. The ISP can also optimize the exposure, color temperature and other parameters of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0082] Video codecs are used to compress or decompress digital videos. The electronic device 100 may support one or more video codecs. Thus, the electronic device 100 may play or record videos in a variety of coding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0083] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement a data storage function, such as storing music, video and other files in the external memory card.

[0084] The internal memory 121 can be used to store one or more computer programs, which include instructions. The processor 110 can enable the electronic device 100 to execute the methods provided in some embodiments of the present application, as well as various functional applications and data processing, etc. by running the above instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system; the program storage area can also store one or more applications (such as a gallery, contacts, etc.). The data storage area can store data (such as photos, contacts, etc.) created during the use of the electronic device 100. In addition, the internal memory 121 may include a high-speed random access memory; for example, a double data rate synchronous dynamic random access memory (DDR SDRAM), a tightly coupled memory (TCM), etc. It may also include a non-volatile memory, such as one or more disk storage devices, flash memory devices, universal flash storage (UFS), etc. In other embodiments, the processor 110 enables the electronic device 100 to execute the methods provided in the embodiments of the present application, as well as various functional applications and data processing by running instructions stored in the internal memory 121 and / or instructions stored in a memory provided in the processor.

[0085] Among them, TCM is included in the address mapping space of the internal memory 121 and can be accessed as a fast memory. TCM is used to provide low-latency memory to the processor 110. Since TCM does not have the unpredictability unique to cache, TCM can be used to store important routines, such as terminal processing routines or real-time tasks that need to avoid cache uncertainty. In addition, TCM can be used to save temporary register data, data types whose local attributes are not suitable for cache, and important data structures such as interrupt stacks.

[0086] Generally, TCM has a smaller capacity and DDR has a larger capacity; the data transmission rate of DDR is greater than that of TCM; and the power consumption of DDR is greater than that of TCM. In the embodiment of the present application, DDR or TCM is used for data caching, which can be set according to the actual situation without limitation.

[0087] The electronic device 100 can implement audio functions such as music playing and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.

[0088] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be arranged in the processor 110, or some functional modules of the audio module 170 can be arranged in the processor 110.

[0089] The speaker 170A, also called a "speaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or listen to a hands-free call through the speaker 170A.

[0090] The receiver 170B, also called a "earpiece", is used to convert audio electrical signals into sound signals. When the electronic device 100 receives a call or voice message, the voice can be received by placing the receiver 170B close to the human ear.

[0091] Microphone 170C, also called "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to microphone 170C to input the sound signal into microphone 170C. The electronic device 100 can be provided with one or more microphones 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the source of sound, realize directional recording function, etc.

[0092] The earphone interface 170D is used to connect a wired earphone and can be a USB interface 130 or a 3.5 mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunication industry association of the USA (CTIA) standard interface.

[0093] The sensor module 180 may include a pressure sensor, a gyroscope sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc., but the embodiments of the present application do not impose any restrictions on this.

[0094] Of course, the electronic device 100 provided in the embodiment of the present application may also include one or more devices such as a positioning module 181, a button 190, a motor 191, an indicator 192 and a SIM card interface 195, and the embodiment of the present application does not impose any restrictions on this.

[0095] The implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0096] Exemplary, combined Figure 1 As shown, in some embodiments of the present application, the camera 193 can be used to collect the air gestures input by the user. For example, the camera 193 can be set to always on to collect the current image in real time. The camera 193 can send each frame of the collected image to the processor 110, and the processor 110 can identify the specific air gesture input by the user at this time based on the image sent by the camera 193.

[0097] For example, Figure 3 As shown, the camera (such as Figure 1 When the camera 193 shown in the figure collects the air gesture input by the user, the reflected light of the photographed object (i.e., the air gesture) can generate a light signal after passing through the lens in the camera. The light signal is projected onto the photosensitive element in the camera. The photosensitive element converts the received light signal into an electrical signal. Then, the photosensitive element in the camera sends the obtained electrical signal to the ISP module for processing, and finally obtains each frame of the image. Optionally, the photosensitive element in the camera can send the obtained electrical signal to the ISP module for processing through the MIPI interface.

[0098] The ISP module processes the electrical signals sent by the photosensitive element in the camera to obtain each frame of image. Figure 3 As shown, the ISP module transfers each image frame to the memory (such as Figure 1 The internal memory 121 shown in the figure). The storage data area of ​​the memory can store each frame of the image sent by the ISP module. In some embodiments, the memory can also store one or more computer programs, and the one or more computer programs include instructions. The processor (such as Figure 1 The processor 110 shown can read each frame of the image from the memory and identify the specific air gesture input by the user at this time by running the above instructions stored in the memory.

[0099] Optionally, the ISP module and the processor may be integrated together, or the ISP module and the processor may be independent components. Figure 3 As shown, the ISP module and the processor can be integrated in the same chip, for example, in a system-on-a-chip (SoC). In the embodiment of the present application, the processor can include a processing unit such as a DSP and a CPU.

[0100] In some embodiments, the ISP module may be integrated into the processor, or the ISP module may be an independent component outside the processor. Figure 3 The ISP module is taken as an example as an independent component outside the processor.

[0101] Exemplarily, the camera can send each frame of the captured image to the processor through the ISP module, and the processor can determine whether the air gesture input by the user this time is a swipe up, swipe down, swipe left, or swipe right, etc. based on the multiple frames of images sent by the camera. It should be noted that the "up", "down", "left", and "right" described in the embodiments of the present application are all relative directions based on the normal use of electronic devices (such as mobile phones) by users, which can be understood in conjunction with the drawings of the specific embodiments, and the embodiments of the present application do not impose any limitations on this.

[0102] Optionally, the processor can recognize the air gesture input by the user through gesture recognition technology. Among them, gesture recognition technology is mainly a technology that recognizes human gestures through algorithms, detects the key point information of the fingers through vision, determines the complex movements of the five fingers of the human body through the displacement information of the key point information of the fingers, and then determines the type of gesture.

[0103] As an example, a first coordinate system may be pre-set in a mobile phone, the first coordinate system including an x-axis and a y-axis, the x-axis being the horizontal direction when the mobile phone is normally displayed, that is, the horizontal direction when the user normally reads the displayed content of the mobile phone. Correspondingly, the y-axis is the vertical direction perpendicular to the x-axis. For example, the x-axis is the horizontal direction along the short side of the mobile phone screen, and the y-axis is the vertical direction along the long side of the mobile phone screen. The processor may determine the movement distance of the user's palm on the x-axis and the y-axis based on the position change of a certain point on the user's palm in multiple consecutive frames of images. If the movement distance of the user's palm on the x-axis is greater than a threshold value 1, the processor may determine that the user has performed an air gesture of sliding to the left (or right). If the movement distance of the user's palm on the y-axis is greater than a threshold value 2, the processor may determine that the user has performed an air gesture of sliding upward (or downward).

[0104] For example, Figure 4 As shown, after the mobile phone enters the air gesture operation mode, the processor can identify the pattern of the user's palm in the image every time it obtains a frame of image. Then, the processor can determine the position of the user's palm in the image. For example, the processor can use the coordinates of the center point A(m) of the user's palm as the position of the user's palm in each frame of image. In this way, starting from the coordinates of the starting position A1, the processor can calculate the movement distance S1 of the user's palm in the negative direction of the y-axis according to the coordinates of the center point A(m) of each frame of image (such as A2, A3). When the movement distance S1 of the user's palm in the negative direction of the y-axis is greater than the threshold 2, the processor determines that the user has performed an air gesture of sliding downward.

[0105] In other embodiments, the camera may send each captured frame of image to the processor, and the processor may determine that the air gesture input by the user this time is a grasping gesture based on the multiple frames of images sent by the camera.

[0106] For example, Figure 5 As shown, after the mobile phone enters the air gesture operation mode, the processor can identify the pattern of the user's palm in each frame of the image after acquiring the image. Then, the processor can determine the air gesture performed by the user by identifying the state of the user's palm. For example, the processor recognizes that the user's palm changes from being opened to being closed, and determines that the user has performed a grasping air gesture.

[0107] Accordingly, when the processor recognizes a specific air gesture input by the user, it can respond to the air gesture to perform a corresponding operation. This allows the user to control the electronic device to implement the corresponding function through the air gesture. For example, when the processor recognizes that the air gesture input by the user is swiping up, swiping down, swiping left or swiping right, it can respond to the air gesture to perform operations such as turning pages, returning, and the next step. For example, Figure 4 As shown, when the processor recognizes that the air gesture input by the user is a downward slide, a page turning operation can be performed in response to the air gesture.

[0108] Another example is Figure 5 As shown, when the processor recognizes that the air gesture input by the user is a grasping gesture, it can perform operations such as screenshots in response to the air gesture. The embodiment of the present application does not specifically limit the operation corresponding to the air gesture, which is subject to the actual setting.

[0109] It should be noted that in addition to collecting air gestures input by users through cameras, electronic devices can also collect air gestures input by users through one or more devices such as infrared sensors and ultrasonic sensors. The embodiments of the present application do not impose any restrictions on this.

[0110] In other embodiments of the present application, the camera can also be used to collect the user's face image. For example, the camera can be set to a normally open state to collect the current image in real time. The camera can send each frame of the collected image to the processor, and the processor can identify whether the user is looking at the screen based on the multiple frames of images sent by the camera.

[0111] For example, the processor can identify whether the user is looking at the screen through eye tracking technology. Among them, eye tracking technology refers to tracking eye movement by detecting the position of the eye's gaze point or the movement of the eyeball relative to the head. Accordingly, when the processor recognizes that the user is looking at the screen, it can execute a corresponding function in response to the user's operation of looking at the screen.

[0112] For example, Figure 6 As shown, when the processor recognizes that the user is looking at the screen, it can respond to the user's operation of looking at the screen to execute the "looking at the screen without turning off the screen" function. Among them, "looking at the screen without turning off the screen" means that when the user looks at the screen, the mobile phone does not enter the screen-off state. The screen-off state means that the electronic device displays the screen-off interface in part of the screen without lighting up the entire screen. Among them, the screen-off interface is used to display time, date, incoming call information, push messages, weather and other content.

[0113] For example, Figure 7 As shown, when the processor recognizes that the user is looking at the screen, it can respond to the user's operation of looking at the screen and execute the function of "looking at the ringtone automatically becomes smaller". Among them, "looking at the ringtone automatically becomes smaller" means that when the user looks at the screen, the incoming call ringtone (or SMS notification ringtone) of the mobile phone automatically becomes smaller.

[0114] In the embodiment of the present application, when the AO function of the electronic device is activated, that is, the camera is set to the normally open state, the camera collects the current image in real time, which will inevitably increase the functions of the electronic device and affect the performance of the electronic device. Based on this, controlling the power consumption caused by the activation of the AO function has become a technical problem that needs to be solved urgently.

[0115] In some embodiments, the power consumption can be reduced by reducing the frame rate of the camera. The frame rate of the camera refers to the number of images collected per second; for example, if the frame rate of the camera is 25fps, it means that the camera collects 25 images per second. The higher the frame rate, the more images the camera collects per second, and the faster the image recognition rate. Correspondingly, the lower the frame rate, the fewer images the camera collects per second, and the slower the image recognition rate.

[0116] In other embodiments, the purpose of reducing power consumption can be achieved by reducing the image resolution. For example, after the image resolution is reduced, the camera will capture images at a lower resolution. The image resolution represents the number of pixels in the length and width of the image, and the image resolution determines the quality of the image. The higher the image resolution, the more pixels in the length and width of the image, and the clearer the image.

[0117] The image resolution can be expressed as m×n, where m represents the number of pixels in the length direction of the image, and n represents the number of pixels in the width direction of the image. m and n can be set according to actual needs.

[0118] Exemplarily, the image resolution may be 160×120, that is, the camera captures each frame of the image at a fixed resolution (such as 160×120). In this case, if the distance between the user and the electronic device is within a preset range (such as 0 to 60 cm), the electronic device can accurately identify the user's operation (such as the user's air gesture input, the user looking at the screen, etc.). When the distance between the user and the electronic device is outside the preset range, there will be a problem of low accuracy in identifying the user's operation, that is, inaccurate long-distance recognition.

[0119] In the above two solutions, reducing the frame rate of the camera may reduce the sensitivity of the electronic device in responding to user operations (such as user-input gestures, user staring at the screen, etc.). Reducing the image resolution may result in inaccurate recognition results at a long distance.

[0120] Based on this, an embodiment of the present application provides a method for recognizing body language. The method can dynamically adjust the image resolution, that is, the camera captures images at different resolutions, thereby achieving the purpose of reducing power consumption without affecting the long-distance recognition results.

[0121] The following is a detailed description of the body language recognition method provided in the embodiment of the present application in conjunction with the accompanying drawings of the specification. The body language recognition method provided in the embodiment of the present application can be performed by an electronic device equipped with a camera and supporting the AO function, such as a mobile phone, a tablet computer, a laptop computer, etc., or by a chip, a chip system or a processor that can implement the body language recognition method provided in the embodiment of the present application, or a logic module or software that can implement all or part of the functions of the electronic device, and the embodiment of the present application does not impose specific restrictions on this. The following is a detailed description of the scheme provided in the embodiment of the present application with the electronic device as the execution subject.

[0122] Figure 8 A flow chart of a method for recognizing body language provided in an embodiment of the present application is shown as follows: Figure 8 As shown, the method may include the following steps.

[0123] Step 201: The electronic device starts the AO function and captures the Nth frame image at a first resolution; N≥1.

[0124] Optionally, the electronic device can start the AO function after entering the screen-off state; or, the electronic device can start the AO function in the screen-on state; or, the electronic device can start the AO function when receiving an incoming call (or text message); or, the electronic device can start the AO function after turning on; or, the electronic device can start the AO function in response to user operations, etc., and the embodiments of the present application do not limit this.

[0125] Optionally, when N=1, the first resolution may be pre-set by the electronic device; or, when N is greater than 1, the first resolution may be determined by the electronic device based on the N-1th frame image; the embodiment of the present application does not limit the size of the first resolution, which shall be subject to actual conditions.

[0126] Exemplarily, the first resolution may be 160×120; or, the first resolution may be 320×240; or, the first resolution may be 480×360; or, the first resolution may be 640×480; or, the first resolution may be other specifications, which will not be described in detail.

[0127] In some embodiments, in combination Figure 3As shown, when the electronic device captures the Nth frame image at the first resolution, the photosensitive element in the camera sends the obtained electrical signal to the ISP module for processing to obtain the Nth frame image of the first resolution. Wherein, when the first resolution is 160×120, the ISP module outputs the image according to the QQVGA (quarter-quarter video graphics array) specification, that is, the ISP module outputs the Nth frame image of the QQVGA specification to the memory. Alternatively, when the first resolution is 320×240, the ISP module outputs the image according to the QVGA (quarter video graphics array) specification, that is, the ISP module outputs the Nth frame image of the QVGA specification to the memory. Wherein, QVGA is the standard resolution, and QQVGA is 1 / 4 screen of QVGA.

[0128] It should be noted that, in actual implementation, the camera can capture the first frame of image at the lowest resolution (e.g., 160×120), which is preset by the electronic device. In this way, when the camera captures the first frame of image at a lower resolution, it can not only reduce power consumption, but also will not affect subsequent image recognition.

[0129] Step 202: If the electronic device recognizes that the Nth frame image includes a first target object, a second resolution is determined based on the first target object.

[0130] In some embodiments, the electronic device may identify the Nth frame image based on model A to determine whether the Nth frame image includes the first target object. Exemplarily, the model A may be a face detection module, or the model A may be a hand detection model, etc., without limitation.

[0131] Optionally, the first target object may be the user's palm; or, the first target object may be the user's face; or, the first target object may also be the user's torso (such as the user's waist, legs, arms, etc.), etc., which will not be elaborated herein.

[0132] Exemplarily, the first target object is used to indicate the user's body language. For example, when the first target object is the user's palm, it can indicate the user's air gestures; for another example, when the first target object is the user's face, it can indicate the user's facial expressions (such as looking at the screen); for another example, when the first target object is the user's torso, it can indicate the user's posture, etc., which will not be elaborated herein.

[0133] In some embodiments, the resolution of the electronic device when capturing an image is related to the distance between the electronic device and the captured object. Exemplarily, the electronic device can detect feature information of the first target object, determine the distance between the first target object and the electronic device, and determine the second resolution based on the distance.

[0134] Optionally, the farther the distance between the first target object and the electronic device is, the higher the second resolution is. Correspondingly, the closer the distance between the first target object and the electronic device is, the lower the second resolution is.

[0135] Exemplarily, the characteristic information may be the pixel height of the first target object; or, the characteristic information may be the area ratio of the first target object in the Nth frame image; or, the characteristic information may include the pixel height of the first target object and the area ratio of the first target object in the Nth frame image; or, the characteristic information may also be the first distance between the first target object and the electronic device, which is not elaborated herein.

[0136] The pixel height refers to the number of pixels of the first target object in the longitudinal direction (ie, vertical direction). The area ratio refers to the proportion of the first target object in the Nth frame image.

[0137] In some embodiments, the electronic device determines the pixel height of the first target object, and determines the second resolution based on the pixel height. The smaller the pixel height, the farther the distance between the first target object and the electronic device, and the larger the second resolution. Correspondingly, the larger the pixel height, the closer the distance between the first target object and the electronic device, and the smaller the second resolution.

[0138] Exemplarily, the electronic device may determine the image resolution of the first target object based on the first resolution, and further determine the pixel height of the first target object. The image resolution of the first target object refers to the number of pixels occupied by the length and width of the first target object in the Nth frame image. In other words, the electronic device may calculate the number of pixels occupied by the length and width of the first target object in the Nth frame image based on the first resolution, and determine the pixel height of the first target object based on the number of pixels occupied by the first target object in the width direction (i.e., longitudinal direction) in the Nth frame image.

[0139] For example, Fig. 9 As shown, the electronic device determines the second resolution based on the pixel height, including: if the pixel height is less than the first height threshold, increasing the first resolution to obtain the second resolution, that is, the second resolution is greater than the first resolution; if the pixel height is greater than the second height threshold, reducing the first resolution to obtain the second resolution, that is, the second resolution is less than the first resolution; if the pixel height is greater than or equal to the first height threshold and less than or equal to the second height threshold, maintaining the first resolution unchanged, that is, the second resolution is equal to the first resolution.

[0140] It should be noted that the first height threshold and the second height threshold corresponding to different first target objects are different. That is to say, in the embodiment of the present application, the first height threshold and the second height threshold can be flexibly set based on the first target object. For example, under normal circumstances, the size of the user's face is larger than the size of the user's palm. When the first target object is the user's face, the first height threshold and the second height threshold can be set higher. Correspondingly, when the first target object is the user's palm, the first height threshold and the second height threshold can be set lower. The specific setting is subject to the actual setting and is not limited.

[0141] In some other embodiments, the electronic device determines the area ratio of the first target object in the Nth frame image, and determines the second resolution based on the area ratio. The smaller the area ratio, the farther the distance between the first target object and the electronic device, and the larger the second resolution. Correspondingly, the larger the area ratio, the closer the distance between the first target object and the electronic device, and the smaller the second resolution.

[0142] Exemplarily, the electronic device may determine the area of ​​the Nth frame image, for example, by multiplying the length and width of the Nth frame image. Similarly, the electronic device may determine the area of ​​the first target object by multiplying the length and width of the first target object. Then, the electronic device determines the area ratio based on the area of ​​the first target object and the area of ​​the Nth frame image.

[0143] For example, Fig.10 As shown, the electronic device determines the second resolution based on the area ratio, including: if the area ratio is less than the first preset value, the first resolution is increased to obtain the second resolution, that is, the second resolution is greater than the first resolution. If the area ratio is greater than the second preset value, the first resolution is reduced to obtain the second resolution, that is, the second resolution is less than the first resolution. If the area ratio is greater than or equal to the first preset value and less than or equal to the second preset value, the first resolution is maintained unchanged, that is, the second resolution is equal to the first resolution.

[0144] It should be noted that the first preset value and the second preset value corresponding to different first target objects are different. That is to say, in the embodiment of the present application, the first preset value and the second preset value can be flexibly set based on the first target object. For example, under normal circumstances, the user's face size is larger than the user's palm size. When the first target object is the user's face, the first preset value and the second preset value can be set higher. Correspondingly, when the first target object is the user's palm, the first preset value and the second preset value can be set lower. The specific setting is subject to the actual setting and is not limited.

[0145] In some other embodiments, the electronic device determines a first distance between the first target object and the electronic device, and determines the second resolution based on the first distance. Optionally, the larger the first distance, the larger the second resolution; the smaller the first distance, the smaller the second resolution.

[0146] For example, the electronic device may first determine the pixel height of the first target object, and then determine the first distance based on the pixel height. Fig.11 As shown, the positional relationship between the camera and the first target object can be similar to an isosceles triangle, wherein the vertex angle of the isosceles triangle can be approximately understood as the field of view (FOV) of the camera, the base of the isosceles triangle can be approximately understood as the pixel height of the first target object, and the height of the isosceles triangle can be approximately understood as the first distance.

[0147] Optionally, the first distance may satisfy the following expression: L=(h / 2) / tan(FOV / 2); wherein L represents the first distance, h represents the pixel height, and FOV represents the field of view of the camera.

[0148] For example, Fig.12 As shown, the electronic device determines the second resolution based on the first distance, including: if the first distance is less than the first distance threshold, the first resolution is reduced to obtain the second resolution, that is, the second resolution is less than the first resolution. If the first distance is greater than the second distance threshold, the first resolution is increased, that is, the second resolution is greater than the first resolution, to obtain the second resolution. If the first distance is greater than or equal to the first distance threshold, and less than or equal to the second distance threshold, the first resolution is maintained unchanged, that is, the second resolution is equal to the first resolution.

[0149] In other embodiments of the present application, the electronic device may also determine the second resolution based on the pixel height of the first target object and the area ratio of the first target object in the Nth frame image. The specific implementation method may refer to the above embodiment and will not be described in detail here.

[0150] In some embodiments, in combination Figure 3 As shown, the ISP module processes the electrical signal sent by the photosensitive element in the camera to obtain the Nth frame image. The ISP module transmits the Nth frame image to the memory, and the processor can read the Nth frame image from the memory, and identify the first target object included in the Nth frame image by running the instructions stored in the memory, and determine the second resolution based on the first target object.

[0151] Step 203: The electronic device captures the N+1th frame image at a second resolution; wherein the N+1th frame image includes the second target object, and the first target object is the same as the second target object.

[0152] It should be noted that, for an example of the second resolution, reference may be made to the above step 201, which will not be described in detail here. In addition, for an example of the second target object, reference may be made to the above step 202, which will not be described in detail here.

[0153] In some embodiments, if the electronic device does not recognize that the Nth frame image includes the first target object, the electronic device captures the N+1th frame image at a preset resolution. The preset resolution is related to QVGA. For example, the preset resolution is QVGA, that is, the preset resolution is 320×240; or the preset resolution is QQVGA, that is, the preset resolution is 160×120.

[0154] In some embodiments, in combination Figure 3 As shown, after the processor determines the second resolution, the processor sends a first instruction to the camera, where the first instruction is used to instruct the camera to capture the N+1th frame image at the second resolution.

[0155] Optionally, the processor sends the first instruction to the camera through the ISP module; for example, the processor sends the first instruction to the ISP module, and after the ISP module receives the first instruction sent by the processor, it calls an interface such as I2C to send the first instruction to the camera. Alternatively, the processor calls an interface such as I2C to directly send the first instruction to the camera. No limitation is given.

[0156] For example, Figure 3 As shown, after the camera receives the first instruction sent by the processor, the camera captures the N+1th frame image at the second resolution. For example, the photosensitive element in the camera sends the obtained electrical signal to the ISP module for processing to obtain the N+1th frame image at the second resolution.

[0157] Among them, when the second resolution is 640×480, the ISP module outputs the image according to the VGA (video graphics array) specification, that is, the ISP module outputs the N+1th frame image of the VGA specification to the memory. Or, when the second resolution is 480×360, the ISP module outputs the image according to the 360P specification, that is, the ISP module outputs the N+1th frame image of the 360P specification to the memory.

[0158] Optionally, after the electronic device captures the N+1th frame image at the second resolution, if the electronic device recognizes that the N+1th frame image includes the second target object, the electronic device determines the third resolution based on the second target object, and captures the N+2th frame image at the third resolution. Further, if the electronic device recognizes that the N+2th frame image includes the third target object, the electronic device determines the fourth resolution based on the third target object, and captures the N+3th frame image at the fourth resolution, and so on.

[0159] That is to say, in the embodiment of the present application, when the electronic device starts the AO function, the electronic device dynamically adjusts the resolution and captures each frame of image at different resolutions; and the dynamic adjustment of the resolution is based on the distance between the target object and the electronic device, that is, the closer the distance between the target object and the electronic device, the smaller the resolution; and the farther the distance between the target object and the electronic device, the larger the resolution. In this way, by dynamically adjusting the resolution based on the distance between the target object and the electronic device, the purpose of reducing power consumption can be achieved without affecting the long-distance recognition results.

[0160] It should be noted that, for each frame of image, the specific implementation method of the electronic device determining the resolution can refer to the above step 202 and will not be described in detail.

[0161] Optionally, after the electronic device captures each frame of image at different resolutions, the electronic device identifies the body language corresponding to the target object based on the continuous Z frame images. Wherein, Z ≥ 2. For example, the electronic device identifies the air gesture input by the user based on the continuous Z frame images; or, the electronic device identifies whether the user is looking at the screen based on the continuous Z frame images. Wherein, the continuous Z frame images include the Nth frame image and the N+1th frame image. Exemplarily, after the electronic device captures each frame of image at different resolutions, the electronic device may continue to perform the following steps, such as executing step 204.

[0162] Step 204: The electronic device identifies the body language corresponding to the first target object and the second target object based on the Nth frame image and the N+1th frame image.

[0163] Taking the example of the electronic device identifying the body language corresponding to the first target object based on the Nth frame image, illustratively, before the electronic device identifies the body language corresponding to the first target object, the electronic device determines whether the image size of the Nth frame image is larger than a preset size. If the image size is larger than the preset size, the Nth frame image is cropped or scaled down; if the image size is smaller than the preset size, the Nth frame image is enlarged.

[0164] For example, Fig.13 As shown, the image size of the Nth frame image is larger than the preset size, and the electronic device crops the Nth frame image to obtain the cropped Nth frame image. For example, the image size of the Nth frame image is 360×480, and the image size of the cropped Nth frame image is 320×240.

[0165] It should be noted that when the electronic device recognizes the body language corresponding to the first target object based on the Nth frame image, if the image size of the Nth frame image is too large, it will lead to excessive calculation time and occupy more computing resources and memory; if the image size of the Nth frame image is too small, it will cause the image to be blurred or distorted. Based on this, the embodiment of the present application determines the relationship between the image size of the Nth frame image and the preset size, and adopts the method of cropping and scaling (resizing) to ensure the resolution of the Nth frame image, thereby improving the recognition accuracy.

[0166] It should be noted that, for the N+1th frame image of the second resolution, the electronic device will also determine the relationship between the image size of the N+1th frame image and the preset size. The specific implementation method can be referred to the above description and will not be repeated here.

[0167] In some embodiments, the electronic device can identify the body language corresponding to the first target object and the second target object through model B; wherein model B can be an eye movement detection model; or, model B can be a gesture detection model, without limitation.

[0168] It is understandable that when the electronic device identifies the body language corresponding to the target object based on the continuous Z-frame images, for each frame of the image, the electronic device can use the implementation shown in step 204 to perform identification, which will not be described in detail.

[0169] In addition, when the body language corresponding to the target object is an air gesture, for an example of a specific implementation method of recognizing the air gesture input by the user, refer to the above Figure 4 and Figure 5 When the body language corresponding to the target object is looking at the screen, the specific implementation method of identifying whether the user is looking at the screen can be described in detail in the above-mentioned Figure 6 and Figure 7 The embodiments shown are not described in detail.

[0170] The following embodiments are combined Figure 3 and Fig.14 As shown, the solution provided in the embodiment of the present application is described in detail.

[0171] For example, Figure 3 and Fig.14 As shown, the camera captures the Nth frame image at the first resolution. The processor obtains the Nth frame image at the first resolution; on the one hand, the processor identifies whether the Nth frame image includes the first target object; on the other hand, if the Nth frame image includes the first target object, the processor determines the second resolution based on the first target object.

[0172] Optionally, if the Nth frame image includes the first target object, the processor determines whether the image size of the Nth frame image is larger than a preset size. If the image size is larger than the preset size, the processor crops (or reduces) the Nth frame image, and identifies the body language corresponding to the first target object based on the cropped (or reduced) Nth frame image. Correspondingly, if the image size is smaller than the preset size, the processor enlarges the Nth frame image, and identifies the body language corresponding to the first target object based on the enlarged Nth frame image.

[0173] Optionally, after the processor determines the second resolution, the processor sends a first instruction to the camera, instructing the camera to capture the N+1th frame image at the second resolution. Then, the camera captures the N+1th frame image at the second resolution, and after the processor obtains the N+1th frame image at the second resolution, it continues to execute Fig.14 The process shown is repeated in this way.

[0174] To sum up, by adopting the solution provided in the embodiment of the present application, after the electronic device starts the AO function, the electronic device collects each frame of image at different resolutions. By dynamically adjusting the resolution, not only can the power consumption be reduced, but also the accuracy of the long-distance recognition results can be ensured.

[0175] It should be noted that in the embodiments of the present application, any collection and processing of user data shall comply with applicable privacy laws, and the user's personal information shall be collected only for legitimate and reasonably applicable purposes. Such data collection / sharing shall only occur with the user's consent or other legal basis prescribed by applicable law. In other words, in the embodiments of the present application, the electronic device activates the AO function and collects the user's image in real time with the user's consent.

[0176] It can be understood that the contents recorded in each embodiment of the present application can explain the technical solutions in other embodiments of the embodiments of the present application, and the technical features recorded in each embodiment can also be applied in other embodiments to form new solutions by combining the technical features in other embodiments. The present application only exemplarily lists several embodiments for illustration, and does not mean that the present application is limited to this.

[0177] The present application embodiment provides an electronic device, which may include a memory and one or more processors; the memory stores computer program code, which includes computer instructions. When the computer instructions are executed by the processor, the electronic device performs the functions or steps in the above embodiment. The structure of the electronic device can refer to the above Figure 1 The structure of the electronic device 100 is shown.

[0178] The present application also provides a chip system for use in electronic devices. Fig.15 As shown, the chip system 1100 includes at least one processor 1101 and at least one interface circuit 1102. The processor 1101 may be the processor of the above embodiment. Figure 1 The processor 110 is shown. The interface circuit 1102 may be, for example, an interface circuit between the processor and an external memory; or an interface circuit between the processor and an internal memory.

[0179] The processor 1101 and the interface circuit 1102 can be interconnected via lines. For example, the interface circuit 1102 can be used to receive signals from other devices (such as the memory of the electronic device 100). For another example, the interface circuit 1102 can be used to send signals to other devices (such as the processor 1101). Exemplarily, the interface circuit 1102 can read the instructions stored in the memory and send the instructions to the processor 1101. When the instructions are executed by the processor 1101, the electronic device can execute the various functions or steps executed by the mobile phone in the above embodiment. Of course, the chip system can also include other discrete devices, which is not specifically limited in the embodiments of the present application.

[0180] An embodiment of the present application also provides a computer-readable storage medium, which includes computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes each function or step executed by the electronic device in the above method embodiment.

[0181] The embodiment of the present application also provides a computer program product. When the computer program product is run on a computer, the computer is enabled to execute each function or step executed by the electronic device in the above method embodiment.

[0182] It should be noted that the terms "first" and "second" in the specification, claims and drawings of this application are used to distinguish different objects rather than to describe a specific order. "First" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of this embodiment, unless otherwise specified, "multiple" means two or more.

[0183] In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device comprising a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products or devices.

[0184] It should be understood that in the present application, "at least one (item)" refers to one or more. "Multiple" refers to two or more. "At least two (items)" refers to two or three and more than three. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple. “When” and “if” both mean that corresponding measures will be taken under certain objective circumstances. It does not limit the time, nor does it require any judgment when it is implemented, nor does it mean that there are other limitations.

[0185] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way for easy understanding.

[0186] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules according to the system, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0187] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0188] The units described as separate components may or may not be physically separated, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple different places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0189] In addition, each functional unit in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of software functional units.

[0190] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.

[0191] The above contents are only specific implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application shall be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A method for recognizing body language, characterized in that: Applied to an electronic device, the electronic device starts an always-online AO ​​function, and the method includes: The electronic device collects the Nth frame of image at a first resolution; N≥1; If the electronic device recognizes that the Nth frame image includes a first target object, a second resolution is determined based on the first target object; wherein the resolution of the electronic device when capturing an image is related to the distance between the electronic device and the captured object; The electronic device collects an N+1th frame of image at the second resolution; the N+1th frame of image includes a second target object, and the first target object is the same as the second target object; The electronic device recognizes body language corresponding to the first target object and the second target object based on the Nth frame image and the N+1th frame image.

2. The method according to claim 1, characterized in that: The farther the distance between the first target object and the electronic device is, the higher the second resolution is; and the closer the distance between the electronic device and the first target object is, the lower the second resolution is.

3. The method according to claim 1 or 2, characterized in that: The determining a second resolution based on the first target object comprises: The electronic device determines a pixel height of the first target object, and determines the second resolution based on the pixel height; the pixel height is used to indicate the number of pixels of the first target object in a vertical direction, and the smaller the pixel height is, the farther the distance between the first target object and the electronic device is; and / or, The electronic device determines the area ratio of the first target object in the Nth frame image, and determines the second resolution based on the area ratio; the smaller the area ratio, the farther the distance between the first target object and the electronic device is; or The electronic device determines a first distance between the first target object and the electronic device, and determines the second resolution based on the first distance.

4. The method according to claim 3, characterized in that The determining the second resolution based on the pixel height comprises: If the pixel height is less than a first height threshold, the electronic device determines that the second resolution is greater than the first resolution; or, If the pixel height is greater than a second height threshold, the electronic device determines that the second resolution is smaller than the first resolution; or, If the pixel height is greater than or equal to the first height threshold and less than or equal to the second height threshold, the electronic device determines that the second resolution is equal to the first resolution.

5. The method according to claim 3, characterized in that: The determining the second resolution based on the area ratio includes: If the area ratio is less than a first preset value, the electronic device determines that the second resolution is greater than the first resolution; or, If the area ratio is greater than a second preset value, the electronic device determines that the second resolution is smaller than the first resolution; or, If the area ratio is greater than or equal to the first preset value and less than or equal to the second preset value, the electronic device determines that the second resolution is equal to the first resolution.

6. The method according to claim 3, characterized in that The determining the second resolution based on the first distance includes: If the first distance is less than a first distance threshold, the electronic device determines that the second resolution is less than the first resolution; or, If the first distance is greater than a second distance threshold, the electronic device determines that the second resolution is greater than the first resolution; or, If the first distance is greater than or equal to the first distance threshold and less than or equal to the second distance threshold, the electronic device determines that the second resolution is equal to the first resolution.

7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: If the electronic device does not recognize that the Nth frame image includes the first target object, the electronic device captures the N+1th frame image at a preset resolution; The preset resolution is related to QVGA.

8. An electronic device, characterized in that: include: memory and one or more processors; The memory stores computer program code, which includes computer instructions. When the computer instructions are executed by the processor, the electronic device executes the method as claimed in any one of claims 1 to 7.

9. A chip system, characterized in that: The chip system includes: at least one processor and an interface; The interface is used to receive instructions and transmit them to the at least one processor; the at least one processor executes the instructions so that the electronic device executes the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The method comprises computer instructions, which, when executed on an electronic device, cause the electronic device to execute the method as claimed in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method for identifying identification code and mobile terminal

    CN109033913A

  • Moving target identification method and shooting equipment

    CN115731258A

  • Adaptive object detection and recognition

    US20190213420A1