A method of human-computer interaction and an electronic device
By detecting the position and time interval of the wake word in the voice command, and adjusting the device posture in combination with the direction of the sound source and the direction of the gaze, the problem of the wake word interrupting the task is solved, and the continuity of human-computer interaction and user experience are improved.
Patent Information
- Application Number
- CN202110381295.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-08
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2041-04-08
AI Technical Summary
In existing technologies, wake words may interrupt the current task during human-computer interaction, resulting in discontinuous human-computer dialogue and affecting user experience.
By detecting the position and time interval of the wake word in the voice command, the system can determine the user's intent, avoid the wake word from interrupting the current interaction flow, and adjust the device posture according to the direction of the sound source and the direction of the gaze to improve the continuity of the interaction.
This technology avoids interrupting the current task during voice interaction by using wake words, improving the coherence of human-computer dialogue and user experience, and enhancing the interaction process between the device and the user to meet user expectations.
Smart Images

Figure CN115206308B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electronics, and in particular to a human-computer interaction method and an electronic device. BACKGROUND
[0002] With the development of technology, more and more electronic devices support "human-computer interaction", or "voice interaction". Human-computer interaction has gradually become a way for users to convey intentions and control electronic devices. Human-computer interaction is mainly controlled by voice instructions of a user, thereby liberating the user's hands and facilitating the user to control the electronic device.
[0003] Before the user and the electronic device perform human-computer interaction, the electronic device can be woken up by a "wake-up word". After the electronic device is woken up, a response of successful wake-up can be provided for the user, and voice instructions of the user are collected and automatic speech recognition (ASR) is performed. In the speech recognition process after the electronic device is woken up, if the voice instructions obtained include the wake-up word, the wake-up word can interrupt the current human-computer interaction process, and the collection of voice instructions of the user and the speech recognition are restarted. The process of interrupting the current human-computer interaction can not be expected by the user, that is, the wake-up word directly interrupts the task being executed, so that the electronic device restarts the collection of voice instructions of the user. This can cause the human-computer dialogue to be incoherent, affect the use process of the user, and reduce the experience of human-computer interaction. SUMMARY
[0004] The present application provides a human-computer interaction method and an electronic device. The electronic device can include a mobile phone, a robot, a tablet, a computer, and the like, which have a voice recognition function. The method can provide a coherent and immersive experience for the user, and improve the visual experience of the user.
[0005] In a first aspect, a human-computer interaction method is provided. The method includes receiving a wake-up word issued by a user, starting a voice recognition function of an electronic device in response to the wake-up word, obtaining a first voice instruction of the user, determining a first time period occupied by the wake-up word in the first voice instruction when it is detected that the first voice instruction includes the wake-up word, removing the wake-up word in the first time period, recognizing a target voice instruction in the first voice instruction except the wake-up word, and responding to the target voice instruction.
[0006] In a possible scenario, taking the user waking up the mobile phone by the wake-up word "Xiaoyi Xiaoyi" as an example, after the mobile phone is woken up, the mobile phone enters a state of listening to a voice instruction, if the voice instruction issued by the user again includes the wake-up word "Xiaoyi Xiaoyi", the wake-up word can interrupt the current human-computer interaction process and re-enter the next human-computer interaction process, which may not be expected by the user, that is, the wake-up word directly interrupts the task being executed, so that the mobile phone needs to start collecting the voice instruction of the user again, which causes the human-computer dialogue to be incoherent, affects the use process of the user, and reduces the experience of human-computer interaction.
[0007] Through the method, in the voice interaction process between the user and the electronic device, after the user wakes up the electronic device by the wake-up word, if the voice instruction issued by the user again includes the wake-up word, the method can avoid the wake-up word in the voice instruction interrupting the current interaction process, thereby avoiding directly interrupting the task being executed by the electronic device and starting the process of collecting the voice instruction of the user again, ensuring the coherence of the human-computer dialogue and improving the user experience.
[0008] It should be understood that the automatic speech recognition (ASR) module of the mobile phone is not always in a working state, when the user issues a voice instruction, the ASR module of the mobile phone is closed; or when the mobile phone is answering the user, the ASR module is closed, to avoid collecting the voice of the mobile phone and interfering with the collection and recognition of the voice instruction of the user. Through the wake-up word, after the mobile phone is woken up, whether the ASR module is in an open state can be detected first, if the ASR is in a closed state of dormancy or non-working, the ASR module can be triggered to be opened, that is, the voice recognition function of the electronic device is opened.
[0009] Optionally, when the mobile phone first acquires and recognizes the wake-up word "Xiaoyi Xiaoyi", if it is determined that the mobile phone is currently in a state of opening the ASR module, the current wake-up can be ignored, and the current dialogue process is continued.
[0010] With reference to the first aspect, in some implementations of the first aspect, the first time period is a tail time period, a middle time period, or a start time period of the time period corresponding to the first voice instruction.
[0011] After the mobile phone is woken up, the first voice instruction of the user is monitored. When it is detected that the first voice instruction again includes the wake-up word "Xiaoyi Xiaoyi", it can be first judged that the position of the wake-up word "Xiaoyi Xiaoyi" in the first voice instruction. The position can mainly include the first position of the first voice instruction, the middle of the first voice instruction, and the end of the first voice instruction. For example, the first voice instruction issued by the user in the case of including the wake-up word can be "imitate the call of a cow, Xiaoyi Xiaoyi" (the wake-up word is at the end of the first voice instruction), "imitate the call of an animal, Xiaoyi Xiaoyi, imitate the call of a cow" (the wake-up word is in the middle of the first voice instruction), or "Xiaoyi Xiaoyi, imitate the call of a cow" (the wake-up word is at the first position of the first voice instruction).
[0012] In combination with the first aspect and the above implementation manners, in some implementation manners of the first aspect, when the first time period is the end time period of the time period corresponding to the first voice instruction, the method further includes: detecting the time interval from the wake-up word to the voice instruction closest to the wake-up word in the first voice instruction; and when the time interval is greater than or equal to a first preset value, pausing the current dialogue process and re-enabling the voice recognition function of the electronic device in response to the wake-up word, so that the electronic device acquires a second voice instruction.
[0013] It should be understood that the first preset value can be used to judge whether the current user wants to interrupt the dialogue process. For example, when the first voice instruction issued by the user is "imitate the call of a cow, Xiaoyi Xiaoyi", the wake-up word is at the end of the voice instruction. The voice instruction closest to the wake-up word "Xiaoyi Xiaoyi" is "imitate the call of a cow". The mobile phone can judge the mother of the user who issues the wake-up word "Xiaoyi Xiaoyi" according to the time interval between "imitate the call of a cow" and "Xiaoyi Xiaoyi". When the time interval between "imitate the call of a cow" and the first "Xiao" of "Xiaoyi Xiaoyi" is less than the first preset value, it can be judged that the user may only take the wake-up word "Xiaoyi Xiaoyi" as part of a catchword and want to continue the current dialogue process without switching to a new dialogue process.
[0014] Optionally, the mobile phone can record the time information of the wake-up word "Xiaoyi Xiaoyi" in the first voice instruction according to the first voice instruction. The recording and marking rules of the time information are not limited in the embodiments of the present application. For example, if the initial wake-up of the mobile phone is taken as the starting time, the time period in which the wake-up word appears again in the first voice instruction is t1-t2; if the initial wake-up of the mobile phone is taken as the starting time, the time period in which the wake-up word appears again in the first voice instruction is T1-T2. The position of the wake-up word in the first voice instruction can be determined according to the time information.
[0015] In a second aspect, a method for human-computer interaction is provided, which includes: obtaining a first voice instruction of a user, detecting a sound source direction of the first voice instruction according to the first voice instruction; determining a first angle between the sound source direction of the first voice instruction and a first line-of-sight direction currently faced by an electronic device; when the first angle is greater than or equal to a first preset angle, determining a second angle between the sound source direction of the first voice instruction and a sound source direction of a second voice instruction, the second voice instruction being a voice instruction closest to the first voice instruction and issued by the user before the first voice instruction; and when the second angle is less than or equal to a second preset angle, the electronic device responding to the first voice instruction.
[0016] In another possible scenario, some electronic devices can have the capability of sound source positioning or the function of image acquisition by a camera, such as robots, etc. When the robot is woken up by a wake-up word, the direction where the user is located can be determined according to the sound source positioning function, and the camera with the function of image acquisition is turned to directly go to the direction or position where the user is located according to the sound source positioning. In this process, the direction where the user is located can have a large error in judgment due to the reflection of sound by a wall, etc. When such a large error occurs, the phenomenon that the device is not facing the person after being turned can occur.
[0017] Through the above method, the wake-up process of the robot is more in line with the expectation of the person. When the included angle θ between the sound source direction of the voice instruction of the user and the line-of-sight direction currently faced by the robot is greater than or equal to the first preset angle and the interactive intention of the user is strong, the robot can be determined to automatically turn to the user; when the included angle θ between the sound source direction of the voice instruction of the user and the line-of-sight direction currently faced by the robot is greater than or equal to the first preset angle and the interactive intention of the user is low, the robot can also turn back, and in this process, the interactive process between the user and the robot will not be interrupted, bringing a better human-computer interaction experience to the user.
[0018] It should be understood that when the included angle θ between the sound source direction of the first voice instruction and the line-of-sight direction is greater than or equal to the first preset angle, it can be considered that the user who issues the voice instruction and the robot are not in a face-to-face position relationship, or in other words, the user who issues the voice instruction is not within the central area range of the image collected by the robot. The range corresponding to the central area is not limited in the embodiments of the present application.
[0019] It should also be understood that the "previous voice instruction" here is the closest voice instruction before the first voice instruction. Alternatively, the "previous voice instruction" can be a wake-up word instruction of the user, for example: Xiaoyi Xiaoyi. Or the "previous voice instruction" is other voice instructions after the wake-up word, for example: please imitate the mooing sound of a cow. The embodiments of the present application do not limit this.
[0020] With reference to the second aspect, in some implementations of the second aspect, the method further includes: detecting a time interval between the first voice instruction and the second voice instruction; and calling a turning execution function to turn the electronic device to face or infinitely approach a sound source direction of the first voice instruction when the time interval is less than or equal to a second preset value.
[0021] Optionally, the display conditions of "an included angle between the sound source direction of the first voice instruction and a sound source direction of a previous voice instruction is greater than or equal to a second preset angle" and "a time interval between two voice instructions is greater than or equal to a second preset value" can meet any one or both, and the turning execution function is called to change the direction of the robot. The embodiments of the present application do not limit this.
[0022] With reference to the second aspect and the above implementations, in some implementations of the second aspect, the method further includes: collecting a first image of the electronic device in the first line-of-sight direction; and calling the turning execution function to turn the electronic device to face or infinitely approach the sound source direction of the first voice instruction when the first image includes the user and a third angle between a line-of-sight direction of the user and the sound source direction of the first voice instruction is less than or equal to a third preset angle.
[0023] Optionally, the robot can also collect images through a camera and detect a direction at which eyes of the user gaze in the collected images to estimate the interaction intention of the user.
[0024] With reference to the second aspect and the above implementations, in some implementations of the second aspect, the method further includes: the electronic device collects a second image facing or infinitely approaching the sound source direction of the first voice instruction; and turning the electronic device to restore to the first line-of-sight direction when the second image does not include the user or a fourth angle between the line-of-sight direction of the user and a current second line-of-sight direction of the electronic device is greater than a fourth preset angle.
[0025] In summary, in the process of voice interaction between the user and the electronic device, after the user wakes up the electronic device through the wake-up word, if the wake-up word is included in the voice instruction or the answer of the user to the electronic device again, the method can avoid the wake-up word in the voice instruction from interrupting the current interaction process, thereby avoiding directly interrupting the task being executed by the electronic device, restarting the process of collecting the voice instruction of the user, ensuring the coherence of the human-computer dialogue, and improving the user experience.
[0026] In addition, for electronic devices such as robots having the capability of sound source positioning, the method provided by the embodiments of the present application can determine whether to deflect according to the sound source direction of the voice instruction, and estimate the interaction intention of the user according to the collected images and the like, and then perform voice interaction with the user more accurately. Specifically, when the included angle θ between the sound source direction of the voice instruction of the user and the line-of-sight direction in which the robot currently faces is greater than or equal to a first preset angle and the interaction intention of the user is strong, the robot can be determined to automatically turn to the user; when the included angle θ between the sound source direction of the voice instruction of the user and the line-of-sight direction in which the robot currently faces is greater than or equal to the first preset angle and the interaction intention of the user is low, the robot can also turn back, and in this process, the interaction process between the user and the robot will not be interrupted, and a better human-computer interaction experience is brought to the user.
[0027] In a third aspect, an electronic device is provided, comprising: one or more processors; one or more memories; a module installed with a plurality of application programs; the memory stores one or more programs, when the one or more programs are executed by the processor, the electronic device is caused to perform the following steps: receiving a wake-up word issued by a user, in response to the wake-up word, starting the voice recognition function of the electronic device; obtaining a first voice instruction of the user, when it is detected that the first voice instruction includes the wake-up word, determining a first time period occupied by the wake-up word in a time period corresponding to the first voice instruction; removing the wake-up word in the first time period, recognizing a target voice instruction in the first voice instruction except the wake-up word; and responding to the target voice instruction.
[0028] In combination with the third aspect, in some implementations of the third aspect, the first time period is the end time period, the middle time period or the start time period of the time period corresponding to the first voice instruction.
[0029] In combination with the third aspect and the above implementations, in some implementations of the third aspect, when the first time period is the end time period of the time period corresponding to the first voice instruction, the electronic device can further perform the following steps: detecting the voice instruction in the first voice instruction closest to the wake-up word, and the time interval to the wake-up word; when the time interval is greater than or equal to a first preset value, pausing the current dialogue process and re-starting the voice recognition function of the electronic device in response to the wake-up word, so that the electronic device obtains a second voice instruction.
[0030] In a fourth aspect, an electronic device is provided, comprising: a camera; one or more processors; one or more memories; a module having a plurality of applications installed; the memory stores one or more programs, which when executed by the processor, cause the electronic device to perform the following steps: obtaining a first voice instruction of a user, detecting a sound source direction of the first voice instruction according to the first voice instruction; determining a first angle between the sound source direction of the first voice instruction and a first line-of-sight direction in which the electronic device currently faces; when the first angle is greater than or equal to a first preset angle, determining a second angle between the sound source direction of the first voice instruction and a sound source direction of a second voice instruction, the second voice instruction being a voice instruction closest to the first voice instruction and issued by the user before the first voice instruction; when the second angle is less than or equal to a second preset angle, the electronic device responds to the first voice instruction.
[0031] With reference to the fourth aspect, in some implementations of the fourth aspect, when the one or more programs are executed by the processor, the electronic device is caused to perform the following steps: detecting a time interval between the first voice instruction and the second voice instruction; when the time interval is greater than or equal to a second preset value, calling a turning execution function to turn the electronic device to face or infinitely approach the sound source direction of the first voice instruction.
[0032] With reference to the fourth aspect and the above implementation, in some implementations of the fourth aspect, when the one or more programs are executed by the processor, the electronic device is caused to perform the following steps: capturing a first image of the electronic device in the first line-of-sight direction; when the first image includes the user and a third angle between a line-of-sight direction of the user and the sound source direction of the first voice instruction is less than or equal to a third preset angle, calling a turning execution function to turn the electronic device to face or infinitely approach the sound source direction of the first voice instruction.
[0033] With reference to the fourth aspect and the above implementation, in some implementations of the fourth aspect, when the one or more programs are executed by the processor, the electronic device is caused to perform the following steps: capturing a second image facing or infinitely approaching the sound source direction of the first voice instruction; when the second image does not include the user or a fourth angle between a line-of-sight direction of the user and a second line-of-sight direction of the electronic device is greater than a fourth preset angle, turning the electronic device back to the first line-of-sight direction.
[0034] In a fifth aspect, the present application provides an apparatus, which is included in an electronic device, and the apparatus has functions to implement the behaviors of the electronic device in the above aspects and possible implementation manners of the above aspects. The functions can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions. For example, a display module or unit, a detection module or unit, a processing module or unit, and the like.
[0035] In a sixth aspect, the present application provides an electronic device, comprising: a touch display screen, wherein the touch display screen comprises a touch-sensitive surface and a display; one or more audio devices; a camera; one or more processors; a memory; a plurality of application programs; and one or more computer programs. The one or more computer programs are stored in the memory, and the one or more computer programs comprise instructions. When the instructions are executed by the electronic device, the electronic device performs the method of human-computer interaction in any one of the above aspects and any one of the possible implementations.
[0036] In a seventh aspect, the present application provides an electronic device, comprising one or more processors and one or more memories. The one or more memories are coupled to the one or more processors, and the one or more memories are configured to store computer program codes, and the computer program codes comprise computer instructions. When the one or more processors execute the computer instructions, the electronic device performs the method of human-computer interaction in any one of the above aspects and any one of the possible implementations.
[0037] In an eighth aspect, the present application provides a computer-readable storage medium, comprising computer instructions. When the computer instructions are run on an electronic device, the electronic device performs the method of human-computer interaction in any one of the above aspects and any one of the possible implementations.
[0038] In a ninth aspect, the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device performs the method of human-computer interaction in any one of the above aspects and any one of the possible implementations. BRIEF DESCRIPTION OF DRAWINGS
[0039] Figure 1 FIG. 1 is a structural schematic diagram of an example of an electronic device provided by an embodiment of the present application.
[0040] Figure 2 FIG. 2 is a software structural block diagram of the electronic device of the embodiment of the present application.
[0041] Figure 3 FIG. 3 is a schematic diagram of a graphical user interface of an example of a human-computer interaction process.
[0042] Figure 4 FIG. 4 is a schematic flowchart of an example of a method of human-computer interaction provided by an embodiment of the present application.
[0043] Figure 5 is a schematic diagram of a scene of human-computer interaction provided by an embodiment of the present application.
[0044] Figure 6 is a schematic flow chart of a method of human-computer interaction provided by an embodiment of the present application. DETAILED DESCRIPTION
[0045] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, “ / ” represents the meaning of or, for example, A / B can represent A or B; “and / or” in this document only describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, in the description of the embodiments of the present application, “multiple” means two or more than two.
[0046] Hereinafter, the terms “first” and “second” are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with “first” and “second” can explicitly or implicitly include one or more of the features.
[0047] The method of human-computer interaction provided by the embodiments of the present application can be applied to electronic devices such as mobile phones, tablet computers, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, notebook computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), etc. The embodiments of the present application do not make any limitation on the specific type of electronic device.
[0048] Exemplarily, Figure 1Fig. 1 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. The electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a loudspeaker 170A, a receiver 170B, a microphone 170C, a headset interface 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0049] It can be understood that the structure illustrated by the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 can include more or fewer components than illustrated, or combine certain components, or split certain components, or different arrangement of components. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.
[0050] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices, or can be integrated in one or more processors.
[0051] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching and executing instructions.
[0052] The processor 110 can also have a memory that stores instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can hold instructions or data that the processor 110 has just used or is using repeatedly. If the processor 110 needs to use the instructions or data again, it can call them directly from the memory. This avoids repeated access and reduces the latency of the processor 110, thus improving the efficiency of the system.
[0053] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0054] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 can include multiple sets of I2C buses. The processor 110 can be coupled to the touch sensor 180K, the charger, the flash, the camera 193, etc. through different I2C bus interfaces. For example, the processor 110 can be coupled to the touch sensor 180K through an I2C interface, so that the processor 110 and the touch sensor 180K communicate through the I2C bus interface to realize the touch function of the electronic device 100.
[0055] The I2S interface can be used for audio communication. In some embodiments, the processor 110 can include multiple sets of I2S buses. The processor 110 can be coupled with the audio module 170 through the I2S buses to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can deliver audio signals to the wireless communication module 160 through the I2S interface to enable the function of answering a phone call through a Bluetooth earphone.
[0056] The PCM interface can also be used for audio communication to sample, quantize, and encode analog signals. In some embodiments, the audio module 170 can be coupled with the wireless communication module 160 through a PCM bus interface. In some embodiments, the audio module 170 can also deliver audio signals to the wireless communication module 160 through the PCM interface to enable the function of playing music through a Bluetooth earphone. Both the I2S interface and the PCM interface can be used for audio communication.
[0057] The UART interface is a universal serial bus for asynchronous communication. The bus can be a bidirectional communication bus. It converts data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is usually used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 through the UART interface to enable Bluetooth functionality. In some embodiments, the audio module 170 can deliver audio signals to the wireless communication module 160 through the UART interface to enable the function of playing music through a Bluetooth earphone.
[0058] The MIPI interface can be used to connect the processor 110 and peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), and the like. In some embodiments, the processor 110 and the camera 193 communicate through the CSI interface to enable the camera function of the electronic device 100. The processor 110 and the display screen 194 communicate through the DSI interface to enable the display function of the electronic device 100.
[0059] The GPIO interface can be configured through software. The GPIO interface can be configured as a control signal or as a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 and the camera 193, the display screen 194, the wireless communication module 160, the audio module 170, the sensor module 180, and the like. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, and the like.
[0060] The USB interface 130 is an interface conforming to the USB standard specification, and can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device 100, and can also be used to transmit data between the electronic device 100 and a peripheral device. It can also be used to connect a headset to play audio through the headset. The interface can also be used to connect other electronic devices, such as AR devices, etc.
[0061] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a structural limitation on the electronic device 100. In some other embodiments of the present application, the electronic device 100 can also use different interface connection modes or combinations of multiple interface connection modes in the above embodiments.
[0062] The charging management module 140 is used to receive charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from a wired charger through the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input through the wireless charging coil of the electronic device 100. The charging management module 140 can charge the battery 142 while also supplying power to the electronic device through the power management module 141.
[0063] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to supply power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, and the wireless communication module 160, etc. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, battery health status (leakage, impedance), etc. In some other embodiments, the power management module 141 can also be disposed in the processor 110. In some other embodiments, the power management module 141 and the charging management module 140 can also be disposed in the same device.
[0064] The wireless communication function of the electronic device 100 can be realized through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, etc.
[0065] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in combination with a tuning switch.
[0066] The mobile communication module 150 can provide a solution for wireless communication including 2G / 3G / 4G / 5G, etc. applied to the electronic device 100. The mobile communication module 150 can include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves by the antenna 1, and perform filtering, amplification, etc. on the received electromagnetic waves, and transfer the same to the modem processor for demodulation. The mobile communication module 150 can also amplify signals modulated by the modem processor, and radiate the same as electromagnetic waves through the antenna 1. In some embodiments, at least part of the functional modules of the mobile communication module 150 can be disposed in the processor 110. In some embodiments, at least part of the functional modules of the mobile communication module 150 can be disposed in the same device as at least part of the modules of the processor 110.
[0067] The modem processor can include a modulator and a demodulator. The modulator is configured to modulate a low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is configured to demodulate a received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. The low-frequency baseband signal processed by the baseband processor is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the microphone 170B, etc.), or displays an image or a video through the display screen 194. In some embodiments, the modem processor can be a separate device. In other embodiments, the modem processor can be independent of the processor 110, and disposed in the same device as the mobile communication module 150 or other functional modules.
[0068] The wireless communication module 160 can provide a solution for wireless communication including wireless local area networks (WLAN) (e.g., wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc. applied to the electronic device 100. The wireless communication module 160 can be one or more devices that integrate at least one communication processing module. The wireless communication module 160 receives an electromagnetic wave via the antenna 2, frequency-modulates and filters the electromagnetic wave signal, and transmits the processed signal to the processor 110. The wireless communication module 160 can also receive a signal to be transmitted from the processor 110, frequency-modulate it, amplify it, and radiate it as an electromagnetic wave via the antenna 2.
[0069] In some embodiments, the antenna 1 and the mobile communication module 150 of the electronic device 100 are coupled, and the antenna 2 and the wireless communication module 160 are coupled, so that the electronic device 100 can communicate with a network and other devices through wireless communication technology. The wireless communication technology can include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The GNSS can include a global positioning system (GPS), a global navigation satellite system (GLONASS), a beidu navigation satellite system (BDS), a quasi-zenith satellite system (QZSS), and / or a satellite based augmentation systems (SBAS).
[0070] The electronic device 100 implements a display function through a GPU, a display screen 194, and an application processor, etc. The GPU is a microprocessor for image processing, which is connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs, which execute program instructions to generate or change display information.
[0071] The display screen 194 is configured to display images, videos, and the like. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light emitting diode (QLED), or the like. In some embodiments, the electronic device 100 can include one or N display screens 194, where N is a positive integer greater than 1.
[0072] The electronic device 100 can implement the photographing function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor.
[0073] The ISP is configured to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing to convert it into an image visible to the naked eye. The ISP can also optimize the noise, brightness, and skin color of the image. The ISP can also optimize the exposure, color temperature, and other parameters of the shooting scene. In some embodiments, the ISP can be disposed in the camera 193.
[0074] The camera 193 is configured to capture still images or videos. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV, or the like format. In some embodiments, the electronic device 100 can include one or N cameras 193, where N is a positive integer greater than 1.
[0075] The digital signal processor is used to process digital signals, in addition to being able to process digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0076] The video codec is used to compress or decompress digital video. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple encoding formats, such as: moving picture experts group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.
[0077] The NPU is a neural-network (NN) calculation processor, which can quickly process input information by drawing on the structure of a biological neural network, such as drawing on the transmission mode between human brain neurons, and can also constantly self-learn. Through the NPU, the electronic device 100 can realize intelligent cognition applications such as image recognition, face recognition, voice recognition, text understanding, etc.
[0078] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to realize data storage functions. For example, music, video, etc. Files are saved in the external memory card.
[0079] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various function applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), etc. The data storage area can store data created during the use of the electronic device 100 (such as audio data, a phonebook, etc.), etc. In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0080] The electronic device 100 can realize audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the earphone interface 170D, and the application processor, etc. For example, music playing, recording, etc.
[0081] The audio module 170 is configured to convert digital audio information into an analog audio signal output, and to convert an analog audio input into a digital audio signal. The audio module 170 can also be configured to encode and decode audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some of the functions of the audio module 170 can be disposed in the processor 110.
[0082] The speaker 170A, also referred to as a "loudspeaker", is configured to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or listen to a hands-free call through the speaker 170A.
[0083] The receiver 170B, also referred to as a "earpiece", is configured to convert an audio electrical signal into a sound signal. When the electronic device 100 receives a call or a voice message, the user can listen to the voice by holding the receiver 170B close to the ear.
[0084] The microphone 170C, also referred to as a "microphone", "voice microphone", is configured to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak into the microphone 170C by holding the mouth close to the microphone 170C, and input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, in addition to collecting sound signals, the noise reduction function can also be realized. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C, in addition to collecting sound signals, noise reduction, it can also identify the source of the sound, realize the function of directional recording, etc.
[0085] The earphone interface 170D is configured to connect a wired earphone. The earphone interface 170D can be a USB interface 130, or a 3.5mm open mobile terminal platform (OMTP) standard interface, a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0086] The pressure sensor 180A is configured to sense a pressure signal and convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be disposed on the display screen 194. The pressure sensor 180A can be of various types, such as a resistive pressure sensor, an inductive pressure sensor, a capacitive pressure sensor, etc. The capacitive pressure sensor can include at least two parallel plates of conductive material. When a force is applied to the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device 100 determines the intensity of the pressure according to the change in capacitance. When a touch operation is applied to the display screen 194, the electronic device 100 detects the intensity of the touch operation according to the pressure sensor 180A. The electronic device 100 can also calculate the position of the touch according to the detection signal of the pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation instructions. For example, when a touch operation with a touch operation intensity less than a first pressure threshold is applied to a short message application icon, an instruction to view short messages is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold is applied to the short message application icon, an instruction to create a new short message is executed.
[0087] The gyroscope sensor 180B can be configured to determine the motion attitude of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., x, y, and z axes) can be determined by the gyroscope sensor 180B. The gyroscope sensor 180B can be used for anti-shake photography. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of shaking of the electronic device 100, calculates the distance that the lens module needs to compensate according to the angle, and lets the lens offset the shaking of the electronic device 100 by reverse movement to achieve anti-shake. The gyroscope sensor 180B can also be used for navigation and motion sensing game scenarios.
[0088] The barometric pressure sensor 180C is configured to measure air pressure. In some embodiments, the electronic device 100 calculates the altitude, assists positioning and navigation by using the air pressure value measured by the barometric pressure sensor 180C.
[0089] The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can detect the opening and closing of a flip cover by using the magnetic sensor 180D. In some embodiments, when the electronic device 100 is a flip phone, the electronic device 100 can detect the opening and closing of the flip cover according to the magnetic sensor 180D. Further, according to the detected opening and closing state of the cover or the flip cover, the electronic device 100 can set a feature such as automatic unlocking of the flip cover.
[0090] The acceleration sensor 180E can detect the acceleration of the electronic device 100 in various directions (generally three axes). When the electronic device 100 is stationary, the acceleration sensor 180E can detect the magnitude and direction of gravity. The acceleration sensor 180E can also be used to identify the attitude of the electronic device, and can be applied to landscape / portrait screen switching and pedometer applications.
[0091] Distance sensor 180F is configured to measure distance. Electronic device 100 can measure distance by infrared or laser. In some embodiments, electronic device 100 can utilize distance sensor 180F to measure distance for fast focusing when taking a picture.
[0092] Proximity light sensor 180G can include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The light emitting diode can be an infrared light emitting diode. Electronic device 100 emits infrared light outwardly through the light emitting diode. Electronic device 100 detects infrared reflected light from nearby objects using the photodiode. When sufficient reflected light is detected, electronic device 100 can determine that there is an object near electronic device 100. When insufficient reflected light is detected, electronic device 100 can determine that there is no object near electronic device 100. Electronic device 100 can utilize proximity light sensor 180G to detect that a user is holding electronic device 100 close to the ear for a phone call, so as to automatically turn off the screen to save power. Proximity light sensor 180G can also be used for automatic unlocking and locking of the screen in a holster mode or a pocket mode.
[0093] Ambient light sensor 180L is configured to sense ambient light brightness. Electronic device 100 can adaptively adjust the brightness of display 194 according to the sensed ambient light brightness. Ambient light sensor 180L can also be used to automatically adjust white balance when taking a picture. Ambient light sensor 180L can also cooperate with proximity light sensor 180G to detect whether electronic device 100 is in a pocket to prevent accidental touch.
[0094] Fingerprint sensor 180H is configured to collect a fingerprint. Electronic device 100 can utilize the collected fingerprint characteristics to implement fingerprint unlocking, access application lock, fingerprint photographing, fingerprint answering a call, and the like.
[0095] Temperature sensor 180J is configured to detect temperature. In some embodiments, electronic device 100 utilizes the temperature detected by temperature sensor 180J to implement a temperature handling strategy. For example, when the temperature reported by temperature sensor 180J exceeds a threshold, electronic device 100 implements a performance reduction of a processor located near temperature sensor 180J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is lower than another threshold, electronic device 100 heats battery 142 to avoid abnormal shutdown of electronic device 100 caused by low temperature. In other embodiments, when the temperature is lower than yet another threshold, electronic device 100 implements a boost of output voltage of battery 142 to avoid abnormal shutdown caused by low temperature.
[0096] Touch sensor 180K, also referred to as "touch panel". Touch sensor 180K can be disposed on display screen 194, and touch sensor 180K and display screen 194 together form a touch screen, also referred to as "touch panel". Touch sensor 180K is configured to detect touch operations applied to or near the touch sensor 180K. The touch sensor 180K can transmit the detected touch operation to the application processor to determine the touch event type. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K can also be disposed on the surface of electronic device 100, which is different from the position of display screen 194.
[0097] Bone conduction sensor 180M can obtain vibration signals. In some embodiments, bone conduction sensor 180M can obtain vibration signals of the human body's vocal vibration bone block. Bone conduction sensor 180M can also contact the human body pulse to receive blood pressure pulsation signals. In some embodiments, bone conduction sensor 180M can also be disposed in a headset to form a bone conduction headset. Audio module 170 can analyze voice signals based on the vibration signals of the vocal vibration bone block obtained by the bone conduction sensor 180M to realize voice functions. The application processor can analyze heart rate information based on the blood pressure pulsation signals obtained by the bone conduction sensor 180M to realize heart rate detection functions.
[0098] Keys 190 include power on / off keys, volume keys, and the like. Keys 190 can be mechanical keys. They can also be touch keys. Electronic device 100 can receive key input and generate key signal input related to user settings and function control of electronic device 100.
[0099] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts and also for touch vibration feedback. For example, touch operations applied to different applications (such as taking pictures, playing audio, etc.) can correspond to different vibration feedback effects. Touch operations applied to different regions of display screen 194 can also correspond to different vibration feedback effects. Different application scenarios (such as time reminders, received messages, alarms, games, etc.) can also correspond to different vibration feedback effects. Touch vibration feedback effects can also be customizable.
[0100] Indicator 192 can be an indicator light, which can be used to indicate charging status, power changes, and also to indicate messages, missed calls, notifications, and the like.
[0101] The SIM card interface 195 is configured to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation with the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, and the like. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 195 can be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external storage cards. The electronic device 100 interacts with a network through the SIM card to achieve functions such as call and data communication. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.
[0102] The software system of the electronic device 100 can use a layered architecture, an event-driven architecture, a microkernel architecture, a microservice architecture, or a cloud architecture. In this embodiment of this application, the software structure of the electronic device 100 is exemplarily described by taking a layered architecture as an example.
[0103] Figure 2 FIG. 1 is a software structure block diagram of the electronic device 100 according to an embodiment of this application. The layered architecture divides the software into several layers, each of which has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the system is divided into four layers, from top to bottom, an application layer, an application framework layer, an Android runtime and system library, and a kernel layer. The application layer can include a series of application packages. As shown in FIG. 1, the application packages can include camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, and the like.
[0104] As shown in FIG. 1, the application framework layer can include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like. Figure 2
[0105] As shown in FIG. 1, the application framework layer can include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.
[0106] As shown in FIG. 1, the application framework layer can include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like. Figure 2
[0107] The window manager is used to manage windows programs. The window manager can acquire the display screen size, determine whether the screen has a status bar, or participate in executing a lock screen, screen capture, and the like.
[0108] The content provider is used to store and acquire data, and make the data accessible to the application program. The stored data can include video data, image data, audio data, and the like, and can also include call record data of outgoing and incoming calls, user browsing history and bookmark data, and the like, which will not be described herein.
[0109] The view system includes visual controls, such as a control for displaying text, a control for displaying pictures, and the like. The view system can be used to build an application program. A display interface can be composed of one or more views. For example, a display interface including a short message notification icon can include a view for displaying text and a view for displaying pictures.
[0110] The phone manager is used to provide the communication function of the electronic device 100. For example, management of a call state (including connection and hang-up of a phone, and the like).
[0111] The resource manager provides various resources for the application program, such as localized strings, icons, pictures, layout files, video files, and the like.
[0112] The notification manager enables the application program to display notification information in the status bar of the screen, and can be used to convey a message to the user. The notification information can automatically disappear after a short stay in the status bar without the need for the user to perform a closing operation and the like. For example, the notification manager can inform the user of a download completion message and the like. The notification manager can also be a notification in the form of a chart or a scroll bar text appearing in the top status bar of the system, such as a notification of an application program running in the background, or can be a notification in the form of a dialog window appearing on the screen, such as a text information prompt in the status bar, or can control the electronic device to emit a prompt sound, vibrate, and flash an indicator light, and the like, which will not be described herein.
[0113] The runtime includes a core library and a virtual machine. The runtime is responsible for scheduling and management of the Android system.
[0114] The core library includes two parts: one part is a function function that the java language needs to call, and the other part is the core library of Android.
[0115] The application program layer and the application program framework layer run in the virtual machine. The virtual machine executes the java files of the application program layer and the application program framework layer into binary files. The virtual machine is used to perform functions such as management of the life cycle of an object, stack management, thread management, security and exception management, and garbage collection.
[0116] The system library can include a plurality of functional modules. For example, a surface manager, media libraries, a three dimensional (3D) graphics processing library (e.g., OpenGL ES), a two dimensional (2D) graphics engine, etc.
[0117] The surface manager is used to manage the display subsystem of the electronic device, and provides a fusion of 2D and 3D layers for a plurality of application programs.
[0118] The media libraries support a plurality of commonly used audio, video format playback and recording, and static image files, etc. The media libraries can support a plurality of audio and video encoding formats, for example, MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0119] The three dimensional graphics processing library is used to implement three dimensional graphics drawing, image rendering, synthesis, and layer processing, etc.
[0120] The two dimensional graphics engine is a drawing engine for two dimensional drawing.
[0121] The kernel layer is a layer between hardware and software. The kernel layer at least includes a display driver, a camera driver, an audio driver, and a sensor driver.
[0122] For ease of understanding, the following embodiments of the present application will take an electronic device having the structure shown in Figure 1 and Figure 2 as an example, and the method of human-computer interaction provided by the embodiments of the present application will be specifically described in combination with the drawings and application scenarios.
[0123] First, before introducing the method of human-computer interaction provided by the embodiments of the present application, several possible application scenarios are listed.
[0124] In one possible scenario, the method of human-computer interaction provided by the embodiments of the present application can be applied to a scenario including a single electronic device. For example, the electronic device can be different electronic devices such as a mobile phone, a tablet, a smart screen, etc. introduced above in combination with the structure shown in Figure 1 The embodiments of the present application do not limit this. In the following, the method of displaying a human-computer interaction instruction provided by the present application will be introduced in detail taking a mobile phone as an example.
[0125] Figure 3 is a schematic diagram of a graphical user interface (GUI) of a human-computer interaction process, wherein Figure 3Figure (a) shows the phone in unlocked mode, with the phone's screen display system showing the currently displayed interface content 301, which is the phone's main interface. This interface content 301 displays multiple applications (Apps), such as email, calculator, settings, and music. It should be understood that interface content 301 may also include other applications, and this application does not limit this.
[0126] In one possible implementation, during the use of the voice assistant, the user can enable the phone's smart voice function through the settings application. For example, such as... Figure 3 As shown in Figure (a), the user can click the icon of the settings application. In response to the user's click, the phone displays the following: Figure 3 Figure (b) shows the main interface 302 of the settings application. This main interface 302 can include multiple menus, such as WLAN, Bluetooth, desktop and wallpaper, display and brightness, sound, and smart assistant. Users can click the smart assistant menu on interface 302; in response to the user's click, the phone displays... Figure 3 The smart assistant interface 303 shown in Figure (c) includes options such as smart voice, smart vision, smart screen recognition, contextual intelligence, and smart search. In addition, the smart assistant interface 303 also displays the wake word "Xiao Yi Xiao Yi". The wake word "Xiao Yi Xiao Yi" can be used by the user to wake up the phone, so that the phone enters a state of listening to and collecting the user's voice commands.
[0127] like Figure 3 As shown in Figure (c), when the user clicks the smart voice option on the smart assistant interface 303, the phone displays the following in response to the user's click: Figure 3 The intelligent voice interface 304 is shown in Figure (d). This intelligent voice interface 304 may include a voice wake-up switch, a power button toggle switch, an artificial intelligence (AI) letter switch, a driving scenario switch, etc. In this embodiment, the user can click the voice wake-up switch to activate the phone's voice interaction function. In other words, after the voice wake-up switch is activated, the phone can be woken up by the wake-up phrase "Xiaoyi Xiaoyi," and will begin collecting the user's voice commands and entering the voice recognition stage.
[0128] When a user activates their phone's voice interaction function, uttering the wake-up phrase "Hey Celia," a floating window can appear on the phone screen to prompt the user that voice commands are being collected and the voice interaction process has begun. For example, such as... Figure 3As shown in Figure (e), after the user utters the wake-up word "Xiaoyi Xiaoyi", the mobile phone is woken up by the wake-up word and a floating window 10 can be displayed on the screen. The floating window 10 includes the dialogue content between the user and the mobile phone (e.g., Hi, I'm listening...), and a listening icon 10-1 when the mobile phone is listening to the user's voice commands. The listening icon 10-1 can be dynamically flashed to indicate that the user's voice commands are currently being listened to. This application embodiment does not limit this.
[0129] like Figure 3 As shown in Figure (e), the mobile phone listens to the user's voice command: "Imitate the sound of a cow, Cryo, Cryo." The mobile phone can recognize the content of the voice command and display the recognized voice command in the floating window 10. In the existing solution, the mobile phone can respond according to the voice command, for example, by imitating the sound of a cow.
[0130] However, when the voice command includes the wake word "Xiaoyi Xiaoyi" again, the phone may be interrupted by the wake word "Xiaoyi Xiaoyi" in the voice command, thus interrupting the current human-computer interaction process and restarting the collection of user voice commands and speech recognition. For example... Figure 4 As shown in Figure (f), after the mobile phone recognizes that the user's voice command includes the wake word "Xiaoyi Xiaoyi", in response to the voice command including the wake word "Xiaoyi Xiaoyi", the mobile phone will re-enter the next human-computer interaction process and respond in the floating window 10: Hi, I am listening... and display a dynamically flashing listening icon 10-1 to indicate that the mobile phone has started listening to the user's voice command again.
[0131] In the above scenario, if the user's voice command includes the wake word "Xiaoyi Xiaoyi", the wake word can interrupt the current human-computer interaction process and re-enter the next human-computer interaction process. This process may not be what the user expects. That is, the wake word directly interrupts the currently executing task, causing the phone to need to start collecting the user's voice command again. This will result in a disjointed human-computer dialogue, affect the user's usage process, and reduce the human-computer interaction experience.
[0132] This application provides a method for human-computer interaction that can prevent the human-computer interaction process from being interrupted by wake words in voice commands, so as to bring users a better human-computer interaction experience.
[0133] Figure 1 This is a schematic flowchart illustrating an example of a human-computer interaction method provided in an embodiment of this application. It should be understood that this method 400 can be applied to mobile phones, PCs, in-vehicle devices, and other devices with... Figure 2 and Figure 4 On the electronic device with the structure shown. For example... Figure 3 As shown, method 400 includes:
[0134] 401, acquire the first voice instruction of the user, and detect that the first voice instruction comprises the wake-up word.
[0135] For example, in the scenario shown in (e) of FIG. 1, if the user currently expects the dialogue with the mobile phone to be the following: Figure 1
[0136] User: Xiao Yi Xiao Yi.
[0137] Mobile phone: Hi, I am listening…
[0138] User: Imitate the mooing of a cow, Xiao Yi Xiao Yi.
[0139] Mobile phone: Moooo…
[0140] When the mobile phone detects that the user says the wake-up word "Xiao Yi Xiao Yi" for the first time, the mobile phone is woken up, and the mobile phone enters a state of listening to the voice instruction of the user.
[0141] 402, determine whether the ASR module is in an open state.
[0142] It should be understood that the ASR module of the mobile phone is not always in an open state and in a working state. When the user issues a voice instruction, the mobile phone closes the automatic speech recognition (ASR) function, that is, closes the ASR module; or when the mobile phone is answering the user, the ASR module is also closed to avoid collecting the voice of the mobile phone itself and interfering with the collection and recognition of the voice instruction of the user. Through step 402, the mobile phone first detects whether the ASR module is in an open state. If the ASR is in a closed state of dormancy or non-working, the ASR module can be triggered to be opened.
[0143] Optionally, when the mobile phone acquires and recognizes the wake-up word "Xiao Yi Xiao Yi" for the first time, if it is determined that the mobile phone is currently in a state of opening the ASR module, the current wake-up can be ignored, and the current dialogue process is continued.
[0144] 403, when the mobile phone determines that the ASR module is in an open state, the position of the wake-up word in the first voice instruction is determined.
[0145] In a possible implementation manner, after the mobile phone is woken up, the first voice instruction of the user is monitored, and when it is detected that the wake-up word "Xiaoyi Xiaoyi" is included again in the first voice instruction, it can be determined that the position of the wake-up word "Xiaoyi Xiaoyi" in the first voice instruction, which mainly includes the first position of the first voice instruction, the middle of the first voice instruction, and the end of the first voice instruction. For example, the first voice instruction issued by the user in the case of including the wake-up word can be "imitate the call of a cow, Xiaoyi Xiaoyi" (the wake-up word is located at the end of the first voice instruction), "imitate the call of an animal, Xiaoyi Xiaoyi, imitate the call of a cow" (the wake-up word is located in the middle of the first voice instruction), or "Xiaoyi Xiaoyi, imitate the call of a cow" (the wake-up word is located at the beginning of the first voice instruction).
[0146] 404-1, when the position of the wake-up word in the first voice instruction is the end, step 405 is performed to determine whether the time interval of the wake-up word from the closest voice instruction is less than a first preset value.
[0147] 406, when the time interval of the wake-up word from the closest voice instruction is less than the first preset value, the time information corresponding to the wake-up word is recorded.
[0148] It should be understood that the first preset value can be used to determine whether the current user wants to interrupt the dialogue process. For example, when the first voice instruction issued by the user is "imitate the call of a cow, Xiaoyi Xiaoyi", the wake-up word is located at the end of the voice instruction. According to step 406, the closest voice instruction of the wake-up word "Xiaoyi Xiaoyi" is "imitate the call of a cow", and the mobile phone can determine the time when the user issues the wake-up word "Xiaoyi Xiaoyi" according to the time interval between "imitate the call of a cow" and "Xiaoyi Xiaoyi". When the time interval between "imitate the call of a cow" and the first "Xiao" of "Xiaoyi Xiaoyi" is less than the first preset value, it can be determined that the user may only use the wake-up word "Xiaoyi Xiaoyi" as part of a catchphrase, and wants to continue the current dialogue process without switching to the next new dialogue process.
[0149] Optionally, the mobile phone can record the time information of the wake-up word "Xiaoyi Xiaoyi" in the first voice instruction according to the first voice instruction. The recording and marking rules of the time information are not limited in the embodiments of the present application. For example, if the initial wake-up of the mobile phone is taken as the starting time, the time period when the wake-up word appears again in the first voice instruction is t1-t2; if the initial wake-up of the mobile phone is taken as the starting time, the time period when the wake-up word appears again in the first voice instruction is T1-T2, and the position of the wake-up word in the first voice instruction can be determined according to the time information.
[0150] 407, according to the time information corresponding to the wake-up word, the wake-up word is ignored, and the first voice instruction is recognized.
[0151] 408, normal response. Alternatively, the normal response here can include feedback of the mobile phone according to the user's question, and the user continues the conversation, or can also include voice responses such as "um", "good", etc. The embodiments of the present application are not limited to this.
[0152] 409, when the time length of the closest voice instruction to the wake-up word is greater than or equal to the first preset value, the current dialogue box is suspended, and a new dialogue box is opened.
[0153] 410, open ASR, recognize the second voice instruction of the user of the new dialogue box, and again respond normally according to the second voice instruction of the user, or return to step 401, and again detect whether the second voice instruction includes the wake-up word, and repeat the above process. For the sake of simplicity, this will not be repeated here.
[0154] For example, when the first voice instruction issued by the user is "imitate the sound of a cow, Xiaoyi Xiaoyi", the wake-up word is located at the end of the first voice instruction. When the time interval between "sound" of "imitate the sound of a cow" and the first "small" of "Xiaoyi Xiaoyi" is greater than or equal to the first preset value, it can be judged that the user may wish to interrupt the current dialogue process and enter the next new dialogue process. In other words, the mobile phone can take the wake-up word "Xiaoyi Xiaoyi" included again in the first voice instruction as the wake-up word of the next dialogue process, and the mobile phone is re-awakened to interrupt the previous "imitate the sound of a cow" dialogue process. Alternatively, the mobile phone can reply "Hi, I'm listening…", and the embodiments of the present application are not limited to this.
[0155] Alternatively, the first preset value can be 1 second, 2 seconds, etc. The embodiments of the present application are not limited to this.
[0156] 411, when the mobile phone determines that the ASR module is in an unopened state, the ASR module is opened, and the listening function is started. After the ASR listening function is opened, the process of obtaining the voice instruction of the user in step 401 is continued, and this will not be repeated here.
[0157] For step 403, when it is determined that the wake-up word is located at the beginning or in the middle of the first voice instruction, that is, 404-3 when the wake-up word is at the beginning of the first voice instruction, or 404-2 when the wake-up word is in the middle of the first voice instruction, steps 406-408 are executed, the time information corresponding to the wake-up word is recorded, the wake-up word is ignored according to the time information corresponding to the wake-up word, and the first voice instruction is recognized, and a normal response is given. For the sake of simplicity, this will not be repeated here.
[0158] In one possible scenario, if the voice recognition is restarted shortly after it has just ended, the mobile phone can determine whether the user continues to speak, and if the user does not continue to speak, the mobile phone can continue to dialogue with the user using the previous voice recognition result.
[0159] By the above method, in the process of voice interaction between the user and the electronic device, after the user wakes up the electronic device by the wake-up word, if the voice instruction issued by the user again includes the wake-up word, the method can avoid the wake-up word in the voice instruction from interrupting the current interaction process, thereby avoiding directly interrupting the task being performed by the current electronic device, restarting the process of collecting the voice instruction of the user, ensuring the coherence of the human-computer dialogue, and improving the user experience.
[0160] In addition, in another possible scenario, some electronic devices can have the capability of sound source positioning or the function of image acquisition by a camera, such as robots and the like. When the robot is woken up by the wake-up word, the direction where the user is located can be determined according to the sound source positioning function, and the camera with the function of image acquisition is rotated to directly turn to the direction or position where the user is located according to the sound source positioning. In this process, the direction where the user is located can have a large error in judgment due to the reflection of sound by a wall or the like. When such a large error occurs, the phenomenon that the device is not facing the person after being rotated can occur.
[0161] It should be understood that the robot can have Figure 2 part or all of the structures shown, or have Figure 5 the software architecture shown, and the embodiments of the present application do not limit this.
[0162] An exemplary Figure 5 is a schematic diagram of a scene of human-computer interaction provided by an embodiment of the present application. As Figure 6 shown, it is assumed that the robot has the capability of sound source positioning and the function of image acquisition. The robot can determine the direction of the sound source according to the voice instruction of the user, and can determine the gaze estimation direction of itself according to the image acquired by the camera. The angle between the gaze direction and the sound source direction is denoted as θ.
[0163] Optionally, in the process of determining the gaze estimation direction of itself according to the image acquired by the camera, a camera coordinate system can be established, and the gaze target and the position coordinates of the eyes of the user can be transformed to the camera coordinates by a three-dimensional six-key-point algorithm and the like based on the public parameters of the camera. The calculation process can refer to the prior art, and will not be described here.
[0164] The electronic device with the sound source positioning capability, such as the robot, provided by the embodiments of the present application also provides a method for human-computer interaction, which can avoid the human-computer interaction process being interrupted by the wake-up word in the voice instruction, so as to bring better human-computer interaction experience to the user.
[0165] Figure 6 FIG. 6 is a schematic flowchart of an example of the method for human-computer interaction provided by the embodiments of the present application. It should be understood that the method 600 can be applied to the electronic device with the sound source positioning capability, such as the robot. As shown in FIG. 6, the method 600 comprises the following steps. Range of angles between direction of user gaze and direction of robot line of sight
[0166] 601, the robot acquires the first voice instruction of the user.
[0167] 602, the robot detects the sound source direction of the first voice instruction according to the first voice instruction.
[0168] 603, the robot judges whether the included angle θ between the sound source direction of the first voice instruction and the current line-of-sight direction of the robot is greater than or equal to the first preset angle.
[0169] 604, when the included angle θ between the sound source direction of the first voice instruction and the line-of-sight direction is greater than or equal to the first preset angle, the robot judges whether the interaction intention of the user is less than the preset value.
[0170] It should be understood that when the included angle θ between the sound source direction of the first voice instruction and the line-of-sight direction is greater than or equal to the first preset angle, it can be considered that the user who issues the voice instruction and the robot are not in the face-to-face position relationship, or in other words, the user who issues the voice instruction is not within the central region range of the image collected by the robot. The range corresponding to the central region is not limited in the embodiments of the present application.
[0171] Optionally, in step 604, the robot can collect the image through the camera and detect the direction at which the eyes of the user in the collected image are gazing to estimate the interaction intention of the user. For example, Table 1 lists an example of the possible user interaction intention range.
[0172] Table 1
[0173] Range of interaction intent estimation Figure 4 0°-30° 0.8-1.0 30°-60° 0.5-0.8 60°-90° 0.1-0.5
[0174] As shown in Table 1, when the interaction intention estimation range is determined according to the included angle range between the direction of the user's gaze and the direction of the robot's line of sight, the robot can determine that the user's current interaction intention is strong when the interaction intention estimation range is 0.8-1.0, the robot can determine that the user's current interaction intention is general when the interaction intention estimation range is 0.5-0.8, and the robot can determine that the user's current interaction intention is low when the interaction intention estimation range is 0.1-0.5. The embodiments of the present application are not limited in this regard.
[0175] Optionally, the preset value can be set to 0.5, and when the estimated current interaction intention of the user is greater than or equal to the preset value, the following step 605 is continued to be executed.
[0176] 605, the robot determines whether the included angle between the sound source direction of the first voice instruction and the sound source direction of the previous voice instruction is less than a second preset angle, and whether the time interval between the two voice instructions is less than a second preset value.
[0177] 606, when the included angle between the sound source direction of the first voice instruction and the sound source direction of the previous voice instruction is less than the second preset angle, and the time interval between the two voice instructions is less than the second preset value, the robot performs normal response.
[0178] It should be understood that the "previous voice instruction" here is the closest voice instruction before the first voice instruction. Optionally, the "previous voice instruction" can be a wake-up word instruction of the user, for example: Xiaoyi Xiaoyi. Or the "previous voice instruction" is other voice instructions after the wake-up word, for example: please imitate the sound of a cow. The embodiments of the present application are not limited in this regard.
[0179] It should also be understood that the normal response here can be understood as the robot recognizing the first voice instruction of the user and making corresponding feedback according to the first voice instruction, which will not be described here.
[0180] 607, when the included angle between the sound source direction of the first voice instruction and the sound source direction of the previous voice instruction is greater than or equal to the second preset angle, and the time interval between the two voice instructions is greater than or equal to the second preset value, the robot calls a turning execution function to change the direction of the robot.
[0181] Optionally, the display conditions of "the included angle between the sound source direction of the first voice instruction and the sound source direction of the previous voice instruction is greater than or equal to the second preset angle" and "the time interval between the two voice instructions is greater than or equal to the second preset value" can meet any one or both at the same time, and the turning execution function is called to change the direction of the robot. The embodiments of the present application are not limited in this regard.
[0182] 608, in response to the turning execution function, the robot turns to the direction, and determines the user interaction intention after the turning. Optionally, the process of determining the user interaction intention can be achieved by collecting images and judging the direction of the user's gaze in the images. For details, refer to the foregoing description of step 604, which will not be repeated here.
[0183] In one possible implementation, in step 608, when the robot turns to the direction in response to the turning execution function, and the user interaction intention after the turning is determined to be relatively low, the robot can turn back to the original line-of-sight direction, and simultaneously perform step 606 to make a corresponding feedback to the user's voice instruction and give a normal response.
[0184] In another possible scenario, if the first voice instruction of the user can contain a wake-up word, the method shown in FIG. 6 can be combined with the method shown in FIG. 7, and images are collected and the user interaction intention is estimated according to the direction of the user's gaze in the images. When the robot does not detect a person in the images, it can be determined that the user's interaction intention is very low, or the current human-robot interaction process is interrupted. In other words, this scenario can be considered as a false wake-up of the robot. Figure 1
[0185] In another possible scenario, if the first voice instruction is a wake-up word, and the included angle θ between the sound source direction of the wake-up word and the line-of-sight direction currently faced by the robot is greater than or equal to a first preset angle, the robot can determine whether to respond to the current wake-up according to whether there is a user in the currently collected images and whether the user's interaction intention is strong.
[0186] For example, if the included angle θ between the sound source direction of the wake-up word and the line-of-sight direction currently faced by the robot is greater than or equal to the first preset angle, and the current user's interaction intention is strong, the robot can be set to need two consecutive wake-up words in the same sound source direction to wake up the robot, that is, the robot will respond to the user's wake-up word.
[0187] Alternatively, if the included angle θ between the sound source direction of the wake-up word and the line-of-sight direction currently faced by the robot is greater than or equal to the first preset angle, the robot turns to the sound source direction of the wake-up word, and does not detect a user, and can turn back to the angle before the wake-up, and continue the voice interaction with the person before the wake-up.
[0188] By the above method, the wake-up process of the robot is more in line with the expectation of the person, when the included angle θ between the sound source direction of the voice instruction of the user and the line of sight direction currently faced by the robot is greater than or equal to the first preset angle and the interaction intention of the user is strong, the robot can determine to automatically turn to the user; when the included angle θ between the sound source direction of the voice instruction of the user and the line of sight direction currently faced by the robot is greater than or equal to the first preset angle and the interaction intention of the user is low, the robot can also turn back, and in this process, the interaction process between the user and the robot will not be interrupted, bringing a better human-computer interaction experience to the user.
[0189] In summary, in the process of voice interaction between the user and the electronic device, after the user wakes up the electronic device through the wake-up word, if the voice instruction or the answer to the electronic device of the user again includes the wake-up word, the method can avoid the wake-up word in the voice instruction from interrupting the current interaction process, thereby avoiding directly interrupting the task being executed by the current electronic device, restarting the process of collecting the voice instruction of the user, ensuring the coherence of human-computer dialogue, and improving the user experience.
[0190] In addition, for electronic devices such as robots with sound source positioning capability, the method provided in the embodiments of the present application can determine whether to occur deflection according to the sound source direction of the voice instruction, and estimate the interaction intention of the user according to the collected images, and then more accurately perform voice interaction with the user. Specifically, when the included angle θ between the sound source direction of the voice instruction of the user and the line of sight direction currently faced by the robot is greater than or equal to the first preset angle and the interaction intention of the user is strong, the robot can determine to automatically turn to the user; when the included angle θ between the sound source direction of the voice instruction of the user and the line of sight direction currently faced by the robot is greater than or equal to the first preset angle and the interaction intention of the user is low, the robot can also turn back, and in this process, the interaction process between the user and the robot will not be interrupted, bringing a better human-computer interaction experience to the user.
[0191] It can be understood that, in order to realize the above functions, the electronic device contains hardware and / or software modules corresponding to each function. The algorithm steps of each example described in conjunction with the embodiments disclosed herein can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is realized in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to realize the described functions for each specific application in conjunction with the embodiments, but such implementation should not be considered beyond the scope of the present application.
[0192] The embodiments can divide the function modules of the electronic device according to the method examples described above. For example, each function module can be divided according to each function, or two or more functions can be integrated into one processing module. The integrated module can be implemented in the form of hardware. It should be noted that the division of the modules in the embodiments is illustrative, and is only a logical function division. In actual implementation, another division manner can be used.
[0193] In the case of dividing each function module according to each function, the electronic device such as a robot or a mobile phone involved in the above embodiments can include a collection unit, a detection unit and a processing unit.
[0194] The collection unit, the detection unit and the processing unit can cooperate with each other to support the electronic device such as a robot or a mobile phone to perform the above steps and / or other processes of the technology described herein.
[0195] It should be noted that all related contents of each step involved in the above method embodiments can be cited in the function description of the corresponding function module, which will not be described here.
[0196] The electronic device provided by the embodiments can be used to perform the above video playing method, and thus the same effects as the implementation method can be achieved.
[0197] In the case of using the integrated unit, the electronic device can include a processing module, a storage module and a communication module. The processing module can be used to control and manage the actions of the electronic device, for example, to support the electronic device to perform the steps performed by the collection unit, the detection unit and the processing unit. The storage module can be used to support the electronic device to store program codes and data. The communication module can be used to support the communication between the electronic device and other devices.
[0198] The processing module can be a processor or a controller. It can implement or execute various exemplary logical blocks, modules and circuits described in combination with the disclosure of the present application. The processor can also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of digital signal processing (DSP) and microprocessors, and the like. The storage module can be a memory. The communication module can be a device for interacting with other electronic devices, such as a radio frequency circuit, a Bluetooth chip, a Wi-Fi chip, and the like.
[0199] In one embodiment, when the processing module is a processor and the storage module is a memory, the electronic device involved in the embodiments can be a device with the structure as shown in the figure.
[0200] The embodiment further provides a computer readable storage medium, which stores computer instructions, and when the computer instructions run on an electronic device, the electronic device executes the related method steps to realize the method of human-computer interaction in the above embodiment.
[0201] The embodiment further provides a computer program product, which, when running on a computer, causes the computer to execute the related steps to realize the method of human-computer interaction in the above embodiment.
[0202] In addition, the embodiment of the present application further provides an apparatus, which can be a chip, a component or a module, and the apparatus can include a processor and a memory connected to each other; and the memory is used to store computer execution instructions, and when the apparatus runs, the processor can execute the computer execution instructions stored in the memory to enable the chip to execute the method of human-computer interaction in the above method embodiments.
[0203] The electronic device, the computer readable storage medium, the computer program product or the chip provided by the embodiment are used to execute the corresponding method provided above, and thus the beneficial effects achieved by the electronic device, the computer readable storage medium, the computer program product or the chip can refer to the beneficial effects of the corresponding method provided above, which will not be described herein.
[0204] From the above description of the embodiments, those skilled in the art can understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the apparatus is divided into different functional modules to complete all or part of the functions described above.
[0205] In the several embodiments provided by the present application, it should be understood that the disclosed apparatus and method can be implemented in other ways. For example, the apparatus embodiments described above are only schematic. The division of the modules or units is only a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another apparatus, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling or direct coupling or communication connection between the units can be indirect coupling or communication connection through some interfaces, and can be electrical, mechanical or other forms.
[0206] The units described as separate components can or can not be physically separate, and the components displayed as units can be one physical unit or multiple physical units, that is, can be located in one place, or can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0207] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0208] When the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such understanding, the technical solutions of the embodiments of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The software product is stored in a storage medium and includes a number of instructions for causing an apparatus (which can be a single-chip microcomputer, a chip, etc.) or a processor to perform all or part of the steps of the method of each embodiment of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0209] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for human-computer interaction, characterized in that, The method includes: Obtain the user's first voice command, and detect the direction of the sound source of the first voice command based on the first voice command; Determine the first angle between the sound source direction of the first voice command and the first line of sight currently facing the electronic device; When the first angle is greater than or equal to the first preset angle, the second angle between the sound source direction of the first voice command and the sound source direction of the second voice command is determined. The second voice command is the voice command issued by the user before the first voice command and is closest to the first voice command. When the second angle is less than or equal to the second preset angle, the electronic device responds to the first voice command and provides an answer.
2. The method according to claim 1, characterized in that, The method further includes: Detect the time interval between the first voice command and the second voice command; When the time interval is greater than or equal to the second preset value, the steering execution function is invoked to rotate the electronic device to face or infinitely approach the sound source direction of the first voice command.
3. The method according to claim 1 or 2, characterized in that, The method further includes: Acquire a first image of the electronic device in the first line of sight direction; When the first image includes the user and the third angle between the user's line of sight and the sound source direction of the first voice command is less than or equal to a third preset angle, the steering execution function is invoked to rotate the electronic device to face or infinitely approach the sound source direction of the first voice command.
4. The method according to claim 3, characterized in that, The method further includes: The electronic device acquires a second image in the direction of the sound source facing or infinitely close to the first voice command; When the second image does not include the user or the fourth angle between the user's line of sight and the current second line of sight of the electronic device is greater than a fourth preset angle, the electronic device is rotated to return to the first line of sight.
5. An electronic device, characterized in that, include: One or more processors; One or more memory units; A module with multiple applications installed; The memory stores one or more programs that, when executed by the processor, cause the electronic device to perform the method as described in any one of claims 1 to 4.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 4.
7. A computer program product, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Voice awakening processing method and device, electronic equipment and storage medium
CN110349579A
Speech recognition method, device and equipment and computer readable storage medium
CN112185388A