Voice assistant interaction method and electronic device
By detecting the target posture and movement of the electronic device, determining whether to wake up the voice assistant, and using sound pickup methods and feedback methods that adapt to the device posture, the problem of voice assistants being easily awakened by mistake and high power consumption in the prior art is solved, and a more natural and easy-to-use voice interaction experience is achieved.
Patent Information
- Application Number
- PCT/CN2024/101163
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-31
- Filing Date
- 2024-06-25
- Publication Date
- 2025-05-08
AI Technical Summary
While improving the ease of use of voice assistants, the prior art is difficult to effectively control the power consumption of electronic devices during voice interaction, and the voice assistant is easily awakened by mistake, affecting the user experience.
By detecting the target posture and movement of the electronic device, determine whether to wake up the voice assistant, and adopt sound pickup methods and feedback methods that adapt to the attitude of the device to reduce the probability of false wake-up and optimize power consumption.
It effectively reduces the probability of voice assistant being accidentally awakened, reduces the power consumption of electronic devices, and improves the naturalness and ease of use of users and voice assistants.
Smart Images

Figure CN2024101163_08052025_PF_FP_ABST
Abstract
Description
Voice assistant interaction method and electronic device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on October 31, 2023, with application number 202311439204.0 and invention name “Voice Assistant Interaction Method and Electronic Device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of terminal device software, and more specifically, to a method and electronic device for voice assistant interaction. Background Art
[0003] Voice interaction is a crucial entry point for human-computer interaction between users and electronic devices. Power consumption, fluency, usability, response speed, and intelligence are key factors influencing whether and how often users use voice assistants. How to improve the usability of voice assistants while controlling the power consumption of electronic devices during voice interaction is a worthy consideration.
[0004] Summary of the Invention
[0005] This application provides a method for interacting with a voice assistant. An electronic device can determine whether to wake up a voice assistant based on its own device posture and device action, and use a sound pickup method and feedback method that are compatible with the device posture of the electronic device. This reduces the probability of the voice assistant being mistakenly woken up, reduces the power consumption of the electronic device, makes the user's voice interaction process more natural, and improves the usability of the voice assistant.
[0006] In a first aspect, a method for voice interaction is provided, which is applied to an electronic device, the method comprising: detecting a target action of the electronic device, determining that the electronic device is in a target posture, the target posture comprising the posture of a lower microphone of the electronic device close to the user's mouth, the target action comprising an action for converting the electronic device from a non-target posture to a target posture; displaying and / or playing first feedback information in response to the target posture of the electronic device, the first feedback information being used to indicate that a voice assistant is successfully awakened; recording a voice instruction at least through the lower microphone of the electronic device; and displaying and / or playing second feedback information in response to the voice instruction, the second feedback information being used to indicate an execution result of the voice instruction.
[0007] It should be understood that before displaying and / or playing the first feedback information, the electronic device can open or call a voice assistant application so that the application can interact with the user, such as obtaining the user's voice commands.
[0008] Before recording the sound of the user's voice interaction, the electronic device first provides feedback to the user that the voice assistant has been successfully awakened. The speaker of the electronic device does not need to be in a normally open state to record audio information. When the user is using the voice assistant, the power consumption of the electronic device is lower. In addition, when the user's mouth is close to the lower microphone of the electronic device, the electronic device can set the method of recording voice commands to a method that is compatible with the device posture, which is conducive to enabling the electronic device to obtain the user's voice commands more clearly and improve the efficiency of the electronic device in calling the voice assistant. From another perspective, the implementation of this method is also conducive to reducing the power consumption of the electronic device during the recording process.
[0009] In combination with the first aspect, in certain implementations of the first aspect, the target posture includes a posture in which the lower microphone is close to the user's mouth and the ambient light sensor and / or the proximity sensor are not blocked. In response to the target posture of the electronic device, a prompt screen is displayed and / or a first prompt audio is played through the lower speaker of the electronic device, and the first feedback information includes the prompt screen and / or the first prompt audio.
[0010] Here, the target posture may also be referred to as the first posture.
[0011] In some scenarios, the lower microphone is close to the user's mouth, and the ambient light sensor and / or proximity sensor are not blocked. This can also be understood as the electronic device's position when the lower microphone is close to the user's mouth and the user is not making a call. In this position, the side of the electronic device where the front camera is located is away from the user's head.
[0012] This technical solution specifically provides a method for an electronic device to send feedback information to a user when the electronic device is in a first posture. For an electronic device in the first posture, if the user is using a voice assistant, there is a high probability that the user is looking at the screen of the electronic device. Displaying a prompt screen through the display screen is more intuitive and easier for the user to notice, and the electronic device is more efficient in sending feedback information. Similarly, when the user's mouth is close to the lower microphone, the lower speaker of the electronic device is closer to the user. Playing the prompt audio through the lower speaker is also conducive to the feedback information being noticed by the user, and the electronic device is more efficient in sending feedback information. In short, the implementation of this technical solution is conducive to making the interaction between the user and the voice assistant more natural, improving the usability of the voice assistant, and increasing the frequency of users using the voice assistant.
[0013] In combination with the first aspect, in certain implementations of the first aspect, a first ultrasonic signal is sent in response to the target posture of the electronic device; a second ultrasonic signal returned based on the first ultrasonic signal is received; and when the second ultrasonic signal indicates that the electronic device is in the target posture, a prompt screen is displayed and / or a first prompt audio is played through the lower speaker of the electronic device.
[0014] The electronic device can confirm whether the electronic device's lower microphone is close to the user's mouth by sending an ultrasonic signal and recording the returned ultrasonic signal. If this confirmation method also indicates that the electronic device is in a posture where the lower microphone is close to the user's mouth, feedback information indicating that the voice assistant has been successfully awakened is sent to the user. Because the results determined using ultrasound are more reliable, the implementation of this technical solution helps reduce the probability of the electronic device's voice assistant being mistakenly awakened, and helps reduce the power consumption of the electronic device caused by mistaken awakening of the voice assistant.
[0015] In combination with the first aspect, in certain implementations of the first aspect, a first ultrasonic signal is sent through the lower speaker in response to the target posture of the electronic device; and a second ultrasonic signal is received through the lower microphone in response to the first ultrasonic signal contacting the user and returning the second ultrasonic signal.
[0016] This technical solution specifically uses a lower speaker near the user's mouth to transmit a first ultrasonic signal, and a lower microphone near the user's mouth to record a second ultrasonic signal. This technical solution improves the accuracy and reliability of ultrasonic measurement results, reduces the probability of the voice assistant being falsely awakened, and reduces the power consumption of electronic devices.
[0017] In combination with the first aspect, in certain implementations of the first aspect, the target posture includes a posture in which the lower microphone of the electronic device is close to the user's mouth and the ambient light sensor and / or the proximity sensor are blocked. In response to the target posture of the electronic device, a second prompt audio is played through the upper speaker of the electronic device, and the first feedback information includes the second prompt audio.
[0018] Here, the target posture may also be referred to as a second posture.
[0019] In some scenarios, the lower microphone is close to the user's mouth, and the ambient light sensor and / or proximity sensor are blocked. This can also be understood as the electronic device's posture when the lower microphone is close to the user's mouth and the user is making a phone call. In this posture, the side of the electronic device with the front camera is close to or in contact with the user's head.
[0020] This technical solution specifically provides a method for an electronic device to send feedback information to a user when the electronic device is in a second posture. In the posture of making or receiving a call, the display screen of the electronic device is close to the user's ear, and the user cannot see the information on the display screen of the electronic device. The upper speaker (or receiver) of the electronic device is close to the user's ear, and the upper speaker is used to play the prompt audio, which helps the feedback information to be perceived by the user, and the efficiency of the electronic device in sending feedback information is higher. The implementation of this technical solution is conducive to making the interaction between the user and the voice assistant more natural, improving the usability of the voice assistant, and increasing the frequency of users using the voice assistant.
[0021] In combination with the first aspect, in certain implementations of the first aspect, touch information of the display screen of the electronic device is obtained in response to the target posture of the electronic device; when the touch information indicates that the display screen of the electronic device is touched, a second prompt audio is played through the upper speaker of the electronic device.
[0022] In combination with the first aspect, in some implementations of the first aspect, the touch information includes information of the outer ear contour.
[0023] In a possible implementation, the touch information in this technical solution may include a capacitance value of a display screen of the electronic device.
[0024] The electronic device can use touch information from the display to confirm whether the electronic device is in the posture for making or receiving a call. If the electronic device is in the posture for making or receiving a call, it can then send feedback to the user indicating that the voice assistant has been successfully awakened. Using touch information from the display to confirm the posture of the electronic device is more reliable. The implementation of this technical solution helps reduce the probability of the electronic device's voice assistant being mistakenly awakened, and helps reduce the power consumption of the electronic device caused by mistaken awakening of the voice assistant.
[0025] In combination with the first aspect, in certain implementations of the first aspect, the lower microphone of the electronic device is used as a main recording unit, and the upper microphone of the electronic device is used as an auxiliary recording unit to record voice instructions.
[0026] The voice command recording method provided by this technical solution is adapted to the device posture, and using multiple microphones to record voice commands is beneficial for the electronic device to obtain the user's voice commands more clearly, thereby improving the efficiency of the electronic device in calling the voice assistant.
[0027] In combination with the first aspect, in certain implementations of the first aspect, before recording voice instructions using the lower microphone of the electronic device as the main recording unit and the upper microphone of the electronic device as the auxiliary recording unit, the method also includes: recording first audio through the upper microphone and the lower microphone; and determining that the user is speaking close to the lower microphone based on the first audio.
[0028] Since the electronic device in the target posture does not directly correspond to the user currently using or intending to use the voice assistant, determining that the user is speaking close to the lower microphone of the electronic device through the upper and lower microphones of the electronic device is beneficial to improving the accuracy of judging the user's intention to use the voice assistant, reducing the probability of the voice assistant being mistakenly awakened, and reducing the power consumption of the electronic device.
[0029] In a second aspect, a method for voice assistant interaction is provided, which is applied to an electronic device, including: detecting a target action of the electronic device, determining that the electronic device is in a target posture, the target posture including the posture of the upper microphone of the electronic device close to the user's mouth, and the target action including an action for converting the electronic device from a non-target posture to a target posture; in response to the target posture of the electronic device, playing first feedback information, the first feedback information is used to indicate that the voice assistant is successfully awakened; recording a voice instruction at least through the upper microphone of the electronic device; in response to the voice instruction, playing second feedback information, the second feedback information is used to indicate the execution result of the voice instruction.
[0030] It should be understood that before playing the first feedback information, the electronic device can open or call a voice assistant application so that the application can interact with the user, such as obtaining the user's voice commands.
[0031] Before recording the sound of the user's voice interaction, the electronic device first provides feedback to the user that the voice assistant has been successfully awakened. The speaker of the electronic device does not need to be in a normally open state to record audio information. When the user is using the voice assistant, the power consumption of the electronic device is lower. In addition, when the user's mouth is close to the upper microphone of the electronic device, the electronic device can set the recording method to a method that is adapted to the device posture, which is conducive to enabling the electronic device to obtain the user's voice command more clearly and improve the efficiency of the electronic device in calling the voice assistant. From another perspective, the implementation of this method is also conducive to reducing the power consumption of the electronic device during the recording process.
[0032] In combination with the second aspect, in some implementations of the second aspect, in response to the target gesture of the electronic device, the first feedback information is played through an upper speaker of the electronic device.
[0033] Here, the target posture may also be referred to as the third posture.
[0034] This technical solution specifically provides a method for an electronic device to send feedback information to a user when the electronic device is in the third posture. For an electronic device in the third posture, if the user is using a voice assistant, the probability of the image on the electronic device's display screen being noticed by the user is low. The upper speaker of the electronic device is close to the user, and playing prompt audio through the upper speaker is conducive to the feedback information being noticed by the user, and the efficiency of the electronic device in sending feedback information is higher. The implementation of this technical solution is conducive to making the interaction between the user and the voice assistant more natural, improving the usability of the voice assistant, and increasing the frequency of users using the voice assistant.
[0035] In combination with the second aspect, in certain implementations of the second aspect, a first ultrasonic signal is sent in response to the target posture of the electronic device; a second ultrasonic signal returned based on the first ultrasonic signal is received; and when the second ultrasonic signal indicates that the electronic device is in the target posture, first feedback information is played through the upper speaker of the electronic device.
[0036] The electronic device can confirm whether the electronic device's upper microphone is close to the user's mouth by sending an ultrasonic signal and recording the returned ultrasonic signal. If this confirmation method also indicates that the electronic device is in a posture where the upper microphone is close to the user's mouth, feedback information indicating that the voice assistant has been successfully awakened is sent to the user. Because the results determined using ultrasound are more reliable, the implementation of this technical solution helps reduce the probability of the electronic device's voice assistant being mistakenly awakened, and helps reduce the power consumption of the electronic device caused by mistaken awakening of the voice assistant.
[0037] In combination with the second aspect, in certain implementations of the second aspect, a first ultrasonic signal is sent through the upper speaker in response to the target posture of the electronic device; and a second ultrasonic signal is received through the upper microphone in response to the first ultrasonic signal contacting the user and returning the second ultrasonic signal.
[0038] This technical solution specifically uses an upper speaker near the user's mouth to transmit a first ultrasonic signal, and an upper microphone near the user's mouth to record a second ultrasonic signal. This technical solution improves the accuracy and reliability of ultrasonic measurement results, reduces the probability of the voice assistant being falsely awakened, and reduces the power consumption of electronic devices.
[0039] In combination with the second aspect, in some implementations of the second aspect, the upper microphone of the electronic device is used as the main recording unit, and the lower microphone of the electronic device is used as the auxiliary recording unit to record voice instructions.
[0040] The voice command recording method provided by this technical solution is adapted to the device posture, and using multiple microphones to record voice commands is beneficial for the electronic device to obtain the user's voice commands more clearly, thereby improving the efficiency of the electronic device in calling the voice assistant.
[0041] In combination with the second aspect, in certain implementations of the second aspect, before recording voice commands using the upper microphone of the electronic device as the main recording unit and the lower microphone of the electronic device as the auxiliary recording unit, the method also includes: recording first audio through the upper microphone and the lower microphone; and determining that the user is speaking close to the upper microphone based on the first audio.
[0042] Since the electronic device in the target posture does not directly correspond to the user currently using or intending to use the voice assistant, determining that the user is speaking close to the upper microphone of the electronic device through the upper and lower microphones of the electronic device is beneficial to improving the accuracy of judging the user's intention to use the voice assistant, reducing the probability of the voice assistant being mistakenly awakened, and reducing the power consumption of the electronic device.
[0043] According to a third aspect, a device for interacting with a voice assistant is provided, which includes an acquisition module and a processing module. The processing module is used to: detect a target action of the device and determine that the device is in a target posture, the target posture including the posture of the lower microphone of the device close to the user's mouth, and the target action including the action for converting the device from a non-target posture to a target posture; in response to the target posture of the device, display and / or play first feedback information, the first feedback information is used to indicate that the voice assistant is successfully awakened; the acquisition module is used to: record voice instructions at least through the lower microphone of the device; the processing module is also used to: in response to the voice instruction, display and / or play second feedback information, the second feedback information is used to indicate the execution result of the voice instruction.
[0044] In combination with the third aspect, in certain implementations of the third aspect, the target posture includes a posture in which the lower microphone is close to the user's mouth and the ambient light sensor and / or the proximity sensor are not blocked. The processing module is specifically used to: in response to the target posture of the device, display a prompt screen and / or play a first prompt audio through the lower speaker of the device, and the first feedback information includes the prompt screen and / or the first prompt audio.
[0045] In combination with the third aspect, in certain implementations of the third aspect, the processing module is specifically used to: send a first ultrasonic signal in response to the target posture of the device; the acquisition module is also used to: receive a second ultrasonic signal returned based on the first ultrasonic signal; when the second ultrasonic signal indicates that the device is in the target posture, display a prompt screen and / or play a first prompt audio through the lower speaker of the device.
[0046] In combination with the third aspect, in certain implementations of the third aspect, the processing module is further used to: send a first ultrasonic signal through the lower speaker in response to the target posture of the device; the acquisition module is further used to: receive a second ultrasonic signal through the lower microphone in response to a second ultrasonic signal returned when the first ultrasonic signal contacts the user.
[0047] In combination with the third aspect, in certain implementations of the third aspect, the target posture includes a posture in which the lower microphone of the device is close to the user's mouth and the ambient light sensor and / or the proximity sensor are blocked, and the processing module is specifically used to: in response to the target posture of the device, play a second prompt audio through the upper speaker of the device, and the first feedback information includes the second prompt audio.
[0048] In combination with the third aspect, in certain implementations of the third aspect, the processing module is specifically used to: obtain touch information of the display screen of the device in response to the target posture of the device; and play a second prompt audio through the upper speaker of the device when the touch information indicates that the display screen of the device is touched.
[0049] In combination with the third aspect, in certain implementations of the third aspect, the touch information includes information of the outer ear contour.
[0050] In combination with the third aspect, in certain implementations of the third aspect, the acquisition module is specifically used to: use the lower microphone of the device as the main recording unit and the upper microphone of the device as the auxiliary recording unit to record voice instructions.
[0051] In combination with the third aspect, in certain implementations of the third aspect, before the lower microphone of the device is used as the main recording unit and the upper microphone of the device is used as the auxiliary recording unit to record voice instructions, the acquisition module is also used to: record the first audio through the upper microphone and the lower microphone; the processing module is also used to: determine that the user is speaking close to the lower microphone based on the first audio.
[0052] In a fourth aspect, a device for interacting with a voice assistant is provided, which includes an acquisition module and a processing module. The processing module is used to: detect a target action of the device and determine that the device is in a target posture, the target posture including the posture of the upper microphone of the device close to the user's mouth, and the target action including the action for converting the device from a non-target posture to a target posture; in response to the target posture of the device, play first feedback information, and the first feedback information is used to indicate that the voice assistant is successfully awakened; the acquisition module is used to: record voice instructions at least through the upper microphone of the device; the processing module is also used to: in response to the voice instruction, play second feedback information, and the second feedback information is used to indicate the execution result of the voice instruction.
[0053] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is specifically configured to: play the first feedback information through an upper speaker of the device in response to the target posture of the device.
[0054] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is further used to: send a first ultrasonic signal in response to the target posture of the device; the acquisition module is further used to: receive a second ultrasonic signal returned based on the first ultrasonic signal; the processing module is further used to: play first feedback information through the upper speaker of the device when the second ultrasonic signal indicates that the device is in the target posture.
[0055] In combination with the fourth aspect, in certain implementations of the fourth aspect, the processing module is specifically used to: send a first ultrasonic signal through the upper speaker in response to the target posture of the device; the acquisition module is specifically used to: receive a second ultrasonic signal through the upper microphone in response to a second ultrasonic signal returned when the first ultrasonic signal contacts the user.
[0056] In combination with the fourth aspect, in certain implementations of the fourth aspect, the acquisition module is specifically used to: use the upper microphone of the device as the main recording unit and the lower microphone of the device as the auxiliary recording unit to record voice instructions.
[0057] In combination with the fourth aspect, in certain implementations of the fourth aspect, before recording voice instructions using the upper microphone of the device as the main recording unit and the lower microphone of the device as the auxiliary recording unit, the acquisition module is also used to: record the first audio through the upper microphone and the lower microphone; the processing module is also used to: determine that the user is speaking close to the upper microphone based on the first audio.
[0058] In a fifth aspect, an electronic device is provided, which includes a processor and a memory, the memory being used to store program instructions, the processor being used to: detect a target action of the electronic device, determine that the electronic device is in a target posture, the target posture including the posture of the lower microphone of the electronic device close to the user's mouth, and the target action including an action for converting the electronic device from a non-target posture to a target posture; display and / or play first feedback information in response to the target posture of the electronic device, the first feedback information being used to indicate that the voice assistant is successfully awakened; record voice instructions at least through the lower microphone of the electronic device; display and / or play second feedback information in response to the voice instruction, the second feedback information being used to indicate the execution result of the voice instruction.
[0059] In combination with the fifth aspect, in certain implementations of the fifth aspect, the target posture includes a posture in which the lower microphone is close to the user's mouth and the ambient light sensor and / or the proximity sensor are not blocked, and the processor is further used to: display a prompt screen and / or play a first prompt audio through the lower speaker of the electronic device in response to the target posture of the electronic device, and the first feedback information includes the prompt screen and / or the first prompt audio.
[0060] In combination with the fifth aspect, in certain implementations of the fifth aspect, the processor is specifically used to: send a first ultrasonic signal in response to the target posture of the electronic device; receive a second ultrasonic signal returned based on the first ultrasonic signal; and display a prompt screen and / or play a first prompt audio through the lower speaker of the electronic device when the second ultrasonic signal indicates that the electronic device is in the target posture.
[0061] In combination with the fifth aspect, in certain implementations of the fifth aspect, the processor is specifically used to: send a first ultrasonic signal through the lower speaker in response to the target posture of the electronic device; and receive a second ultrasonic signal through the lower microphone in response to the second ultrasonic signal returned by the first ultrasonic signal contacting the user.
[0062] In combination with the fifth aspect, in certain implementations of the fifth aspect, the target posture includes a posture in which the lower microphone of the electronic device is close to the user's mouth and the ambient light sensor and / or the proximity sensor are blocked, and the processor is specifically used to: play a second prompt audio through the upper speaker of the electronic device in response to the target posture of the electronic device, and the first feedback information includes the second prompt audio.
[0063] In combination with the fifth aspect, in certain implementations of the fifth aspect, the processor is specifically used to: obtain touch information of the display screen of the electronic device in response to the target posture of the electronic device; and play a second prompt audio through the upper speaker of the electronic device when the touch information indicates that the electronic device is in a state where the display screen is touched.
[0064] In combination with the fifth aspect, in certain implementations of the fifth aspect, the touch information includes information of the outer ear contour.
[0065] In combination with the fifth aspect, in certain implementations of the fifth aspect, the processor is further used to: record voice instructions using the lower microphone of the electronic device as a main recording unit and the upper microphone of the electronic device as an auxiliary recording unit.
[0066] In combination with the fifth aspect, in certain implementations of the fifth aspect, before recording voice instructions using the lower microphone of the electronic device as the main recording unit and the upper microphone of the electronic device as the auxiliary recording unit, the processor is also used to: record first audio through the upper microphone and the lower microphone; and determine that the user is speaking close to the lower microphone based on the first audio.
[0067] In a sixth aspect, an electronic device is provided, which includes a processor and a memory, the memory being used to store program instructions, the processor being used to: detect a target action of the electronic device, determine that the electronic device is in a target posture, the target posture including the posture of the upper microphone of the electronic device close to the user's mouth, and the target action including the action for converting the electronic device from a non-target posture to a target posture; in response to the target posture of the electronic device, play first feedback information, the first feedback information being used to indicate that the voice assistant is successfully awakened; record voice instructions at least through the upper microphone of the electronic device; in response to the voice instruction, play second feedback information, the second feedback information being used to indicate the execution result of the voice instruction.
[0068] In combination with the sixth aspect, in certain implementations of the sixth aspect, the processor is specifically configured to: play the first feedback information through an upper speaker of the electronic device in response to a target posture of the electronic device.
[0069] In combination with the sixth aspect, in certain implementations of the sixth aspect, the processor is specifically used to: send a first ultrasonic signal in response to the target posture of the electronic device; receive a second ultrasonic signal returned based on the first ultrasonic signal; and play first feedback information through the upper speaker of the electronic device when the second ultrasonic signal indicates that the electronic device is in the target posture.
[0070] In combination with the sixth aspect, in certain implementations of the sixth aspect, the processor is specifically used to: send a first ultrasonic signal through the upper speaker in response to the target posture of the electronic device; and receive a second ultrasonic signal through the upper microphone in response to the second ultrasonic signal returned by the first ultrasonic signal contacting the user.
[0071] In combination with the sixth aspect, in certain implementations of the sixth aspect, the processor is specifically used to: use the upper microphone of the electronic device as the main recording unit and the lower microphone of the electronic device as the auxiliary recording unit to record voice instructions.
[0072] In combination with the sixth aspect, in certain implementations of the sixth aspect, before recording voice instructions using the upper microphone of the electronic device as the main recording unit and the lower microphone of the electronic device as the auxiliary recording unit, the processor is also used to: record first audio through the upper microphone and the lower microphone; and determine that the user is speaking close to the upper microphone based on the first audio.
[0073] In a seventh aspect, a computer program product is provided, which includes a computer program code, and when the computer program code runs on a computer, the method in the first aspect and any possible implementation thereof or the method in the second aspect and any possible implementation thereof is executed.
[0074] In an eighth aspect, a computer-readable storage medium is provided, which stores a computer program code. When the computer program code runs on a computer, the method in the first aspect and any possible implementation thereof or the method in the second aspect and any possible implementation thereof is executed.
[0075] In the ninth aspect, a chip is provided, comprising a processor for reading instructions stored in a memory, wherein when the processor executes the instructions, the chip implements the method in the first aspect and any possible implementation thereof or the method in the second aspect and any possible implementation thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] FIG1 is a schematic diagram of the hardware architecture of an electronic device applicable to an embodiment of the present application.
[0077] FIG2 is a schematic diagram of a software architecture of an electronic device applicable to an embodiment of the present application.
[0078] FIG3 is a method for interacting with a voice assistant provided in an embodiment of the present application.
[0079] FIG4 is a schematic diagram of device postures of several electronic devices provided in an embodiment of the present application.
[0080] FIG5 is another voice assistant interaction method provided in an embodiment of the present application.
[0081] FIG6 is a schematic structural diagram of an electronic device provided in an embodiment of the present application.
[0082] FIG7 is a schematic diagram of a method for an electronic device to determine whether a target action is detected, provided by an embodiment of the present application.
[0083] FIG8 is a schematic diagram of a method for an electronic device to detect the distance between a microphone and a user's mouth using ultrasound, provided in an embodiment of the present application.
[0084] FIG9 is a schematic diagram of a method for an electronic device to send feedback information provided by an embodiment of the present application.
[0085] FIG10 is a schematic diagram of a method for an electronic device to determine a positional relationship between a user and the electronic device provided in an embodiment of the present application.
[0086] FIG11 is another voice assistant interaction method provided in an embodiment of the present application.
[0087] FIG12 is a schematic diagram of another method for an electronic device to determine whether a target action is detected, provided by an embodiment of the present application.
[0088] FIG13 is another voice assistant interaction method provided in an embodiment of the present application.
[0089] FIG14 is a schematic diagram of another method for an electronic device to determine whether a target action is detected, provided by an embodiment of the present application.
[0090] FIG15 is a device for voice assistant interaction provided in an embodiment of the present application.
[0091] FIG16 is an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0092] The technical solution in this application will be described below with reference to the accompanying drawings.
[0093] The terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of this application and the appended claims, the singular expressions "a", "an", "said", "above", "the" and "this" are intended to also include expressions such as "one or more", unless there is a clear contrary indication in the context. It should also be understood that in the following embodiments of the present application, "at least one", "one or more" refer to one, two or more. The term "and / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist; for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship.
[0094] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0095] 1 shows a schematic structural diagram of an electronic device 100. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identification module (SIM) card interface 195. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0096] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0097] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0098] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0099] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.
[0100] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.
[0101] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C bus lines. The processor 110 may be coupled to the touch sensor 180K, the charger, the flash, the camera 193, and the like via different I2C bus interfaces. For example, the processor 110 may be coupled to the touch sensor 180K via the I2C interface, enabling communication between the processor 110 and the touch sensor 180K via the I2C bus interface, thereby implementing the touch function of the electronic device 100.
[0102] The I2S interface can be used for audio communication. In some embodiments, the processor 110 can include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface, enabling the function of answering calls through a Bluetooth headset.
[0103] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via a PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering calls via a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0104] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display 194 and the camera 193. MIPI interfaces include the camera serial interface (CSI) and the display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to implement the camera function of the electronic device 100. The processor 110 and the display 194 communicate via the DSI interface to implement the display function of the electronic device 100.
[0105] The GPIO interface can be configured via software. The GPIO interface can be configured as either a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to the camera 193, display 194, wireless communication module 160, audio module 170, sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0106] The USB interface 130 is an interface that complies with USB standards and may be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. The USB interface 130 can be used to connect a charger to charge the electronic device 100, or to transfer data between the electronic device 100 and peripheral devices. It can also be used to connect headphones to play audio. This interface can also be used to connect other electronic devices, such as augmented reality devices.
[0107] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.
[0108] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0109] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0110] Display screen 194 is used to display images, videos, and the like. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED). In some embodiments, electronic device 100 may include one or N display screens 194, where N is a positive integer greater than one.
[0111] Electronic device 100 can implement a camera function using an ISP, camera 193, a video codec, a GPU, a display 194, and an application processor. The ISP processes data fed back by camera 193. Camera 193 is used to capture still images or video. The digital signal processor processes digital signals, and can process not only digital image signals but also other digital signals. The video codec compresses or decompresses digital video.
[0112] NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission mode between neurons in the human brain, it can quickly process input information and can also continuously self-learn.
[0113] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement a data storage function.
[0114] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area can store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0115] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.
[0116] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.
[0117] The speaker 170A, also called a "speaker", is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or listen to hands-free calls through the speaker 170A.
[0118] In some examples, the electronic device 100 may include one or more speakers 170A. If there are multiple speakers 170A, the multiple speakers 170A may be distributed at different locations of the electronic device 100.
[0119] Exemplarily, at least one of the plurality of speakers 170A may be located on the same side of the display screen of the electronic device 100 , and at least one may be located on the other side opposite to the display screen of the electronic device 100 .
[0120] For example, with the direction of the longer side of the electronic device 100 being the length direction and the direction of the shorter side of the electronic device 100 being the width direction, the multiple speakers 170A may be distributed at two opposite upper and lower locations in the length direction or at two opposite left and right locations in the width direction of the electronic device 100. For example, with the location where the front camera of the electronic device 100 is set being the upper portion and the location on the same surface opposite to the upper portion and away from the front camera being the lower portion, the multiple speakers 170A may include an upper speaker and a lower speaker, with the upper speaker being located at the upper portion of the electronic device 100, or near the front camera of the electronic device 100, and the lower speaker being located at the lower portion of the electronic device 100, or away from the front camera of the electronic device 100.
[0121] It should be noted that the upper and lower parts, up and down, and left and right are relative positions. The above description is only for illustrative purposes and should not be construed as limiting the present application.
[0122] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the electronic device 100 receives a call or a voice message, the user can place the receiver 170B close to the ear to hear the voice.
[0123] Microphone 170C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the source of sound, realize directional recording function, etc.
[0124] In the case that there are multiple microphones 170C, the multiple microphones 170C may be distributed at different locations of the electronic device 100 .
[0125] Exemplarily, at least one of the plurality of microphones 170C may be located on the same side of the display screen of the electronic device 100 , and at least one may be located on the other side opposite to the display screen of the electronic device 100 .
[0126] For example, with the direction of the longer side of the electronic device 100 being the length direction and the direction of the shorter side of the electronic device 100 being the width direction, the multiple microphones 170C may be distributed at two opposite upper and lower locations in the length direction or at two opposite left and right locations in the width direction of the electronic device 100. For example, with the location where the front camera of the electronic device 100 is set being the upper portion and the location on the same surface opposite to the upper portion and away from the front camera being the lower portion, the multiple microphones 170C may include an upper speaker and a lower speaker, with the upper speaker being located at the upper portion of the electronic device 100, or near the front camera of the electronic device 100, and the lower speaker being located at the lower portion of the electronic device 100, or away from the front camera of the electronic device 100.
[0127] It should be noted that the upper and lower parts, up and down, and left and right are relative positions. The above description is only for illustrative purposes and should not be construed as limiting the present application.
[0128] The headphone jack 170D is used to connect a wired headphone and can be the USB interface 130 or a 3.5mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0129] The buttons 190 include a power button, a volume button, and the like. The buttons 190 may be mechanical buttons or touch buttons. The electronic device 100 may receive key inputs and generate key signal inputs related to user settings and function control of the electronic device 100.
[0130] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts and can also be used for touch vibration feedback.
[0131] The indicator 192 may be an indicator light, which may be used to indicate the charging status, power level changes, messages, missed calls, notifications, etc.
[0132] The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. In the embodiment of the present application, the Android system with a layered architecture is used as an example to illustrate the software structure of the electronic device 100.
[0133] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.
[0134] Figure 2 is a software structure diagram of the electronic device 100 according to an embodiment of the present application. The layered architecture divides the software into several layers, each with clear roles and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers: the application (app) layer, the application framework layer, the Android runtime (Android runtime) and system libraries, and the kernel layer. The application layer may include a series of application packages.
[0135] As shown in FIG2 , the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and short message.
[0136] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0137] As shown in FIG2 , the application framework layer may include a window manager, an activity manager, a package manager, a resource manager, a view system, a telephony manager, a notification manager, and the like.
[0138] Resource Manager, also known as Resource Management Service (RMS), provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.
[0139] A window manager, also known as a window management service (WMS), manages windowed programs. It can determine the display size, determine whether a status bar is present, lock the screen, and take screenshots.
[0140] The activity manager, also known as the activity manager service (AMS), manages all application processes in the system.
[0141] The package manager, also known as the package management service (PMS), is responsible for functions such as application installation and uninstallation, component query and matching, and permission management.
[0142] The view system includes visual controls, such as those for displaying text and images. The view system is used to build applications. A display interface can consist of one or more views. For example, a display interface containing a text notification icon might include a view for displaying text and a view for displaying images.
[0143] The Notification Manager allows applications to display notifications in the status bar. These messages can be displayed briefly and then disappear automatically without user interaction. For example, the Notification Manager is used to notify users of completed downloads and message reminders. The Notification Manager can also display notifications in the top status bar of the system as icons or scrolling text, such as notifications from background applications, or as dialog windows on the screen. Examples include text messages in the status bar, beeps, vibrations on electronic devices, and flashing indicator lights.
[0144] Android runtime includes core libraries and virtual machines, which are responsible for scheduling and management of the Android system.
[0145] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.
[0146] The application layer and application framework layer run in a virtual machine. The virtual machine executes Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.
[0147] The system library can include multiple functional modules, such as a surface manager, media libraries, a 3D graphics processing library (such as OpenGL ES), and a 2D graphics engine (such as SGL).
[0148] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.
[0149] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0150] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0151] A 2D graphics engine is a drawing engine for 2D drawings.
[0152] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.
[0153] Before formally introducing the embodiments of the present application, some terms that may be used in the following content are first explained.
[0154] The k-nearest neighbors algorithm (kNN), also known as the nearest neighbor method, is a nonparametric statistical method for classification and regression. It uses a vector space model for classification. The concept is that cases of the same category are highly similar to each other, and the algorithm can estimate the possible classification of unknown cases by calculating their similarity to known cases.
[0155] Support vector machines (SVMs), also known as support vector networks, are supervised learning models and related learning algorithms used to analyze data in classification and regression analysis. SVM models represent instances as points in space, mapping them so that instances of different classes are separated by as wide a gap as possible. New instances are then mapped into the same space, and their class is predicted based on which side of the gap they fall.
[0156] Random forest refers to a classifier that contains multiple decision trees, and its output category is determined by the mode of the categories output by individual trees.
[0157] A convolutional neural network (CNN) is a feedforward neural network whose artificial neurons can respond to a subset of surrounding cells within their coverage area, making it an excellent choice for processing large images. In at least one layer of a CNN, a mathematical operation called convolution is used instead of standard matrix multiplication.
[0158] A recurrent neural network (RNN) is an artificial neural network that uses sequence data or time series data. It can be used for sequential or time problems such as language translation, natural language processing, and speech recognition.
[0159] An autoencoder, also known as an automatic encoder, is an artificial neural network used to learn efficient encodings for unlabeled data, a form of unsupervised learning. An autoencoder consists of two main components: an encoder, which encodes the input, and a decoder, which uses the encoding to reconstruct the input.
[0160] A chirp signal is a signal whose frequency changes (increases or decreases) over time.
[0161] Automatic speech recognition (ASR), also known as speech-to-text recognition or speech recognition technology, aims to automatically convert human speech into text responses using a computer. Applications of speech recognition technology include voice dialing, voice navigation, in-home device control, and voice document retrieval.
[0162] A proximity sensor is a non-contact sensor that can detect the presence of nearby objects. It typically emits an electromagnetic field or a beam of electromagnetic radiation (such as infrared) and observes changes in the electric field or return signal to achieve its function.
[0163] Voice interaction is an important method for human-computer interaction between users and electronic devices. In the process of user interaction with a voice assistant, in order to improve the efficiency of voice interaction and reduce the power consumption of electronic devices, this application provides a method for voice assistant interaction, which is introduced as follows.
[0164] Figure 3 shows a flow chart of a voice assistant interaction method provided by an embodiment of the present application. Before recording the sound of the user's voice interaction, the electronic device first provides feedback to the user that the voice assistant has been successfully awakened. The way the electronic device sends feedback information to the user is adapted to the device posture of the electronic device. The user's interaction process with the voice assistant is more natural, and the voice assistant is easier to use.
[0165] S110 , in response to detecting the target posture and target action, sending first feedback information according to a target feedback method.
[0166] Here, the target action may refer to an action of converting the electronic device from a non-target posture to a target posture, and the target action may be completed by a user operation.
[0167] In some examples, the target posture may include a first posture, a second posture, or a third posture, wherein the first posture is a posture in which the lower microphone of the electronic device is close to the user's mouth and the display screen of the electronic device is not close to the user's ear; the second posture is a posture in which the lower microphone of the electronic device is close to the user's mouth and the display screen of the electronic device is close to the user's ear; and the third posture is a posture in which the upper microphone of the electronic device is close to the user's mouth.
[0168] FIG4 exemplarily provides schematic diagrams of the aforementioned first posture, second posture, and third posture. FIG401 on the left side is a schematic diagram of the first posture, FIG402 in the middle is a schematic diagram of the second posture, and FIG403 on the right side is a schematic diagram of the third posture. It should be noted that FIG4 is merely for ease of understanding, and the exemplary description of the aforementioned postures should not be regarded as a limitation on the embodiments of the present application. In other words, the aforementioned first posture, second posture, and third posture may also be postures other than those shown in the schematic diagram of FIG4 , and the embodiments of the present application do not impose any limitation on this.
[0169] In some scenarios, the first posture described above can also be understood as: the electronic device's lower microphone is close to the user's mouth and the posture is not for making or receiving a call; or, the electronic device's lower microphone is close to the user's mouth and the ambient light sensor and / or proximity sensor are not blocked. The second posture described above can also be understood as: the electronic device's lower microphone is close to the user's mouth and the posture is for making or receiving a call; or, the electronic device's lower microphone is close to the user's mouth and the ambient light sensor and / or proximity sensor are blocked.
[0170] In some examples, the electronic device can obtain measurement data from one or more of its own sensors, and the measurement data of these sensors can be used to indicate the device posture and device motion of the electronic device. In other words, the electronic device can determine whether a target posture or target motion is detected based on the measurement data of one or more of its own sensors.
[0171] For example, the electronic device can determine whether the electronic device is in the first posture, the second posture, or the third posture based on the measurement data of the gravity sensor. For example, the electronic device can determine the components of gravity acceleration on the three axes of the electronic device based on the measurement data of the gravity sensor, and then determine the angle between the three axes of the electronic device and the horizontal plane based on the components of gravity acceleration on the three axes, and then determine the posture of the electronic device based on the angle.
[0172] In some examples, before determining the device posture of the electronic device, the electronic device can also determine the acceleration of the electronic device based on measurement data of an inertial measurement unit (IMU), and determine whether the electronic device is in a relatively stable state based on the acceleration. When the electronic device is in a relatively stable state, the electronic device determines its own device posture.
[0173] In some examples, when a target posture and target action are detected, the electronic device can use a target verification method to verify whether the electronic device is in the target posture, thereby reducing the probability of the voice assistant being mistakenly awakened or mistakenly touched, and improving the energy utilization efficiency of the electronic device.
[0174] Exemplarily, when the target posture includes the aforementioned first posture or third posture, the electronic device can use a speaker to send a first ultrasonic signal and use a microphone to receive a second ultrasonic signal returned based on the first ultrasonic signal, and verify whether the electronic device is in the target posture based on the recorded second ultrasonic signal and the sent first ultrasonic signal.
[0175] Here, the speaker may refer to the speaker 170A of the electronic device 100 described above, such as the upper speaker or the lower speaker. The microphone may refer to the microphone 170C of the electronic device 100 described above, such as the upper microphone or the lower microphone. For a detailed description of the speaker or microphone positions of the electronic device, please refer to the relevant description of FIG. 6 below, which will not be elaborated here.
[0176] For example, when an electronic device detects that it is in a first posture according to the aforementioned method, the electronic device can use its lower speaker to send a first ultrasonic signal, use its lower microphone to record a second ultrasonic signal, and verify whether the electronic device is in the first posture based on the first ultrasonic signal and the second ultrasonic signal.
[0177] For another example, when an electronic device detects that it is in the third posture according to the aforementioned method, the electronic device can use its upper speaker to send a first ultrasonic signal, use its upper microphone to record a second ultrasonic signal, and verify whether the electronic device is in the second posture based on the first ultrasonic signal and the second ultrasonic signal.
[0178] The method of using the transmitted ultrasonic signal and the recorded ultrasonic signal to verify whether the electronic device is in the target posture will be described in detail in the following embodiments and will not be expanded here.
[0179] The power consumption of sending ultrasound and recording audio is higher than that of the aforementioned IMU and gravity sensor, but the accuracy of the measurement data in determining the posture of the electronic device is higher. Using this method to verify the posture of the electronic device is beneficial to controlling the power consumption of the electronic device while improving the accuracy of target posture detection, and is beneficial to reducing the false touch rate during the use of the voice assistant.
[0180] Also exemplarily, when the target posture includes the aforementioned second posture, the electronic device may obtain touch information of the display screen and verify whether the electronic device is in the second posture based on the touch information.
[0181] In some examples, the touch information may include information about the user's auricle. For example, the touch information may include a capacitance value of a display screen of an electronic device, and the capacitance value may be used to indicate information about the user's auricle.
[0182] When the electronic device is in the second posture, the user's ear is relatively close to the electronic device display screen. Touch information from the electronic device display screen can be used to determine whether the user's ear is relatively close to the electronic device display screen, thereby verifying whether the electronic device is in the second posture. The method for verifying whether the electronic device is in the second posture using touch information from the display screen will be described in detail in the following embodiments and will not be elaborated here.
[0183] The power consumption of obtaining touch information from the display screen is higher than that of the aforementioned IMU and gravity sensor, but the accuracy of its measurement data in determining the posture of the electronic device is higher. Using this method to verify the posture of the electronic device is beneficial to controlling the power consumption of the electronic device on the basis of improving the accuracy of target posture detection, and is beneficial to reducing the false touch rate during the use of the voice assistant.
[0184] In some examples, when a target posture and target motion are detected, the electronic device may also verify whether the electronic device is in the target posture by reading measurement data from a proximity sensor and / or measurement data from an ambient light sensor. When the electronic device is in the first, second, and third postures described above, the corresponding measurement data from the proximity sensor and measurement data from the ambient light sensor have different values. The method for using the measurement data from the proximity sensor and / or measurement data from the ambient light sensor to verify whether the electronic device is in the target posture will be described in detail below and will not be elaborated on here.
[0185] In some examples, the electronic device may combine multiple of the aforementioned target posture determination and verification methods to determine whether the electronic device is in the target posture, such as combining IMU measurement data and ultrasonic signals, combining IMU measurement data, proximity sensor measurement data, and touch information from the display screen, or combining gravity sensor measurement data, ultrasonic signals, proximity sensor measurement data, and ambient light sensor measurement data.
[0186] Determining whether an electronic device is in a target posture through multiple methods and measurement data from multiple sensors of the electronic device is beneficial to reducing the probability of accidental touches during the use of the voice assistant and improving the energy utilization efficiency of the electronic device.
[0187] The first feedback information is used to indicate that the voice assistant is successfully awakened. For electronic devices in different postures, the method for the electronic device to send the first feedback information may be different, or in other words, the target feedback method for sending the first feedback information may be determined according to the device posture of the electronic device (i.e., the target posture).
[0188] In some examples, the target posture may include the aforementioned first posture, and the first feedback information may be a prompt screen and / or a first prompt audio. The electronic device may display a prompt screen on a display screen and / or play a first prompt audio through a speaker, and the prompt screen and / or the first prompt audio are used to indicate that the voice assistant is successfully awakened. For example, the electronic device may play the first prompt audio through a lower speaker.
[0189] In some examples, the target gesture may include the aforementioned second gesture, and the first feedback information may be a second prompt audio. The electronic device may play the second prompt audio through a speaker, and the second prompt audio may be used to indicate that the voice assistant is successfully awakened. For example, the electronic device may play the second prompt audio through an upper speaker.
[0190] In some examples, the target posture may include the aforementioned third posture, and the first feedback information may be a first prompt audio. The electronic device may play the first prompt audio through a speaker, and the first prompt audio may be used to indicate that the voice assistant is successfully awakened. For example, the electronic device may play the first prompt audio through an upper speaker.
[0191] The method of sending the first feedback information can adapt to different postures of the electronic device, or in other words, the method of sending the first feedback information is more natural and more in line with the user's operating habits. This feedback method is conducive to making the user's human-computer interaction process through the voice assistant more natural.
[0192] It should be understood that the electronic device can also use more other ways to send the above-mentioned first feedback information, such as one or more of vibration, display light effect, flash light effect, etc., and this application does not limit this.
[0193] S120: Record the voice command according to the target sound pickup method.
[0194] After sending the feedback information, the voice command recording begins. The microphone of the electronic device does not need to be in a constantly open state, which is beneficial to reducing the power consumption of the electronic device, improving the energy utilization efficiency of the electronic device, and reducing the probability of the voice assistant being woken up by mistake, making the interaction process between the user and the voice assistant more natural.
[0195] The target sound pickup method refers to a method or strategy for the electronic device to record voice commands, and the target sound pickup method can be determined according to the target posture of the electronic device in S110.
[0196] In some examples, the target gesture includes the aforementioned first gesture or second gesture, and the target sound pickup method may be: recording the voice command using at least the lower microphone of the electronic device. For example, the lower microphone of the electronic device serves as the primary recording unit, and the upper microphone of the electronic device serves as the secondary recording unit for recording the voice command.
[0197] In some examples, the target posture includes the aforementioned third posture, and the target sound pickup method may be: recording the voice command at least through the upper microphone of the electronic device. For example, the upper microphone of the electronic device is used as the main recording unit and the lower microphone of the electronic device is used as the auxiliary recording unit to record the voice command.
[0198] Exemplarily, here, the auxiliary recording unit can be used to perform noise reduction processing on the audio recorded by the primary recording unit.
[0199] The target pickup method is determined according to the device posture of the electronic device. On the one hand, it is beneficial to improve the accuracy of the voice commands recorded by the electronic device, and to improve the efficiency of the voice assistant in executing the user's voice commands. To a certain extent, it is also beneficial to improve the energy utilization efficiency of the electronic device.
[0200] S130: In response to the voice command, send second feedback information according to the target feedback method.
[0201] Based on the recorded voice commands, the electronic device can determine the voice commands sent by the user, and can convert the voice commands into text information through automatic voice recognition. Based on the text information, the electronic device can call a voice assistant to complete the command indicated by the text information.
[0202] Here, the method of sending the second feedback information, ie, the target feedback method, may be determined according to the device posture of the electronic device.
[0203] In some examples, the target gesture may include the aforementioned first gesture, and the second feedback information may be a result prompt screen and / or a result prompt audio. The electronic device may display a result prompt screen on a display screen and / or play a result prompt audio through a speaker, where the result prompt screen and / or result prompt audio are used to indicate the execution result of the voice command. For example, the electronic device may play the result prompt audio through a lower speaker.
[0204] In some examples, the target gesture may include the aforementioned second gesture, and the second feedback information may be a result prompt audio. The electronic device may play the result prompt audio through a speaker, and the result prompt audio may be used to indicate the execution result of the voice command. For example, the electronic device may play the result prompt audio through an upper speaker.
[0205] In some examples, the target gesture may include the aforementioned third gesture, and the second feedback information may be a result prompt audio. The electronic device may play the result prompt audio through a speaker, and the result prompt audio may be used to indicate the execution result of the voice command. For example, the electronic device may play the result prompt audio through a speaker.
[0206] It should be understood that the execution result here can include a successful execution result, an unsuccessful execution result, or a result that cannot be executed. In addition, the execution result can also be an execution result that indicates that the electronic device cannot detect the user's voice within a preset time period and instructs to stop or pause the voice assistant interaction.
[0207] The method of sending the second feedback information can adapt to different postures of the electronic device, or in other words, the method of sending the second feedback information is more natural and more in line with the user's operating habits. This feedback method is conducive to making the user's human-computer interaction process through the voice assistant more natural.
[0208] It should be understood that the electronic device can also use more other ways to send the above-mentioned second feedback information, such as: one or more of vibration, display light effect, flash light effect, indicator light effect, etc., and this application does not limit this.
[0209] For different device postures, the electronic device can adopt different wake-up schemes, sound pickup schemes and feedback schemes. As shown in Figure 5, for the electronic device in a situation where the lower microphone is close to the user's mouth and the user is not making or receiving a call (the posture of the electronic device 100 in Figure 401), the electronic device can execute the following interaction scheme.
[0210] S201 , determining that the electronic device is in a posture where the lower microphone is close to the user's mouth.
[0211] In some examples, while in a relatively stable state, the electronic device determines whether it is in a posture where the lower microphone is close to the user's mouth.
[0212] Whether the electronic device is in a relatively stable state can be determined based on one or more sensors on the electronic device.
[0213] For example, the electronic device can determine whether it is in a relatively stable state based on the measurement data of the IMU. For example, the electronic device can determine the current acceleration of the electronic device based on the measurement data of the IMU. If the acceleration of the electronic device is less than or equal to an acceleration threshold, the electronic device can determine that it is in a relatively stable state. Conversely, if the acceleration of the electronic device is greater than the acceleration threshold, the electronic device can determine that it is in motion.
[0214] It should be noted that the relatively stable state here can also be understood as a relatively static state. Stability, static and movement are all relative descriptions.
[0215] In some examples, the posture of the electronic device can be determined based on one or more sensors on the electronic device.
[0216] For example, an electronic device can determine its own device posture based on measurement data from a gravity sensor. For example, the electronic device can determine the components of the electronic device's gravitational acceleration on the electronic device's three axes based on the measurement data from the gravity sensor, and determine the angle between the three axes of the electronic device and the horizontal plane based on the components of the gravitational acceleration on the aforementioned three axes. This angle can be used to indicate the device posture of the electronic device. For ease of explanation, the following content refers to "the electronic device in a posture where the lower microphone is close to the user's mouth and is not in a call-making state" as the first target posture.
[0217] Taking the three-dimensional schematic diagram of an electronic device shown in Figure 6 as an example, the three axes of the electronic device can be referred to as the x-axis, y-axis, and z-axis, respectively. The x-axis direction can be considered the width direction of the electronic device, the y-axis direction can be considered the height direction of the electronic device, and the z-axis direction can be considered the thickness direction of the electronic device. The quadrilateral region 601 in Figure 6 can be considered a horizontal plane, and the angles between the x-axis, y-axis, and z-axis of the electronic device and the horizontal plane can be referred to as a first angle (α1), a second angle (α2), and a third angle (α3), respectively.
[0218] Electronic device 100 may include an upper microphone 201A, a lower microphone 201B, an upper speaker 202A, a lower speaker 202B, a proximity sensor 203, and an ambient light sensor 204. In FIG6 , upper microphone 201A, upper speaker 202A, proximity sensor 203, and ambient light sensor 204 may all be located in the upper portion of electronic device 100, i.e., in the region of the electronic device on the positive y-axis. Lower microphone 201B and lower speaker 202B may be located in the lower portion of electronic device 100, i.e., in the region of the electronic device on the negative y-axis.
[0219] It should be understood that the position distribution of the multiple electronic components of the electronic device 100 in FIG6 is only an example, and the present application does not impose any limitation thereto.
[0220] In some examples, when the aforementioned first angle, second angle, and third angle simultaneously meet the preset conditions, the electronic device can determine that it is in the first target posture; on the contrary, when any of the aforementioned first angle, second angle, and third angle does not meet the preset conditions, the electronic device can determine that it is not in the first target posture.
[0221] S202: Determine whether the user has moved the lower microphone of the electronic device close to the mouth.
[0222] In order to improve the accuracy of electronic devices in determining the user's intention to wake up the voice assistant, reduce the probability of the electronic device's voice assistant being mistakenly woken up, and reduce the increase in power consumption caused by the use of the voice assistant, when determining that the electronic device is in the first target posture, the electronic device can determine whether it has undergone a change from a non-first target posture to the first target posture. In other words, the electronic device can determine whether the user has adjusted the electronic device from the non-first target posture to the first target posture. For ease of explanation, the action of "adjusting the electronic device from the non-first target posture to the first target posture" is hereinafter referred to as the first target action.
[0223] In some examples, the electronic device may store sensor data for a period of time and determine whether the user has performed the first target action based on the sensor data within the period of time. For example, the electronic device may continuously store IMU measurement data for a certain period of time (e.g., 2 seconds or 1 second, etc.) and determine whether the user has performed the first target action based on the IMU measurement data within the period of time.
[0224] For example, based on the IMU measurement data within the time interval, the electronic device can use a machine learning-based classification algorithm and / or a rule-based judgment algorithm to determine whether the user has performed the first target action.
[0225] Classification algorithms based on machine learning may include one or more of: traditional machine learning methods such as nearest neighbor algorithms, support vector machines, random forests, and deep learning methods such as convolutional neural networks or autoencoders.
[0226] In some examples, the electronic device uses one or more of the above-mentioned classification algorithms to obtain a probability value of "detecting the first target action". If the probability value is greater than or equal to a preset threshold, the electronic device determines that the user has performed the first target action and continues to perform S203 and subsequent operations. If the probability value is less than the preset threshold, the electronic device determines that the user has not performed the first target action and can restart S201.
[0227] In some examples, rule-based judgment algorithms can first classify actions into large and small movements based on the magnitude of acceleration and the duration of the movement. For example, a large movement might involve lifting the electronic device from a tabletop or trouser pocket to the mouth, while a small movement might involve rotating the electronic device from a position close to the face to a certain angle so that the lower microphone of the electronic device is close to the user's mouth.
[0228] For example, the electronic device can determine its own displacement based on acceleration and motion time. If the displacement is greater than or equal to a displacement threshold, it is determined that a significant motion has occurred; if the displacement is less than the displacement threshold, it is determined that a small motion has occurred. For different motion situations, the electronic device can use different methods to determine whether the first target action has been detected. Figure 7 shows one determination method provided by this application.
[0229] For large movements, the distance of vertical upward movement can identify the action of "lifting up", such as lifting the electronic device from near the user's thigh to the mouth. The distance of horizontal movement can identify the action of "pulling the electronic device towards the user", or in other words, the action of moving the electronic device from a position away from the user to a position close to the user, such as moving the electronic device from a position away from the user on a table in front of the user to a position close to the user.
[0230] In some examples, in the event of a substantial movement, the electronic device may determine whether the electronic device has detected the first target action based on its own displacement component in the direction of gravity or the displacement component in the horizontal direction.
[0231] For example, the electronic device can determine the component of acceleration in the direction of gravity based on the angle between the direction of its own acceleration and the direction of gravity. The distance the electronic device has moved in the direction of gravity can be obtained by performing a quadratic integration of the acceleration in the direction of gravity over the movement time. If the distance moved in the direction of gravity is greater than or equal to a vertical displacement threshold, the electronic device determines that the first target action has been detected; if the distance moved in the direction of gravity is less than the vertical displacement threshold, the electronic device determines that the first target action has not been detected.
[0232] Similarly, by way of example, the electronic device can determine its first horizontal acceleration component and second horizontal acceleration component on the horizontal plane based on its own acceleration along the negative direction of the y-axis and the positive direction of the z-axis, and determine the composite acceleration of the electronic device on the horizontal plane based on the first horizontal acceleration component and the second horizontal acceleration component. By performing a quadratic integration of the composite acceleration over the motion time, the movement distance of the electronic device on the horizontal plane can be obtained. If the movement distance on the horizontal plane is greater than or equal to the horizontal displacement threshold, the electronic device determines that the first target action has been detected; if the movement distance on the horizontal plane is less than the horizontal displacement threshold, the electronic device determines that the first target action has not been detected.
[0233] Here, since the negative y-axis direction and the positive z-axis direction of the electronic device are usually pointing towards the user, when calculating the horizontal distance moved, only the acceleration in the negative y-axis direction and the positive z-axis direction can be considered.
[0234] For small movements, the action can be identified by the change in the angle between the electronic device's three axes and the horizontal plane. The magnitude of the angle change should be within a certain range. In some examples, because rotating the electronic device close to the user's mouth generally does not involve reciprocating motion, the change in the angle between the y-axis and the horizontal plane and the change in the angle between the x-axis and the horizontal plane are monotonically increasing or monotonically decreasing.
[0235] In some examples, when a small movement occurs, the electronic device can determine whether a target action is detected based on the angles (first angle α1, second angle α2, and third angle α3) between the three axes of the electronic device and the horizontal plane during the movement.
[0236] For example, the electronic device can determine, during its own motion, a first angle change value (β1) corresponding to α1, a second angle change value (β2) corresponding to α2, and a third angle change value (β3) corresponding to α3. Based on the first angle change value, the second angle change value, and the third angle change value, the electronic device can determine whether a target action is detected. For ease of explanation, the first angle threshold (δ1), the second angle threshold (δ2), and the third angle threshold (δ3) are used here to represent the angle thresholds corresponding to β1, β2, and β3, respectively.
[0237] For example, in the case that β1≤δ1, β2≤δ3, and β3≤δ3, the electronic device may determine that the first target action is detected; otherwise, the electronic device may determine that the first target action is not detected.
[0238] Also exemplarily, the electronic device can continuously record the first angle and the second angle during its own movement, and then determine the first angle increment (γ1) of the first angle and the second angle increment (γ2) of the second angle at different times. Based on the multiple first angle increments and multiple second angle increments at different times during the movement, the electronic device can determine whether the target action is detected.
[0239] For example, during the movement, if all γ1s are positive or negative, and all γ2s are positive or negative, the electronic device can determine that the first target action is detected; otherwise, the electronic device can determine that the first target action is not detected.
[0240] Still illustratively, the electronic device may also combine the above two determination methods to determine whether a target operation has been detected. That is, the electronic device may determine whether a target action has been detected based on the first angle change value, the second angle change value, and the third angle change value corresponding to the start and end times of the motion, as well as multiple first angle increments and multiple second angle increments at different times during the motion.
[0241] For example, when β1≤δ1, β2≤δ3 and β3≤δ3, and all γ1s are positive or negative, and all γ2s are positive or negative, the electronic device can determine that the first target action is detected; otherwise, the electronic device can determine that the first target action is not detected.
[0242] S203 , determining whether the measurement data of the proximity sensor and the measurement data of the ambient light sensor meet the requirements.
[0243] The lower microphone of the electronic device is close to the user's mouth and is not in the posture for making or receiving calls, and the upper part of the electronic device is not blocked. In this case, the distance determined based on the measurement data of the proximity sensor of the electronic device is greater than or equal to the distance threshold, that is, the measurement data of the proximity sensor will not indicate that the upper part of the electronic device is blocked, and the illuminance determined based on the measurement of the ambient light sensor will generally be greater than or equal to the illuminance threshold. In other words, when the reading of the proximity sensor and the measurement data of the ambient light sensor both meet the aforementioned preset conditions, the electronic device is more likely to be in the first target posture. On the contrary, if one or both of the detection data of the proximity sensor and the measurement data of the ambient light sensor do not meet the preset conditions, the electronic device is less likely to be in the first target posture. For example, in this case, the electronic device may be in a pocket or handbag.
[0244] If it is determined that the measurement data of the proximity sensor and the measurement data of the ambient light sensor meet the requirements, the electronic device can perform S204 and subsequent operations. If it is determined that the measurement data of the proximity sensor and / or the measurement data of the ambient light sensor do not meet the requirements, the electronic device can re-execute S201.
[0245] S204: The lower speaker of the electronic device sends an ultrasonic signal, and the lower microphone detects the ultrasonic signal.
[0246] When it is determined that the electronic device is in the first target posture and the electronic device detects the first target action, the electronic device may execute S204 , ie, control the lower speaker to send an ultrasonic signal and the lower microphone to detect the ultrasonic signal.
[0247] When it is determined that the electronic device is not in the first target posture, or when the electronic device does not detect the first target action, the electronic device may re-execute the operation of S201.
[0248] S205 , determining the distance between the lower microphone and the user's mouth and obtaining feature information of the user's mouth.
[0249] Based on the ultrasonic signal detection result of the lower microphone in S204, the distance between the lower microphone of the current electronic device and the user's mouth is determined, and characteristic information of the user's mouth is obtained. Because different points on the user's mouth are at different distances from the microphone, different points have different effects on the reflection or absorption of ultrasonic signals. This information is included in the audio clip of the mobile phone after the ultrasonic signal is reflected by the user's mouth and then recorded by the microphone. In other words, the characteristic information of the user's mouth can refer to the characteristic information of the user's mouth contained in the spectrum of the ultrasonic signal detected by the microphone after the ultrasonic signal is blocked by the user's mouth and reflected.
[0250] FIG8 shows a processing method for determining the device posture of an electronic device using ultrasonic signals provided in an embodiment of the present application.
[0251] The electronic device emits a chirp signal of a preset frequency and duration, records audio for a certain duration, filters the recorded audio, and extracts a spectrogram of the audio segment from the recorded audio as input to a classifier. The classifier's output can be used to determine the distance between the electronic device's lower microphone and the user's mouth. The classifier can be an RNN, CNN, or SVM, among others.
[0252] Specifically, in some examples, the electronic device can generate a chirp signal of 0.02s with a frequency range of 18kHz to 23kHz, and send it for 0.2s at a time interval of 0.02s. The electronic device can record 100ms of audio and apply a bandpass filter of 18kHz to 23kHz to the recorded audio signal. Matched filtering is performed using the original chirp signal and the recorded audio signal, and the maximum position in the result of the matched filtering can be regarded as the starting position of the chirp signal. A 0.02s audio segment is cut from the starting position of the recorded audio, and the spectrum of the cut audio segment is calculated using short-time Fourier transform or fast Fourier transform as the input of the classifier. The classifier can determine the classification result of each complete signal based on the input, and the classification result can be used to determine the distance between the lower microphone and the user's mouth.
[0253] If the ultrasonic signal determines that the lower microphone of the electronic device is close to the user's mouth, the electronic device may execute S206, that is, after opening the voice assistant application, feedback that the voice assistant is successfully woken up. Conversely, if the ultrasonic signal does not indicate that the lower microphone of the electronic device is close to the user's mouth, the electronic device may re-execute the operation of S201.
[0254] S206: Feedback indicating successful voice assistant wake-up.
[0255] In some examples, as shown in Figure 9, the electronic device can display a first prompt message 210 through a display screen. The first prompt message 210 can also be called a first prompt screen 210. The first prompt message 210 is used to prompt that the voice assistant is successfully awakened. The first prompt message 210 can include text information and / or light information (or light effect information), etc.
[0256] In some examples, as shown in FIG9 , the electronic device may play a first prompt audio 220 through the lower speaker, where the first prompt audio 220 is used to prompt that the voice assistant is successfully woken up.
[0257] In some examples, the electronic device can simultaneously send a prompt message through the display screen and the lower speaker to indicate that the voice assistant has been successfully woken up.
[0258] The electronic device can also prompt the successful wake-up of the voice assistant through other methods, such as vibration information, etc. The electronic device can also change the feedback method according to the user's customized operation. This application does not impose any restrictions on this.
[0259] When the electronic device is in the first target posture, the prompt information can be more easily perceived by the user by displaying the prompt information on the display screen and / or playing the audio prompt information using the lower speaker. This feedback method makes the user's interaction with the voice assistant more natural, which is conducive to improving the usability of the voice assistant.
[0260] S207 , using the dual-channel audio feature, determine that the user is speaking close to the lower microphone.
[0261] In order to determine that the audio recorded by the microphone comes from the user of the electronic device, or in other words, to determine that it is the user of the electronic device who is speaking, in some examples, while providing feedback that the voice assistant has been successfully woken up, the electronic device can record through the upper microphone and the lower microphone, and use the dual-channel audio characteristics to determine that the user is speaking close to the lower microphone.
[0262] In some examples, the electronic device can combine the audio information recorded by the upper microphone and the lower microphone to determine that the user is speaking close to the lower microphone. Figure 10 shows a method provided by an embodiment of the present application to combine the audio information recorded by the upper microphone and the lower microphone to determine that the user is speaking close to the lower microphone.
[0263] For example, the electronic device can save an audio signal with a length of 200ms and make a judgment every 100ms. The multiple judgment results can be saved as a historical cache. In some examples, the electronic device can separate the audio recorded by the upper microphone and the lower microphone respectively, and use a low-pass filter with a cutoff frequency of 17kHz to remove the signal in the ultrasonic frequency band. The dual-channel audio signal is then used as the input of a CNN-based classifier, and the classification result output by the classifier is used to determine whether the user is speaking close to the lower microphone of the electronic device.
[0264] For example, the output results of the aforementioned CNN-based classifier can give the probabilities of the following three situations, and the situation where the probability is greater than or equal to the preset threshold is considered to be detected.
[0265] Category 1: The upper microphone of the electronic device is close to the user's mouth. For example, the distance between the upper microphone of the electronic device and the user's mouth is within 5 cm, and the user holds the electronic device and speaks.
[0266] Category 2: The lower microphone of the electronic device is close to the user's mouth. For example, the distance between the lower microphone of the electronic device and the user's mouth is within 5 cm, and the user holds the electronic device and speaks.
[0267] Category 3: Situations where the electronic device is not close to the user's mouth, or other situations, for example, the user uses the electronic device normally, such as talking while looking at the screen, talking in an environment with other people around, talking in a quiet environment, or talking in a noisy environment.
[0268] In some examples, the electronic device may also save the most recent n judgment results, vote on different types of results in the most recent n judgment results through a voting mechanism, and regard the result of the majority vote as the final output, where n is an integer greater than or equal to 1.
[0269] S208: Select a microphone according to the sound pickup strategy and perform automatic speech recognition.
[0270] When the lower microphone of an electronic device is close to the user's mouth, the electronic device can record the user's voice commands through at least the lower microphone, for example, implementing a sound pickup strategy with the lower microphone as the primary and the upper microphone as the secondary. For example, the electronic device can use the audio recorded by the upper microphone to perform auxiliary noise reduction processing on the audio recorded by the lower microphone.
[0271] S209: The voice assistant of the electronic device executes the command and generates feedback according to the corresponding feedback method.
[0272] After the electronic device calls the language assistant to execute the command, it can feedback the command execution result in a similar manner to the feedback of successful awakening of the language assistant in S206. For details, please refer to the content of S206. For the sake of brevity, it will not be repeated here.
[0273] In some examples, the electronic device may use different feedback methods for different execution results, such as different types of feedback methods, or feedback methods of the same type with different contents. This application does not impose any restrictions on this.
[0274] FIG11 shows a method for a user to interact with a voice assistant when the upper microphone of the electronic device is close to the user's mouth (the posture of the electronic device 100 in FIG403 ).
[0275] S301 , determining that the electronic device is in a posture where the upper microphone is close to the user's mouth.
[0276] In some examples, when in a relatively stable state, the electronic device determines whether its posture is a state in which the upper microphone is close to the user's mouth.
[0277] Whether the electronic device is in a relatively stable state can be determined based on one or more sensors on the electronic device. For example, the electronic device can determine based on measurement data from an IMU.
[0278] In some examples, the posture of the electronic device can be determined based on one or more sensors on the electronic device. For example, the electronic device can determine the posture of the electronic device based on measurement data from a gravity sensor. For example, the electronic device can determine the components of the electronic device's gravitational acceleration on the electronic device's three axes based on the measurement data from the gravity sensor, and determine the angle between the three axes of the electronic device and the horizontal plane based on the components of the gravitational acceleration on the aforementioned three axes. This angle can be used to indicate the posture of the electronic device. For ease of explanation, the following content will refer to "the electronic device being in a posture where the lower microphone is close to the user's mouth" as the third target posture.
[0279] Taking the three-dimensional schematic diagram of an electronic device shown in Figure 6 as an example, the angles between the x-axis, y-axis, and z-axis of the electronic device and the horizontal plane can be referred to as a first angle (α1), a second angle (α2), and a third angle (α3), respectively. In some examples, when the first angle, the second angle, and the third angle all meet a preset condition, the electronic device can determine that it is in the third target posture; conversely, when any of the first angle, the second angle, and the third angle do not meet the preset condition, the electronic device can determine that it is not in the third target posture.
[0280] The specific execution method of S301 is similar to that of S201. For details, please refer to the relevant content in S201. For the sake of brevity, it will not be described here.
[0281] S302: Determine whether the user has moved the upper microphone of the electronic device close to the mouth.
[0282] In order to improve the accuracy of electronic devices in determining the user's intention to wake up the voice assistant, reduce the probability of the electronic device's voice assistant being mistakenly woken up, and reduce the increase in power consumption caused by the use of the voice assistant, when it is determined that the electronic device is in the third target posture, the electronic device can determine whether it has undergone a change from a non-third target posture to a third target posture. In other words, the electronic device can determine whether the user has adjusted the electronic device from a non-third target posture to a third target posture. For ease of explanation, the action of "adjusting the electronic device from a non-third target posture to a third target posture" is hereinafter referred to as the third target action.
[0283] In some examples, the electronic device may store sensor data for a period of time and determine whether the user has performed the third target action based on the sensor data within the period. For example, the electronic device may continuously store IMU measurement data for a certain period of time (e.g., 2 seconds or 1 second, etc.) and determine whether the user has performed the aforementioned action based on the IMU measurement data within the period of time.
[0284] For example, based on the IMU measurement data within the time interval, the electronic device can use a machine learning-based classification algorithm and / or a rule-based judgment algorithm to determine whether the user has performed the third target action.
[0285] Classification algorithms based on machine learning may include one or more of: traditional machine learning methods such as nearest neighbor algorithms, support vector machines, random forests, and deep learning methods such as convolutional neural networks or autoencoders.
[0286] In some examples, the electronic device uses one or more of the above-described classification algorithms to obtain a probability value for "detecting the third target action." If the probability value is greater than or equal to a preset threshold, the electronic device determines that the user has performed the third target action and continues to perform S203 and subsequent operations. If the probability value is less than the preset threshold, the electronic device determines that the user has not performed the third target action and may restart S301.
[0287] For rule-based judgment algorithms, the electronic device can determine whether the third target action is detected based on its vertical movement distance and / or rotation angle around the x-axis. FIG12 exemplarily provides a method for an electronic device to determine whether the third target action is detected.
[0288] In some examples, the electronic device may determine that a third target action has been detected based on whether its own vertical movement distance is greater than a vertical displacement threshold. If it is determined that the vertical movement distance is greater than or equal to the vertical displacement threshold, the electronic device may determine that the third target action has been detected; if it is determined that the vertical movement distance is less than the vertical displacement threshold, the electronic device may determine that the third target action has not been detected.
[0289] In some examples, the electronic device may determine that a third target action has been detected based on whether the angle through which it has rotated around the x-axis is greater than an angle threshold. If it is determined that the angle through which it has rotated around the x-axis is greater than or equal to the angle threshold, the electronic device may determine that a third target action has not been detected. If it is determined that the angle through which it has rotated around the x-axis is less than the angle threshold, the electronic device may determine that a third target action has not been detected.
[0290] S303 , determining whether the measurement data of the proximity sensor and the measurement data of the ambient light sensor meet the requirements.
[0291] In some examples, when the upper microphone of the electronic device is close to the user's mouth, the proximity sensor is blocked and the reading of the ambient light sensor is reduced, or in other words, the distance detected by the proximity sensor is less than or equal to a distance threshold, and the illuminance detected by the ambient light sensor is less than or equal to an illuminance threshold. In other words, if both the reading of the proximity sensor and the measurement data of the ambient light sensor meet preset conditions, it is more likely that the electronic device is in a posture where the upper microphone is close to the user's mouth. Conversely, if one or both of the detection data of the proximity sensor and the measurement data of the ambient light sensor do not meet the preset conditions, it is less likely that the electronic device is in a posture where the upper microphone is close to the user's mouth. For example, in this case, the electronic device may be in a pocket or handbag.
[0292] If it is determined that the measurement data of the proximity sensor and the measurement data of the ambient light sensor meet the requirements, the electronic device can execute S304 and subsequent operations. If it is determined that the measurement data of the proximity sensor and / or the measurement data of the ambient light sensor do not meet the requirements, the electronic device can re-execute S301.
[0293] S304: The upper speaker of the electronic device sends an ultrasonic signal, and the upper microphone detects the ultrasonic signal.
[0294] When it is determined that the electronic device is in the third target posture and the electronic device detects the third target action, the electronic device may execute S304 , ie, control the upper speaker to send an ultrasonic signal and the upper microphone to detect the ultrasonic signal.
[0295] When it is determined that the electronic device is not in the third target posture, or when the electronic device does not detect the third target action, the electronic device may re-execute the operation of S301.
[0296] S305 : Determine the distance between the upper microphone and the user's mouth and obtain feature information of the user's mouth.
[0297] Based on the ultrasonic signal detection result of the upper microphone in S304, the distance between the upper microphone of the current electronic device and the user's mouth is determined, and characteristic information of the user's mouth is obtained. Because different points on the user's mouth are at different distances from the microphone, different points have different effects on the reflection or absorption of ultrasonic signals. This information is included in the audio clip of the ultrasonic signal reflected by the user's mouth and then recorded by the microphone. In other words, the characteristic information of the user's mouth can refer to the characteristic information of the user's mouth contained in the spectrum of the ultrasonic signal detected by the microphone after the ultrasonic signal is blocked by the user's mouth and reflected.
[0298] The electronic device emits a chirp signal of a preset frequency and duration, records audio for a certain duration, filters the recorded audio, and extracts a spectrogram of the audio segment from the recorded audio as input to a classifier. The classifier's output can be used to determine the distance between the electronic device's lower microphone and the user's mouth. The classifier can be an RNN, CNN, or SVM, among others.
[0299] The specific execution method of S305 is similar to that of S205. For details, please refer to the relevant content in S205. For the sake of brevity, it will not be described here.
[0300] If the ultrasonic signal determines that the upper microphone of the electronic device is close to the user's mouth, the electronic device may execute S306, that is, after opening the voice assistant application, feedback that the voice assistant is successfully woken up. Conversely, if the ultrasonic signal does not indicate that the upper microphone of the electronic device is close to the user's mouth, the electronic device may re-execute the operation of S301.
[0301] S306: Feedback that the voice assistant is successfully awakened.
[0302] In some examples, the electronic device can display a second prompt message through the display screen. The second prompt message can also be called a second prompt screen. The second prompt message is used to prompt that the voice assistant has been successfully awakened. The second prompt message can be text information or light information (or light effect information), etc.
[0303] In some examples, the electronic device can play a third prompt audio through the upper speaker, and the third prompt audio is used to prompt that the voice assistant is successfully woken up.
[0304] In some examples, the electronic device can simultaneously send a prompt message through the display screen and the upper speaker to indicate that the voice assistant has been successfully awakened.
[0305] The electronic device can also prompt the successful wake-up of the voice assistant through other methods, such as vibration information, etc. The electronic device can also change the feedback method according to the user's customized operation. This application does not impose any restrictions on this.
[0306] When the electronic device is in the third target posture, the prompt information can be more easily perceived by the user by playing audio prompt information through the upper speaker. This feedback method makes the user's interaction with the voice assistant more natural, which is conducive to improving the usability of the voice assistant.
[0307] S307 , using the dual-channel audio feature, determine that the user is speaking close to the lower microphone.
[0308] In order to determine that the audio recorded by the microphone comes from the user of the electronic device, or in other words, to determine that the person speaking is the user of the electronic device, in some examples, while providing feedback that the voice assistant has been successfully woken up, the electronic device can simultaneously record through the upper microphone and the lower microphone, and use the dual-channel audio feature to determine that the user is speaking close to the upper microphone.
[0309] In some examples, the electronic device can combine audio information recorded by the upper microphone and the lower microphone to determine that the user is speaking close to the lower microphone.
[0310] For example, the electronic device can save an audio signal with a length of 200ms and make a judgment every 100ms. The multiple judgment results can be saved as a historical cache. In some examples, the electronic device can separate the audio recorded by the upper microphone and the lower microphone respectively, and use a low-pass filter with a cutoff frequency of 17kHz to remove the signal in the ultrasonic frequency band. The dual-channel audio signal is then used as the input of a CNN-based classifier, and the classification result output by the classifier is used to determine whether the user is speaking close to the lower microphone of the electronic device.
[0311] In some examples, the electronic device may also save the most recent n output results, vote on different types of results in the most recent n output results through a voting mechanism, and regard the result of the majority vote as the final output, where n is an integer greater than or equal to 1.
[0312] The specific execution method of S307 is similar to that of S207. For details, please refer to the relevant content in S207. For the sake of brevity, it will not be described here.
[0313] S308: Select a microphone according to the sound pickup strategy and perform automatic speech recognition.
[0314] When the upper microphone of an electronic device is close to the user's mouth, the electronic device can record the user's voice commands through at least the upper microphone, for example, implementing a sound pickup strategy with the upper microphone as the primary and the lower microphone as the secondary. For example, the electronic device can use the audio recorded by the lower microphone to perform auxiliary noise reduction processing on the audio recorded by the upper microphone.
[0315] S309: The voice assistant of the electronic device executes the command and generates feedback according to the corresponding feedback method.
[0316] After the electronic device calls the language assistant to execute the command, it can feedback the completion of the command execution in a similar manner to the feedback of the successful awakening of the language assistant in S306. For details, please refer to the content of S306. For the sake of brevity, it will not be repeated here.
[0317] In some examples, the electronic device may use different feedback methods for different execution results, such as different types of feedback methods, or feedback methods of the same type with different contents. This application does not impose any restrictions on this.
[0318] Figure 13 shows another voice assistant interaction method provided by an embodiment of the present application. In this method, the electronic device can determine that the electronic device is in a posture for making or receiving a call (the posture of the electronic device 100 in Figure 402) based on its own sensors, and thus execute corresponding wake-up schemes, sound pickup schemes, and feedback schemes based on the posture.
[0319] S401: Determine whether the electronic device is in a posture for making or receiving a call.
[0320] In some examples, when in a relatively stable state, the electronic device can determine whether its posture is that of making or receiving a phone call.
[0321] Whether an electronic device is in a relatively stable state can be determined based on one or more sensors on the electronic device. For example, the electronic device can determine this based on measurement data from an IMU. For example, the electronic device can determine the current acceleration of the electronic device based on the IMU measurement data. If the acceleration of the electronic device is less than or equal to an acceleration threshold, the electronic device can determine that it is in a relatively stable state. Conversely, if the acceleration of the electronic device is greater than the acceleration threshold, the electronic device can determine that it is in motion.
[0322] It should be noted that the relatively stable state here can also be understood as a relatively static state. Stability, static and movement are all relative descriptions.
[0323] In some examples, the posture of an electronic device can be determined based on one or more sensors on the electronic device. For example, the electronic device can determine the posture of the electronic device based on measurement data from a gravity sensor. For example, the electronic device can determine the components of the electronic device's gravitational acceleration on three axes of the electronic device based on the measurement data from the gravity sensor, and determine the angle between the three axes of the electronic device and the horizontal plane based on the components of the gravitational acceleration on the aforementioned three axes. This angle can be used to indicate the posture of the electronic device. For ease of explanation, the "electronic device in the posture of receiving or answering a phone call" is referred to as the second target posture in the following text.
[0324] Taking the three-dimensional schematic diagram of an electronic device shown in FIG6 as an example, the angles between the x-axis, y-axis, and z-axis of the electronic device and a horizontal plane 501 can be referred to as a first angle (α1), a second angle (α2), and a third angle (α3), respectively. In some examples, when the first angle, the second angle, and the third angle all meet a preset condition, the electronic device can determine that it is in the second target posture; conversely, when any of the first angle, the second angle, and the third angle do not meet the preset condition, the electronic device can determine that it is not in the second target posture.
[0325] S402: Determine whether the user has brought the electronic device close to the ear.
[0326] In order to improve the accuracy of electronic devices in determining the user's intention to wake up the voice assistant, reduce the probability of the electronic device's voice assistant being mistakenly woken up, and reduce the increase in power consumption caused by the use of the voice assistant, when it is determined that the electronic device is in the posture for making or receiving a call, the electronic device can determine whether it has undergone a change from a non-second target posture to a second target posture. In other words, the electronic device can determine whether the user has performed an action to adjust the electronic device from a non-second target posture to the second target posture. For ease of explanation, the action of "adjusting the electronic device from a non-second target posture to the second target posture" will be referred to as the second target action below.
[0327] In some examples, the electronic device may store sensor data for a period of time and determine whether the user has performed the second target action based on the sensor data within the period. For example, the electronic device may continuously store IMU measurement data for a certain period of time (e.g., 2 seconds or 1 second, etc.) and determine whether the user has performed the above-mentioned action based on the IMU measurement data within the period of time.
[0328] For example, based on the IMU measurement data within the time interval, the electronic device can use a machine learning-based classification algorithm and / or a rule-based judgment algorithm to determine whether the user has performed the second target action.
[0329] Classification algorithms based on machine learning may include one or more of: traditional machine learning methods such as nearest neighbor algorithms, support vector machines, random forests, and deep learning methods such as convolutional neural networks or autoencoders.
[0330] In some examples, the electronic device uses one or more of the above-mentioned classification algorithms to obtain a probability value of "detecting the second target action". If the probability value is greater than or equal to a preset threshold, the electronic device determines that the user has performed the second target action and continues to perform S403 and subsequent operations. If the probability value is less than the preset threshold, the electronic device determines that the user has not performed the second target action and can restart S401.
[0331] For rule-based judgment algorithms, the electronic device can determine whether the target action is detected based on its vertical movement distance and / or the angles rotated around three axes (x-axis, y-axis, and z-axis). Figure 14 exemplifies a method for an electronic device to determine whether a second target action is detected.
[0332] In some examples, the electronic device may determine that the second target action has been detected based on whether its own vertical movement distance is greater than a vertical displacement threshold. If it is determined that the vertical movement distance is greater than or equal to the vertical displacement threshold, the electronic device may determine that the second target action has been detected; if it is determined that the vertical movement distance is less than the vertical displacement threshold, the electronic device may determine that the second target action has not been detected.
[0333] In some examples, the electronic device can determine that a second target action has been detected based on whether the angle through which it has rotated around the three axes is greater than an angle threshold. If it is determined that the angle through which it has rotated around the three axes is greater than or equal to the angle threshold, the electronic device can determine that a second target action has not been detected. If it is determined that the angle through which it has rotated around the three axes is less than the angle threshold, the electronic device can determine that a second target action has not been detected.
[0334] Exemplarily, the rotation angles of the electronic device around the x-axis, y-axis, and z-axis are respectively a first rotation angle (θ1), a second rotation angle (θ2), and a third rotation angle (θ3). The rotation angle thresholds corresponding to the first rotation angle, the second rotation angle, and the third rotation angle are respectively a first rotation angle threshold (σ1), a second rotation angle threshold (σ2), and a third rotation angle threshold (σ3). When θ1 ≥ σ1, θ2 ≥ σ2, and θ3 ≥ σ3, the electronic device determines that the second target action is detected; if any of the three rotation angles does not meet the aforementioned conditions, the electronic device determines that the second target action is not detected.
[0335] For example, when θ1≥σ1, θ2≥σ2, and θ3≥σ3, the electronic device can determine that it is near the user's left ear, or in other words, the electronic device can determine that the user has performed the action of placing the electronic device near the left ear.
[0336] For another example, when θ1≥σ1, θ2≤-σ2, and θ3≤-σ3, the electronic device can determine that it is near the user's right ear, or in other words, the electronic device can determine that the user has placed the electronic device near the right ear.
[0337] S403 : Determine whether the measurement data of the proximity sensor and / or the measurement data of the ambient light sensor meet the requirements.
[0338] In most cases, the proximity sensor triggers when a phone call is made. In rare cases, the proximity sensor doesn't trigger. In these cases, you can determine whether the phone is close to your ear by observing a decrease in the ambient light sensor reading. If these conditions are not met, the phone call was not detected.
[0339] In some examples, the electronic device can determine that the second target action has been detected based on an event triggered by a proximity sensor. For example, during the movement of the electronic device, if the distance determined based on the proximity sensor's measurement data is less than or equal to a distance threshold, the electronic device determines that the second target action has been detected; if the distance determined based on the proximity sensor's measurement data is greater than the distance threshold, the electronic device determines that the second target action has not been detected.
[0340] In some examples, the electronic device can determine whether the electronic device is close to the ear, that is, whether the second target action is detected, based on whether the measurement value of the ambient light sensor decreases during the electronic device's movement. For example, if the measurement value of the ambient light sensor decreases during the electronic device's movement, it is determined that the second target action is detected. Conversely, if the measurement value of the ambient light sensor does not decrease, it is determined that the second target action is not detected.
[0341] In some examples, the electronic device may also combine the measurement data of the proximity sensor and the measurement data of the ambient light sensor to determine whether the second target action is detected. For example, if the measurement data of the proximity sensor and the measurement data of the ambient light sensor are both within a preset threshold range, the electronic device determines that the second target action is detected. Otherwise, the electronic device determines that the second target action is not detected.
[0342] S404: Determine whether the electronic device is close to the user's ear.
[0343] When an electronic device is placed close to the ear, the capacitive screen reading of the electronic device will change due to the contact between the external auricle and the electronic device screen. At this time, the shape of the auricle can be identified through the capacitive screen reading.
[0344] In some examples, the capacitive screen readings can be represented by a two-dimensional matrix, where generally, points with larger values represent areas in contact with the skin. The capacitive screen readings can be used to determine whether the electronic device is close to the user's ear. For example, the electronic device can use machine learning-based methods such as kNN, SVM, and CNN to analyze the capacitive screen readings to determine whether the capacitive screen readings include target data, which is used to indicate the shape of the auricle, or in other words, the target data includes information about the user's auricle, thereby determining whether the electronic device is close to the user's ear.
[0345] In some scenarios, determining whether the electronic device is close to the user's ear based on whether the capacitive screen reading includes target data can also be understood as determining whether the electronic device is close to the user's ear based on touch information on the electronic device's display. For example, the touch information may include the capacitance value of the electronic device's display, which can be used to indicate information about the user's auricle.
[0346] If the touch information determines that the electronic device is close to the user's ear, the electronic device can execute S405, that is, after opening the voice assistant application, feedback that the voice assistant is successfully woken up. Conversely, if the touch information does not indicate that the electronic device is close to the user's ear, the electronic device can re-execute the operation of S401.
[0347] S405: Feedback that the voice assistant is successfully awakened.
[0348] In some examples, the electronic device can play a second prompt audio through the lower speaker, and the second prompt audio is used to prompt that the voice assistant is successfully woken up.
[0349] The electronic device can also prompt the successful wake-up of the voice assistant through other methods, such as vibration information, etc. The electronic device can also change the feedback method according to the user's customized operation. This application does not impose any restrictions on this.
[0350] When the electronic device is in the second target posture, the prompt information can be more easily perceived by the user by playing audio prompt information through the lower speaker. This feedback method makes the user's interaction with the voice assistant more natural, which is conducive to improving the usability of the voice assistant.
[0351] S406: Determine using the dual-channel audio feature that the user is speaking close to the lower microphone.
[0352] In order to determine that the audio recorded by the microphone comes from the user of the electronic device, or in other words, to determine that it is the user of the electronic device who is speaking, in some examples, while providing feedback that the voice assistant has been successfully woken up, the electronic device can simultaneously record through the upper microphone and the lower microphone, and use the dual-channel audio characteristics to determine that the user is speaking close to the lower microphone.
[0353] In some examples, the electronic device can combine the audio information recorded by the upper microphone and the lower microphone to determine that the user is speaking close to the lower microphone. Figure 10 above shows a method provided by an embodiment of the present application to combine the audio information recorded by the upper microphone and the lower microphone to determine that the user is speaking close to the lower microphone.
[0354] For example, the electronic device can save an audio signal with a length of 200ms and make a judgment every 100ms. The multiple judgment results can be saved as a historical cache. In some examples, the electronic device can separate the audio recorded by the upper microphone and the lower microphone respectively, and use a low-pass filter with a cutoff frequency of 17kHz to remove the signal in the ultrasonic frequency band. The dual-channel audio signal is then used as the input of a CNN-based classifier, and the classification result output by the classifier is used to determine whether the user is speaking close to the lower microphone of the electronic device.
[0355] In some examples, the electronic device may also save the most recent n output results, vote on different types of results in the most recent n output results through a voting mechanism, and regard the result of the majority vote as the final output, where n is an integer greater than or equal to 1.
[0356] The specific execution method of S406 is similar to that of S207. For details, please refer to the relevant content in S207. For the sake of brevity, it will not be described here.
[0357] S407: Select a microphone according to the sound pickup strategy and perform automatic speech recognition.
[0358] When the electronic device is in a call-making position, the electronic device can record the user's voice commands through at least the lower microphone, for example, implementing a sound pickup strategy in which the lower microphone is primary and the upper microphone is secondary. For example, the electronic device can use the audio recorded by the upper microphone to perform auxiliary noise reduction processing on the audio recorded by the lower microphone.
[0359] S408: The voice assistant executes the command and generates feedback according to the corresponding feedback method.
[0360] After the electronic device calls the language assistant to execute the command, it can feedback the completion of the command execution in a similar manner to the feedback of the successful awakening of the language assistant in S405. For details, please refer to the content of S405. For the sake of brevity, it will not be repeated here.
[0361] In some examples, the electronic device may use different feedback methods for different execution results, such as different types of feedback methods, or feedback methods of the same type with different contents. This application does not impose any restrictions on this.
[0362] The method embodiment of the present application is described in detail above in conjunction with Figures 1 to 14. The device embodiment of the present application is described below in conjunction with Figures 15 and 16. It should be understood that the description of the method embodiment corresponds to the description of the device embodiment. Therefore, for parts not described in detail, reference can be made to the above method embodiment.
[0363] Figure 15 is a device 1500 for voice assistant interaction provided in an embodiment of the present application. The device 1500 may have the functions of the electronic device in the above method embodiment and may be used to execute the steps performed by the functions of the electronic device in the above method embodiment. The functions may be implemented by hardware, or by software or hardware executing corresponding software implementations. The hardware or software includes one or more modules corresponding to the above functions.
[0364] In one possible implementation, the voice assistant interaction device 1500 may include an acquisition module 1510 and a processing module 1520, and the acquisition module 1510 and the processing module 1520 are coupled to each other.
[0365] The acquisition module 1510 can be used to support the electronic device in acquiring user input, such as performing operations such as recording audio information in FIG. 3 .
[0366] The processing module 1520 is used to support the electronic device in executing the processing actions in the above method embodiments, such as sending feedback information, confirming the device posture, device action, etc.
[0367] Optionally, the voice assistant interaction device 1500 may further include a storage unit 1530 for storing program codes and data of the voice assistant interaction device 1500 .
[0368] FIG16 illustrates an electronic device 1600 provided in an embodiment of the present application. As shown in the figure, the electronic device 1600 includes at least one processor 1610 and a transceiver 1620. The processor 1610 is coupled to a memory and configured to execute instructions stored in the memory to control the transceiver 1620 to send and / or receive signals.
[0369] Optionally, the electronic device 1600 further includes a memory 1630 for storing instructions.
[0370] In some embodiments, the processor 1610 and memory 1630 may be combined into a processing device, and the processor 1610 is configured to execute program codes stored in the memory 1630 to implement the above functions. In specific implementations, the memory 1630 may also be integrated into the processor 1610 or independent of the processor 1610.
[0371] In some embodiments, the transceiver 1620 may include a receiver (or receiver) and a transmitter (or transmitter).
[0372] The transceiver 1620 may further include an antenna, and the number of antennas may be one or more. The transceiver 1620 may be a communication interface or an interface circuit.
[0373] When the electronic device 1600 is a chip, the chip includes a transceiver module and a processing module. The transceiver module may be an input / output circuit or a communication interface; and the processing module may be a processor, microprocessor, or integrated circuit integrated on the chip.
[0374] This embodiment also provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed on an electronic device, the electronic device executes the above-mentioned related method steps to implement the method of preloading shaders in the above-mentioned embodiment.
[0375] This embodiment further provides a computer program product. When the computer program product is run on a computer, the computer is caused to execute the above-mentioned related steps to implement the method for preloading shaders in the above-mentioned embodiment.
[0376] In addition, embodiments of the present application further provide a device, which may be a chip, component, or module, and may include a processor and a memory connected thereto. The memory is configured to store computer-executable instructions, and when the device is running, the processor may execute the computer-executable instructions stored in the memory to cause the chip to perform the shader preloading method described in each of the above method embodiments.
[0377] Among them, the electronic device, computer-readable storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be repeated here.
[0378] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0379] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0380] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0381] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0382] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0383] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0384] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A method for voice assistant interaction, applied to electronic equipment, characterized in that: include: Detecting a target action of the electronic device, determining that the electronic device is in a target posture, the target posture comprising a posture of a lower microphone of the electronic device close to a user's mouth, and the target action comprising an action for converting the electronic device from a non-target posture to the target posture; In response to the target posture of the electronic device, display and / or play first feedback information, where the first feedback information is used to indicate that the voice assistant is successfully awakened; recording a voice command via at least a lower microphone of the electronic device; In response to the voice instruction, second feedback information is displayed and / or played, where the second feedback information is used to indicate an execution result of the voice instruction.
2. The method according to claim 1, characterized in that The target posture includes a posture in which the lower microphone is close to the user's mouth and the ambient light sensor and / or the proximity sensor are not blocked, The displaying and / or playing first feedback information in response to the target posture of the electronic device includes: In response to the target posture of the electronic device, a prompt screen is displayed and / or a first prompt audio is played through a lower speaker of the electronic device, and the first feedback information includes the prompt screen and / or the first prompt audio.
3. The method according to claim 2, characterized in that The step of displaying a prompt screen and / or playing a first prompt audio through a lower speaker of the electronic device in response to the target gesture of the electronic device includes: In response to a target posture of the electronic device, sending a first ultrasonic signal; receiving a second ultrasonic signal returned based on the first ultrasonic signal; When the second ultrasonic signal indicates that the electronic device is in the target posture, the prompt screen is displayed and / or the first prompt audio is played through the lower speaker of the electronic device.
4. The method according to claim 3, characterized in that The sending of a first ultrasonic signal in response to a target posture of the electronic device comprises: In response to a target posture of the electronic device, sending the first ultrasonic signal through the lower speaker; The receiving a second ultrasonic signal returned based on the first ultrasonic signal comprises: The second ultrasonic signal is returned in response to the first ultrasonic signal contacting the user, and the second ultrasonic signal is received by the lower microphone.
5. The method according to claim 1, characterized in that The target posture includes a posture in which the lower microphone of the electronic device is close to the user's mouth and the ambient light sensor and / or the proximity sensor are blocked. The displaying and / or playing first feedback information in response to the target posture of the electronic device includes: In response to the target gesture of the electronic device, a second prompt audio is played through an upper speaker of the electronic device, and the first feedback information includes the second prompt audio.
6. The method according to claim 5, characterized in that The step of playing a second prompt audio through an upper speaker of the electronic device in response to the target gesture of the electronic device comprises: In response to the target posture of the electronic device, acquiring touch information of a display screen of the electronic device; When the touch information indicates that the display screen of the electronic device is touched, the second prompt audio is played through the upper speaker of the electronic device.
7. The method according to claim 6, characterized in that The touch information includes information of the outer ear contour.
8. The method according to any one of claims 1 to 7, characterized in that The recording of the voice command at least through the lower microphone of the electronic device includes: using the lower microphone of the electronic device as a main recording unit and the upper microphone of the electronic device as an auxiliary recording unit to record the voice command.
9. The method according to claim 8, characterized in that Before recording the first voice instruction using the lower microphone of the electronic device as a main recording unit and the upper microphone of the electronic device as an auxiliary recording unit, the method further includes: recording first audio through the upper microphone and the lower microphone; It is determined based on the first audio that the user is speaking close to the lower microphone.
10. A method for voice assistant interaction, applied to electronic equipment, characterized in that: include: Detecting a target action of the electronic device, determining that the electronic device is in a target posture, the target posture comprising a posture of an upper microphone of the electronic device close to a user's mouth, and the target action comprising an action for converting the electronic device from a non-target posture to the target posture; In response to the target posture of the electronic device, playing first feedback information, the first feedback information is used to indicate that the voice assistant is successfully awakened; recording a voice command via at least an upper microphone of the electronic device; In response to the voice instruction, second feedback information is played, where the second feedback information is used to indicate an execution result of the voice instruction.
11. The method according to claim 10, characterized in that The step of playing first feedback information in response to the target posture of the electronic device includes: In response to the target gesture of the electronic device, the first feedback information is played through an upper speaker of the electronic device.
12. The method according to claim 11, characterized in that The step of playing first feedback information through an upper speaker of the electronic device in response to the target posture of the electronic device comprises: In response to a target posture of the electronic device, sending a first ultrasonic signal; receiving a second ultrasonic signal returned based on the first ultrasonic signal; When the second ultrasonic signal indicates that the electronic device is in the target posture, the first feedback information is played through an upper speaker of the electronic device.
13. The method according to claim 12, characterized in that The sending of a first ultrasonic signal in response to a target posture of the electronic device comprises: In response to a target posture of the electronic device, sending the first ultrasonic signal through the upper speaker; The receiving a second ultrasonic signal returned based on the first ultrasonic signal comprises: The second ultrasonic signal is returned in response to the first ultrasonic signal contacting the user, and the second ultrasonic signal is received by the upper microphone.
14. The method according to any one of claims 10 to 13, characterized in that The recording of the voice command at least through the upper microphone of the electronic device includes: using the upper microphone of the electronic device as a main recording unit and the lower microphone of the electronic device as an auxiliary recording unit to record the voice command.
15. The method according to any one of claims 10 to 14, characterized in that Before recording the voice command using the upper microphone of the electronic device as a main recording unit and the lower microphone of the electronic device as an auxiliary recording unit, the method further includes: recording first audio through the upper microphone and the lower microphone; It is determined based on the first audio that the user is speaking close to the upper microphone.
16. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory is used to store program instructions, and the processor is used to call the program instructions to execute the method according to any one of claims 1 to 9 or claims 10 to 15.
17. A memory management device, characterized in that: Comprising modules for implementing the method of any one of claims 1 to 9 or claims 10 to 15.
18. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a computer, the method according to any one of claims 1 to 9 or claims 10 to 15 is implemented.
19. A chip, characterized in that: The chip comprises a processor and a memory, wherein the processor is used to read instructions stored in the memory, and when the processor executes the instructions, the chip implements the method according to any one of claims 1 to 9 or claims 10 to 15.
Citation Information
Patent Citations
Voice assistant interaction method and electronic equipment
CN119920247A
Application wake method and mobile terminal
CN104571529A
Method and device for waking up voice assistant
CN104820556A
Voice interaction awakening electronic device based on microphone signal, method and medium
CN110097875A
Voice interaction method and related equipment
CN116229953A
Cited By
Livestock and poultry behavior identification method and system based on brain-like network
CN121881010A