Method and electronic device for voice assistant interaction
By detecting the posture and movement of electronic devices, and using multiple microphones and ultrasonic signals to confirm that the microphone is close to the user's mouth, adaptive feedback is provided, solving the problems of ease of use and power consumption of voice assistants, and achieving low false wake-up and low power consumption voice interaction.
Patent Information
- Application Number
- CN202311439204.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-31
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2043-10-31
AI Technical Summary
How to improve the usability of voice assistants while controlling the power consumption of electronic devices during voice interaction, and reducing the probability of false wake-ups and power consumption.
By detecting the posture and movement of electronic devices, using multiple microphones and ultrasonic signals to confirm that the microphone is close to the user's mouth, the recording method is adaptively set, feedback information is provided to wake up the voice assistant, and feedback audio is played when necessary.
It improves the ease of use of the voice assistant, reduces the probability of false wake-ups and power consumption, and enhances the naturalness and efficiency of voice interaction.
Smart Images

Figure CN119920247B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of terminal device software, and more particularly, to a voice assistant interaction method and an electronic device. BACKGROUND
[0002] Voice interaction is an important entrance for human-computer interaction between a user and an electronic device. Power consumption, fluency, ease of use, response speed, and intelligence of voice interaction are important factors affecting whether a user uses a voice assistant and the frequency of use. How to improve the ease of use of a voice assistant while controlling the power consumption of an electronic device in a voice interaction process is a problem worth considering. SUMMARY
[0003] The present application provides a voice assistant interaction method. An electronic device can determine whether to wake up a voice assistant according to a device posture and a device action of the electronic device, and use a pickup method and a feedback method that are adapted to the device posture of the electronic device. The probability of false wake-up of the voice assistant is lower, the power consumption of the electronic device is lower, the process of voice interaction of the user is more natural, and the ease of use of the voice assistant is higher.
[0004] In a first aspect, a voice interaction method is provided, applied to an electronic device, and the method includes: detecting a target action of the electronic device, determining that the electronic device is in a target posture, the target posture including a posture in which a lower microphone of the electronic device is close to a user's mouth, and the target action including an action for converting the electronic device from a non-target posture to the target posture; in response to the target posture of the electronic device, displaying and / or playing first feedback information, the first feedback information being used to indicate that a voice assistant is successfully woken up; recording a voice instruction through at least the lower microphone of the electronic device; and in response to the voice instruction, displaying and / or playing second feedback information, the second feedback information being used to indicate an execution result of the voice instruction.
[0005] It should be understood that before the first feedback information is displayed and / or played, the electronic device can open or call a voice assistant application, so that the application can interact with the user, for example, to obtain a voice command of the user.
[0006] The electronic device feeds back to the user that the voice assistant has been successfully woken up before recording the sound of the user's voice interaction. The speaker of the electronic device does not need to be in a constant open state to record audio information. The power consumption of the electronic device is lower during the use of the voice assistant by the user. In addition, in the case where the user's mouth is close to the lower microphone of the electronic device, the electronic device can set the method for recording the voice instruction to a method that is adapted to the device posture, which is conducive to enabling the electronic device to more clearly obtain the voice command of the user, improving the efficiency of the electronic device in calling the voice assistant, and in another aspect, the implementation of the method is also conducive to reducing the power consumption of the electronic device during recording.
[0007] With reference to the first aspect, in some implementations of the first aspect, the target posture includes a posture in which a lower microphone is close to a user's mouth and an ambient light sensor and / or a proximity sensor is not blocked, in response to the target posture of the electronic device, a prompt picture is displayed and / or a first prompt audio is played through a lower speaker of the electronic device, and the first feedback information includes the prompt picture and / or the first prompt audio.
[0008] Here, the target posture can also be referred to as a first posture.
[0009] In some scenarios, the posture in which the lower microphone is close to the user's mouth and the ambient light sensor and / or the proximity sensor is not blocked can also be understood as a posture in which the lower microphone of the electronic device is close to the user's mouth and the user is not in a situation of making or receiving a call. In this posture, the side of the electronic device where the front camera is arranged is away from the user's head.
[0010] The technical solution specifically provides a method for an electronic device to send feedback information to a user in a first posture. For an electronic device in a first posture, if the user is using a voice assistant, the user has a high probability of looking at the screen of the electronic device. The method of displaying a prompt picture through a display screen is more intuitive and easier for the user to perceive, and the efficiency of the electronic device sending feedback information is higher. Similarly, in the case where the user's mouth is close to the lower microphone, the lower speaker of the electronic device is closer to the user, and playing a prompt audio through the lower speaker is also conducive to making the feedback information perceived by the user, and the efficiency of the electronic device sending feedback information is higher. In summary, the implementation of the technical solution is conducive to making the interaction between the user and the voice assistant more natural, improving the ease of use of the voice assistant, and increasing the frequency of the user using the voice assistant.
[0011] With reference to the first aspect, in some implementations of the first aspect, in response to the target posture of the electronic device, a first ultrasonic signal is sent; a second ultrasonic signal returned based on the first ultrasonic signal is received; and in a case where the second ultrasonic signal indicates that the electronic device is in the target posture, a prompt picture is displayed and / or a first prompt audio is played through a lower speaker of the electronic device.
[0012] The electronic device can confirm whether the lower microphone of the electronic device is close to the user's mouth by sending an ultrasonic signal and recording the returned ultrasonic signal, and in a case where the confirmation also indicates that the electronic device is in a posture in which the lower microphone is close to the user's mouth, feedback information that the voice assistant is successfully woken up is sent to the user. Since the result determined by using ultrasonic waves is more reliable, the implementation of the technical solution is conducive to reducing the probability of the voice assistant of the electronic device being mistakenly woken up and reducing the power consumption caused by the mistaken wake-up of the voice assistant of the electronic device.
[0013] In some implementations of the first aspect, in response to the target posture of the electronic device, the first ultrasonic signal is transmitted through a lower loudspeaker; and in response to a second ultrasonic signal returned by the user in contact with the first ultrasonic signal, the second ultrasonic signal is received through a lower microphone.
[0014] In the technical solution, the first ultrasonic signal is transmitted through the lower loudspeaker close to the user's mouth, and the second ultrasonic signal is recorded through the lower microphone close to the user's mouth. The technical solution is beneficial to improve the accuracy and reliability of the ultrasonic measurement result, reduce the probability of the voice assistant being mistakenly woken up, and reduce the power consumption of the electronic device.
[0015] In some implementations of the first aspect, the target posture includes a posture in which a lower microphone of the electronic device is close to a user's mouth, and an ambient light sensor and / or a proximity sensor are blocked, in response to the target posture of the electronic device, a second prompt audio is played through an upper loudspeaker of the electronic device, and the first feedback information includes the second prompt audio.
[0016] Here, the target posture can also be referred to as a second posture.
[0017] In some scenarios, the posture in which the lower microphone is close to the user's mouth and the ambient light sensor and / or the proximity sensor are blocked can also be understood as a posture in which the lower microphone of the electronic device is close to the user's mouth, and the user is in a situation of making or receiving a call, and the electronic device is in the posture. In this posture, the side of the electronic device where the front camera is arranged is close to or in contact with the user's head.
[0018] The technical solution specifically provides a method for an electronic device to send feedback information to a user in a second posture. In the posture of making or receiving a call, the display screen of the electronic device is close to the user's ear, and the user cannot see the information on the display screen of the electronic device. The upper loudspeaker (or earpiece) of the electronic device is close to the user's ear, and the prompt audio is played through the upper loudspeaker, which is beneficial to make the feedback information be perceived by the user, and the efficiency of the electronic device sending the feedback information is higher. The implementation of the technical solution is beneficial to make the interaction between the user and the voice assistant more natural, improve the ease of use of the voice assistant, and improve the frequency of the user using the voice assistant.
[0019] In some implementations of the first aspect, in response to the target posture of the electronic device, touch information of a display screen of the electronic device is obtained; and in response to the touch information indicating that the display screen of the electronic device is touched, a second prompt audio is played through an upper loudspeaker of the electronic device.
[0020] In some implementations of the first aspect, the touch information includes information of an outer ear contour.
[0021] In a possible implementation, the touch information in the technical solution can include a capacitance value of the display screen of the electronic device.
[0022] The electronic device can use the touch information of the display screen to determine whether the electronic device is in a posture for making or receiving a call, and in a case where it is determined that the electronic device is in the posture for making or receiving a call, send feedback information indicating that the voice assistant is successfully woken up to the user. The reliability of the result of the touch information of the display screen in determining the posture of the electronic device is higher, and the implementation of the technical solution is beneficial to reduce the probability of the voice assistant of the electronic device being mistakenly woken up and reduce the power consumption of the electronic device caused by the mistaken wake-up of the voice assistant.
[0023] With reference to the first aspect, in some implementations of the first aspect, the lower microphone of the electronic device is used as a main recording unit, and the upper microphone of the electronic device is used as an auxiliary recording unit to record the voice instruction.
[0024] The recording method and the device for the voice command provided in the technical solution are adapted to the posture of the device, and the use of multiple microphones to record the voice command is beneficial to enable the electronic device to more clearly obtain the voice command of the user and improve the efficiency of the electronic device in invoking the voice assistant.
[0025] With reference to the first aspect, in some implementations of the first aspect, before the lower microphone of the electronic device is used as a main recording unit and the upper microphone of the electronic device is used as an auxiliary recording unit to record the voice instruction, the method further includes: recording a first audio through the upper microphone and the lower microphone; and determining that the user is speaking close to the lower microphone according to the first audio.
[0026] Since the electronic device in the target posture does not directly correspond to the user being or intending to use the voice assistant, the determination that the user is speaking close to the lower microphone of the electronic device through the upper microphone and the lower microphone of the electronic device is beneficial to improve the accuracy of determining that the user intends to use the voice assistant, reduce the probability of the voice assistant being mistakenly woken up, and reduce the power consumption of the electronic device.
[0027] In a second aspect, a method for voice assistant interaction is provided, applied to an electronic device, including: detecting a target action of the electronic device, determining that the electronic device is in a target posture, the target posture including a posture in which an upper microphone of the electronic device is close to a mouth of a user, and the target action including an action for converting the electronic device from a non-target posture to the target posture; playing first feedback information in response to the target posture of the electronic device, the first feedback information being used to indicate that the voice assistant is successfully woken up; recording a voice instruction at least through the upper microphone of the electronic device; and playing second feedback information in response to the voice instruction, the second feedback information being used to indicate an execution result of the voice instruction.
[0028] It should be understood that before playing the first feedback information, the electronic device can open or invoke the voice assistant application, so that the application can interact with the user, for example, obtain the voice command of the user, and the like.
[0029] The electronic device feeds back to the user that the voice assistant has been successfully woken up before recording the sound of the user's voice interaction. The speaker of the electronic device does not need to be in an always-on state to record audio information, and the power consumption of the electronic device is lower during the user's use of the voice assistant. In addition, in the case that the mouth of the user is close to the upper microphone of the electronic device, the electronic device can set the recording method to a method suitable for the device posture, which is conducive to enabling the electronic device to more clearly obtain the voice command of the user, improving the efficiency of the electronic device invoking the voice assistant. In another aspect, the implementation of the method is also conducive to reducing the power consumption of the electronic device during recording.
[0030] In combination with the second aspect, in some implementations of the second aspect, the first feedback information is played through the upper speaker of the electronic device in response to the target posture of the electronic device.
[0031] Here, the target posture can also be referred to as a third posture.
[0032] The technical solution specifically provides a method for the electronic device to send feedback information to the user in the case of the third posture. For the electronic device in the third posture, if the user is using the voice assistant, the probability that the user perceives the screen on the display screen of the electronic device is small, and the upper speaker of the electronic device is close to the user. Playing prompt audio through the upper speaker is conducive to enabling the feedback information to be perceived by the user, and the efficiency of the electronic device sending the feedback information is higher. The implementation of the technical solution is conducive to making the interaction between the user and the voice assistant more natural, improving the ease of use of the voice assistant, and improving the frequency of the user using the voice assistant.
[0033] In combination with the second aspect, in some implementations of the second aspect, the first feedback information is played through the upper speaker of the electronic device in response to the target posture of the electronic device.
[0034] The electronic device can confirm whether the upper microphone of the electronic device is close to the mouth of the user by sending an ultrasonic signal and recording a returned ultrasonic signal. In the case that the confirmation also indicates that the electronic device is in a posture in which the upper microphone is close to the mouth of the user, the electronic device sends feedback information to the user that the voice assistant has been successfully woken up. Since the result determined by the ultrasonic method is more reliable, the implementation of the technical solution is conducive to reducing the probability of the voice assistant of the electronic device being mistakenly woken up, and reducing the power consumption of the electronic device caused by the mistaken wake-up of the voice assistant.
[0035] With reference to the second aspect, in some implementations of the second aspect, the first ultrasonic signal is transmitted through the upper loudspeaker in response to the target posture of the electronic device; and the second ultrasonic signal is received through the upper microphone in response to the second ultrasonic signal returned by the user in contact with the first ultrasonic signal.
[0036] In the technical solution, the first ultrasonic signal is transmitted by the upper loudspeaker close to the user's mouth, and the second ultrasonic signal is recorded by the upper microphone close to the user's mouth. The technical solution is conducive to improving the accuracy and reliability of the ultrasonic measurement result, reducing the probability of false wake-up of the voice assistant, and reducing the power consumption of the electronic device.
[0037] With reference to the second aspect, in some implementations of the second aspect, the voice instruction is recorded by taking the upper microphone of the electronic device as a main recording unit and the lower microphone of the electronic device as an auxiliary recording unit.
[0038] The recording method and device of the voice command provided in the technical solution are adapted to the posture of the electronic device, and the use of multiple microphones to record the voice command is conducive to enabling the electronic device to more clearly obtain the voice command of the user and improving the efficiency of the electronic device in calling the voice assistant.
[0039] With reference to the second aspect, in some implementations of the second aspect, before the voice instruction is recorded by taking the upper microphone of the electronic device as a main recording unit and the lower microphone of the electronic device as an auxiliary recording unit, the method further comprises: recording a first audio through the upper microphone and the lower microphone; and determining that the user is speaking close to the upper microphone according to the first audio.
[0040] Since the electronic device in the target posture does not directly correspond to the user using or intending to use the voice assistant, the determination that the user is speaking close to the upper microphone of the electronic device through the upper microphone and the lower microphone of the electronic device is conducive to improving the accuracy of determining that the user intends to use the voice assistant, reducing the probability of false wake-up of the voice assistant, and reducing the power consumption of the electronic device.
[0041] The third aspect provides a device for voice assistant interaction, which comprises an acquisition module and a processing module. The processing module is configured to: detect a target action of the device, determine that the device is in a target posture, the target posture comprising a posture in which a lower microphone of the device is close to a user's mouth, and the target action comprising an action for converting the device from a non-target posture to the target posture; display and / or play first feedback information in response to the target posture of the device, the first feedback information being used to indicate successful wake-up of the voice assistant; and the acquisition module is configured to record a voice instruction at least through the lower microphone of the device. The processing module is further configured to: display and / or play second feedback information in response to the voice instruction, the second feedback information being used to indicate an execution result of the voice instruction.
[0042] With reference to the third aspect, in some implementations of the third aspect, the target posture includes a posture in which a lower microphone of the device is close to a mouth of the user and an ambient light sensor and / or a proximity sensor of the device is not blocked, and the processing module is specifically configured to: in response to the target posture of the device, display a prompt picture and / or play a first prompt audio through a lower speaker of the device, and the first feedback information includes the prompt picture and / or the first prompt audio.
[0043] With reference to the third aspect, in some implementations of the third aspect, the processing module is specifically configured to: in response to the target posture of the device, send a first ultrasonic signal; and the acquisition module is further configured to: receive a second ultrasonic signal returned based on the first ultrasonic signal; and in a case where the second ultrasonic signal indicates that the device is in the target posture, display a prompt picture and / or play a first prompt audio through a lower speaker of the device.
[0044] With reference to the third aspect, in some implementations of the third aspect, the processing module is further configured to: in response to the target posture of the device, send a first ultrasonic signal through a lower speaker; and the acquisition module is further configured to: receive a second ultrasonic signal returned in response to the first ultrasonic signal contacting the user through a lower microphone.
[0045] With reference to the third aspect, in some implementations of the third aspect, the target posture includes a posture in which a lower microphone of the device is close to a mouth of the user and an ambient light sensor and / or a proximity sensor of the device is blocked, and the processing module is specifically configured to: in response to the target posture of the device, play a second prompt audio through an upper speaker of the device, and the first feedback information includes the second prompt audio.
[0046] With reference to the third aspect, in some implementations of the third aspect, the processing module is specifically configured to: in response to the target posture of the device, acquire touch information of a display screen of the device; and in a case where the touch information indicates that the device is in a posture in which the display screen is touched, play a second prompt audio through an upper speaker of the device.
[0047] With reference to the third aspect, in some implementations of the third aspect, the touch information includes information of an outer ear contour.
[0048] With reference to the third aspect, in some implementations of the third aspect, the acquisition module is specifically configured to: record a voice instruction by taking a lower microphone of the device as a main recording unit and taking an upper microphone of the device as an auxiliary recording unit.
[0049] In some implementations of the third aspect, before recording the voice instruction with the upper microphone of the device as a main recording unit and the lower microphone of the device as an auxiliary recording unit, the obtaining module is further configured to: record first audio through the upper microphone and the lower microphone; and the processing module is further configured to: determine, according to the first audio, that the user is speaking close to the lower microphone.
[0050] In the fourth aspect, a device for voice assistant interaction is provided, which includes an obtaining module and a processing module. The processing module is configured to: detect a target action of the device, determine that the device is in a target posture, the target posture including a posture in which an upper microphone of the device is close to a user's mouth, and the target action including an action for converting the device from a non-target posture to the target posture; in response to the target posture of the device, play first feedback information, the first feedback information being used to indicate that the voice assistant is successfully woken up; the obtaining module is configured to: record a voice instruction through at least the upper microphone of the device; and the processing module is further configured to: in response to the voice instruction, play second feedback information, the second feedback information being used to indicate an execution result of the voice instruction.
[0051] In some implementations of the fourth aspect, the processing module is specifically configured to: in response to the target posture of the device, play the first feedback information through an upper loudspeaker of the device.
[0052] In some implementations of the fourth aspect, the processing module is further configured to: in response to the target posture of the device, send a first ultrasonic signal; the obtaining module is further configured to: receive a second ultrasonic signal returned based on the first ultrasonic signal; and the processing module is further configured to: in a case where the second ultrasonic signal indicates that the device is in the target posture, play the first feedback information through the upper loudspeaker of the device.
[0053] In some implementations of the fourth aspect, the processing module is specifically configured to: in response to the target posture of the device, send the first ultrasonic signal through the upper loudspeaker; and the obtaining module is specifically configured to: in response to a second ultrasonic signal returned when the first ultrasonic signal contacts a user, receive the second ultrasonic signal through the upper microphone.
[0054] In some implementations of the fourth aspect, the obtaining module is specifically configured to: record the voice instruction with the upper microphone of the device as a main recording unit and the lower microphone of the device as an auxiliary recording unit.
[0055] In some implementations of the fourth aspect, before recording the voice instruction with the upper microphone of the device as a main recording unit and the lower microphone of the device as an auxiliary recording unit, the obtaining module is further configured to: record first audio through the upper microphone and the lower microphone; and the processing module is further configured to: determine, according to the first audio, that the user is speaking close to the upper microphone.
[0056] In a fifth aspect, an electronic device is provided, including a processor and a memory storing program instructions, the processor configured to: detect a target action of the electronic device, determine that the electronic device is in a target posture, the target posture including a posture in which a lower microphone of the electronic device is close to a user's mouth, the target action including an action for converting the electronic device from a non-target posture to the target posture; in response to the target posture of the electronic device, display and / or play first feedback information, the first feedback information being used to indicate that a voice assistant is successfully woken up; record a voice instruction through at least the lower microphone of the electronic device; in response to the voice instruction, display and / or play second feedback information, the second feedback information being used to indicate an execution result of the voice instruction.
[0057] With reference to the fifth aspect, in some implementations of the fifth aspect, the target posture includes a posture in which the lower microphone is close to the user's mouth and an ambient light sensor and / or a proximity sensor are not blocked, and the processor is further configured to: in response to the target posture of the electronic device, display a prompt picture and / or play a first prompt audio through a lower speaker of the electronic device, the first feedback information including the prompt picture and / or the first prompt audio.
[0058] With reference to the fifth aspect, in some implementations of the fifth aspect, the processor is specifically configured to: in response to the target posture of the electronic device, send a first ultrasonic signal; receive a second ultrasonic signal returned based on the first ultrasonic signal; in a case where the second ultrasonic signal indicates that the electronic device is in the target posture, display a prompt picture and / or play a first prompt audio through a lower speaker of the electronic device.
[0059] With reference to the fifth aspect, in some implementations of the fifth aspect, the processor is specifically configured to: in response to the target posture of the electronic device, send a first ultrasonic signal through a lower speaker; in response to a second ultrasonic signal returned in response to the first ultrasonic signal contacting a user, receive the second ultrasonic signal through a lower microphone.
[0060] With reference to the fifth aspect, in some implementations of the fifth aspect, the target posture includes a posture in which the lower microphone of the electronic device is close to the user's mouth and an ambient light sensor and / or a proximity sensor are blocked, and the processor is specifically configured to: in response to the target posture of the electronic device, play a second prompt audio through an upper speaker of the electronic device, the first feedback information including the second prompt audio.
[0061] With reference to the fifth aspect, in some implementations of the fifth aspect, the processor is specifically configured to: in response to the target posture of the electronic device, obtain touch information of a display screen of the electronic device; in a case where the touch information indicates that the display screen of the electronic device is touched, play a second prompt audio through an upper speaker of the electronic device.
[0062] With reference to the fifth aspect, in some implementations of the fifth aspect, the touch information includes information of an outer ear contour.
[0063] With reference to the fifth aspect, in some implementations of the fifth aspect, the processor is further configured to: record the voice instruction with a lower microphone of the electronic device as a main recording unit and an upper microphone of the electronic device as an auxiliary recording unit.
[0064] With reference to the fifth aspect, in some implementations of the fifth aspect, before recording the voice instruction with the lower microphone of the electronic device as the main recording unit and the upper microphone of the electronic device as the auxiliary recording unit, the processor is further configured to: record a first audio through the upper microphone and the lower microphone; and determine that the user is speaking close to the lower microphone according to the first audio.
[0065] A sixth aspect provides an electronic device, including a processor and a memory, the memory being configured to store program instructions, and the processor being configured to: detect a target action of the electronic device, determine that the electronic device is in a target posture, the target posture including a posture in which an upper microphone of the electronic device is close to a user's mouth, and the target action including an action for converting the electronic device from a non-target posture to the target posture; in response to the target posture of the electronic device, play first feedback information, the first feedback information being configured to indicate that a voice assistant is successfully woken up; record a voice instruction at least through the upper microphone of the electronic device; and in response to the voice instruction, play second feedback information, the second feedback information being configured to indicate an execution result of the voice instruction.
[0066] With reference to the sixth aspect, in some implementations of the sixth aspect, the processor is specifically configured to: in response to the target posture of the electronic device, play the first feedback information through an upper loudspeaker of the electronic device.
[0067] With reference to the sixth aspect, in some implementations of the sixth aspect, the processor is specifically configured to: in response to the target posture of the electronic device, send a first ultrasonic signal; receive a second ultrasonic signal returned based on the first ultrasonic signal; and in a case where the second ultrasonic signal indicates that the electronic device is in the target posture, play the first feedback information through the upper loudspeaker of the electronic device.
[0068] With reference to the sixth aspect, in some implementations of the sixth aspect, the processor is specifically configured to: in response to the target posture of the electronic device, send a first ultrasonic signal through the upper loudspeaker; and in response to a second ultrasonic signal returned in response to the first ultrasonic signal contacting a user, receive the second ultrasonic signal through the upper microphone.
[0069] With reference to the sixth aspect, in some implementations of the sixth aspect, the processor is specifically configured to: record the voice instruction with the upper microphone of the electronic device as a main recording unit and a lower microphone of the electronic device as an auxiliary recording unit.
[0070] In combination with the sixth aspect, in some implementations of the sixth aspect, before recording the voice instruction with the upper microphone of the electronic device as the main recording unit and the lower microphone of the electronic device as the auxiliary recording unit, the processor is further configured to: record a first audio through the upper microphone and the lower microphone; and determine that the user is close to the upper microphone to speak according to the first audio.
[0071] In a seventh aspect, a computer program product is provided, which includes computer program code, when the computer program code is run on a computer, causes the method in the first aspect and any possible implementation manner thereof or the method in the second aspect and any possible implementation manner thereof to be executed.
[0072] In an eighth aspect, a computer readable storage medium is provided, which stores computer program code, when the computer program code is run on a computer, causes the method in the first aspect and any possible implementation manner thereof or the method in the second aspect and any possible implementation manner thereof to be executed.
[0073] In a ninth aspect, a chip is provided, which includes a processor configured to read instructions stored in a memory, when the processor executes the instructions, causes the chip to implement the method in the first aspect and any possible implementation manner thereof or the method in the second aspect and any possible implementation manner thereof. BRIEF DESCRIPTION OF DRAWINGS
[0074] Figure 1 FIG. 1 is a hardware architecture schematic diagram of an electronic device suitable for embodiments of the present application.
[0075] Figure 2 FIG. 2 is a software architecture schematic diagram of an electronic device suitable for embodiments of the present application.
[0076] Figure 3 FIG. 3 is an interaction method of a voice assistant provided by embodiments of the present application.
[0077] Figure 4 FIG. 4 is a schematic diagram of device postures of several electronic devices provided by embodiments of the present application.
[0078] Figure 5 FIG. 5 is another interaction method of a voice assistant provided by embodiments of the present application.
[0079] Figure 6 FIG. 6 is a structural schematic diagram of an electronic device provided by embodiments of the present application.
[0080] Figure 7 FIG. 7 is a method schematic diagram of an electronic device determining whether a target action is detected provided by embodiments of the present application.
[0081] Figure 8Fig. 1 is a schematic diagram of a method for an electronic device to detect a distance between a microphone and a user's mouth using ultrasound, according to an embodiment of the present application.
[0082] Figure 9 Fig. 2 is a schematic diagram of a method for an electronic device to send feedback information, according to an embodiment of the present application.
[0083] Figure 10 Fig. 3 is a schematic diagram of a method for an electronic device to determine a positional relationship between a user and the electronic device, according to an embodiment of the present application.
[0084] Figure 11 Fig. 4 is another method for a voice assistant to interact, according to an embodiment of the present application.
[0085] Figure 12 Fig. 5 is a schematic diagram of a method for an electronic device to determine whether a target action is detected, according to an embodiment of the present application.
[0086] Figure 13 Fig. 6 is another method for a voice assistant to interact, according to an embodiment of the present application.
[0087] Figure 14 Fig. 7 is a schematic diagram of a method for an electronic device to determine whether a target action is detected, according to an embodiment of the present application.
[0088] Figure 15 Fig. 8 is a device for a voice assistant to interact, according to an embodiment of the present application.
[0089] Figure 16 Fig. 9 is an electronic device, according to an embodiment of the present application. DETAILED DESCRIPTION
[0090] The technical solutions in the present application will be described below with reference to the drawings.
[0091] The terms used in the following embodiments are for the purpose of describing particular embodiments only and are not intended to be limiting of the present application. As used in this specification and the appended claims, the singular forms "a," "an" and "the" are intended to include both singular and plural forms, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, objects, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, objects, and / or components thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. The term "at least one of' followed by a list of two or more items, means that at least one of the listed items is present at any given occurrence of the term. However, in the list of two or more items, one or more of the listed items can be present, with the rest omitted. The terms "comprise," "comprising," "include," "including," and "includes" are used in the specification to indicate the presence of the stated feature(s) but do not preclude the presence or addition of one or more other feature(s). The term "plurality" is intended to indicate a quantity of two or more.
[0092] Reference within the specification to "one embodiment" or "an embodiment" or "some embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrase "in one embodiment" or "in some embodiments" within the specification are not necessarily all referring to the same embodiment, however, it is contemplated that the features, structures, or characteristics of one embodiment can be combined with those of another embodiment.
[0093] Figure 1 A structural diagram of the electronic device 100 is shown. The electronic device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headset jack 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 can include a pressure sensor 180A, a gyro sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0094] It can be understood that the structure shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In some other embodiments of the present application, the electronic device 100 can include more or fewer components than shown, or combine certain components, or split certain components, or different component arrangements. The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0095] The processor 110 can include one or more processing units, for example: the processor 110 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices or integrated in one or more processors.
[0096] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching and executing instructions.
[0097] The processor 110 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can save instructions or data that the processor 110 has just used or repeatedly uses. If the processor 110 needs to use the instructions or data again, it can directly call from the memory. This avoids repeated access and reduces the waiting time of the processor 110, thereby improving the efficiency of the system.
[0098] In some embodiments, the processor 110 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0099] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 can include multiple sets of I2C buses. The processor 110 can be coupled to the touch sensor 180K, the charger, the flash, the camera 193, etc. through different I2C bus interfaces respectively. For example, the processor 110 can be coupled to the touch sensor 180K through an I2C interface, so that the processor 110 and the touch sensor 180K communicate through the I2C bus interface, and the touch function of the electronic device 100 is realized.
[0100] The I2S interface can be used for audio communication. In some embodiments, the processor 110 can include multiple sets of I2S buses. The processor 110 can be coupled to the audio module 170 through the I2S bus, and communication between the processor 110 and the audio module 170 is realized. In some embodiments, the audio module 170 can deliver audio signals to the wireless communication module 160 through the I2S interface, and the function of answering a phone through a Bluetooth headset is realized.
[0101] The PCM interface can also be used for audio communication, sampling, quantizing and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled through the PCM bus interface. In some embodiments, the audio module 170 can also deliver audio signals to the wireless communication module 160 through the PCM interface, and the function of answering a phone through a Bluetooth headset is realized. Both the I2S interface and the PCM interface can be used for audio communication.
[0102] The MIPI interface can be used to connect the processor 110 and peripheral devices such as the display 194 and the camera 193. The MIPI interface includes the camera serial interface (CSI), the display serial interface (DSI), etc. In some embodiments, the processor 110 and the camera 193 communicate through the CSI interface, and the shooting function of the electronic device 100 is realized. The processor 110 and the display 194 communicate through the DSI interface, and the display function of the electronic device 100 is realized.
[0103] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or as a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 and the camera 193, the display 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0104] The USB interface 130 is an interface conforming to the USB standard specification, and can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device 100, and can also be used to transmit data between the electronic device 100 and a peripheral device. It can also be used to connect a headset to play audio through the headset. The interface can also be used to connect other electronic devices, such as AR devices, etc.
[0105] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a structural limitation of the electronic device 100. In some other embodiments of the present application, the electronic device 100 can also use different interface connection methods or a combination of multiple interface connection methods.
[0106] The wireless communication function of the electronic device 100 can be realized by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, and the baseband processor, etc.
[0107] The electronic device 100 realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs that execute program instructions to generate or change display information.
[0108] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can use a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diode (QLED), etc. In some embodiments, the electronic device 100 can include 1 or N display screens 194, and N is a positive integer greater than 1.
[0109] The electronic device 100 can implement a photographing function through an ISP, a camera 193, a video codec, a GPU, a display 194, and an application processor, etc. The ISP is used to process data fed back by the camera 193. The camera 193 is used to capture a still image or a video. The digital signal processor is used to process a digital signal, which can process not only a digital image signal but also other digital signals. The video codec is used to compress or decompress a digital video.
[0110] The NPU is a neural-network (NN) computing processor, which can quickly process input information by referring to a biological neural network structure, for example, referring to a transmission mode between human brain neurons, and can also constantly self-learn.
[0111] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement a data storage function.
[0112] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various function applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), etc. The data storage area can store data created during use of the electronic device 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0113] The electronic device 100 can implement an audio function through an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, and an application processor, etc. For example, music playing, recording, etc.
[0114] The audio module 170 is used to convert digital audio information into an analog audio signal output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used to encode and decode an audio signal. In some embodiments, the audio module 170 can be arranged in the processor 110, or part of the function modules of the audio module 170 can be arranged in the processor 110.
[0115] The speaker 170A, also known as a "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or listen to a hands-free call through the speaker 170A.
[0116] In some examples, the electronic device 100 can include one or more speakers 170A, and in the case of multiple speakers 170A, the multiple speakers 170A can be distributed at different positions of the electronic device 100.
[0117] For example, at least one of the multiple speakers 170A can be located on the same side of the display screen of the electronic device 100, and at least one can be located on the opposite side of the display screen of the electronic device 100.
[0118] For example, in the direction of the longer side of the electronic device 100 as the length direction, and the direction of the shorter side of the electronic device 100 as the width direction, the multiple speakers 170A can be distributed at opposite upper and lower positions in the length direction or opposite left and right positions in the width direction of the electronic device 100. For example, taking the position where the front camera of the electronic device 100 is arranged as the upper part, and the position opposite to the upper part on the same side and away from the front camera of the electronic device 100 as the lower part, the multiple speakers 170A can include an upper speaker and a lower speaker, the upper speaker is located at the upper part of the electronic device 100, or close to the front camera of the electronic device 100, and the lower speaker is located at the lower part of the electronic device 100, or away from the front camera of the electronic device 100.
[0119] It should be noted that the upper part and the lower part, the upper and lower, and the left and right here are relative positional relationships. The above is only illustrative and should not be construed as limiting the present application.
[0120] The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device 100 answers a call or a voice message, the receiver 170B can be held close to the ear to listen to the voice.
[0121] The microphone 170C, also known as the "microphone", "sound collector", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak into the microphone 170C through the mouth to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, in addition to collecting sound signals, it can also achieve noise reduction function. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C, to achieve the functions of collecting sound signals, noise reduction, and identifying sound sources, and achieving directional recording functions, etc.
[0122] When the number of the microphone 170C is plural, the plural microphones 170C can be distributed at different positions of the electronic device 100.
[0123] For example, at least one of the plural microphones 170C can be located at the same side of the display screen of the electronic device 100, and at least one can be located at the opposite side of the display screen of the electronic device 100.
[0124] For example, taking the direction of the longer side of the electronic device 100 as the length direction, and the direction of the shorter side of the electronic device 100 as the width direction, the plural microphones 170C can be distributed at the upper and lower positions of the length direction of the electronic device 100, or at the left and right positions of the width direction of the electronic device 100. For example, taking the position where the front camera of the electronic device 100 is arranged as the upper part, and the position opposite to the upper part on the same side and away from the front camera of the electronic device 100 as the lower part, the plural microphones 170C can include an upper speaker and a lower speaker, the upper speaker is located at the upper part of the electronic device 100, or close to the front camera of the electronic device 100, and the lower speaker is located at the lower part of the electronic device 100, or away from the front camera of the electronic device 100.
[0125] It should be noted that the upper part and the lower part, the upper and the lower, and the left and the right here are relative positional relationships. The above are only illustrative, and should not be understood as a limitation on the present application.
[0126] The earphone interface 170D is used to connect a wired earphone. The earphone interface 170D can be a USB interface 130, or a 3.5mm open mobile terminal platform (OMTP) standard interface, a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0127] The keys 190 include a power-on key, a volume key, etc. The keys 190 can be mechanical keys. They can also be touch keys. The electronic device 100 can receive key inputs, and generate key signal inputs related to user settings and function control of the electronic device 100.
[0128] The motor 191 can generate a vibration prompt. The motor 191 can be used for incoming call vibration prompt, and also can be used for touch vibration feedback.
[0129] The indicator 192 can be an indicator light, and can be used to indicate a charging state, a power change, and also can be used to indicate a message, a missed call, a notification, etc.
[0130] The software system of the electronic device 100 can employ a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. Embodiments of the present application take an Android system with a layered architecture as an example to illustrate the software structure of the electronic device 100.
[0131] It can be understood that the interface connection relationship between the modules shown in the embodiments of the present application is only illustrative and does not constitute a limitation on the structure of the electronic device 100. In other embodiments of the present application, the electronic device 100 can also employ different interface connection manners or combinations of multiple interface connection manners in the above embodiments.
[0132] Figure 2 is a software structure block diagram of the electronic device 100 in the embodiments of the present application. The layered architecture divides the software into several layers, each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, the application (app) layer, the application framework layer, the Android runtime and system library, and the kernel layer. The application layer can include a series of application packages.
[0133] As shown in Figure 2 , the application package can include camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.
[0134] The application framework layer provides the application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some pre-defined functions.
[0135] As shown in Figure 2 , the application framework layer can include window manager, activity manager, package manager, resource manager, view system, phone manager, notification manager, etc.
[0136] The resource manager, also known as resource management service (RMS), provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, etc.
[0137] The window manager, also known as window management service (WMS), is used to manage window programs. The window manager can obtain the size of the display screen, determine whether there is a status bar, lock the screen, and take screenshots, etc.
[0138] The activity manager, also referred to as the activity manager service (AMS), manages all application processes in the system.
[0139] The package manager, also referred to as the package manager service (PMS), is responsible for application installation and uninstallation, component query and matching, permission management, and the like.
[0140] The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, and the like. The view system can be used to build an application. A display interface can be composed of one or more views. For example, a display interface including a short message notification icon can include a view for displaying text and a view for displaying a picture.
[0141] The notification manager enables an application to display notification information in the status bar, and can be used to convey messages of the notification type, which can automatically disappear after a short stay without user interaction. For example, the notification manager is used to notify of a download completion, a message reminder, and the like. The notification manager can also be a notification appearing in the form of a chart or a scroll bar text in the top status bar of the system, such as a notification of an application running in the background, and can also be a notification appearing in the form of a dialog window on the screen. For example, a text information is prompted in the status bar, a prompt sound is emitted, the electronic device is vibrated, an indicator light flashes, and the like.
[0142] The Android runtime includes a core library and a virtual machine, and is responsible for scheduling and management of the Android system.
[0143] The core library includes two parts: one part is a function function called by the java language, and the other part is the core library of Android.
[0144] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the java files of the application layer and the application framework layer into binary files. The virtual machine is used to perform functions of management of the object life cycle, stack management, thread management, security and exception management, and garbage collection.
[0145] The system library can include multiple functional modules. For example: a surface manager, media libraries, a three-dimensional graphics processing library (for example: OpenGL ES), a 2D graphics engine (for example: SGL), and the like.
[0146] The surface manager is used to manage the display subsystem, and provides fusion of 2D and 3D layers for multiple applications.
[0147] The media library supports a variety of commonly used audio, video format playback and recording, and static image files, etc. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0148] The three-dimensional graphics processing library is used to realize three-dimensional graphics drawing, image rendering, synthesis, and layer processing, etc.
[0149] The 2D graphics engine is a drawing engine for 2D drawing.
[0150] The kernel layer is a layer between hardware and software. The kernel layer at least includes display drivers, camera drivers, audio drivers, and sensor drivers.
[0151] Before formally introducing the embodiments of the present application, first, some terms that may be used in the following content are explained.
[0152] K-nearest neighbors algorithm (kNN), also known as nearest neighbor method, is a non-parametric statistical method for classification and regression. The algorithm uses a vector space model for classification, and the concept is that cases of the same category have high similarity to each other, and the similarity of unknown category cases can be evaluated by calculating the similarity with known category cases.
[0153] Support vector machine (SVM), also known as support vector network, is a supervised learning model and related learning algorithm for analyzing data in classification and regression analysis. The SVM model represents instances as points in space, so that the instances of individual classes are separated by as wide a margin as possible. Then, new instances are mapped to the same space, and the class to which they belong is predicted based on which side of the margin they fall on.
[0154] Random forest refers to a classifier containing multiple decision trees, and the output class is determined by the mode of the class output by individual trees.
[0155] Convolutional neural network (CNN) refers to a kind of feedforward neural network, and the artificial neurons of which can respond to a part of the surrounding units within the coverage range, and have excellent performance for large image processing. In at least one layer of the convolutional neural network, a mathematical operation called convolution is used instead of general matrix multiplication.
[0156] Recurrent neural network (RNN) refers to an artificial neural network that uses sequence data or time series data, which can be used for sequential or time problems, such as language translation, natural language processing, and speech recognition, etc.
[0157] An autoencoder, also known as an auto-encoder, is an artificial neural network used to learn an efficient encoding of unlabeled data, belonging to unsupervised learning. An autoencoder has two main parts: an encoder to encode the input, and a decoder to reconstruct the input using the encoding.
[0158] A chirp signal refers to a signal whose frequency changes (increases or decreases) over time.
[0159] Automatic speech recognition (ASR), also known as speech-to-text recognition or speech recognition technology, aims to automatically convert human speech content into responsive text using a computer. Applications of speech recognition technology include voice dialing, voice navigation, indoor device control, and voice document retrieval.
[0160] A proximity sensor is a sensor that can detect the presence of objects in the vicinity without contact. Proximity sensors typically emit electromagnetic fields or electromagnetic radiation beams (such as infrared) and observe changes in the electric field or return signal to achieve functionality.
[0161] Voice interaction is an important method for users to interact with electronic devices. In the process of user interaction with a voice assistant, in order to improve the efficiency of voice interaction and reduce the power consumption of the electronic device, the present application provides a voice assistant interaction method, which is introduced as follows.
[0162] Figure 3 The flowchart shows a voice assistant interaction method provided by an embodiment of the present application. Before recording the user's voice interaction sound, the electronic device first feeds back to the user that the voice assistant has been successfully woken up. The electronic device sends feedback information to the user in a manner adapted to the device posture of the electronic device. The user's interaction process with the voice assistant is more natural, and the voice assistant is more user-friendly.
[0163] S110, in response to detecting the target posture and the target action, sending first feedback information according to the target feedback method.
[0164] Here, the target action can refer to an action of converting the electronic device from a non-target posture to a target posture, which can be completed by user operation.
[0165] In some examples, the target posture can include a first posture, a second posture or a third posture, where the first posture is a posture in which a lower microphone of the electronic device is close to a user's mouth and a display screen of the electronic device is not close to the user's ear; the second posture is a posture in which the lower microphone of the electronic device is close to the user's mouth and the display screen of the electronic device is close to the user's ear; and the third posture is a posture in which an upper microphone of the electronic device is close to the user's mouth.
[0166] Figure 4 Exemplarily, a schematic diagram of the aforementioned first posture, second posture and third posture is provided, Figure 4 The diagram 401 on the left is a schematic diagram of the first posture, the diagram 402 in the middle is a schematic diagram of the second posture, and the diagram 403 on the right is a schematic diagram of the third posture. It should be noted that, Figure 4 The exemplarily description of the aforementioned postures is only for the convenience of understanding, and should not be regarded as a limitation on the embodiments of the present application. In other words, the aforementioned first posture, second posture and third posture can also be other postures in addition to the schematic diagrams shown in the above, Figure 4 The embodiments of the present application do not limit this.
[0167] In some scenarios, the aforementioned first posture can also be understood as a posture in which the lower microphone of the electronic device is close to the user's mouth and the electronic device is not in a state of making or receiving a call; or a posture in which the lower microphone of the electronic device is close to the user's mouth and the ambient light sensor and / or the proximity sensor are not blocked. The aforementioned second posture can also be understood as a posture in which the lower microphone of the electronic device is close to the user's mouth and the electronic device is in a state of making or receiving a call; or a posture in which the lower microphone of the electronic device is close to the user's mouth and the ambient light sensor and / or the proximity sensor are blocked.
[0168] In some examples, the electronic device can obtain measurement data of one or more sensors of the electronic device, and the measurement data of the sensors can be used to indicate a device posture and a device action of the electronic device. In other words, the electronic device can determine whether a target posture or a target action is detected according to the measurement data of one or more sensors of the electronic device.
[0169] Exemplarily, the electronic device can determine whether the electronic device is in the first posture, the second posture or the third posture according to the measurement data of the gravity sensor. For example, the electronic device can determine the components of the gravitational acceleration on the three axes of the electronic device according to the measurement data of the gravity sensor, determine the included angle between the three axes of the electronic device and the horizontal plane according to the components of the gravitational acceleration on the three axes, and determine the posture of the electronic device according to the included angle.
[0170] In some examples, before determining the device posture of the electronic device, the electronic device can also determine the acceleration of the electronic device according to the measurement data of the inertial measurement unit (IMU), and determine whether the electronic device is in a relatively stable state according to the acceleration. If the electronic device is in a relatively stable state, the electronic device determines the device posture of the electronic device.
[0171] In some examples, in the case of detecting the target posture and the target action, the electronic device can use a target verification method to verify whether the electronic device is in the target posture, thereby reducing the probability of false wake-up and false touch of the voice assistant, and improving the energy utilization efficiency of the electronic device.
[0172] For example, in the case where the target posture includes the first posture or the third posture, the electronic device can use the loudspeaker to send a first ultrasonic signal and use the microphone to receive a second ultrasonic signal returned based on the first ultrasonic signal, and verify whether the electronic device is in the target posture according to the recorded second ultrasonic signal and the sent first ultrasonic signal.
[0173] Here, the loudspeaker can refer to the loudspeaker 170A of the electronic device 100 in the foregoing, such as the upper loudspeaker or the lower loudspeaker. The microphone can refer to the microphone 170C of the electronic device 100 in the foregoing, such as the upper microphone and the lower microphone. For specific descriptions of the loudspeaker position or the microphone position of the electronic device, please refer to the related descriptions in the following Figure 6 .
[0174] For example, the electronic device detects that it is in the first posture according to the foregoing method, and the electronic device can use the lower loudspeaker of the electronic device to send a first ultrasonic signal, use the lower microphone of the electronic device to record a second ultrasonic signal, and verify whether the electronic device is in the first posture according to the first ultrasonic signal and the second ultrasonic signal.
[0175] For another example, the electronic device detects that it is in the third posture according to the foregoing method, and the electronic device can use the upper loudspeaker of the electronic device to send a first ultrasonic signal, use the upper microphone of the electronic device to record a second ultrasonic signal, and verify whether the electronic device is in the second posture according to the first ultrasonic signal and the second ultrasonic signal.
[0176] For the method of verifying whether the electronic device is in the target posture by using the sent ultrasonic signal and the recorded ultrasonic signal, specific descriptions will be given in the following embodiments, and will not be expanded here.
[0177] The power consumption of sending ultrasound and recording audio is higher than that of the aforementioned IMU and gravity sensor, but the accuracy of measuring data for determining the posture of the electronic device is higher, and the method is used to verify the posture of the electronic device, which is beneficial to control the power consumption of the electronic device on the basis of improving the accuracy of target posture detection, and is beneficial to reduce the false touch rate in the use process of the voice assistant.
[0178] Similarly, in the case where the target posture includes the aforementioned second posture, the electronic device can obtain touch information of the display screen, and verify whether the electronic device is in the second posture according to the touch information.
[0179] In some examples, the touch information can include information of the user's pinna. Illustratively, the touch information can include a capacitance value of the display screen of the electronic device, which can be used to indicate the information of the user's pinna.
[0180] When the electronic device is in the second posture, the user's ear is close to the display screen of the electronic device, and the touch information of the display screen of the electronic device can be used to determine whether the user's ear is close to the display screen of the electronic device, that is, to verify whether the electronic device is in the second posture. The method of how to use the touch information of the display screen to verify whether the electronic device is in the second posture will be described in the embodiments below, and will not be expanded here.
[0181] The power consumption of obtaining the touch information of the display screen is higher than that of the aforementioned IMU and gravity sensor, but the accuracy of measuring data for determining the posture of the electronic device is higher, and the method is used to verify the posture of the electronic device, which is beneficial to control the power consumption of the electronic device on the basis of improving the accuracy of target posture detection, and is beneficial to reduce the false touch rate in the use process of the voice assistant.
[0182] In some examples, in the case where the target posture and the target action are detected, the electronic device can also verify whether the electronic device is in the target posture by reading the measurement data of the proximity sensor and / or the measurement data of the ambient light sensor. The measurement values of the measurement data of the corresponding proximity sensor and the measurement data of the ambient light sensor are different when the electronic device is in the aforementioned first posture, second posture and third posture, and the method of how to use the measurement data of the proximity sensor and / or the measurement data of the ambient light sensor to verify whether the electronic device is in the target posture will be described in detail below, and will not be expanded here.
[0183] In some examples, the electronic device can determine whether the electronic device is in the target posture in combination with multiple of the aforementioned determination, verification methods of the target posture. For example, in combination with the measurement data of the IMU and the ultrasonic signal, in combination with the measurement data of the IMU, the measurement data of the proximity sensor and the touch information of the display screen, or in combination with the measurement data of the gravity sensor, the ultrasonic signal, the measurement data of the proximity sensor and the measurement data of the ambient light sensor, etc.
[0184] Determining whether the electronic device is in the target posture by multiple methods, measurement data of multiple sensors of the electronic device is conducive to reducing the probability of being mistaken in the process of using the voice assistant, and is conducive to improving the energy utilization efficiency of the electronic device.
[0185] The first feedback information is used to indicate that the voice assistant is successfully woken up. For different postures of the electronic device, the method of sending the first feedback information by the electronic device can be different, or in other words, the target feedback method of sending the first feedback information can be determined according to the device posture (i.e., the target posture) of the electronic device.
[0186] In some examples, the target posture can include the aforementioned first posture, and the first feedback information can be a prompt picture and / or a first prompt audio. The electronic device can display the prompt picture through the display screen and / or play the first prompt audio through the loudspeaker, and the prompt picture and / or the first prompt audio are used to indicate that the voice assistant is successfully woken up. Exemplarily, the electronic device can play the first prompt audio through the lower loudspeaker.
[0187] In some examples, the target posture can include the aforementioned second posture, and the first feedback information can be a second prompt audio. The electronic device can play the second prompt audio through the loudspeaker, and the second prompt audio can be used to indicate that the voice assistant is successfully woken up. Exemplarily, the electronic device can play the second prompt audio through the upper loudspeaker.
[0188] In some examples, the target posture can include the aforementioned third posture, and the first feedback information can be a first prompt audio. The electronic device can play the first prompt audio through the loudspeaker, and the first prompt audio can be used to indicate that the voice assistant is successfully woken up. Exemplarily, the electronic device can play the first prompt audio through the upper loudspeaker.
[0189] The method of sending the first feedback information can be adapted to different postures of the electronic device, or in other words, the method of sending the first feedback information is more natural and more in line with the operation habits of the user, and the feedback method is conducive to making the process of human-computer interaction of the user through the voice assistant more natural.
[0190] It should be understood that the electronic device can also use more other ways to send the aforementioned first feedback information, for example, one or more of vibration, display screen light effect, flash light effect, etc., which are not limited by the present application.
[0191] S120, record the voice command according to the target pickup method.
[0192] Starting to record the voice command after sending the feedback information, the microphone of the electronic device does not need to be in an always-on state, which is conducive to reducing the power consumption of the electronic device, improving the energy utilization efficiency of the electronic device, and also conducive to reducing the probability of the voice assistant being mistakenly awakened, making the interaction process between the user and the voice assistant more natural.
[0193] The target pickup method refers to the method or strategy of recording the voice command by the electronic device, which can be determined according to the target posture of the electronic device in S110.
[0194] In some examples, the target posture includes the first posture or the second posture described above, and the target pickup method can be to record the voice command at least through the lower microphone of the electronic device. Illustratively, the voice command is recorded with the lower microphone of the electronic device as the main recording unit and the upper microphone of the electronic device as the auxiliary recording unit.
[0195] In some examples, the target posture includes the third posture described above, and the target pickup method can be to record the voice command at least through the upper microphone of the electronic device. Illustratively, the voice command is recorded with the upper microphone of the electronic device as the main recording unit and the lower microphone of the electronic device as the auxiliary recording unit.
[0196] Illustratively, here, the auxiliary recording unit can be used to perform noise reduction processing on the audio recorded by the main recording unit.
[0197] The target pickup method is determined according to the device posture of the electronic device, which is conducive to improving the accuracy of the voice command recorded by the electronic device, improving the efficiency of the voice assistant executing the user voice instruction, and to a certain extent, improving the energy utilization efficiency of the electronic device.
[0198] S130, in response to the voice command, sending second feedback information according to the target feedback method.
[0199] According to the recorded voice command, the electronic device can determine the voice command sent by the user, and through automatic speech recognition, the voice command can be converted into text information. According to the text information, the electronic device can call the voice assistant to complete the command indicated by the text information.
[0200] Here, the method of sending the second feedback information, i.e. the target feedback method, can be determined according to the device posture of the electronic device.
[0201] In some examples, the target posture can include the first posture described above, and the second feedback information can be a result prompt picture and / or a result prompt audio. The electronic device can display the result prompt picture through the display screen and / or play the result prompt audio through the loudspeaker, the result prompt picture and / or the result prompt audio being used to indicate the execution result of the voice command. Exemplarily, the electronic device can play the result prompt audio through the lower loudspeaker.
[0202] In some examples, the target posture can include the second posture described above, and the second feedback information can be a result prompt audio. The electronic device can play the result prompt audio through the loudspeaker, which can be used to indicate the execution result of the voice command. Exemplarily, the electronic device can play the result prompt audio through the upper loudspeaker.
[0203] In some examples, the target posture can include the third posture described above, and the second feedback information can be a result prompt audio. The electronic device can play the result prompt audio through the loudspeaker, which can be used to indicate the execution result of the voice command. Exemplarily, the electronic device can play the result prompt audio through the loudspeaker.
[0204] It should be understood that the execution result herein can include both a successful execution result and an unsuccessful execution result or a result of failure to execute. In addition, the execution result can also be a result of failure to detect the user's voice within a preset time period of the electronic device, indicating that the execution of the voice assistant interaction is stopped or paused.
[0205] The method of sending the second feedback information can be adapted to different postures of the electronic device, or in other words, the method of sending the second feedback information is more natural and conforms to the operation habits of the user, and the feedback manner is conducive to making the user's human-computer interaction process through the voice assistant more natural.
[0206] It should be understood that the electronic device can also use more other ways to send the second feedback information described above, for example, one or more of vibration, display screen light effect, flash light effect, indicator light effect, etc., which are not limited in the present application.
[0207] For different device postures, the electronic device can adopt different wake-up schemes, sound pickup schemes and feedback schemes, such as Figure 5 As shown in FIG. 401, for the case that the electronic device is in a posture that the lower microphone is close to the user's mouth and is not in a state of making a call (the posture of the electronic device 100 in FIG. 401), the electronic device can execute the following interaction scheme.
[0208] S201, determine that the electronic device is in a posture that the lower microphone is close to the user's mouth.
[0209] In some examples, the electronic device determines whether it is in a posture in which the lower microphone is close to the user's mouth while being in a relatively stable state.
[0210] Whether the electronic device is in a relatively stable state can be determined according to one or more sensors on the electronic device.
[0211] Exemplarily, the electronic device can determine whether it is in a relatively stable state according to the measurement data of the IMU. For example, the electronic device can determine the magnitude of the current acceleration of the electronic device according to the measurement data of the IMU, and in the case where the acceleration of the electronic device is less than or equal to an acceleration threshold, the electronic device can determine that it is in a relatively stable state. Conversely, in the case where the acceleration of the electronic device is greater than the acceleration threshold, the electronic device can determine that it is in a motion state.
[0212] It should be noted that the relatively stable state here can also be understood as a relatively stationary state, and stability, stationarity and motion are all relative descriptions.
[0213] In some examples, the posture of the electronic device can be determined according to one or more sensors on the electronic device.
[0214] Exemplarily, the electronic device can determine its device posture according to the measurement data of the gravity sensor. For example, the electronic device can determine the components of the gravitational acceleration of the electronic device on the three axes of the electronic device according to the measurement data of the gravity sensor, and determine the included angle between the three axes of the electronic device and the horizontal plane according to the components of the gravitational acceleration on the three axes, which can be used to indicate the device posture of the electronic device. For ease of description, the following content refers to the first target posture as "the electronic device is in a posture in which the lower microphone is close to the user's mouth and is not in a state of making a phone call".
[0215] In the example of the perspective view of the electronic device shown in Figure 6 The three axes of the electronic device can be referred to as the x-axis, the y-axis and the z-axis respectively, where the x-axis direction can be regarded as the width direction of the electronic device, the y-axis direction can be regarded as the height direction of the electronic device, and the z-axis direction can be regarded as the thickness direction of the electronic device. Figure 6 The quadrilateral region 601 in the figure can be regarded as a horizontal plane, and the included angles between the x-axis, the y-axis and the z-axis of the electronic device and the horizontal plane can be referred to as the first included angle (α1), the second included angle (α2) and the third included angle (α3) respectively.
[0216] The electronic device 100 can include an upper microphone 201A, a lower microphone 201B, an upper speaker 202A, a lower speaker 202B, a proximity sensor 203 and an ambient light sensor 204. Figure 6In some examples, the upper microphone 201A, the upper speaker 202A, the proximity sensor 203, and the ambient light sensor 204 can all be located in the upper portion of the electronic device 100, i.e., the region of the electronic device in the positive direction of the y-axis. The lower microphone 201B and the lower speaker 202B can be located in the lower portion of the electronic device 100, i.e., the region of the electronic device in the negative direction of the y-axis.
[0217] It should be understood that, Figure 6 In some examples, the distribution of the positions of the plurality of electronic components of the electronic device 100 is merely an example, and the present application is not limited in this regard.
[0218] In some examples, when the first angle, the second angle, and the third angle all satisfy the preset condition, the electronic device can determine that it is in the first target posture. Conversely, when any one of the first angle, the second angle, and the third angle does not satisfy the preset condition, the electronic device can determine that it is not in the first target posture.
[0219] S202, determining that the user performs an action of placing the lower microphone of the electronic device close to the mouth.
[0220] In order to improve the accuracy of the electronic device in determining the user's intention to wake up the voice assistant, reduce the probability of the voice assistant of the electronic device being mistakenly woken up, and reduce the increase in power consumption caused by using the voice assistant, in a case where the electronic device is determined to be in the first target posture, the electronic device can determine whether it has experienced a change from a non-first target posture to a first target posture. In other words, the electronic device can determine whether the user has adjusted the electronic device from a non-first target posture to a first target posture. For ease of description, the action of "adjusting the electronic device from a non-first target posture to a first target posture" is referred to as a first target action.
[0221] In some examples, the electronic device can save the data of the sensor for a period of time, and determine whether the user has performed the first target action according to the data of the sensor in the period of time. For example, the electronic device can continuously save the measurement data of the IMU for a certain time period (e.g., 2 seconds or 1 second, etc.), and determine whether the user has performed the first target action according to the measurement data of the IMU in the time period.
[0222] For example, according to the measurement data of the IMU in the time period, the electronic device can use a machine learning-based classification algorithm and / or a rule-based judgment algorithm to determine whether the user has performed the first target action.
[0223] The machine learning-based classification algorithm can include one or more of traditional machine learning methods such as a nearest neighbor algorithm, a support vector machine, a random forest, and deep learning methods such as a convolutional neural network or an autoencoder.
[0224] In some examples, the electronic device obtains a probability value of "detecting the first target action" by using one or more of the above classification algorithms, and in a case where the probability value is greater than or equal to a preset threshold, the electronic device determines that the user performs the first target action and continues to perform S203 and the operations thereafter. In a case where the probability value is less than the preset threshold, the electronic device determines that the user does not perform the first target action and can restart performing S201.
[0225] For the rule-based judgment algorithm, in some examples, the electronic device can first divide the action into large motion and small motion according to the size of the acceleration and the motion time. For example, the large motion can be that the electronic device is picked up from the desktop, the pocket of the trousers to the mouth, and the small motion can be that the electronic device is rotated by a certain angle from a position close to the face and then the lower microphone of the electronic device is close to the user's mouth.
[0226] Exemplarily, the electronic device can determine the displacement of the electronic device according to the acceleration and the motion time, and in a case where the displacement is greater than or equal to a displacement threshold, it is determined that the large motion occurs; in a case where the displacement is less than the displacement threshold, it is determined that the small motion occurs. For different motion conditions, the electronic device can use different methods to determine whether the first target action is detected, Figure 7 A determination method provided by the present application is shown.
[0227] For the large motion, the distance of the vertical upward motion can identify the "lifting up" action, such as lifting the electronic device from the user's thigh to the mouth. The distance of the motion in the horizontal direction can identify the action of "pulling the electronic device to the user", or the action of moving the electronic device from a position away from the user to a position close to the user, such as moving the electronic device on the desktop in front of the user away from the user to close to the user.
[0228] In some examples, in a case where the large motion occurs, the electronic device can determine whether the electronic device detects the first target action according to the displacement component of the electronic device in the gravity direction or the displacement component of the electronic device in the horizontal direction.
[0229] Exemplarily, the electronic device can determine the component of the acceleration in the gravity direction according to the included angle between the direction of the acceleration of the electronic device and the gravity direction, and the moving distance of the electronic device in the gravity direction can be obtained by twice integrating the acceleration in the gravity direction within the motion time. In a case where the moving distance in the gravity direction is greater than or equal to a vertical displacement threshold, the electronic device determines that the first target action is detected; in a case where the moving distance in the gravity direction is less than the vertical displacement threshold, the electronic device determines that the first target action is not detected.
[0230] Similarly, the electronic device can determine, according to the acceleration of the electronic device in the negative direction of the y-axis and the positive direction of the z-axis, a first horizontal plane acceleration component and a second horizontal plane acceleration component of the electronic device in the horizontal plane, and determine, according to the first horizontal plane acceleration component and the second horizontal plane acceleration component, a resultant acceleration of the electronic device in the horizontal plane, and the movement distance of the electronic device in the horizontal plane can be obtained by twice integrating the resultant acceleration in the movement time. In a case where the movement distance in the horizontal plane is greater than or equal to the horizontal displacement threshold, the electronic device determines that the first target action is detected; in a case where the movement distance in the horizontal plane is less than the horizontal displacement threshold, the electronic device determines that the first target action is not detected.
[0231] Here, since the negative direction of the y-axis and the positive direction of the z-axis of the electronic device generally point to the user, when calculating the distance moved in the horizontal direction, only the acceleration in the negative direction of the y-axis and the positive direction of the z-axis can be considered.
[0232] For small amplitude motion, the motion can be identified by the change of the angle between the three axes of the electronic device and the horizontal plane. The amplitude of the angle change should be within a certain range. In some examples, since the process of rotating the electronic device close to the user's mouth generally does not contain reciprocating motion, the change of the angle between the y-axis and the horizontal plane and the change of the angle between the x-axis and the horizontal plane are monotonously increasing or monotonously decreasing.
[0233] In some examples, in a case where small amplitude motion occurs, the electronic device can determine whether the target action is detected according to the angles between the three axes of the electronic device and the horizontal plane (the first angle α1, the second angle α2 and the third angle α3) during the motion.
[0234] Exemplarily, the electronic device can determine a first angle change value (β1) corresponding to α1, a second angle change value (β2) corresponding to α2 and a third angle change value (β3) corresponding to α3 during the motion of the electronic device, and according to the first angle change value, the second angle change value and the third angle change value, the electronic device can determine whether the target action is detected. For ease of illustration, here the first angle threshold (δ1), the second angle threshold (δ2) and the third angle threshold (δ3) represent the angle thresholds corresponding to β1, β2 and β3, respectively.
[0235] For example, in a case where β1≤δ1, β2≤δ3 and β3≤δ3, the electronic device can determine that the first target action is detected, otherwise, the electronic device can determine that the first target action is not detected.
[0236] Similarly, the electronic device can continuously record the first included angle and the second included angle in the process of the motion of the electronic device, and further determine a first angle increment (γ1) of the first included angle and a second angle increment (γ2) of the second included angle at different time points. According to the plurality of first angle increments and the plurality of second angle increments at different time points in the process of the motion, the electronic device can determine whether the target action is detected.
[0237] For example, in the process of the motion, all γ1 are positive or negative, and all γ2 are positive or negative. The electronic device can determine that the first target action is detected. Otherwise, the electronic device can determine that the first target action is not detected.
[0238] Still exemplarily, the electronic device can also determine whether the target operation is detected in combination with the above two determination manners. That is, according to the first angle change value, the second angle change value, and the third angle change value corresponding to the start and end time points of the motion of the electronic device, and the plurality of first angle increments and the plurality of second angle increments at different time points in the process of the motion, the electronic device can determine whether the target action is detected.
[0239] For example, in the process of the motion, all γ1 are positive or negative, and all γ2 are positive or negative. The electronic device can determine that the first target action is detected. Otherwise, the electronic device can determine that the first target action is not detected.
[0240] S203, determine that the measurement data of the proximity sensor and the measurement data of the ambient light sensor meet the requirements.
[0241] In the case that the lower microphone of the electronic device is close to the mouth of the user and does not belong to the posture of making or receiving a call, the upper part of the electronic device is not blocked. In this case, the distance determined based on the measurement data of the proximity sensor of the electronic device is greater than or equal to the distance threshold, that is, the measurement data of the proximity sensor does not indicate that the upper part of the electronic device is blocked, and the illuminance determined based on the measurement of the ambient light sensor is also generally greater than or equal to the illuminance threshold. In other words, in the case that the reading of the proximity sensor and the measurement data of the ambient light sensor both meet the aforementioned preset conditions, the electronic device is more likely to be in the first target posture. On the contrary, if one or both of the detection data of the proximity sensor and the measurement data of the ambient light sensor do not meet the preset conditions, the electronic device is less likely to be in the first target posture, for example, in this case, the electronic device can be in a pocket or a handbag, etc.
[0242] In the case that the measurement data of the proximity sensor and the measurement data of the ambient light sensor meet the requirements, the electronic device can perform S204 and subsequent operations. In the case that the measurement data of the proximity sensor and / or the measurement data of the ambient light sensor do not meet the requirements, the electronic device can re-perform the operation of S201.
[0243] S204, the lower speaker of the electronic device sends an ultrasonic signal, and the lower microphone detects the ultrasonic signal.
[0244] When the electronic device is determined to be in the first target posture and the electronic device detects the first target action, the electronic device can execute S204, that is, control the lower speaker to send an ultrasonic signal and the lower microphone to detect the ultrasonic signal.
[0245] If it is determined that the electronic device is not in the first target posture, or if the electronic device does not detect the first target action, the electronic device may re-execute the operation of S201.
[0246] S205, determine the distance between the microphone and the user's mouth and obtain the feature information of the user's mouth.
[0247] Based on the detection results of the ultrasonic signal by the lower microphone in S204, the distance between the lower microphone of the current electronic device and the user's mouth is determined, and the characteristic information of the user's mouth is acquired. Since the distance between different points of the user's mouth and the microphone varies, and these different points have different effects on the reflection or absorption of the ultrasonic signal, this information is contained in the audio segment of the microphone after the ultrasonic signal is reflected by the user's mouth. In other words, the characteristic information of the user's mouth can refer to the fact that the spectrum of the ultrasonic signal detected by the microphone after the ultrasonic signal is obstructed and reflected by the user's mouth contains the characteristic information of the user's mouth.
[0248] Figure 8 The image shows a processing method for determining the device posture of an electronic device using ultrasonic signals, as provided in an embodiment of this application.
[0249] An electronic device emits a chirp signal of a preset frequency and duration, and records audio for a certain duration. The recorded audio is filtered, and the spectrogram of an audio segment is extracted and used as input to a classifier. The classifier's output can be used to determine the distance between the lower microphone of the electronic device and the user's mouth. Here, the classifier can be an RNN, CNN, or SVM, etc.
[0250] Specifically, in some examples, the electronic device can generate a chirp signal of 0.02s, with a frequency range of 18kHz-23kHz, and send 0.2s at a time interval of 0.02s. The electronic device can record 100ms of audio, and apply a band-pass filter of 18kHz to 23kHz to the recorded audio signal. Matched filtering is performed using the original chirp signal and the recorded audio signal, and the maximum position in the result of the matched filtering can be considered as the position of the start of the chirp signal. A 0.02s audio segment is cut from the start position in the recorded audio, and the spectrum of the cut audio segment is calculated using short-time Fourier transform or fast Fourier transform as the input of the classifier, and the classifier can determine the classification result of each complete signal according to the input, which can be used to determine the distance between the lower microphone of the electronic device and the user's mouth.
[0251] For the case of determining that the lower microphone of the electronic device is close to the user's mouth by using the ultrasonic signal, the electronic device can perform S206, i.e., after opening the voice assistant application, feedback that the voice assistant is successfully woken up. Conversely, if the ultrasonic signal does not indicate that the lower microphone of the electronic device is close to the user's mouth, the electronic device can re-perform the operation of S201.
[0252] S206, feedback that the voice assistant is successfully woken up.
[0253] In some examples, as shown in FIG. 2A, the electronic device can display first prompt information 210 on the display screen, which can also be referred to as first prompt screen 210, the first prompt information 210 being used to prompt that the voice assistant is successfully woken up, and the first prompt information 210 can include text information and / or light information (or light effect information) and the like. Figure 9 In some examples, as shown in FIG. 2A, the electronic device can play first prompt audio 220 through the lower speaker, which is used to prompt that the voice assistant is successfully woken up.
[0254] Figure 9 In some examples, the electronic device can simultaneously issue prompt information through the display screen and the lower speaker, which is used to prompt that the voice assistant is successfully woken up.
[0255] In some examples, the electronic device can simultaneously issue prompt information through the display screen and the lower speaker, which is used to prompt that the voice assistant is successfully woken up.
[0256] The electronic device can also prompt for the voice assistant to be successfully woken up in other ways, such as vibration information and the like, and the electronic device can also change the feedback mode according to the user's custom operation, which is not limited by the present application.
[0257] In the case that the electronic device is in the first target posture, the prompt information can be more easily perceived by the user in the manner of displaying prompt information through the display screen and / or playing audio prompt information through the lower loudspeaker. This feedback manner makes the interaction between the user and the voice assistant more natural, and is conducive to improving the ease of use of the voice assistant.
[0258] S207, determine that the user is close to the lower microphone by using the dual-channel audio feature.
[0259] In order to determine that the audio recorded by the microphone comes from the user of the electronic device, or in other words, to determine that the user of the electronic device is speaking, in some examples, the electronic device can record audio through the upper microphone and the lower microphone at the same time as feeding back that the voice assistant is successfully awakened, and determine that the user is close to the lower microphone by using the dual-channel audio feature.
[0260] In some examples, the electronic device can determine that the user is close to the lower microphone by combining the audio information recorded by the upper microphone and the lower microphone. Figure 10 A manner of determining that the user is close to the lower microphone by combining the audio information recorded by the upper microphone and the lower microphone is shown.
[0261] Exemplarily, the electronic device can save an audio signal with a length of 200 ms, and make a judgment every 100 ms, and the results of multiple judgments can be saved as a history cache. In some examples, the electronic device can separate the audio recorded by the upper microphone and the lower microphone respectively, and use a low-pass filter with a cutoff frequency of 17 kHz to remove signals in the ultrasonic frequency band. Then, the dual-channel audio signal is used as the input of a CNN-based classifier, and the classification result output by the classifier is used to determine whether the user is close to the lower microphone of the electronic device.
[0262] For example, the output result of the aforementioned CNN-based classifier can give the probability of the following three cases, and the case with a probability greater than or equal to a preset threshold is regarded as detection.
[0263] Classification one: the upper microphone of the electronic device is close to the mouth of the user, for example, the distance between the upper microphone of the electronic device and the mouth of the user is within 5 cm, and the user holds the electronic device and speaks.
[0264] Classification two: the lower microphone of the electronic device is close to the mouth of the user, for example, the distance between the lower microphone of the electronic device and the mouth of the user is within 5 cm, and the user holds the electronic device and speaks.
[0265] Category 3: the case where the electronic device is not close to the user's mouth, or other cases, such as the user normally uses the electronic device, for example, looks at the screen and speaks, or the user speaks in an environment where there are other people around, or the user speaks in a quiet environment, or the user speaks in a noisy environment, and the like.
[0266] In some examples, the electronic device can also save the results of the last n times of determination, vote for different types of results in the last n times of determination through a voting mechanism, and regard the result of the majority vote as the final output, where n is an integer greater than or equal to 1.
[0267] In S208, a microphone is selected according to the sound pickup strategy, and automatic speech recognition is performed.
[0268] In the case where the lower microphone of the electronic device is close to the user's mouth, the electronic device can at least record the user's voice command through the lower microphone, for example, the electronic device can execute the sound pickup strategy of taking the lower microphone as the main microphone and the upper microphone as the auxiliary microphone. Illustratively, the electronic device can use the audio recorded by the upper microphone to assist in noise reduction processing of the audio recorded by the lower microphone, and the like.
[0269] In S209, the voice assistant of the electronic device executes the command and generates feedback according to the corresponding feedback mode.
[0270] After the electronic device invokes the language assistant to execute the command, the electronic device can feed back the execution result of the command in a manner similar to the feedback of the successful wake-up of the language assistant in S206. For brevity, the details are not repeated here.
[0271] In some examples, the electronic device can adopt different feedback modes for different execution results. For example, different types of feedback modes, or feedback modes with different contents of the same type. The present application does not make any limitation in this regard.
[0272] Figure 11 A method for a user to interact with a voice assistant when the electronic device is in a state where the upper microphone is close to the user's mouth is shown.
[0273] In S301, it is determined that the electronic device is in a state where the upper microphone is close to the user's mouth.
[0274] In some examples, in a relatively stable state, the electronic device determines whether its posture is in a state where the upper microphone is close to the user's mouth.
[0275] Whether the electronic device is in a relatively stable state can be determined according to one or more sensors on the electronic device. Illustratively, the electronic device can determine according to the measurement data of the IMU.
[0276] In some examples, the posture of the electronic device can be determined according to one or more sensors on the electronic device. Illustratively, the electronic device can determine the posture of the electronic device according to the measurement data of the gravity sensor. For example, the electronic device can determine the components of the gravitational acceleration of the electronic device on three axes of the electronic device according to the measurement data of the gravity sensor, and determine the included angle between the three axes of the electronic device and the horizontal plane according to the components of the gravitational acceleration on the three axes, which can be used to indicate the posture of the electronic device. For the sake of illustration, the following content will refer to the posture of the electronic device in which the lower microphone is close to the user's mouth as the third target posture.
[0277] In some examples, the posture of the electronic device can be determined according to one or more sensors on the electronic device. Illustratively, the electronic device can determine the posture of the electronic device according to the measurement data of the gravity sensor. For example, the electronic device can determine the components of the gravitational acceleration of the electronic device on three axes of the electronic device according to the measurement data of the gravity sensor, and determine the included angle between the three axes of the electronic device and the horizontal plane according to the components of the gravitational acceleration on the three axes, which can be used to indicate the posture of the electronic device. For the sake of illustration, the following content will refer to the posture of the electronic device in which the lower microphone is close to the user's mouth as the third target posture. Figure 6 As shown in the perspective view of the electronic device, the included angles between the x-axis, y-axis and z-axis of the electronic device and the horizontal plane can be referred to as the first included angle (α1), the second included angle (α2) and the third included angle (α3), respectively. In some examples, when the first included angle, the second included angle and the third included angle all satisfy the preset condition, the electronic device can determine that it is in the third target posture; on the contrary, when any one of the first included angle, the second included angle and the third included angle does not satisfy the preset condition, the electronic device can determine that it is not in the third target posture.
[0278] The specific execution method of S301 is similar to that of S201, and for the sake of brevity, the relevant content in S201 will not be repeated here.
[0279] In S302, it is determined that the user has performed an action of moving the upper microphone of the electronic device close to the mouth.
[0280] In order to improve the accuracy of the electronic device in judging the user's intention to wake up the voice assistant, reduce the probability of the voice assistant of the electronic device being mistakenly woken up, and reduce the increase in power consumption caused by using the voice assistant, in the case where the electronic device is in the third target posture, the electronic device can determine whether it has experienced a change from a non-third target posture to a third target posture. In other words, the electronic device can determine whether the user has adjusted the electronic device from a non-third target posture to a third target posture. For the sake of illustration, the following will refer to the action of adjusting the electronic device from a non-third target posture to a third target posture as the third target action.
[0281] In some examples, the electronic device can save the data of the sensor for a period of time, and determine whether the user has performed the third target action according to the data of the sensor in the time period. Illustratively, the electronic device can continuously save the measurement data of the IMU for a certain time length (for example, 2 seconds or 1 second, etc.), and determine whether the user has performed the above-mentioned action according to the measurement data of the IMU in the time length interval.
[0282] For example, according to the measurement data of the IMU in the time interval, the electronic device can determine whether the user performs the third target action by using a classification algorithm based on machine learning and / or a judgment algorithm based on rules.
[0283] The classification algorithm based on machine learning can include one or more of traditional machine learning methods such as a nearest neighbor algorithm, a support vector machine, a random forest, and deep learning methods such as a convolutional neural network or an autoencoder.
[0284] In some examples, the electronic device obtains a probability value of "detecting the third target action" by using one or more of the above classification algorithms, and in a case where the probability value is greater than or equal to a preset threshold, the electronic device determines that the user performs the third target action and continues to perform S203 and subsequent operations. In a case where the probability value is less than the preset threshold, the electronic device determines that the user does not perform the third target action and can restart performing S301.
[0285] For the judgment algorithm based on rules, the electronic device can determine whether the third target action is detected according to the movement distance of itself in the vertical direction and / or the angle of rotation around the x-axis. Figure 12 An example method for an electronic device to determine whether a third target action is detected is provided.
[0286] In some examples, the electronic device can determine whether the movement distance of itself in the vertical direction is greater than a vertical displacement threshold, and in a case where the movement distance in the vertical direction is greater than or equal to the vertical displacement threshold, determine that the third target action is detected, and in a case where the movement distance in the vertical direction is less than the vertical displacement threshold, determine that the third target action is not detected.
[0287] In some examples, the electronic device can determine whether the angle of rotation of itself around the x-axis is greater than an angle threshold, and in a case where the angle of rotation around the x-axis is greater than or equal to the angle threshold, determine that the third target action is detected, and in a case where the angle of rotation around the x-axis is less than the angle threshold, determine that the third target action is not detected.
[0288] S303, determine that the measurement data of the proximity sensor and the measurement data of the ambient light sensor meet the requirements.
[0289] In some examples, in a case where the upper microphone of the electronic device is close to the mouth of the user, the proximity sensor is blocked, the reading of the ambient light sensor is reduced, or in other words, based on the distance detected by the proximity sensor being less than or equal to the distance threshold, the illuminance detected by the ambient light sensor being less than or equal to the illuminance threshold. In other words, in a case where both the reading of the proximity sensor and the measurement data of the ambient light sensor meet the preset condition, the electronic device is more likely to be in a posture in which the upper microphone is close to the mouth of the user, and conversely, if one or both of the detection data of the proximity sensor and the measurement data of the ambient light sensor do not meet the preset condition, the electronic device is less likely to be in a posture in which the upper microphone is close to the mouth of the user, for example, in which case the electronic device can be in a pocket or a handbag, etc.
[0290] In a case where it is determined that the measurement data of the proximity sensor and the measurement data of the ambient light sensor meet the requirements, the electronic device can perform the operation of S304 and thereafter, and in a case where it is determined that the measurement data of the proximity sensor and / or the measurement data of the ambient light sensor do not meet the requirements, the electronic device can re-perform the operation of S301.
[0291] S304, the upper speaker of the electronic device sends an ultrasonic signal, and the upper microphone detects the ultrasonic signal.
[0292] In a case where it is determined that the electronic device is in the third target posture and the electronic device detects the third target action, the electronic device can perform S304, that is, control the upper speaker to send an ultrasonic signal and the upper microphone to detect the ultrasonic signal.
[0293] In a case where it is determined that the electronic device is not in the third target posture, or in a case where the electronic device does not detect the third target action, the electronic device can re-perform the operation of S301.
[0294] S305, determining the distance between the upper microphone and the mouth of the user and obtaining feature information of the mouth of the user.
[0295] The distance between the upper microphone of the electronic device and the mouth of the user is determined according to the detection result of the ultrasonic signal by the upper microphone in S304, and the feature information of the mouth of the user is obtained. Since different point positions of the mouth of the user have different distances from the microphone, different point positions have different effects on reflection or absorption of the ultrasonic signal, and the audio segment detected by the microphone after the ultrasonic signal is reflected by the mouth of the user contains these information. In other words, the feature information of the mouth of the user can refer to the feature information of the mouth of the user contained in the frequency spectrum of the ultrasonic signal detected by the microphone after the ultrasonic signal is reflected by the mouth of the user.
[0296] The electronic device sends a chirp signal of a preset frequency and a preset duration, and records audio for a certain duration, filters the recorded audio, and intercepts a frequency spectrum of an audio segment in the recorded audio as an input of a classifier. An output result of the classifier can be used to determine the distance between the lower microphone of the electronic device and the user's mouth. Here, the classifier can be an RNN, a CNN, or an SVM, etc.
[0297] The specific implementation method of S305 is similar to that of S205. For brevity, details are not repeated here.
[0298] For the case of determining that the upper microphone of the electronic device is close to the user's mouth by using the ultrasonic signal, the electronic device can perform S306, that is, after the voice assistant application is opened, feedback that the voice assistant is successfully woken up. Conversely, if the ultrasonic signal does not indicate that the upper microphone of the electronic device is close to the user's mouth, the electronic device can re-perform the operation of S301.
[0299] S306, feedback that the voice assistant is successfully woken up.
[0300] In some examples, the electronic device can display second prompt information on the display screen, which can also be referred to as a second prompt screen. The second prompt information is used to prompt that the voice assistant is successfully woken up. The second prompt information can be text information or light information (or light effect information), etc.
[0301] In some examples, the electronic device can play third prompt audio through the upper loudspeaker, which is used to prompt that the voice assistant is successfully woken up.
[0302] In some examples, the electronic device can simultaneously send prompt information through the display screen and the upper loudspeaker, which is used to prompt that the voice assistant is successfully woken up.
[0303] The electronic device can also use other ways to prompt that the voice assistant is successfully woken up, such as vibration information, etc. The electronic device can also change the feedback mode according to the user's custom operation, which is not limited in the present application.
[0304] When the electronic device is in the third target posture, the prompt information can be more easily perceived by the user by playing audio prompt information through the upper loudspeaker. This feedback mode makes the interaction between the user and the voice assistant more natural, which is conducive to improving the ease of use of the voice assistant.
[0305] S307, determine that the user is close to the lower microphone by using the dual-channel audio feature.
[0306] To determine that the audio recorded by the microphones is from the user of the electronic device, or in other words, to determine that the user is speaking into the electronic device, in some examples, the electronic device can record audio through both the upper microphone and the lower microphone at the same time while the feedback voice assistant is successfully woken up, and determine that the user is speaking close to the upper microphone using the binaural audio features.
[0307] In some examples, the electronic device can determine that the user is speaking close to the lower microphone in combination with the audio information recorded by the upper microphone and the lower microphone.
[0308] Exemplarily, the electronic device can save an audio signal with a length of 200 ms, and make a judgment every 100 ms, and the results of multiple judgments can be saved as a history cache. In some examples, the electronic device can separate the audio recorded by the upper microphone and the lower microphone respectively, and use a low-pass filter with a cutoff frequency of 17 kHz to remove signals in the ultrasonic frequency band. Then, the binaural audio signals are used as the input of a CNN-based classifier, and the classification result output by the classifier is used to determine whether the user is speaking close to the lower microphone of the electronic device.
[0309] In some examples, the electronic device can also save the output results of the last n times, vote for different types of results in the output results of the last n times through a voting mechanism, and regard the result of the majority vote as the final output, where n is an integer greater than or equal to 1.
[0310] The specific implementation method of S307 is similar to that of S207, and specifically can refer to the related content in S207. For the sake of brevity, no further description is given here.
[0311] S308, selecting a microphone according to the sound pickup strategy, and performing automatic speech recognition.
[0312] In the case that the upper microphone of the electronic device is close to the mouth of the user, the electronic device can at least record the voice command of the user through the upper microphone, for example, execute the sound pickup strategy in which the upper microphone is the main microphone and the lower microphone is the auxiliary microphone. Exemplarily, the electronic device can use the audio recorded by the lower microphone to assist in noise reduction processing of the audio recorded by the upper microphone, etc.
[0313] S309, the voice assistant of the electronic device executes the command, and generates feedback according to the corresponding feedback mode.
[0314] After the electronic device invokes the language assistant to execute the command, the electronic device can feed back that the command execution is completed in a manner similar to the feedback that the language assistant is successfully woken up in S306. For the sake of brevity, the content in S306 is specifically referred to here.
[0315] In some examples, the electronic device can employ different feedback manners for different execution results. For example, different types of feedback manners, or feedback manners with different contents. The present application does not limit this.
[0316] Figure 13 Another voice assistant interaction method provided by the embodiments of the present application is shown. In this method, the electronic device can determine that the electronic device is in a posture for making or receiving a call (the posture of the electronic device 100 in FIG. 402) according to the sensors of the electronic device, and accordingly execute a corresponding wake-up scheme, a sound pickup scheme, and a feedback scheme.
[0317] S401, determining that the electronic device is in a posture for making or receiving a call.
[0318] In some examples, in a relatively stable state, the electronic device can determine whether the posture of the electronic device is a posture for making or receiving a call.
[0319] Whether the electronic device is in a relatively stable state can be determined according to one or more sensors on the electronic device. For example, the electronic device can determine according to the measurement data of the IMU. For example, the electronic device can determine the magnitude of the current acceleration of the electronic device according to the measurement data of the IMU, and in the case that the acceleration of the electronic device is less than or equal to an acceleration threshold, the electronic device can determine that it is in a relatively stable state. On the contrary, in the case that the acceleration of the electronic device is greater than the acceleration threshold, the electronic device can determine that it is in a motion state.
[0320] It should be noted that the relatively stable state here can also be understood as a relatively stationary state, and stability, stationarity and motion are all relative descriptions.
[0321] In some examples, the posture of the electronic device can be determined according to one or more sensors on the electronic device. For example, the electronic device can determine the posture of the electronic device according to the measurement data of the gravity sensor. For example, the electronic device can determine the components of the gravitational acceleration of the electronic device on the three axes of the electronic device according to the measurement data of the gravity sensor, and determine the included angle between the three axes of the electronic device and the horizontal plane according to the components of the gravitational acceleration on the three axes, which can be used to indicate the posture of the electronic device. For ease of description, the following content will refer to the "posture of the electronic device for making or receiving a call" as the second target posture.
[0322] For example, the electronic device can determine the components of the gravitational acceleration of the electronic device on the three axes of the electronic device according to the measurement data of the gravity sensor, and determine the included angle between the three axes of the electronic device and the horizontal plane according to the components of the gravitational acceleration on the three axes, which can be used to indicate the posture of the electronic device. Figure 6The angles between the x-axis, y-axis and z-axis of the electronic device and the horizontal plane 501 can be referred to as a first angle (a1), a second angle (a2) and a third angle (a3) respectively. In some examples, when the first angle, the second angle and the third angle all satisfy a preset condition, the electronic device can determine that it is in the second target posture; on the contrary, when any one of the first angle, the second angle and the third angle does not satisfy the preset condition, the electronic device can determine that it is not in the second target posture.
[0323] S402, it is determined that the user performs the action of holding the electronic device close to the ear.
[0324] In order to improve the accuracy of the electronic device in judging the user's intention to wake up the voice assistant, reduce the probability of the voice assistant of the electronic device being mistakenly woken up, and reduce the increase in power consumption caused by using the voice assistant, in the case where it is determined that the electronic device is in the posture of making or receiving a call, the electronic device can determine whether it has experienced a change from a non-second target posture to a second target posture. In other words, the electronic device can determine whether the user has performed an action of adjusting the electronic device from a non-second target posture to a second target posture. For ease of description, the action of "adjusting the electronic device from a non-second target posture to a second target posture" is referred to as a second target action.
[0325] In some examples, the electronic device can save the data of the sensor within a period of time, and determine whether the user has performed the second target action according to the data of the sensor within the period of time. For example, the electronic device can continuously save the measurement data of the IMU for a certain period of time (e.g., 2 seconds or 1 second, etc.), and determine whether the user has performed the above-mentioned action according to the measurement data of the IMU within the period of time.
[0326] For example, according to the measurement data of the IMU within the period of time, the electronic device can use a machine learning-based classification algorithm and / or a rule-based judgment algorithm to determine whether the user has performed the second target action.
[0327] The machine learning-based classification algorithm can include one or more of traditional machine learning methods such as a nearest neighbor algorithm, a support vector machine, a random forest, and deep learning methods such as a convolutional neural network or an autoencoder.
[0328] In some examples, the electronic device obtains a probability value of "detecting the second target action" using one or more of the above-mentioned classification algorithms, and in the case where the probability value is greater than or equal to a preset threshold, the electronic device determines that the user has performed the second target action and continues to perform S403 and subsequent operations. In the case where the probability value is less than the preset threshold, the electronic device determines that the user has not performed the second target action, and can restart performing S401.
[0329] For the rule-based determination algorithm, the electronic device can determine whether the target action is detected according to the moving distance of itself in the vertical direction and / or the angle turned by itself around the three axes (x-axis, y-axis and z-axis). Figure 14 An example is provided for a method for an electronic device to determine whether a second target action is detected.
[0330] In some examples, the electronic device can determine whether the moving distance of itself in the vertical direction is greater than a vertical displacement threshold value, determine that the second target action is detected in a case where the moving distance in the vertical direction is greater than or equal to the vertical displacement threshold value, and determine that the second target action is not detected in a case where the moving distance in the vertical direction is less than the vertical displacement threshold value.
[0331] In some examples, the electronic device can determine whether the angle turned by itself around the three axes is greater than an angle threshold value, determine that the second target action is detected in a case where the angle turned around the three axes is greater than or equal to the angle threshold value, and determine that the second target action is not detected in a case where the angle turned around the three axes is less than the angle threshold value.
[0332] For example, the angles turned by the electronic device around the x-axis, the y-axis and the z-axis are a first angle (θ1), a second angle (θ2) and a third angle (θ3) respectively. The angle threshold values corresponding to the first angle, the second angle and the third angle are a first angle threshold value (σ1), a second angle threshold value (σ2) and a third angle threshold value (σ3) respectively. In a case where θ1≥σ1, θ2≥σ2 and θ3≥σ3, the electronic device determines that the second target action is detected; for a case where any one of the three angles does not satisfy the foregoing condition, the electronic device determines that the second target action is not detected.
[0333] For example, in a case where θ1≥σ1, θ2≥σ2 and θ3≥σ3, the electronic device can determine that it is near the left ear of the user, or in other words, the electronic device can determine that the user has performed an action of placing the electronic device near the left ear.
[0334] For another example, in a case where θ1≥σ1, θ2≤-σ2 and θ3≤-σ3, the electronic device can determine that it is near the right ear of the user, or in other words, the electronic device can determine that the user has performed an action of placing the electronic device near the right ear.
[0335] S403, determine that the measurement data of the proximity sensor and / or the measurement data of the ambient light sensor meet the requirements.
[0336] In most cases, the proximity sensor will trigger during a phone call action. In a few cases, the proximity sensor does not trigger, and in this case, it can be determined whether the phone is close to the ear by whether the reading of the ambient light sensor is reduced. If the above conditions are not met, it is determined that the phone call action is not detected.
[0337] In some examples, the electronic device can determine that the second target action is detected according to the event that the proximity sensor triggers. Illustratively, during the movement of the electronic device, if the distance determined based on the measurement data of the proximity sensor is less than or equal to the distance threshold, the electronic device determines that the second target action is detected; if the distance determined based on the measurement data of the proximity sensor is greater than the distance threshold, the electronic device determines that the second target action is not detected.
[0338] In some examples, the electronic device can determine whether the electronic device is close to the ear, i.e., whether the second target action is detected, according to whether the measurement value of the ambient light sensor decreases during the movement of the electronic device. Illustratively, during the movement of the electronic device, if the measurement value of the ambient light sensor decreases, it is determined that the second target action is detected. Conversely, if the measurement value of the ambient light sensor does not decrease, it is determined that the second target action is not detected.
[0339] In some examples, the electronic device can also determine whether the second target action is detected in combination with the measurement data of the proximity sensor and the measurement data of the ambient light sensor. Illustratively, in the case that the measurement data of the proximity sensor and the measurement data of the ambient light sensor are both within the preset threshold range, the electronic device determines that the second target action is detected. Otherwise, the electronic device determines that the second target action is not detected.
[0340] S404, determining that the electronic device is close to the ear of the user.
[0341] When the electronic device is close to the ear, the reading of the capacitive screen of the electronic device changes due to the contact between the pinna and the screen of the electronic device, and the shape of the pinna can be recognized by the reading of the capacitive screen.
[0342] In some examples, the reading of the capacitive screen can be represented by a two-dimensional matrix, and in general cases, the points with larger values represent the areas in contact with the skin. The reading of the capacitive screen can be used to determine whether the electronic device is close to the ear of the user. Illustratively, the electronic device can use machine learning-based methods such as kNN, SVM, CNN, etc. to analyze the reading of the capacitive screen to determine whether the target data is included in the reading of the capacitive screen, the target data being used to indicate the shape of the pinna, or in other words, the target data including the information of the pinna of the user, so as to determine whether the electronic device is close to the ear of the user.
[0343] In some scenarios, the above determining whether the electronic device is close to the user's ear according to whether the target data is included in the capacitance screen reading can also be understood as: determining whether the electronic device is close to the user's ear according to touch information of the display screen of the electronic device. Exemplarily, the touch information can include a capacitance value of the display screen of the electronic device, which can be used to indicate information of the user's external auditory meatus.
[0344] For the case of determining that the electronic device is close to the user's ear by using the touch information, the electronic device can perform S405, that is, after the voice assistant application is opened, feeding back that the voice assistant is successfully woken up. Conversely, if the touch information does not indicate that the electronic device is close to the user's ear, the electronic device can re-perform the operation of S401.
[0345] S405, feeding back that the voice assistant is successfully woken up.
[0346] In some examples, the electronic device can play second prompt audio through the lower speaker, the second prompt audio being used to prompt that the voice assistant is successfully woken up.
[0347] The electronic device can also prompt that the voice assistant is successfully woken up in other manners, such as vibration information and the like, and the electronic device can also change the feedback manner according to the user's custom operation, which is not limited in the present application.
[0348] In the case that the electronic device is in the second target posture, the prompt information can be more easily perceived by the user by using the lower speaker to play the audio prompt information, and this feedback manner makes the interaction between the user and the voice assistant more natural, which is beneficial to improving the ease of use of the voice assistant.
[0349] S406, determining that the user is close to the lower microphone to speak by using a binaural audio feature.
[0350] In order to determine that the audio recorded by the microphone comes from the user of the electronic device, or in other words, in order to determine that the person who is speaking is the user of the electronic device, in some examples, the electronic device can record audio through the upper microphone and the lower microphone at the same time while feeding back that the voice assistant is successfully woken up, and determine that the user is close to the lower microphone to speak by using a binaural audio feature.
[0351] In some examples, the electronic device can determine that the user is close to the lower microphone to speak in combination with audio information recorded by the upper microphone and the lower microphone. In the above Figure 10 Fig. 6 shows a manner of determining that the user is close to the lower microphone to speak in combination with audio information recorded by the upper microphone and the lower microphone according to an embodiment of the present application.
[0352] For example, the electronic device can store an audio signal of 200ms in length and make a judgment every 100ms, with multiple judgment results stored as a history cache. In some examples, the electronic device can separate the audio recorded by the upper and lower microphones and use a low-pass filter with a cutoff frequency of 17kHz to remove signals in the ultrasonic frequency band. The dual-channel audio signal is then used as input to a CNN-based classifier, and the classification result output by the classifier is used to determine whether the user is speaking close to the lower microphone of the electronic device.
[0353] In some examples, electronic devices can also save the most recent n output results, and use a voting mechanism to vote on different types of results in the most recent n output results, and regard the result of the majority vote as the final output, where n is an integer greater than or equal to 1.
[0354] The specific execution method of S406 is similar to that of S207. For details, please refer to the relevant content in S207. For the sake of brevity, it will not be elaborated here.
[0355] The S407 selects a microphone based on a pickup strategy and performs automatic voice recognition.
[0356] When an electronic device is in the position of making or receiving a phone call, it can record the user's voice commands at least through the lower microphone. For example, it can execute a pickup strategy that primarily uses the lower microphone and secondarily uses the upper microphone. For instance, the electronic device can use the audio recorded by the upper microphone to perform auxiliary noise reduction processing on the audio recorded by the lower microphone.
[0357] S408: The voice assistant executes commands and generates feedback based on the corresponding feedback method.
[0358] After an electronic device invokes a voice assistant to execute a command, it can provide feedback that the command execution is complete, similar to the feedback of successful voice assistant activation in S405. For details, please refer to the content in S405; for the sake of brevity, it will not be elaborated here.
[0359] In some examples, electronic devices may employ different feedback methods for different execution results. For example, different types of feedback methods, or feedback methods with different content of the same type. This application does not impose any limitations on this.
[0360] The above text combined Figures 1 to 14 The method embodiments of this application are described in detail below, in conjunction with... Figure 15 and Figure 16 The apparatus embodiments of this application are described below. It should be understood that the descriptions of the method embodiments correspond to the descriptions of the apparatus embodiments; therefore, any parts not described in detail can be referred to the foregoing method embodiments.
[0361] Figure 15 The device 1500 for voice assistant interaction provided by the embodiments of the present application can have the functions of the electronic device in the method embodiments, and can be used to execute the steps executed by the functions of the electronic device in the method embodiments. The functions can be implemented by hardware, or by software or hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.
[0362] In a possible implementation, the device 1500 for voice assistant interaction can include an acquisition module 1510 and a processing module 1520, which are coupled to each other.
[0363] The acquisition module 1510 can be configured to support the electronic device to acquire the input of the user, for example, to execute the operation of recording the audio information in the above method embodiments. Figure 3
[0364] The processing module 1520 is configured to support the electronic device to execute the processing actions in the above method embodiments, for example, to send feedback information, to confirm the device posture, to execute the device action, and the like.
[0365] Optionally, the device 1500 for voice assistant interaction can further include a storage unit 1530 configured to store the program code and data of the device 1500 for voice assistant interaction.
[0366] Figure 16 An electronic device 1600 provided by the embodiments of the present application is shown in the figure. The electronic device 1600 includes at least one processor 1610 and a transceiver 1620. The processor 1610 is coupled to a memory and is configured to execute the instructions stored in the memory to control the transceiver 1620 to send and / or receive signals.
[0367] Optionally, the electronic device 1600 further includes a memory 1630 configured to store instructions.
[0368] In some embodiments, the processor 1610 and the memory 1630 can be combined into a processing device. The processor 1610 is configured to execute the program code stored in the memory 1630 to implement the above functions. In specific implementation, the memory 1630 can be integrated in the processor 1610 or independent of the processor 1610.
[0369] In some embodiments, the transceiver 1620 can include a receiver (or receiver) and a transmitter (or transmitter).
[0370] The transceiver 1620 can further include an antenna, and the number of antennas can be one or more. The transceiver 1620 can be a communication interface or an interface circuit.
[0371] When the electronic device 1600 is a chip, the chip includes a transceiver module and a processing module. The transceiver module can be an input / output circuit or a communication interface, and the processing module can be a processor or a microprocessor integrated on the chip or an integrated circuit.
[0372] The embodiment further provides a computer readable storage medium, which stores computer instructions. When the computer instructions are run on an electronic device, the electronic device executes the related method steps to implement the method for preloading a shader in the above embodiment.
[0373] The embodiment further provides a computer program product. When the computer program product is run on a computer, the computer executes the related steps to implement the method for preloading a shader in the above embodiment.
[0374] In addition, the embodiment of the present application further provides a device, which can be a chip, a component or a module. The device can include a processor and a memory connected to each other. The memory is used for storing computer execution instructions. When the device is running, the processor can execute the computer execution instructions stored in the memory, so that the chip executes the method for preloading a shader in the above method embodiments.
[0375] The electronic device, the computer readable storage medium, the computer program product or the chip provided by the embodiment are used to execute the corresponding method provided above, and thus the beneficial effects achieved thereby can refer to the beneficial effects of the corresponding method provided above, which will not be described herein again.
[0376] Those skilled in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0377] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described herein again.
[0378] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic. The division of the units is merely logical function division. There can be other division manners in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0379] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0380] In addition, each functional unit in the various embodiments of the present application can be integrated into a processing unit, or each unit can be a physically separate unit, or two or more units can be integrated into one unit.
[0381] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.
[0382] The above description is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for voice assistant interaction, applied to electronic devices, characterized in that, Comprising: detecting a target action of the electronic device, determining that the electronic device is in a target posture, the target posture comprising a posture in which a lower microphone of the electronic device is close to a user's mouth, the target action comprising an action for converting the electronic device from a non-target posture to the target posture; in response to the target posture of the electronic device, displaying and / or playing first feedback information, the first feedback information being used to indicate that the voice assistant is successfully woken up; recording a voice instruction at least through the lower microphone of the electronic device; in response to the voice instruction, displaying and / or playing second feedback information, the second feedback information being used to indicate the execution result of the voice instruction.
2. The method of claim 1, wherein, The target posture comprises a posture in which the lower microphone is close to the user's mouth, and the ambient light sensor and / or the proximity sensor are not blocked, The response to the target posture of the electronic device, displaying and / or playing first feedback information, comprises: in response to the target posture of the electronic device, displaying a prompt picture and / or playing a first prompt audio through the lower speaker of the electronic device, the first feedback information comprising the prompt picture and / or the first prompt audio.
3. The method of claim 2, wherein, The response to the target posture of the electronic device, displaying a prompt picture and / or playing a first prompt audio through the lower speaker of the electronic device, comprises: in response to the target posture of the electronic device, sending a first ultrasonic signal; receiving a second ultrasonic signal returned based on the first ultrasonic signal; in a case where the second ultrasonic signal indicates that the electronic device is in the target posture, displaying the prompt picture and / or playing the first prompt audio through the lower speaker of the electronic device.
4. The method of claim 3, wherein The response to the target posture of the electronic device, sending a first ultrasonic signal, comprises: in response to the target posture of the electronic device, sending the first ultrasonic signal through the lower speaker; The receiving a second ultrasonic signal returned based on the first ultrasonic signal, comprises: in response to the second ultrasonic signal returned by the first ultrasonic signal contacting the user, receiving the second ultrasonic signal through the lower microphone.
5. The method of claim 1, wherein, The target posture comprises a posture in which the lower microphone of the electronic device is close to the user's mouth, and the ambient light sensor and / or the proximity sensor are blocked, The response to the target posture of the electronic device, displaying and / or playing first feedback information, comprises: in response to the target posture of the electronic device, playing a second prompt audio through the upper speaker of the electronic device, the first feedback information comprising the second prompt audio.
6. The method of claim 5, wherein, The response to the target posture of the electronic device, playing a second prompt audio through the upper speaker of the electronic device, comprises: in response to the target posture of the electronic device, obtaining touch information of a display screen of the electronic device; in a case where the touch information indicates that the electronic device is in a state in which the display screen is touched, playing the second prompt audio through the upper speaker of the electronic device.
7. The method of claim 6, wherein, The touch information comprises information of an outer ear contour.
8. The method according to any one of claims 1 to 7, characterized in that, The at least through the lower microphone of the electronic device record voice instruction, including: with the lower microphone of the electronic device as the main recording unit, the upper microphone of the electronic device is the auxiliary recording unit record the voice instruction.
9. The method of claim 8, wherein, Before recording the first voice instruction with the lower microphone of the electronic device as the main recording unit, and the upper microphone of the electronic device as the auxiliary recording unit, the method further comprises: Record first audio through the upper microphone and the lower microphone; According to the first audio, it is determined that the user speaks close to the lower microphone. 10.A method of voice assistant interaction, applied to an electronic device, the method comprising: Including: Detecting the target action of the electronic device, determining that the electronic device is in the target posture, the target posture includes the posture of the upper microphone of the electronic device close to the mouth of the user, and the target action includes the action for converting the electronic device from the non-target posture to the target posture; In response to the target posture of the electronic device, play first feedback information, the first feedback information is used to indicate that the voice assistant wakes up successfully; At least through the upper microphone of the electronic device record voice instruction; In response to the voice instruction, play second feedback information, the second feedback information is used to indicate the execution result of the voice instruction.
11. The method of claim 10, wherein, The response to the target posture of the electronic device, play first feedback information, including: In response to the target posture of the electronic device, play the first feedback information through the upper speaker of the electronic device.
12. The method of claim 11, wherein, The response to the target posture of the electronic device, play first feedback information through the upper speaker of the electronic device, including: In response to the target posture of the electronic device, send first ultrasonic signal; Receive the second ultrasonic signal returned based on the first ultrasonic signal; In the case where the second ultrasonic signal indicates that the electronic device is in the target posture, play the first feedback information through the upper speaker of the electronic device.
13. The method of claim 12, wherein The response to the target posture of the electronic device, send first ultrasonic signal, including: In response to the target posture of the electronic device, send the first ultrasonic signal through the upper speaker; The receiving second ultrasonic signal returned based on the first ultrasonic signal, including: In response to the second ultrasonic signal returned by the first ultrasonic signal contacting the user, the second ultrasonic signal is received through the upper microphone.
14. The method according to any one of claims 10 to 13, characterized in that, The at least through the lower microphone of the electronic device record voice instruction, including: with the lower microphone of the electronic device as the main recording unit, the upper microphone of the electronic device is the auxiliary recording unit record the voice instruction.
15. The method according to any one of claims 10 to 13, characterized in that, Before recording the voice instruction with the upper microphone of the electronic device as the main recording unit, and the lower microphone of the electronic device as the auxiliary recording unit, the method further comprises: Record first audio through the upper microphone and the lower microphone; According to the first audio, it is determined that the user speaks close to the upper microphone.
16. An electronic device, comprising: A computer program product comprising a computer readable medium having stored thereon instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9 or claims 10 to 15.
17. A memory management device, comprising: A computer program product comprising a computer readable medium having stored thereon instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9 or claims 10 to 15.
18. A computer-readable storage medium, characterized in that, A computer program product comprising a computer readable medium having stored thereon instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9 or claims 10 to 15.
19. A chip, characterized by A computer program product comprising a computer readable medium having stored thereon instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9 or claims 10 to 15.
Citation Information
Patent Citations
Man-machine interaction method and device, electronic equipment and storage medium
CN110689889A
Audio processing method, electronic equipment and computer readable storage medium
CN113994426A