Voice interaction method, device and terminal
By detecting the user's breath and determining whether the terminal is close to the user's mouth, the problem of the terminal's feedback to other people or noise during extended reception is solved, thereby improving the accuracy of voice interaction and user experience.
Patent Information
- Application Number
- CN202211113419.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-14
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2042-09-14
AI Technical Summary
In the prior art, if the user does not speak during the extended reception time and someone is speaking nearby, the terminal will give feedback to other people or noise, resulting in a poor user experience.
By detecting the user's breath and/or judging whether the terminal is close to the user's mouth, it determines whether to continue receiving sound and providing feedback, including using sensors such as pressure, temperature, humidity, and carbon dioxide sensors, as well as gesture recognition modules to ensure that feedback is targeted at the user's voice interaction.
It effectively reduces incorrect responses to other people or noise, and improves the accuracy of voice interaction and user experience.
Smart Images

Figure CN117746849B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of human-computer interaction technology, and in particular relates to a voice interaction method, device and terminal. Background Art
[0002] Voice interaction is a new generation of interaction mode based on voice input. Based on the voice information input by the user to the terminal, feedback results corresponding to the input voice information can be obtained.
[0003] Before interacting with the terminal by voice, you must first wake up the voice interaction system (such as a voice assistant) on the terminal. For example, you can use a specific wake-up word to wake up the voice assistant. After the voice assistant is awakened, the user can interact with the terminal by voice. In the process of voice interaction between the user and the terminal, generally after the user finishes speaking a voice, the terminal outputs the feedback result corresponding to this voice, and then the user can speak the next voice, thus achieving a continuous dialogue with the terminal.
[0004] However, the current terminal's continuous conversation feature is implemented by extending the terminal's reception time. For example, after outputting feedback corresponding to the first voice message, the terminal continues to receive the message for a period of time, such as 10 seconds. If no voice signal is received within 10 seconds, the terminal stops receiving the message. If a voice signal is received within 10 seconds, the terminal continues to output feedback for the received voice message. Therefore, during this extended reception period, if the user remains silent but others are speaking nearby, the terminal will continue to provide feedback based on the other person's speech. This can cause confusion and annoyance to the user, impacting the user experience. Summary of the Invention
[0005] The present application provides a voice interaction method, device and terminal, which can solve the problem that if the user does not make any sound during the extended reception period of the terminal, but there are other people talking around, the terminal will continue to provide feedback on what other people say, which will cause trouble and disgust to the user and affect the user experience.
[0006] In a first aspect, the present application provides a voice interaction method, which includes: detecting a wake-up indication for initiating voice interaction; entering a working state of voice interaction in response to the wake-up indication; detecting a first voice message; outputting a feedback result for the first voice message; if a second voice message is detected within a preset time length, detecting a user's breath; if a user's breath is detected, outputting a feedback result for the second voice message.
[0007] In this way, through user breath detection, it is possible to identify with a high probability that the user himself intends to continue voice interaction, effectively reducing the terminal's erroneous response to other people or other surrounding noises, and improving the accuracy of voice interaction and user experience.
[0008] In one implementable method, after outputting the feedback result for the first voice information, it also includes: determining whether the terminal is close to the user's mouth; if it is determined that the terminal is close to the user's mouth, extending the working state of the voice interaction for the preset time length; if it is determined that the terminal is not close to the user's mouth, ending the working state of the voice interaction.
[0009] In this way, before performing the user's breath detection, it is first determined whether the terminal is close to the user's mouth. If it is determined that the terminal is close to the user's mouth, the sound reception time is extended. If it is determined that the terminal is not close to the user's mouth, the sound reception is directly terminated. This can greatly reduce the energy consumption caused by sound reception.
[0010] In one implementable manner, if the second voice information is detected within a preset time period, it also includes: determining whether the terminal is close to the user's mouth; if it is determined that the terminal is close to the user's mouth, detecting the user's breath; if it is determined that the terminal is not close to the user's mouth, ending the voice interaction working state.
[0011] Thus, when the second voice information is detected, it is first determined whether the terminal is close to the user's mouth, and then whether to detect the user's breath. If the terminal is not close to the user's mouth, it is considered that the second voice information is not a sound made by the user, and there is no need to detect the user's breath.
[0012] In one implementable manner, if the wake-up indication is the user's breath, determining whether the terminal is close to the user's mouth includes: identifying the user's gesture in the working state of the voice interaction; if the user's gesture is a first gesture, determining that the terminal is close to the user's mouth, and the first gesture is used to represent that the user is holding the terminal in a stationary state; if the user's gesture is a second gesture, determining that the terminal is not close to the user's mouth, and the second gesture is used to represent that the user is holding the terminal away from the user's mouth.
[0013] In this way, if the voice interaction working state is awakened by breath awakening, it means that the terminal was near the user's mouth when waking up the terminal. Therefore, after outputting the feedback result for the first voice information, it can be determined whether the user holds the terminal 100 away from the user's mouth to determine whether the terminal 100 is still near the user's mouth.
[0014] In one implementable embodiment, if the wake-up indication is in a manner other than the user's breath, before determining whether the terminal is close to the user's mouth, it includes: determining whether a third gesture is recognized before outputting the feedback result for the first voice information, the third gesture being used to represent that the user is holding the terminal close to the user's mouth; if the third gesture is recognized, determining whether the terminal is still close to the user's mouth after outputting the feedback result for the first voice information; if the third gesture is not recognized, terminating the working state of the voice interaction.
[0015] In this way, if the working state of voice interaction is not awakened by breath awakening, it means that the terminal was not at the user's mouth when waking up the terminal. In this case, after entering the working state of voice interaction, the present application can first determine whether the user holds the terminal close to the user's mouth before outputting the feedback result for the first voice information. If it is determined that the user holds the terminal close to the user's mouth before outputting the feedback result for the first voice information, it is then determined whether the terminal is still at the user's mouth after outputting the feedback result for the first voice information.
[0016] In one implementable manner, the identifying of the user's gesture in the working state of the voice interaction includes: obtaining the angular velocity and acceleration at different times in the working state of the voice interaction; determining the user's gesture using the angular velocity, acceleration, and gesture recognition module at the different times; wherein the gesture recognition module is used to identify whether the user's handheld terminal is approaching the user's mouth, the user's handheld terminal is moving away from the user's mouth, or the user is holding the terminal in a stationary state.
[0017] In this way, the gesture recognition module can be used to determine the user's gesture based on the angular velocity and acceleration data at different times.
[0018] In one implementable manner, the detecting of the user's breath includes: inputting the second voice information into a breath recognition module, the breath recognition module being used to identify whether the second voice information is a sound emitted from the user's mouth within a preset distance from the terminal; if the breath recognition module identifies that the second voice information is a sound emitted from the user's mouth within a preset distance from the terminal, it is determined that the user's breath is detected; if the breath recognition module identifies that the second voice information is not a sound emitted from the user's mouth within the preset distance from the terminal, it is determined that the user's breath is not detected.
[0019] In this way, the breath recognition module can be used to perform feature recognition on the second voice information to determine whether the second voice information is the sound produced by the user's mouth close to the terminal.
[0020] In one implementation, the terminal includes a pressure sensor, and detecting the user's breath includes: obtaining a pressure value corresponding to the pressure sensor when the second voice information is collected; if the pressure value is greater than a preset pressure threshold, determining that the user's breath is detected; if the pressure value is less than or equal to the preset pressure threshold, determining that the user's breath is not detected;
[0021] In one implementation, the terminal includes a temperature sensor, and the detecting of the user's breath includes: obtaining a first temperature and a second temperature, wherein the first temperature is the temperature corresponding to the temperature sensor before the second voice information is collected, and the second temperature is the temperature corresponding to the temperature sensor when the second voice information is collected; if the second temperature is greater than the first temperature, it is determined that the user's breath is detected; if the second temperature is less than or equal to the first temperature, it is determined that the user's breath is not detected.
[0022] In one implementation, the terminal includes a humidity sensor, and detecting the user's breath includes: obtaining a humidity corresponding to the humidity sensor when the second voice information is collected; if the humidity is greater than a preset humidity threshold, determining that the user's breath is detected; if the humidity is less than or equal to the preset humidity threshold, determining that the user's breath is not detected;
[0023] In one implementation, the terminal includes a carbon dioxide sensor, and the detecting of the user's breath includes: obtaining the carbon dioxide concentration corresponding to the carbon dioxide sensor when the second voice information is collected; if the carbon dioxide concentration is greater than a preset carbon dioxide concentration threshold, it is determined that the user's breath is detected; if the carbon dioxide concentration is less than or equal to the preset carbon dioxide concentration threshold, it is determined that the user's breath is not detected.
[0024] In this way, if the user speaks with his mouth close to the terminal, the airflow generated by speaking will produce a certain pressure on the terminal, and the temperature, humidity, and carbon dioxide concentration near the terminal will also change to a certain extent. In this way, the present application can use pressure sensors, temperature sensors, humidity sensors or carbon dioxide sensors to detect the user's breath.
[0025] In a second aspect, the present application provides a voice interaction method, the method comprising: detecting a wake-up indication for initiating voice interaction; entering a working state of voice interaction in response to the wake-up indication; detecting a first voice message; outputting a feedback result for the first voice message; determining whether the terminal is close to the user's mouth; if it is determined that the terminal is close to the user's mouth, extending the working state of the voice interaction for a preset time length; if a second voice message is detected within the preset time length, outputting a feedback result for the second voice message.
[0026] In a third aspect, the present application provides a voice interaction method, the method comprising: detecting a wake-up indication for initiating voice interaction; entering a working state of voice interaction in response to the wake-up indication; detecting a first voice message; outputting a feedback result for the first voice message; if a second voice message is detected within a preset time length, determining whether the terminal is close to the user's mouth; if it is determined that the terminal is close to the user's mouth, outputting a feedback result for the second voice message.
[0027] In a fourth aspect, the present application provides a voice interaction device, which includes a processor; the processor is used to detect a wake-up indication for initiating voice interaction; enter a working state of voice interaction in response to the wake-up indication; detect a first voice message; output a feedback result for the first voice message; if a second voice message is detected within a preset time length, detect the user's breath; if the user's breath is detected, output a feedback result for the second voice message.
[0028] In one implementation, the processor is further used to determine whether the terminal is close to the user's mouth after outputting the feedback result for the first voice information; if it is determined that the terminal is close to the user's mouth, the working state of the voice interaction is extended by the preset time length; if it is determined that the terminal is not close to the user's mouth, the working state of the voice interaction is terminated.
[0029] In one implementation, the processor is further configured to determine whether the terminal is close to the user's mouth; if it is determined that the terminal is close to the user's mouth, the processor detects the user's breath; if it is determined that the terminal is not close to the user's mouth, the processor terminates the voice interaction working state.
[0030] In one implementation, the processor is further used to identify the user's gesture in the working state of the voice interaction; if the user's gesture is a first gesture, it is determined that the terminal is close to the user's mouth, and the first gesture is used to represent that the user is holding the terminal in a stationary state; if the user's gesture is a second gesture, it is determined that the terminal is not close to the user's mouth, and the second gesture is used to represent that the user is holding the terminal away from the user's mouth.
[0031] In one implementable embodiment, if the wake-up indication is in a manner other than the user's breath, then before determining whether the terminal is close to the user's mouth, the processor is further used to determine whether a third gesture is recognized before outputting the feedback result for the first voice information, and the third gesture is used to represent that the user holds the terminal close to the user's mouth; if the third gesture is recognized, it is determined whether the terminal is still close to the user's mouth after outputting the feedback result for the first voice information; if the third gesture is not recognized, the working state of the voice interaction is terminated.
[0032] In one implementation, the processor is further used to obtain the angular velocity and acceleration at different times in the working state of the voice interaction; and determine the user's gesture using the angular velocity, acceleration, and gesture recognition module at different times; wherein the gesture recognition module is used to identify whether the user's handheld terminal is approaching the user's mouth, the user's handheld terminal is moving away from the user's mouth, or the user is holding the terminal in a stationary state.
[0033] In one implementation, the processor is further used to input the second voice information into a breath recognition module, and the breath recognition module is used to identify whether the second voice information is a sound emitted by the user's mouth within a preset distance from the terminal; if the breath recognition module identifies that the second voice information is a sound emitted by the user's mouth within a preset distance from the terminal, it is determined that the user's breath is detected; if the breath recognition module identifies that the second voice information is not a sound emitted by the user's mouth within the preset distance from the terminal, it is determined that the user's breath is not detected.
[0034] In one implementation, the terminal includes a pressure sensor, and the processor is further used to obtain the pressure value corresponding to the pressure sensor when the second voice information is collected; if the pressure value is greater than a preset pressure threshold, it is determined that the user's breath is detected; if the pressure value is less than or equal to the preset pressure threshold, it is determined that the user's breath is not detected.
[0035] In one implementation, the terminal includes a temperature sensor, and the processor is further used to obtain a first temperature and a second temperature, wherein the first temperature is the temperature corresponding to the temperature sensor before the second voice information is collected, and the second temperature is the temperature corresponding to the temperature sensor when the second voice information is collected; if the second temperature is greater than the first temperature, it is determined that the user's breath is detected; if the second temperature is less than or equal to the first temperature, it is determined that the user's breath is not detected.
[0036] In one implementation, the terminal includes a humidity sensor, and the processor is further used to obtain the humidity corresponding to the humidity sensor when the second voice information is collected; if the humidity is greater than a preset humidity threshold, it is determined that the user's breath is detected; if the humidity is less than or equal to the preset humidity threshold, it is determined that the user's breath is not detected.
[0037] In one implementation, the terminal includes a carbon dioxide sensor, and the processor is further used to obtain the carbon dioxide concentration corresponding to the carbon dioxide sensor when the second voice information is collected; if the carbon dioxide concentration is greater than a preset carbon dioxide concentration threshold, it is determined that the user's breath is detected; if the carbon dioxide concentration is less than or equal to the preset carbon dioxide concentration threshold, it is determined that the user's breath is not detected.
[0038] In a fifth aspect, the present application provides a voice interaction device, which includes a processor; the processor is used to detect a wake-up indication for initiating voice interaction; enter a working state of voice interaction in response to the wake-up indication; detect a first voice message; output a feedback result for the first voice message; determine whether the terminal is close to the user's mouth; if it is determined that the terminal is close to the user's mouth, extend the working state of voice interaction for a preset period of time; if it is determined that the terminal is not close to the user's mouth, end the working state of voice interaction.
[0039] In a sixth aspect, the present application provides a voice interaction device, which includes a processor; the processor is used to detect a wake-up indication for initiating voice interaction; enter a working state of voice interaction in response to the wake-up indication; detect a first voice message; output a feedback result for the first voice message; if a second voice message is detected within a preset time length, determine whether the terminal is close to the user's mouth; if it is determined that the terminal is close to the user's mouth, output a feedback result for the second voice message.
[0040] In the seventh aspect, the present application provides a terminal, which includes a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device executes the method described in any one of the first to third aspects.
[0041] In an eighth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program or instructions. When the computer program or instructions are executed, the method described in any one of the first to third aspects is executed.
[0042] In a ninth aspect, the present application provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are run on a computer, the computer executes the method described in any one of the first to third aspects.
[0043] In summary, the voice interaction method, device and terminal provided in this application can detect the user's breath and / or determine whether the terminal is close to the user's mouth, and can highly likely identify that the user himself has the intention to continue voice interaction, effectively reducing the terminal's erroneous response to other people or other surrounding noises, and improving the accuracy of voice interaction and user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0045] Figure 1 An application scenario diagram of voice interaction provided in an embodiment of the present application;
[0046] Figure 2 A hardware structure block diagram of the terminal 100 provided in an embodiment of the present application;
[0047] Figure 3 A flowchart of a voice interaction method provided in an embodiment of the present application;
[0048] Figure 4 A flowchart of a first implementation method for determining whether a user intends to continue voice interaction provided in an embodiment of the present application;
[0049] Figure 5 A flowchart of a second implementation method for determining whether a user intends to continue voice interaction provided in an embodiment of the present application;
[0050] Figure 6 A flowchart of a third implementation method for determining whether a user intends to continue voice interaction provided in an embodiment of the present application;
[0051] Figure 7 A flowchart of a fourth implementation method for determining whether a user intends to continue voice interaction provided in an embodiment of the present application;
[0052] Figure 8 A flowchart of a fifth implementation method for determining whether a user intends to continue voice interaction provided in an embodiment of the present application;
[0053] Figure 9 A schematic diagram of the structure of a voice interaction device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0054] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0055] Before describing the technical solution of the present application, the application scenario of the present application is described first.
[0056] Figure 1 This is an application scenario diagram of a voice interaction provided by an embodiment of the present application. Figure 1 As shown, the application scenario diagram includes a terminal 100 and a user 200. The terminal 100 has a voice interaction function, and the user 200 can perform voice interaction with the terminal 100. At present, a specific event is required to trigger the voice interaction function of the terminal before the terminal 100 can enter the working state of voice interaction. Generally, we call the triggering of the voice interaction function of the terminal as wake-up voice interaction. The way to wake up the voice interaction can be to wake up with a wake-up word, long press the power button to wake up, click on the desktop voice assistant application, etc., which is not limited in this application.
[0057] After the voice interaction function is activated, user 200 can interact with terminal 100 through voice. During the voice interaction process, after user 200 speaks a voice message, terminal 100 generally outputs feedback corresponding to the voice message. For example, after the voice interaction function is activated, user 200 says "How's the weather today?" After receiving the voice message "How's the weather today?" from user 200, terminal 100 will recognize the voice message and output corresponding feedback, such as "The weather is sunny today" through the speaker.
[0058] Then, if the user 200 wants to continue to perform voice interaction with the terminal 100 , the user 200 can directly speak the next voice message after the terminal 100 has fed back the previous voice message, thus achieving a continuous dialogue with the terminal 100 .
[0059] In one implementation, terminal 100 implements the aforementioned continuous conversation feature by extending the audio reception time after each round of voice interaction with user 200. For example, after terminal 100 outputs feedback corresponding to the first voice message, it does not exit the audio reception but instead continues to monitor the audio for a period of time, such as 10 seconds. If no voice signal is received within 10 seconds, terminal 100 exits the audio reception. If a voice signal is received within 10 seconds, terminal 100 continues to provide feedback for the received voice message.
[0060] However, during the period when the terminal 100 extends the reception, if the user 200 does not make any sound, that is, the user 200 has no intention of continuing the conversation, and there are other people talking or other noises around, the terminal 100 will continue to provide feedback on what other people say or other noises around, which will cause trouble and disgust to the user 200 and affect the user experience.
[0061] To address the above technical issues, this application provides a voice interaction method that can effectively reduce the terminal 100's erroneous responses to other people or surrounding noise, thereby improving the accuracy of voice interaction. This application provides a voice interaction method that can be applied to the terminal 100. In the embodiments of this application, the terminal 100 can be a mobile phone, a remote control, or a smart wearable device such as a watch or a bracelet.
[0062] The following describes the hardware structure of the terminal 100 by taking the terminal 100 being a mobile phone as an example.
[0063] Figure 2 This is a hardware structure diagram of the terminal 100 provided in the embodiment of the present application. Figure 2 As shown, the terminal 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.
[0064] Among them, the above-mentioned sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, a humidity sensor 180N and a carbon dioxide sensor 180P and other sensors.
[0065] It should be understood that the structure illustrated in this embodiment does not constitute a specific limitation on terminal 100. In other embodiments, terminal 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0066] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0067] The controller may be the nerve center and command center of the terminal 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0068] Processor 110 may also be provided with a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of processor 110, and thus improves system efficiency. In some embodiments, processor 110 may include one or more interfaces.
[0069] It is understood that the interface connection relationship between the modules illustrated in this embodiment is merely an illustrative illustration and does not limit the structure of the terminal 100. In other embodiments, the terminal 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.
[0070] The charging management module 140 is used to receive charging input from a charger. The charger can be a wireless charger or a wired charger. While charging the battery 142, the charging management module 140 can also power the terminal through the power management module 141.
[0071] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to provide power to the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, and the wireless communication module 160.
[0072] The wireless communication function of the terminal 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 360, the modem processor and the baseband processor.
[0073] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0074] The mobile communication module 150 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for the terminal 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low-noise amplifier (LNA), and the like. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, filter and amplify the received electromagnetic waves, and transmit them to the modem processor for demodulation. The mobile communication module 150 can also amplify the signals modulated by the modem processor and convert them into electromagnetic waves for radiation via the antenna 1.
[0075] The wireless communication module 160 can provide wireless communication solutions applied on the terminal 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication technology (NFC), infrared technology (IR), etc.
[0076] In some embodiments, antenna 1 of terminal 100 is coupled to mobile communication module 150 , and antenna 2 is coupled to wireless communication module 160 , so that terminal 100 can communicate with the network and other devices through wireless communication technology.
[0077] Terminal 100 implements display functions through a GPU, display screen 194, and an application processor. The GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0078] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. For example, the display screen 194 can be a touch screen.
[0079] The terminal 100 can realize the shooting function through the ISP, camera 193, video codec, GPU, display screen 194 and application processor.
[0080] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.
[0081] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the terminal 100 by running the instructions stored in the internal memory 121. For example, in an embodiment of the present application, the processor 110 can execute instructions stored in the internal memory 121, and the internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area can store data created during the use of the terminal 100 (such as audio data, a phone book, etc.), etc.
[0082] The terminal 100 can implement audio functions through the audio module 370, the speaker 370A, the receiver 370B, the microphone 370C, the headphone jack 370D, and the application processor. For example, the user's voice information can be collected through the microphone 370C, and the feedback result of the user's voice information can be played through the speaker 370A.
[0083] A touch sensor, also known as a "touch panel," can be provided on display screen 194. The touch sensor and display screen 194 form a touch screen, also known as a "touch screen." The touch sensor is used to detect touch operations applied to or near the touch sensor. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided via display screen 194. In other embodiments, the touch sensor can also be provided on the surface of terminal 100, in a location different from that of display screen 194.
[0084] In the embodiment of the present application, the terminal 100 can detect a touch operation input by a user on the touch screen through a touch sensor, and collect one or more of the touch position of the touch operation on the touch screen, the touch time, etc. In some embodiments, the terminal 100 can determine the touch position of the touch operation on the touch screen by combining the touch sensor 180K and the pressure sensor 180A.
[0085] Keys 190 include a power button, a volume button, and the like. Keys 190 can be mechanical or touch-sensitive. Terminal 100 can receive key inputs and generate key signal inputs related to user settings and function control of terminal 100. For example, a long press of the power button can activate the voice interaction function.
[0086] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. For touch operations acting on different areas of the display screen 194, motor 191 can also correspond to different vibration feedback effects. Different application scenarios (for example: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.
[0087] Indicator 192 can be an indicator light that can be used to indicate charging status, battery level changes, messages, missed calls, notifications, etc. SIM card interface 195 is used to connect a SIM card. A SIM card can be connected to and disconnected from terminal 100 by inserting or removing it from SIM card interface 195. Terminal 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc.
[0088] The gyroscope sensor 180B may be a three-axis gyroscope for tracking changes in the state of the terminal 100 in six directions. The acceleration sensor 180E is used to detect the speed, direction, and displacement of the terminal 100. In the embodiment of the present application, the terminal 100 can detect the state and position of the terminal 100 using the gyroscope sensor 180B and the acceleration sensor 180E, and can determine the gesture of the user holding the terminal 100 based on the state and position of the terminal 100 at different times. For example, the user holding the terminal 100 is moving closer to the user's mouth, or the user holding the terminal 100 is moving away from the user's mouth.
[0089] The methods in the following embodiments can all be implemented in the terminal 100 having the above hardware structure.
[0090] The following is an exemplary description of the voice interaction method provided in the embodiment of the present application.
[0091] Figure 3 This is a flow chart of a voice interaction method provided in an embodiment of the present application. Figure 3 As shown, the method may include the following steps:
[0092] Step S1: detecting a wake-up indication for initiating voice interaction.
[0093] The wake-up instruction is used to wake up the terminal 100 and enter the voice interaction working state. The wake-up instruction can be a specific wake-up word said by the user to the terminal 100, the user can press and hold the power button, or the user can click on the desktop voice assistant application.
[0094] The embodiment of the present application also provides a wake-up method of breath wake-up. Breath wake-up refers to a method in which the user generates a breath (such as speaking or blowing) by facing the terminal 100 with his mouth and within a preset distance range from the terminal 100 to wake up the terminal 100 to enter the working state of voice interaction. In this way, the user can put the terminal 100 next to his mouth and speak or blow directly at the terminal 100 to wake up the terminal 100 to enter the working state of voice interaction without using a specific wake-up word or pressing a button. Correspondingly, when the terminal 100 detects the user's breath, it enters the working state of voice interaction.
[0095] In one possible implementation, the method for detecting the user's breath can be implemented as follows: microphone 170C is used to collect voice information. If voice information is collected, a breath recognition module can be used to determine whether the collected voice information is speech or breath from the user speaking to the terminal 100 through the mouth and within a preset distance from the terminal 100. The breath recognition module can be a trained neural network for breath recognition.
[0096] For example, when a user speaks at different distances from microphone 170C, different airflows will form over microphone 170C. For example, when a user speaks close to microphone 170C, and the speech includes consonants such as "b, c, d, f, j, k, l, p, q, r, s, t, v, w, x, y, z," a popping sound will be generated over microphone 170C. Thus, a breath recognition module can be obtained by training and learning the characteristics of the popping sound produced when a user speaks into microphone 170C. The breath recognition module is a trained neural network that can identify whether the input voice information is a sound input close to microphone 170C. For example, the trained neural network can accurately detect human voices within 5 centimeters of the microphone. Thus, when the breath recognition module recognizes that the input voice information is a human voice within 5 centimeters of the microphone, it determines that the user's breath has been detected, thereby waking up terminal 100 to enter the voice interaction working state.
[0097] It should be noted that when the user blows air into the terminal 100 , the microphone 170C can also collect the sound. In this application, the sound generated by blowing air is also referred to as voice information.
[0098] In one implementation, the method for detecting the user's breath can also be implemented as follows: if microphone 170C collects voice information, the pressure value detected by pressure sensor 180A when the voice information is collected is obtained. If the pressure value is greater than a preset pressure threshold, it is determined that the user's breath is detected.
[0099] When a user's mouth is facing the terminal and they are speaking or blowing air within a preset distance from the terminal 100, the airflow generated will exert a certain amount of pressure on the terminal 100. Thus, in embodiments of the present application, the pressure sensor 180A can be used to detect the pressure exerted on the terminal 100 by the user's speech. If the pressure value is greater than a preset pressure threshold, it indicates that the user's mouth is facing the terminal 100 and they are speaking or blowing air within the preset distance from the terminal 100, thus confirming that the user's breath has been detected. Conversely, if the pressure value is less than or equal to the preset pressure threshold, it indicates that the user is not speaking or blowing air within the preset distance from the terminal 100, thus confirming that the user's breath has not been detected.
[0100] It should be noted that in this embodiment of the present application, the parameters of pressure sensor 180A must meet the accuracy requirements for breath detection. For example, when a user speaks or blows air toward terminal 100 within a preset distance from terminal 100, the airflow generated on terminal 100 generates a pressure of 0.07 MPa. The range of pressure sensor 180A is 0-0.3 MPa, and the measurement accuracy is 0.001 MPa.
[0101] It should also be noted that, in order to improve the detection accuracy of the pressure sensor 180A, the pressure sensor 180A can be placed near the microphone 170C. In this way, when the user speaks near the microphone 170C, the pressure sensor 180A near the microphone 170C can detect the pressure exerted on the pressure sensor 180A by the airflow generated by the speech.
[0102] In one possible implementation, the method for detecting user breath can also be implemented as follows: if microphone 170C collects voice information, a first temperature and a second temperature are obtained, where the first temperature is the temperature collected by temperature sensor 180J before microphone 170C collects the voice information, and the second temperature is the temperature collected by temperature sensor 180J when microphone 170C collects the voice information. If the second temperature is greater than the first temperature, it is determined that the user's breath has been detected; if the second temperature is less than or equal to the first temperature, it is determined that the user's breath has not been detected.
[0103] In one possible implementation, the method for detecting user breath can also be implemented as follows: if microphone 170C collects voice information, the humidity detected by humidity sensor 180N at the time the voice information was collected is obtained. If the humidity is greater than a preset humidity threshold, it is determined that the user's breath is detected; if the humidity is less than or equal to the preset humidity threshold, it is determined that the user's breath is not detected.
[0104] In one possible implementation, the method for detecting user breath can also be implemented as follows: if microphone 170C collects voice information, the carbon dioxide concentration detected by carbon dioxide sensor 180P at the time the voice information was collected is obtained. If the carbon dioxide concentration is greater than a preset carbon dioxide concentration threshold, it is determined that the user's breath has been detected; if the carbon dioxide concentration is less than or equal to the preset carbon dioxide concentration threshold, it is determined that the user's breath has not been detected.
[0105] When a user speaks or blows with their mouth facing the terminal 100 and within a preset distance from the terminal 100, the temperature, humidity, and carbon dioxide concentration near the terminal 100 will change. Therefore, embodiments of the present application can determine whether the user's breath is detected based on data collected by the temperature sensor 180J, the humidity sensor 180N, or the carbon dioxide sensor 180P.
[0106] It should be noted that the above embodiments are merely illustrative of the implementation methods for detecting the user's breath, and do not limit the specific implementation methods for detecting the user's breath. For example, the various implementation methods listed in the above embodiments may be combined. For example, the "breath recognition module" and "pressure sensor" scheme may be combined, or the "breath recognition module" and "temperature sensor" scheme may be combined, or the "breath recognition module" and "humidity sensor" scheme may be combined, or the "breath recognition module" and "carbon dioxide sensor" scheme may be combined. It should also be noted that when the microphone is occupied by other applications, the voice interaction method provided in the embodiments of the present application is not available. For example, when a user uses the terminal 100 to make a call, even if the user's mouth is facing the terminal 100 and a breath is generated within a preset distance range from the terminal 100, the terminal 100 will not be awakened to enter the voice interaction working state.
[0107] Step S2: In response to the wake-up instruction, enter the voice interaction working state.
[0108] After entering the voice interaction working state, the terminal 100 will continue to receive the voice to obtain the user's voice information.
[0109] Step S3: Detecting the first voice information.
[0110] Step S4: outputting a feedback result for the first voice information.
[0111] In the embodiment of the present application, the feedback result for the first voice information can be voice, text, image, or entering a certain application, etc., and this application does not limit this.
[0112] For example, after entering the voice interaction working state, the user says a sentence, such as "How's the weather today?" Then, "How's the weather today?" is detected as the first voice message by the terminal 100. The terminal 100 then outputs feedback results for the first voice message, such as the terminal 100 outputting the voice message "Today's weather is sunny" through the speaker 170A. For another example, the terminal 100 may display the text "Today's weather is sunny" on the display screen 194.
[0113] For example, after entering the voice interaction working state, the user says a sentence, such as "Call Zhang San", and then "Call Zhang San" is detected as the first voice message by the terminal 100. Then, the terminal 100 outputs a feedback result for the first voice message, such as the terminal 100 enters the voice call application and dials Zhang San's phone number.
[0114] Step S5: Determine whether the user intends to continue voice interaction.
[0115] If the user intends to continue the voice interaction, the voice interaction working state is maintained; if the user does not intend to continue the voice interaction, the voice interaction working state is ended.
[0116] It should be noted that, in the voice interaction working state, the terminal 100 can continue to receive sound; after the voice interaction working state ends, the terminal stops receiving sound.
[0117] The embodiments of the present application provide the following methods for determining whether the user intends to continue voice interaction.
[0118] Figure 4 A flowchart of a first implementation method for determining whether a user intends to continue voice interaction provided in an embodiment of the present application.
[0119] like Figure 4 As shown, the first implementation method of determining whether the user intends to continue voice interaction may include the following steps:
[0120] Step S51 : determining whether the terminal 100 is close to the user's mouth.
[0121] In the embodiment of the present application, determining whether the terminal 100 is close to the user's mouth refers to determining whether the terminal 100 is near the user's mouth.
[0122] If the voice interaction working state is awakened by breath awakening, it means that the terminal 100 was at the user's mouth when it was awakened. Therefore, after outputting the feedback result for the first voice information, it is possible to determine whether the terminal 100 is still at the user's mouth by judging whether the user held the terminal 100 away from the user's mouth during the period from step S1 to step S4. If so, it is considered that the terminal 100 is no longer at the user's mouth. In this case, it can be assumed that the user has no intention to continue the voice interaction, and the voice interaction working state can be terminated; if not, it is considered that the terminal 100 is still at the user's mouth. In this case, it is considered that the user may have the intention to continue the voice interaction, and the subsequent steps can be continued.
[0123] If the working state of voice interaction is not awakened by breath awakening, it means that when the terminal 100 is awakened, the terminal 100 is not at the user's mouth. In this case, the present application can first determine whether the user holds the terminal 100 close to the user's mouth before outputting the feedback result for the first voice information after entering the working state of voice interaction. If it is determined that the user holds the terminal 100 close to the user's mouth before outputting the feedback result for the first voice information, then it is determined whether the terminal 100 is still at the user's mouth after outputting the feedback result for the first voice information (specifically, it is determined whether the user holds the terminal 100 away from the user's mouth after outputting the feedback result for the first voice information). If it is determined that the user does not hold the terminal 100 close to the user's mouth before outputting the feedback result for the first voice information, it can be considered that the user has no intention to continue the voice interaction, and the working state of voice interaction can be ended.
[0124] In one implementation, the gyroscope sensor 180C and accelerometer 180E on the terminal 100 can be used to collect the angular velocity and acceleration of the terminal 100. The collected angular velocity and acceleration are then used to determine the user's gesture. The user's gesture can include a first gesture, a second gesture, and a third gesture. The first gesture indicates that the user is holding the terminal 100 in a stationary state, the second gesture indicates that the user is holding the terminal 100 away from the user's mouth, and the third gesture indicates that the user is holding the terminal 100 toward the user's mouth.
[0125] For example, gyroscope sensor 180C and accelerometer 180E can be used to collect angular velocity and acceleration during steps S1 to S4. The collected angular velocity and acceleration are then input into a gesture recognition module. The gesture recognition module can be a trained neural network for gesture recognition. After processing, the gesture recognition module outputs the user's gesture. The gesture recognition module can determine the user's hand gesture based on the angular velocity and acceleration of the terminal 100 at different times.
[0126] It should be noted that in the embodiment of the present application, if the user's gesture changes slightly and is within a preset range of change, the user's gesture is considered to be static. For example, if the user holds the terminal 100 from the mouth for 5 cm and then changes to 4 cm from the mouth, the user's gesture is considered to be the first gesture.
[0127] Step S52: If it is determined that the terminal 100 is close to the user's mouth, the voice interaction working state is extended for a preset time.
[0128] If the terminal 100 is still near the user's mouth after outputting the feedback result for the first voice message, it is considered that the user may have the intention to continue the voice interaction. In this case, the present application extends the voice interaction working state for a preset time. During the extended preset time, the terminal 100 continues to receive the sound.
[0129] This application does not limit the preset time length, for example, it can be 5s, 10s, 20s, etc.
[0130] Step S53: Determine whether the second voice information is detected within a preset time period.
[0131] Step S54: If the second voice information is detected within the preset time period, it is determined whether the user's breath is detected.
[0132] If no second voice message is detected within the preset duration, the voice interaction state is terminated. If a second voice message is detected within the preset duration, the second voice message detected by terminal 100 may be the user's speech or someone else's speech. Therefore, the present application further detects the user's breath to determine whether the second voice message is the user's speech to terminal 100.
[0133] It should be noted that in the embodiment of the present application, when voice interaction is in operation, the user's mouth needs to be close to the terminal 100 to perform voice interaction with the terminal 100. Therefore, if the second voice information is spoken by the user, the breath produced by the user when speaking can be detected by the terminal. In other words, the present application can determine whether the second voice information is spoken by the user or by someone else nearby based on whether the user's breath can be detected.
[0134] The method for detecting the user's breath can be referred to the description of step S1 above and will not be repeated here. For example, the user's breath can be detected using a breath recognition module, a pressure sensor 180A, a temperature sensor 180J, a humidity sensor 180N, or a carbon dioxide sensor 180P.
[0135] Step S55: If the user's breath is detected, a feedback result for the second voice information is output.
[0136] If the user's breath is detected, it means that the second voice information was spoken by the user with their mouth facing the terminal 100 and within a preset distance from the terminal 100. In this case, it is assumed that the user intends to continue the voice interaction, and feedback results for the second voice information are output. If the user's breath is not detected, it means that the second voice information was spoken by someone nearby, not the user. In this case, it is assumed that the user has no intention to continue the voice interaction, and the voice interaction working state can be terminated.
[0137] In summary, in the first implementation of determining whether a user intends to continue voice interaction provided by the embodiments of the present application, after outputting feedback regarding a first voice message, the terminal 100 is first determined to be near the user's mouth. If the terminal 100 is not near the user's mouth, the voice interaction state is terminated. If it is determined that the terminal 100 is near the user's mouth, the voice interaction state is extended for a preset duration. Next, if no second voice message is detected within the preset duration, the voice interaction state is terminated. If the second voice message is detected within the preset duration, the user's breath is detected. If no breath is detected, the voice interaction state is terminated. If breath is detected, feedback regarding the second voice message is output. In other words, in the first implementation, if the terminal 100 is near the user's mouth and the user's breath can be detected, the user is determined to have intended to continue voice interaction. Thus, the voice interaction method provided by the embodiments of the present application can highly likely identify the user's intention to continue voice interaction, effectively reducing the terminal 100's erroneous responses to other people or surrounding noise, thereby improving the accuracy of voice interaction and user experience.
[0138] Figure 5 A flowchart of a second implementation method for determining whether a user intends to continue voice interaction provided in an embodiment of the present application.
[0139] like Figure 5 As shown, the second implementation method of determining whether the user intends to continue voice interaction may include the following steps:
[0140] Step S61: After outputting the feedback result for the first voice information, the voice interaction working state is extended for a preset time.
[0141] Step S62, determining whether the second voice information is detected within a preset time period;
[0142] Step S63: If the second voice information is detected within the preset time period, it is determined whether the terminal 100 is close to the user's mouth.
[0143] Step S64: If it is determined that the terminal 100 is close to the user's mouth, it is determined whether the user's breath is detected.
[0144] Step S65: If the user's breath is detected, a feedback result for the second voice information is output.
[0145] In summary, in the second implementation, after outputting the feedback result for the first voice message, the voice interaction working state is directly extended for a preset duration. If the second voice message is not detected within the preset duration, the voice interaction working state is terminated. If the second voice message is detected within the preset duration, it is first determined whether the terminal 100 is near the user's mouth. If the terminal 100 is not near the user's mouth, the voice interaction working state is terminated. If it is determined that the terminal 100 is near the user's mouth, the user's breath is detected. If the user's breath is not detected, the voice interaction working state is terminated. If the user's breath is detected, the feedback result for the second voice message is output.
[0146] It should be noted that the specific implementation method of determining whether the terminal 100 is near the user's mouth in the above step S63 can be found in the description of step S51, the specific implementation method of detecting the user's breath in step S64 can be found in the description of step S54, and the specific implementation method of step S65 can be found in the description of step S55, which will not be repeated here.
[0147] Figure 6 A flowchart of a third implementation method for determining whether a user intends to continue voice interaction provided in an embodiment of the present application.
[0148] like Figure 6 As shown, a third implementation method for determining whether the user intends to continue voice interaction may include the following steps:
[0149] Step S71: After outputting the feedback result for the first voice information, the voice interaction working state is extended for a preset time.
[0150] Step S72, determining whether the second voice information is detected within a preset time period;
[0151] Step S73: If the second voice information is detected within the preset time period, it is determined whether the user's breath is detected.
[0152] Step S74: If the user's breath is detected, a feedback result for the second voice information is output.
[0153] In summary, in the third implementation, after outputting feedback regarding the first voice message, the voice interaction state is directly extended for a preset duration. If no second voice message is detected within the preset duration, the voice interaction state is terminated. If the second voice message is detected within the preset duration, the user's breath is detected. If no user's breath is detected, the voice interaction state is terminated. If the user's breath is detected, feedback regarding the second voice message is output.
[0154] It should be noted that the specific implementation method of detecting the user's breath in the above step S73 can be found in the description of step S54, and the specific implementation method of step S74 can be found in the description of step S55, which will not be repeated here.
[0155] Figure 7 A flowchart of the fourth implementation method of determining whether the user intends to continue voice interaction provided in an embodiment of the present application.
[0156] like Figure 7 As shown, the fourth implementation method of determining whether the user intends to continue voice interaction may include the following steps:
[0157] Step S81: After outputting the feedback result for the first voice information, the voice interaction working state is extended for a preset time.
[0158] Step S82, determining whether the second voice information is detected within a preset time period;
[0159] Step S83: If the second voice information is detected within the preset time period, it is determined whether the terminal 100 is close to the user's mouth.
[0160] Step S84: If it is determined that the terminal 100 is close to the user's mouth, a feedback result for the second voice information is output.
[0161] In summary, in the fourth implementation, after outputting the feedback result for the first voice message, the voice interaction working state is directly extended for a preset duration. If the second voice message is not detected within the preset duration, the voice interaction working state is terminated. If the second voice message is detected within the preset duration, it is determined whether the terminal 100 is near the user's mouth. If the terminal 100 is not near the user's mouth, the voice interaction working state is terminated. If it is determined that the terminal 100 is near the user's mouth, the feedback result for the second voice message is output.
[0162] It should be noted that the specific implementation method of determining whether the terminal 100 is near the user's mouth in the above step S83 can be found in the description of step S51, and the specific implementation method of step S84 can be found in the description of step S55, which will not be repeated here.
[0163] Figure 8 A flowchart of the fifth implementation method for determining whether the user intends to continue voice interaction provided in an embodiment of the present application.
[0164] like Figure 8 As shown, a fifth implementation method for determining whether the user intends to continue voice interaction may include the following steps:
[0165] Step S91: Determine whether the terminal 100 is close to the user's mouth.
[0166] Step S92: If it is determined that the terminal 100 is close to the user's mouth, the working state of the voice interaction is extended for a preset time.
[0167] Step S93: Determine whether the second voice information is detected within a preset time period.
[0168] Step S94: If the second voice information is detected within the preset time period, a feedback result for the second voice information is output.
[0169] In summary, in the fifth implementation, it is first determined whether the terminal 100 is at the user's mouth. If it is determined that the terminal 100 is at the user's mouth, the working state of the voice interaction is extended by a preset time length; if it is determined that the terminal 100 is not at the user's mouth, the working state of the voice interaction is terminated. In this way, the energy consumption of the terminal 100 can be reduced. Furthermore, if a second voice message is detected within the preset time length, a feedback result for the second voice message is output; if no second voice message is detected within the preset time length, the voice interaction working state is terminated. In the fifth implementation, after outputting the feedback result for the first voice message, if the terminal 100 is still at the user's mouth, it is considered that the user has the intention to continue the voice interaction, so that the reception time can be extended.
[0170] Furthermore, to improve recognition of the user's intention to continue voice interaction, after detecting the second voice message within a preset time period, the user's breath can be detected. If the user's breath is detected, feedback results for the second voice message can be output. For details, please refer to the first implementation method above and will not be repeated here.
[0171] In summary, the voice interaction method provided in the embodiment of the present application can recognize with a high probability that the user himself intends to continue the voice interaction, effectively reduce the terminal 100's erroneous response to other people or other surrounding noises, and improve the accuracy of voice interaction and user experience.
[0172] The various method embodiments described herein may be independent solutions or may be combined according to internal logic, and all of these solutions fall within the scope of protection of this application.
[0173] It can be understood that, in the above-mentioned various method embodiments, the methods and operations implemented by the electronic device can also be implemented by components (such as chips or circuits) that can be used in the electronic device.
[0174] The above embodiments introduce the voice interaction method provided by the present application. It is understandable that, in order to implement the above functions, the terminal includes a hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0175] The above describes in detail the method provided by the embodiment of the present application. Figure 9 The apparatus provided in the embodiments of the present application will be described in detail. It should be understood that the description of the apparatus embodiment corresponds to the description of the method embodiment, and therefore, for matters not described in detail, reference can be made to the method embodiment above, and for the sake of brevity, no further description will be given here.
[0176] Figure 9 This is a schematic diagram of the structure of a voice interaction device provided by an embodiment of the present application. In one embodiment, the terminal can Figure 9 The hardware devices shown in the figure implement the corresponding functions. Figure 9 As shown, the device 1000 may include: a processor 1001 and a memory 1002. The processor 1001 may include one or more processing units, for example: the processor 1001 may include an application processor, a modem processor, a graphics processor, an image signal processor, a controller, a video codec, a digital signal processor, a baseband processor, and / or a neural network processor. The different processing units may be independent devices or integrated into one or more processors. The memory 1002 is coupled to the processor 1001 and is used to store various software programs and / or multiple sets of instructions. The memory 1002 may include a volatile memory and / or a non-volatile memory.
[0177] The apparatus 1000 can execute the operations performed in the above method embodiments.
[0178] For example, in an optional embodiment of the present application, the processor 1001 can be used to detect a wake-up indication for initiating voice interaction; enter a working state of voice interaction in response to the wake-up indication; detect a first voice message; output a feedback result for the first voice message; if a second voice message is detected within a preset time length, detect the user's breath; if the user's breath is detected, output a feedback result for the second voice message.
[0179] In one implementation, the processor is further used to determine whether the terminal is close to the user's mouth after outputting the feedback result for the first voice information; if it is determined that the terminal is close to the user's mouth, the working state of the voice interaction is extended by the preset time length; if it is determined that the terminal is not close to the user's mouth, the working state of the voice interaction is terminated.
[0180] In one implementation, the processor is further configured to determine whether the terminal is close to the user's mouth; if it is determined that the terminal is close to the user's mouth, the processor detects the user's breath; if it is determined that the terminal is not close to the user's mouth, the processor terminates the voice interaction working state.
[0181] In one implementation, the processor is further used to identify the user's gesture in the working state of the voice interaction; if the user's gesture is a first gesture, it is determined that the terminal is close to the user's mouth, and the first gesture is used to represent that the user is holding the terminal in a stationary state; if the user's gesture is a second gesture, it is determined that the terminal is not close to the user's mouth, and the second gesture is used to represent that the user is holding the terminal away from the user's mouth.
[0182] In one implementable embodiment, if the wake-up indication is in a manner other than the user's breath, then before determining whether the terminal is close to the user's mouth, the processor is further used to determine whether a third gesture is recognized before outputting the feedback result for the first voice information, and the third gesture is used to represent that the user holds the terminal close to the user's mouth; if the third gesture is recognized, it is determined whether the terminal is still close to the user's mouth after outputting the feedback result for the first voice information; if the third gesture is not recognized, the working state of the voice interaction is terminated.
[0183] In one implementation, the processor is further used to obtain the angular velocity and acceleration at different times in the working state of the voice interaction; and determine the user's gesture using the angular velocity, acceleration, and gesture recognition module at different times; wherein the gesture recognition module is used to identify whether the user's handheld terminal is approaching the user's mouth, the user's handheld terminal is moving away from the user's mouth, or the user is holding the terminal in a stationary state.
[0184] In one implementation, the processor is further used to input the second voice information into a breath recognition module, and the breath recognition module is used to identify whether the second voice information is a sound emitted by the user's mouth within a preset distance from the terminal; if the breath recognition module identifies that the second voice information is a sound emitted by the user's mouth within a preset distance from the terminal, it is determined that the user's breath is detected; if the breath recognition module identifies that the second voice information is not a sound emitted by the user's mouth within the preset distance from the terminal, it is determined that the user's breath is not detected.
[0185] In one implementation, the terminal includes a pressure sensor, and the processor is further used to obtain the pressure value corresponding to the pressure sensor when the second voice information is collected; if the pressure value is greater than a preset pressure threshold, it is determined that the user's breath is detected; if the pressure value is less than or equal to the preset pressure threshold, it is determined that the user's breath is not detected.
[0186] In one implementation, the terminal includes a temperature sensor, and the processor is further used to obtain a first temperature and a second temperature, wherein the first temperature is the temperature corresponding to the temperature sensor before the second voice information is collected, and the second temperature is the temperature corresponding to the temperature sensor when the second voice information is collected; if the second temperature is greater than the first temperature, it is determined that the user's breath is detected; if the second temperature is less than or equal to the first temperature, it is determined that the user's breath is not detected.
[0187] In one implementation, the terminal includes a humidity sensor, and the processor is further used to obtain the humidity corresponding to the humidity sensor when the second voice information is collected; if the humidity is greater than a preset humidity threshold, it is determined that the user's breath is detected; if the humidity is less than or equal to the preset humidity threshold, it is determined that the user's breath is not detected.
[0188] In one implementation, the terminal includes a carbon dioxide sensor, and the processor is further used to obtain the carbon dioxide concentration corresponding to the carbon dioxide sensor when the second voice information is collected; if the carbon dioxide concentration is greater than a preset carbon dioxide concentration threshold, it is determined that the user's breath is detected; if the carbon dioxide concentration is less than or equal to the preset carbon dioxide concentration threshold, it is determined that the user's breath is not detected.
[0189] For another example, in an optional embodiment of the present application, the processor 1001 can be used to detect a wake-up indication for initiating voice interaction; enter a working state of voice interaction in response to the wake-up indication; detect a first voice message; output a feedback result for the first voice message; determine whether the terminal is close to the user's mouth; if it is determined that the terminal is close to the user's mouth, extend the working state of voice interaction for a preset time; if it is determined that the terminal is not close to the user's mouth, end the working state of voice interaction.
[0190] For another example, in an optional embodiment of the present application, the processor 1001 can be used to detect a wake-up indication for initiating voice interaction; enter a working state of voice interaction in response to the wake-up indication; detect a first voice message; output a feedback result for the first voice message; if a second voice message is detected within a preset time length, determine whether the terminal is close to the user's mouth; if it is determined that the terminal is close to the user's mouth, output a feedback result for the second voice message.
[0191] During implementation, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor or by instructions in the form of software. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it will not be described in detail here.
[0192] It should be noted that the processor in the embodiments of the present application can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiment can be completed by an integrated logic circuit of the hardware in the processor or by instructions in the form of software. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The various methods, steps, and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0193] It is understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0194] According to the method provided in the embodiment of the present application, the embodiment of the present application also provides a computer program product, which includes: a computer program or instructions, which, when the computer program or instructions are run on a computer, enables the computer to execute the method of any one of the method embodiments.
[0195] According to the method provided in the embodiments of the present application, the embodiments of the present application also provide a computer-readable storage medium, which stores a computer program or instructions. When the computer program or instructions are run on a computer, the computer executes the method of any one of the embodiments of the method.
[0196] According to the method provided in an embodiment of the present application, an embodiment of the present application also provides a terminal, including a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device executes the method of any one of the embodiments of the method embodiment.
[0197] Those skilled in the art will appreciate that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented using electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0198] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0199] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0200] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network elements. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0201] In addition, the functional modules in the various embodiments of the present application may be integrated into one processing unit, or each module may exist physically separately, or two or more modules may be integrated into one unit.
[0202] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0203] The voice interaction device, chip, computer storage medium, computer program product, and terminal provided in the above-mentioned embodiments of the present application are all used to execute the method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects corresponding to the method provided above, and will not be repeated here.
[0204] It should be understood that in each embodiment of the present application, the execution order of each step should be determined by its function and internal logic. The size of the sequence number of each step does not mean the order of execution and does not limit the implementation process of the embodiment.
[0205] The various sections of this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other. Each embodiment focuses on the differences between the other embodiments. In particular, the embodiments of the voice interaction device, chip, computer storage medium, computer program product, and terminal are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant details, refer to the descriptions of the method embodiments.
[0206] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0207] The above-described embodiments of the present application do not constitute a limitation on the scope of protection of the present application.
Claims
1. A voice interaction method, characterized in that: The method comprises: Detecting a wake-up indication for initiating voice interaction, wherein the wake-up indication is a user's breath; In response to the wake-up instruction, entering a working state of voice interaction; detecting first voice information; Outputting a feedback result for the first voice information; Recognizing the user's gestures during the voice interaction operation; If the user's gesture is a first gesture, determining that the terminal is close to the user's mouth, the first gesture being used to indicate that the user is holding the terminal in a stationary state; If it is determined that the terminal is close to the user's mouth, extending the working state of the voice interaction for a preset time; If a second voice message is detected within a preset time period, determining whether the user's breath is detected based on at least one of a popping sound feature corresponding to the second voice message, a pressure value corresponding to the second voice message, a temperature difference value corresponding to the second voice message, a humidity value corresponding to the second voice message, and a carbon dioxide concentration value corresponding to the second voice message; If the user's breath is detected, a feedback result for the second voice information is output.
2. The method according to claim 1, characterized in that After the recognition of the user's gesture in the voice interaction working state, the method further includes: If the user's gesture is a second gesture, it is determined that the terminal is not close to the user's mouth, and the second gesture is used to represent that the user holds the terminal away from the user's mouth; If it is determined that the terminal is not close to the user's mouth, the working state of the voice interaction is terminated.
3. The method according to claim 2, characterized in that The identifying, in the working state of the voice interaction, the user's gesture includes: Obtaining angular velocity and acceleration at different times in the working state of the voice interaction; The user's gesture is determined by using the angular velocity, acceleration, and gesture recognition module at the different moments; wherein the gesture recognition module is used to identify whether the user is holding the terminal close to the user's mouth, the user is holding the terminal away from the user's mouth, or the user is holding the terminal in a stationary state.
4. The method according to claim 2, characterized in that If the second voice information is detected within the preset time period, determining whether the user's breath is detected based on at least one of a popping sound feature corresponding to the second voice information, a pressure value corresponding to the second voice information, a temperature difference value corresponding to the second voice information, a humidity value corresponding to the second voice information, and a carbon dioxide concentration value corresponding to the second voice information includes: detecting a plosive feature in the second voice information by a breath recognition module; Based on the pop feature, it is identified whether the second voice information is a sound emitted by the user's mouth within a preset distance from the terminal to determine whether the user's breath is detected.
5. The method according to claim 4, characterized in that The terminal includes a pressure sensor, and if the second voice information is detected within a preset time period, determining whether the user's breath is detected based on at least one of a popping sound feature corresponding to the second voice information, a pressure value corresponding to the second voice information, a temperature difference value corresponding to the second voice information, a humidity value corresponding to the second voice information, and a carbon dioxide concentration value corresponding to the second voice information includes: Acquiring the pressure value corresponding to the pressure sensor when the second voice information is collected; Determine whether the pressure value is greater than a preset pressure threshold to determine whether the user's breath is detected.
6. The method according to claim 5, characterized in that The terminal includes a temperature sensor, and if the second voice information is detected within a preset time period, determining whether the user's breath is detected based on at least one of a popping sound feature corresponding to the second voice information, a pressure value corresponding to the second voice information, a temperature difference value corresponding to the second voice information, a humidity value corresponding to the second voice information, and a carbon dioxide concentration value corresponding to the second voice information includes: Acquire a first temperature and a second temperature, wherein the first temperature is the temperature corresponding to the temperature sensor before the second voice information is collected, and the second temperature is the temperature corresponding to the temperature sensor when the second voice information is collected; It is determined whether the temperature difference between the second temperature and the first temperature is greater than 0 to determine whether the user's breath is detected.
7. The method according to claim 6, characterized in that The terminal includes a humidity sensor, and if second voice information is detected within a preset time period, determining whether the user's breath is detected based on at least one of a popping sound feature corresponding to the second voice information, a pressure value corresponding to the second voice information, a temperature difference value corresponding to the second voice information, a humidity value corresponding to the second voice information, and a carbon dioxide concentration value corresponding to the second voice information, includes: Acquiring a humidity value corresponding to the humidity sensor when the second voice information is collected; Determine whether the humidity value is greater than a preset humidity threshold to determine whether the user's breath is detected.
8. The method according to claim 7, characterized in that The terminal includes a carbon dioxide sensor, and if second voice information is detected within a preset time period, determining whether the user's breath is detected based on at least one of a popping sound feature corresponding to the second voice information, a pressure value corresponding to the second voice information, a temperature difference value corresponding to the second voice information, a humidity value corresponding to the second voice information, and a carbon dioxide concentration value corresponding to the second voice information, includes: Acquiring a carbon dioxide concentration value corresponding to the carbon dioxide sensor when the second voice information is collected; Determine whether the carbon dioxide concentration value is greater than a preset carbon dioxide concentration threshold to determine whether the user's breath is detected.
9. A voice interaction device, characterized in that: The apparatus includes a processor; The processor is configured to detect a wake-up indication for initiating voice interaction, the wake-up indication being a user's breath; enter a voice interaction working state in response to the wake-up indication; detect a first voice message; output a feedback result for the first voice message; recognize a gesture of the user in the voice interaction working state; and determine that the terminal is close to the user's mouth if the gesture of the user is a first gesture, the first gesture being used to indicate that the user is holding the terminal in a stationary state; If it is determined that the terminal is close to the user's mouth, extending the working state of the voice interaction for a preset time; If a second voice message is detected within a preset time period, determining whether the user's breath is detected based on at least one of a popping sound feature corresponding to the second voice message, a pressure value corresponding to the second voice message, a temperature difference value corresponding to the second voice message, a humidity value corresponding to the second voice message, and a carbon dioxide concentration value corresponding to the second voice message; If the user's breath is detected, a feedback result for the second voice information is output.
10. A terminal, characterized in that: The terminal includes a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the terminal executes the method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program or instructions. When the computer program or instructions are executed, the method according to any one of claims 1 to 8 is performed.
Citation Information
Patent Citations
Method and device for searching
CN105512125A
Continuous wake-up method and device for intelligent equipment, intelligent equipment and storage medium
CN108733419A
Voice interaction awakening electronic device based on microphone signal, method and medium
CN110097875A
Voice interaction method, device and system
CN111402900A
Method for waking up application program, and electronic device
WO2024103893A1