Voice interaction method and electronic device
By combining audio data and multiple sensing data, the voice interaction function is automatically triggered, which solves the problem of users in the prior art that repeated wake-up operations are needed, and improves the accuracy and user experience of voice interaction.
Patent Information
- Application Number
- CN202310334681.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-27
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2043-03-27
AI Technical Summary
Existing voice interaction applications require users to repeatedly perform wake-up operations when performing voice interaction, resulting in a reduced user experience, especially when frequently used, which is cumbersome and wastes time.
The voice interaction function is automatically triggered by combining audio data and multiple sensing data. The specific steps include acquiring audio data and the first sensing data (characterizing the device movement), detecting a voice signal in the audio data and determining based on the first sensing data that the degree of position change is greater than the threshold, acquiring the sound source direction and the second sensing data (characterizing the user distance change), and automatically starting the voice interactive application when a specific condition is met.
It improves the accuracy of voice interaction recognition, avoids misidentification, reduces user operation time, and improves user experience.
Smart Images

Figure CN117133282B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of voice control technology, and in particular to a voice interaction method and electronic device. Background Art
[0002] With the development of computer technology, voice recognition technology has become increasingly mature, and voice input has become increasingly important due to its high naturalness and effectiveness in interaction. Voice interaction applications in electronic devices have also gradually become a function that people often use. Users can interact with electronic devices (mobile phones, tablets, smart watches, etc.) through voice to complete various operations such as command input, information query, and voice chat.
[0003] Usually, voice interaction applications need to be woken up by users when performing voice interaction. However, the wake-up methods of voice interaction applications mainly include: waking up the voice interaction application by inputting a specific voice wake-up word; or clicking a physical button or screen on an electronic device. In this way, the user must repeat the above wake-up operation every time before performing voice interaction with the voice interaction application. When the user needs to use the voice interaction application frequently, it is cumbersome to repeatedly perform the wake-up operation, which wastes a lot of usage time. This reduces the user experience. Summary of the invention
[0004] In view of this, the present application provides a voice interaction method and an electronic device, which can automatically trigger the voice interaction function according to audio data and different sensor data, thereby saving the user's operation time.
[0005] In a first aspect, the present application provides a voice interaction method, which is applied to an electronic device, and the method includes: the electronic device obtains audio data and first sensor data; the first sensor data is used to characterize the movement of the electronic device. When the electronic device includes a voice signal in the audio data, and determines that the position change of the electronic device is greater than a first preset threshold according to the first sensor data, the sound source direction and second sensor data corresponding to the audio data are obtained; the second sensor data is used to characterize the change in the distance between the electronic device and the user. Afterwards, when the electronic device determines that the electronic device is close to the user according to the sound source direction and the second sensor data, the voice interaction application is started. And, through the voice interaction application, the response operation corresponding to the voice signal is executed.
[0006] It can be seen that the electronic device of the present application includes a voice signal in the audio data, and when it is determined according to the first sensor data that the position change degree of the electronic device is greater than the first preset threshold, the sound source direction and the second sensor data corresponding to the audio data are obtained. And the sound source direction and the second sensor data corresponding to the audio data are judged, and the voice interaction application is automatically started after the sound source direction and the second sensor data corresponding to the audio data meet certain conditions. In other words, the electronic device will only start the voice interaction application after detecting the audio device, the first sensor data and the second sensor data. Among them, the first sensor data and the second sensor data respectively characterize the position of the electronic device and the change in the distance between the electronic device and the user. Compared with the use of sensor data of a single sensor for voice interaction recognition, the accuracy of voice interaction recognition is effectively improved to avoid the occurrence of misrecognition. In addition, the user does not need to repeat a specific wake-up operation for the mobile phone. In addition, the user's sense of use is improved.
[0007] In an implementation method of the first aspect, in the process of determining that the degree of position change of the electronic device is greater than a first preset threshold value based on first sensing data, the electronic device can determine a position characteristic value based on the first sensing data, where the position characteristic value is used to characterize the degree of position change of the electronic device. Thereafter, the electronic device determines that the degree of position change represented by the position characteristic value is greater than the first preset threshold value.
[0008] It can be seen that the electronic device of the present application, in the process of determining that the degree of position change of the electronic device is greater than the first preset threshold value according to the first sensing data, can determine the position characteristic value according to the first sensing data, and then determine that the degree of position change represented by the position characteristic value is greater than the first preset threshold value. In other words, the electronic device can extract the position characteristic value of the first sensing data, and the position characteristic value can represent the degree of position change of the electronic device. When the degree of position change represented by the position characteristic value is greater than the first preset threshold value, it can indicate that the position of the electronic device has changed significantly. In this way, when the position of the electronic device has changed significantly, it can indicate that the user has moved the electronic device. Furthermore, the subsequent voice interaction recognition operation can be triggered. By determining the position of the electronic device, the accuracy of voice interaction recognition is improved, and misrecognition is avoided. Furthermore, the user's usage experience is improved.
[0009] In an implementation of the first aspect, the electronic device includes a display screen, and the electronic device acquires third sensor data in the process of acquiring the sound source direction and the second sensor data corresponding to the audio data, and the third sensor data is used to characterize the posture of the electronic device. Afterwards, when the electronic device determines that the display screen is facing the user according to the third sensor data, it acquires the sound source direction and the second sensor data corresponding to the audio data.
[0010] It can be seen that the electronic device of the present application obtains the third sensor data in the process of obtaining the sound source direction and the second sensor data corresponding to the audio data, and the third sensor data is used to characterize the posture of the electronic device. In other words, before the electronic device obtains the sound source direction and the second sensor data corresponding to the audio data, it will also determine the posture of the display screen according to the third sensor data. And the subsequent voice interaction recognition operation can be performed only when the display screen is determined to be facing the user according to the third sensor data. This avoids the situation where the electronic device display screen misidentifies the voice interaction when it is facing away from the user. It effectively improves the accuracy of voice interaction recognition and improves the user experience.
[0011] In an implementation method of the first aspect, when the electronic device determines that the electronic device is close to the user based on the sound source direction and the second sensor data, during the process of starting the voice interaction application, the target distance characteristic value can be determined based on the sound source angle corresponding to the sound source direction and the first distance characteristic value and the second distance characteristic value corresponding to the second sensor data.
[0012] The first distance feature value is used to characterize the proximity between the electronic device and the user's face, and the second distance feature value is used to characterize the proximity between the electronic device and the sound source; the sound source angle is the angle between the sound source direction and the reference direction, the sound source direction is the direction from the sound receiving position of the electronic device to the sound source, and the reference direction is the direction passing through the sound receiving position of the electronic device, perpendicular to the display screen of the electronic device and extending upward;
[0013] Afterwards, when the electronic device determines that the electronic device is close to the user based on the target distance characteristic value, the voice interaction application is started.
[0014] It can be seen that the electronic device can determine the target distance characteristic value based on the sound source angle, the first distance characteristic value characterizing the proximity between the electronic device and the user's face, and the second distance characteristic value characterizing the proximity between the electronic device and the sound source. Afterwards, when it is determined that the electronic device is close to the user according to the target distance characteristic value, the voice interaction application is started. In other words, the electronic device can characterize the proximity between the electronic device and the user according to the target distance characteristic value, determine that the electronic device is close to the user and automatically start the voice interaction application. In this way, the electronic device can make a comprehensive judgment through the audio data and the second sensor data, and automatically start the voice interaction application after determining that the electronic device is close to the user. Avoid the situation where the information expressed by a single sensor data is relatively limited and misrecognition occurs. In addition, the accuracy of voice interaction recognition is effectively improved, and the user's experience is improved.
[0015] In an implementation manner of the first aspect, in a process of determining the target distance characteristic value according to the sound source angle corresponding to the sound source direction and the first distance characteristic value and the second distance characteristic value corresponding to the second sensing data, the electronic device may determine the first distance characteristic value as the target distance characteristic value when the sound source angle is less than an angle threshold;
[0016] Alternatively, when the sound source angle is greater than or equal to the angle threshold, the second distance feature value is determined as the target distance feature value.
[0017] It can be seen that the present application sets weights for the first distance eigenvalue and the second distance eigenvalue according to the sound source angle corresponding to the sound source direction, and determines the first distance eigenvalue or the second distance eigenvalue as the target distance eigenvalue. In other words, the smaller the sound source angle, the closer the user's face is to the screen of the mobile phone. On the contrary, the larger the sound source angle, the farther the user's face is from the screen of the mobile phone. In this way, the electronic device can accurately determine whether the target distance eigenvalue is the first distance eigenvalue or the second distance eigenvalue based on the sound source angle. Furthermore, the target distance eigenvalue that meets the current user scenario is selected to facilitate the subsequent determination that the electronic device is close to the user based on the target distance eigenvalue. The accuracy of voice interaction recognition is effectively improved to avoid misrecognition.
[0018] In an implementation manner of the first aspect, the method further includes:
[0019] The electronic device can extract the voice airflow feature corresponding to the audio data, and the voice airflow feature is used to characterize the wind noise generated by the voice airflow hitting the microphone. In addition, in the process of determining that the electronic device is close to the user according to the sound source angle corresponding to the sound source direction and the second sensing data, the electronic device determines that the electronic device is close to the user according to the voice airflow feature, the target distance feature value and the first sensing data.
[0020] Specifically, when the sum of the airflow characteristic value, the target distance characteristic value, and the position characteristic value of the voice airflow characteristic is greater than a second preset threshold, it is determined that the electronic device is close to the user;
[0021] The airflow characteristic value is used to characterize the characteristic prominence of the speech airflow feature in the audio data, and the airflow characteristic value is positively correlated with the characteristic prominence of the speech airflow feature.
[0022] It can be seen that the electronic device of the present application can also extract the voice airflow feature corresponding to the audio data. And when the sum of the airflow feature value, the target distance feature value and the position feature value of the voice airflow feature is greater than the second preset threshold value, it is determined that the electronic device is close to the user. That is to say, the electronic device can also determine that the electronic device is close to the user based on the voice airflow feature corresponding to the audio data, the second sensor data and the first sensor data. In other words, when the airflow feature value of the voice airflow feature is larger, it means that the feature of the voice airflow feature in the audio data is more obvious, that is, it indicates that the wind noise generated by the voice airflow hitting the microphone in the audio data is very obvious. The larger the target distance feature value, the closer the electronic device is to the user. The larger the position feature value, the greater the change in the position of the electronic device. Therefore, by setting weights for the above-mentioned airflow feature value, target distance feature value and position feature value, it can be comprehensively determined whether the electronic device is close to the user. It effectively improves the accuracy of voice interaction recognition and improves the user's experience.
[0023] In an implementation method of the first aspect, the electronic device may also obtain the sound pressure difference of different microphones for audio data, and the sound pressure difference is used to characterize the sound pressure change generated by different microphones at the sound source distance. And in the process of obtaining the sound source direction and the second sensor data corresponding to the audio data when it is determined according to the first sensor data that the position change of the electronic device is greater than the first preset threshold, when the airflow characteristic value of the voice airflow characteristic is less than or equal to the third preset threshold or the sound pressure difference is less than or equal to the fourth preset threshold, when it is determined according to the first sensor data that the position change of the electronic device is greater than the first preset threshold, the sound source direction and the second sensor data corresponding to the audio data are obtained.
[0024] It can be seen that the electronic device of the present application can also obtain the sound pressure difference of different microphones for audio data. When the sound pressure difference is larger, it means that the sound pressure change generated by the sound source from different microphones is larger, that is, it means that the sound source is closer to at least one of the different microphones. When the airflow characteristic value is larger, it means that the characteristic of the voice airflow feature in the audio data is more obvious, that is, it means that the wind noise generated by the voice airflow hitting the microphone in the audio data is very obvious. Therefore, the present application can first determine the voice airflow feature and the sound pressure difference in the audio data, and when the airflow characteristic value is less than or equal to the third preset threshold or the sound pressure difference is less than or equal to the fourth preset threshold, trigger the subsequent process of determining according to the sound source direction and the second sensor data corresponding to the audio data. In other words, the present application can not only determine the voice airflow feature and the sound pressure difference in the audio data, but also make a comprehensive determination on the first sensor data, the sound source direction in the audio data, and the second sensor data. It effectively improves the accuracy of voice interaction recognition and avoids the occurrence of misidentification. In turn, it improves the user experience.
[0025] In an implementation manner of the first aspect, the electronic device may further start a voice interaction application when the airflow characteristic value is greater than a third preset threshold and the sound pressure difference value is greater than a fourth preset threshold.
[0026] It can be seen that the electronic device of the present application can also start the voice interaction application when the airflow characteristic value is greater than the third preset threshold and the sound pressure difference value is greater than the fourth preset threshold. That is to say, the electronic device of the present application can directly use audio data to start the voice interaction application. Specifically, the electronic device starts the voice interaction application when the airflow characteristic value is greater than the third preset threshold and the sound pressure difference value is greater than the fourth preset threshold. Among them, the airflow characteristic value represents the characteristic obviousness of the voice airflow characteristics in the audio data. The sound pressure difference value is used to characterize the sound pressure change caused by different microphones at the sound source distance. When the airflow characteristic value is larger, it means that the characteristic obviousness of the voice airflow characteristics in the audio data is greater, that is, it indicates that the wind noise caused by the voice airflow hitting the microphone in the audio data is obvious. When the sound pressure difference value is larger, it means that the sound pressure change caused by different microphones at the sound source distance is greater, that is, it indicates that at least one of the different microphones is closer to the sound source. Therefore, the electronic device can determine that the electronic device is close to the user and the sound source corresponding to the audio data is the user according to the airflow characteristic value and the sound pressure difference value in the collected audio data meet certain conditions, and then automatically start the voice interaction application. Effectively improve the accuracy of voice interaction recognition and avoid the occurrence of misrecognition. Then, improve the user's experience.
[0027] In an implementation of the first aspect, when the electronic device executes a response operation corresponding to a voice signal through a voice interaction application, the electronic device may provide a voice prompt of a response result, where the response result is generated by the server after semantic understanding of the voice signal. And / or display a target interface, where the target interface includes the response result.
[0028] It can be seen that the present application can use a voice interaction application to give a voice prompt of the answer result. The answer result is generated after the server performs semantic understanding of the voice signal. Alternatively, the target interface can be displayed through the voice interaction application, and the target interface includes the answer result. Thus, the electronic device starts the voice interaction application and executes the voice interaction function, so that the user can hear and / or see the required answer result. This improves the interactivity between the user and the electronic device and improves the user's experience.
[0029] In an implementation manner of the first aspect, the electronic device may further prompt first information before displaying the target interface, where the first information is used to prompt that the voice interaction application has been running.
[0030] It can be seen that the present application can also prompt the first information before displaying the target interface, and the first information is used to prompt that the voice interaction application has been run. During the process of the electronic device starting the voice interaction application, the user can perceive that the electronic device has started and run the voice interaction application through the first information, thereby improving the interactivity between the user and the electronic device and improving the user's experience.
[0031] In a second aspect, the present application provides an electronic device, which includes a memory and one or more processors; the memory is coupled to the processor; wherein computer program code is stored in the memory, and the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device executes the voice interaction method described in the first aspect above.
[0032] In a third aspect, the present application provides a computer-readable storage medium, in which instructions are stored. When the instructions are executed on a computer, the computer can execute the voice interaction method as described in the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 A schematic diagram of a scenario for waking up a voice interaction application provided in an embodiment of the present application;
[0034] Figure 2 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;
[0035] Figure 3 A schematic diagram of the software structure of an electronic device provided in an embodiment of the present application;
[0036] Figure 4 A flowchart of a voice interaction method provided in an embodiment of the present application;
[0037] Figure 5 A schematic diagram of a data processing flow provided in an embodiment of the present application;
[0038] Figure 6 A schematic diagram of a scene in which a mobile phone moves to a face provided in an embodiment of the present application;
[0039] Figure 7 A schematic diagram of a sound source angle provided in an embodiment of the present application;
[0040] Figure 8 A schematic diagram of an interface of first information provided in an embodiment of the present application;
[0041] Fig. 9 A schematic diagram of an interaction between an electronic device and a server performing voice interaction provided in an embodiment of the present application;
[0042] Fig.10A schematic diagram of a floating frame interface provided in an embodiment of the present application Figure 1 ;
[0043] Fig.11 A schematic diagram of a floating frame interface provided in an embodiment of the present application Figure 2 ;
[0044] Fig.12 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0045] Various exemplary embodiments, features and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0046] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0047] The term "and / or" herein is only a description of the association relationship of the associated objects, indicating that there may be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the term "at least one" herein represents any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set consisting of A, B, and C.
[0048] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the subject matter of the present disclosure.
[0049] With the popularization of electronic devices, their functions are becoming more and more powerful, especially smart phones. At present, many smart phones are installed with voice interaction applications, which have voice control functions such as making calls, sending text messages, opening software, playing music, and setting memos through voice interaction, and have information query functions such as answering weather, bus routes, map locations and business information, and viewing pictures. It also comes with daily chat dialogue functions. Users can input voice commands to electronic devices to control electronic devices to complete some operations. So that voice interaction applications can provide users with very intelligent and humanized services, bringing great convenience and better service experience to users.
[0050] Usually, voice interaction applications need to be woken up by the user when performing voice interaction. Figure 1 In (A), the user can wake up the voice interaction application by inputting a specific voice wake-up word. The electronic device activates the voice input after detecting the corresponding voice wake-up word. Figure 1 In (B), the user can click a physical button on the electronic device or an interface element on the screen to wake up the voice interaction application. The electronic device activates the voice input after detecting the corresponding click operation.
[0051] However, the above wake-up of the voice interaction application is passive, and the user must repeat the above wake-up operation every time before voice interaction with the voice interaction application. When the user needs to use the voice interaction application frequently, it is cumbersome to repeat the wake-up operation, which wastes a lot of usage time and has low interaction efficiency with the electronic device, thereby reducing the user experience.
[0052] Based on the above content, an embodiment of the present application provides a voice interaction method and an electronic device, and the voice interaction method is applied to an electronic device. The electronic device obtains audio data and first sensor data; the first sensor data is used to characterize the movement of the electronic device. When the electronic device includes a voice signal in the audio data and determines that the position change of the electronic device is greater than a first preset threshold according to the first sensor data, the sound source direction and second sensor data corresponding to the audio data are obtained; the second sensor data is used to characterize the change in the distance between the electronic device and the user. Afterwards, when the electronic device determines that the electronic device is close to the user according to the sound source direction and the second sensor data, the voice interaction application is started. And through the voice interaction application, the response response operation corresponding to the voice signal is executed.
[0053] The embodiments of the present application can fuse the sensor data corresponding to a variety of different sensors into multimodal information, and the multimodal information can express more comprehensive and accurate information. Specifically, not only can the audio characteristic airflow characteristic value and sound pressure difference value corresponding to the audio data be judged separately, but also the position characteristic value can be comprehensively judged in combination with the movement of the electronic device and the change in distance from the user. The voice interaction function is triggered only after the audio data, the first sensor data and the second sensor data meet certain conditions.
[0054] The information expressed by a single sensor data is relatively limited. For example, only the distance sensor is used to determine the distance between the mobile phone and the user's mouth, and the voice interaction function is triggered when the distance between the electronic device and the user's mouth reaches a certain range. In this way, the user may not want to use the voice interaction function in the electronic device and only use the electronic device at a close distance, which will cause misrecognition.
[0055] Thus, compared with using the sensing data of a single sensor for voice interaction recognition, the embodiment of the present application effectively improves the accuracy of voice interaction recognition and avoids the occurrence of misrecognition, thereby improving the user experience.
[0056] Moreover, in the embodiments of the present application, the user does not need to perform specific wake-up operations on the electronic device, such as inputting a voice wake-up word, clicking on a physical button, and clicking on the screen. Of course, the electronic device does not need to trigger the voice interaction function after detecting the wake-up operation. The user only needs to approach the phone and speak without performing the wake-up operation. The electronic device can automatically trigger the voice interaction function after detecting that different sensor data meets certain conditions. The interactivity between the user and the mobile phone is improved, making the interaction more natural. The user can interact directly with the mobile phone every time, without having to repeat the wake-up operation every time the voice is input, saving a lot of usage time and reducing the power consumption of the mobile phone.
[0057] The electronic device provided in the embodiment of the present application may include at least one of a mobile phone, a smart watch, a tablet computer, a foldable electronic device, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (Ultra-Mobile Personal Computer, UMPC), a netbook, a cellular phone, a personal digital assistant (Personal Digital Assistant, PDA), an augmented reality (Augmented Reality, AR) device, a virtual reality (Virtual Reality, VR) device, an artificial intelligence (Artificial Intelligence, AI) device, a wearable device, an in-vehicle device, a vehicle, a smart home device or a smart city device. The embodiment of the present application does not impose any special restrictions on the specific type of the electronic device.
[0058] The operating system installed in the electronic device provided in the embodiment of the present application includes but is not limited to Or other operating systems. This application does not limit the specific type of electronic device and the type of operating system when the operating system is installed.
[0059] For example, taking the electronic device as a mobile phone, Figure 2 A schematic structural diagram of the mobile phone 100 is shown.
[0060] The mobile phone 100 may include a processor 110, an external memory interface 120, an internal memory 121, a Universal Serial Bus (USB) connector 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a Subscriber Identification Module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, a gravity acceleration sensor 180E, an ultrasonic distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, an image sensor 180M, etc.
[0061] It is to be understood that the structure illustrated in the embodiment of the present application does not constitute a specific limitation on the mobile phone 100. In other embodiments of the present application, the mobile phone 100 may include more or fewer components than shown in the figure, or combine some components, or separate some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0062] The processor 110 may include one or more processing units, for example, the processor 110 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0063] The processor can generate operation control signals based on instruction opcodes and timing signals to complete the control of instruction fetching and execution.
[0064] The processor 110 may also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 may be a cache memory (CACHE). The memory may store instructions or data that have been used or are frequently used by the processor 110. If the processor 110 needs to use the instruction or data, it may be directly called from the memory. This avoids repeated access, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0065] In some embodiments, the memory also stores other data besides the computer program. The other data may include data generated after the operating system or application is run. The data includes system data (such as configuration parameters of the operating system) and user data. For example, data cached by an application opened by a user is typical user data. The memory generally includes internal memory and external memory. The internal memory may be RAM, ROM, and CACHE. The external memory may include a hard disk, a floppy disk, an optical disk, a USB flash drive, and a multimedia card.
[0066] It is understandable that the connection relationship between the modules illustrated in the embodiment of the present application is only for illustrative purposes and does not constitute a structural limitation on the mobile phone 100. In other embodiments of the present application, the mobile phone 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0067] The charging management module 140 is used to receive charging input from a charger. The power management module 141 is used to connect the battery 142, the charging management module 140 and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and supplies power to the processor 110, the internal memory 121, the display screen 194, the camera 193, and the wireless communication module 160. The wireless communication function of the mobile phone 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0068] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in mobile phone 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve the utilization of antennas. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0069] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the mobile phone 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LowNoise Amplifier, LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.
[0070] In some embodiments, the antenna 1 of the mobile phone 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the mobile phone 100 can communicate with the network and other electronic devices through wireless communication technology. The wireless communication technology may include Global System For Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include a Global Positioning System (GPS), a Global Navigation Satellite System (GLONASS), a BeiDou Navigation Satellite System (BDS), a Quasi-Zenith Satellite System (QZSS) and / or a Satellite Based Augmentation System (SBAS).
[0071] The mobile phone 100 can realize the display function through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, which connects the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.
[0072] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some embodiments, the mobile phone 100 may include one or more display screens 194.
[0073] The mobile phone 100 can realize the camera function through the camera 193, ISP, video codec, GPU, display screen 194, application processor AP, neural network processor NPU, etc.
[0074] In some embodiments, the CPU or GPU or NPU in the processor 110 can process the color image data and depth data collected by the camera 193 .
[0075] The processor 110 executes various functional methods or data processing of the mobile phone 100 by running instructions stored in the internal memory 121 and / or instructions stored in a memory set in the processor.
[0076] The mobile phone 100 can implement audio functions such as music playing and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the earphone interface 170D, and the application processor.
[0077] The gyro sensor 180B may be used to determine the motion posture of the mobile phone 100. In some embodiments, the angular velocity of the mobile phone 100 around three axes (ie, x, y, and z axes) may be determined by the gyro sensor 180B.
[0078] The gravity acceleration sensor 180E can be used to detect the magnitude of acceleration in various directions (generally three axes) of the mobile phone 100. When the mobile phone 100 is stationary, the magnitude of gravity and the direction of the display screen can be detected.
[0079] The ultrasonic distance sensor 180F can be used to determine the distance data between the mobile phone 100 and a sound source. The image sensor 180M can be used to determine the distance data between the mobile phone 100 and a user's face.
[0080] The software system of the mobile phone 100 may adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. The embodiment of the present application takes the Android system of the layered architecture as an example to exemplify the software structure of the mobile phone 100.
[0081] Figure 3 It is a software structure block diagram of the mobile phone 100 according to an embodiment of the present application.
[0082] The layered architecture divides the software into several layers, each with clear roles and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, from top to bottom: application layer, application framework layer, Android runtime (ART) and native C / C++ library, hardware abstract layer (HAL) and kernel layer.
[0083] The application layer can include a series of application packages.
[0084] like Figure 3 As shown, the application package may include camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message and other applications.
[0085] The application framework layer provides an application programming interface (API) and a programming framework for the applications in the application layer. The application framework layer includes some predefined functions.
[0086] like Figure 3 As shown, the application framework layer may include a window manager, a content provider, a view system, a resource manager, a notification manager, an activity manager, an input manager, and the like.
[0087] The window manager provides window management services (WMS). WMS can be used for window management, window animation management, surface management, and as a transfer station for the input system.
[0088] Content providers are used to store and retrieve data and make it accessible to applications. This data can include video, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.
[0089] The view system includes visual controls, such as controls for displaying text, controls for displaying images, etc. The view system can be used to build applications. A display interface can be composed of one or more views. For example, a display interface including a text notification icon can include a view for displaying text and a view for displaying images.
[0090] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.
[0091] The notification manager enables applications to display notification information in the status bar. It can be used to convey notification-type messages and can disappear automatically after a short stay without user interaction. For example, the notification manager is used to notify download completion, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as notifications of applications running in the background, or a notification that appears on the screen in the form of a dialog window. For example, a text message is displayed in the status bar, a prompt sound is emitted, an electronic device vibrates, an indicator light flashes, etc.
[0092] The activity manager can provide activity management services (AMS), which can be used to start, switch, and schedule system components (such as activities, services, content providers, and broadcast receivers) as well as manage and schedule application processes.
[0093] The input manager can provide input management services (IMS), which can be used to manage system input, such as touch screen input, key input, sensor input, etc. IMS takes events from input device nodes and distributes them to appropriate windows through interaction with WMS.
[0094] The Android runtime includes the core library and the Android runtime. The Android runtime is responsible for converting source code into machine code. The Android runtime mainly uses the ahead of time (AOT) compilation technology and the just in time (JIT) compilation technology.
[0095] The core library is mainly used to provide basic Java class library functions, such as basic data structures, mathematics, IO, tools, databases, networks, etc. The core library provides an API for users to develop Android applications.
[0096] The native C / C++ library can include multiple functional modules, such as surface manager, media framework, libc, OpenGL ES, SQLite, Webkit, etc.
[0097] Among them, the surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications. The media framework supports playback and recording of multiple commonly used audio and video formats, as well as static image files, etc. The media library can support multiple audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc. OpenGL ES provides drawing and operation of 2D graphics and 3D graphics in applications. SQLite provides a lightweight relational database for applications of electronic device 100.
[0098] The hardware abstraction layer runs in user space and may include display modules, audio modules, camera modules, Bluetooth modules, etc. It can encapsulate kernel layer drivers and provide a calling interface to the upper layer.
[0099] The kernel layer is the layer between hardware and software. The kernel layer contains at least display driver, camera driver, audio driver, and sensor driver.
[0100] Figure 4 The flowchart of a voice interaction method shown in the embodiment of the present application is as follows. The above-mentioned electronic device is a mobile phone as an example to specifically describe the voice interaction method provided in the embodiment of the present application. Figure 4 As shown, the method may include the following steps S401-S410.
[0101] Step S401: The mobile phone obtains audio data and first sensor data.
[0102] In actual application scenarios, in order to interact with the mobile phone by voice, the user usually moves the mobile phone. For example, the user picks up the mobile phone and keeps moving it for a period of time until the mobile phone is moved in front of the face. Then the user speaks to the mobile phone (such as "What's the weather like today") to realize the voice interaction function through the mobile phone.
[0103] In some embodiments of the present application, the voice interaction function in the mobile phone can be implemented in the form of an application. For example, a user can install a voice interaction application including a voice interaction function in a mobile phone, and the voice interaction application can be a third-party voice application or a system application. Of course, the voice interaction function in the mobile phone can also be implemented in the form of a service. For example, the voice interaction function is a function in the operating system of the mobile phone. The embodiments of the present application do not specifically limit the implementation method of the voice interaction function.
[0104] At the same time, the mobile phone can detect whether the voice interaction function is triggered when the user interface is displayed or when the user interface is not displayed. The user interface may include the main screen interface of the mobile phone, the interface of the application, the system settings interface, and the status bar display interface, etc. In other words, the mobile phone can detect whether the voice interaction function is triggered when it is used by the user or when the mobile phone is in the lock screen.
[0105] In some embodiments of the present application, the mobile phone can control the sensor path to turn on the first sensor to collect the first sensor data. The first sensor may include a gyroscope sensor, and the first sensor data is used to characterize the movement of the electronic device. At the same time, the mobile phone controls the audio path to turn on the microphone to collect audio data. The microphone is used to collect audio data input by the user and / or the environment to the mobile phone.
[0106] When a user moves a mobile phone, he usually picks up the mobile phone in a stationary state in front of his face, and then keeps the mobile phone in a suitable position to use it. Then, when the mobile phone is moving, the movement state of the mobile phone usually changes from a stationary state to a moving state, and then from a moving state to a stationary state. In this way, the mobile phone can determine the degree of position change of the electronic device based on the first sensor data. Specifically, the mobile phone can obtain acceleration data in real time based on the first sensor, and determine the degree of position change of the mobile phone through changes in the acceleration data.
[0107] Furthermore, when the user uses the voice interaction function, he usually speaks to the mobile phone. Therefore, the mobile phone can determine whether the collected audio data includes the voice output by the user, that is, whether there is voice activity in the audio data.
[0108] In one achievable manner, the mobile phone determines whether there is voice activity according to whether the audio data collected by the microphone includes a voice signal.
[0109] The audio signal in the audio data may include a speech signal and a non-speech signal. The speech signal is used to represent the signal converted from the sound of the user's speech, and the non-speech signal may be a signal generated when the user pauses in speaking, or a signal generated by environmental noise or noise generated by a speech acquisition device.
[0110] Therefore, the embodiment of the present application can facilitate the subsequent determination of whether the current mobile phone needs to start a voice interaction application by combining the first sensor data and audio data collected as mentioned above, that is, combining the position change of the electronic device and the voice activity, thereby avoiding the problem of determination error caused by only sensor data collected by a single sensor.
[0111] Step S402: The mobile phone performs voice activity detection on the audio data to determine whether the audio data includes a voice signal.
[0112] In some embodiments of the present application, see Figure 5 After collecting audio data, the mobile phone needs to perform voice activity detection on the audio data to determine whether the audio data contains voice signals. Among them, voice activity detection (VAD), also known as voice detection, is used to detect the presence or absence of voice in voice processing, thereby separating the voice segments and non-voice segments of the audio signal in the audio data.
[0113] Therefore, if the mobile phone determines that the audio data includes a voice signal, it determines that there is voice activity. In other words, if the mobile phone determines that the audio data includes a voice signal, it determines that the audio data includes the voice input by the user to the mobile phone. If the mobile phone determines that the audio data does not include a voice signal, it determines that there is no user voice activity, that is, the user did not speak to the mobile phone. Then the mobile phone will not need to start the voice interaction function later.
[0114] Exemplarily, the mobile phone can perform frame processing on the audio data to obtain the corresponding audio frame sequence. After that, the acoustic features of the audio frames in the audio frame learning are extracted, and the acoustic features include Fbank features. Next, the acoustic features of each audio frame in the audio frame sequence are input into the voice activity detection (VAD) model for classification to obtain the activity detection value (speech signal / non-speech signal posterior probability) of each audio frame in the audio frame sequence, and the activity detection value is used to indicate that the corresponding audio frame is a speech frame or a noise frame. Finally, the mobile phone can determine whether the audio data includes a speech signal according to the activity detection value of each audio frame in the audio frame sequence. For example, when the average value corresponding to the activity detection value of each audio frame is greater than the activity detection threshold, the mobile phone can determine that the audio data includes a speech signal. When the average value corresponding to the activity detection value of each audio frame is less than or equal to the activity detection threshold, the mobile phone can determine that the audio data does not include a speech signal. For another example, the activity detection value can be a logical variable, that is, it can be represented by "0" and "1", such as "1" to indicate the presence of a speech signal, and "0" to indicate the absence of a speech signal. Of course, "0" can also represent the presence of a voice signal, and "1" can represent the absence of a voice signal.
[0115] Of course, the mobile phone can also extract other time domain features corresponding to the audio data for characterizing the voice signal. For example, short-time energy, short-time zero crossing rate, fundamental frequency period, kurtosis of short-time amplitude spectrum and skewness of short-time amplitude spectrum, etc. The embodiment of the present application does not limit the specific implementation method of determining that the audio data includes the voice signal.
[0116] Step S403: When the audio data includes a voice signal, the mobile phone obtains voice airflow characteristics corresponding to the audio data and sound pressure differences of different microphones with respect to the audio data.
[0117] Among them, the speech airflow characteristics are used to characterize the wind noise generated by the speech airflow hitting the microphone; the sound pressure difference is used to characterize the sound pressure changes generated by microphones at different distances from the sound source.
[0118] In some embodiments of the present application, when the audio data includes a voice signal, it can indicate the presence of voice activity. However, in order to more accurately determine the scenario in which the user speaks to the mobile phone at a close distance, continue to refer to Figure 5 In one implementable manner, the embodiment of the present application may also perform speech airflow detection and sound pressure level difference detection on the audio data to determine the airflow characteristic value and sound pressure difference value corresponding to the audio data.
[0119] In one feasible manner, when the audio data includes a voice signal, the mobile phone can perform voice airflow detection on the audio data and extract the voice airflow features corresponding to the audio data. It is understandable that when a user speaks to a mobile phone at a close distance, the audio data collected by the microphone will have a unique breath and wind noise. Therefore, it is possible to determine whether the audio data includes the wind noise generated by the user's speech by the airflow feature value corresponding to the voice airflow feature. In other words, the airflow feature value is positively correlated with the degree of characteristic prominence of the voice airflow feature. The larger the airflow feature value of the voice airflow feature, the greater the degree of characteristic prominence of the voice airflow feature in the audio data, which means that the wind noise generated by the voice airflow hitting the microphone in the audio data is more obvious.
[0120] Exemplarily, the mobile phone can identify phonemes in the audio data, and for each phoneme in the audio data, determine whether the phoneme is an exhaled phoneme, which is used to characterize that air flows out of the mouth when the user speaks. Afterwards, the mobile phone divides the audio data into an audio frame sequence according to a fixed window length. And use the frequency characteristics to identify whether each audio frame sequence contains wind noise. Finally, the exhaled phonemes in the audio data and the wind noise segments identified in the audio frame sequence are compared for overlap, and the airflow feature value corresponding to the audio data is generated. Furthermore, it is convenient to judge the characteristic obviousness of the speech airflow features in the audio data according to the airflow feature value. It can then be determined whether the audio data includes the wind noise sound generated by the airflow generated by the user speaking hitting the microphone.
[0121] In one achievable manner, the mobile phone can also input the audio data into the trained first model to generate a corresponding airflow feature value. When the characteristic level corresponding to the speech airflow feature represented by the airflow feature value is greater than a certain level, it is determined that the user is close to the electronic device and speaking.
[0122] In some embodiments of the present application, a plurality of microphones such as a microphone array may be provided in a mobile phone. A microphone array is an array formed by a group of omnidirectional microphones located at different positions in space and arranged in a certain shape rule in a mobile phone. The array collects spatially propagated sound input, and the collected signal contains its spatial position information. For example, a microphone is respectively provided at the top and the bottom of the mobile phone. Then, the mobile phone can determine whether the user is close to the electronic device and speaking based on the difference in sound pressure between different microphones. In other words, the larger the difference in sound pressure, the greater the change in sound pressure generated by different microphones at the sound source distance, which means that the sound source distance is closer to at least one of the different microphones.
[0123] Specifically, when the audio data includes a voice signal, the mobile phone can also perform sound pressure level difference detection. It is understandable that when the sound source is close to the bottom microphone in the mobile phone, the sound pressure levels corresponding to the two microphones are different. For example, take the user speaking close to the bottom microphone as an example, that is, the sound source is close to the bottom microphone. The mobile phone can calculate the corresponding sound pressure level difference based on the distance between the sound source and the two microphones using a preset formula. The preset formula is as follows:
[0124] ΔL=20lg10(r1 / r2);
[0125] Among them, △L is the attenuation value caused by the increase in distance, unit is dB; r1 is the distance from the sound source to the top microphone in the mobile phone, unit is m; r2 is the distance from the sound source to the bottom microphone in the mobile phone, unit is m.
[0126] Based on the above calculation of the sound pressure level difference, it can be seen that when the distance from the sound source to the top microphone in the mobile phone is twice the distance from the sound source to the bottom microphone in the mobile phone, the sound pressure level attenuation, that is, the sound pressure level difference is 6dB. In other words, since the distance between the top microphone and the bottom microphone is fixed, the smaller the distance from the sound source to the bottom microphone in the mobile phone, the greater the sound pressure level difference.
[0127] Therefore, the embodiment of the present application can determine that the electronic device is close to the user and the user is speaking to the mobile phone by performing voice airflow detection and sound pressure level difference detection on the audio data when the audio data includes a voice signal.
[0128] At the same time, compared with the information represented by a single type of sensor in the related art, it can only provide information about a certain aspect of user behavior, resulting in limitations in its recognition. The embodiment of the present application performs voice activity detection, voice airflow detection, and sound pressure level difference detection on the audio data, so as to facilitate the subsequent fusion of audio data and different sensor data to obtain corresponding multi-modal information. In addition, it is possible to utilize the different characteristics corresponding to different data to identify user behaviors at various levels and in different dimensions and achieve better speech recognition effects.
[0129] Step S404: The mobile phone detects whether the airflow characteristic value corresponding to the audio data is greater than the third preset threshold, and whether the sound pressure difference is greater than the fourth preset threshold; if so, execute step S409, otherwise, execute step S408.
[0130] In some embodiments of the present application, in order to more accurately detect whether the voice interaction function is triggered, continue to refer to Figure 5 The mobile phone can perform voice detection, voice airflow detection, and sound pressure level difference detection on the audio data to generate the results and perform close-talk voice information fusion. In other words, the voice information obtained by the above detections can be fused to determine whether the user is speaking close to the mobile phone.
[0131] In one possible implementation, the mobile phone can further determine the airflow characteristic value and the sound pressure difference value corresponding to the audio data. If the airflow characteristic value is larger, it can indicate that the possibility of wind noise sound signal in the audio data is greater. Similarly, if the sound pressure difference value is larger, it can indicate that the sound pressure levels received by different microphones are different, that is, the size of the sound is different. And the size of the sound received by different microphones varies greatly. In this way, when there is a difference in the size of the sound received by different microphones in the mobile phone, it can be determined that the sound source is close to a certain microphone.
[0132] Based on the above, the embodiment of the present application sets a third preset threshold and a fourth preset threshold to fuse the airflow characteristic value and the sound pressure difference value for determination.
[0133] When the airflow characteristic value is greater than the third preset threshold and the sound pressure difference value is greater than the fourth preset threshold, the voice interaction application is started. That is to say, when the airflow characteristic value meets the third preset threshold, it means that the characteristic of the voice airflow feature in the audio data is more obvious. And when the sound pressure difference value meets the fourth preset threshold, it means that the sound source is close to the electronic device and close to the microphone in the electronic device. It can then be determined that the user is close to the phone and speaking, triggering the start of the voice interaction application.
[0134] Exemplarily, the fourth preset threshold is 6dB, and the third preset threshold is 0.8. When the overlap between the exhaled phonemes in the audio data and the wind noise segment is greater than 0.8, and the sound pressure level attenuation of the top microphone in the mobile phone is greater than 6dB compared with the bottom microphone, it is determined that the audio data is likely to include the wind noise sound generated by the user's speech, and the sound source is determined to be close to the bottom microphone. It is then determined that the user is speaking close to the mobile phone.
[0135] In some embodiments of the present application, based on the above fusion judgment of the airflow characteristic value and the sound pressure difference value, when both the airflow characteristic value and the sound pressure difference value meet the judgment condition, it is indicated that the judgment condition for the audio data is relatively strict. In other words, when both the airflow characteristic value and the sound pressure difference value do not meet the judgment condition, it cannot be completely determined that the user is not close to the mobile phone and speaking.
[0136] Therefore, when the airflow characteristic value of the voice airflow characteristic is less than or equal to the third preset threshold or the sound pressure difference is less than or equal to the fourth preset threshold, the mobile phone can also obtain the sound source direction and the second sensor data corresponding to the audio data when it is determined according to the first sensor data that the position change degree of the electronic device is greater than the first preset threshold. This is to facilitate the subsequent combination of the audio data and the sensor data of other sensors to comprehensively determine whether to trigger the voice interaction function.
[0137] Step S405: The mobile phone determines whether the degree of change in the position of the mobile phone is greater than a first preset threshold value based on the first sensing data.
[0138] Step S406: When the mobile phone determines, based on the first sensor data, that the degree of change in the position of the mobile phone is greater than a first preset threshold, the mobile phone obtains the sound source direction corresponding to the audio data and the second sensor data.
[0139] In some embodiments of the present application, the embodiments of the present application can comprehensively judge whether to start the voice interaction application based on multi-mode information such as the user's voice information and posture information. In other words, the first sensor and microphone in the mobile phone are always on, and the first sensor data collected by the first sensor and the audio data collected by the microphone are used to comprehensively judge whether to start the voice interaction application. Among them, the audio data can be used to characterize the voice information corresponding to the user, and the first sensor data can be used to characterize the posture information corresponding to the user.
[0140] Therefore, continue to see Figure 5 In the process of determining that the degree of change in the location of the mobile phone is greater than the first preset threshold value according to the first sensor data, the mobile phone can determine a location characteristic value according to the first sensor data, and the location characteristic value is used to characterize the degree of change in the location of the electronic device. Afterwards, the mobile phone determines that the degree of change in the location represented by the location characteristic value is greater than the first preset threshold value. In other words, the larger the location characteristic value, the greater the degree of change in the location of the mobile phone, which in turn indicates that the user is more likely to move the mobile phone.
[0141] Wherein, the first sensing data is collected by a first sensor, and the first sensor may include a gyroscope sensor. Specifically, the mobile phone can obtain the acceleration data at the first moment and the acceleration data at the second moment in the first sensing data; according to the acceleration data at the first moment and the acceleration data at the second moment, the position characteristic value corresponding to the acceleration change data is calculated; if the position characteristic value is greater than the first preset threshold, it is determined that the position change degree of the mobile phone is large. Wherein, the first moment is the starting moment when the mobile phone is picked up, and the second moment is the ending moment of the mobile phone in this picking up process. The embodiment of the present application does not specifically limit the first moment and the second moment. The process of the mobile phone being picked up may also include multiple first moments and second moments.
[0142] In the above process of determining whether the position feature value is greater than the first preset threshold, the mobile phone can analyze the movement angle of the mobile phone in the three-axis coordinate directions in the acceleration change data; if the movement angles in the three-axis coordinate directions are all greater than the first preset threshold, it is determined that the position feature value is greater than the first preset threshold.
[0143] Exemplarily, the mobile phone monitors and receives acceleration data collected by the gyroscope sensor at the first moment and the second moment in real time. The acceleration data includes acceleration data on three axis coordinates (X, Y and Z axis). Since the acceleration data represents the acceleration data collected in the three axis coordinate directions, when the acceleration data on one of the axes changes, it means that the mobile phone has moved in the axis coordinate direction. Therefore, it is possible to determine whether the mobile phone has moved by judging the change data of the acceleration data in the three axis directions. Then, the acceleration data in the three axis coordinate directions are converted into angular velocity data, and the movement angle of the mobile phone in the three axis coordinate directions is determined by the angular velocity data.
[0144] The embodiments of the present application do not impose any numerical limitation on the preset acceleration threshold, and those skilled in the art can determine the value of the preset acceleration threshold by themselves according to actual needs, such as 0.1, 0.2, 0.3, etc. These designs do not exceed the protection scope of the embodiments of the present application.
[0145] At the same time, since the acceleration data is collected in real time by the gyroscope sensor, the process of acquiring multiple groups of acceleration data at different times can be performed multiple times to more accurately determine whether the mobile phone has actually moved. The embodiment of the present application does not specifically limit the number of executions of the acquisition and judgment process, and can be set multiple times according to the device status corresponding to the specific mobile phone, and all are within the protection scope of the embodiment of the present application.
[0146] As another example, the mobile phone can also input the first sensor data into the trained second model to generate a corresponding position feature value. When the position feature value is greater than the first preset threshold, it is determined that the position change degree of the electronic device meets the preset degree, and then it can be determined that the position change degree of the electronic device is large, indicating that the user has lifted the mobile phone. This application does not specifically limit the model types of the first model and the second model.
[0147] Afterwards, in some embodiments of the present application, when the audio data includes a voice signal and it is determined based on the first sensor data that the position change of the electronic device is greater than a first preset threshold, the sound source direction and the second sensor data corresponding to the audio data are obtained. The second sensor data is used to characterize the change in the distance between the electronic device and the user.
[0148] In some implementations of the present application, before obtaining the second sensor data and the direction of the sound source, the mobile phone may obtain third sensor data, and the third sensor data is used to characterize the posture of the electronic device. When determining the direction of the display screen toward the user according to the third sensor data, the sound source direction and the second sensor data corresponding to the audio data are obtained.
[0149] The third sensor data is collected by the third sensor, and the third sensor may include a gravity acceleration sensor. Specifically, the gravity acceleration sensor is used to obtain the posture information of the mobile phone. The gravity acceleration sensor can sense the front and back data changes of the device in the three directions of the X-axis, Y-axis, and Z-axis. For example, if the origin of the mobile phone is the lower left corner, the first gravity acceleration data of the X-axis, Y-axis, and Z-axis are (0,0,9). At T1, the mobile phone is placed horizontally, the screen is facing up, and the second gravity acceleration data of the X-axis, Y-axis, and Z-axis are (0,0,9). At T2, the user picks up the mobile phone, the mobile phone is upright on a vertical plane, the screen is facing the user, and the third acceleration data of the X-axis, Y-axis, and Z-axis are (0,9,0). In this way, the mobile phone can determine the orientation of the display screen according to the third sensor data. When determining the orientation of the display screen to the user according to the third sensor data, the subsequent steps of obtaining the sound source direction and the second sensor data corresponding to the audio data are performed. Avoid the false triggering of the voice interaction function when the display screen is facing away from the user.
[0150] In some embodiments of the present application, when the audio data includes a voice signal, the mobile phone can also detect the direction of the sound source and determine the sound source angle corresponding to the sound source direction. The sound source angle is the angle between the sound source direction and the reference direction. The sound source direction is the direction from the sound receiving position of the mobile phone (such as the microphone position) to the sound source. The reference direction is the direction passing through the sound receiving position of the mobile phone, perpendicular to the display screen of the mobile phone and extending upward; the sound source angle can be greater than or equal to 0° and less than or equal to 180°.
[0151] That is to say, the mobile phone can locate the sound source based on the audio data collected by multiple microphones in the microphone array to determine the position information corresponding to the user who is speaking. The determined position information can be the two-dimensional position coordinates of the user, or the azimuth and distance of the user relative to the multiple microphones. The azimuth is the azimuth of the user in the coordinate system where the multiple microphones are located, and the distance is the distance between the user and the center position of the multiple microphones.
[0152] Exemplarily, the mobile phone obtains audio data collected by multiple microphones and converts each audio data into a corresponding frequency domain signal. Afterwards, a cross-spectral calculation is performed on the frequency domain signal corresponding to each audio data to obtain the time difference between the multiple microphones collecting the audio data. Next, a cross-spectral calculation is performed on each frequency domain signal obtained by the conversion to obtain the time difference (t2-t1)-(t2-t1) between the time when the second microphone to the nth microphone collects the sound source and the time when the first microphone collects the sound source. n -t1). Finally, the two-dimensional coordinates corresponding to the sound source can be obtained based on the time difference between the audio data collected by different microphones and the distance between each microphone. Finally, the sound source angle is determined based on the two-dimensional coordinates corresponding to the sound source.
[0153] Therefore, the embodiment of the present application can determine the position information of the sound source relative to the electronic device, that is, the position information of the user relative to the mobile phone, by detecting the direction of the sound source, so that the mobile phone can accurately determine whether to start the voice interaction application based on the position information corresponding to the user.
[0154] In some embodiments of the present application, the mobile phone can characterize the posture information corresponding to the user based on the first sensor data and / or the second sensor data. Of course, the mobile phone can also characterize the posture information corresponding to the user based on the first sensor data, the second sensor data and the third sensor data. It can be understood that the first sensor data is used to characterize the movement of the mobile phone, that is, it can characterize whether the user moves the mobile phone. The second sensor data is used to characterize the change in the distance between the mobile phone and the user, that is, it can characterize whether the user is close to the mobile phone. The third sensor data can characterize the orientation of the mobile phone display, that is, it can characterize whether the user is facing the mobile phone display towards himself. In other words, the mobile phone can combine the audio data and the first sensor data, the second sensor data and the third sensor data to comprehensively determine whether to start the voice interaction application.
[0155] In a feasible way, the mobile phone can further obtain data collected by different sensors to accurately identify the posture information corresponding to the user. Figure 5 The mobile phone obtains the second sensing data collected by the second sensor, and performs face distance and approach process detection and ultrasonic distance and approach process detection on the second sensing data, so as to further identify the user's posture information.
[0156] The second sensor may include an image sensor and a distance sensor, and the image sensor and the distance sensor are not always on. When the mobile phone detects that the position characteristic value corresponding to the first sensor data is greater than the first preset threshold, it triggers the image sensor and the distance sensor to be turned on, thereby obtaining the corresponding second sensor data.
[0157] Specifically, see Figure 6 , taking the scenario where the user holds the phone in front of the face as an example, see Figure 6 (A) to Figure 6 In (B), the user picks up the phone and holds it horizontally. Figure 6 (B) to Figure 6 In (C), the user changes from holding the phone horizontally to holding the phone 45° from the horizontal. In other words, when the user picks up the phone, the distance between the phone and the user changes from far to near. Therefore, the phone can characterize the change in the distance between the phone and the user through the second sensor data, and then characterize the corresponding posture information of the user.
[0158] The second sensing data may include first distance data and first distance change data, the first distance data is used to characterize the distance between the mobile phone and the user's face, and the first distance change data is used to characterize the distance change process when the mobile phone moves to the user's face. The first distance data may be greater than or equal to 5 cm and less than or equal to 15 cm.
[0159] In one implementation, the image sensor may include a first color (red green blue, RGB) image sensor, a second color (red yellow blue, RYB) image sensor, a black and white image sensor, a time of flight (TOF) sensor, an infrared image sensor, and the like.
[0160] Taking the TOF sensor as an example, a mobile phone can be provided with a TOF sensor, and the TOF sensor can be placed close to the camera. The TOF sensor usually measures the depth data of the measured target through the time-of-flight method. Specifically, the time-of-flight method measures the time interval T from the emission to the reception of the pulse signal actively emitted by the measuring instrument (often called the pulse ranging method) or the phase difference generated by the laser going back and forth to the measured object once (phase difference ranging method), and converts it into the distance of the photographed target.
[0161] For example, when a user is speaking on a mobile phone, the user needs to bring the mobile phone close to the face. The TOF sensor in the mobile phone collects the corresponding depth image. Afterwards, the mobile phone identifies the face area in the depth image, and obtains the average depth value of the face area through the area corresponding to the face area in the depth image. Then, the average depth value of the face area is determined as the first distance data.
[0162] Similarly, the TOF sensor in the mobile phone can collect different depth images corresponding to two different moments to obtain a first average depth value and a second average depth value of the face area, and then determine the difference between the first average depth value and the second average depth value as the first distance change data.
[0163] Afterwards, a first distance characteristic value is determined based on the first distance data and the first distance change data, and the first distance characteristic value is used to characterize the proximity between the mobile phone and the user's face. Exemplarily, different weights can be set for the first distance data and the first distance change data to determine the first distance characteristic value. When the first distance data is smaller and the first distance change data is larger, the obtained first distance characteristic value is smaller, that is, the closer the mobile phone is to the user's face.
[0164] The distance sensor may include an ultrasonic distance sensor, a laser distance sensor, an infrared distance sensor, and the like.
[0165] Taking an ultrasonic distance sensor as an example, a mobile phone may be provided with an ultrasonic distance sensor, which may be provided close to a speaker. The ultrasonic distance sensor may play an ultrasonic signal and use the ultrasonic echo ranging principle to detect the distance between the ultrasonic distance sensor and the user's sound source.
[0166] For example, when a user is talking on a mobile phone, the user needs to bring the mobile phone close to his mouth. The ultrasonic distance sensor in the mobile phone transmits ultrasonic waves into the air, and immediately returns when it encounters an obstacle, that is, the user. After receiving the reflected wave, the ultrasonic distance sensor calculates the distance between the ultrasonic distance sensor and the obstacle based on the time difference between the emission and reception of the echo and the propagation speed of the ultrasonic wave in the air, that is, calculates the distance between the user and the mobile phone (second distance data).
[0167] Similarly, the ultrasonic distance sensor can collect the distances from the obstacle at two different times to obtain a first obstacle distance and a second obstacle distance, and then determine the difference between the first obstacle distance and the second obstacle distance as the second distance change data.
[0168] Afterwards, a second distance characteristic value is determined according to the second distance data and the second distance change data, and the second distance characteristic value is used to characterize the proximity between the mobile phone and the sound source. The embodiment of the present application does not specifically limit the specific implementation method of how to determine the first distance characteristic value and the second distance characteristic value.
[0169] It can be seen that in addition to processing audio data, the embodiment of the present application also processes sensor data collected by other different sensors. The user's corresponding voice information and posture information are used to determine whether to subsequently start the voice interaction application.
[0170] The second sensor is activated to collect data only when the position characteristic value corresponding to the first sensor data is greater than the first preset threshold. If the second sensor and the third sensor are always controlled to collect data during the process of the first sensor collecting the first sensor data, a large amount of mobile phone resources will be wasted.
[0171] If the position characteristic value corresponding to the first sensor data is less than the first preset threshold, the second sensor data is no longer needed. Therefore, the embodiment of the present application can reduce the power consumption of the mobile phone and avoid unnecessary waste of resources such as power.
[0172] Step S407: The mobile phone determines the target distance characteristic value according to the sound source angle corresponding to the sound source direction and the first distance characteristic value and the second distance characteristic value corresponding to the second sensor data.
[0173] In the embodiment of the present application, when the mobile phone determines that the electronic device is close to the user according to the sound source direction and the second sensor data, in the process of starting the voice interaction application, the target distance characteristic value can be determined according to the sound source angle corresponding to the sound source direction and the first distance characteristic value and the second distance characteristic value corresponding to the second sensor data. Afterwards, the mobile phone will start the voice interaction application only when it determines that the electronic device is close to the user according to the target distance characteristic value.
[0174] Specifically, when the sound source angle is less than the angle threshold, the mobile phone determines the first distance characteristic value as the target distance characteristic value; or, when the sound source angle is greater than or equal to the angle threshold, the mobile phone determines the second distance characteristic value as the target distance characteristic value.
[0175] In some embodiments of the present application, the weights corresponding to the first distance eigenvalue and the second distance eigenvalue may be set based on the sound source angle obtained in step S406, thereby completing the fusion of information for characterizing the user's posture.
[0176] See also Figure 7, taking the mobile phone placed horizontally as an example, when the user is speaking, if the angle of the sound source between the user and the mobile phone is smaller, it means that the user's face is closer and facing the screen of the mobile phone. On the contrary, when the user is speaking, if the angle of the sound source between the user and the mobile phone is larger, it means that the sound source is closer to the sound receiving position of the mobile phone, such as the microphone. In this way, when the angle of the sound source is less than the angle threshold, the weight of the first distance eigenvalue corresponding to the image sensor can be set to 1, and the weight of the second distance eigenvalue corresponding to the ultrasonic distance sensor can be set to 0. Conversely, when the angle of the sound source is greater than or equal to the angle threshold, the weight of the first distance eigenvalue corresponding to the image sensor can be set to 0, and the weight of the second distance eigenvalue corresponding to the ultrasonic distance sensor can be set to 1. Of course, the embodiment of the present application does not specifically limit the weights corresponding to the first distance eigenvalue and the second distance eigenvalue.
[0177] Exemplarily, the angle threshold may be 45°. When the mobile phone determines that the angle of the sound source is less than the angle threshold, such as 45°, the weight of the first distance characteristic value corresponding to the image sensor is set to 1, that is, the target distance characteristic value is the first distance characteristic value. Conversely, the weight of the second distance characteristic value corresponding to the ultrasonic distance sensor is set to 1, that is, the target distance characteristic value is the second distance characteristic value.
[0178] It can be seen that the first distance characteristic value and the second distance characteristic value proposed in the embodiment of the present application both characterize the distance between the user and the mobile phone and the changing process of the user moving the mobile phone to the face in different dimensions. It is also possible to further characterize the posture information corresponding to the user and determine the specific scenario when the user uses the mobile phone. It is convenient to determine whether to start the voice interaction application based on the user's posture information, voice information and other multi-modal information in the future, and avoid determining the start of the voice interaction application based on the sensor information of a single sensor, which may cause the occurrence of false start.
[0179] Step S408: The mobile phone detects whether the sum of the target distance characteristic value, the position characteristic value and the airflow characteristic value is greater than a second preset threshold value. If yes, step S409 is executed, otherwise, step S410 is executed.
[0180] Step S409: The mobile phone starts the voice interaction application and executes the answer response operation corresponding to the voice signal through the voice interaction application.
[0181] Step S410: The mobile phone does not start the voice interaction application.
[0182] In some embodiments of the present application, see Figure 5 After obtaining the above-mentioned voice information and posture information used to characterize the user, the mobile phone can also continue to perform fusion judgment in combination with the above-mentioned voice information and posture information to comprehensively characterize the information of different aspects of the user to further determine whether it is necessary to start the voice interaction application.
[0183] In the process of determining that the mobile phone is close to the user according to the sound source angle corresponding to the sound source direction and the second sensor data, the mobile phone can determine that the mobile phone is close to the user according to the voice airflow feature, the target distance feature value and the first sensor data. Specifically, when the sum of the airflow feature value, the target distance feature value and the position feature value of the voice airflow feature is greater than the second preset threshold, it is determined that the mobile phone is close to the user.
[0184] That is to say, the larger the target distance characteristic value is, the closer the user is to the phone. The larger the position characteristic value is, the more the phone itself has moved and the greater the degree of position change is. The larger the airflow characteristic value corresponding to the audio data is, the greater the possibility that the audio data contains wind noise signals, that is, it can be determined that the audio data is input by the user and the user is close to the phone. Therefore, the embodiment of the present application can set weights for the airflow characteristic value, the target distance characteristic value, and the position characteristic value.
[0185] In one practicable manner, the mobile phone starts the voice interaction application and sends the audio data to the server. Otherwise, the mobile phone does not start the voice interaction application when the sum of the detected target distance characteristic value, the position characteristic value and the airflow characteristic value is less than or equal to the second preset threshold.
[0186] In some embodiments of the present application, during the process of executing the response operation corresponding to the voice signal through the voice interaction application, the mobile phone can voice prompt the response result, which is generated by the server after semantic understanding of the voice signal; and / or display the target interface, which includes the response result. In some embodiments of the present application, before displaying the target interface, the mobile phone can prompt the first information, which is used to prompt that the voice interaction application has been run.
[0187] For example, the first information may be displayed in a text form in a user interface. Figure 8 In (A), the first information is: "The voice interaction application is running". In another exemplary embodiment, the first information can also be displayed in the user interface in the form of an animated image. Figure 8 In (B), for example, the dynamic effect image includes a heart-shaped image, and the heart-shaped image rotates in a clockwise or counterclockwise direction.
[0188] In summary, the embodiment of the present application can comprehensively perform voice interaction recognition in dimensions such as user voice information and posture information. Specifically, not only can the collected audio data used to represent the user's voice be judged separately, but also the collected sensor data used to represent the user's posture can be combined for fusion judgment.
[0189] Among them, in the process of using only audio data for judgment, the judgment conditions of audio data need to be strict before the voice interaction application can be started. In the process of fusion judgment, by identifying two types of data of different dimensions, the accuracy of voice interaction recognition is effectively improved to avoid misidentification. In turn, the user experience is improved.
[0190] After some embodiments of the present application, Fig. 9 This is a schematic diagram of an electronic device performing voice interaction with a server according to an embodiment of the present application. Fig. 9 As shown, the voice interaction process between the mobile phone and the server may include the following steps S501-S504.
[0191] Step S501: The mobile phone starts the voice interaction application and sends the audio data to the server.
[0192] Step S502: The server performs semantic understanding on the audio data and generates semantic information corresponding to the audio data.
[0193] Step S503: If the server detects a response result that matches the semantic information, the server returns the response result to the mobile phone.
[0194] In some embodiments of the present application, after the mobile phone starts the voice interaction application, it sends the audio data to the server. The audio data includes the voice information input by the user. Figure 5 , the server can perform semantic understanding on the received audio data and generate semantic information corresponding to the audio data. After that, the server searches the semantic database for a response result that matches the semantic information.
[0195] In one achievable manner, if the server finds a response result that matches the semantic information, the response result is returned to the mobile phone.
[0196] Specifically, in the embodiment of the present application, the server uses automatic speech recognition to first convert the received audio data into text information, i.e., voice information. Among them, automatic speech recognition (automatic speech recognition, ASR) technology allows the computer to "dictate" the continuous speech spoken by different users, which is commonly known as a "voice dictation machine". It is a technology that realizes the conversion of "sound" to "text". Exemplarily, the words spoken by the user to the mobile phone are "Please open the settings application". The mobile phone sends the audio data to the server, and the server performs automatic speech recognition on the audio data in advance to generate the corresponding text information, i.e., "Please open the settings application".
[0197] After that, the server needs to perform semantic understanding on the text information and generate semantic information. And find out whether there is a response result that matches the semantic information. Exemplarily, the server can perform natural language understanding on the text information and generate a corresponding query statement, i.e., semantic information. The server searches the semantic database according to the query statement, finds and outputs the response result that matches the query statement. In other words, the server needs to extract the user's service needs from the semantic information and return the corresponding response result to the mobile phone according to the user's service needs.
[0198] In one achievable manner, the query statement matches the response result means that the two have a corresponding relationship, that is, the user's service demand is associated with the mobile phone. Of course, the embodiment of the present application does not limit the specific implementation of how the server queries the response result.
[0199] For example, the query statement is: "Please open the settings app", and the user's service demand is to open the settings app. The settings app is a system application installed in the mobile phone, that is, the service demand of the user is associated with the mobile phone. As a result, the server can detect the response result corresponding to the semantic information, for example: "OK, the settings app will be opened for you", and return the response result to the mobile phone. The mobile phone can also give a voice prompt for the response result, and / or display the response result in the target interface.
[0200] In one achievable manner, the server can perform semantic understanding of text information based on a semantic understanding model during the semantic understanding process. Specifically, the text information can be input into the semantic understanding model to identify the user's intention, so as to provide corresponding services to the user in the future.
[0201] In some embodiments of the present application, the server may also return text information corresponding to the audio data to the mobile phone in the process of returning the response result to the mobile phone. The text information corresponding to the audio data is the corresponding speech recognition text obtained after the server performs speech recognition on the audio data. The response result is the semantically understood text obtained by the server performing semantic understanding on the speech recognition text. For example, the text information corresponding to the audio data is: "Please open the settings app", and the response result is: "OK, the settings app will be opened for you soon."
[0202] Step S504: the mobile phone voice prompts the answer result and / or displays the target interface, and the target interface includes the answer result.
[0203] In some embodiments of the present application, after the voice interaction application in the mobile phone is started, the response result sent by the server is received, and the display screen can be controlled to display the target interface, which includes the response result. In the process of the mobile phone displaying the response result, the text information corresponding to the audio data sent by the server can also be displayed.
[0204] For example, when displaying the answer result and text information, the mobile phone may display them in the form of a floating box and voice broadcast. Fig.10 , the mobile phone can control the floating box to be displayed at a preset position of the user interface, and voice play the voice text displayed in the floating box. For example, close to the top end face or the bottom end face of the screen. Among them, the floating box includes a first display area and a second display area. The first display area is used to display the text information corresponding to the audio data, such as "Please open the settings application", and the second display area is used to display the response result, such as "OK, the settings application will be opened for you soon."
[0205] In one possible implementation, the mobile phone continues to refer to the Fig.10 , the above floating frame can be displayed on the top of the user interface in a certain proportion, so that the user can perceive it while avoiding blocking a large area of the user interface.
[0206] In one possible implementation, the mobile phone may also display the above floating frame on the unlocking interface in a certain proportion without displaying the user interface. Fig.11 , when the user does not currently operate the phone, the phone currently displays a black screen. Afterwards, the user interacts with the phone using voice. The phone can trigger the display of the unlock interface and display a floating box.
[0207] In another possible implementation, the mobile phone can give a voice prompt for the answer result. For example, the query sentence corresponding to the audio data is: "Please dial Ms. Li", and the user's service demand is to communicate with Ms. Li. Then, the answer result can be given a voice prompt: "OK, dialing Ms. Li".
[0208] It can be seen that when the mobile phone is displaying the text information corresponding to the answer result and the audio data, if the voice interaction function is detected to be triggered while the user interface is displayed, the user interface is kept displayed. And the text information corresponding to the answer result and the audio data is displayed at the preset position of the user interface. Avoid affecting the original user interface and make the voice interaction function better integrated with the mobile phone. If the mobile phone detects that the voice interaction function is triggered when the user interface is not displayed, it controls the display of the unlocking interface. And the text information corresponding to the answer result and the audio data is displayed at the preset position of the unlocking interface.
[0209] The embodiment of the present application can detect that the voice interaction function is triggered in both scenarios where the mobile phone displays the user interface and does not display the user interface, and respond to the user quickly to improve the user's usage experience.
[0210] In some schemes, multiple embodiments of the present application can be combined, and the combined scheme can be implemented. Optionally, some operations in the process of each method embodiment are optionally combined, and / or the order of some operations is optionally changed. In addition, the execution order between the steps of each process is only exemplary and does not constitute a restriction on the execution order between the steps. There can also be other execution orders between the steps. It is not intended to show that the execution order is the only order in which these operations can be performed. A person of ordinary skill in the art will think of a variety of ways to reorder the operations described in the embodiments of the present application. In addition, it should be noted that the process details involved in a certain embodiment of the present application are also applicable to other embodiments in a similar manner, or different embodiments can be used in combination.
[0211] In addition, some steps in the method embodiment may be equivalently replaced by other possible steps. Alternatively, some steps in the method embodiment may be optional and may be deleted in certain usage scenarios. Alternatively, other possible steps may be added to the method embodiment.
[0212] The embodiment of the present application also provides a speech recognition device, which includes an acquisition module, an audio analysis module, a sensor data analysis module, an audio fusion module, and a comprehensive posture fusion module. The acquisition module is used to acquire audio data, first sensor data, second sensor data, and third sensor data. The audio analysis module is used to detect whether the audio data includes a voice signal, extract the voice airflow characteristics and sound source direction corresponding to the audio data, and the sound pressure difference of different microphones for the audio data. The audio fusion module is used to fuse the airflow characteristic value and the sound pressure difference of the voice airflow characteristic to determine whether to start the voice interaction application. The sensor data analysis module is used to determine the position characteristic value according to the first sensor data, determine the target distance characteristic value according to the second sensor data, and determine the orientation direction of the display screen according to the third sensor data. The comprehensive posture fusion module is used to fuse the airflow characteristic value, the target distance characteristic value, and the position characteristic value of the voice airflow characteristic to determine whether to start the voice interaction application.
[0213] The present application also provides an electronic device, such as a mobile phone, see Fig.12 The mobile phone includes: a memory 1020 and one or more processors 1010 and a communication interface 1030 .
[0214] The memory 1020 and the communication interface 1030 are coupled to the processor 1010. For example, the memory 1020, the communication interface 1030 and the processor 1010 may be coupled together via a bus 1040. The memory also stores a computer program code, which includes computer instructions. When the computer instructions are executed by the processor, the electronic device may execute the various functions or steps executed by the mobile phone 100 in the above method embodiment. The structure of the electronic device may refer to Figure 2 The structure of the mobile phone 100 is shown.
[0215] The embodiment of the present application also provides a chip system, which includes at least one processor and at least one interface circuit. The processor and the interface circuit can be interconnected by lines. For example, the interface circuit can be used to receive signals from other devices (such as a memory of an electronic device). For another example, the interface circuit can be used to send signals to other devices (such as a processor). Exemplarily, the interface circuit can read instructions stored in the memory and send the instructions to the processor. When the instructions are executed by the processor, the electronic device can perform the various steps in the above embodiments. Of course, the chip system can also include other discrete devices, which is not specifically limited in the embodiment of the present application.
[0216] An embodiment of the present application also provides a computer-readable storage medium, which includes computer instructions. When the computer instructions are executed on the above-mentioned electronic device, the electronic device executes each function or step executed by the mobile phone in the above-mentioned method embodiment.
[0217] The embodiment of the present application also provides a computer program product. When the computer program product is run on a computer, the computer is enabled to execute each function or step executed by the mobile phone in the above method embodiment.
[0218] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0219] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0220] The units described as separate components may or may not be physically separated, and the components shown as units may be one physical unit or multiple physical units, which may be located in one place or distributed in multiple different places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0221] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0222] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
[0223] The above contents are only specific implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application shall be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A voice interaction method, characterized in that: Applied to electronic equipment, the method comprises: Acquire audio data and first sensor data, wherein the first sensor data is used to characterize the movement of the electronic device; When the audio data includes a voice signal and it is determined according to the first sensing data that the position change degree of the electronic device is greater than a first preset threshold, third sensing data is acquired, and the third sensing data is used to characterize the posture of the electronic device; When it is determined that the display screen of the electronic device is facing the user according to the third sensor data, the sound source direction and second sensor data corresponding to the audio data are obtained; the second sensor data is used to characterize the change of the distance between the electronic device and the user; Determine a target distance characteristic value according to a sound source angle corresponding to the sound source direction and a first distance characteristic value and a second distance characteristic value corresponding to the second sensing data; the first distance characteristic value is used to characterize the proximity between the electronic device and the user's face, and the second distance characteristic value is used to characterize the proximity between the electronic device and the sound source; When it is determined according to the target distance characteristic value that the electronic device is close to the user, starting the voice interaction application; Executing a response operation corresponding to the voice signal through the voice interaction application; The method further comprises: Extracting speech airflow features corresponding to the audio data, wherein the speech airflow features are used to characterize wind noise generated by the speech airflow hitting the microphone; Acquire sound pressure difference values of different microphones for the audio data, where the sound pressure difference value is used to characterize the sound pressure changes generated by different microphones at the distance of the sound source; When it is determined according to the first sensor data that the position change degree of the electronic device is greater than a first preset threshold, acquiring the sound source direction and second sensor data corresponding to the audio data; comprising: When the airflow characteristic value of the speech airflow characteristic is less than or equal to the third preset threshold or the sound pressure difference value is less than or equal to the fourth preset threshold, when it is determined according to the first sensor data that the position change degree of the electronic device is greater than the first preset threshold, acquiring the sound source direction and the second sensor data corresponding to the audio data; When the airflow characteristic value is greater than the third preset threshold and the sound pressure difference value is greater than the fourth preset threshold, a voice interaction application is started.
2. The method according to claim 1, characterized in that The determining, according to the first sensing data, that the position change degree of the electronic device is greater than a first preset threshold value includes: Determine a position characteristic value according to the first sensing data, wherein the position characteristic value is used to characterize a degree of position change of the electronic device; It is determined that the position change degree represented by the position feature value is greater than the first preset threshold.
3. The method according to claim 1, characterized in that The sound source angle is the angle between the sound source direction and the reference direction. The sound source direction is the direction from the sound receiving position of the electronic device to the sound source. The reference direction is the direction passing through the sound receiving position of the electronic device, perpendicular to the display screen of the electronic device and extending upward.
4. The method according to claim 3, characterized in that: The determining the target distance characteristic value according to the sound source angle corresponding to the sound source direction and the first distance characteristic value and the second distance characteristic value corresponding to the second sensing data includes: When the sound source angle is less than an angle threshold, determining the first distance feature value as the target distance feature value; Alternatively, when the sound source angle is greater than or equal to the angle threshold, the second distance feature value is determined as the target distance feature value.
5. The method according to claim 1, characterized in that The determining, according to the target distance characteristic value, that the electronic device is close to the user comprises: The electronic device is determined to be close to the user according to the voice airflow characteristics, the target distance characteristic value and the first sensing data.
6. The method according to claim 5, characterized in that The determining that the electronic device is close to the user according to the voice airflow feature, the target distance feature value and the first sensing data includes: When the sum of the airflow characteristic value of the voice airflow characteristic, the target distance characteristic value and the position characteristic value is greater than a second preset threshold, determining that the electronic device is close to the user; The airflow characteristic value is used to characterize the characteristic prominence of the speech airflow feature in the audio data, and the airflow characteristic value is positively correlated with the characteristic prominence of the speech airflow feature.
7. The method according to any one of claims 1 to 6, characterized in that: The performing, by the voice interaction application, a response operation corresponding to the voice signal, includes: A voice prompt response result, wherein the response result is generated by the server after semantic understanding of the voice signal; And / or, displaying a target interface, wherein the target interface includes the response result.
8. An electronic device, characterized in that: The electronic device includes a memory and one or more processors; the memory is coupled to the processor; wherein computer program code is stored in the memory, and the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device executes the voice interaction method as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed on a computer, the computer can execute the voice interaction method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Voice interaction wake-up electronic equipment based on microphone signals, method and medium
CN110428806A
Voice interaction auxiliary operation method and device and computer readable storage medium
CN114038461A
Radio method of handheld device and handheld device
CN114283798A