Speech Recognition Method, Apparatus, Electronic Device, and Storage Medium
Through ultrasonic detection of the motion state and orientation of the sound source, combined with multiple device detection points, the ASR drop caused by the change of the speaker's position after the smart device wakes up is solved, and the accuracy of speech recognition is improved.
Patent Information
- Application Number
- CN202110600951.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-31
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2041-05-31
AI Technical Summary
In the prior art, the automatic speech recognition rate (ASR) decreases when the speaker's position changes after the smart device wakes up.
The motion state of the sound source is captured through ultrasonic detection, the direction after the sound source is obtained, and the ultrasonic detection points of multiple electronic devices are combined to determine the direction of the sound source, thereby performing speech recognition.
Without significantly increasing the calculation amount, the automatic speech recognition rate (ASR) is improved and the recognition accuracy of sound source position changes is improved.
Smart Images

Figure CN115480250B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition technology, and in particular, to a speech recognition method, apparatus, electronic device, and storage medium.
Background Art
[0002] With the development of technology, intelligent devices have gradually entered users' homes, forming an intelligent home environment and enriching users' lives. For example, intelligent speakers, as an intelligent device in smart homes, can help users check the weather, answer encyclopedic questions, play music, etc.; these intelligent devices are all set with default wake-up words, and the intelligent devices can be in a sleep state when not working. When a user needs the intelligent device to work, the user can call the wake-up word of the intelligent speaker (such as "Xiaoyi Xiaoyi") by voice. After the intelligent device detects its own wake-up word, it enters the working state. It receives the user's voice request, then performs speech recognition and intent recognition, then generates a speech according to the user's intent, and converts the speech into TTS (text to speech) to broadcast feedback information to the user by voice.
[0003] When an intelligent device wakes up and performs speech recognition, a microphone array is used to collect the speaker's voice signal. Among them, the microphone array includes a linear microphone array and a circular microphone array. After the speaker's voice signal is collected by the microphone array, wake-up word recognition is performed. Before performing wake-up word recognition, audio preprocessing is performed on the voice signal, and this process includes echo cancellation, beamforming, automatic gain, and noise reduction, etc. After the audio preprocessing is completed, wake-up word recognition is performed, and the orientation of the speaker is confirmed through sound source localization during the wake-up word recognition process.
[0004] After waking up the intelligent device, the intelligent device will continue to use the microphone array to collect the speaker's audio for subsequent speech recognition, and then perform corresponding actions after recognizing the user's intent. When the intelligent device collects ASR (automatic speech recognition) audio, it usually uses the orientation obtained through sound source localization during wake-up word recognition to collect the voice signal.
[0005] After the speaker wakes up the intelligent device by voice, the sound pickup of the intelligent device still uses the orientation of the speaker estimated through sound source localization during wake-up detection, and there is a possibility that the ASR recognition rate will decrease. Because after waking up the intelligent device, the audio preprocessing algorithm of the intelligent device cannot perceive the change in the speaker's position, and still uses the orientation of the speaker obtained through sound source localization during the wake-up word recognition process, that is, the orientation information of the speaker when waking up the intelligent device, resulting in a deviation in the audio preprocessing algorithm for ASR audio processing, and thus a decrease in the ASR recognition rate.
Summary of the Invention
[0006] In view of this, embodiments of the present application provide a voice recognition method, apparatus, electronic device, and storage medium, which are used to solve the technical problem that the ASR accuracy rate decreases when the position of the speaker changes after the intelligent electronic device is awakened and enters the working state in the prior art.
[0007] In a first aspect, embodiments of the present application provide a voice recognition method, and the method includes:
[0008] Receiving a wake-up signal from a sound source and obtaining a first direction of the wake-up signal;
[0009] Capturing the motion state of the sound source through ultrasonic detection and obtaining a second direction after the motion of the sound source;
[0010] Determining the orientation of the sound source through ultrasonic detection according to the first direction and the second direction, and performing voice recognition on the sound source.
[0011] Through the solution provided in this embodiment, the movement of the speaker at the sound source is detected by ultrasonic waves, and the orientation is determined by determining the position of the speaker's movement. Without significantly increasing the calculation amount, the ASR recognition rate is improved to a certain extent.
[0012] In a preferred implementation, between the step of capturing the motion state of the sound source through ultrasonic detection and obtaining a second direction after the motion of the sound source, and the step of determining the orientation of the sound source through ultrasonic detection according to the first direction and the second direction, and performing voice recognition on the sound source, the voice recognition method further includes the following steps:
[0013] Using multiple electronic devices to obtain the orientation of the sound source relative to each of the electronic devices through ultrasonic detection.
[0014] Through the solution provided in this embodiment, in the application scenario of multiple electronic devices, multiple ultrasonic detection points are formed to more accurately locate the orientation information of the sound source.
[0015] In a preferred implementation, in the step of using multiple electronic devices to obtain the orientation of the sound source relative to each of the electronic devices through ultrasonic detection, the following steps are further included:
[0016] Measuring a first distance between two electronic devices through ultrasonic detection;
[0017] According to the second direction, respectively measuring a second distance and a third distance between the sound source and the two electronic devices through ultrasonic detection;
[0018] Calculating the orientation of the sound source relative to the two electronic devices according to the first distance, the second distance, and the third distance;
[0019] Repeat the measurement of the first distance, the second distance, and the third distance, and calculate the azimuth of the sound source relative to each of the electronic devices.
[0020] Through the solution provided in this embodiment, the cosine theorem is used to determine the azimuth of the sound source relative to each electronic device. The number of terms required for calculation is small and easy to obtain. While increasing the accuracy of sound source azimuth determination, the required amount of calculation is small.
[0021] In a preferred embodiment, in the step of receiving the wake-up signal of the sound source and obtaining the first direction of the wake-up signal, the following steps are further included:
[0022] Detect the sound source in real time;
[0023] Capture the wake-up signal that triggers the wake-up event in the sound source;
[0024] Obtain the first direction of the wake-up signal.
[0025] Through the solution provided in this embodiment, only the wake-up signal that triggers the wake-up event will activate the electronic device and the voice recognition function, and other non-wake-up signals can be excluded at this step, playing a role in signal filtering.
[0026] In a preferred embodiment, in the step of detecting the motion state of the sound source by ultrasonic detection and obtaining the second direction after the sound source moves, the following steps are further included:
[0027] Send a first ultrasonic signal in the first direction;
[0028] The first receiving unit receives the first echo signal of the first ultrasonic signal;
[0029] Calculate the frequency shift between the first ultrasonic signal and the first echo signal;
[0030] Determine the motion state of the sound source according to the frequency shift;
[0031] When the motion state of the sound source is moving, the second receiving unit at a first distance from the first receiving unit receives the second echo signal of the first ultrasonic signal with a first wavelength;
[0032] Calculate the phase difference between the first echo signal and the second echo signal;
[0033] According to the first distance, the first wavelength, and the phase difference, calculate the azimuth angle of the sound source relative to the first receiving unit and the second receiving unit;
[0034] Wherein, the first distance is perpendicular to the first direction.
[0035] Through the solution provided by this embodiment, by using the phase differences of different echo signals of the same ultrasonic signal at different positions and the differences in the positions of the receiving points, the azimuth angles of the sound source relative to different receiving points can be accurately calculated.
[0036] In a preferred implementation, in the step of determining the azimuth of the sound source by ultrasonic detection according to the first direction and the second direction and performing speech recognition on the sound source, the following steps are further included:
[0037] Send ultrasonic waves in the first direction to obtain a first beam;
[0038] Send ultrasonic waves in the second direction to obtain a second beam;
[0039] Discriminate between the first beam and the second beam to confirm the azimuth of the sound source;
[0040] Perform speech recognition on the sound source.
[0041] Through the solution provided by this embodiment, by using the first beam and the second beam and discriminating various parameters in two different scenarios of the presence and absence of a sound source, the direction corresponding to the beam close to or being the human voice can be confirmed, thereby facilitating speech recognition and identifying the intention of the speaker.
[0042] In a second aspect, an embodiment of the present application provides a speech recognition device, which includes: an ultrasonic transceiver module, a processing module, and an identification module that communicate with each other;
[0043] The ultrasonic transceiver module is used to receive the wake-up signal of the sound source and obtain the first direction of the wake-up signal;
[0044] The processing module is used to capture the motion state of the sound source by ultrasonic detection and obtain the second direction after the sound source moves;
[0045] The identification module is used to determine the azimuth of the sound source by ultrasonic detection according to the first direction and the second direction and perform speech recognition on the sound source.
[0046] Through the solution provided by this embodiment, by detecting the movement of the speaker at the sound source according to ultrasonic waves and determining the azimuth by determining the position of the speaker's movement, the ASR recognition rate can be improved to a certain extent without significantly increasing the calculation amount.
[0047] In a preferred implementation, the ultrasonic transceiver module includes an ultrasonic unit and at least two receiving units. The ultrasonic unit is used to send ultrasonic signals, and the receiving unit is used to receive the wake-up signal and the echo signal of the ultrasonic signal.
[0048] Through the solution provided in this embodiment, using multiple receiving units to receive echo signals is beneficial to more accurately calculate the azimuth of the sound source and improve the recognition rate of ASR.
[0049] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor: the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory so that the electronic device executes the method described in the first aspect.
[0050] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, including a program or instruction, when the program or instruction runs on a computer, the method described in the first aspect is executed.
[0051] Compared with the prior art, the technical solution of the present application has at least the following beneficial effects:
[0052] The speech recognition method, device, electronic device and storage medium disclosed in the embodiments of the present application can effectively identify the position of the sound source in a scenario where the position of the sound source changes, accurately locate the sound source, confirm the azimuth of the sound source relative to the electronic device, and then better perform speech recognition and improve the recognition rate of ASR.
Description of the Drawings
[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0054] Figure 1 It is a schematic structural diagram of the electronic device provided in Embodiment 1 of the present application;
[0055] Figure 2 It is a basic step flowchart of the speech recognition method provided in Embodiment 2 of the present application;
[0056] Figure 3 It is a step flowchart of the speech recognition method provided in Embodiment 2 of the present application, including the step of using multiple electronic devices to measure the azimuth of the sound source;
[0057] Figure 4 It is a step flowchart of step Step200' in the speech recognition method provided in Embodiment 2 of the present application;
[0058] Figure 5 It is a step flowchart of step Step100 in the speech recognition method provided in Embodiment 2 of the present application;
[0059] Figure 6 It is the flowchart of step Step200 in the voice recognition method provided in Embodiment 2 of this application;
[0060] Figure 7 It is the flowchart of step Step300 in the voice recognition method provided in Embodiment 2 of this application;
[0061] Figure 8 It is the schematic diagram of detecting the motion state of the sound source by ultrasonic waves in the voice recognition method provided in Embodiment 2 of this application;
[0062] Figure 9 It is the schematic diagram of measuring the azimuth angle of the sound source relative to the electronic device after the sound source moves by ultrasonic waves in the voice recognition method provided in Embodiment 2 of this application;
[0063] Figure 10 It is the schematic diagram of measuring the azimuth of the sound source by two electronic devices in the voice recognition method provided in Embodiment 2 of this application;
[0064] Figure 11 It is the module schematic diagram of the voice recognition device provided in Embodiment 3 of this application;
[0065] Figure 12 It is the module schematic diagram of the ultrasonic transceiver module in the voice recognition device provided in Embodiment 3 of this application.
[0066] Reference numerals:
[0067] 1 - Antenna;
[0068] 2 - Antenna;
[0069] 100 - Electronic device; 110 - Processor; 120 - External memory interface; 121 - Internal memory; 130 - Universal serial bus interface; 140 - Charge management module; 141 - Power management module; 142 - Battery; 150 - Mobile communication module; 160 - Wireless communication module; 170 - Audio module; 170A - Speaker; 170B - Receiver; 170C - Microphone; 170D - Headphone jack; 180 - Sensor module; 180A - Pressure sensor; 180B - Gyroscope sensor; 180C - Barometric pressure sensor; 180D - Magnetic sensor; 180E - Acceleration sensor; 180F - Distance sensor; 180G - Proximity light sensor; 180H - Fingerprint sensor; 180J - Temperature sensor; 180K - Touch sensor; 18oL - Ambient light sensor; 180M - Bone conduction sensor; 190 - Button; 191 - Motor; 192 - Indicator; 193 - Camera; 194 - Display screen; 195 - User identification module card interface;
[0070] 10 - Ultrasonic transceiver module; 11 - Ultrasonic unit; 12 - Receiving unit; 20 - Processing module; 30 - Identification module.
Detailed implementation manners
[0071] For a better understanding of the technical solutions of this application, the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0072] It should be clear that the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.
[0073] The following describes embodiments of an electronic device and a method for implementing the electronic device. Among them, the electronic device may be a mobile phone (also known as a smart electronic device), a tablet personal computer, a personal digital assistant, an e-book reader, or a virtual reality interactive device, etc. The electronic device can be connected to various types of communication systems, such as: Long Term Evolution (LTE) system, future 5th Generation (5G) system, New Radio Access Technology (NR), and future communication systems, such as 6G system; it can also be a Wireless Local Area Networks (WLAN), etc.
[0074] For the convenience of description, in the following embodiments, a smart electronic device is taken as an example for illustration.
[0075] Embodiment 1
[0076] As Figure 1The following is a schematic structural diagram of an electronic device disclosed in Embodiment 1 of the present application. Among them, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0077] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0078] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modulation and demodulation processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0079] The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.
[0080] A memory may also be provided in the processor 110 for storing instructions and data. In one embodiment, the memory in the processor 110 is a cache memory. This memory may hold instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses and reduces the waiting time of the processor 110, thus improving the efficiency of the system.
[0081] In one embodiment, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0082] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In one embodiment, the processor 110 may include multiple groups of I2C buses. The processor 110 may be respectively coupled to the touch sensor 180K, the charger, the flash, the camera 193, etc. through different I2C bus interfaces. For example: The processor 110 may be coupled to the touch sensor 180K through the I2C interface, enabling the processor 110 to communicate with the touch sensor 180K through the I2C bus interface to implement the touch function of the electronic device 100.
[0083] The I2S interface can be used for audio communication. In one embodiment, the processor 110 may include multiple groups of I2S buses. The processor 110 may be coupled to the audio module 170 through the I2S bus to implement communication between the processor 110 and the audio module 170. In one embodiment, the audio module 170 may transmit an audio signal to the wireless communication module 160 through the I2S interface to implement the function of answering a call through a Bluetooth headset.
[0084] The PCM interface can also be used for audio communication to sample, quantize, and encode analog signals. In one embodiment, the audio module 170 and the wireless communication module 160 can be coupled through a PCM bus interface. In one embodiment, the audio module 170 can also transmit audio signals to the wireless communication module 160 through the PCM interface to implement the function of answering a call through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0085] The UART interface is a general-purpose serial data bus for asynchronous communication. This bus can be a two-way communication bus. It converts the data to be transmitted between serial communication and parallel communication. In one embodiment, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 through the UART interface to implement the Bluetooth function. In one embodiment, the audio module 170 can transmit audio signals to the wireless communication module 160 through the UART interface to implement the function of playing music through a Bluetooth headset.
[0086] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), etc. In one embodiment, the processor 110 and the camera 193 communicate through the CSI interface to implement the shooting function of the electronic device 100. The processor 110 and the display screen 194 communicate through the DSI interface to implement the display function of the electronic device 100.
[0087] The GPIO interface can be configured by software. The GPIO interface can be configured as a control signal or a data signal. In one embodiment, the GPIO interface can be used to connect the processor 110 to the camera 193, the display screen 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0088] The USB interface 130 is an interface that complies with the USB standard specification and can specifically be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device 100, and can also be used for data transmission between the electronic device 100 and peripheral devices. It can also be used to connect a headset to play audio. This interface can also be used to connect other electronic devices, such as AR devices, etc.
[0089] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are only illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0090] The charging management module 140 is used to receive a charging input from a charger. Among them, the charger can be a wireless charger or a wired charger. In an embodiment of wired charging, the charging management module 140 can receive the charging input of the wired charger through the USB interface 130. In an embodiment of wireless charging, the charging management module 140 can receive the wireless charging input through the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device through the power management module 141.
[0091] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives the inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the display screen 194, the camera 193, the wireless communication module 160, etc. The power management module 141 can also be used to monitor parameters such as the battery capacity, the number of battery charge cycles, and the battery health status (leakage, impedance). In one embodiment, the power management module 141 can also be disposed in the processor 110. In another embodiment, the power management module 141 and the charging management module 140 can also be disposed in the same device.
[0092] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.
[0093] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example: The antenna 1 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0094] The mobile communication module 150 may provide a solution for wireless communication including 2G / 3G / 4G / 5G, etc., which is applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 may receive electromagnetic waves through the antenna 1, filter and amplify the received electromagnetic waves, and then transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 may also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through the antenna 1 for radiation. In one embodiment, at least some functional modules of the mobile communication module 150 may be provided in the processor 110. In one embodiment, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be provided in the same device.
[0095] The modulation and demodulation processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, receiver 170B, etc.), or displays an image or video through the display screen 194. In one embodiment, the modulation and demodulation processor may be an independent device. In some other embodiments, the modulation and demodulation processor may be independent of the processor 110 and be provided in the same device as the mobile communication module 150 or other functional modules.
[0096] The wireless communication module 160 may provide solutions for wireless communications applied to the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The wireless communication module 160 may be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 may also receive signals to be sent from the processor 110, perform frequency modulation and amplification on them, and convert them into electromagnetic waves through the antenna 2 for radiation.
[0097] In one embodiment, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, such that electronic device 100 can communicate with a network and other devices through wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS), and / or satellite based augmentation systems (SBAS).
[0098] Electronic device 100 implements a display function through a GPU, display screen 194, and an application processor, etc. The GPU is a microprocessor for image processing, and is connected to display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.
[0099] The display screen 194 is used to display images, videos, etc. Among them, the display screen 194 includes a display panel. Specifically, the display screen can include a foldable screen, a special-shaped screen, etc. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In one embodiment, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.
[0100] The electronic device 100 can implement the shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor, etc.
[0101] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and light passes through the lens and is transmitted to the camera sensor. The light signal is converted into an electrical signal, and the camera sensor transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In one embodiment, the ISP can be set in the camera 193.
[0102] The camera 193 is used to capture static images or videos. An object generates an optical image through the lens and projects it onto the sensor. The sensor can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The sensor converts the light signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV, etc. format. In one embodiment, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0103] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0104] The video codec is used to compress or decompress digital videos. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple coding formats, such as: Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0105] The NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission pattern between human brain neurons, it can quickly process input information and can also continuously learn by itself. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.
[0106] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to achieve the data storage function. For example, files such as music and videos are saved in the external memory card.
[0107] The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.). The data storage area can store the data created during the use of the electronic device 100 (such as audio data, phone book, etc.). In addition, the internal memory 121 can include high-speed random access memory and can also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121 and / or the instructions stored in the memory provided in the processor.
[0108] The electronic device 100 can implement audio functions through the audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor, etc. For example, music playback, recording, etc.
[0109] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used for encoding and decoding audio signals. In one embodiment, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.
[0110] The speaker 170A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or hands-free calls through the speaker 170A.
[0111] The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device 100 answers a call or a voice message, the voice can be listened to by placing the receiver 170B close to the human ear.
[0112] The microphone 170C, also known as the "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak by placing the mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In some other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and implement functions such as directional recording.
[0113] The headphone jack 170D is used to connect a wired headphone. The headphone jack 170D can be a USB interface 130, or a 3.5 mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0114] The pressure sensor 180A is used to sense pressure signals and can convert pressure signals into electrical signals. In one embodiment, the pressure sensor 180A can be disposed on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor can include at least two parallel plates having conductive materials. When a force acts on the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device 100 determines the intensity of the pressure based on the change in capacitance. When a touch operation acts on the display screen 194, the electronic device 100 detects the intensity of the touch operation according to the pressure sensor 180A. The electronic device 100 can also calculate the position of the touch based on the detection signal of the pressure sensor 180A. In one embodiment, touch operations acting on the same touch position but with different touch operation intensities can correspond to different operation instructions. For example: When a touch operation with a touch operation intensity less than the first pressure threshold acts on the short message application icon, the instruction to view the short message is executed. When a touch operation with a touch operation intensity greater than or equal to the first pressure threshold acts on the short message application icon, the instruction to create a new short message is executed.
[0115] The gyroscope sensor 180B can be used to determine the motion posture of the electronic device 100. In one embodiment, the angular velocity of the electronic device 100 around three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 180B. The gyroscope sensor 180B can be used for anti-shake shooting. Exemplarily, when the shutter is pressed, the gyroscope sensor 180B detects the angle of jitter of the electronic device 100, calculates the distance that the lens module needs to compensate according to the angle, and enables the lens to offset the jitter of the electronic device 100 through reverse movement to achieve anti-shake. The gyroscope sensor 180B can also be used for navigation and somatosensory game scenarios.
[0116] The barometric pressure sensor 180C is used to measure barometric pressure. In one embodiment, the electronic device 100 calculates the altitude based on the barometric pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.
[0117] The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip leather case. In one embodiment, when the electronic device 100 is a flip phone, the electronic device 100 can detect the opening and closing of the flip according to the magnetic sensor 180D. Furthermore, according to the detected opening and closing state of the leather case or the opening and closing state of the flip, features such as automatic flip unlocking are set.
[0118] The acceleration sensor 180E can detect the magnitude of the acceleration of the electronic device 100 in various directions (generally three axes). When the electronic device 100 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers.
[0119] A distance sensor 180F for measuring distance. The electronic device 100 can measure distance through infrared or laser. In one embodiment, when shooting a scene, the electronic device 100 can use the distance sensor 180F to measure distance to achieve fast focusing.
[0120] The proximity light sensor 180G may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The light emitting diode may be an infrared light emitting diode. The electronic device 100 emits infrared light outward through the light emitting diode. The electronic device 100 uses the photodiode to detect the infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 can determine that there is no object near the electronic device 100. The electronic device 100 can use the proximity light sensor 180G to detect that the user holds the electronic device 100 close to the ear for a call, so as to automatically turn off the screen to save power. The proximity light sensor 180G can also be used for automatic unlocking and locking of the holster mode and pocket mode.
[0121] The ambient light sensor 180L is used to sense the ambient light brightness. The electronic device 100 can adaptively adjust the brightness of the display screen 194 according to the sensed ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 180L can also cooperate with the proximity light sensor 180G to detect whether the electronic device 100 is in the pocket to prevent accidental touch.
[0122] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can use the collected fingerprint characteristics to achieve fingerprint unlocking, access application locks, fingerprint taking pictures, fingerprint answering calls, etc.
[0123] The temperature sensor 180J is used to detect temperature. In one embodiment, the electronic device 100 uses the temperature detected by the temperature sensor 180J to execute a temperature processing strategy. For example, when the temperature reported by the temperature sensor 180J exceeds the threshold, the electronic device 100 reduces the performance of the processor near the temperature sensor 180J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is lower than another threshold, the electronic device 100 heats the battery 142 to avoid abnormal shutdown of the electronic device 100 caused by low temperature. In other embodiments, when the temperature is lower than yet another threshold, the electronic device 100 boosts the output voltage of the battery 142 to avoid abnormal shutdown caused by low temperature.
[0124] The touch sensor 180K, also known as the "touch control device". The touch sensor 180K can be disposed on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, also known as the "touch control screen". The touch sensor 180K is used to detect a touch operation acting thereon or nearby. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In some other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, at a different position from where the display screen 194 is located.
[0125] In one embodiment, the touch control screen formed by the touch sensor 180K and the display screen 194 can be located in the side area or the folding area of the electronic device 100, and is used to determine the position and gesture of the user's touch when the user's hand touches the touch control screen; for example, when the user holds the electronic device, the user can click on any position on the touch control screen with the thumb, then the touch sensor 180K can detect the user's click operation and transmit the click operation to the processor, and the processor determines that the click operation is used to wake up the screen according to the click operation.
[0126] The bone conduction sensor 180M can acquire vibration signals. In one embodiment, the bone conduction sensor 180M can acquire the vibration signals of the vibrating bone mass of the human vocal part. The bone conduction sensor 180M can also contact the human pulse to receive the blood pressure pulsation signal. In one embodiment, the bone conduction sensor 180M can also be disposed in the earphone to form a bone conduction earphone. The audio module 170 can analyze the voice signal based on the vibration signals of the vibrating bone mass of the human vocal part acquired by the bone conduction sensor 180M to implement the voice function. The application processor can analyze the heart rate information based on the blood pressure pulsation signal acquired by the bone conduction sensor 180M to implement the heart rate detection function.
[0127] The button 190 includes a power-on button, a volume button, etc. The button 190 can be a mechanical button. It can also be a touch button. The electronic device 100 can receive button inputs and generate key signal inputs related to the user settings and function controls of the electronic device 100.
[0128] The motor 191 can generate vibration prompts. The motor 191 can be used for incoming call vibration prompts and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, playing audio, etc.) can correspond to different vibration feedback effects. For touch operations acting on different areas of the display screen 194, the motor 191 can also correspond to different vibration feedback effects. Different application scenarios (such as: time reminder, receiving information, alarm clock, game, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.
[0129] The indicator 192 can be an indicator light, which can be used to indicate the charging status, power change, and can also be used to indicate messages, missed calls, notifications, etc.
[0130] The SIM card interface 195 is used to connect the SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation from the electronic device 100. The electronic device 100 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to implement functions such as calls and data communication. In one embodiment, the electronic device 100 uses an eSIM, that is, an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.
[0131] When the electronic device uses a special-shaped screen or a folding screen, the touch display screen of the electronic device can include multiple touch display areas. For example, the folding area of the folding screen of the electronic device includes a folding area in the folded state, and this folding area can also achieve touch response. However, in the prior art, the operation of the electronic device on a specific touch display area has great limitations, and there are no related operations specifically for a specific touch display area. Based on this, the embodiment of the present application provides a gesture interaction method. In this gesture interaction method, there are touch response areas in the side area or folding area of the electronic device. The electronic device can obtain the input events of the touch response area and, in response to the input events, trigger the electronic device to execute the operation instructions corresponding to the input events to implement gesture operations on the side area or folding area of the electronic device and improve the control experience of the electronic device.
[0132] In the electronic device disclosed in Embodiment 1 of the present application, the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the electronic device executes the method described in Embodiment 2 of the present application.
[0133] Embodiment 2
[0134] A speech recognition method provided in Embodiment 2 of the present application combines ultrasonic movement detection to improve the ASR accuracy rate, so as to solve the problem of the decrease in the ASR accuracy rate caused by the change of the speaker's position after the electronic device is awakened and enters the working state. In Embodiment 2, the electronic device adds an ultrasonic detection module with an ultrasonic frequency of 20KHz. If the movement of the speaker is not detected according to the Doppler effect of ultrasonic waves, the original ASR process is used for the speech recognition of the speaker. If the movement of the speaker is detected, the relative position of the speaker is detected according to the Doppler effect of ultrasonic waves. In the application scenarios of multiple electronic devices, two or more devices are used as multiple ultrasonic detection points to more accurately locate the azimuth information of the speaker. Finally, the prior information of the speaker's azimuth is added to improve the accuracy of sound processing in the electronic device, thereby improving the ASR recognition rate to a certain extent without significantly increasing the amount of calculation.
[0135] As Figure 2 shown, the speech recognition method of Embodiment 2 includes the following steps:
[0136] Step100: Receive the wake-up signal of the sound source and obtain the first direction of the wake-up signal.
[0137] Step200: Capture the movement state of the sound source through ultrasonic detection and obtain the second direction after the movement of the sound source.
[0138] Step300: Determine the azimuth of the sound source through ultrasonic detection according to the first direction and the second direction, and perform speech recognition on the sound source.
[0139] The speech recognition method of Embodiment 2 detects the movement of the speaker at the sound source according to ultrasonic waves, determines the azimuth by determining the position of the speaker's movement, and improves the ASR recognition rate to a certain extent without significantly increasing the amount of calculation.
[0140] As Figure 3 shown, in the speech recognition method of Embodiment 2, between step Step200 "Capture the movement state of the sound source through ultrasonic detection and obtain the second direction after the movement of the sound source" and step Step300 "Determine the azimuth of the sound source through ultrasonic detection according to the first direction and the second direction, and perform speech recognition on the sound source", the speech recognition method of Embodiment 2 further includes the following steps:
[0141] Step200’: Use multiple electronic devices to obtain the azimuth of the sound source relative to each electronic device through ultrasonic detection.
[0142] The speech recognition method of Embodiment 2 forms multiple ultrasonic detection points in the application scenarios of multiple electronic devices to more accurately locate the azimuth information of the sound source.
[0143] AsFigure 4 As shown, in the voice recognition method of this Embodiment 2, in step Step200’ “Use multiple electronic devices to obtain the orientation of the sound source relative to each electronic device through ultrasonic detection”, the following steps are further included:
[0144] Step201’: Measure the first distance between two electronic devices through ultrasonic detection.
[0145] Step202’: According to the second direction, measure the second distance and the third distance between the sound source and the two electronic devices respectively through ultrasonic detection.
[0146] Step203’: Calculate the orientation of the sound source relative to the two electronic devices according to the first distance, the second distance and the third distance.
[0147] Step204’: Repeat measuring the first distance, the second distance and the third distance, and calculate the orientation of the sound source relative to each electronic device.
[0148] The voice recognition method of this Embodiment 2 uses the cosine theorem to determine the orientation of the sound source relative to each electronic device, with fewer terms required for calculation and easy to obtain. While increasing the accuracy of sound source orientation determination, the required computational amount is small.
[0149] As Figure 5 shown, in the voice recognition method of this Embodiment 2, in step Step100 “Receive the wake-up signal of the sound source and obtain the first direction of the wake-up signal”, the following steps are further included:
[0150] Step101: Detect the sound source in real time.
[0151] Step102: Capture the wake-up signal that triggers the wake-up event in the sound source.
[0152] Step103: Obtain the first direction of the wake-up signal.
[0153] In the voice recognition method of this Embodiment 2, only the wake-up signal that triggers the wake-up event will activate the electronic device and the voice recognition function, and other non-wake-up signals can be excluded at this step, playing a role in signal filtering.
[0154] As Figure 6 shown, in the voice recognition method of this Embodiment 2, in step Step200 “Capture the motion state of the sound source through ultrasonic detection and obtain the second direction after the sound source moves”, the following steps are further included:
[0155] Step201: Send a first ultrasonic signal in the first direction.
[0156] Step202: The first receiving unit receives the first echo signal of the first ultrasonic signal.
[0157] Step203: Calculate the frequency shift between the first ultrasonic signal and the first echo signal.
[0158] Step204: Determine the motion state of the sound source according to the frequency shift.
[0159] Step205: When the motion state of the sound source is moving, the second receiving unit at a first distance from the first receiving unit receives the second echo signal of the first ultrasonic signal with a first wavelength.
[0160] Step206: Calculate the phase difference between the first echo signal and the second echo signal.
[0161] Step207: Calculate the azimuth angle of the sound source relative to the first receiving unit and the second receiving unit according to the first distance, the first wavelength and the phase difference.
[0162] Wherein, the first distance is perpendicular to the first direction.
[0163] For the voice recognition method of this Embodiment 2, by using the phase differences of different echo signals of the same ultrasonic signal at different positions and the differences in the receiving point positions, the azimuth angle of the sound source relative to different receiving points can be calculated more accurately.
[0164] Such as Figure 7 shown, in the voice recognition method of this Embodiment 2, in step Step300 "Determine the azimuth of the sound source through ultrasonic detection according to the first direction and the second direction, and perform voice recognition on the sound source", the following steps are further included:
[0165] Step301: Send ultrasonic waves in the first direction to obtain a first beam.
[0166] Step302: Send ultrasonic waves in the second direction to obtain a second beam.
[0167] Step303: Discriminate the first beam and the second beam to confirm the azimuth of the sound source.
[0168] Step304: Perform voice recognition on the sound source.
[0169] For the voice recognition method of this Embodiment 2, by using the first beam and the second beam and discriminating various parameters in two different scenarios with and without a sound source, the direction corresponding to the beam close to or being the human voice can be confirmed, thereby facilitating voice recognition and identifying the intention of the speaker.
[0170] In addition, in Figure 6In [the method], when performing step Step204, if the motion state of the sound source is stationary, that is, the sound source does not move, it directly jumps to step Step304 to perform speech recognition on the sound source in the first direction.
[0171] See Figures 8 to 10 , taking a smart speaker product as an example of the electronic device 100 below, the speech recognition method of this Embodiment 2 is described.
[0172] When the electronic device 100 is in a non-awakened state, the electronic device 100 will perform step Step101 to detect the sound source in real time, capture the wake-up signal through step Step102. Once the wake-up signal triggering the wake-up event is captured, it will perform step Step103 to obtain the first direction of the wake-up signal. Then it performs step Step201 to send a first ultrasonic signal of 20KHz in the first direction for detection, and the microphone array (first receiving unit) in the electronic device 100 performs step Step202 to receive the first echo signal of the first ultrasonic signal. By measuring the periodicity of the frequency shift Δf through echo autocorrelation, it determines whether there is human movement at the sound source, and performs step Step203. The frequency shift Δf is calculated by the frequency f of the sound source detected by the first ultrasonic signal and the echo frequency F of the first echo signal, and the method is Δf = F - f; then according to the result of whether there is human movement at the sound source determined by the frequency shift Δf, it performs step Step204 to determine the motion state of the sound source. If the motion state is stationary, it directly performs step Step304 to perform speech recognition on the sound source. If the motion state is moving, it then performs step Step205 and uses the second echo signal collected by the microphone (second receiving unit) perpendicular to the first direction to participate in the calculation. The wavelength of the first ultrasonic signal is λ, and the two microphones are separated by a first distance d. By performing step Step206, the phase difference between the first echo signal and the second echo signal is calculated. Then, by performing step Step207, the azimuth angle θ of the sound source relative to the two microphones is calculated. The ultrasonic angle measurement method with two microphones as a group of ultrasonic receivers is See Figure 9 , the two echo signals in the target direction are received by the ultrasonic receiver composed of two microphones separated by a first distance, and the azimuth angle of the sound source in the target direction is measured by the ultrasonic angle measurement method. According to the azimuth angle of the speaker at the sound source obtained by the above steps, through multiple electronic devices 100 (here are multiple speakers, such as Figure 10The ultrasonic unit (as shown) measures the distance between the speaker and the electronic device 100 through the ultrasonic ranging method, and obtains the azimuth of the speaker relative to each electronic device 100 through the cosine theorem. Among them, the ultrasonic ranging method is D = vt / 2, where D is the speaker distance, t is the echo time difference between the transmitted ultrasonic signal and the echo signal, and v is the wave speed. Execute step Step201' to obtain the first distance a between the two electronic devices 100, and execute step Step202' to obtain the second distance b and the third distance c between the sound source and the two electronic devices 100. Combining Figure 10 , by executing step Step203', substitute the first distance a, the second distance b, and the third distance c into the ultrasonic ranging formula and the cosine theorem respectively. According to the three formulas cosC = (a^2 + b^2 - c^2) / (2·a·b), cosB = (a^2 + c^2 - b^2) / (2·a·c), and cosA = (c^2 + b^2 - a^2) / (2·b·c), obtain the azimuth A of the sound source relative to the two electronic devices 100. Execute step Step204' to repeat steps Step201', Step202', and Step203' until all azimuths A are obtained. After determining the azimuth of the sound source, it is necessary to perform the ASR process on the speaker at the sound source. Through step Step301, the microphone array audio preprocessing algorithm of the electronic device 100 obtains the first beam in the first direction. Through step Step302, when the microphone array of the electronic device 100 picks up sound, it adds a second beam in the azimuth of the speaker at the sound source estimated by ultrasound. Through step Step303, the first beam and the second beam are discriminated in terms of amplitude, signal-to-noise ratio, zero-crossing rate parameters, and the probability of being speech, etc., to confirm the beam that is more likely to be human voice. Finally, through step Step304, the microphone array audio preprocessing algorithm of the electronic device 100 picks up sound in the corrected direction of the speaker at the sound source, and then performs ASR to identify the speaker's intention.
[0173] Embodiment 3
[0174] As Figure 11 shown is a voice recognition device provided in Embodiment 3 of the present application. The device includes: an ultrasonic transceiver module 10, a processing module 20, and an identification module 30 that communicate with each other. Among them, the ultrasonic transceiver module 10 is used to receive the wake-up signal of the sound source and obtain the first direction of the wake-up signal; the processing module 20 is used to capture the motion state of the sound source through ultrasonic detection and obtain the second direction after the sound source moves; the identification module 30 is used to determine the azimuth of the sound source through ultrasonic detection according to the first direction and the second direction, and perform voice recognition on the sound source.
[0175] The voice recognition device of this Embodiment 3 detects the movement of the speaker at the sound source according to ultrasonic waves, determines the orientation by determining the position of the speaker's movement, and improves the ASR recognition rate to a certain extent without significantly increasing the computational complexity.
[0176] As Figure 12 shown, in the voice recognition device of this Embodiment 3, the ultrasonic transceiver module 10 includes an ultrasonic unit 11 and at least two receiving units 12. The ultrasonic unit 11 is used to send ultrasonic signals, and the receiving unit 12 is used to receive wake-up signals and echo signals of the ultrasonic signals.
[0177] The voice recognition device of this Embodiment 3 uses multiple receiving units 12 to receive echo signals, which is beneficial to more accurately calculate the orientation of the sound source and improve the ASR recognition rate.
[0178] Embodiment 4
[0179] Embodiment 4 of this application provides a computer-readable storage medium, including programs or instructions. When the programs or instructions run on a computer, the voice recognition method disclosed in Embodiment 2 of this application is executed.
[0180] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a Digital Video Disc (DVD)), or a semiconductor medium (such as a Solid State Disk (SSD)), etc.
[0181] The voice recognition method, device, electronic device, and storage medium disclosed in the embodiments of the present application can effectively identify the position of the sound source in a scenario where the position of the sound source changes, accurately locate the sound source, confirm the orientation of the sound source relative to the electronic device, and thus better perform voice recognition and improve the recognition rate of ASR.
[0182] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0183] The above are only the preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application shall be included within the scope of protection of this application.
Claims
1. A voice recognition method, characterized in that, The method includes: Receiving a wake-up signal of a sound source and obtaining a first direction of the wake-up signal; Capturing a motion state of the sound source through ultrasonic detection and obtaining a second direction after the sound source moves; Determining the orientation of the sound source through ultrasonic detection according to the first direction and the second direction, and performing speech recognition on the sound source; In the step of capturing the motion state of the sound source through ultrasonic detection and obtaining a second direction after the sound source moves, the following steps are further included: Sending a first ultrasonic signal in the first direction; A first receiving unit receives a first echo signal of the first ultrasonic signal; Calculating a frequency shift between the first ultrasonic signal and the first echo signal; Determining the motion state of the sound source according to the frequency shift; When the motion state of the sound source is moving, a second receiving unit at a first distance from the first receiving unit receives a second echo signal of the first ultrasonic signal having a first wavelength; Calculating a phase difference between the first echo signal and the second echo signal; Calculating an azimuth angle of the sound source relative to the first receiving unit and the second receiving unit according to the first distance, the first wavelength and the phase difference; Wherein, the first distance is perpendicular to the first direction.
2. The voice recognition method according to claim 1, wherein, Between the step of capturing the motion state of the sound source through ultrasonic detection and obtaining a second direction after the sound source moves, and the step of determining the orientation of the sound source through ultrasonic detection according to the first direction and the second direction and performing speech recognition on the sound source, the speech recognition method further includes the following steps: Using multiple electronic devices to obtain the orientation of the sound source relative to each of the electronic devices through ultrasonic detection.
3. The speech recognition method according to claim 2, characterized in that, In the step of using multiple electronic devices to obtain the orientation of the sound source relative to each of the electronic devices through ultrasonic detection, the following steps are further included: Measuring a first distance between two electronic devices through ultrasonic detection; According to the second direction, respectively measuring a second distance and a third distance between the sound source and the two electronic devices through ultrasonic detection; Calculating the orientation of the sound source relative to the two electronic devices according to the first distance, the second distance and the third distance; Repeatedly measuring the first distance, the second distance and the third distance, and calculating the orientation of the sound source relative to each of the electronic devices.
4. The voice recognition method according to claim 1, wherein In the step of receiving a wake-up signal of a sound source and obtaining a first direction of the wake-up signal, the following steps are further included: Real-time detecting the sound source; Capturing a wake-up signal that triggers a wake-up event in the sound source; Obtaining a first direction of the wake-up signal.
5. The voice recognition method according to claim 1, wherein In the step of determining the orientation of the sound source through ultrasonic detection according to the first direction and the second direction and performing speech recognition on the sound source, the following steps are further included: Sending ultrasonic waves in the first direction to obtain a first beam; Sending ultrasonic waves in the second direction to obtain a second beam; Discriminating the first beam and the second beam to confirm the orientation of the sound source; Performing speech recognition on the sound source.
6. A voice recognition device, characterized in that, The device includes: an ultrasonic transceiver module, a processing module and an identification module that communicate with each other; The ultrasonic transceiver module is used to receive the wake-up signal of the sound source and obtain the first direction of the wake-up signal; The processing module is used to capture the motion state of the sound source through ultrasonic detection and obtain the second direction after the sound source moves; The recognition module is used to determine the orientation of the sound source through ultrasonic detection according to the first direction and the second direction, and perform voice recognition on the sound source; The ultrasonic transceiver module includes an ultrasonic unit, a first receiving unit and a second receiving unit. In the step of capturing the motion state of the sound source through ultrasonic detection and obtaining the second direction after the sound source moves, the following steps are further included: Sending a first ultrasonic signal in the first direction through the ultrasonic unit; Receiving a first echo signal of the first ultrasonic signal through the first receiving unit; Calculating the frequency shift between the first ultrasonic signal and the first echo signal; Determining the motion state of the sound source according to the frequency shift; When the motion state of the sound source is moving, receiving a second echo signal of the first ultrasonic signal with a first wavelength through the second receiving unit that is at a first distance from the first receiving unit; Calculating the phase difference between the first echo signal and the second echo signal; Calculating the azimuth angle of the sound source relative to the first receiving unit and the second receiving unit according to the first distance, the first wavelength and the phase difference; Wherein, the first distance is perpendicular to the first direction.
7. An electronic device, characterized in that, Comprising: A memory and a processor: The memory is used to store a computer program; The processor is used to execute the computer program stored in the memory, so that the electronic device executes the method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, Including a program or instruction, when the program or instruction runs on a computer, the method according to any one of claims 1 to 5 is executed.
Citation Information
Patent Citations
Method and system for improving remote field voice recognition rate, and readable storage medium
CN110085258A