Speech recognition method and electronic device
By detecting the voice environment and adopting an adaptive speech recognition method, the problem of voice assistants having difficulty recognizing in noisy environments has been solved, achieving efficient speech recognition and interaction in different environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-31
- Publication Date
- 2026-04-14
AI Technical Summary
Existing voice assistants cannot effectively recognize user input in noisy environments, resulting in poor voice recognition performance and an inability to provide effective voice interaction services.
By detecting the voice environment when the user inputs audio, the voice interaction mode is determined, and corresponding voice recognition methods are used for audio preprocessing and recognition, including filtering or retaining interfering audio in different environments and using specific voice recognition models to improve recognition accuracy.
It improves the accuracy of speech recognition, ensuring effective recognition of user speech in noisy or quiet environments, thus enhancing the user experience.
Smart Images

Figure CN120412566B_ABST
Abstract
Description
[0001] This application is a divisional application. The original application has the application number 202411219770.5 and the original application date is August 31, 2024. The entire contents of the original application are incorporated herein by reference. Technical Field
[0002] This application relates to the field of terminal technology, and in particular to a speech recognition method and electronic device. Background Technology
[0003] A voice assistant is a voice interaction function in electronic devices. This function can collect the user's voice input, perform speech recognition, and then respond via voice. Currently, voice assistant functions can only recognize normal voice input. In some scenarios, voice assistants cannot effectively recognize speech, resulting in poor recognition quality and thus failing to provide voice interaction services to users. Summary of the Invention
[0004] This application provides a speech recognition method and electronic device for improving the accuracy of speech recognition.
[0005] In a first aspect, this application provides a speech recognition method that can be executed by an electronic device. In this method, the electronic device acquires audio input by a user; the electronic device detects the speech environment when the user inputs the audio based on the audio, and determines the speech interaction mode corresponding to the audio; the electronic device performs speech recognition on the audio according to the speech recognition method corresponding to the speech interaction mode, and obtains a speech recognition result.
[0006] In the above methods, electronic devices can detect the environment when a user inputs audio, thereby determining whether the user is in a noisy or quiet voice environment, and then determining the voice interaction mode. The electronic devices can then perform corresponding voice recognition on the acquired audio through the voice recognition method corresponding to the voice interaction mode, thereby improving the accuracy of voice recognition.
[0007] In one possible design, the electronic device supports multiple voice interaction modes, each corresponding to a different voice recognition method. This design allows the electronic device to perform different voice recognition methods on audio acquired under different voice interaction modes, thereby employing a more suitable approach to the user's current voice environment and improving voice recognition accuracy.
[0008] In one possible design, detecting the voice environment when the user inputs the audio and determining the corresponding voice interaction mode includes: determining the voice interaction mode corresponding to the audio as a first mode when the characteristics of the audio satisfy a first condition; wherein the first condition includes at least one of the following: the distance between the sound source of the audio and the microphone of the electronic device is less than a first distance threshold; the audio characteristics of the user's audio include the vocal cords not producing sound; and the ambient noise intensity in the audio is less than a third sound intensity threshold. With this design, the electronic device can determine that the user is currently in a relatively quiet environment based on the distance between the sound source of the audio and the microphone of the electronic device, the audio characteristics, and the ambient noise intensity. The user may speak in a breathy voice to avoid disturbing others, ensuring that the determined voice interaction mode matches the user's actual voice environment.
[0009] In one possible design, detecting the voice environment when the user inputs the audio and determining the corresponding voice interaction mode includes: determining the voice interaction mode corresponding to the audio as a second mode when the characteristics of the audio satisfy a second condition; wherein the second condition includes at least one of the following: the distance between the audio source and the microphone of the electronic device is less than a first distance threshold; the user's voice intensity in the audio is greater than a first sound intensity threshold; and the ambient noise intensity in the audio is greater than a second sound intensity threshold. With this design, when the electronic device determines that the user may be in a noisy environment based on the distance between the audio source and the microphone of the electronic device, the user's voice intensity, and the ambient noise intensity, the user can bring the microphone of the electronic device closer to their mouth to input audio, ensuring that the determined voice interaction mode matches the user's actual voice environment.
[0010] In one possible design, the step of performing speech recognition on the audio according to the speech recognition method corresponding to the voice interaction mode to obtain a speech recognition result includes: performing audio preprocessing on the audio according to the voice interaction mode; and performing speech recognition on the preprocessed audio data according to the speech recognition method corresponding to the voice interaction mode to obtain the speech recognition result. Through this design, the electronic device can perform audio preprocessing on the acquired audio before performing speech recognition on the preprocessed audio data, further improving the accuracy of speech recognition.
[0011] In one possible design, when the voice interaction mode is in the first mode, the audio preprocessing based on the voice interaction mode includes filtering out other voices besides the user's voice in the audio, but not filtering out breath sounds in the audio. With this design, even when the voice interaction mode is in the first mode, the user may use breath sounds to input audio; therefore, the electronic device does not filter out breath sounds during audio preprocessing, thus preserving as much of the audio input by the user through breath sounds as possible, further improving the accuracy of voice recognition.
[0012] In one possible design, when the voice interaction mode is the second mode, the audio preprocessing based on the voice interaction mode includes filtering out other voices besides the user's voice and breath sounds in the audio. Through this design, the electronic device can filter out other voices and breath sounds in the audio that may interfere with speech recognition, thereby preserving as much of the user's own audio as possible and further improving the accuracy of speech recognition.
[0013] In one possible design, performing speech recognition on the pre-processed audio data according to the speech recognition method corresponding to the voice interaction mode includes: performing speech recognition on the pre-processed audio data using a speech recognition model corresponding to the voice interaction mode, wherein the speech recognition models corresponding to different voice interaction modes are models trained based on training sample sets corresponding to different voice interaction modes. Through this design, electronic devices can use different speech recognition models for different voice interaction modes. Since different speech recognition models are trained based on training sample sets corresponding to different voice interaction modes, the speech recognition model corresponding to the voice interaction mode has a higher recognition accuracy for the audio input by the user under that voice interaction mode.
[0014] In one possible design, after obtaining the speech recognition result, the method further includes: determining the response content corresponding to the speech recognition result; determining a target playback method from multiple playback methods corresponding to the speech interaction mode based on information from an audio device, wherein the audio device is a device for playing the response content, and the information of the audio device includes at least one of the following: device type, device function, and connection method between the audio device and the electronic device; and playing the response content corresponding to the speech recognition result through the audio device according to the target playback method. With this design, the electronic device can support multiple speech playback methods for the same speech interaction mode. The electronic device can determine the target playback method based on the information of the audio device used for playing the speech, thereby flexibly using audio devices controlled by the electronic device to interact with the user via voice, improving the user experience.
[0015] In one possible design, when the voice interaction mode is the first mode and the audio device is a speaker within the electronic device or connected to the electronic device, the target playback method is a breathy playback method, the audio characteristics of which include no vocal cord output. When the voice interaction mode is the first mode and the audio device is an earphone connected to the electronic device, the target playback method is a default playback method, the audio characteristics of which include vocal cord output. With this design, if the audio device is a speaker within the electronic device or connected to the electronic device, the electronic device can also respond to the user using a breathy playback method, thus avoiding loud voice output disturbing the user in a quiet environment. If the audio device is an earphone connected to the electronic device, the electronic device can respond to the user using the default playback method through the earphone, thus ensuring audio playback quality when the user is wearing earphones.
[0016] In one possible design, acquiring the user-input audio includes controlling the electronic device's microphone to capture the audio input from the user located at a preset position and preset distance from the microphone. This design allows the electronic device to control the microphone to capture the sound emitted by the user at the preset position and preset distance, thereby effectively capturing the user's voice during the audio acquisition process.
[0017] In one possible design, the method further includes: determining that the user activates the voice interaction function of the electronic device via breath, wherein the audio is the audio input by the user when activating the voice interaction function of the electronic device via breath. With this design, since the user's voice environment when activating the voice assistant via breath may be a noisy public place or a quiet environment requiring whispering, the electronic device can detect the user's voice environment and then perform adaptive voice recognition, thereby improving the accuracy of voice recognition.
[0018] In one possible design, the audio is the audio input by the user when activating the voice interaction function of the electronic device, or the audio input by the user during voice interaction with the electronic device.
[0019] Secondly, this application provides a speech recognition method, which can be executed by an electronic device. In this method, the electronic device acquires audio input by a user; the electronic device detects the speech environment when the user inputs the audio based on the audio, and determines whether the speech interaction mode corresponding to the audio is a first mode, wherein the first mode is a speech interaction mode in which the user inputs audio using breathy sounds without vocal cords. When the speech interaction mode is the first mode, the audio is speech recognized using the speech recognition method corresponding to the first mode to obtain a speech recognition result; a target speech playback method is determined from multiple speech playback methods corresponding to the first mode, and the response content corresponding to the speech recognition result is played according to the target speech playback method.
[0020] In one possible design, before performing speech recognition on the audio using the speech recognition method corresponding to the whispering mode to obtain the speech recognition result, the audio is preprocessed according to the audio preprocessing method corresponding to the first mode. The audio preprocessing method corresponding to the first mode includes filtering out other voices in the audio besides the user's voice, but not filtering out breath sounds in the audio.
[0021] In one possible design, detecting the voice environment when the user inputs the audio based on the audio, and determining whether the voice interaction mode corresponding to the audio is a first mode, includes: determining whether the features of the audio satisfy a first condition based on the audio; and determining that the voice interaction mode corresponding to the audio is a first mode when the features of the audio satisfy the first condition; wherein the first condition includes at least one of the following: the distance between the sound source of the audio and the microphone of the electronic device is less than a first distance threshold; the audio features of the user's audio in the audio include the vocal cords not producing sound; and the ambient noise intensity in the audio is less than a first sound intensity threshold.
[0022] In one possible design, determining the target voice broadcast method from multiple voice broadcast methods corresponding to the first mode includes: determining the target broadcast method from multiple broadcast methods corresponding to the first mode based on information from an audio device, wherein the audio device is a device for broadcasting the reply content, and the information of the audio device includes at least one of the following: the device type of the audio device, the device function, and the connection method between the audio device and the electronic device.
[0023] In one possible design, when the audio device is a speaker within the electronic device or established in communication with the electronic device, the target broadcast mode is a breathy broadcast mode, the audio characteristics of which include no vocal cord output; when the audio device is an earphone established in communication with the electronic device, the target broadcast mode is a default broadcast mode, the audio characteristics of which include vocal cord output.
[0024] In one possible design, performing speech recognition on the audio using the speech recognition method corresponding to the first mode includes: performing speech recognition on the processed audio data using the speech recognition model corresponding to the first mode, wherein the speech recognition model corresponding to the first mode is a model trained based on the training sample set corresponding to the first mode.
[0025] In one possible design, the method further includes: detecting the voice environment when the user inputs the audio based on the audio, and determining whether the voice interaction mode corresponding to the audio is a second mode; when the voice interaction mode is the second mode, performing voice recognition on the audio using the voice recognition method corresponding to the second mode to obtain a voice recognition result; and broadcasting the response content corresponding to the voice recognition result.
[0026] In one possible design, before performing speech recognition on the audio using the speech recognition method corresponding to the second mode, the method further includes: performing audio preprocessing on the audio according to the audio preprocessing method corresponding to the second mode, wherein the audio preprocessing method corresponding to the second mode includes filtering out other voices in the audio besides the user's voice and breath sounds in the audio.
[0027] In one possible design, detecting the voice environment when the user inputs the audio based on the audio, and determining whether the voice interaction mode corresponding to the audio is a second mode, includes: determining whether the characteristics of the audio satisfy a second condition based on the audio; and determining that the voice interaction mode corresponding to the audio is a second mode when the characteristics of the audio satisfy the second condition; wherein the second condition includes at least one of the following: the distance between the sound source of the audio and the microphone of the electronic device is less than a first distance threshold; the user's voice intensity in the audio is greater than a second sound intensity threshold; and the ambient noise intensity in the audio is greater than a third sound intensity threshold.
[0028] In one possible design, the step of performing speech recognition on the audio using the speech recognition method corresponding to the second mode includes: performing speech recognition on the processed audio data using the speech recognition model corresponding to the second mode, wherein the speech recognition model corresponding to the second mode is a model trained based on the training sample set corresponding to the second mode.
[0029] In one possible design, prior to acquiring the user's audio input, the method further includes determining that the user activates the voice interaction function of the electronic device by breathing.
[0030] Thirdly, this application provides an electronic device comprising multiple functional modules; the multiple functional modules interact to implement the methods performed by the electronic device in any of the above aspects and their respective embodiments. The multiple functional modules can be implemented based on software, hardware, or a combination of software and hardware, and the multiple functional modules can be arbitrarily combined or divided based on specific implementations.
[0031] Fourthly, this application provides an electronic device including at least one processor and at least one memory, wherein the at least one memory stores computer program instructions, and when the electronic device is running, the at least one processor executes any of the above aspects and the methods executed by the electronic device in its various embodiments.
[0032] Fifthly, this application also provides a computer program product containing instructions that, when the computer program product is run on a computer, cause the computer to perform the method executed by the server or electronic device in any of the above aspects and their embodiments.
[0033] Sixthly, this application also provides a computer-readable storage medium storing a computer program that, when executed by a computer, causes the computer to perform the method executed by the server or electronic device in any of the above aspects and embodiments.
[0034] In a seventh aspect, this application also provides a chip for reading a computer program stored in a memory and executing the method performed by a server or electronic device in any of the above aspects and their embodiments.
[0035] Eighthly, this application also provides a chip system including a processor for supporting a computer device in implementing the methods performed by a server or electronic device in any of the above aspects and their embodiments. In one possible design, the chip system further includes a memory for storing programs and data necessary for the computer device. The chip system may be composed of chips or may include chips and other discrete devices. Attached Figure Description
[0036] Figure 1 A schematic diagram illustrating the applicable scenarios for the speech recognition method provided in the embodiments of this application;
[0037] Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0038] Figure 3 A software structure block diagram of an electronic device provided in an embodiment of this application;
[0039] Figure 4 A schematic diagram of the settings interface of a voice assistant provided in an embodiment of this application;
[0040] Figure 5 A schematic diagram of a breath-activated voice assistant provided in an embodiment of this application;
[0041] Figure 6 A flowchart illustrating a speech recognition method provided in an embodiment of this application;
[0042] Figure 7 A flowchart illustrating a speech recognition method provided in an embodiment of this application;
[0043] Figure 8 A flowchart of a speech recognition method provided in an embodiment of this application. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the embodiments of this application will be further described in detail below with reference to the accompanying drawings. In the description of the embodiments of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include one or more of that feature.
[0045] It should be understood that in the embodiments of this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, a and b, a and c, b and c, or a, b, and c, where a, b, and c can be single or multiple.
[0046] A voice assistant, also known as a smart assistant, is a voice interaction function in electronic devices. This function collects the user's voice input, performs speech recognition, and then responds to the user via voice based on the recognition results. Voice assistants can provide users with services such as voice dialogue, image and text recognition, service suggestions, and device interconnection management. Optionally, a voice assistant can be a system-level service within an electronic device, or it can be an application installed on the electronic device.
[0047] Current voice assistant functions can only recognize normal voice input. However, when users are in special environments such as noisy public places and input voice, the voice assistant cannot effectively obtain the user's audio input and respond, thus failing to provide voice interaction services.
[0048] Based on the above problems, this application provides a speech recognition method. Figure 1 This is a schematic diagram illustrating the applicable scenarios for the speech recognition method provided in the embodiments of this application. (Refer to...) Figure 1 The scenario includes electronic devices. In the speech recognition method provided in this application embodiment, the electronic device acquires the audio input by the user, detects the speech environment when the user inputs the audio based on the audio input, determines the speech interaction mode corresponding to the audio, and performs speech recognition on the audio according to the speech recognition method corresponding to the speech interaction mode to obtain the speech recognition result.
[0049] Optionally, Figure 1 The scenario shown can also include a server. After the electronic device obtains the user's audio input and determines the voice interaction mode, it can send the user's audio and the identifier of the voice interaction mode to the server. The server performs voice recognition according to the voice recognition method corresponding to the voice interaction mode, obtains the voice recognition result, and sends the voice recognition result to the electronic device.
[0050] The speech recognition method provided in this application allows electronic devices to detect the environment when a user inputs audio, thereby determining whether the user is in a noisy or quiet environment, and then determining the speech interaction mode. The electronic device can then perform speech recognition on the acquired audio using the speech recognition method corresponding to the speech interaction mode, thereby improving the accuracy of speech recognition.
[0051] The following describes an electronic device and embodiments for using such an electronic device. The electronic device in this application embodiment can be a tablet computer, mobile phone, in-vehicle device, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), wearable device, etc. This application embodiment does not limit the specific type of electronic device.
[0052] In some embodiments of this application, the electronic device may also be a portable terminal device that includes other functions such as a personal digital assistant and / or a music player. Exemplary embodiments of the portable terminal device include, but are not limited to, devices equipped with... Or portable terminal devices with other operating systems.
[0053] Figure 2 This is a schematic diagram of the structure of an electronic device 100 provided in an embodiment of this application. Figure 2 As shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.
[0054] The processor 110 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU).
[0055] The display screen 194 is used to display the display interface of an application, such as the display page of an application installed on the electronic device 100. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc.
[0056] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0057] The sensor module 180 may include a pressure sensor 180A, an acceleration sensor 180B, a touch sensor 180C, a gravity sensor 180D, a gyroscope sensor 180E, etc.
[0058] Gravity sensor 180D is used to measure gravity, and gyroscope sensor 180E is used to measure the rotation angle of electronic device. In this embodiment, the posture of the user holding the electronic device can be determined based on the data measured by acceleration sensor 180B, gravity sensor 180D and gyroscope sensor 180E.
[0059] Understandable, Figure 2 The components shown do not constitute a specific limitation on the electronic device 100. The electronic device may also include more or fewer components than shown, or combine some components, or separate some components, or have different component arrangements. Furthermore, Figure 2 The combination / connection relationships between the components can also be adjusted and modified.
[0060] Figure 3 This is a software structure block diagram of an electronic device provided in an embodiment of this application. For example... Figure 3As shown, the software architecture of an electronic device can be a layered architecture. For example, the software can be divided into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the operating system is divided into four layers, from top to bottom: the application layer, the application framework layer (framework, FWK), the runtime and system libraries, and the kernel layer.
[0061] The application layer can include a series of application packages. For example... Figure 3 As shown, the application layer can include a camera, settings, skin modules, user interface (UI), and third-party applications. Third-party applications can include gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, SMS, and voice assistants. The voice assistant provides users with voice interaction services. After activating the voice assistant, the user can input voice through the electronic device's microphone. The voice assistant can respond to the user's voice input and reply to the user via voice prompts.
[0062] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer can include some predefined functions. For example... Figure 3 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, and notification manager.
[0063] The window manager is used to manage windowed applications. It can obtain the screen size, determine if a status bar is present, lock the screen, and capture screenshots. The content provider stores and retrieves data, making this data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.
[0064] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0065] A phone manager is used to provide communication functions for electronic devices. For example, it manages call status (including connection and disconnection).
[0066] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0067] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.
[0068] The runtime includes the core libraries and the virtual machine. The runtime is responsible for the scheduling and management of the operating system.
[0069] The core library consists of two parts: one part contains the functionalities that the Java language needs to call, and the other part contains the core libraries of the operating system. The application layer and application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0070] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), image processing libraries, etc.
[0071] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.
[0072] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0073] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0074] A 2D graphics engine is a graphics engine for 2D drawing.
[0075] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.
[0076] The hardware layer can include various types of sensors, such as accelerometers, gyroscopes, and touch sensors.
[0077] It should be noted that, Figure 2 and Figure 3 The structure shown is merely an example of an electronic device provided in this application embodiment and is not intended to limit the electronic device provided in this application embodiment in any way. In specific implementations, the electronic device may have more than Figure 2 or Figure 3 The structure shown may contain more or fewer devices or modules.
[0078] The speech recognition method provided in the embodiments of this application is described below.
[0079] In the speech recognition method provided in this application embodiment, the electronic device has a speech interaction function, which can be referred to as a speech assistant, smart assistant, etc. In the following description, a speech assistant is used as an example. The speech assistant can be a system-level service in the electronic device, or it can be an application installed on the electronic device. Users can choose whether to enable the speech assistant. For example, the electronic device's settings interface may include a speech assistant settings option, allowing users to enable the speech assistant. When a user uses the speech assistant function of the electronic device, the user can input audio. The electronic device recognizes the audio input based on the speech interaction function and responds to the user based on the speech recognition result.
[0080] In this embodiment, the electronic device can detect the voice environment when the user inputs audio, determine the voice interaction mode, and perform audio preprocessing corresponding to the detected voice interaction mode. The electronic device then performs speech recognition corresponding to the voice interaction mode on the preprocessed audio data to obtain the speech recognition result.
[0081] Optionally, the voice interaction mode in the embodiments of this application may include a first mode, a second mode, or a third mode. The first mode may also be called a whisper mode, the second mode may also be called a whisper mode, and the third mode may also be called a normal mode. Of course, the various voice interaction modes involved in the embodiments of this application may also have other names. The embodiments of this application do not limit this. In the following description, the whisper mode, the whisper mode, and the normal mode are used as examples for illustration.
[0082] In this embodiment, the electronic device can detect the distance between the sound source and the microphone based on the acquired audio. For example, the electronic device can determine the distance between the user's mouth and the microphone. The electronic device can also determine the user's voice intensity and the ambient noise intensity in the audio. Optionally, the electronic device can also detect audio features, such as vocalization features. The electronic device can determine the voice interaction mode based on at least one of the following: the distance between the sound source and the microphone, the user's voice intensity, the ambient noise intensity, and audio features.
[0083] In some examples, electronic devices can determine that the voice interaction mode is whisper mode when the acquired user audio meets the following conditions:
[0084] 1. The distance between the sound source and the microphone is less than the first distance threshold.
[0085] 2. The user's voice intensity is greater than the second sound intensity threshold.
[0086] 3. The ambient noise intensity is greater than the third sound intensity threshold.
[0087] Understandably, when a user is in a noisy environment, they can bring the microphone of their electronic device close to their mouth to input audio, ensuring that the device can effectively capture the user's audio input. For example, the aforementioned first distance threshold could be 20 centimeters, or it could be 5 centimeters.
[0088] In other examples, electronic devices can determine that the voice interaction mode is whisper mode when the acquired user audio meets the following conditions:
[0089] 1. The distance between the sound source and the microphone is less than the first distance threshold.
[0090] 2. The audio characteristics of user audio include that the user's vocal cords do not produce sound.
[0091] 3. The ambient noise intensity is less than the first sound intensity threshold.
[0092] In this example, when the user is in a relatively quiet environment, the user may speak in a whisper to avoid disturbing others. In this case, the sound produced by the user is not produced by the vibration of the vocal cords, but by breath. The electronic device can detect the user's audio characteristics and then determine whether the current voice interaction mode is the whisper mode.
[0093] Optionally, for audio that does not meet the whisper mode and whisper mode described above, the electronic device may determine the voice interaction mode as normal mode.
[0094] In this embodiment, after determining the voice interaction mode, the electronic device can perform voice recognition on the audio using the voice recognition method corresponding to the voice interaction mode to obtain a voice recognition result. Optionally, when performing voice recognition on the audio, the electronic device can perform audio preprocessing corresponding to the voice interaction mode on the acquired audio, and then perform voice recognition on the preprocessed audio data using the voice recognition method corresponding to the voice interaction mode to obtain a voice recognition result.
[0095] In one optional implementation, the audio preprocessing corresponding to different voice interaction modes includes different processing methods. The electronic device can perform audio preprocessing corresponding to the voice interaction mode on the acquired audio to obtain clearer human voice. Optionally, the audio preprocessing corresponding to the normal mode includes noise reduction processing, which can filter out environmental noise other than human voice; the audio preprocessing corresponding to the whisper mode can include directional and fixed-distance sound pickup and noise reduction processing, wherein the directional and fixed-distance sound pickup can be used by the electronic device to control the microphone to collect the sound emitted by the user at a preset location and preset distance, thereby effectively collecting the user's human voice during the audio acquisition process; the noise reduction processing corresponding to the whisper mode can include filtering out other human voices and filtering out breath audio; the audio processing corresponding to the whisper mode can include directional and fixed-distance sound pickup and noise reduction processing, wherein the directional and fixed-distance sound pickup can be used by the electronic device to control the microphone to collect the sound emitted by the user at a preset location and preset distance, thereby effectively collecting the user's human voice during the audio acquisition process; the noise reduction processing corresponding to the whisper mode does not filter out breath audio, thereby preserving the user's breath audio.
[0096] After preprocessing the acquired user audio, the electronic device can perform speech recognition on the preprocessed audio data using the speech recognition method corresponding to the voice interaction mode, and obtain the speech recognition result. In this embodiment, the electronic device can perform speech recognition on the audio data using automatic speech recognition (ASR). For example, in this embodiment, the electronic device can store multiple speech recognition models corresponding to multiple voice interaction modes. For audio acquired under different voice interaction modes, the electronic device can perform speech recognition on the audio using the speech recognition model corresponding to the voice interaction mode, and obtain the speech recognition result output by the speech recognition model.
[0097] Optionally, the electronic device can also send the pre-processed audio data and the identifier of the voice interaction mode to the server. The server can store multiple voice recognition models corresponding to multiple voice interaction modes. The server performs voice recognition on the received audio data through the voice recognition model corresponding to the voice interaction mode, obtains the voice recognition result output by the voice recognition model, and the server can send the voice recognition result to the electronic device.
[0098] In some examples, different speech recognition models can be models trained on different training sample sets. For instance, the training sample set used when training the speech recognition model corresponding to the whispering mode may include high-noise frequencies, while the training sample set used when training the speech recognition model corresponding to the soft-spoken mode may include breath audio with less noise. For example, the speech recognition model in this application embodiment can be a neural network model. The model training process can be executed by a server, and after training is completed, the server can publish multiple speech recognition models to electronic devices. Alternatively, the electronic device can also execute the training process of multiple speech recognition models. This application embodiment does not limit this approach.
[0099] After receiving the speech recognition results, the electronic device can determine the content to reply to the user based on the speech recognition results and broadcast it to the user via voice. Optionally, the electronic device can determine the function of the voice assistant that the user wants to use based on the speech recognition results and provide interactive services to the user based on the function, such as providing voice dialogue, image and text recognition, service suggestions, device interconnection management and other services.
[0100] In some embodiments, among the multiple voice interaction modes supported by the electronic device, different voice interaction modes correspond to different voice playback methods, and the same voice interaction mode can correspond to at least one voice broadcasting method. When the electronic device replies to the user via voice based on the voice recognition result, it can determine the target broadcasting method from the multiple broadcasting methods corresponding to the voice interaction mode according to the information of the audio device. The audio device is a device used to broadcast the reply content. For example, the audio device can be a speaker inside the electronic device or a speaker connected to the electronic device, a speaker connected to the electronic device, an earphone connected to the electronic device, etc. The information of the audio device includes at least one of the following: device type, function, and connection method with the electronic device. The connection method with the electronic device includes Bluetooth connection, Starlink connection, wired connection, etc.
[0101] For example, when the voice interaction mode is in whisper mode, if the audio device is a speaker within the electronic device or connected to the electronic device, the electronic device can also respond to the user using a breathy tone. The audio characteristics of the breathy tone include no vocal cord output, thus avoiding disturbing the user in a quiet environment with a loud voice. If the audio device is a headset connected to the electronic device, the electronic device can respond to the user using a default tone through the headset. The audio characteristics of the default tone include vocal cord output, thus ensuring audio quality when the user is wearing the headset. In some examples, the volume of the default tone tone can be higher than that of the breathy tone tone.
[0102] In one alternative implementation, before acquiring the user-inputted audio, the electronic device may determine the voice interaction function that the user will activate, such as activating the device's voice assistant. When determining the voice interaction mode based on the audio, the audio used by the electronic device may be the audio input by the user when activating the device, or it may be the audio input by the user during voice interaction with the device.
[0103] Optionally, the electronic device can support multiple ways to wake up the voice assistant. Users can choose which wake-up methods to enable in the voice assistant's settings interface. For example, the voice assistant wake-up methods supported by the electronic device may include: voice wake-up, power button wake-up, Bluetooth device wake-up, desktop shortcut wake-up, breath wake-up, etc. Among them, voice wake-up allows the user to wake up the voice assistant after saying a preset wake-up word. Breath wake-up means that the electronic device can detect the user's breath near the microphone and wake up the voice assistant after detecting the user's breath. In the breath wake-up method, the user can say a wake-up word to wake up the voice assistant, or directly say a voice interaction command to wake up the voice assistant and perform voice interaction.
[0104] For example, Figure 4 A schematic diagram of the settings interface of a voice assistant provided in an embodiment of this application, with reference to... Figure 4 The voice assistant settings interface of an electronic device can include options for various wake-up methods supported by the electronic device, such as... Figure 4 The interface shown may include options for voice wake-up, power button wake-up, headphone wire wake-up, Bluetooth device wake-up, desktop shortcut, and breath wake-up. In the display area for the breath wake-up option, the electronic device can also display instructions to guide the user in using breath wake-up, such as... Figure 4 As shown, the electronic device can display text instructions in the display area corresponding to the breath-activated option: "Lift your phone and bring the bottom of the phone close to your mouth (about 5cm away) so that the phone can sense your breath and start a voice conversation." Optionally, the electronic device can also display operation diagrams, operation animations, and other multimedia content in the voice assistant's settings interface to guide users in using the breath-activated function. When using the electronic device, users can operate according to the relevant instructions for breath-activated activation, thereby activating the voice assistant of the electronic device through breath activation. Optionally, such as... Figure 4 As shown, the voice assistant's settings interface can also include voice settings options, such as selecting a female voice for playback, and allowing users to adjust the volume of the broadcast.
[0105] In one optional implementation, the electronic device can detect the user's posture while holding the device. When the electronic device detects that the user's posture matches a preset posture, it initiates the breath wake-up recognition process. Optionally, the electronic device can acquire posture data collected by devices such as gravity sensors, accelerometers, and gyroscopes within the device, and detect whether the user's posture matches a preset posture based on the posture data. The preset posture could be, for example, raising a hand or raising a wrist.
[0106] When an electronic device detects whether a user has woken up a voice assistant using breath, it can acquire the user's input audio through multiple microphones and determine whether the user has woken up the voice assistant using breath based on the acquired audio. For example, if the electronic device can detect the user's breath based on the acquired audio and the audio includes human voice with an intensity greater than a preset threshold, then the electronic device can determine that the user has woken up the voice assistant using breath.
[0107] For example, Figure 5 This is a schematic diagram of a breath-activated voice assistant provided as an embodiment of this application. (Reference) Figure 5 When a user raises their hand and brings the bottom of their phone close to their mouth, they input voice into the microphone of the electronic device. The electronic device can detect that the user is holding the device and raising their hand. The device then acquires the audio input through the microphone and performs breath detection on the audio. If the device determines that the user has activated the voice assistant through breath, it can then start the voice interaction function of the voice assistant.
[0108] In some examples, when an electronic device determines that a user has activated a voice assistant via breath, it can detect the speech environment of the user's input audio based on the speech recognition method provided in the above embodiments of this application, determine the speech interaction mode, and perform speech recognition based on the speech interaction mode to obtain the speech recognition result. In this way, since the user's speech environment when activating the voice assistant via breath may be a noisy public place or a quiet environment requiring whispering, the electronic device can detect the user's speech environment and then perform adaptive speech recognition, which can improve the accuracy of speech recognition.
[0109] Based on the above embodiments, Figure 6 A flowchart illustrating a speech recognition method provided in this application embodiment, which can be executed by an electronic device, is shown below. Figure 6 The method includes the following steps:
[0110] S601: The electronic device detects that the user's posture matches the preset posture.
[0111] Among them, the preset posture can be, for example, the user raising their hand or wrist after holding the electronic device.
[0112] S602: The electronic device acquires the audio input by the user and determines, based on the audio, that the user wakes up the voice assistant by breathing.
[0113] Optionally, breath wake-up means that the electronic device can detect the user's breath near the microphone and wake up the voice assistant after detecting the user's breath. The audio input by the user obtained by the electronic device can be the audio input by the user when waking up the voice assistant with breath.
[0114] It should be noted that S601-S602 are optional steps. In practice, after the electronic device obtains the audio input by the user, it can directly execute S603. At this time, the audio input by the user obtained by the electronic device can be the audio input by the user when waking up the voice assistant, or the audio input by the user during voice interaction with the electronic device.
[0115] S603: The electronic device detects the voice environment based on the audio input by the user and determines the voice interaction mode.
[0116] Based on the different results of the voice interaction mode detected by the electronic device, the electronic device can perform different voice recognition methods on the audio input by the user. For example, when the electronic device performs voice recognition, it can include any one of the following S604a, S604b and S604c:
[0117] S604a: When the electronic device determines that the voice interaction mode is normal mode, it performs voice recognition on the audio according to the voice recognition method corresponding to normal mode to obtain the voice recognition result.
[0118] Optionally, the speech recognition method corresponding to the normal mode may include: performing audio preprocessing on the audio according to the normal mode, and performing speech recognition on the preprocessed audio data based on a general speech recognition model.
[0119] S604b: When the electronic device determines that the voice interaction mode is whisper mode, it performs voice recognition on the audio according to the voice recognition method corresponding to whisper mode to obtain the voice recognition result.
[0120] Optionally, the speech recognition method corresponding to the whispering mode may include: performing audio preprocessing on the audio corresponding to the whispering mode, and performing speech recognition on the preprocessed audio data based on the whispering speech recognition model.
[0121] S604c: When the electronic device determines that the voice interaction mode is the whisper mode, it performs voice recognition on the audio according to the voice recognition method corresponding to the whisper mode to obtain the voice recognition result.
[0122] Optionally, the speech recognition method corresponding to the whisper mode may include: performing audio preprocessing on the audio corresponding to the whisper mode, performing speech recognition on the preprocessed audio data based on the whisper speech recognition model, and obtaining the speech recognition result.
[0123] S605: The electronic device determines the content of the reply to the user based on the speech recognition result, determines the target broadcast method from multiple broadcast methods corresponding to the voice interaction mode according to the information of the audio device, and broadcasts the reply content to the user through the audio device according to the target broadcast method.
[0124] The audio device is a device used to broadcast reply content. For example, the audio device can be a speaker inside an electronic device or a speaker connected to the electronic device, a speaker connected to the electronic device, or headphones connected to the electronic device. The information of the audio device includes at least one of the following: device type, function, and connection method with the electronic device.
[0125] It should be noted that the step S605, "determining the target playback method from multiple playback methods corresponding to the voice interaction mode based on the information from the audio device, and playing the reply content to the user through the audio device according to the target playback method," is optional. The electronic device may also provide the reply content to the user without using audio playback. For example, the electronic device can provide the reply content to the user through a display interface. If the electronic device determines, based on the voice recognition result, that the audio input by the user is used to request the display of the application interface or to execute the functions provided by the application, then the electronic device can respond to the audio input by the user by displaying the interface or executing the application's functions. This application embodiment does not limit this.
[0126] Optionally, Figure 6 In the illustrated embodiment, S604a, S604b, and S604c can also be executed by a server. In this case, the electronic device can send the pre-processed audio data to the server. After obtaining the speech recognition result based on the speech recognition method corresponding to the speech interaction mode, the server can send the speech recognition result to the electronic device. Figure 6 The methods by which the electronic device determines the voice interaction mode, performs audio preprocessing, and performs voice recognition in the illustrated embodiment can be found in the foregoing embodiments, and will not be repeated here.
[0127] In some embodiments of this application, during voice interaction with a user, the electronic device can detect the user's surrounding voice environment based on the acquired user-input audio to determine whether the voice interaction mode is a whisper mode. For example, the electronic device can detect the distance between the audio source and the microphone based on the acquired user audio, such as the distance between the user's mouth and the microphone. The electronic device can also determine the sound intensity of ambient noise in the user audio and detect audio features, such as vocalization features. The electronic device can determine whether the voice interaction mode is a whisper mode based on at least one of the following: the distance between the sound source and the microphone, the sound intensity of ambient noise, and audio features.
[0128] For example, when an electronic device determines that the acquired user audio meets the following conditions, it can determine that the voice interaction mode is whisper mode:
[0129] 1. The distance between the sound source and the microphone is less than the first distance threshold.
[0130] 2. The audio characteristics of user audio include that the user's vocal cords do not produce sound.
[0131] 3. The ambient noise intensity is less than the first sound intensity threshold.
[0132] When an electronic device determines that the voice interaction mode during a voice interaction with a user is a whisper mode, it can perform a speech recognition method corresponding to the whisper mode on the acquired user-input audio. Optionally, the electronic device can perform audio preprocessing corresponding to the whisper mode on the audio, and then perform speech recognition on the audio data obtained from the audio preprocessing based on the speech recognition method corresponding to the whisper mode to obtain a speech recognition result. For example, the audio processing corresponding to the whisper mode can include directional and distance-based sound acquisition and noise reduction processing. The directional and distance-based sound acquisition allows the electronic device to control the microphone to collect the user's voice from a preset location and distance, thereby effectively acquiring the user's voice during the audio acquisition process. The noise reduction processing corresponding to the whisper mode does not filter breath audio, thus preserving the user's breath audio.
[0133] In this embodiment, when the electronic device performs speech recognition on the pre-processed audio data based on the speech recognition method corresponding to the whisper mode, the electronic device can perform speech recognition on the pre-processed audio data based on the speech recognition model corresponding to the whisper mode to obtain the speech recognition result. The speech recognition model corresponding to the whisper mode can be a neural network model trained by the electronic device based on the training sample set corresponding to the whisper mode.
[0134] The electronic device determines the content of the reply to the user based on the speech recognition result and broadcasts the reply to the user through the voice broadcast method corresponding to the whisper mode. The electronic device can determine the target broadcast method from multiple broadcast methods corresponding to the whisper mode based on the information of the audio device. The audio device is a device used to broadcast the reply content. For example, the audio device can be a speaker inside the electronic device or a speaker connected to the electronic device, a speaker connected to the electronic device, or headphones connected to the electronic device. The information of the audio device includes at least one of the following: device type, function, and connection method with the electronic device. The connection method with the electronic device includes Bluetooth connection, satellite connection, wired connection, etc.
[0135] For example, when the voice interaction mode is whisper mode, if the audio device is a speaker within the electronic device or a speaker connected to the electronic device, the electronic device can also respond to the user using a whispered tone. The audio characteristics of a whispered tone include the vocal cords not producing sound. This allows the electronic device to determine in real-time whether it is in whisper mode during voice interaction with the user, and thus instantly switch the voice recognition method and play the audio using a whispered tone, responding to the user in a manner similar to when the user inputs audio, preventing loud voice broadcasts from disturbing the user. If the audio device is a headset connected to the electronic device, the electronic device can respond to the user using a default broadcast mode through the headset. The audio characteristics of the default broadcast mode include the vocal cords producing sound. Therefore, when the user is wearing headsets, the default broadcast mode can be used to ensure audio playback quality.
[0136] It should be noted that, in the embodiments of this application, after the electronic device obtains the audio input by the user, it can also detect the voice environment in which the user is located to determine whether the voice interaction mode is whisper mode or normal mode. Furthermore, the electronic device can perform the voice recognition method corresponding to the voice interaction mode on the obtained user audio to obtain the voice recognition result. For specific implementation, please refer to the foregoing embodiments, and the repeated parts will not be described again.
[0137] Based on the above embodiments, Figure 7 A flowchart illustrating a speech recognition method provided in this application embodiment, which can be executed by an electronic device, is shown below. Figure 7 The method includes the following steps:
[0138] S701: Electronic devices acquire audio input from users during voice interaction.
[0139] S702: The electronic device detects the voice environment based on the audio input by the user and determines that the voice interaction mode is the whisper mode.
[0140] S703: The electronic device performs speech recognition on audio based on the speech recognition method corresponding to the whisper mode, and obtains the speech recognition result.
[0141] Optionally, the electronic device can perform audio preprocessing corresponding to the whisper mode on the audio, and perform speech recognition on the audio data obtained from the audio preprocessing based on the speech recognition method corresponding to the whisper mode.
[0142] For example, the audio processing corresponding to the whisper mode can include directional and distance sound pickup and noise reduction processing. The directional and distance sound pickup allows the electronic device to control the microphone to collect the sound emitted by the user at a preset location and distance, so as to effectively collect the user's voice during the audio acquisition process. The noise reduction processing corresponding to the whisper mode does not filter the breath audio, thus preserving the user's breath audio.
[0143] S704: The electronic device determines the content of the reply to the user based on the speech recognition result, and determines the target broadcast method from the multiple voice broadcast methods corresponding to the whisper mode, and broadcasts the reply content to the user according to the target broadcast method.
[0144] Optionally, the electronic device can determine the target playback method from multiple playback methods corresponding to the voice interaction mode based on information from the audio device. The information from the audio device includes at least one of the following: device type, device function, and connection method between the audio device and the electronic device. For example, if the audio device is a speaker within the electronic device or established in communication with the electronic device, the electronic device can also respond to the user using a breathy playback method. The audio characteristics of the breathy playback method include no vocal cord output, thus preventing the electronic device's loud voice playback from disturbing the user when the user needs to speak softly. If the audio device is a headset established in communication with the electronic device, the electronic device can respond to the user using a default playback method through the headset. The audio characteristics of the default playback method include vocal cord output, thus ensuring audio playback quality when the user is wearing the headset.
[0145] Based on the same concept, embodiments of this application also provide a speech recognition method, which can be executed by an electronic device, such as... Figure 8 A flowchart of a speech recognition method provided in an embodiment of this application is shown below. Figure 8 The method includes the following steps:
[0146] S801: Electronic device acquires audio input from user.
[0147] Optionally, the audio input by the user can be the audio input when the user wakes up the voice assistant of the electronic device, or the audio input by the user during voice interaction with the electronic device.
[0148] S802: The electronic device detects the voice environment when the user inputs the audio based on the audio input by the user, and determines the voice interaction mode corresponding to the audio.
[0149] For example, in this embodiment, the electronic device supports multiple voice interaction modes, including a first mode, a second mode, and a third mode, with each mode corresponding to a different voice recognition method. In this embodiment, the first mode is also called the whisper mode. When the characteristics of the audio meet a first condition, the voice interaction mode corresponding to the audio is determined to be the first mode. The first condition includes at least one of the following: the distance between the audio source and the microphone of the electronic device is less than a first distance threshold; the audio characteristics of the user's audio in the audio include the vocal cords not producing sound; and the ambient noise intensity in the audio is less than a third sound intensity threshold. The second mode is also called the whisper mode. When the characteristics of the audio meet a second condition, the voice interaction mode corresponding to the audio is determined to be the second mode. The second condition includes at least one of the following: the distance between the audio source and the microphone of the electronic device is less than a first distance threshold; the user's voice intensity in the audio is greater than a first sound intensity threshold; and the ambient noise intensity in the audio is greater than a second sound intensity threshold. For audio whose characteristics do not meet the first and second conditions, the electronic device can determine the voice interaction mode corresponding to the audio as the third mode, also called the normal mode.
[0150] S803: The electronic device performs voice recognition on the audio according to the voice recognition method corresponding to the voice interaction mode, and obtains the voice recognition result.
[0151] Optionally, in this embodiment, the electronic device performs a speech recognition method corresponding to the voice interaction mode on the audio, including: the electronic device performs audio preprocessing on the audio according to the voice interaction mode; the electronic device performs speech recognition on the preprocessed audio data according to the speech recognition method corresponding to the voice interaction mode to obtain the speech recognition result. Different voice interaction modes correspond to different voice preprocessing methods. For example, the voice preprocessing corresponding to the first mode includes not filtering breath sounds in the audio, while the audio preprocessing corresponding to the second mode includes filtering breath sounds in the audio. This allows the electronic device to perform adaptive audio preprocessing for user-input audio in different voice environments, preserving as much of the user's voice as possible, thereby improving the accuracy of speech recognition.
[0152] It should be noted that this application Figure 8 The speech recognition method shown can be referred to in the above embodiments of this application for specific implementation, and repeated parts will not be described again.
[0153] Based on the above embodiments, this application also provides an electronic device, which includes multiple functional modules; the multiple functional modules interact to realize the functions performed by the electronic device in the methods described in the embodiments of this application. For example, [the following is an example of implementation]. Figures 6-8 The speech recognition method provided in the illustrated embodiment. The plurality of functional modules can be implemented based on software, hardware, or a combination of software and hardware, and the plurality of functional modules can be arbitrarily combined or divided based on specific implementation.
[0154] Based on the above embodiments, this application also provides an electronic device, which includes at least one processor and at least one memory, wherein the at least one memory stores computer program instructions. When the electronic device is running, the at least one processor executes the functions performed by the electronic device in the various methods described in the embodiments of this application. For example, when executing... Figures 6-8 The speech recognition method provided in the illustrated embodiment.
[0155] Based on the above embodiments, this application also provides a computer program product containing instructions, which, when run on a computer, causes the computer to execute the methods described in the embodiments of this application.
[0156] Based on the above embodiments, this application also provides a computer-readable storage medium storing a computer program, which, when executed by a computer, causes the computer to perform the methods described in the embodiments of this application.
[0157] Based on the above embodiments, this application also provides a chip for reading computer programs stored in a memory to implement the methods described in the embodiments of this application.
[0158] Based on the above embodiments, this application provides a chip system including a processor for supporting a computer device in implementing the methods described in the embodiments of this application. In one possible design, the chip system further includes a memory for storing necessary programs and data of the computer device. This chip system may be composed of chips or may include chips and other discrete devices.
[0159] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0160] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0161] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0162] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0163] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of protection of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A speech recognition method, characterized in that, Applied to an electronic device that supports multiple voice interaction modes, the method includes: Get the audio input from the user; Based on the audio, the voice environment when the user inputs the audio is detected, and the corresponding voice interaction mode of the audio is determined; Based on information from the audio device, a target playback method is determined from multiple playback methods corresponding to the voice interaction mode. The audio device is a device used to play the response content corresponding to the audio. The information of the audio device includes at least one of the following: device type, device function, and connection method between the audio device and the electronic device. The audio device broadcasts the response content corresponding to the audio according to the target broadcast method; The step of detecting the voice environment when the user inputs the audio based on the audio, and determining the voice interaction mode corresponding to the audio, includes: When the features of the audio satisfy a first condition based on the audio, the voice interaction mode corresponding to the audio is determined to be the first mode; or When the characteristics of the audio satisfy the second condition based on the audio, the voice interaction mode corresponding to the audio is determined to be the second mode; or Based on the audio, it is determined that the characteristics of the audio do not satisfy either the first condition or the second condition, and the voice interaction mode corresponding to the audio is determined to be the third mode. The first condition includes that the distance between the sound source of the audio and the microphone of the electronic device is less than a first distance threshold, the audio characteristics of the user's audio in the audio include that the vocal cords are not producing sound, and the ambient noise intensity in the audio is less than a third sound intensity threshold; the second condition includes that the distance between the sound source of the audio and the microphone of the electronic device is less than the first distance threshold, the user's voice intensity in the audio is greater than a first sound intensity threshold, and the ambient noise intensity in the audio is greater than a second sound intensity threshold.
2. The method as described in claim 1, characterized in that, When the audio device is a speaker in the electronic device or a speaker that has established a communication connection with the electronic device, the target broadcast mode is a breathy broadcast mode, and the audio characteristics of the breathy broadcast mode include that the vocal cords do not produce sound; when the audio device is an earphone that has established a communication connection with the electronic device, the target broadcast mode is a default broadcast mode, and the audio characteristics of the default broadcast mode include that the vocal cords produce sound.
3. The method as described in claim 1 or 2, characterized in that, The response content corresponding to the audio is the response content corresponding to the speech recognition result of the audio; the method further includes: When the voice interaction mode is the first mode, the audio is recognized by the voice recognition method corresponding to the first mode to obtain the voice recognition result.
4. The method as described in claim 3, characterized in that, Before performing speech recognition on the audio using the speech recognition method corresponding to the first mode to obtain the speech recognition result, the method further includes: The audio is preprocessed according to the audio preprocessing method corresponding to the first mode, wherein the audio preprocessing method corresponding to the first mode includes filtering out other voices in the audio besides the user's voice, and not filtering out breath sounds in the audio.
5. The method as described in claim 3, characterized in that, The step of performing speech recognition on the audio using the speech recognition method corresponding to the first mode includes: The audio is speech-recognized using the speech recognition model corresponding to the first mode, wherein the speech recognition model corresponding to the first mode is a model trained based on the training sample set corresponding to the first mode.
6. The method as described in claim 1 or 2, characterized in that, The response content corresponding to the audio is the response content corresponding to the speech recognition result of the audio; the method further includes: When the voice interaction mode is the second mode, the audio is recognized by the voice recognition method corresponding to the second mode to obtain the voice recognition result.
7. The method as described in claim 6, characterized in that, Before performing speech recognition on the audio using the speech recognition method corresponding to the second mode, the method further includes: The audio is preprocessed according to the audio preprocessing method corresponding to the second mode, wherein the audio preprocessing method corresponding to the second mode includes filtering out other voices besides the user's voice and breath sounds in the audio.
8. The method according to any one of claims 1-2, 4-5, and 7, characterized in that, Prior to acquiring the user-input audio, the method further includes: It is determined that the user activates the voice interaction function of the electronic device by breathing.
9. An electronic device, characterized in that, It includes at least one processor coupled to at least one memory, the at least one processor being configured to read a computer program stored in the at least one memory to perform the method as described in any one of claims 1-8.
10. A computer program product containing instructions, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1-8.
Citation Information
Patent Citations
Method and device for adjusting playing volume, equipment and storage medium
CN114429766A