Speech processing methods, apparatus, devices, storage media, and computer program products

By setting up a device detection and adaptation module in smart devices, the access and parameter adaptation of external sound card devices are supported, solving the problem that traditional smart devices cannot expand external microphones, realizing a wider range of voice interaction capabilities and application scenarios, and improving the user experience.

CN113157240BActive Publication Date: 2025-10-31BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202110463143.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-27
Publication Date
2025-10-31
Estimated Expiration
2041-04-27

AI Technical Summary

Technical Problem

Traditional smart devices cannot support external microphone expansion, making it impossible to simultaneously achieve voice wake-up recognition and singing functions, and the voice wake-up recognition system does not support external sound card devices.

Method used

By setting up a device detection module, a low-level adaptation module, and a voice recognition service module in the electronic device, the system can detect the access of an external sound card device, close the voice wake-up recognition link, perform parameter adaptation, and restart the voice wake-up recognition link to use the external sound card device for voice processing.

Benefits of technology

It enables the expansion of external sound card devices for smart devices, enhances adaptive voice interaction capabilities, supports more application scenarios, and improves user stickiness and the reputation of electronic products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113157240B_ABST
    Figure CN113157240B_ABST
Patent Text Reader

Abstract

This disclosure discloses speech processing methods, apparatus, devices, storage media, and computer program products, relating to the field of artificial intelligence, particularly speech technology and deep learning. Specifically, in response to the detection of an external sound card device being connected, the following operations are performed: disabling the voice wake-up recognition link; adapting the parameters of the connected external sound card device; and, in response to the completion of parameter adaptation of the external sound card device, restarting the voice wake-up recognition link to enable speech processing via the external sound card device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence, and more particularly to speech technology and deep learning. Specifically, it relates to a speech processing method, a speech processing device, an electronic device, a non-transitory computer-readable storage medium storing computer instructions, and a computer program product. Background Technology

[0002] Currently, smart devices with voice wake-up and recognition capabilities typically use a built-in microphone as the input, with the built-in sound card preprocessing the audio data collected by the microphone. Then, the hardware adaptation layer reads the preprocessed audio data from the built-in sound card node in real time and transmits it to the ASR (Automatic Speech Recognition) module for further audio data processing. Summary of the Invention

[0003] This disclosure provides a speech processing method, apparatus, device, storage medium, and computer program product.

[0004] According to one aspect of this disclosure, a voice processing method is provided, comprising, in response to sensing the access of an external sound card device, performing the following operations: disabling the voice wake-up recognition link; performing parameter adaptation on the accessed external sound card device; and, in response to the completion of parameter adaptation of the external sound card device, restarting the voice wake-up recognition link so as to process the voice through the external sound card device.

[0005] According to another aspect of this disclosure, a voice processing apparatus is provided, comprising, in response to sensing the access of an external sound card device, performing corresponding operations through the following modules: a first link shutdown module for shutting down a voice wake-up recognition link; a first parameter adaptation module for adapting parameters to the accessed external sound card device; and a first link restart module for restarting the voice wake-up recognition link in response to the completion of parameter adaptation of the external sound card device, so as to process voice through the external sound card device.

[0006] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the methods described in the embodiments of this disclosure.

[0007] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method described according to embodiments of this disclosure.

[0008] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method described according to embodiments of this disclosure.

[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0011] Figure 1A An exemplary embodiment of a speech processing method and apparatus system architecture suitable for embodiments of this disclosure is shown;

[0012] Figure 1B and Figure 1C An exemplary diagram illustrates a scenario in which the speech processing method and apparatus of the present disclosure embodiments can be implemented;

[0013] Figure 2 An exemplary flowchart of a speech processing method according to an embodiment of the present disclosure is shown;

[0014] Figures 3A-3C A schematic diagram of a speech processing method according to an embodiment of the present disclosure is shown as an example;

[0015] Figure 4 A block diagram of a voice processing apparatus according to an embodiment of the present disclosure is shown as an example; and

[0016] Figure 5 A block diagram of an electronic device for implementing the speech processing method and apparatus of the present disclosure is shown as an example. Detailed Implementation

[0017] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0018] The recording, storage, and application of relevant data (such as audio data) involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0019] It should be understood that traditional smart devices with voice wake-up recognition capabilities (such as smart speakers) typically use a built-in microphone as the input. Furthermore, the modules within the voice wake-up recognition system of these traditional smart devices are usually pre-configured by the manufacturer and do not support external expansion. Therefore, these traditional smart devices can only perform voice wake-up and recognition based on their built-in sound card.

[0020] It should also be understood that traditional smart speakers cannot simultaneously support voice wake-up recognition and singing (i.e., recording). For example, when a user wants to connect their handheld microphone to a traditional smart speaker to sing, the smart speaker's existing hardware and software cannot support this external microphone.

[0021] In response, this disclosure provides a novel voice processing solution for electronic devices (smart devices), aiming to comprehensively improve the adaptive voice interaction capabilities of electronic devices in various application scenarios in terms of versatility, ease of use, practicality, and stability.

[0022] The present disclosure will be described in detail below with reference to specific embodiments.

[0023] The system architecture of the speech processing method and apparatus suitable for embodiments of this disclosure is described below.

[0024] Figure 1A An exemplary system architecture suitable for the speech processing method and apparatus of this disclosure is illustrated. It should be noted that... Figure 1A The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other environments or scenarios.

[0025] like Figure 1A As shown, the system architecture 100 may include: electronic device 101 and external sound card devices such as Bluetooth headset 102, USB (Universal Serial Bus) handheld microphone 103, and wired headset 104. Among them, electronic device 101 has a built-in microphone 1011 and a built-in sound card 1012.

[0026] It should be noted that, in this embodiment, a device detection module 1013, a low-level adaptation module 1014, and a voice recognition service module 1015 are configured in the electronic device 101 to support external expansion of various modules in its voice wake-up and recognition system. For example, external sound card devices such as Bluetooth headsets 102, USB (Universal Serial Bus) handheld microphones 103, and wired headphones 104 can be connected to the electronic device 101 for use (including for voice wake-up and recognition, and for recording songs, etc.), thereby comprehensively improving the adaptive voice interaction capabilities of the electronic device 101 in various application scenarios.

[0027] For example, electronic device 101 uses its built-in microphone 1011 and built-in sound card 1012 for voice processing by default. Simultaneously, electronic device 101 can detect whether an external sound card device is connected via device detection module 1013. When device detection module 1013 detects an external sound card device, it can notify voice recognition service module 1015 to close the voice wake-up recognition link. It can also notify the underlying adaptation module 1014 to perform parameter adaptation for the connected underlying device, i.e., the external sound card device. After completing the parameter adaptation for the external sound card device, the underlying adaptation module 1014 can notify voice recognition service module 1015 to restart the voice wake-up recognition link. After completing the above operations, electronic device 101 can perform voice processing through the newly connected external sound card device. For example, electronic device 101 can use an external Bluetooth headset 102 for voice wake-up and recognition. For example, a user can use an external USB handheld microphone 103 connected to electronic device 101 to sing, etc.

[0028] It should be understood that Figure 1A The types and number of external sound card devices shown are merely illustrative. Depending on implementation needs, any type and number of external sound card devices can be included.

[0029] The application scenarios of the speech processing method and apparatus suitable for the embodiments of this disclosure are described below.

[0030] It should be noted that the voice processing solution provided in this disclosure can be used in all electronic devices with voice wake-up recognition function (i.e., voice interaction function), such as smart speakers, and is not limited to them in this disclosure.

[0031] It should be understood that the voice processing solution provided by the embodiments of this disclosure can be used to externally expand electronic devices with voice wake-up recognition function, thereby broadening the application scenarios of electronic products, increasing user stickiness to electronic products using this solution, and enhancing the reputation of electronic products.

[0032] For example, such as Figure 1B As shown, Bluetooth headsets can be connected to electronic devices, enabling the electronic devices to be woken up and recognized by voice through the external Bluetooth headset.

[0033] Or, for example, such as Figure 1C As shown, a USB handheld microphone can also be connected to an electronic device, allowing the device to be woken up and recognized via voice using the external USB handheld microphone. Furthermore, after waking up the electronic device, the user can also sing using the external USB handheld microphone; that is, after the electronic device is woken up, it can be used to record songs.

[0034] According to embodiments of this disclosure, this disclosure provides a speech processing method.

[0035] Figure 2 A flowchart of a speech processing method according to an embodiment of the present disclosure is shown as an example.

[0036] like Figure 2 As shown, the voice processing method 200 is applied to electronic devices (such as smart speakers). In response to sensing that an external sound card device is connected, the electronic device can perform the following operations: S210 to S230.

[0037] When operating S210, disable the voice wake-up recognition link.

[0038] When operating the S220, parameter adaptation is performed on the connected external sound card device.

[0039] When operating S230, in response to the completion of parameter adaptation of the external sound card device, the voice wake-up recognition link is restarted so that the voice can be processed through the external sound card device.

[0040] It should be noted that, in this embodiment, the electronic device can use its built-in microphone and sound card for voice wake-up and recognition by default. Once an external sound card is detected connected to the electronic device, the external sound card can be used first for voice wake-up and recognition, as well as for recording songs, etc.

[0041] In some embodiments of this disclosure, a device detection module, an underlying adaptation module, and a voice recognition service module may be provided in the electronic device.

[0042] The electronic device can detect the presence of an external sound card via a device detection module. Upon detecting an external sound card, the detection module can instruct the voice recognition service module to disable the voice wake-up recognition link. Simultaneously, it can instruct the underlying adaptation module to adapt the parameters of the connected device, i.e., the external sound card. After completing the parameter adaptation, the underlying adaptation module can instruct the voice recognition service module to restart the voice wake-up recognition link. After these operations, the electronic device can perform voice processing through the newly connected external sound card. For example, the electronic device can use an external Bluetooth headset for voice wake-up and recognition. Or, a user can use an external USB handheld microphone connected to the electronic device to sing, etc.

[0043] It should be noted that, in the embodiments disclosed herein, when no external sound card device is connected to the electronic device, the voice wake-up recognition link may include a voice wake-up recognition link consisting of the built-in microphone, built-in sound card and hardware adapter layer of the electronic device and the ASR module in the cloud.

[0044] Furthermore, in this embodiment, the electronic device can connect to only one external sound card device at a time, or it can connect to multiple external sound card devices simultaneously. However, regardless of how many external sound card devices are connected at once, the electronic device can choose to use only one of them.

[0045] Furthermore, in this embodiment of the disclosure, when an external sound card device is connected to the electronic device, the electronic device can generate a dedicated external sound card for each external sound card device. In this case, i.e., when an external sound card device is connected to the electronic device, the voice wake-up recognition link can include a voice wake-up recognition link consisting of an external sound card device (such as an external Bluetooth headset), the corresponding external sound card and hardware adaptation layer, and an ASR module in the cloud.

[0046] It should be understood that, in the absence of an external sound card connected to the electronic device, the built-in microphone of the electronic device serves as the sound input, and its output audio data can be preprocessed by the built-in sound card. For example, the built-in sound card can convert the audio data input from the built-in microphone into a specific format. For example, the built-in sound card can integrate multiple audio data input from the built-in microphone. For example, the built-in sound card can convert the analog audio signal input from the built-in microphone into a corresponding digital audio signal. The hardware adaptation layer can read the preprocessed audio data from the built-in sound card node and transmit the read audio data to the ASR module in the cloud. The ASR module processes the audio data to ultimately complete voice wake-up and recognition. For example, the ASR module can extract features from the input audio data to ultimately complete voice wake-up and recognition. Furthermore, in this embodiment, different speech recognition models can be trained and configured for different sound card devices. The ASR module can be used to train various reference models and create corresponding reference model libraries. Moreover, the ASR module can also be used to match the corresponding speech recognition model to the sound card device when performing voice wake-up and recognition based on different sound card devices.

[0047] It should also be understood that when an external sound card is connected to an electronic device, the external sound card acts as a sound input terminal, and its output audio data can be preprocessed by the external sound card. The content and processing logic of the external sound card are the same as or similar to those of the built-in sound card, and will not be repeated here. Furthermore, when an external sound card is connected to an electronic device, the processing content and logic of the hardware adaptation layer and the cloud-based ASR module in the voice wake-up recognition link are the same as or similar to those when no external sound card is connected to the electronic device, and will not be repeated here either.

[0048] It should be understood that the ASR module is the core module for realizing human-computer dialogue interaction. Its main function is to convert the user's audio stream data into text data for the analysis and matching of user intent. The ASR module provides a standard recognition interface based on audio stream data, which can acquire audio stream data through the built-in microphone of electronic devices, external microphones connected to external USB sound cards, or Bluetooth headset microphones. Within this module, it performs noise reduction, feature extraction, speech decoding, and text conversion of the audio stream data, converting the user's dialogue information into precise text for further semantic analysis.

[0049] In this embodiment of the disclosure, the external sound card device may include, but is not limited to, an external sound card, wired headphones, Bluetooth headphones, and USB handheld microphones.

[0050] Through the embodiments of this disclosure, upon detecting the connection of an external sound card device, the parameters of the external sound card device can be adapted first, and then the external sound card device can be used for recording, playing recordings (including playing music), voice wake-up, and recognition. Therefore, the technical solution provided by the embodiments of this disclosure can enable electronic devices to adapt to more application scenarios by expanding the external sound card device.

[0051] As an optional embodiment, the method may further include: before performing parameter adaptation on the accessed external sound card devices, in response to sensing that multiple external sound card devices are accessed simultaneously, determining the external sound card device with the highest priority among the multiple external sound card devices.

[0052] Correspondingly, parameter adaptation for the connected external sound card device can include: parameter adaptation for the external sound card device with the highest priority.

[0053] In some embodiments of this disclosure, the electronic device can support connecting only one external sound card device at a time. With this approach, there is no need to determine the type or priority of the external sound card device; the appropriate parameters can be directly adapted to it.

[0054] In other embodiments of this disclosure, the electronic device can support the simultaneous connection of multiple external sound card devices. In this approach, the electronic device can select only one of the multiple connected external sound card devices for use at a time. Therefore, before adapting the parameters, the type of each connected external sound card device can be determined, and the priority of each external sound card device can be determined based on its type. Then, only the external sound card device with the highest priority is selected, and parameter adaptation is performed only on it.

[0055] It should be understood that, in the embodiments disclosed herein, users can customize the priority of various types of external sound card devices according to actual application scenarios or usage habits. For example, the priority of Bluetooth external sound card devices can be defined as: priority of USB external sound card devices > priority of wired external sound card devices > priority of internal sound card devices.

[0056] For example, the device detection module can monitor changes in the kernel device tree, such as whether kernel device nodes have been added, deleted, or modified. When multiple related external sound card devices are detected being inserted into or connected to an electronic device (i.e., multiple related kernel device nodes corresponding to external sound card devices are added to the kernel device tree), the module first reads the attribute information of these kernel device nodes. Then, based on the read attribute information, it determines the type of each currently connected external sound card device. Based on the type of these external sound card devices, it determines their priority and identifies the external sound card device with the highest priority. After identifying the highest priority external sound card device, it notifies the upper-layer speech recognition service module to shut down the currently used speech wake-up recognition link and simultaneously notifies the lower-layer adaptation module to configure the relevant parameters of the newly added external sound card device. Once the relevant parameters are configured, the lower-layer adaptation module can configure the speech recognition service module to restart the speech recognition wake-up link. Alternatively, after the lower-layer adaptation module completes the relevant parameter configuration, the device detection module can notify the upper-layer speech recognition service module to restart the speech recognition wake-up link.

[0057] Through the embodiments of this disclosure, when multiple external sound card devices are connected to an electronic device, only the one with the highest priority can be selected for use according to the priority of each external sound card device.

[0058] Furthermore, as an optional embodiment, the method may also include: performing the following operations in response to sensing that the highest priority external sound card device has been disconnected from the electronic device.

[0059] Disable the voice wake-up recognition link.

[0060] Identify the external sound card device with the second highest priority among multiple external sound card devices.

[0061] Perform parameter adaptation for external sound card devices with the second highest priority.

[0062] Once the parameters of the second-highest priority external sound card device have been adapted, the voice wake-up recognition link is restarted so that the voice can be processed through the second-highest priority external sound card device.

[0063] In this embodiment, when the device detection module detects that the highest-priority external sound card device has been disconnected from the electronic device, and other external sound card devices are still connected to the electronic device, the voice wake-up recognition link can be shut down first. Then, the second-highest-priority external sound card device among the multiple external sound card devices mentioned in the previous embodiment (i.e., the highest-priority one among the multiple remaining external sound card devices currently connected to the electronic device) can be selected for parameter adaptation. After the parameter adaptation is completed, the voice wake-up recognition link is restarted so that the second-highest-priority external sound card device can be used for subsequent voice processing.

[0064] According to the embodiments of this disclosure, when multiple external sound card devices are simultaneously connected to an electronic device, if the external sound card device with the highest priority disconnects, the device with the highest priority can still be selected from the remaining at least one external sound card device for continued use. Therefore, the embodiments of this disclosure can support hardware expansion for smart electronic devices. Furthermore, it can also support plug-and-play functionality for external sound card devices such as USB, wired (e.g., wired headphones), or Bluetooth (e.g., Bluetooth headphones).

[0065] Alternatively, as an optional embodiment, the method may further include: in response to sensing that multiple external sound card devices have been disconnected, performing the following operations.

[0066] Disable the voice wake-up recognition link.

[0067] The speech recognition model was switched from the second model associated with the external sound card device to the first model associated with the internal sound card device.

[0068] Once the speech recognition model has switched, the speech wake-up recognition link is restarted so that the speech can be processed through the built-in sound card.

[0069] In this embodiment of the disclosure, after the device detection module detects that all external sound card devices connected to the electronic device have been disconnected from the electronic device, the built-in sound card device (such as built-in microphone, built-in speaker, and built-in sound card) of the electronic device can be reactivated.

[0070] Specifically, the voice wake-up recognition link can be shut down, and the voice recognition model used by the ASR module can be switched from a model adapted to the external sound card device to a model adapted to the internal sound card device. After completing the model switch, the voice wake-up recognition link can be restarted so that the built-in sound card device of the electronic device can be used for subsequent voice processing.

[0071] Furthermore, in this embodiment, different speech recognition models can be trained and configured for different sound card devices (including built-in sound card devices and various external sound card devices) so that in practical applications, the electronic device can achieve better sound quality by adapting the corresponding speech recognition model according to the selected sound card device.

[0072] In some embodiments, the ASR module can be used to train various reference models (each speech recognition model) and create a corresponding reference model library. Furthermore, the ASR module can also be used to match the corresponding speech recognition model to the sound card device when performing voice wake-up and recognition based on different sound card devices.

[0073] For example, if a first model is configured for a built-in sound card device and a second model is configured for an external sound card device, then when the external sound card device is disconnected from the electronic device and no other external sound card device is connected to the electronic device, the built-in sound card device of the electronic device can be restarted, and the speech recognition model used by the ASR module can be switched from the second model previously used with the external sound card device to the first model required for the built-in sound card device.

[0074] Through the embodiments of this disclosure, the ASR module can simultaneously support multiple speech recognition models and can match the corresponding speech recognition model with different sound card devices. This overcomes the shortcomings of related technologies where ASR modules do not support speech recognition model expansion and only support a single speech recognition model. This allows electronic devices to achieve high sound quality in various application scenarios. Furthermore, through the embodiments of this disclosure, after all external sound card devices are disconnected from the electronic device, the built-in sound card of the electronic device can be automatically reactivated for subsequent voice processing by the user.

[0075] As an optional embodiment, the method may further include: switching the voice recognition model from a first model associated with the built-in sound card device to a second model associated with the external sound card device before restarting the voice wake-up recognition link.

[0076] In this embodiment, after an electronic device detects the connection of an external sound card device, in addition to disabling the voice wake-up recognition link and adapting the parameters of the connected external sound card device, it can also switch the voice recognition model from the original model adapted to the built-in sound card device to the model adapted to the currently connected external sound card device. Furthermore, after completing the parameter adaptation of the external sound card device and the switching of the voice recognition model, the voice wake-up recognition link is restarted to allow audio stream data to be input through the external sound card device, and to perform voice processing on the relevant audio stream data using the voice recognition model matched to the external sound card device.

[0077] Through the embodiments disclosed herein, electronic devices can achieve high sound quality levels regardless of the application scenario. This is because different speech recognition models trained and configured for different sound card devices (including built-in sound card devices and various external sound card devices) are used to process audio stream data collected by different hardware microphones, which can improve the parsing accuracy of audio stream data collected by different hardware microphones.

[0078] As an optional embodiment, the method may further include performing at least one of the following operations after restarting the voice wake-up recognition link.

[0079] Record voice information using an external sound card device.

[0080] Output voice information through an external sound card device.

[0081] The voice wake-up operation is performed based on the voice information input through an external sound card device.

[0082] Speech recognition is performed based on the voice information input through an external sound card device.

[0083] For example, when an external Bluetooth headset is connected and used, it can be used to wake up an electronic device via voice. After waking the electronic device, it can also perform voice recognition processing on the audio stream data input from the external Bluetooth headset. Furthermore, when the external Bluetooth headset is enabled on the electronic device, it can also output audio stream data, such as playing music.

[0084] For example, when an external handheld microphone is connected and used, it can be used to wake up an electronic device via voice. After waking the electronic device, it can also perform speech recognition processing on the audio stream data input through the external handheld microphone. Furthermore, when the external handheld microphone is enabled on the electronic device, it can also be used to input audio stream data, such as for recording songs.

[0085] Through the embodiments of this disclosure, an external sound card device can be used for voice wake-up and recognition. Furthermore, the external sound card device can also be used for recording (e.g., recording songs) and playing back recordings (e.g., playing music). Therefore, by expanding the range of external sound card devices, electronic devices can be adapted to more application scenarios.

[0086] As an optional embodiment, the method further includes: sensing whether an external sound card device is connected based on the kernel device tree.

[0087] For example, the device detection module can monitor changes in the kernel device tree, such as whether kernel device nodes have been added, deleted, or modified. If it detects the addition of one or more related external sound card device kernel device nodes in the kernel device tree, it assumes that one or more related external sound card devices are currently inserted into or connected to the electronic device. If it detects the deletion of one or more related external sound card device kernel device nodes in the kernel device tree, it assumes that one or more related external sound card devices have been unplugged from the electronic device or disconnected from the electronic device. If it detects no addition, deletion, or modification of corresponding kernel device nodes in the kernel device tree, it assumes that no related external sound card devices have been inserted into or unplugged from the electronic device or have established or disconnected from the electronic device.

[0088] Furthermore, as an optional embodiment, parameter adaptation of the connected external sound card device may include the following operations.

[0089] Determine the kernel device node in the kernel device tree that corresponds to the external sound card device.

[0090] Based on the kernel device node, obtain the attribute information of the external sound card device.

[0091] Based on attribute information, parameters are adapted for external sound card devices.

[0092] For example, the device detection module can monitor changes in the kernel device tree, such as whether kernel device nodes have been added, deleted, or modified. When multiple related external sound card devices are detected being inserted into or connected to an electronic device (i.e., multiple related kernel device nodes corresponding to external sound card devices are added to the kernel device tree), the module first reads the attribute information of these kernel device nodes. Then, based on the read attribute information, it determines the type of each currently connected external sound card device. Furthermore, based on the type of these external sound card devices, it determines their priority and identifies the external sound card device with the highest priority. After identifying the highest priority external sound card device, it notifies the upper-layer speech recognition service module to disable the currently used speech wake-up recognition link, and simultaneously notifies the lower-layer adaptation module to configure the relevant parameters of the newly added external sound card device.

[0093] In this embodiment, the attribute information of the external sound card device (including but not limited to its sampling rate, bit width, buffer size, and number of channels (i.e., how many channels it has, such as mono or multi-channel)) can be read from the kernel device node corresponding to the newly added external sound card device in the kernel device count. Based on the read attribute information, the external sound card device is adapted to match its parameters such as sampling rate, bit width, buffer size, and number of channels with the parameters of the ASR module.

[0094] After the relevant parameters are configured, the underlying adaptation module can notify the speech recognition service module to restart the speech recognition wake-up link. Alternatively, after the underlying adaptation module has completed the relevant parameter configuration, the device detection module can notify the upper-layer speech recognition service module to restart the speech recognition wake-up link.

[0095] It should be noted that, in this embodiment of the disclosure, in addition to adapting the parameters of the connected external sound card device, the underlying adaptation module can also integrate the audio stream data input from the external sound card device, such as converting the audio stream data input from the external sound card device into a unified data format.

[0096] When the device detection module detects the insertion or connection of a new external sound card device, it can notify the underlying adaptation module to read the attribute information of the corresponding kernel device node and configure the parameters of the new external sound card device. Once the voice recognition service module restarts the voice wake-up recognition link, the electronic device can activate the new external sound card device to record voice data and integrate the recorded multi-channel voice data, such as integrating voice data recorded from multiple microphone arrays, to meet the requirements of the ASR module.

[0097] The following will combine Figures 3A-3C The method logic and basic principles of the embodiments of this disclosure are explained in detail with reference to specific examples.

[0098] First, we can define the priority as follows: Bluetooth external sound card devices > USB external sound card devices > wired external sound card devices > internal sound card devices.

[0099] like Figure 3AAs shown, electronic device 101 uses its built-in microphone 1011 as the input terminal by default, and the built-in sound card 1012 preprocesses the audio stream data output by the built-in microphone 1011. Once the device detection module 1013 detects that an external sound card device, such as a Bluetooth headset 102, a USB handheld microphone 103, or a wired headset 104, is connected to electronic device 101, it notifies the voice recognition service module 1015 to close the original voice wake-up recognition link, and simultaneously notifies the underlying adaptation module 1014 to perform parameter adaptation for the connected underlying external sound card device. Since the Bluetooth headset 102 has the highest priority among the three currently connected external sound card devices, the underlying adaptation module 1014 can currently only perform parameter adaptation for the Bluetooth headset 102. After the parameter adaptation is completed, the underlying adaptation module 1014 or the device detection module 1013 can notify the voice recognition service module 1015 to switch the voice recognition model matched with the built-in microphone 1011 to the voice recognition model matched with the Bluetooth headset 102, and restart the voice wake-up recognition link after the switch is completed. In this system, the voice wake-up recognition link after restarting uses the Bluetooth headset 102 as the voice input terminal, and the external sound card 1016 preprocesses the audio stream data output by the Bluetooth headset 102. The external sound card 1016 is automatically generated after the electronic device 101 detects the Bluetooth headset 102.

[0100] like Figure 3B As shown, once the device detection module 1013 detects that the Bluetooth headset 102 has disconnected from the electronic device 101, it notifies the voice recognition service module 1015 again to close the currently used voice wake-up recognition link. Simultaneously, it notifies the underlying adaptation module 1014 to perform parameter adaptation on the USB handheld microphone 103, which is the second-priority external sound card device previously connected. After parameter adaptation is complete, the underlying adaptation module 1014 or the device detection module 1013 can again notify the voice recognition service module 1015 to switch the voice recognition model matched with the Bluetooth headset 102 to the voice recognition model matched with the USB handheld microphone 103, and restart the voice wake-up recognition link after the switch is complete. The restarted voice wake-up recognition link uses the USB handheld microphone 103 as the voice input terminal, and the external sound card 1017 preprocesses the audio stream data output by the USB handheld microphone 103. The external sound card 1017 is automatically generated after the electronic device 101 detects the USB handheld microphone 103.

[0101] like Figure 3CAs shown, once the device detection module 1013 detects that all previously connected external sound card devices (such as Bluetooth headset 102, USB handheld microphone 103, and wired headset 104) have been disconnected from the electronic device 101, it notifies the voice recognition service module 1015 to close the currently used voice wake-up recognition link. Simultaneously, it notifies the voice recognition service module 1015 to switch the voice recognition model matched with the external sound card device, such as the USB handheld microphone 103, to the voice recognition model matched with the built-in microphone 1011, and restarts the voice wake-up recognition link after the switch is complete. The restarted voice wake-up recognition link uses the built-in microphone 1011 as the voice input terminal, and the built-in sound card 1012 preprocesses the audio stream data output by the built-in microphone 1011.

[0102] According to embodiments of this disclosure, this disclosure also provides a voice processing apparatus.

[0103] Figure 4 A block diagram of a speech processing apparatus according to an embodiment of the present disclosure is shown as an example.

[0104] like Figure 4 As shown, in response to sensing the connection of an external sound card device, the voice processing device 400 can perform corresponding operations through the following modules: first link shutdown module 410, first parameter adaptation module 420, and first link restart module 430.

[0105] Specifically, the first link shutdown module 410 is used to shut down the voice wake-up recognition link.

[0106] The first parameter adaptation module 420 is used to adapt the parameters of the connected external sound card device.

[0107] The first link restart module 430 is used to restart the voice wake-up recognition link in response to the completion of parameter adaptation of the external sound card device, so that the voice can be processed through the external sound card device.

[0108] As an optional embodiment, the above-described apparatus further includes: a first determining module, configured to, in response to sensing that multiple external sound card devices are connected simultaneously, determine the external sound card device with the highest priority among the multiple external sound card devices before the first parameter adaptation module performs parameter adaptation on the connected external sound card device. Correspondingly, the first parameter adaptation module is further configured to: perform parameter adaptation on the external sound card device with the highest priority.

[0109] As an optional embodiment, the above-described apparatus further includes: in response to sensing that the highest-priority external sound card device has been disconnected, performing corresponding operations through the following modules: a second link shutdown module for shutting down the voice wake-up recognition link; a second determination module for determining the second-highest-priority external sound card device among the plurality of external sound card devices; a second parameter adaptation module for performing parameter adaptation on the second-highest-priority external sound card device; and a second link restart module for restarting the voice wake-up recognition link in response to the completion of parameter adaptation of the second-highest-priority external sound card device, so that voice can be processed through the second-highest-priority external sound card device.

[0110] As an optional embodiment, the above-mentioned device further includes: in response to sensing that all of the plurality of external sound card devices have been disconnected, performing corresponding operations through the following modules: a third link shutdown module, used to shut down the voice wake-up recognition link; a first switching module, used to switch the voice recognition model from a second model associated with the external sound card device to a first model associated with the built-in sound card device; and a third link restart module, used to restart the voice wake-up recognition link in response to the completion of the voice recognition model switching, so as to process the voice through the built-in sound card device.

[0111] As an optional embodiment, the above-mentioned device further includes: a second switching module, used to switch the voice recognition model from a first model associated with the built-in sound card device to a second model associated with the external sound card device before the first link restart module restarts the voice wake-up recognition link.

[0112] As an optional embodiment, the above-described apparatus further includes at least one of the following modules: an input module for performing corresponding operations after the first link restart module restarts the voice wake-up recognition link; an output module for inputting voice information through the external sound card device; an output module for outputting voice information through the external sound card device; a voice wake-up module for performing a voice wake-up operation based on the voice information input through the external sound card device; and a voice recognition module for performing a voice recognition operation based on the voice information input through the external sound card device.

[0113] As an optional embodiment, the above-described apparatus further includes: a sensing module, used to sense whether an external sound card device is connected based on the kernel device tree.

[0114] As an optional embodiment, the first parameter adaptation module includes: a determining unit, used to determine the kernel device node in the kernel device tree corresponding to the external sound card device; an obtaining unit, used to obtain the attribute information of the external sound card device based on the kernel device node; and a parameter adaptation unit, used to perform parameter adaptation on the external sound card device based on the attribute information.

[0115] It should be understood that the embodiments of the apparatus portion of this disclosure correspond to the same or similar embodiments of the method portion of this disclosure, and the technical problems solved and the technical effects achieved are also the same or similar, which will not be repeated here.

[0116] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0117] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0118] like Figure 5 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. The RAM 503 may also store various programs and data required for the operation of the electronic device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0119] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0120] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as speech processing methods. For example, in some embodiments, the speech processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the speech processing method described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform speech processing methods by any other suitable means (e.g., by means of firmware).

[0121] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0122] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0123] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0124] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0125] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0126] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0127] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0128] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A speech processing method, comprising: Based on the kernel device tree, detect whether an external sound card device is connected; In response to the detection of multiple external sound card devices being connected, perform the following operations: Disable the voice wake-up recognition link; Parameter adaptation is performed on the external sound card device with the highest priority among the multiple external sound card devices connected. The priority of the external sound card device is determined according to the type of the external sound card device. In response to the completion of parameter adaptation of the external sound card device, the voice wake-up recognition link is restarted so that the voice can be processed through the external sound card device; In response to the detection that the highest priority external sound card device has been disconnected, the voice wake-up recognition link is closed; the parameters of the second highest priority external sound card device among the multiple connected external sound card devices are adapted to adapt the parameter size of the second highest priority external sound card device to be compatible with the parameter size of the automatic voice recognition module. In response to the completion of parameter adaptation for the second-highest priority external sound card device, the speech recognition model used by the automatic speech recognition module is switched from the second model associated with the highest priority external sound card device to the second-highest priority external sound card device. This is to use the second model associated with the second-highest priority external sound card device to perform speech processing on the unified format audio stream data input through the second-highest priority external sound card device and obtained by integrating multiple voice data. After the switch is completed, the voice wake-up recognition link is restarted, and a speech recognition operation is performed based on the voice information input through the second-highest priority external sound card device to convert the audio stream data into text data. The parameters of the second-highest priority external sound card device are adapted based on the attribute information of the second-highest priority external sound card device, which is obtained from the kernel device node corresponding to the second-highest priority external sound card device in the kernel device tree.

2. The method according to claim 1, wherein: The method further includes: before performing parameter adaptation on the connected external sound card devices, in response to sensing that multiple external sound card devices are connected at the same time, determining the external sound card device with the highest priority among the multiple external sound card devices.

3. The method according to claim 2, further comprising: In response to the detection that the highest-priority external sound card device has been disconnected, the following operations are also performed: The external sound card device with the second highest priority among the plurality of external sound card devices is determined.

4. The method according to claim 1, further comprising, after restarting the voice wake-up recognition link, performing at least one of the following operations: Voice information is recorded using the external sound card device; Voice information is output through the external sound card device; The voice wake-up operation is performed based on the voice information input through the external sound card device.

5. A voice processing device, comprising, The perception module is used to detect whether an external sound card device is connected based on the kernel device tree. Upon detecting the connection of an external sound card device, the following operations are performed through the following modules: The first link shutdown module is used to shut down the voice wake-up recognition link; The first parameter adaptation module is used to adapt parameters to the external sound card device with the highest priority among the multiple external sound card devices connected. The priority of the external sound card device is determined according to the type of the external sound card device. The first link restart module is used to restart the voice wake-up recognition link in response to the completion of parameter adaptation of the external sound card device, so that the voice can be processed through the external sound card device; In response to the detection that the highest priority external sound card device has been disconnected, the following modules are used to perform corresponding operations: a second link shutdown module, used to shut down the voice wake-up recognition link; and a second parameter adaptation module, used to adapt the parameters of the second highest priority external sound card device among the multiple external sound card devices, so as to adapt the parameter size of the second highest priority external sound card device to be compatible with the parameter size of the automatic voice recognition module. The second link restart module is used to switch the speech recognition model used by the automatic speech recognition module from the second model associated with the highest priority external sound card device to the second model associated with the second highest priority external sound card device in response to the completion of parameter adaptation of the second highest priority external sound card device. This is to use the second model associated with the second highest priority external sound card device to perform speech processing on the unified format audio stream data input through the second highest priority external sound card device and obtained by integrating multiple voice data. After the switching is completed, the speech wake-up recognition link is restarted, and speech recognition operation is performed based on the speech information input through the second highest priority external sound card device to convert the audio stream data into text data. The parameters of the second-highest priority external sound card device are adapted based on the attribute information of the second-highest priority external sound card device, which is obtained from the kernel device node corresponding to the second-highest priority external sound card device in the kernel device tree.

6. The apparatus according to claim 5, wherein: The device further includes: a first determining module, configured to, in response to sensing that multiple external sound card devices are connected simultaneously, determine the external sound card device with the highest priority among the multiple external sound card devices before the first parameter adaptation module performs parameter adaptation on the connected external sound card devices.

7. The apparatus according to claim 6, further comprising: In response to the detection that the highest-priority external sound card device has been disconnected, the following operations are also performed through the following modules: The second determining module is used to determine the external sound card device with the second highest priority among the plurality of external sound card devices.

8. The apparatus according to claim 5 further comprises at least one of the following modules: configured to perform a corresponding operation after the first link restart module restarts the voice wake-up recognition link. The input module is used to input voice information through the external sound card device; The output module is used to output voice information through the external sound card device; The voice wake-up module is used to perform a voice wake-up operation based on the voice information input through the external sound card device; The speech recognition module is used to perform speech recognition operations based on the speech information input through the external sound card device.

9. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-4.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-4.

11. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Task processing method and apparatus

    CN107766156A

  • Voice processing method, device and electronic equipment

    CN110197658A

  • Voice control method and device, electronic equipment and storage medium

    CN111681654A

  • Voice transmission method, intelligent terminal and computer readable storage medium

    CN112216279A

  • Bluetooth management method and mobile terminal

    CN112218385A