Multi-language switching-free interaction method, device, and electronic device
Through an end-to-end multilingual speech recognition model, audio features are extracted and common and differential features are obtained. Language decoding is performed in combination with a language routing network, which solves the problem of high switching and maintenance costs in multilingual interaction and achieves an efficient and seamless voice interaction experience.
Patent Information
- Application Number
- CN202310081913.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-16
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2043-01-16
AI Technical Summary
Existing technologies cannot achieve seamless multi-language interaction, require manual switching, have high maintenance costs, and have poor speech recognition and semantic understanding between different languages.
An end-to-end multilingual speech recognition model is used to extract audio features and obtain common and differential features. Language decoding is performed in combination with a language routing network, and transcription text and language labels are output to achieve seamless voice interaction.
It achieves seamless multi-language interaction, improves the accuracy of speech recognition and semantic understanding, reduces maintenance costs, and enhances the human-computer interaction experience.
Smart Images

Figure CN116486784B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of voice interaction technology, and in particular to a multi-language switching-free interaction method, device, and electronic device. Background Art
[0002] Current human-computer voice interaction systems generally only support single-language interaction, such as Chinese, and do not support multilingual interaction. Even if the system supports multilingual interaction, manual language switching is required. The differences between different languages far exceed those between Mandarin and dialects. Mandarin-dialect recognition cannot adapt to cross-lingual recognition, especially in languages with significant differences such as Chinese, English, and Arabic. This is a challenge for Mandarin-dialect conversion technology.
[0003] Specifically, most human-computer voice interaction solutions use a separate voice interaction system for each language, and each language system is independent, making it impossible for users of different languages to use the same voice interaction system simultaneously. Currently, the approach to switching between Mandarin Chinese and dialects is to either use separate recognition systems for Mandarin and dialects and then compare the results, or to fine-tune and optimize the dialect based on Mandarin. However, due to the significant differences between languages, especially those with less data, the performance of voice recognition and semantic understanding differs significantly from that of languages with relatively large data volumes. For example, Chinese speech recognition rates can reach 95% or higher, while Arabic speech recognition rates may only reach 80% or even lower. Switching between these two languages using the aforementioned Mandarin and dialect technology approach will not achieve the desired results. Furthermore, the Mandarin-to-dialect conversion solution is difficult to migrate and adapt.
[0004] In summary, the existing technical solutions have the following shortcomings:
[0005] First, the system cannot support multilingual voice interaction, or multilingual voice interaction requires manual switching, and there is no stable and reliable adaptive switching adaptation strategy;
[0006] Second, each language's voice interaction system is an independent one, resulting in high maintenance costs.
[0007] Third, the core interactive capabilities of different languages, such as speech recognition, semantic understanding, and speech synthesis, will have greatly different effects due to differences in data volume, which in turn leads to a poor interactive experience when switching languages. Summary of the Invention
[0008] In view of the above, the present invention aims to provide a multi-language switching-free interaction method, device, and electronic device to solve the aforementioned problems arising from multi-language interaction.
[0009] The technical solution adopted in the present invention is as follows:
[0010] In a first aspect, the present invention provides a multi-language switching-free interaction method, comprising:
[0011] Correspondingly extracting audio features of input speech in different languages;
[0012] Inputting the audio features into a pre-trained multilingual speech recognition model, wherein the multilingual speech recognition model adopts an end-to-end modeling mechanism;
[0013] The multilingual speech recognition model obtains common features and differential features of multiple languages from the audio features, and converts the audio feature sequence into a unified modeling unit sequence by combining the common features and the differential features to obtain feature-enhanced acoustic information;
[0014] Based on the acoustic information, the multilingual speech recognition model performs language decoding for different languages and outputs transcribed text and language labels corresponding to each language;
[0015] The transcribed text and the language label are used to perform semantic understanding and perform interactive operations.
[0016] In at least one possible implementation, the multilingual speech recognition model is trained based on a transfer learning mechanism, and a trained and converged model of a language with higher resources is used as an initialization model for a language with lower resources to obtain similar characteristics of pronunciations of different languages in terms of time-frequency features.
[0017] In at least one possible implementation, different unified modeling units are used for multiple languages with different resource amounts.
[0018] In at least one possible implementation, in the multilingual speech recognition model, a language routing network is set before language decoding. The language routing network is used to automatically match the language difference information contained in the acoustic information to the corresponding language channel and output the corresponding language label.
[0019] In at least one possible implementation, the training process of the multilingual speech recognition model includes:
[0020] Use data from all languages to train a universal model;
[0021] Inserting the language routing network into the trained universal model;
[0022] The general model parameters are fixed, and only the language routing network is trained according to each language.
[0023] In at least one possible implementation, during the language decoding process, constrained decoding is performed for each language based on the language label output by the language routing network.
[0024] In at least one possible implementation, a wake-up voice provided by a user before inputting a mixed multilingual voice is used to obtain a wake-up tag containing language information, and the wake-up tag is used to assist the multilingual voice recognition model in outputting the language tag.
[0025] In a second aspect, the present invention provides a multi-language switching-free interactive device, comprising:
[0026] Audio feature extraction module, used to extract audio features of input speech in different languages;
[0027] a feature input module, configured to input the audio features into a pre-trained multilingual speech recognition model, wherein the multilingual speech recognition model adopts an end-to-end modeling mechanism;
[0028] an acoustic information processing module, configured to obtain common features and differential features of multiple languages from the audio features by the multilingual speech recognition model, and convert the audio feature sequence into a unified modeling unit sequence by combining the common features and the differential features to obtain feature-enhanced acoustic information;
[0029] A language decoding processing module, configured to perform language decoding for different languages using the multilingual speech recognition model based on the acoustic information, and output a transcribed text and language label corresponding to each language;
[0030] The semantic understanding and interaction module is used to use the transcribed text and the language label to perform semantic understanding and perform interactive operations.
[0031] In a third aspect, the present invention provides an electronic device, comprising:
[0032] One or more processors, a memory, and one or more computer programs, wherein the memory may adopt a non-volatile storage medium, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions, and when the instructions are executed by the device, the device performs the method as described in the first aspect or any possible implementation of the first aspect.
[0033] The main idea of the present invention is to avoid training the recognition model for each language separately, but to jointly train the multilingual speech recognition model with data of multiple languages, and to achieve seamless multilingual switching-free voice interaction based on the mixed modeling of multilingual common features. Specifically, the input mixed language speech audio features are fed into the end-to-end multilingual speech recognition model, from which the common features and difference features of the multilinguals are obtained, and the two are combined for acoustic modeling and language decoding, and the transcribed text and language labels corresponding to each language are output. Finally, the transcribed text and language labels are used for semantic understanding and to perform interactive operations. The present invention does not need to rely on manual switching, and eliminates the differences between different languages in speech recognition, semantic understanding, and speech synthesis. In particular, it does not require switching, and directly performs comprehensive recognition and understanding of the mixed language voice interaction, thereby significantly improving the human-computer interaction experience.
[0034] Furthermore, the present invention also proposes a language routing concept, which uses multilingual tags for constrained decoding, which can greatly improve the accuracy of cross-language speech recognition and avoid deviations in subsequent semantic understanding and interactive operations caused by crosstalk. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be further described below with reference to the accompanying drawings, in which:
[0036] Figure 1 A flowchart of an embodiment of the multi-language switching-free interaction method provided by the present invention;
[0037] Figure 2 A schematic diagram of an embodiment of the multi-language switching-free interaction device provided by the present invention;
[0038] Figure 3 A schematic diagram of an embodiment of an electronic device provided by the present invention. DETAILED DESCRIPTION
[0039] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.
[0040] In order to solve the current drawbacks of switching between multiple languages, especially languages with large differences in data volume and recognition rate, in interactive scenarios, the present invention proposes at least one embodiment of the following multi-language switching-free interaction method, such as Figure 1 Specifically, it may include:
[0041] Step S1: extracting audio features of input speech in different languages;
[0042] It is understandable that in actual operation, the input mixed language speech can also be acoustically preprocessed, such as but not limited to noise reduction. The speech noise reduction process can adopt existing mature solutions, such as noise reduction, echo cancellation, sound source localization, etc.; similarly, extracting conventional audio feature sequences is also a conventional method in the field of speech processing, and the present invention will not elaborate on this.
[0043] Step S2: inputting the audio features into a pre-trained multilingual speech recognition model, wherein the multilingual speech recognition model adopts an end-to-end modeling mechanism;
[0044] Those skilled in the art can understand that the speech recognition model mainly includes an acoustic model for calculating the probability of speech to acoustic units, and a language model for further decoding into corresponding text according to the probability of acoustic units. Based on this, the present invention proposes that the multilingual speech recognition model adopts an end-to-end modeling mechanism, which is to merge the traditional acoustic model and language model, and complete the tasks that the two traditional models handle separately within one model. It can be understood that the processing logic executed inside the "black box" of the end-to-end model in this embodiment is similar to the traditional acoustic model and language model, and the advantage is that after merging the two into an end-to-end model, it is convenient to share acoustic and language features. The following text will explain it in conjunction with the model training process. It should be pointed out here that the acoustic modeling and language decoding processing processes in the following links can be implemented in a unified end-to-end model architecture.
[0045] Step S3: The multilingual speech recognition model obtains common features and differential features of multiple languages from the audio features, and converts the audio feature sequence into a unified modeling unit sequence by combining the common features and the differential features to obtain feature-enhanced acoustic information;
[0046] The unified modeling unit mentioned here can be different unified modeling units for different languages. Specifically, for languages with certain language research knowledge, relatively rich resources (which can be distinguished by expert experience, or by setting quantitative values to represent the amount of resources), and for which Global Phone pronunciation dictionaries can be quickly constructed (such as "major languages" in the popular sense), Global Phone can be used as the unified modeling unit, but is not limited to achieving effective sharing within the multilingual pronunciation space, and reducing the amount of speech data annotation and dependence on expert knowledge; correspondingly, for languages with relatively insufficient language research and fewer resources, Unicode basic spelling elements can be used as the unified modeling unit, but is not limited to reducing the dependence of small languages on pronunciation dictionaries and language expert knowledge, thereby improving the construction efficiency of a large number of resource-scarce languages.
[0047] Furthermore, with regard to the acquisition of common features, the present invention also proposes in some preferred embodiments the idea of adopting transfer learning to improve the information interaction of multilingual speech data, which has the significant effect of improving the recognition effect of relatively low-resource languages. Specifically, in the process of training the multilingual speech recognition model based on the transfer learning mechanism, the trained and converged model of the language with higher resources is used as the initialization model of the language with lower resources. In this way, the similar characteristics of the pronunciation of different languages in time-frequency features can be fully utilized to achieve cross-language information sharing. In actual operation, it can be implemented through a cross-language common feature extraction network. That is, by "replacing small with large", the mature technology and data accumulation of large languages such as Chinese and English can be effectively overflowed and generalized to other small languages, thereby improving the model's effect on speech recognition of low-resource languages.
[0048] Step S4: Based on the acoustic information, the multilingual speech recognition model performs language decoding for different languages and outputs a transcribed text and a language label corresponding to each language;
[0049] Here, regarding the output language label, the following concept can be referred to: in the multilingual speech recognition model, a language routing network is set before language decoding (more preferably, it is also possible to consider adding language losses corresponding to each language to make language attent ion learning more comprehensive and accurate). The language routing network is used to automatically match the language difference information contained in the acoustic information to the corresponding language channel and output the corresponding language label.
[0050] The language routing network mentioned here can perform language adaptive training on multilingual speech recognition models to improve the recognition effect of the corresponding languages. Specifically, the multilingual speech recognition model training process can include two parts. The first part is to use data from all languages to train a general model, insert the language routing network into the general model, and fix the parameters of the trained general model. Only the language routing networks corresponding to different languages are trained, thereby realizing plug-and-play model adaptive training, and further strengthening the information of different languages.
[0051] Based on the above, those skilled in the art can also understand that the character range supported by each language is fixed, so each language is provided with a character encoding set belonging to this language. When performing language decoding, if character decoding is performed in the global space, it is bound to cause crosstalk in the recognition results between languages. Therefore, in order to solve the problem of crosstalk in the recognition results between languages, the present invention proposes, in some preferred embodiments, the use of a language-constrained decoding mechanism, which constrains the recognition results to be within the language based on the language label provided in the previous step, thereby improving the accuracy of speech recognition. Of course, in the training stage, the corresponding constrained encoding and decoding is introduced. The acoustic information provided by the aforementioned acoustic modeling and the language label given by the language routing constrain the encoding and decoding process to a single language to avoid the mixing of multiple languages in the encoding and decoding process.
[0052] In addition, in conjunction with the role of the language tag proposed in this preferred embodiment, to ensure the accuracy and efficiency of the output language tag, in other preferred embodiments of the present invention, in combination with mature voice wake-up mechanisms in human-computer interaction scenarios, a default primary language is pre-determined based on the wake-up voice provided by the user before inputting mixed multilingual voice, and a wake-up tag containing language information is obtained. This wake-up tag is then input into the multilingual voice recognition model to assist in outputting the language tag. For example, but not limited to, the wake-up tag is used to participate in language routing determination and verify the output language tag. It will be understood that this is merely an illustrative example, and its purpose is to make the language tag determination process more reliable and accurate through the additional information provided by the wake-up operation. This is because, generally, when a user uses a certain language to wake up the interactive object during voice interaction, it can be basically determined that the multilingual mixed voice input by the user in subsequent interactions will likely contain the language used in the wake-up operation, thereby improving the accuracy and processing efficiency of language recognition.
[0053] Step S5: Using the transcribed text and the language label to perform semantic understanding and execute interactive operations.
[0054] The semantic understanding process described in the present invention itself can draw on mature methods in the field, and specifically, it can also train corresponding semantic understanding models and engines for each language, and automatically select and call corresponding semantic understanding services based on the language labels adaptively output by the recognition model, including but not limited to obtaining interaction intentions, semantic slots, etc. Then, the semantic understanding results and language labels can be sent to the interactive end together. The interactive end can complete corresponding interactive display and other actions based on the language labels and semantic understanding, such as but not limited to calling speech synthesis in different languages, displaying interfaces in different languages, etc. The present invention does not elaborate on or limit the application of the subsequent output results.
[0055] In summary, the main idea of the present invention is to avoid training a recognition model for each language separately, but to jointly train a multilingual speech recognition model with data from multiple languages, and to achieve seamless multilingual switching-free voice interaction based on mixed modeling of multilingual common features. Specifically, the input mixed language speech audio features are fed into the end-to-end multilingual speech recognition model, from which the common features and difference features of the multilinguals are obtained, and the two are combined for acoustic modeling and language decoding, and the transcribed text and language labels corresponding to each language are output, and finally the transcribed text and language labels are used for semantic understanding and to perform interactive operations. The present invention does not need to rely on manual switching, and eliminates the differences between different languages in speech recognition, semantic understanding, and speech synthesis. In particular, it does not require switching, and directly performs comprehensive recognition and understanding of the mixed language voice interaction, thereby significantly improving the human-computer interaction experience.
[0056] Corresponding to the above embodiments and preferred solutions, the present invention also provides an embodiment of a multi-language switching-free interactive device, such as Figure 2 As shown, it may specifically include the following components:
[0057] Audio feature extraction module 1, used to extract audio features of input speech in different languages;
[0058] Feature input module 2, configured to input the audio features into a pre-trained multilingual speech recognition model, wherein the multilingual speech recognition model adopts an end-to-end modeling mechanism;
[0059] an acoustic information processing module 3, configured to obtain common features and differential features of multiple languages from the audio features by the multilingual speech recognition model, and convert the audio feature sequence into a unified modeling unit sequence by combining the common features and the differential features to obtain feature-enhanced acoustic information;
[0060] A language decoding processing module 4 is configured to perform language decoding for different languages using the multilingual speech recognition model based on the acoustic information, and output a transcribed text and a language label corresponding to each language;
[0061] The semantic understanding and interaction module 5 is used to use the transcribed text and the language label to perform semantic understanding and perform interactive operations.
[0062] It should be understood that the above Figure 2The division of the various components in the multi-language switching-free interactive device shown is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. And these components can all be implemented in the form of software calling through processing elements; they can also all be implemented in the form of hardware; some components can also be implemented in the form of software calling through processing elements, and some components can be implemented in the form of hardware. For example, one of the above modules can be a separately established processing element, or it can be integrated in a chip of an electronic device. The implementation of other components is similar. In addition, all or part of these components can be integrated together, or they can be implemented independently. In the implementation process, each step of the above method or the above components can be completed by the hardware integrated logic circuit in the processor element or the instructions in the form of software.
[0063] For example, the above components may be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs). For another example, these components may be integrated together to form a system-on-a-chip (SOC).
[0064] Based on the above embodiments and their preferred solutions, those skilled in the art will appreciate that, in actual operation, the technical concepts involved in the present invention can be applied to a variety of implementations. The present invention uses the following carriers as schematic illustrations:
[0065] (1) An electronic device. The device may specifically include: one or more processors, a memory, and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions. When the instructions are executed by the device, the device performs the steps / functions of the aforementioned embodiment or an equivalent embodiment.
[0066] The electronic device may specifically be an electronic device related to a computer, such as but not limited to various interactive terminals, electronic products, mobile terminals, and the like.
[0067] Figure 3This is a schematic structural diagram of an embodiment of an electronic device provided by the present invention. Specifically, the electronic device 900 includes a processor 910 and a memory 930. The processor 910 and the memory 930 can communicate with each other through an internal connection path to transmit control and / or data signals. The memory 930 is used to store computer programs, and the processor 910 is used to call and run the computer program from the memory 930. The above-mentioned processor 910 and the memory 930 can be combined into a processing device, or more commonly, they are independent components. The processor 910 is used to execute the program code stored in the memory 930 to implement the above-mentioned functions. In specific implementation, the memory 930 can also be integrated into the processor 910, or be independent of the processor 910.
[0068] In addition, to further improve the functionality of the electronic device 900, the device 900 may further include one or more of an input unit 960, a display unit 970, an audio circuit 980, a camera 990, and a sensor 901. The audio circuit may further include a speaker 982, a microphone 984, etc. The display unit 970 may include a display screen.
[0069] Furthermore, the device 900 may further include a power supply 950 for providing electrical energy to various devices or circuits in the device 900 .
[0070] It should be understood that the operation and / or function of each component in the device 900 can be specifically referred to the description of the embodiments of the method, system, etc. in the above text. To avoid repetition, the detailed description is appropriately omitted here.
[0071] It should be understood that Figure 3 The processor 910 in the electronic device 900 shown can be a system on a chip SOC, which can include a central processing unit (CPU) and can further include other types of processors, such as a graphics processing unit (GPU), etc., which will be described in detail below.
[0072] In summary, the various processors or processing units within the processor 910 can work together to implement the previous method flow, and the corresponding software programs of the various processors or processing units can be stored in the memory 930.
[0073] (2) A computer data storage medium having a computer program or the aforementioned apparatus stored thereon, which, when executed, causes a computer to execute the steps / functions of the aforementioned embodiments or equivalent implementations.
[0074] In the several embodiments provided herein, any function, if implemented as a software functional unit and sold or used as an independent product, may be stored on a computer data storage medium. Based on this understanding, certain technical solutions of the present invention, or portions that contribute to the prior art, or portions of such solutions, may be embodied in the form of software products as described below.
[0075] It should be particularly noted that the storage medium may refer to a server or a similar computer device, specifically, a storage device in the server or a similar computer device storing the aforementioned computer program or the aforementioned apparatus.
[0076] (3) A computer program product (which may include the above-mentioned apparatus) which, when running on a terminal device, enables the terminal device to execute the multi-language switching-free interaction method of the aforementioned embodiment or an equivalent implementation.
[0077] Through the description of the above implementation methods, it can be seen that those skilled in the art can clearly understand that all or part of the steps in the above implementation method can be implemented by means of software plus the necessary general hardware platform. Based on this understanding, the above computer program product may include but is not limited to an APP.
[0078] Continuing from the above, the aforementioned device / terminal may be a computer device, and the hardware structure of the computer device may further specifically include: at least one processor, at least one communication interface, at least one memory, and at least one communication bus; the processor, communication interface, and memory may all communicate with each other via the communication bus. The processor may be a central processing unit (CPU), a DSP, a microcontroller, or a digital signal processor, and may also include a GPU, an embedded neural network processor (NPU), and an image signal processor (ISP). The processor may also include a specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. Furthermore, the processor may have the function of operating one or more software programs, which may be stored in a storage medium such as a memory. The aforementioned memory / storage medium may include: non-volatile memory (such as a non-removable disk, a USB flash drive, a mobile hard drive, an optical disk), read-only memory (ROM), random access memory (RAM), and the like.
[0079] In the embodiment of the present invention, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can be represented by: a, b, c, a and b, a and c, b and c, or a, b and c, where a, b, c can be single or multiple.
[0080] Those skilled in the art will appreciate that the various modules, units, and method steps described in the embodiments disclosed in this specification can be implemented using a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0081] Furthermore, the modules and units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple locations, such as nodes in a system network. Part or all of the modules and units may be selected based on actual needs to achieve the objectives of the above-described embodiments. Those skilled in the art can understand and implement the above-described embodiments without inventive effort.
[0082] The above describes in detail the structure, features and effects of the present invention based on the embodiments shown in the drawings, but the above is only a preferred embodiment of the present invention. It should be noted that the technical features involved in the above embodiments and their preferred modes can be reasonably combined and matched into a variety of equivalent schemes by those skilled in the art without departing from or changing the design ideas and technical effects of the present invention; therefore, the scope of implementation of the present invention is not limited to what is shown in the drawings. Any changes made in accordance with the concept of the present invention, or modifications to equivalent embodiments with equivalent changes, which still do not exceed the spirit covered by the description and drawings, should be within the scope of protection of the present invention.
Claims
1. A multi-language switching-free interaction method, characterized in that: include: Correspondingly extracting audio features of input speech in different languages; Inputting the audio features into a pre-trained multilingual speech recognition model, wherein the multilingual speech recognition model adopts an end-to-end modeling mechanism; The multilingual speech recognition model obtains common features and differential features of multiple languages from the audio features, and converts the audio feature sequence into a unified modeling unit sequence based on the common features and the differential features to obtain feature-enhanced acoustic information; wherein different unified modeling units are used for multiple languages with different resource amounts; Based on the acoustic information, the multilingual speech recognition model performs language decoding for different languages and outputs transcribed text and language labels corresponding to each language; The transcribed text and the language label are used to perform semantic understanding and perform interactive operations.
2. The multi-language switching-free interaction method according to claim 1, characterized in that: The multilingual speech recognition model is trained based on a transfer learning mechanism, and a trained and converged model of a language with higher resources is used as an initialization model for a language with lower resources to obtain similar characteristics of pronunciations of different languages in terms of time-frequency features.
3. The multi-language switching-free interaction method according to claim 1, characterized in that: In the multilingual speech recognition model, a language routing network is set before language decoding. The language routing network is used to automatically match the corresponding language channel based on the language difference information contained in the acoustic information and output the corresponding language label.
4. The multi-language switching-free interaction method according to claim 3, characterized in that: The training process of the multilingual speech recognition model includes: Use data from all languages to train a universal model; Inserting the language routing network into the trained universal model; The general model parameters are fixed, and only the language routing network is trained according to each language.
5. The multi-language switching-free interaction method according to claim 3, characterized in that: During the language decoding process, constraint decoding is performed for each language based on the language label output by the language routing network.
6. The multi-language switching-free interaction method according to any one of claims 1 to 5, characterized in that: A wake-up tag containing language information is obtained by using the wake-up speech provided by the user before inputting the mixed multilingual speech. The wake-up tag is used to assist the multilingual speech recognition model in outputting the language tag.
7. A multi-language switching-free interactive device, characterized in that: include: Audio feature extraction module, used to extract audio features of input speech in different languages; a feature input module, configured to input the audio features into a pre-trained multilingual speech recognition model, wherein the multilingual speech recognition model adopts an end-to-end modeling mechanism; an acoustic information processing module configured to obtain common features and differential features of multiple languages from the audio features using the multilingual speech recognition model, and convert the audio feature sequence into a unified modeling unit sequence based on the common features and differential features to obtain feature-enhanced acoustic information; wherein different unified modeling units are used for multiple languages with different resource amounts; A language decoding processing module, configured to perform language decoding for different languages using the multilingual speech recognition model based on the acoustic information, and output a transcribed text and language label corresponding to each language; The semantic understanding and interaction module is used to use the transcribed text and the language label to perform semantic understanding and perform interactive operations.
8. An electronic device, characterized in that: include: One or more processors, a memory, and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the electronic device, cause the electronic device to execute the multilingual switching-free interaction method according to any one of claims 1 to 6.
9. A computer data storage medium, characterized in that The computer data storage medium stores a computer program, and when the computer program is run on a computer, the computer is caused to execute the multi-language switching-free interaction method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Multilingual speech recognition model training method and device thereof, equipment and storage medium
CN111833845A
Hybrid speech recognition method and device, electronic equipment and storage medium
CN114694637A