Multi-modal language conversion method and related product

Through the multimodal language conversion method, the input data is encoded and decoded by using the multimodal encoding unit and the decoding unit, which solves the problem of difficult to take into account both real-time and accuracy in traditional speech translation systems, and realizes efficient and accurate multimodal language conversion, which is suitable for multi-person and multi-language dialogue scenarios.

CN120375833APending Publication Date: 2025-07-25HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510872931.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing voice translation system is difficult to ensure real-time and translation accuracy at the same time in multi-person conversations or multi-lingual scenarios. Traditional chain processing methods lead to information loss and delay problems.

Method used

The multimodal language conversion method is adopted, and the input data is encoded by the multimodal encoding unit, vector representation is generated, and decoded by multiple decoding units. Each decoding unit is trained by a different basic decoding unit, responsible for specific language conversion tasks, and supports any combination of text and audio input.

Benefits of technology

It improves the accuracy and consistency of translation and pronunciation synthesis, reduces information loss and cumulative errors, is suitable for multi-person multi-language dialogue environments, taking into account real-time and translation quality, and is especially suitable for real-time communication and communication needs in multi-lingual and multi-scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375833A_ABST
    Figure CN120375833A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal language conversion method and a related product. The input data is coded through a multi-mode coding unit in the multi-mode decoding unit, vector representation of the input data is obtained, the vector representation is decoded through a plurality of decoding units in the multi-mode decoding unit, and target data obtained after language conversion is conducted on the input data is obtained and output. Wherein the plurality of decoding units are respectively obtained by carrying out association training on different basic decoding units, and each decoding unit is related to a multi-modal cross-language understanding function. By adopting the multi-mode coding unit and the multi-mode decoding unit, the problems of information loss and time delay caused by a chained processing mode in a traditional speech translation system are solved, and the balance between real-time performance and accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a multi-modal language conversion method and related products. Background Art

[0002] Current speech translation systems usually include multi-modal inputs in different forms or types, such as text and audio information, to fully capture language and speech features. However, these systems generally rely on a chain processing method, which usually requires multiple steps (such as voiceprint recognition, translation, speech synthesis, etc.), resulting in information loss and latency problems. Especially in multi-person conversations or multi-language scenarios, traditional chain processing solutions are difficult to ensure sufficient real-time performance and output accuracy.

[0003] In addition, multi-language translation problems may also occur in multi-person conversation scenarios. Current technologies often face a trade-off between translation accuracy and real-time performance, that is, improving real-time performance may sacrifice translation accuracy, and vice versa. Summary of the Invention

[0004] Based on the above problems, this application provides a multi-modal language conversion method and related products, aiming to improve the real-time performance of translation while ensuring translation accuracy.

[0005] The embodiments of this application disclose the following technical solutions: A multi-modal language conversion method is applied to a multi-modal conversion model, and the multi-modal conversion model includes a multi-modal encoding unit and a multi-modal decoding unit. The method includes: In response to receiving input data, use the multi-modal encoding unit to encode the input data to obtain a vector representation of the input data; the type of the input data includes at least one of text data and audio data; Use multiple decoding units in the multi-modal decoding unit to respectively perform their own decoding processes on the vector representation to obtain target data after language conversion of the input data and output it; the multiple decoding units are respectively obtained by associative training of different basic decoding units, and each decoding unit is related to the multi-modal cross-language understanding function.

[0006] A multi-modal language conversion device, the device includes: An encoding processing unit, in response to receiving input data, is configured to use the multi-modal encoding unit to encode the input data to obtain a vector representation of the input data; the type of the input data includes at least one of text data and audio data; A decoding processing unit is used to use multiple decoding units in the multimodal decoding unit to perform respective decoding processing on the vector representation, obtain target data after the input data is language converted and output; the multiple decoding units are obtained by performing associated training on different basic decoding units, and each of the decoding units is related to the multimodal cross-language understanding function.

[0007] An electronic device comprises a memory, a processor and a machine executable program stored in the memory and running on the processor, wherein when the processor executes the machine executable program, the electronic device executes the multimodal language conversion method as described above.

[0008] Compared with the prior art, this application has the following beneficial effects: When receiving input data in text form or audio form, the embodiment of the present application firstly performs preliminary processing through the multimodal encoding unit when receiving the input data, whether it is text data or audio data, and converts it into a vector representation. This process can effectively capture the semantic and structural information of the input data. Subsequently, multiple decoding units in the multimodal decoding unit are used to decode these vector representations respectively. Each decoding unit is based on different basic decoding units and is obtained through special association training. They are each responsible for a specific language conversion task and are closely related to the multimodal cross-language understanding function. Finally, after decoding processing, the converted target data is obtained and output, thereby completing the entire language conversion process. In the present application, multiple basic decoding units are trained through association, and each undertakes different decoding tasks, so as to realize the multi-angle interpretation of the same vector representation, enhance the cross-language understanding ability and conversion performance of the model, and improve the accuracy and consistency of translation and synthesis. At the same time, by supporting any combination of text and audio input, the multimodal conversion model has a wider scope of application, is particularly suitable for multi-person multi-language dialogue environments, effectively takes into account real-time and translation quality, and alleviates the problem that real-time and accuracy cannot be taken into account in traditional solutions. In addition, multiple input types such as text and audio are directly encoded uniformly through multimodal encoding units, avoiding multi-step conversion in traditional chain processing and reducing information loss and cumulative errors. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor. Figure 1 A schematic diagram of an application scenario of a modal language conversion method provided in an embodiment of the present application; Figure 2 Flow chart of a multimodal language conversion method provided by an embodiment of the present application; Figure 3 Structural schematic diagram of a multimodal conversion model provided by an embodiment of the present application; Figure 4 Flow chart of a speech synthesis training method provided by an embodiment of the present application; Figure 5 Flow chart of a language translation training method provided by an embodiment of the present application; Figure 6 Flow chart of a cross - language conversion training method provided by an embodiment of the present application; Figure 7 Flow chart of a text - audio alignment training method provided by an embodiment of the present application; Figure 8 Flow chart of a timbre alignment training method provided by an embodiment of the present application; Figure 9 Flow chart of a language alignment training method provided by an embodiment of the present application; Figure 10 Application scenario schematic diagram of a multimodal decoding unit provided by an embodiment of the present application; Figure 11 Application scenario schematic diagram of a multimodal encoding unit provided by an embodiment of the present application; Figure 12 Schematic diagram of a multimodal language conversion device provided by an embodiment of the present application; Figure 13 Structural schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0010] To facilitate understanding of the technical solutions provided by the embodiments of the present application, the background technologies related to the embodiments of the present application will be described first below.

[0011] Current speech translation systems usually involve multimodal input forms, such as text and audio information, aiming to comprehensively capture language and speech features. However, these systems usually adopt a chain - type processing method, including multiple links such as voiceprint recognition, translation, and speech synthesis, which are prone to information loss and high processing delays during information transmission. In a multi - person conversation or multilingual environment, this traditional method is difficult to meet the requirements of real - time response and high accuracy at the same time, and often faces the contradiction that the translation quality decreases when improving real - time performance, or the response speed is sacrificed when pursuing high accuracy.

[0012] Based on this, the embodiment of the present application provides a multimodal language conversion method and related products, first using a multimodal encoding unit to encode the received input data to obtain a vector representation of the input data. The input data can be text data, audio data, or a combination of the two. Then, multiple decoding units in the multimodal decoding unit are used to perform respective decoding processes on the vector representation, and after decoding, the target data after the input data is language converted is generated and output. Each decoding unit is obtained by association training of different basic decoding units, and each decoding unit is related to a multimodal cross-language understanding function. In this process, not only is information loss and latency reduced, but the overall quality of translation and speech generation is also significantly improved, which is particularly suitable for complex scenarios with real-time, multilingual, and multiple speakers.

[0013] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0014] See also Figure 1 , Figure 1 A schematic diagram of an application scenario of a modal language conversion method provided in an embodiment of the present application. Figure 1 As shown, the application scenario includes three users A, B and C, each of whom is equipped with a multimodal language conversion device 110. The device integrates multiple key modules, among which the communication unit 111 is responsible for voice collection and information transmission, the preprocessing unit 112 is responsible for the normalization of input data, the multimodal conversion model 113 realizes the conversion function between languages, and the output display unit 114 is used to present the converted text and audio.

[0015] In this embodiment, user A expresses in spoken Japanese "私は果物を食べるのが好きです" (which means "I love to eat fruit"), and at the same time, user B inputs the Chinese text "我爱大模" through the device. At this time, the communication unit 111 in the device of user B and user C receives and obtains the Japanese voice signal spoken by user A, ensuring real-time sharing of information. At the same time, the communication unit 111 of user B transmits the Chinese content it inputs to the devices of user A and user C, so that the three parties can achieve effective information exchange and understanding in a multilingual and multimodal environment, demonstrating the powerful cross-language collaboration capability and efficient data interaction mechanism of the multimodal language conversion device 110.

[0016] This embodiment first explains the working process of the multimodal language conversion device 110 held by user C: Specifically, the main task of the preprocessing unit 112 is to normalize and structure the received input data, so as to provide accurate and model-friendly input information for subsequent multimodal conversion. The input data in this embodiment is explained by taking the text "I love the big model" and the audio "I also like eating fruits" as examples. The preprocessing unit 112 first performs a word segmentation operation on this continuous text to split it into more fine-grained word units, namely "I / love / big / model". At the same time, speech feature extraction (such as Mel Frequency Cepstral Coefficients (MFCC) feature extraction) is performed on "I also like eating fruits" to obtain a Mel spectrum, that is, "a two-dimensional matrix representing the energy distribution of the outputs of each Mel filter group of the speech signal at different time frames, which is used to reflect the spectral characteristics and time-varying characteristics of the speech."

[0017] Afterwards, the multimodal encoding unit in the multimodal conversion model 113 can be used to jointly process the two types of preprocessed data, namely, the word segmentation sequence "I / love / big / model" obtained in the preprocessing stage and the corresponding Mel spectrum. The encoding unit not only captures the semantic information of the text, but also combines the speech characteristics in the audio through in-depth extraction and fusion of text and audio features, thereby generating a comprehensive vector representation. This vector representation effectively integrates the language content and sound features, provides a rich and accurate input basis for subsequent multimodal decoding and language conversion, and improves the performance of the multimodal language conversion device 110 in a cross-modal and multi-language environment. In addition, the multimodal decoding unit continuously optimizes its adaptability to speech synthesis, language translation, and cross-language conversion tasks in the process of gradual training of the basic decoding unit, ensuring that the multimodal encoding unit can accurately convert the fused vector information into text and natural and fluent audio output in the target language. This decoding mechanism based on multi-stage training not only improves the conversion quality, but also enhances the generalization ability and real-time response performance of the system in complex multimodal and multi-language environments.

[0018] Furthermore, the multimodal decoding unit in the multimodal conversion model 113 is used to perform decoding processing on the previously generated vector representation. This decoding unit takes the vector that fuses text and audio features as input and can accurately analyze the semantic information contained therein, thereby generating the conversion result in the target language. In this embodiment, the multimodal decoding unit not only successfully maps the vector of the tokenized sequence to the English text "I love large models", but also synchronously generates a natural and fluent speech signal file "audio_I_love_the_large_model.wav" that highly matches the text content; at the same time, it also accurately converts the vector representation of the mel spectrogram to the English text "I love fruits" and generates the corresponding speech file "audio_I_love_fruits.wav". This dual output mode of text and speech not only ensures the accuracy of language conversion, but also significantly improves the naturalness and expressiveness of speech synthesis, making the conversion result more vivid and credible. In addition, the multimodal encoding unit itself is obtained through a series of training processes gradually completed by the basic encoding unit, including text-audio alignment training, timbre alignment training, and language alignment training. These training steps effectively enhance the model's ability to synchronously understand and fuse different modal information, making the finally generated vector representation both accurate and rich in multimodal features, providing a solid foundation for the subsequent decoding stage, and thus achieving high-quality cross-language and multimodal conversion effects.

[0019] Finally, the target language conversion result generated by the multimodal decoding unit will be presented by the output display unit 114. Specifically, the output display unit 114 can not only play the generated target language audio files "audio_I_love_the_large_model.wav" and "audio_I_love_fruits.wav", enabling users to intuitively hear the natural and fluent speech synthesis content with accurate semantics, but also clearly display the corresponding conversion texts "I love large models" and "I love fruits" on the interface for users to read and check. This dual output mode not only meets the dual needs of users for visual and auditory information, but also enhances the overall interaction experience and operation convenience. By presenting text and audio synchronously, the output display unit effectively completes the complete closed-loop from input data to target language expression, and transmits the conversion results of the model to the final user in an intuitive and easy-to-understand form, ensuring the accuracy and real-time nature of information transmission, and is particularly suitable for real-time communication and communication needs in multiple languages and scenarios.

[0020] Those skilled in the art can understand, Figure 1The schematic diagram of the framework shown is only an example in which the embodiments of the present application can be implemented. The scope of application of the embodiments of the present application is not limited by any aspect of this framework.

[0021] To facilitate the understanding of the present application, a multi-modal language conversion method provided by an embodiment of the present application will be described below with reference to the accompanying drawings.

[0022] See Figure 2 As shown, this figure is a flowchart of a multi-modal language conversion method provided by an embodiment of the present application. This method is applied to a multi-modal conversion model including a multi-modal encoding unit and a multi-modal decoding unit. This model uniformly encodes the input multi-modal data and combines the collaborative work of multiple decoding units to achieve efficient conversion and understanding of different language and modal information.

[0023] As Figure 2 shown, this method may include S201 - S203: S201: In response to receiving the input data, use the multi-modal encoding unit to perform encoding processing on the input data to obtain a vector representation of the input data; the type of the input data includes at least one of text data and audio data.

[0024] In step S201, in response to the received input data, perform in-depth feature extraction and fusion encoding processing (i.e., encoding processing) on the input data, so as to generate a unified vector representation containing rich semantic and acoustic information. The input data may include text data, audio data, or text data + audio data. Through the processing of the multi-modal encoding unit, different types of modal information are uniformly mapped into the same feature space, realizing the effective fusion and expression of multi-source information, providing rich and structured semantic representations for the subsequent multi-modal decoding unit, and thus supporting efficient cross-language conversion and understanding.

[0025] S202: Use multiple decoding units in the multi-modal decoding unit to perform respective decoding processing on the vector representation to obtain the target data after language conversion of the input data and output it; the multiple decoding units are respectively obtained by associatively training different basic decoding units, and each decoding unit is related to the multi-modal cross-language understanding function.

[0026] In step S202, multiple decoding units in the multimodal decoding unit respectively perform their own decoding processes on the vector representation, thereby generating corresponding language conversion texts and audio (i.e., target data). Specifically, the multimodal decoding unit includes multiple base decoding units obtained through associated training. These decoding units respectively perform independent decoding processes on the vector representation obtained by encoding the input data. Through this parallel and collaborative decoding method, each decoding unit can deeply understand the input information from different perspectives and modalities, and jointly complete the language conversion task. Each decoding unit focuses on a certain aspect of multimodal cross-language understanding (such as speech decoding, text decoding, language decoding, or timbre decoding). Through associated training, information sharing and functional complementarity are achieved, thereby improving the conversion accuracy of the target data. Finally, these conversion texts and audio are output for the user to view or further use. This step not only completes the conversion task from the source language to the target language, but also provides a complete dual output of text and audio, meeting the requirements of multimodal information processing.

[0027] Based on the descriptions of S201 - S203, an embodiment of the present application provides a multimodal language conversion method, which is applied to a multimodal conversion model including a multimodal encoding unit and a multimodal decoding unit. Specifically, when receiving input data including text data or audio data, first, the multimodal encoding unit encodes the input data to generate a corresponding vector representation. Subsequently, multiple decoding units in the model respectively use these vector representations for independent decoding processes, and each decoding unit is obtained through associated training by different base decoding units. In this way, each decoding unit can achieve cross-language understanding and conversion for information of different modalities, and finally output the corresponding target language data, realizing efficient and accurate multimodal language conversion. The present application obtains multiple decoding units related to the multimodal cross-language understanding function through associated training of different base decoding units respectively, strengthens the fusion and conversion ability between different modalities, significantly improves the translation accuracy and the naturalness of speech synthesis, effectively solves the contradiction between real-time performance and accuracy in multi-person and multi-language conversations, reduces information loss and latency, and is applicable to complex real-time multi-language scenarios.

[0028] In a possible implementation manner, the multiple decoding units include a speech decoding unit and a text decoding unit. The speech decoding unit is responsible for decoding the vector representation of the input data to generate and output speech conversion data; the text decoding unit independently decodes the same vector representation to generate and output corresponding text conversion data. This way of multiple decoding units working together can achieve efficient and diverse language conversion of the input multimodal data.

[0029] In a possible implementation, step S202 uses multiple decoding units in the multimodal decoding unit to separately perform respective decoding processes on the vector representation, obtain the target data after language conversion of the input data, and output it, including: A1: Use the speech decoding unit to perform decoding processing on the vector representation, obtain the speech conversion data in the target data, and output it; A2: Use the text decoding unit to perform decoding processing on the vector representation, obtain the text conversion data in the target data, and output it.

[0030] The specific process is as follows: Speech decoding processing: During the speech decoding process, the speech decoding unit converts the encoded vector representation into speech data in a specific language. To complete this conversion, the speech decoding unit uses the directly set default target language information and conversion tone color information. These information ensure that the generated speech data not only meets the requirements of the target language but also can simulate specific tone color characteristics. Finally, the converted speech data is output.

[0031] Text decoding processing: Use the text decoding unit to perform decoding processing on the vector representation. Similarly, the text decoding unit converts the encoded vector representation into text data in a specific language. Similarly, the text decoding unit also uses the directly set default target language information to ensure that the generated text data meets the requirements of the target language. The output result is the converted text data.

[0032] It should be noted that here, the speech decoding unit and the text decoding unit can not rely on externally dynamically adjusted target language or tone color parameters, but use internally preset default values to ensure the stability and consistency of the decoding process. This design enables the two decoding units to process multimodal information in a coordinated manner, realizes the unified conversion requirements for multiple languages and multiple tone colors, and at the same time supports multimodal output forms, thereby effectively improving the accuracy of language conversion and the flexibility of applications.

[0033] Through the above steps, the multimodal decoding unit can separately process speech and text data, thereby generating and outputting the target data after language conversion. The speech decoding unit and the text decoding unit respectively use the directly set default target language information and conversion tone color information to ensure the accuracy and consistency of the conversion results. This method not only improves the flexibility and practicality of the system but also ensures the efficient processing of different modal data during cross-language conversion.

[0034] In a possible implementation, the plurality of decoding units further include: a language decoding unit. The language decoding unit is responsible for decoding the vector representation of the input data to generate target language information. The target language can assist the speech decoding unit and the text decoding unit to perform more accurate decoding processing on the vector representation in combination with the specific language context, so as to output the corresponding speech conversion data and text conversion data respectively. By introducing the language decoding unit, the model can better support the language conversion requirements in a multilingual environment and improve the flexibility of the model and the accuracy of the conversion effect.

[0035] In a possible implementation, step A1 uses the speech decoding unit to decode the vector representation to obtain and output the speech conversion data in the target data, including: First, obtain the target language information provided by the language decoding unit; subsequently, in combination with the target language information, the speech decoding unit decodes the input vector representation, thereby generating speech conversion data containing target language features and using it as the output result. By introducing the target language information, the speech decoding unit can more accurately complete cross-language speech conversion and improve the naturalness and accuracy of the conversion.

[0036] In a possible implementation, step A2 uses the text decoding unit to decode the vector representation to obtain and output the text conversion data in the target data, including: First, obtain the target language information provided by the language decoding unit; subsequently, in combination with the target language information, the text decoding unit decodes the input vector representation, thereby generating the corresponding text conversion data and using it as the output result. By introducing the target language information, the text decoding unit can more accurately complete cross-language text conversion and improve the accuracy of translation and semantic consistency.

[0037] In a possible implementation, the plurality of multimodal decoding units further include: a timbre decoding unit. The timbre decoding unit is responsible for decoding the input vector representation to extract or generate the timbre information required for target conversion. The timbre information can be used to guide the decoding processes of the speech decoding unit and the text decoding unit, so that the output speech and text conversion data not only accurately reflect the target language content, but also maintain or adjust specific timbre characteristics, thereby enhancing the personalized performance ability and natural interaction experience of the multimodal conversion model.

[0038] In a possible implementation, step A1 uses the speech decoding unit to decode the vector representation to obtain and output the speech conversion data in the target data, including: First, obtain the converted timbre information provided by the timbre decoding unit; subsequently, in combination with this converted timbre information, the speech decoding unit decodes the input vector representation, thereby generating speech conversion data with specified timbre characteristics and using it as the output result. By introducing the converted timbre information, the speech decoding unit can more accurately control the timbre of speech synthesis, making the converted speech more in line with the expected sound style and user requirements.

[0039] In a possible implementation manner, step A2 uses the text decoding unit to decode the vector representation to obtain the text conversion data in the target data and output it, including: First, obtain the converted timbre information in the timbre decoding unit; subsequently, in combination with this converted timbre information, the text decoding unit decodes the input vector representation, thereby generating text conversion data containing corresponding language content and outputting it. By introducing the converted timbre information, the text decoding unit can more accurately reflect the correlation between speech features and text content, improving the naturalness and expression consistency of the text conversion result.

[0040] In a possible implementation manner, the embodiment of the present application also provides a structural schematic diagram of a multimodal conversion model, as Figure 3 shown. The multimodal conversion model includes a multimodal encoding unit 301 and a multimodal decoding unit 302. The multimodal decoding unit 302 includes a timbre decoding unit 3021, a speech decoding unit 3022, a text decoding unit 3023, and a language decoding unit 3024. The input ends of the speech decoding unit 3022 and the text decoding unit 3023 are both connected to the output end of the multimodal encoding unit; the output end of the timbre decoding unit 3021 is respectively connected to the input ends of the speech decoding unit 3022 and the text decoding unit 3023; the output end of the language decoding unit 3024 is respectively connected to the input ends of the speech decoding unit 3022 and the text decoding unit 3023.

[0041] Each sub-unit included in the multimodal decoding unit 302 is responsible for specific information processing tasks to ensure that the final conversion result is both accurate and natural. Specifically: Timbre decoding unit 3021: This unit is used to provide the converted timbre information for the speech decoding unit 3022 and the text decoding unit 3023. The timbre decoding unit 3021 generates timbre characteristics for speech synthesis to ensure that the converted audio has the required timbre characteristics.

[0042] Language decoding unit 3024: This unit is used to provide the target language information for the speech decoding unit 3022 and the text decoding unit 3023. The language decoding unit 3024 can transfer the grammar and vocabulary characteristics of the target language to ensure that the converted text and audio conform to the norms of the target language.

[0043] Voice decoding unit 3022: Based on the converted timbre information provided by the timbre decoding unit 3021 and the target language information provided by the language decoding unit 3024, this unit performs voice synthesis on the vector representation to generate converted audio. This process comprehensively considers the characteristics of timbre and language to generate natural and fluent voice signals.

[0044] Text decoding unit 3023: Similarly based on the converted timbre information provided by the timbre decoding unit 3021 and the target language information provided by the language decoding unit 3024, this unit performs text decoding on the vector representation to generate converted text. This process not only ensures the accurate translation of the text content but also considers the influence of timbre and language to improve the naturalness and accuracy of the output text.

[0045] Through the collaborative work of these sub-units, it is possible to efficiently generate high-quality target language text and natural and fluent audio output, meeting the requirements of multi-modal information processing.

[0046] In a possible implementation, the basic decoding unit includes a basic timbre decoding unit, a basic voice decoding unit, a basic text decoding unit, and a basic language decoding unit. Multiple decoding units are obtained by performing associated training on these basic decoding units. The associated training process mainly includes voice synthesis training, language translation training, and cross-language conversion training. Voice synthesis training aims to enhance the accurate expression ability of the basic timbre decoding unit and the basic voice decoding unit for timbre and voice features; language translation training focuses on strengthening the multi-language understanding and translation ability of the basic text decoding unit; cross-language conversion training promotes the information interaction and collaborative optimization between the basic language decoding unit and other decoding units to achieve efficient conversion between different languages. Through the above-mentioned associated training, multiple decoding units can complement each other and work together, thus significantly improving the overall performance and adaptability of the multi-modal language conversion model.

[0047] Among them, the basic timbre decoding unit: Its main function is to process and generate timbre information in the audio. During the training process, it needs to learn how to predict or generate corresponding timbre information according to the specific features of the input.

[0048] Basic voice decoding unit: This decoding unit is mainly responsible for converting the input vector representation into voice output. During the training stage, it optimizes the quality of voice synthesis by learning timbre and language information.

[0049] Basic text decoding unit: The task of this decoding unit is to decode the input vector into text format. During the training process, it needs to combine timbre and language information to accurately decode the text.

[0050] Basic language decoding unit: This unit is specifically used to process information in different languages, that is, to identify and generate text or speech in the target language. During training, it needs to learn how to extract the correct language information from the input and apply it to the output.

[0051] These basic decoding units are continuously optimized during the training process. They are specifically trained through different training data sets and loss functions respectively, and finally integrated into a multi-modal decoding unit that can process cross-modal inputs and output target language text and audio.

[0052] In one possible implementation, the present application provides a flowchart of a speech synthesis training method. Refer to Figure 4 , Figure 4 which is a flowchart of a speech synthesis training method provided by an embodiment of the present application. Specifically, the speech synthesis training process of the basic decoding unit can be implemented through steps S401 - S403: S401: Construct a first training data set.

[0053] In the stage of the speech synthesis training, it is crucial to construct a first training data set. This training data set covers speech vector samples with text vector labels. Each speech vector sample in this data set is equipped with a corresponding text vector label, and each speech vector sample and its corresponding text vector label belong to the same language. This is done to ensure the consistency and accuracy of the training data, so that the basic decoding unit can better learn the correspondence between speech and text and the characteristics of a specific language during the training process. In this way, the accuracy and naturalness of the speech synthesis system can be effectively improved.

[0054] S402: Input the first training data set into the basic timbre decoding unit and the basic language decoding unit respectively to obtain a first timbre feature representation and a first language feature representation.

[0055] When processing the first training data, the basic timbre decoding unit extracts feature information related to timbre by analyzing the sound signals in the training data and generates a first timbre feature representation for describing timbre attributes; at the same time, the basic language decoding unit identifies and extracts the language category and its related features for the same training data set to form a corresponding first language feature representation. This process lays the foundation for subsequent multi-modal association training and realizes independent and effective encoding of timbre and language information.

[0056] S403: Use the first training data set, the first voice feature representation, and the first language feature representation as inputs simultaneously and input them into the basic speech decoding unit and the basic text decoding unit respectively for joint decoding and error calculation. When it is determined that the text output error of the basic text decoding unit is less than or equal to the first error threshold, stop the training to obtain the speech synthesis decoding unit.

[0057] After obtaining the first voice feature representation and the first language feature representation, use the first training data set, the first voice feature representation, and the first language feature representation as inputs for the basic speech decoding unit and the basic text decoding unit simultaneously, and perform joint decoding and error calculation on the basic speech decoding unit and the basic text decoding unit. During this process, the basic speech decoding unit and the basic text decoding unit cooperate to generate corresponding speech and text outputs according to the input information, and compare them with the expected results to calculate the error. When it is detected that the text output error of the basic text decoding unit is less than or equal to the preset first error threshold, it indicates that the model has reached the desired accuracy level. At this time, stop the training process to finally obtain an optimized and improved speech synthesis decoding unit. Through this joint training method, the consistency and conversion quality of speech and text decoding are effectively improved.

[0058] It should be noted that the training objective of this speech synthesis decoding unit (that is, the training objective in the speech synthesis training stage, as Figure 10 shown) is to further minimize the text output error of the text decoding unit (that is, make the text output error of the basic text decoding unit less than or equal to the first error threshold) by minimizing the voice prediction error of the basic voice decoding unit and the language prediction error of the basic language decoding unit. Specifically, by optimizing these prediction errors, the quality and accuracy of the speech generated by the decoding unit can be improved.

[0059] Among them, the text output error, the voice prediction error, and the language prediction error are all measured by the same first loss function to ensure that each part can be optimized cooperatively during the training process.

[0060] In a possible implementation, the first loss function is: Looss1 = L text + λ1L voice + λ2L lang , L text is the text output error, L voice is the voice prediction error, L lang is the language prediction error, and λ1 and λ2 are hyperparameters.

[0061] It should also be noted that this application does not specifically limit the size of the first error threshold, and users can adjust the size of the first error threshold according to actual needs.

[0062] Through the above steps, the basic decoding unit has completed learning the mapping relationship from text to speech in the multi-modal data, and has been effectively constrained and optimized in terms of timbre and language, laying a solid foundation for subsequent more complex language translation and cross-language conversion training.

[0063] In a possible implementation, the present application provides a flowchart of a language translation training method. Refer to Figure 5 , Figure 5 which is a flowchart of a language translation training method provided by an embodiment of the present application. Specifically, the language translation training process of the speech synthesis decoding unit can be implemented through steps S501 - S503: S501: Construct a second training dataset.

[0064] In the language translation training process, constructing a second training dataset is crucial. This training dataset covers text vector samples with speech vector labels. Each text vector sample in this dataset is equipped with a corresponding speech vector label, and each text vector sample and its corresponding speech vector label belong to the same language. This is done to ensure the consistency and accuracy of the training data, so that the speech synthesis decoding unit can better learn the correspondence between text and speech and the characteristics of a specific language during the training process. In this way, the translation accuracy and naturalness of the speech synthesis decoding unit can be effectively improved.

[0065] S502: Input the second training dataset into the timbre decoding unit and the language decoding unit of the speech synthesis decoding unit respectively to obtain a second timbre feature representation and a second language feature representation.

[0066] When the timbre decoding unit of the speech synthesis decoding unit processes the second training data, it analyzes the sound signals in the training data, extracts the feature information related to timbre, and generates a second timbre feature representation for describing the timbre attributes; at the same time, the language decoding unit of the speech synthesis decoding unit identifies and extracts the language category and its related features for the same second training dataset, forming a corresponding second language feature representation. This process lays a foundation for subsequent language translation training and realizes independent and effective encoding of timbre and language information.

[0067] S503: Simultaneously input the second training dataset, the second timbre feature representation, and the second language feature representation as inputs into the speech decoding unit and the text decoding unit of the speech synthesis decoding unit for joint decoding and error calculation. When it is determined that the speech output error of the speech decoding unit in the speech synthesis decoding unit is less than or equal to the second error threshold, stop training to obtain a language translation decoding unit.

[0068] After obtaining the second tone feature representation and the second language feature representation, use the second training dataset, the second tone feature representation, and the second language feature representation as the inputs of the speech synthesis decoding unit, and perform joint decoding and error calculation on the tone decoding unit and the language decoding unit of the speech synthesis decoding unit. During this process, the tone decoding unit and the language decoding unit of the speech synthesis decoding unit cooperate to generate corresponding speech and text outputs according to the input information, and compare them with the expected results to calculate the error. When it is detected that the speech output error of the speech decoding unit in the speech synthesis decoding unit is less than or equal to the preset second error threshold, it indicates that the model has reached the desired accuracy level. At this time, stop the training process, and finally obtain an optimized and improved speech synthesis decoding unit. Through this joint training method, the translation accuracy and speech naturalness of the speech synthesis decoding unit in multi-modal language conversion are effectively improved, the adaptability of the model to different languages and tone features is enhanced, and thus a more fluent and high-quality language conversion effect is achieved.

[0069] It should be noted that the training objective of this language translation decoding unit (that is, the training objective in the language translation training stage, as Figure 10 shown) is to further minimize the speech output error of the speech decoding unit in the speech synthesis decoding unit (that is, to make the speech output error of the speech decoding unit in the speech synthesis decoding unit less than or equal to the second error threshold) by minimizing the tone prediction error of the tone decoding unit and the language prediction error of the language decoding unit in the speech synthesis decoding unit. Specifically, by optimizing these prediction errors, the quality and accuracy of the speech decoding unit in generating language translation speech can be improved.

[0070] Among them, the text output error, the tone prediction error, and the speech output error are all measured by the same second loss function to ensure that all parts can be jointly optimized during the training process.

[0071] In a possible implementation manner, the second loss function is: , y output is the speech output of the speech synthesis decoding unit for the text sample during the model training process, y target is the speech label of the text sample in the second training dataset, L voice is the tone prediction error, is the language prediction error, and λ3 and λ4 are hyperparameters.

[0072] It should also be noted that this application does not specifically limit the size of the second error threshold, and users can adjust the size of the second error threshold according to actual needs.

[0073] Through the above steps, the speech synthesis decoding unit has completed the learning of the mapping relationship from text to speech, and has been effectively constrained and optimized in terms of timbre and language, laying a solid foundation for subsequent more complex cross-lingual conversion training.

[0074] In a possible implementation manner, the present application provides a flowchart of a cross-lingual conversion training method. Refer to Figure 6 , Figure 6 which is a flowchart of a cross-lingual conversion training method provided by an embodiment of the present application. Specifically, the cross-lingual conversion training process of the language translation decoding unit can be implemented through steps S601 - S603: S601: Construct a third training data set.

[0075] In the cross-lingual conversion training process, constructing a third training data set plays a key role. This training data set contains two types of samples: one is a text vector sample with a speech vector label, and the other is a speech vector sample with a text vector label. Among them, each text vector sample in the third training data set and its corresponding speech vector label belong to different languages, and similarly, each speech vector sample and its corresponding text vector label also belong to different languages. Specifically, these speech vector labels and text vector labels are both in the target language, while the corresponding speech vector samples and text vector samples belong to non-target languages. By designing such multi-lingual and cross-lingual training data, it can effectively prompt the model to learn the conversion relationship from non-target languages to the target language, thereby improving the accuracy and robustness of cross-lingual conversion.

[0076] S602: Input the third training data set into the timbre decoding unit and the language decoding unit of the language translation decoding unit respectively to obtain a third timbre feature representation and a third language feature representation.

[0077] When the timbre decoding unit of the language translation decoding unit processes the third training data, it analyzes the sound signals in the training data, extracts feature information related to timbre, and generates a third timbre feature representation for describing timbre attributes; at the same time, the language decoding unit of the speech synthesis decoding unit identifies and extracts the language category and its related features for the same third training data set to form a corresponding third language feature representation. This process lays a foundation for subsequent cross-lingual conversion training and realizes independent and effective encoding of timbre and language information.

[0078] S603: The third training data set, the third timbre feature representation, and the third language feature representation are simultaneously used as inputs and respectively input into the speech decoding unit and the text decoding unit of the language translation decoding unit for joint decoding and error calculation. When it is determined that the multi-language translation text error of the text decoding unit in the language translation decoding unit is less than or equal to the fourth error threshold and the speech output error of the speech decoding unit is less than or equal to the fifth error threshold, the training is stopped to obtain the multi-modal decoding unit.

[0079] After obtaining the third timbre feature representation and the third language feature representation, the third training data set, the third timbre feature representation, and the third language feature representation are used as inputs to the language translation decoding unit, and joint decoding and error calculation are performed on the speech decoding unit and the text decoding unit of the speech synthesis decoding unit. During this process, the speech decoding unit and the text decoding unit of the language translation decoding unit jointly generate corresponding speech and text outputs according to the input information, and compare them with the expected cross-language target results to calculate the multi-modal error. When it is detected that the multi-language translation text error generated by the text decoding unit is less than or equal to the preset fourth error threshold, and the speech output error of the speech decoding unit is less than or equal to the preset fifth error threshold, it indicates that the model has achieved the expected training goal. At this time, the training process is stopped, and finally a multi-modal decoding unit with excellent performance and cross-language conversion ability is obtained. Through this training method, the conversion accuracy and robustness of the model in a multi-language and multi-modal environment are effectively enhanced, and high-quality cross-language speech and text conversion effects are achieved.

[0080] It should be noted that the training goal of this multi-modal decoding unit (that is, the training goal of the cross-language conversion training stage, as Figure 10 shown) is to further reduce the multi-language translation text error in the text decoding unit of the language translation decoding unit and the speech output error of the speech decoding unit by minimizing the timbre prediction error of the timbre decoding unit and the language prediction error of the language decoding unit (that is, making the multi-language translation text error of the text decoding unit in the language translation decoding unit less than or equal to the fourth error threshold, and making the speech output error of the speech decoding unit less than or equal to the fifth error threshold). Specifically, by uniformly optimizing these error metrics, the performance of the multi-modal decoding unit in cross-language text and speech conversion can be significantly improved.

[0081] Among them, the timbre prediction error, the language prediction error, the multi-language translation text error, and the speech output error are all comprehensively measured by the third loss function, ensuring that each decoding unit part is jointly optimized during the training process to achieve better cross-language conversion performance.

[0082] In a possible implementation manner, the third loss function is: Looss3 = L text+L speech +L lang +λ5L align ,L text is the text output error, L speech is the language output error, L lang is the language prediction error, L align The audio-text alignment error, and λ5 is a hyperparameter.

[0083] It should also be noted that this application does not specifically limit the magnitudes of the third and fourth error thresholds, and users can adjust the magnitudes of the third and fourth error thresholds according to actual needs.

[0084] Through the above two steps, the multi-modal decoding unit has successfully achieved cross-language mapping and conversion between non-target languages and target languages. Combining the constraints of timbre and language information, it has achieved high-quality multi-language speech-text mutual conversion, providing a solid technical foundation for complex multi-modal multi-language applications.

[0085] In a possible implementation manner, this application provides a flowchart of a text-audio alignment training method. Refer to Figure 7 , Figure 7 is a flowchart of a text-audio alignment training method provided by an embodiment of this application. Specifically, the text-audio alignment training process of the basic encoding unit can be implemented through steps S701 - S702: S701: Construct a fourth training dataset.

[0086] In the text-audio alignment training process, constructing a fourth training dataset is a crucial step. This dataset includes multiple pairs of first sample pairs, and each pair of first sample pairs contains a text sample and an audio sample. These sample pairs can be further divided into two categories: The first category: The sample pair consists of a text sample and an audio sample with the same language but different semantics (i.e., a negative sample pair).

[0087] The second category: The sample pair consists of a text sample and an audio sample with the same language and the same semantics (i.e., a positive sample pair).

[0088] This structured dataset design aims to help the basic encoding unit better distinguish the relationships between text and audio under different semantics and the same semantics, so as to achieve more accurate text-audio alignment.

[0089] S702: Use the fourth training dataset as the input of the basic encoding unit, perform encoding and error calculation. When it is determined that the text-audio alignment error of the positive sample pairs in the basic encoding unit is less than or equal to the fifth error threshold, and the text-audio alignment error of the negative sample pairs is greater than or equal to the sixth error threshold, stop the training to obtain the audio alignment encoding unit.

[0090] Use the fourth training dataset as the input of the basic encoding unit, and perform encoding and calculation of the text-audio alignment error on the basic encoding unit. During this process, the basic encoding unit generates corresponding encoding representations according to the input training samples respectively, and calculates the text-audio alignment errors of the positive sample pairs and the negative sample pairs respectively. When it is detected that the text-audio alignment error of the positive sample pairs is less than or equal to the preset fifth error threshold, and the text-audio alignment error of the negative sample pairs is greater than or equal to the preset sixth error threshold, it indicates that the model has achieved the expected training goal. At this time, stop the training process, and finally obtain an audio alignment encoding unit with excellent performance and accurate text-audio alignment ability. Through this training method, the alignment accuracy and robustness of the basic encoding unit in distinguishing positive and negative samples are effectively improved, and a high-quality text-audio alignment effect is achieved.

[0091] It should be noted that the training goal of this basic encoding unit (that is, the training goal in the text-audio alignment training stage, as Figure 11 shown) is as follows: Error maximization: For the first type of sample pairs (same language but different semantics), the goal is to maximize the text-audio alignment error between the text sample and the audio sample (that is, to make the text-audio alignment error of the negative sample pairs in the basic encoding unit greater than or equal to the sixth error threshold).

[0092] Error minimization: For the second type of sample pairs (same language and same semantics), the goal is to minimize the text-audio alignment error between the text sample and the audio sample (that is, to make the text-audio alignment error of the positive sample pairs in the basic encoding unit less than or equal to the fifth error threshold).

[0093] It should also be noted that this application does not specifically limit the magnitudes of the fifth and sixth error thresholds, and users can adjust the magnitudes of the fifth and sixth error thresholds according to actual needs.

[0094] In this way, the basic encoding unit can learn how to effectively distinguish audio and text samples with different semantics and achieve precise alignment when the semantics are the same.

[0095] Among them, the measurement and optimization of the text-audio alignment error are both completed through the same fourth loss function to ensure that all parts in the training process can be optimized collaboratively.

[0096] In a possible implementation, the fourth loss function is as follows: , where N1 is the number of the first sample pairs, etext1i is the encoding vector of the inner text sample of the i-th pair of the first sample pairs, and eaudio1i is the encoding vector of the inner audio sample of the i-th pair of the first sample pairs.

[0097] Through the above two steps, the basic encoding unit can successfully learn the correspondence and its subtle differences between text and audio from the multimodal data, thus laying a solid foundation for subsequent more complex tasks (such as timbre alignment training and language alignment training).

[0098] In a possible implementation, the present application provides a flowchart of a timbre alignment training method. Refer to Figure 8 , Figure 8 which is a flowchart of a timbre alignment training method provided by an embodiment of the present application. Specifically, the timbre alignment training process of the audio alignment encoding unit can be implemented through steps S801 - S802: S801: Construct a fifth training dataset.

[0099] In the timbre alignment training process, constructing a fifth training dataset is a crucial step. This dataset includes multiple pairs of second sample pairs, and each pair of second sample pairs contains two audio samples. These sample pairs can be further divided into two categories: The first category: The sample pair consists of two audio samples with the same semantics but different timbres (i.e., negative sample pairs).

[0100] The second category: The sample pair consists of two audio samples with the same semantics and the same timbre (i.e., positive sample pairs).

[0101] This structured dataset design aims to help the audio alignment encoding unit better distinguish the relationships between audio with different timbres and audio with the same timbre, so as to achieve more accurate timbre alignment.

[0102] S802: Use the fifth training dataset as the input of the audio alignment encoding unit, perform encoding and error calculation. When it is determined that the timbre alignment error of the positive sample pairs in the audio alignment encoding unit is less than or equal to the seventh error threshold, and the timbre alignment error of the negative sample pairs is greater than or equal to the eighth error threshold, stop the training to obtain the audio alignment encoding unit.

[0103] Use the fifth training dataset as the input of the audio alignment encoding unit, and perform encoding on the audio alignment encoding unit and calculation of the timbre alignment error. During this process, the audio alignment encoding unit generates corresponding encoding representations according to the input training samples respectively, and calculates the timbre alignment errors of the positive sample pairs and the negative sample pairs respectively. When it is detected that the timbre alignment error of the positive sample pair is less than or equal to the preset seventh error threshold, and the timbre alignment error of the negative sample pair is greater than or equal to the preset eighth error threshold, it indicates that the model has achieved the expected training goal. At this time, stop the training process, and finally obtain an audio alignment encoding unit with excellent performance and accurate timbre alignment ability. Through this training method, the alignment accuracy and robustness of the audio alignment encoding unit in distinguishing positive and negative samples are effectively improved, and a high-quality timbre alignment effect is achieved.

[0104] It should be noted that the training goal of this audio alignment encoding unit (that is, the timbre alignment training stage, as Figure 11 shown) is as follows: Error maximization: For the first type of sample pairs (with the same semantics but different timbres), the goal is to maximize the timbre alignment error between the two audio samples (that is, to make the timbre alignment error of the negative sample pairs in the audio alignment encoding unit greater than or equal to the eighth error threshold).

[0105] Error minimization: For the second type of sample pairs (with the same semantics and the same timbre), the goal is to minimize the timbre alignment error between the two audio samples (that is, to make the timbre alignment error of the positive sample pairs in the audio alignment encoding unit less than or equal to the seventh error threshold).

[0106] In this way, the audio alignment encoding unit can learn how to effectively distinguish audio samples with different timbres and achieve precise alignment when the timbres are the same.

[0107] Among them, the measurement and optimization of the audio timbre alignment error are both completed through the same fifth loss function to ensure that all parts in the training process can be optimized collaboratively.

[0108] In a possible implementation manner, the fifth loss function is: , N2 is the number of the second sample pairs, and e audio3i and e audio4i are the encoding vectors of the two audio samples in the i-th pair of the second sample pairs.

[0109] It should also be noted that this application does not specifically limit the magnitudes of the seventh and eighth error thresholds, and users can adjust the magnitudes of the seventh and eighth error thresholds according to actual needs.

[0110] Through the above two steps, the audio alignment encoding unit can successfully learn the corresponding relationships and subtle differences between audios from multi-modal data, thus laying a solid foundation for subsequent more complex tasks (such as language alignment training).

[0111] In a possible implementation manner, the present application provides a flowchart of a language alignment training method. Refer to Figure 9 , Figure 9 which is a flowchart of a language alignment training method provided by an embodiment of the present application. Specifically, the language alignment training process of the audio timbre alignment encoding unit can be implemented through steps S901 - S902: S901: Construct a sixth training dataset.

[0112] During the language alignment training process, constructing a sixth training dataset is crucial. This training dataset covers multiple groups of first sample groups. Each group of first sample groups includes text samples and audio samples of one language and text samples and audio samples of another language under the same semantics. This is done to ensure that the training data can cover the corresponding relationships between different languages, so that the audio timbre alignment encoding unit can better learn the features and differences between different languages during the training process. In this way, the accuracy and naturalness of the cross-language conversion system can be effectively improved.

[0113] It should be noted that S902: Use the sixth training dataset as the input of the audio timbre alignment encoding unit, perform encoding and error calculation. When it is determined that the language alignment error of samples of the same language in the audio timbre alignment encoding unit is less than or equal to a ninth error threshold, and the language alignment error of samples of different languages is greater than or equal to a tenth error threshold, stop the training to obtain a multi-modal encoding unit.

[0114] Use the sixth training dataset as the input of the audio timbre alignment encoding unit, perform encoding on the audio timbre alignment encoding unit and calculate the language alignment error. During this process, the audio timbre alignment encoding unit generates corresponding encoding representations according to the input training samples respectively, and calculates the language alignment errors of samples of the same language and samples of different languages respectively. When it is detected that the language alignment error of samples of the same language is less than or equal to a preset ninth error threshold, and the language alignment error of samples of different languages is greater than or equal to a preset tenth error threshold, it indicates that the model has achieved the expected training goal. At this time, stop the training process, and finally obtain a multi-modal encoding unit with excellent performance and multi-modal alignment ability. Through this training method, the performance of the audio timbre alignment encoding unit in language discrimination and alignment accuracy is effectively improved, and a high-quality multi-modal alignment effect is achieved.

[0115] It should be noted that the training objectives of this audio timbre alignment encoding unit are as follows: Error maximization: For text samples and audio samples in different languages within the first sample group, the goal is to maximize the language alignment error between them (that is, to make the language alignment error of samples in different languages in the audio tone alignment coding unit greater than or equal to the tenth error threshold).

[0116] Error minimization: For text samples and audio samples in the same language within the first sample group, the goal is to minimize the language alignment error between them (that is, to make the language alignment error of samples in the same language in the audio tone alignment coding unit less than or equal to the ninth error threshold).

[0117] In this way, the audio tone alignment coding unit can learn how to effectively distinguish text and audio samples in different languages and achieve precise alignment in the case of the same language.

[0118] Among them, the language alignment error is measured by the same sixth loss function to ensure that all parts in the training process can be optimized collaboratively.

[0119] In a possible implementation, the sixth loss function is: , N3 is the number of the first sample group, etext2(i, L1) and eaudio5(i, L1) are the coding vectors of the text sample and the audio sample in the same language within the i-th group of the first sample group; e text3 (i,L1) and e audio6 (i,L1) are the coding vectors of the text sample and the audio sample in different languages within the i-th group of the first sample group; λ is a hyperparameter.

[0120] It should also be noted that this application does not specifically limit the magnitudes of the ninth and tenth error thresholds, and users can adjust the magnitudes of the ninth and tenth error thresholds according to actual needs.

[0121] Through the above two steps, the multimodal coding unit has successfully achieved cross-language text and audio alignment, improved the ability to fuse and understand multi-language and multimodal data, and laid a solid foundation for subsequent multimodal decoding and cross-language conversion tasks.

[0122] In a possible implementation, in order to ensure that the input data can be effectively used for training and subsequent processing of the model, it is necessary to preprocess the input data to obtain preprocessed data, so as to use the multimodal coding unit to perform coding processing on the preprocessed data to obtain the vector representation of the input data. Specifically, the method further includes: Text tokenization: Tokenize the text data in the input data, splitting the continuous text string into a series of meaningful lexical units (words or phrases) to form a token sequence. For example, for the input text "Helloworld", the token sequence ["Hello", "world"] can be obtained after tokenization.

[0123] Speech feature extraction: Extract speech features from the audio data in the input data, extracting the key feature information in the audio signal and converting it into a numerical vector form. Speech feature extraction can include, but is not limited to, features extraction such as MFCC feature extraction, spectral feature extraction, energy feature extraction, etc., and finally form a speech feature vector.

[0124] Through the above preprocessing steps, the original text and audio data are converted into forms suitable for model processing, namely the token sequence and the speech feature vector, providing a data basis for the subsequent training and inference processes.

[0125] In a possible implementation, the encoding process of the input data using the multimodal encoding unit to obtain the vector representation of the input data includes: Use the multimodal encoding unit to encode the preprocessed data to obtain the vector representation of the input data.

[0126] This application provides a flowchart of a language alignment training method, which can be specifically implemented through steps S1201 - S1202: S1201: Construct an autoregressive sample set.

[0127] In the process of constructing the autoregressive sample set, it is necessary to ensure that the sample set includes autoregressive samples in both text and speech forms. Specifically, the autoregressive sample set consists of N text autoregressive samples and M speech autoregressive samples. To maintain the balance and rationality of the sample set, the ratio of the number of text samples to the number of speech samples is set as N / M = X, where X is a positive integer. This ratio design helps to ensure that the basic encoding unit can evenly process text and speech data during the learning process, thereby improving the training effect and the performance of the basic encoding unit.

[0128] S1202: Use the autoregressive sample set to perform autoregressive training on the basic encoding unit to obtain a pre-trained encoding unit.

[0129] During the training process, the constructed autoregressive sample set is used to perform autoregressive training on the basic encoding unit. Through this process, the basic encoding unit can learn the sequence dependencies and pattern features in the sample set. After sufficient training, the basic encoding unit can effectively process and understand the input text and speech sequences, and finally form a pre-trained encoding unit, which has preliminary sequence modeling ability and feature extraction ability.

[0130] Through the above steps, not only the balance and diversity of text and speech samples are ensured, but also the understanding ability of the basic encoding unit for sequence data is improved through autoregressive training. The finally formed pre-trained encoding unit has good sequence modeling ability and generalization performance, providing a basis for subsequent complex training.

[0131] In a possible implementation, the multimodal encoding unit is obtained by gradually performing text-audio alignment training, timbre alignment training, and language alignment training on the basic encoding unit, including: The multimodal decoding unit is obtained by gradually performing text-audio alignment training, timbre alignment training, and language alignment training on the pre-trained encoding unit.

[0132] Using the pre-trained encoding unit as the initial basis for training can provide good parameter initialization and feature representation ability for subsequent text-audio alignment training, timbre alignment training, and language alignment training. The pre-trained encoding unit has mastered the basic temporal relationship and expression structure between text and audio sequences through autoregressive training, which lays a solid foundation for the multimodal alignment task. Based on this, the gradually carried out alignment training at each stage can more efficiently optimize the corresponding relationships of the encoding unit at different levels, accelerate the convergence speed, and improve the accuracy and robustness of the final multimodal encoding unit, so as to achieve more refined and comprehensive multimodal fusion.

[0133] Based on the multimodal language conversion method provided in the above method embodiment, the embodiment of the present application also provides a multimodal language conversion device, which will be described below with reference to the accompanying drawings.

[0134] See Figure 12 As shown Figure 12 in the figure, it is a schematic diagram of a multimodal language conversion device provided by the embodiment of the present application. As Figure 12 shown in the figure, the multimodal language conversion device includes: An encoding processing unit 1201, in response to input data, is configured to use the multimodal encoding unit to perform encoding processing on the input data to obtain a vector representation of the input data; the type of the input data includes at least one of text data and audio data; A decoding processing unit 1202, configured to use multiple decoding units in the multimodal decoding unit to perform respective decoding processes on the vector representation, obtain target data after language conversion of the input data, and output the target data; the multiple decoding units are respectively obtained by associative training with different basic decoding units, and each decoding unit is related to the multimodal cross-language understanding function.

[0135] In a possible implementation manner, the multiple decoding units include a speech decoding unit and a text decoding unit.

[0136] In a possible implementation manner, the decoding processing unit 1202 is specifically configured to: Use the speech decoding unit to perform decoding processing on the vector representation, obtain speech conversion data in the target data, and output the speech conversion data; Use the text decoding unit to perform decoding processing on the vector representation, obtain text conversion data in the target data, and output the text conversion data.

[0137] In a possible implementation manner, the multiple decoding units further include: a language decoding unit.

[0138] In a possible implementation manner, the decoding processing unit 1202 is specifically configured to: Obtain target language information in the language decoding unit; Use the speech decoding unit to combine the target language information to perform decoding processing on the vector representation, obtain speech conversion data in the target data, and output the speech conversion data.

[0139] In a possible implementation manner, the decoding processing unit 1202 is specifically configured to: Obtain target language information in the language decoding unit; Use the text decoding unit to combine the target language information to perform decoding processing on the vector representation, obtain text conversion data in the target data, and output the text conversion data.

[0140] In a possible implementation manner, the multiple multimodal decoding units further include: a timbre decoding unit.

[0141] In a possible implementation manner, the decoding processing unit 1202 is specifically configured to: Obtain conversion timbre information in the timbre decoding unit; Use the speech decoding unit to combine the conversion timbre information to perform decoding processing on the vector representation, obtain speech conversion data in the target data, and output the speech conversion data.

[0142] In a possible implementation manner, the decoding processing unit 1202 is specifically configured to: Obtain the converted timbre information in the timbre decoding unit; Use the text decoding unit to combine the converted timbre information to perform decoding processing on the vector representation, obtain the text conversion data in the target data, and output it.

[0143] In a possible implementation manner, the basic decoding unit includes a basic timbre decoding unit, a basic speech decoding unit, a basic text decoding unit, and a basic language decoding unit; the associated training of the multiple decoding units includes speech synthesis training, language translation training, and cross-language conversion training.

[0144] In a possible implementation manner, the training process of the speech synthesis training includes: In the stage of the speech synthesis training, input the first training data set into the basic timbre decoding unit and the basic language decoding unit respectively to obtain the first timbre feature representation and the first language feature representation, and use the first training data set, the first timbre feature representation, and the first language feature representation as inputs simultaneously and input them into the basic speech decoding unit and the basic text decoding unit respectively for joint decoding and error calculation. When it is determined that the text output error of the basic text decoding unit is less than or equal to the first error threshold, stop the training to obtain the speech synthesis decoding unit; Wherein, the first training data set includes speech vector samples with text vector labels, and the languages of the respective speech vector samples in the first training data set are the same as the languages of their corresponding text vector labels.

[0145] In a possible implementation manner, the training process of the language translation training includes: In the stage of the language translation training, input the second training data set into the timbre decoding unit and the language decoding unit of the speech synthesis decoding unit respectively to obtain the second timbre feature representation and the second language feature representation, and use the second training data set, the second timbre feature representation, and the second language feature representation as inputs simultaneously and input them into the speech decoding unit and the text decoding unit of the speech synthesis decoding unit respectively for joint decoding and error calculation. When it is determined that the speech output error of the speech decoding unit in the speech synthesis decoding unit is less than or equal to the second error threshold, stop the training to obtain the language translation decoding unit; Wherein, the second training data set includes text vector samples with speech vector labels, and the languages of the respective text vector samples in the fifth training set are the same as the languages of their corresponding speech vector labels.

[0146] In a possible implementation manner, the training process of the cross-language conversion training includes: In the stage of cross-lingual conversion training, the third training dataset is respectively input into the timbre decoding unit and the language decoding unit of the language translation decoding unit to obtain the third timbre feature representation and the third language feature representation. The third training dataset, the third timbre feature representation, and the third language feature representation are simultaneously used as inputs and respectively input into the speech decoding unit and the text decoding unit of the language translation decoding unit for joint decoding and error calculation. When it is determined that the multi-lingual translation text error of the text decoding unit in the language translation decoding unit is less than or equal to the fourth error threshold and the speech output error of the speech decoding unit is less than or equal to the fifth error threshold, the training is stopped to obtain the multi-modal decoding unit; Among them, the third training dataset includes text vector samples with speech vector labels and speech vector samples with text vector labels; the languages of each text vector sample and its corresponding speech vector label in the third training dataset are all different, the languages of each speech vector sample and its corresponding text vector label are all different, the languages of the speech vector label and the text vector label are both the target languages, and the languages of the speech vector sample and the text vector sample are both non-target languages.

[0147] In a possible implementation manner, the multi-modal encoding unit is obtained by gradually performing text-audio alignment training, timbre alignment training, and language alignment training through a basic encoding unit; In the stage of text-audio alignment training, with the training objective of making the text-audio alignment error of positive sample pairs in the basic encoding unit less than or equal to the fifth error threshold and the text-audio alignment error of negative sample pairs greater than or equal to the sixth error threshold, the fourth training dataset is used to perform the text-audio alignment training on the basic encoding unit to obtain an audio alignment encoding unit as the training result of the text-audio alignment training; In the stage of timbre alignment training, with the training objective of making the timbre alignment error of positive sample pairs in the audio alignment encoding unit less than or equal to the seventh error threshold and the text-audio alignment error of negative sample pairs greater than or equal to the eighth error threshold, the fifth training dataset is used to perform the timbre alignment training on the audio alignment encoding unit to obtain an audio-timbre alignment encoding unit as the training result of the timbre alignment training; In the stage of language alignment training, with the training objective of making the language alignment error of samples of the same language in the audio-timbre alignment encoding unit less than or equal to the ninth error threshold and the language alignment error of samples of different languages greater than or equal to the tenth error threshold, the sixth training dataset is used to perform the language alignment training on the audio-timbre alignment encoding unit to obtain the multi-modal encoding unit as the training result of the language alignment training; Among them, the fourth training data set includes multiple pairs of first sample pairs, and each first sample pair includes a text sample and an audio sample; the first sample pairs are divided into two categories. The first category of the first sample pairs includes a text sample and an audio sample with the same language but different semantics, and the second category of the first sample pairs includes a text sample and an audio sample with the same language and the same semantics; The fifth training data set includes multiple pairs of second sample pairs, and each second sample pair includes two audio samples; the second sample pairs are divided into two categories. The first category of the second sample pairs includes two audio samples with the same semantics but different timbres, and the second category of the second sample pairs includes two audio samples with the same semantics and the same timbre; The sixth training data set includes multiple groups of first sample groups, and each first sample group includes, under the same semantics, a text sample and an audio sample in one language and a text sample and an audio sample in another language.

[0148] See Figure 13 , based on the same concept, the present application also provides an electronic device 1300. The electronic device 1300 may include a memory 1301, a processor 1302, and a machine-executable program stored on the memory 1301 and running on the processor 1302. When the processor 1302 executes the machine-executable program, the electronic device 1300 is enabled to execute the multi-modal language conversion method as described above.

[0149] In a possible implementation manner, a multi-modal conversion model is loaded in the processor 1302 to execute the multi-modal language conversion method as described above through the multi-modal conversion model.

[0150] In the embodiment of the present application, after receiving text data or audio data as input, the input data is uniformly encoded by a multi-modal encoding unit to generate a corresponding vector representation. Subsequently, a multi-modal decoding unit composed of multiple associated trained basic decoding units respectively executes their respective decoding tasks on the vector representation, realizing multi-angle understanding and processing of the input data, and finally outputting target data that completes language conversion. The multi-modal language conversion method in the present application, by adopting a multi-modal encoding unit and a multi-modal decoding unit, not only solves the problems of information loss and time delay caused by the chain processing method in the traditional speech translation system, but also improves the balance between real-time performance and accuracy.

[0151] It should be noted that the various embodiments in this specification are described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the relevant content. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components referred to as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0152] As described above, this is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A multimodal language conversion method, characterized in that, Applied to a multimodal conversion model, the multimodal conversion model includes a multimodal encoding unit and a multimodal decoding unit, and the method includes: In response to receiving input data, use the multimodal encoding unit to encode the input data to obtain a vector representation of the input data; the type of the input data includes at least one of text data and audio data; Use multiple decoding units in the multimodal decoding unit to perform respective decoding processes on the vector representation respectively, to obtain target data after language conversion of the input data and output it; the multiple decoding units are respectively obtained by associative training with different basic decoding units, and each of the decoding units is related to the multimodal cross-language understanding function.

2. The method according to claim 1, characterized in that, The multiple decoding units include a speech decoding unit and a text decoding unit; The step of using multiple decoding units in the multimodal decoding unit to perform respective decoding processes on the vector representation respectively, to obtain target data after language conversion of the input data and output it, includes: Use the speech decoding unit to decode the vector representation to obtain and output the speech conversion data in the target data; Use the text decoding unit to decode the vector representation to obtain and output the text conversion data in the target data.

3. The method according to claim 2, characterized in that, The multiple decoding units further include: a language decoding unit; The step of using the speech decoding unit to decode the vector representation to obtain and output the speech conversion data in the target data includes: Obtain the target language information in the language decoding unit; Use the speech decoding unit to combine the target language information to decode the vector representation to obtain and output the speech conversion data in the target data; The step of using the text decoding unit to decode the vector representation to obtain and output the text conversion data in the target data includes: Obtain the target language information in the language decoding unit; Use the text decoding unit to combine the target language information to decode the vector representation to obtain and output the text conversion data in the target data.

4. The method according to claim 2, wherein The multiple multimodal decoding units further include: a timbre decoding unit; The step of using the speech decoding unit to decode the vector representation to obtain and output the speech conversion data in the target data includes: Obtain the converted timbre information in the timbre decoding unit; Use the speech decoding unit to combine the converted timbre information to decode the vector representation to obtain and output the speech conversion data in the target data; The step of using the text decoding unit to decode the vector representation to obtain and output the text conversion data in the target data includes: Obtain the converted timbre information in the timbre decoding unit; Use the text decoding unit to combine the converted timbre information to decode the vector representation to obtain and output the text conversion data in the target data.

5. The method according to claim 1, wherein The basic decoding unit includes a basic timbre decoding unit, a basic speech decoding unit, a basic text decoding unit, and a basic language decoding unit; the associated training of the multiple decoding units includes speech synthesis training, language translation training, and cross-lingual conversion training; The training process of the speech synthesis training includes: In the stage of the speech synthesis training, the first training dataset is respectively input into the basic timbre decoding unit and the basic language decoding unit to obtain a first timbre feature representation and a first language feature representation, and the first training dataset, the first timbre feature representation, and the first language feature representation are simultaneously used as inputs and respectively input into the basic speech decoding unit and the basic text decoding unit for joint decoding and error calculation. When it is determined that the text output error of the basic text decoding unit is less than or equal to the first error threshold, the training is stopped to obtain a speech synthesis decoding unit; Among them, the first training dataset includes speech vector samples with text vector labels, and the languages of the respective speech vector samples in the first training dataset and their corresponding text vector labels are the same.

6. The method according to claim 5, characterized in that, The training process of the language translation training includes: In the stage of the language translation training, the second training dataset is respectively input into the timbre decoding unit and the language decoding unit of the speech synthesis decoding unit to obtain a second timbre feature representation and a second language feature representation, and the second training dataset, the second timbre feature representation, and the second language feature representation are simultaneously used as inputs and respectively input into the speech decoding unit and the text decoding unit of the speech synthesis decoding unit for joint decoding and error calculation. When it is determined that the speech output error of the speech decoding unit in the speech synthesis decoding unit is less than or equal to the second error threshold, the training is stopped to obtain a language translation decoding unit; Among them, the second training dataset includes text vector samples with speech vector labels, and the languages of the respective text vector samples in the fifth training set and their corresponding speech vector labels are the same.

7. The method according to claim 6, wherein The training process of the cross-lingual conversion training includes: In the stage of the cross-lingual conversion training, the third training dataset is respectively input into the timbre decoding unit and the language decoding unit of the language translation decoding unit to obtain a third timbre feature representation and a third language feature representation, and the third training dataset, the third timbre feature representation, and the third language feature representation are simultaneously used as inputs and respectively input into the speech decoding unit and the text decoding unit of the language translation decoding unit for joint decoding and error calculation. When it is determined that the multi-lingual translation text error of the text decoding unit in the language translation decoding unit is less than or equal to the fourth error threshold and the speech output error of the speech decoding unit is less than or equal to the fifth error threshold, the training is stopped to obtain a multi-modal decoding unit; Among them, the third training dataset includes text vector samples with speech vector labels and speech vector samples with text vector labels; for each text vector sample in the third training dataset, the language of the corresponding speech vector label is different, and for each speech vector sample, the language of the corresponding text vector label is different. The languages of the speech vector labels and the text vector labels are both the target language, and the languages of the speech vector samples and the text vector samples are both non-target languages.

8. The method according to claim 1, wherein The multimodal encoding unit is obtained by gradually performing text-audio alignment training, timbre alignment training, and language alignment training on the basic encoding unit; In the stage of text-audio alignment training, taking the text-audio alignment error of the positive sample pairs in the basic encoding unit being less than or equal to the fifth error threshold and the text-audio alignment error of the negative sample pairs being greater than or equal to the sixth error threshold as the training objectives, using the fourth training dataset to perform the text-audio alignment training on the basic encoding unit, and obtaining the audio alignment encoding unit as the training result of the text-audio alignment training; In the stage of timbre alignment training, taking the timbre alignment error of the positive sample pairs in the audio alignment encoding unit being less than or equal to the seventh error threshold as the training objective and the text-audio alignment error of the negative sample pairs being greater than or equal to the eighth error threshold as the training objective, using the fifth training dataset to perform the timbre alignment training on the audio alignment encoding unit, and obtaining the audio timbre alignment encoding unit as the training result of the timbre alignment training; In the stage of language alignment training, taking the language alignment error of the same-language samples in the audio timbre alignment encoding unit being less than or equal to the ninth error threshold and the language alignment error of different-language samples being greater than or equal to the tenth error threshold as the training objectives, using the sixth training dataset to perform the language alignment training on the audio timbre alignment encoding unit, and obtaining the multimodal encoding unit as the training result of the language alignment training; Among them, the fourth training dataset includes multiple pairs of first sample pairs, and the first sample pair includes a text sample and an audio sample; the first sample pair is divided into two categories. The first category of the first sample pair includes text samples and audio samples with the same language but different semantics, and the second category of the first sample pair includes text samples and audio samples with the same language and the same semantics; The fifth training dataset includes multiple pairs of second sample pairs, and the second sample pair includes two audio samples; the second sample pair is divided into two categories. The first category of the second sample pair includes two audio samples with the same semantics but different timbres, and the second category of the second sample pair includes two audio samples with the same semantics and the same timbre; The sixth training dataset includes multiple groups of first sample groups, and the first sample group includes, under the same semantics, text samples and audio samples in one language and text samples and audio samples in another language.

9. A multimodal language conversion device, characterized in that, The device includes: An encoding processing unit, in response to receiving input data, is configured to perform encoding processing on the input data by using a multimodal encoding unit to obtain a vector representation of the input data; the type of the input data includes at least one of text data and audio data; A decoding processing unit is configured to perform respective decoding processing on the vector representation by using a plurality of decoding units in a multimodal decoding unit to obtain target data after language conversion of the input data and output the target data; the plurality of decoding units are respectively obtained by associative training of different basic decoding units, and each of the decoding units is related to a multimodal cross-language understanding function.

10. An electronic device, characterized in that, It includes a memory, a processor, and a machine-executable program stored on the memory and running on the processor, and when the processor executes the machine-executable program, the electronic device executes the multimodal language conversion method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Translation method, device and equipment and storage medium

    CN112668346A

  • Speech generation and understanding system and method based on large language model, and electronic equipment

    CN118155638A

Cited By

  • Cross-modal conversion method and system suitable for multiple languages and multiple voices, and medium

    CN120766659A

  • A method, system, and medium for cross-modal conversion applicable to multiple languages ​​and multiple speech languages.

    CN120766659B