Cross-modal cross-language voice large model training method and system

Through the cross-modal cross-language alignment strategy, the trained speech model aligns speech and text within the same language, and aligning different languages through text modality, solving the difficulties of cross-modal and cross-language alignment in the existing technology, and achieving efficient completion of single-language and cross-language speech dialogue tasks.

CN120279897APending Publication Date: 2025-07-08INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510381792.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing speech model has difficulties in cross-modal and cross-language alignment, and cannot effectively complete cross-language speech tasks, and has disadvantages in the quality and diversity of the speech generation end.

Method used

The cross-modal cross-language alignment strategy is adopted, and the multilingual voice text parallel data set and text instruction data set are collected, and the large language model is pre-trained and word list expansion is performed. The alignment method of connection timing classification is used to perform cross-modal alignment within the same language, and the text modal alignment is aligned between different languages through text modality, and a single language or cross-language voice instruction data set is constructed, and supervised fine-tuning is performed to train the speech model.

Benefits of technology

The efficient alignment of the speech model in single-language and cross-language dialogue tasks is realized, improving the accuracy of the output language and the performance of cross-language dialogue tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279897A_ABST
    Figure CN120279897A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal cross-language voice large model training method and system, and the method comprises the steps: collecting a multi-language voice text parallel data set and a text instruction data set, and obtaining voice recognition and voice synthesis data; merging the data sets, and performing pre-training and word list expansion of a large language model; the method comprises the following steps: carrying out cross-modal alignment on voice and texts in the same language by adopting an alignment method of connecting time sequence classification, carrying out cross-language alignment on different languages through texts, constructing and generating a single-language or cross-language voice instruction data set, and training to obtain a large voice model for completing a single-language or cross-language voice conversation task; and performing supervised fine tuning by adopting voice dialogue instruction data, and reasoning and applying a pre-trained voice large model. According to the method and the system, cross-modal and cross-language alignment is achieved on the voice large model, so that errors in the language output by the voice large model are less, and meanwhile, the method and the system have better performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of large language models, and particularly relates to a cross-modal and cross-language speech large model training method and system thereof. Background Art

[0002] In recent years, the related technologies of large language models have developed rapidly, making it possible for chatbots, intelligent assistants, etc. based on the text modality. However, as the application of large language model technology gradually deepens, people hope that large models can interact with humans more conveniently and explore more diverse possibilities of large models in other modalities. Large language models supporting the speech modality have emerged as the times require.

[0003] However, building a speech large model is not easy. Speech resources are limited compared to the text modality. The speech data used in a specific language and the language types contained in the speech data are insufficient. Therefore, it is necessary to combine with a text large model and utilize the knowledge of the text modality. The simplest combination method is to cascade a text large model with a speech recognition model and a speech synthesis model, but this method brings problems such as high latency and error accumulation.

[0004] To solve the above problems, the current mainstream technology is to use a single speech large model to process speech input and output end-to-end. This new method still faces challenges in terms of implementation methods and model performance. There are currently different ways to implement an end-to-end speech large model, but each has some problems.

[0005] Some researchers train speech models in the same way as previous text language models, but the amount of data in the speech modality is naturally much less than that in the text modality; some researchers hope to establish a mapping between an existing speech encoder and a text large model, but it has disadvantages at the speech generation end; there are also some researchers who try to perform unified modeling of speech and text, and this method requires a large amount of data.

[0006] In addition, in order to enable the speech large model to have a broader application space, people also hope that it can support multiple languages. How to enable the large model to learn different languages well while modeling speech information and establish a mapping between different languages is also an urgent problem to be solved.

[0007] In summary, the following disadvantages exist in the prior art:

[0008] 1) For a large model for pure speech, although it has the ability to model the context relationship of speech, the data in the speech modality is naturally of a different order of magnitude from that in the text modality, and a pure speech model cannot have the powerful general task ability of a text large model;

[0009] 2) For the method of connecting a pre-trained speech encoder to a pre-trained large text model, a large speech model with a speech encoder and a large model structure is obtained, but it has certain disadvantages in terms of the quality and diversity at the speech generation end.

[0010] 3) For the attempt to jointly model discrete speech units and text, the model requires a large amount of training data for training and has achieved good results in single-language speech tasks. However, they lack the ability of cross-language alignment and cannot complete some cross-language speech tasks.

[0011] In summary, the existing large speech model technologies are only trained on single-language speech tasks and do not perform cross-modal and cross-language alignment simultaneously, resulting in the inability to complete some cross-language speech tasks.

[0012] Therefore, based on the above analysis, it is urgent to explore a cross-language and cross-modal alignment strategy under a suitable large speech model architecture to train a multilingual large speech model. Research and develop a set of solutions that can simultaneously achieve cross-modal and cross-language alignment, collect relevant training data, and endow the large speech model with alignment capabilities. At the same time, a single-language and cross-language speech instruction dataset is constructed to further endow the model with the ability to complete single-language and cross-language speech dialogue tasks. Summary of the Invention

[0013] To solve the defect that the large speech model in the above existing technologies cannot simultaneously achieve cross-modal and cross-language alignment, based on the existing large speech model structure, the present invention proposes a cross-modal and cross-language alignment strategy, which aligns the speech and text modalities within the same language and aligns different languages through the text modality.

[0014] In a first aspect, an embodiment of the present application provides a cross-modal and cross-language large speech model training method, and the method includes:

[0015] Data collection step: Collect a multilingual speech-text parallel dataset and a text instruction dataset to obtain speech recognition and speech synthesis data;

[0016] Model pre-training step: Combine the text instruction dataset with the speech recognition and speech synthesis data to perform pre-training of the large language model and vocabulary expansion;

[0017] Speech dialogue instruction data construction step: Based on the text instruction dataset, adopt the connectionist temporal classification alignment method to perform cross-modal alignment of speech and text within the same language and cross-language alignment through text between different languages, construct and generate a single-language or cross-language speech instruction dataset, and based on the speech instruction dataset, train a large speech model that can complete single-language or cross-language speech dialogue tasks;

[0018] Model fine-tuning step: Based on the pre-trained large speech model, supervised fine-tuning is performed using speech dialogue instruction data.

[0019] In the embodiments of the present invention, for the above-mentioned cross-modal and cross-language large speech model training method, the method further includes:

[0020] Model inference step: Input a piece of speech to be inferred into the pre-trained large speech model, and infer and reconstruct it into an audible speech waveform.

[0021] In the embodiments of the present invention, the above data collection step further includes:

[0022] Collect speech recognition and speech synthesis data sets within the same language or between multiple languages, perform cross-modal or cross-language alignment to obtain cross-modal or cross-language aligned data;

[0023] Select a speech tokenizer to discretize the speech part of the speech recognition and speech synthesis data sets into discrete speech units;

[0024] Process the cross-modal or cross-language aligned data into a speech instruction format, where the instruction format includes: speech recognition instructions, speech synthesis instructions, and machine translation instructions.

[0025] In the embodiments of the present invention, the above speech dialogue instruction data construction step further includes:

[0026] Cross-modal alignment step: Collect speech recognition and speech synthesis data sets in the same language, perform cross-modal alignment; select a speech tokenizer to discretize the speech part of the speech recognition and speech synthesis data sets into discrete speech units; process the cross-modal aligned data into a speech instruction format, where the instruction form includes speech recognition instructions and speech synthesis instructions;

[0027] Cross-language alignment step: Collect machine translation data between different languages, perform cross-language alignment; combine the machine translation data with speech data to construct a cross-language speech instruction data set; perform alignment through the text modality between different languages: for bilingual data, align the speech and text of one language with the text of another language; use the text modality to map the speech and text of different languages to the same semantic space; establish an association between the speech and text of different languages through the shared text modality.

[0028] In the embodiments of the present invention, the above supervised instruction fine-tuning step of the large model further includes:

[0029] Initially input a speech instruction to the large speech model, and the large speech model outputs a corresponding text instruction;

[0030] Continue to let the large speech model alternately output a text response and a speech response corresponding to the text response until a termination symbol is generated;

[0031] The cross-entropy loss method is used as the supervision method until the speech large model is fine-tuned.

[0032] In a second aspect, an embodiment of the present application provides a cross-modal and cross-language speech large model training system. Using the cross-modal and cross-language speech large model training method as described above, the system includes:

[0033] Data collection module: used to collect multi-language speech text parallel datasets and text instruction datasets to obtain speech recognition and speech synthesis data;

[0034] Model pre-training module: used to merge the text instruction dataset with the speech recognition and speech synthesis data for pre-training of the large language model and vocabulary expansion;

[0035] Speech dialogue instruction data construction module: used to perform cross-modal alignment of speech and text within the same language and cross-language alignment through text between different languages based on the text instruction dataset by using the connectionist temporal classification alignment method, construct and generate a single-language or cross-language speech instruction dataset, and train a speech large model to complete single-language or cross-language speech dialogue tasks based on the speech instruction dataset;

[0036] Model fine-tuning module: used to perform supervised fine-tuning based on the pre-trained speech large model using speech dialogue instruction data.

[0037] Model inference module: used to input a speech to be inferred into the pre-trained speech large model and infer and reconstruct it into an audible speech waveform.

[0038] In an embodiment of the present invention, the above data collection module further includes:

[0039] Collect speech recognition and speech synthesis datasets within the same language or between multiple languages, perform cross-modal or cross-language alignment to obtain cross-modal or cross-language aligned data;

[0040] Select a speech tokenizer to discretize the speech part of the speech recognition and speech synthesis datasets into discrete speech units;

[0041] Process the cross-modal or cross-language aligned data into a speech instruction format, where the instruction format includes: speech recognition instruction, speech synthesis instruction, and machine translation instruction.

[0042] In an embodiment of the present invention, the above speech dialogue instruction data construction module further includes:

[0043] Cross-modal alignment module: Collect speech recognition and speech synthesis datasets in the same language for cross-modal alignment; select a speech tokenizer to discretize the speech part of the speech recognition and speech synthesis datasets into discrete speech units; process the cross-modal alignment data into a speech instruction format, where the instruction forms include speech recognition instructions and speech synthesis instructions.

[0044] Cross-language alignment module: Collect machine translation data between different languages for cross-language alignment; combine the machine translation data with speech data to construct a cross-language speech instruction dataset; perform alignment between different languages through the text modality: for bilingual data, align the speech and text of one language with the text of another language; use the text modality to map the speech and text of different languages to the same semantic space; establish an association between the speech and text of different languages through the shared text modality.

[0045] Thirdly, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and the steps of the cross-modal and cross-language speech large model training method are implemented when the program is executed by a processor.

[0046] Fourthly, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the cross-modal and cross-language speech large model training method as described are implemented.

[0047] Compared with the related prior art, it has the following prominent beneficial effects:

[0048] 1) The present invention proposes a cross-modal and cross-language alignment scheme for a speech large model, aligning the speech and text modalities within the same language, and aligning between different languages through the text modality (specific implementation steps need to be added in the corresponding specific embodiments), expanding the cross-modal and cross-language alignment capabilities of the large model;

[0049] 2) The present invention designs a scheme for constructing single-language and cross-language speech instruction datasets, obtaining a sentence-level speech-text interleaved speech instruction dataset, including speech-text pairs, and training a speech large model capable of completing single-language and cross-language speech dialogue tasks;

[0050] Compared with the baseline model, the method of the present invention achieves cross-modal and cross-language alignment simultaneously on the speech large model, making the speech large model make fewer mistakes in the output language when performing single-language or cross-language dialogue tasks, and having better performance in cross-language dialogue tasks. Brief Description of the Drawings

[0051] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the accompanying drawings:

[0052] Figure 1 It is a schematic diagram of the method for training a cross-modal and cross-language speech large model of the present invention;

[0053] Figure 2 It is a schematic diagram of the training process of the speech large model according to an embodiment of the present invention;

[0054] Figure 3 It is a schematic diagram of the cross-modal and cross-language speech large model training system of the present invention;

[0055] Figure 4 It is a schematic diagram of the computer hardware of the present invention. Detailed implementation manners

[0056] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one (item)" or a similar expression thereof refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.

[0057] It should also be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.

[0058] It should also be understood that in various embodiments of the present invention, the magnitudes of the serial numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0059] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0060] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0061] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0062] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0063] To make the above features and effects of the present invention more clearly understandable, specific embodiments are hereinafter given and detailed descriptions are made in conjunction with the accompanying drawings of the specification as follows. This specification discloses one or more embodiments including the features of the present invention. The disclosed embodiments are only for illustrative purposes. The protection scope of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.

[0064] The following is a system embodiment corresponding to the method embodiment above. This embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.

[0065] The method of the present invention aims to explore a cross - language and cross - modality alignment strategy under a suitable speech large - model architecture to train a multilingual speech large - model.

[0066] The present invention attempts to train a multilingual speech large - model on the existing speech large - model architecture based on discrete speech units. The present invention attempts to design a scheme for simultaneous cross - modality and cross - language alignment, collect relevant training data, and endow the speech large - model with alignment capabilities. The present invention also constructs monolingual and cross - language speech instruction data sets to further endow the model with the ability to complete monolingual and cross - language speech dialogue tasks.

[0067] Based on the existing speech large - model structure, the present invention proposes a cross - modality and cross - language alignment strategy. Within the same language, the speech and text modalities are aligned, and between different languages, they are aligned through the text modality. At the same time, a scheme for constructing monolingual and cross - language speech instruction data sets is designed, and speech instruction data with sentence - level speech - text interleaving is obtained, and then a speech large - model capable of completing monolingual and cross - language speech dialogue tasks is trained.

[0068] The following will explain the method of the embodiment of the present application in detail with specific embodiments:

[0069] Embodiment 1

[0070] As Figure 1 shown, the present invention proposes a cross - language and cross - modality alignment strategy to train a multilingual speech large - model. On the existing speech large - model based on discrete speech units, the present invention proposes to align the speech and text modalities within the same language, and align different languages through the text modality, and use the corresponding data to endow the large - model with relevant alignment capabilities. The present invention collects open - source speech - text parallel data and conducts the first - stage continued pre - training on the open - source text large - model. The present invention proposes a set of schemes for constructing monolingual and cross - language speech instruction data sets, constructs a speech dialogue instruction data set, and conducts supervised instruction fine - tuning on the first - stage continued pre - training model to train a multilingual speech large - model with cross - modality and cross - language alignment.

[0071] The working process of the multilingual speech large - model with cross - modality and cross - language alignment is as Figure 2 shown. The embodiment of the present application provides a method for training a cross - modality and cross - language speech large - model. The method includes:

[0072] Data collection step 101: Collect a parallel dataset of multilingual speech texts and a text instruction dataset to obtain speech recognition and speech synthesis data;

[0073] Model pre-training step 102: Merge the text instruction dataset with the speech recognition and speech synthesis data, and perform pre-training of the large language model and vocabulary expansion;

[0074] Specifically, in the specific embodiments of the present invention, the continued pre-training of the large model includes: constructing a speech large model based on an open-source large language model. First, expand the vocabulary of the large language model by merging the vocabulary of the speech tokenizer with the vocabulary of the large language model. Through the constructed speech-text parallel data, additional machine translation data and text instruction data, continue to pre-train the large language model to achieve cross-language and cross-modal alignment. Supervise this process with cross-entropy loss. In the vocabulary expansion stage of the large language model, we deeply fuse the vocabulary of the speech tokenizer with the vocabulary of the large language model. Specifically, the speech tokenizer can quantize continuous speech waveforms into discrete speech units, and the vocabulary composed of these units is unified with the vocabulary originally used by the large language model to process text in terms of dimension and semantic understanding. Through mapping and conversion operations, we make it correspond to the text vocabulary of the large language model in the same semantic space, so that the model can process both speech and text inputs simultaneously. In the data integration and pre-training stage, we innovatively construct a comprehensive training dataset, which not only includes speech-text parallel data in multiple languages, but also innovatively introduces machine translation data and text instruction data. The speech-text parallel data is used to achieve alignment between the speech and text modalities, the machine translation data helps with cross-language alignment between different languages, and the text instruction data can further enhance the model's ability in text understanding and services. During the model training process, we use cross-entropy loss as a supervision signal to optimize the model. Finally, the model has the ability of cross-language alignment and cross-modal alignment at the same time.

[0075] Speech dialogue instruction data construction step 103: Based on the text instruction dataset, use the alignment method of connectionist temporal classification to perform cross-modal alignment of speech and text within the same language, and cross-language alignment through text between different languages, construct and generate a monolingual or cross-language speech instruction dataset, and based on the speech instruction dataset, train a speech large model to complete monolingual or cross-language speech dialogue tasks;

[0076] Model fine-tuning step 104: Based on the pre-trained speech large model, perform supervised fine-tuning using speech dialogue instruction data.

[0077] Model inference step 105: Input a speech to be inferred into the pre-trained speech large model, and infer and reconstruct it into an audible speech waveform.

[0078] Specifically, the inference application in the specific embodiments of the present invention includes: during model inference, first, a piece of speech is input. The speech tokenizer converts it into discrete speech units and inputs them to the trained speech large model. The speech large model will output text instructions, and then alternately output a text reply and a corresponding speech reply. The generated speech replies are still discrete speech units, and they are input to the speech decoder module, which reconstructs these discrete speech units into audible speech waveforms.

[0079] In the embodiments of the present invention, the above data collection step 101 further includes:

[0080] Collect speech recognition and speech synthesis data sets within the same language or between multiple languages, perform cross-modal or cross-language alignment, and obtain cross-modal or cross-language aligned data;

[0081] Select a speech tokenizer to discretize the speech part of the speech recognition and speech synthesis data sets into discrete speech units;

[0082] Process the cross-modal or cross-language aligned data into a speech instruction format, where the instruction format includes: (1) a speech recognition instruction, that is, an instruction to convert speech data into written text, such as "Convert this Chinese speech into text."; (2) a speech synthesis instruction, that is, an instruction to convert text data into speech, such as "Convert the following Chinese text into speech."; (3) a machine translation instruction, that is, an instruction to convert text data in one language into text data in another semantically equivalent language, such as "Please translate this Chinese into English."

[0083] Specifically, the collection of speech-text parallel data in the embodiments of the present invention includes: collecting open-source Chinese and English speech recognition and speech synthesis data sets for cross-modal alignment. Select an open-source speech tokenizer to discretize the speech part therein into discrete speech units. Finally, process these cross-modal aligned data into an instruction format.

[0084] In the embodiments of the present invention, the above speech dialogue instruction data construction step 103 further includes:

[0085] Cross-modal alignment step:

[0086] 1) Collect speech recognition and speech synthesis data sets in the same language for cross-modal alignment.

[0087] 2) Select a speech tokenizer to discretize the speech part of the speech recognition and speech synthesis data sets into discrete speech units.

[0088] 3) Process the cross-modal aligned data into a speech instruction format, where the instruction form includes a speech recognition instruction and a speech synthesis instruction.

[0089] Cross - language alignment steps:

[0090] 1) Collect machine translation data between different languages for cross - language alignment.

[0091] 2) Combine the machine translation data with speech data to construct a cross - language speech instruction dataset.

[0092] 3) Align through the text modality between different languages: For bilingual data, align the speech and text of one language with the text of another language; use the text modality as a bridge to map the speech and text of different languages to the same semantic space; through the shared text modality,

[0093] Establish the association between the speech and text of different languages.

[0094] Specifically, in the embodiments of the present invention, a cross - modal and cross - language alignment scheme for a speech large model is proposed. The speech and text modalities are aligned within the same language, and the alignment between different languages is achieved through the text modality. In the data collection stage, the present invention collects multilingual speech - text parallel datasets and text instruction datasets, obtains speech recognition and synthesis data, discretizes the speech data into discrete speech units convenient for model processing using a speech tokenizer, and then processes the cross - modal or cross - language alignment data into a speech instruction format including speech recognition, synthesis, and machine translation instructions to better guide the model to execute tasks. In the model pre - training stage, the present invention combines the text instruction dataset with the speech recognition and synthesis datasets, performs pre - training of the large language model and vocabulary expansion, and by combining the vocabulary of the speech tokenizer and the large language model, enables the model to process both speech and text inputs simultaneously, and optimizes the parameters using cross - entropy loss during training, enabling the model to have cross - language and cross - modal alignment capabilities.

[0095] Specifically, in the embodiments of the present invention, a scheme for constructing monolingual and cross - language speech instruction datasets is proposed, obtaining a sentence - level speech - text interleaved speech instruction dataset, including speech - text pairs, and training a speech large model capable of completing monolingual and cross - language speech dialogue tasks;

[0096] In the embodiments of the present invention, the existing text instruction dataset is rewritten using a large language model to make its content more suitable for the speech dialogue scenario. The specific steps include: collecting the existing text instruction dataset and speech dialogue scenario corpus; selecting an open - source pre - trained large language model with good Chinese - English capabilities; during the rewriting process, adopting strategies such as colloquial expression, simplicity principle, natural fluency, and semantic accuracy to make the instructions more in line with the speech dialogue habit; finally, organizing the rewritten instruction data to form a new instruction dataset to improve the interaction effect and user experience of the speech large model.

[0097] The obtained instructions and response texts are respectively synthesized into speech using a speech synthesis model. In the present invention, the process of synthesizing the obtained instructions and response texts into speech is as follows: First, a suitable open-source speech synthesis model is selected, which can generate relatively natural and fluent speech. Then, the text instructions rewritten by the large model previously are selected. Next, the selected speech synthesis model is used to perform speech synthesis on the text. The text is input into the model to generate the corresponding speech waveform.

[0098] For the obtained text response and the corresponding speech response, with the help of an alignment module based on connectionist temporal classification technology, the sentence-level alignment relationship between them is obtained, and thus the response composition text and speech are interleaved at the sentence level. When aligning the obtained text response and the corresponding speech response, the present invention adopts an alignment module based on connectionist temporal classification (CTC) technology. The specific implementation steps are as follows: First, the speech data is converted into a speech representation through a speech encoder, and the text data is tokenized. Then, using the dynamic programming algorithm of the CTC model, the optimal alignment path between the speech and the text is established. Specifically, for a given speech-text pair, through the CTC alignment module, the optimal alignment path between the speech and the text on the time axis can be obtained. Then, applying the CTC dynamic programming algorithm, the best path for aligning the speech and the text is found, so as to obtain the time boundaries of each text word in the speech. Finally, based on this alignment relationship, the text and the speech are interleaved at the sentence level to form a format in which the text and the speech alternate at the sentence level. For example, a sentence of text appears first, followed by the corresponding speech, then the next sentence of text and its corresponding speech, and so on, thus realizing the interleaving of the text and the speech at the sentence level.

[0099] In the embodiments of the present invention, using the above methods, data for Chinese speech conversations and data for English speech conversations are respectively constructed. In addition, using the existing Chinese-English parallel text instructions, speech instruction data is obtained in the above manner, and cross-language speech instruction data is respectively formed according to two types: Chinese input and English output, and English input and Chinese output. The present invention is not limited to this, and other languages such as French, German, etc. can also be used.

[0100] In the embodiments of the present invention, the above-mentioned supervised instruction fine-tuning step 104 of the large model further includes:

[0101] Initially, a speech instruction is input to the speech large model, and the speech large model outputs the corresponding text instruction;

[0102] Continue to make the speech large model alternately output a sentence of text response and the speech response corresponding to the text response until a termination symbol is generated;

[0103] Use the cross-entropy loss method for supervision until the speech large model is fine-tuned.

[0104] Specifically, the supervised instruction fine-tuning of the large model in the specific embodiment of the present invention includes: continuing with the pre-trained model and performing supervised fine-tuning with speech dialogue instruction data at this stage. Input a speech instruction to the speech large model. First, let the large model output the corresponding text instruction, and then let it alternately output a text response and the corresponding speech response until a termination symbol is generated. Supervise this process with cross-entropy loss.

[0105] Embodiment 2

[0106] As Figure 3 shown, the embodiment of the present application provides a cross-modal and cross-language speech large model training system. Using the above-mentioned cross-modal and cross-language speech large model training method, the system includes:

[0107] Data collection module 201: used to collect multi-language speech-text parallel data sets and text instruction data sets to obtain speech recognition and speech synthesis data;

[0108] Model pre-training module 202: used to merge the text instruction data set with the speech recognition and speech synthesis data for pre-training of the large language model and vocabulary expansion;

[0109] Speech dialogue instruction data construction module 203: based on the text instruction data set, using the alignment method of connectionist temporal classification to perform cross-modal alignment of speech and text within the same language and cross-language alignment through text between different languages, construct and generate a single-language or cross-language speech instruction data set, and based on the speech instruction data set, train to obtain a speech large model that completes single-language or cross-language speech dialogue tasks;

[0110] Model fine-tuning module 204: used to perform supervised fine-tuning based on the pre-trained speech large model with speech dialogue instruction data.

[0111] Model inference module 205: used to input a speech to be inferred into the pre-trained speech large model and infer and reconstruct it into an audible speech waveform.

[0112] In the embodiment of the present invention, the above data collection module 201 further includes:

[0113] Collect speech recognition and speech synthesis data sets within the same language or between multiple languages, perform cross-modal or cross-language alignment to obtain cross-modal or cross-language aligned data;

[0114] Select a speech tokenizer to discretize the speech part of the speech recognition and speech synthesis data set into discrete speech units;

[0115] Process cross-modal or cross-language alignment data into speech instruction formats, where the instruction formats include: (1) Speech recognition instruction, which is an instruction to convert speech data into written text, such as "Convert this Chinese speech into text."; (2) Speech synthesis instruction, which is an instruction to convert text data into speech, such as "Convert the following Chinese text into speech."; (3) Machine translation instruction, which is an instruction to convert text data in one language into text data in another semantically equivalent language, such as "Please translate this Chinese into English."

[0116] In the embodiments of the present invention, the above-mentioned speech dialogue instruction data construction module 203 further includes:

[0117] Cross-modal alignment module:

[0118] 1) Collect speech recognition and speech synthesis data sets in the same language and perform cross-modal alignment.

[0119] 2) Select a speech tokenizer to discretize the speech part of the speech recognition and speech synthesis data sets into discrete speech units.

[0120] 3) Process the cross-modal alignment data into speech instruction formats, where the instruction forms include speech recognition instructions and speech synthesis instructions.

[0121] Cross-language alignment module:

[0122] 1) Collect machine translation data between different languages and perform cross-language alignment.

[0123] 2) Combine the machine translation data with the speech data to construct a cross-language speech instruction data set.

[0124] 3) Align through the text modality between different languages: For bilingual data, align the speech and text in one language with the text in another language; use the text modality as a bridge to map the speech and text in different languages to the same semantic space; through the shared text modality,

[0125] Establish the association between the speech and text in different languages.

[0126] Embodiment III

[0127] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the cross-modal and cross-language speech large model training method are implemented.

[0128] Embodiment IV

[0129] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the cross-modal and cross-language speech large model training method as described are implemented.

[0130] In addition, the cross-modal and cross-language speech large model training method described in conjunction with Figure 1 the embodiments of the present application can be implemented by an electronic device, such as a computer device. Figure 4 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present application.

[0131] In some of these embodiments, the computer device may further include a communication interface 83 and a bus 80. Among them, as Figure 4 shown, the processor 81, the memory 82, and the communication interface 83 are connected through the bus 80 and complete communication with each other.

[0132] Specifically, the above-mentioned processor 81 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits implementing the embodiments of the present application.

[0133] The memory 82 may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 81.

[0134] The processor 81 reads and executes the computer program instructions stored in the memory 82 to implement any one of the above-mentioned hyper-realistic digital human video detection methods.

[0135] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0136] The above-described embodiments only express several implementation manners of the present application, and their descriptions are relatively specific and detailed, but should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A training method for a cross-modal and cross-language speech large model, characterized in that, The method includes: Data collection step: Collect a parallel dataset of multilingual speech texts and a text instruction dataset to obtain speech recognition and speech synthesis data; Model pre-training step: Combine the text instruction dataset with the speech recognition and speech synthesis data to perform pre-training of the large language model and vocabulary expansion; Speech dialogue instruction data construction step: Based on the text instruction dataset, adopt the alignment method of connectionist temporal classification to perform cross-modal alignment between speech and text within the same language, and cross-language alignment through text between different languages, construct and generate a monolingual or cross-language speech instruction dataset, and based on the speech instruction dataset, train a speech large model to complete monolingual or cross-language speech dialogue tasks; Model fine-tuning step: Based on the pre-trained speech large model, perform supervised fine-tuning using speech dialogue instruction data.

2. The cross-modal and cross-language speech large model training method according to claim 1, wherein, The method further includes: Model inference step: Input a speech to be inferred into the pre-trained speech large model, and infer and reconstruct it into an audible speech waveform.

3. The cross-modal and cross-language speech large model training method according to claim 1 or 2, characterized in that The data collection step further includes: Collect speech recognition and speech synthesis datasets within the same language or between multiple languages, perform cross-modal or cross-language alignment to obtain cross-modal or cross-language aligned data; Select a speech tokenizer to discretize the speech part of the speech recognition and speech synthesis dataset into discrete speech units; Process the cross-modal or cross-language aligned data into a speech instruction format, where the instruction format includes: speech recognition instruction, speech synthesis instruction, and machine translation instruction.

4. The cross-modal and cross-language speech large model training method according to claim 1, wherein The speech dialogue instruction data construction step further includes: Cross-modal alignment step: Collect speech recognition and speech synthesis datasets in the same language, perform cross-modal alignment; select a speech tokenizer to discretize the speech part of the speech recognition and speech synthesis dataset into discrete speech units; process the cross-modal aligned data into a speech instruction format, where the instruction form includes speech recognition instruction and speech synthesis instruction; Cross-language alignment step: Collect machine translation data between different languages, perform cross-language alignment; combine the machine translation data with speech data to construct a cross-language speech instruction dataset; perform alignment between different languages through the text modality: for bilingual data, align the speech and text of one language with the text of another language; use the text modality to map the speech and text of different languages to the same semantic space; establish an association between the speech and text of different languages through the shared text modality.

5. The cross-modal and cross-language speech large model training method according to claim 1, wherein The supervised instruction fine-tuning step of the large model further includes: Initially input a speech instruction to the speech large model, and the speech large model outputs a corresponding text instruction; Continue to let the speech large model alternately output a text response and a speech response corresponding to the text response until a termination symbol is generated; Adopt the cross-entropy loss method as a supervision method for supervision until the speech large model completes fine-tuning.

6. A cross-modal and cross-language speech large model training system, which adopts the cross-modal and cross-language speech large model training method described in any one of claims 1-5, characterized in that, The system includes: Data collection module: Used to collect a parallel dataset of multilingual speech texts and a text instruction dataset to obtain speech recognition and speech synthesis data; Model pre-training module: used to merge the text instruction dataset and the speech recognition and speech synthesis data, and perform pre-training of the large language model and vocabulary expansion; Speech dialogue instruction data construction module: used to perform cross-modal alignment of speech and text within the same language and cross-language alignment through text between different languages based on the text instruction dataset, construct and generate a mono-language or cross-language speech instruction dataset, and train a speech large model to complete mono-language or cross-language speech dialogue tasks based on the speech instruction dataset; Model fine-tuning module: used to perform supervised fine-tuning based on the pre-trained speech large model using speech dialogue instruction data. Model inference module: used to input a piece of speech to be inferred into the pre-trained speech large model and infer and reconstruct it into an audible speech waveform.

7. The cross-modal and cross-language speech large model training system according to claim 6, wherein The data collection module further includes: Collect speech recognition and speech synthesis datasets within the same language or between multiple languages, perform cross-modal or cross-language alignment, and obtain cross-modal or cross-language aligned data; Select a speech tokenizer to discretize the speech part of the speech recognition and speech synthesis dataset into discrete speech units; Process the cross-modal or cross-language aligned data into a speech instruction format, where the instruction format includes: speech recognition instruction, speech synthesis instruction, and machine translation instruction.

8. The cross-modal and cross-language speech large model training system according to claim 6, wherein The speech dialogue instruction data construction module further includes: Cross-modal alignment module: collect speech recognition and speech synthesis datasets in the same language, perform cross-modal alignment; select a speech tokenizer to discretize the speech part of the speech recognition and speech synthesis dataset into discrete speech units; process the cross-modal aligned data into a speech instruction format, where the instruction form includes speech recognition instruction and speech synthesis instruction; Cross-language alignment module: collect machine translation data between different languages, perform cross-language alignment; combine the machine translation data with speech data to construct a cross-language speech instruction dataset; perform alignment through the text modality between different languages: for bilingual data, align the speech and text of one language with the text of another language; use the text modality to map the speech and text of different languages to the same semantic space; establish an association between the speech and text of different languages through the shared text modality.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the cross-modal and cross-language speech large model training method described in any one of claims 1-5.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the cross-modal and cross-language speech large model training method described in any one of claims 1 to 5.