A multi-modal speech interaction large model training method and system based on speech acoustic feature regulation, a terminal device, and a medium

By constructing a speech token that includes semantic and acoustic features, and combining multimodal input samples and a cross-modal feature alignment architecture, the problem of insufficient speech diversity and long speech generation in multimodal speech interaction models is solved, thereby improving the naturalness and adaptability of speech interaction.

CN120954388BActive Publication Date: 2026-01-16HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511461447.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2026-01-16
Estimated Expiration
2045-10-14

AI Technical Summary

Technical Problem

Existing multimodal voice interaction models have shortcomings such as insufficient voice diversity, difficulty in generating long speech, lack of datasets, and mismatch between text and speech lengths, resulting in poor interactive experience and high resource consumption.

Method used

By constructing speech tokens that include semantic and acoustic features, and combining them with multimodal input samples, a large-scale multimodal speech interaction model is built. A fine-grained pre-trained speech processing model and a cross-modal feature alignment architecture are adopted, and a deep learning network is designed for multi-stage training to improve the ability to understand and generate long speech.

Benefits of technology

It improves the capabilities of multimodal voice interaction models in long speech processing and diverse speech generation, enhances the naturalness and adaptability of voice interaction, reduces resource consumption, and supports emotion recognition and applications in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954388B_ABST
    Figure CN120954388B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-modal speech interaction big model training method, system, terminal equipment and medium of voice acoustic feature regulation, it is related to multi-modal speech interaction technical field, the method includes: obtaining the text token of text training sample and constructs corresponding voice token, obtains the pre-training data for converting text token into voice token;Combining multi-modal input sample and pre-training data, construct fine-tuning training data for speech understanding and dialogue generation;Using pre-training data constructs and pre-trains basic model;Based on pre-training basic model, build multi-modal speech interaction big model, with fine-tuning data training, so that it can be based on multi-modal input regulation voice acoustic feature and output voice.The application is trained by the alignment and stage of text token and voice token, realize the fine regulation of voice acoustic feature, improve long speech coherence and interactive naturalness, efficiently give model controllable timbre, the speech interaction ability of emotion.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of multi-modal voice interaction, and in particular to a multi-modal voice interaction large model training method and system based on voice acoustic feature regulation, a terminal device and a medium. BACKGROUND

[0002] In the field of multi-modal voice interaction, multi-modal large model GPT-4o realizes real-time multi-modal voice interaction, promoting the development of end-to-end voice-to-voice systems. Such systems can directly process the voice modality without generating intermediate text, improving the naturalness of interaction and response speed. At the same time, the emergence of open-source large language model frameworks such as Moshi, Mini-Omni, and LLaMA-Omni enables multi-modal large language models to incorporate voice capabilities, reducing the latency of traditional voice-text-voice systems and solving the alignment problem of audio and text modalities.

[0003] However, the existing technology still has deficiencies. First, the answer voice of the voice interaction large model lacks diversity, and most voice codecs excessively compress audio, resulting in monotonous tone and lack of emotional and other secondary semantic expression capabilities, making it difficult to meet individualized interaction needs. These problems affect the practicality and experience of multi-modal voice interaction systems. Second, the existing voice interaction large model has weak long voice processing capabilities, and the long voice understanding and generation dataset is scarce. Moreover, after the voice codec encodes the long voice, the number of processing units far exceeds that of long text, requiring high training resources and model context capabilities, and the mismatch between text and voice length also limits generation.

[0004] Therefore, there is an urgent need for a multi-modal voice interaction large model training method that can efficiently process long voice, generate diverse voice, and adapt to commonly used multi-modal models to fill the gap in existing technology. SUMMARY

[0005] The technical problem to be solved by the present application is that in the field of multi-modal voice interaction, the existing model voice codec excessively compresses, resulting in insufficient diversity of generated voice, difficulty in generating long voice, lack of related dataset, and mismatch between text and voice length. The multi-modal model voice interaction capability consumes a lot of migration resources and is difficult to accurately capture modal semantic associations. There are also problems such as low training efficiency and poor scalability. Therefore, an effective solution is needed to solve the above technical problems.

[0006] To solve the above technical problems, the technical solution adopted by the present application is as follows:

[0007] In a first aspect, the present application provides a multi-modal voice interaction large model training method based on voice acoustic feature regulation, comprising:

[0008] obtaining text tokens of text training samples, and constructing speech tokens containing semantic features and speech acoustic features expressed by the text tokens based on the text tokens, to obtain pre-training data for converting text tokens into speech tokens;

[0009] constructing fine-tuning training data for speech understanding and dialogue generation based on the multimodal input samples and the pre-training data, wherein the fine-tuning training data contains speech semantic understanding labels and dialogue generation logic labels;

[0010] constructing a base model and pre-training the base model based on the pre-training data, wherein the base model is used to convert input text tokens into speech tokens;

[0011] constructing a multimodal speech interaction large model based on the pre-trained base model, and training the multimodal speech interaction large model using the fine-tuning training data, wherein the multimodal speech interaction large model is used to regulate speech acoustic features for generating speech answers based on multimodal input, and output the speech answers.

[0012] In an implementation manner, the obtaining text tokens of text training samples, and constructing speech tokens containing semantic features and speech acoustic features expressed by the text tokens based on the text tokens, to obtain pre-training data for converting text tokens into speech tokens, comprises:

[0013] obtaining single-person voice training initial data, the single-person voice training initial data including speech format dialogue pure text data and long text data set, wherein the speech format dialogue pure text data includes filtered general natural dialogue text data set and reconstructed dialogue text data set, and the reconstructed dialogue text data set is obtained by reconstructing pure text instruction fine-tuning data based on a preset prompt word through a general large model;

[0014] filtering special characters in reply round text content of the speech format dialogue pure text data and special characters in the long text data set based on a preset rule, and performing symbol division to obtain preprocessed single-person voice training initial data;

[0015] performing token conversion on the single-person voice training initial data to obtain the text tokens, and extracting semantic features expressed by the text tokens;

[0016] performing speech synthesis on the single-person voice training initial data to generate matched single-person voice samples, and extracting speech acoustic features from the single-person voice samples;

[0017] Correlate the text token, the semantic feature expressed by the text token, and the speech acoustic feature to obtain the speech token.

[0018] In an implementation manner, the text token of the text training sample is obtained, and based on the text token, a speech token containing a semantic feature expressed by the text token and a speech acoustic feature is constructed to obtain pre-training data for converting the text token into the speech token, and the pre-training data further includes:

[0019] Obtain multi-person voice training initial data, the multi-person voice training initial data including dialogue text, corresponding multi-person voice data, and corresponding speech acoustic feature annotation;

[0020] Perform data cleaning and token conversion on the multi-person voice training initial data to obtain the text token and extract the semantic feature expressed by the text token;

[0021] Correlate the text token, the semantic feature expressed by the text token, and the speech acoustic feature annotation to obtain the speech token.

[0022] In an implementation manner, the fine-tuning training data includes speech understanding capability alignment data, fine-tuning data of multi-round dialogue of speech interaction, fine-tuning data of text-to-speech generation polishing, and pre-training capability preservation data set, the fine-tuning training data for speech understanding and dialogue generation is constructed based on the multi-modal input sample and the pre-training data, and the fine-tuning training data contains speech semantic understanding annotation and dialogue generation logic annotation, and includes:

[0023] Obtain speech-to-text sample data, perform speech semantic annotation on the speech-to-text sample data, and construct the speech understanding capability alignment data based on the speech-to-text sample data and the corresponding speech semantic annotation, wherein the speech understanding capability alignment data is used for learning of speech understanding capability of the model;

[0024] Based on the multi-round dialogue data in the pre-training data, obtain context association logic of the multi-round dialogue data, and construct the fine-tuning data of multi-round dialogue of speech interaction based on the multi-round dialogue data and the corresponding context association logic, wherein the multi-round dialogue of speech interaction includes dialogue generation logic annotation, and the fine-tuning data of multi-round dialogue of speech interaction is used for learning of context association logic of the model;

[0025] Obtain spoken language voice output sample data, and construct the fine-tuning data of text-to-speech generation polishing, wherein the fine-tuning data of text-to-speech generation polishing is used for learning of spoken language of the model;

[0026] Obtaining a multi-modal input sample, constructing the pre-training ability maintaining dataset, wherein the multi-modal input sample at least includes a visual input sample, a text input sample, and a speech input sample, and the pre-training ability maintaining dataset is used for learning multi-modal recognition ability of a model.

[0027] In an implementation manner, the constructing and pre-training of the base model based on the pre-training data comprises:

[0028] Inputting the text token, the previous speech token, and the preset prompt word in the pre-constructed speech autoregressive module of the base model to obtain a predicted speech token, wherein the speech autoregressive module comprises 24 layers of Transformer blocks;

[0029] Outputting the predicted speech token to a 4096-dimensional prediction head to obtain a predicted synthesized speech;

[0030] Iteratively calculating a loss function and back-propagating to optimize parameters of the base model.

[0031] In an implementation manner, when the base model generates a speech reply, an arbitrary-length synthesized speech is generated by block synthesis, comprising:

[0032] Dividing the input text corresponding to the speech to be generated into several text blocks with complete semantics by taking a period as the only truncation position;

[0033] According to the order of the text blocks, inputting the text token corresponding to the first text block into the pre-trained base model to generate the speech token corresponding to the first text block and a synthesized speech segment through the speech autoregressive module;

[0034] Taking the speech token corresponding to the text block as a speech information prompt, and inputting the text token of the next text block into the base model to generate the corresponding speech token and a synthesized speech segment;

[0035] According to the order of the text blocks, generating synthesized speech segments corresponding to all text blocks, and splicing all synthesized speech segments to obtain an arbitrary-length synthesized speech.

[0036] In an implementation manner, based on the pre-trained base model, a multi-modal speech interaction large model is constructed, and the multi-modal speech interaction large model is trained using the fine-tuning training data, comprising:

[0037] Using a Qformer architecture to modify a Whisper decoder to construct a multi-modal speech interaction large model, and an expression of a structure of the multi-modal speech interaction large model is:

[0038]

[0039]

[0040]

[0041]

[0042]

[0043] wherein, is a fixed-length query vector, is a vector, is a hidden layer output, is an audio, is a query query, is the i-th hidden layer output component constituting the fixed-length query vector , is the top output of the pre-trained audio encoder Whisper encoder, is a multi-head self-attention layer, is a multi-head cross-attention layer, is a layer normalization layer, is a multi-layer perceptron, is the corresponding output content of the Whisper encoder, is the corresponding output content of , is the corresponding output content of , is a multi-head self-attention calculation on the fixed-length query vector , represents the total number of initial vectors, represents the output of the cross-attention module, used to extract the main content of the input audio, is a hidden layer output after multi-layer perceptron processing on the cross-attention module output , after four layers of the same operation, a learnable linear layer is applied to project the final output to the representation space of the large language model;

[0044] using the fine-tuning training data, training the multi-modal voice interaction large model.

[0045] In a second aspect, the embodiments of the present application also provide a multi-modal voice interaction large model training system based on voice acoustic feature regulation, the system comprising:

[0046] The pre-training data construction module is configured to obtain text tokens of a text training sample, and construct speech tokens containing semantic features and speech acoustic features expressed by the text tokens based on the text tokens, to obtain pre-training data for converting the text tokens into the speech tokens.

[0047] The fine-tuning training data construction module is configured to construct fine-tuning training data for speech understanding and dialogue generation based on the multimodal input sample and the pre-training data, wherein the fine-tuning training data contains speech semantic understanding labels and dialogue generation logic labels.

[0048] The base model construction and pre-training module is configured to construct a base model and pre-train the base model based on the pre-training data, wherein the base model is used to convert input text tokens into speech tokens.

[0049] The multimodal speech interaction large model construction and training module is configured to construct a multimodal speech interaction large model based on the pre-trained base model, and train the multimodal speech interaction large model using the fine-tuning training data, wherein the multimodal speech interaction large model is used to regulate speech acoustic features for generating a speech answer based on a multimodal input, and output the speech answer.

[0050] In a third aspect, an embodiment of the present application further provides a terminal device, which comprises a memory, a processor, and a multimodal speech interaction large model training program based on speech acoustic feature regulation stored in the memory and executable on the processor. When the processor executes the multimodal speech interaction large model training program based on speech acoustic feature regulation, the steps of the multimodal speech interaction large model training method based on speech acoustic feature regulation in any of the above solutions are implemented.

[0051] In a fourth aspect, an embodiment of the present application further provides a computer readable storage medium, which stores a multimodal speech interaction large model training program based on speech acoustic feature regulation. When the multimodal speech interaction large model training program based on speech acoustic feature regulation is executed by a processor, the steps of the multimodal speech interaction large model training method based on speech acoustic feature regulation in any of the above solutions are implemented.

[0052] Beneficial effects: A multi-modal voice interaction large model training method, system, terminal device and medium based on voice acoustic feature regulation, related to the technical field of multi-modal voice interaction, the method comprises: first, obtaining the text token of the text training sample, and based on the text token, constructing the voice token containing the semantic features and voice acoustic features expressed by the text token, to obtain the pre-training data for converting the text token into the voice token. Then, based on the multi-modal input sample and the pre-training data, construct the fine-tuning training data for voice understanding and dialogue generation, wherein the fine-tuning training data contains voice semantic understanding annotation and dialogue generation logic annotation. Next, based on the pre-training data, construct the base model and pre-train it, wherein the base model is used to convert the input text token into a voice token. Finally, based on the pre-trained base model, construct a multi-modal voice interaction large model, and use the fine-tuning training data to train the multi-modal voice interaction large model, wherein the multi-modal voice interaction large model is used to regulate the voice acoustic features for generating voice answers based on multi-modal input, and output the voice answers. The present application solves the problems of insufficient voice diversity of existing models, difficulty in long voice generation, and large resource consumption of multi-modal adaptation, realizes acoustic features such as tone and emotion, supports controllable long voice coherent generation, quickly activates the voice interaction ability of the model, while retaining the multi-modal core ability, and improves the voice interaction naturalness and generalization of the multi-modal large model. BRIEF DESCRIPTION OF DRAWINGS

[0053] Figure 1 The flowchart of the specific implementation mode of the multi-modal voice interaction large model training method based on voice acoustic feature regulation provided by the embodiments of the present application is provided.

[0054] Figure 2 The flowchart of the pre-training data construction of the multi-modal voice interaction large model training method based on voice acoustic feature regulation provided by the embodiments of the present application is provided.

[0055] Figure 3 The flowchart of the fine-tuning training data construction in the multi-modal voice interaction large model training method based on voice acoustic feature regulation provided by the embodiments of the present application is provided.

[0056] Figure 4 The schematic diagram of each training stage of the multi-modal voice interaction large model in the multi-modal voice interaction large model training method based on voice acoustic feature regulation provided by the embodiments of the present application is provided.

[0057] Figure 5 The principle block diagram of the multi-modal voice interaction large model training device based on voice acoustic feature regulation provided by the embodiments of the present application is provided.

[0058] Figure 6 is an internal structure principle block diagram of a terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0059] To make the objectives, technical solutions and effects of the present application clearer and more explicit, the present application is further described in detail below with reference to the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and do not limit the present application.

[0060] The flowchart shown in the drawings is only an example and does not necessarily include all contents and operations or steps, nor does it necessarily execute in the order described. For example, some operations or steps can be further divided, combined or partially merged, so the actual execution order can be changed according to the actual situation.

[0061] It should be understood that the terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0062] It should be understood that, in order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, the terms "first", "second" and the like are used to distinguish the same or similar items with basically the same function and role. For example, the first control information and the second control information are only used to distinguish different control information and do not limit the order.

[0063] Those skilled in the art can understand that the terms "first", "second" and the like do not limit the quantity and execution order, and the terms "first", "second" and the like do not necessarily mean different.

[0064] It should also be understood that the term "and / or" used in the present application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.

[0065] In the field of multi-modal speech interaction technology, multi-modal large model GPT-4o realizes real-time multi-modal speech interaction, promoting the development of end-to-end speech-to-speech systems. Such systems do not require intermediate text generation and can directly process speech modalities, improving interaction naturalness and response speed. At the same time, the emergence of open source large language model frameworks such as Moshi, Mini-Omni, and LLaMA-Omni enables multi-modal large language models to incorporate speech capabilities, reducing the latency of traditional speech-text-speech systems and solving the alignment problem of audio-text modalities.

[0066] However, there are still some problems in the prior art. First, the current speech dialogue model is difficult to effectively generate long speech. From the perspective of data set construction, the existing data set for long speech understanding and generation is severely insufficient, which cannot fully activate the comprehensive understanding and generation ability of the model for long speech; from the perspective of model generation, most models using speech codec encode a large amount of long speech, resulting in that the number of speech coding units that the model needs to process is ten to one hundred times that of equal-length text, which puts high requirements on training resources and the context processing ability of the model. In addition, the mismatch between text and speech length also limits the generation ability of many models. Second, the speech codec used by many existing speech models compresses the audio content too much, which finally leads to the generated speech lacking diversity and being monotonous in form, and cannot realize the expression ability beyond semantics, such as realizing the change of timbre and emotion.

[0067] Therefore, in view of the deficiencies of the existing speech dialogue model in long text processing, diversified speech synthesis and multi-modal adaptation, an end-to-end aligned speech dialogue structure is proposed to enhance the long speech interaction ability of the large language model. Specifically, at the speech understanding level, a fine-grained pre-trained speech and audio processing model Whisper model is used as an input module, and the encoder thereof is innovatively used as a speech coding structure, and the decoder combined with the length of 200 query tokens is used to construct a cross-modal feature alignment architecture QFormer projection architecture based on query vectors and attention mechanism. A variable-length coding strategy of 200 speech tokens (cross-modal feature unit) per 30 seconds is adopted to balance information acquisition and coding efficiency, and modal mapping from speech to text is realized. At the same time, a speech generation module composed of 24 attention mechanism-based deep learning network architectures Transformer layers containing cross-attention layers is designed, and a pre-trained large language model Qwen model is used for initialization to give it the ability of text understanding priori knowledge, and the processing of text tokens (the smallest semantic carrying unit) to speech tokens is completed. In the generation stage, text and speech generation are synchronized at a rate of 40 speech tokens per second, and the speech synthesis duration is dynamically controlled according to the text length. Through multi-stage training, the ability of the speech synthesis module is first strengthened by using large-scale text-speech pair data, and then the speech and audio understanding ability of the model is activated by constructing data and alignment modules to standardize the form of speech output.

[0068] The scheme enables the model to have strong long speech understanding and diversified generation capability, which can efficiently capture fine-grained speech information and process super-long speech content and accurately respond. The synchronous generation mechanism ensures the consistency of speech and text, and the multi-stage training strategy endows the model with robust speech interaction performance. Finally, the model significantly improves the application performance in complex life scenarios such as emotion recognition and empathetic dialogue based on the preservation of the text capability of large language models, realizes natural and fluent speech dialogue with variable length, and can quickly adapt to different multi-modal large model frameworks, maximizing the preservation of the core dialogue capability of the original model.

[0069] The embodiment provides a multi-modal speech interaction large model training method based on voice acoustic feature regulation, as shown in the formula (I), and specifically includes the following steps: Figure 1

[0070] Step S100, obtaining a text token of a text training sample, and based on the text token, constructing a voice token containing semantic features and voice acoustic features expressed by the text token to obtain pre-training data for converting the text token into the voice token.

[0071] In the embodiment, the pre-training data is constructed with the goal of high-quality speech generation. Specifically, through the design of a professional speech text dataset, the cooperation of text semantics, prompt word information and voice features is realized, and after the model outputs the text token, it can be accurately recognized and processed by the text-to-speech module, and the voice token with high semantic and prosody matching can be generated, providing training data support for subsequent basic model pre-training and multi-modal speech interaction large model training. In order to achieve the above goal, the construction of the dataset in the embodiment follows three core principles.

[0072] The first is the natural dialogue paradigm adaptation principle. The training data focuses on daily speech interaction scenarios, only retains text content conforming to human oral communication logic, strictly eliminates non-dialogue elements such as mathematical symbols, code fragments, tables, tree diagrams, etc., avoids special formats or non-communication information interference with the semantic coherence and prosodic naturalness of speech synthesis, and ensures the purity and interaction adaptability of text expression.

[0073] The second is the single language purity guarantee principle. The training data adopts a single language data construction mode, and establishes independent English and Chinese datasets to avoid the loss of speech synthesis quality caused by pronunciation rule differences and prosodic rhythm conflicts when multiple languages are mixed, and to avoid the semantic mapping deviation problem of different languages, thereby improving the accuracy and naturalness of speech synthesis in single language scenarios.

[0074] ​Thirdly, the multi-dimensional data coverage principle. The training data set covers multiple fields, multiple scenarios, and diversified text lengths. Specifically, the multiple fields are, for example, daily casual conversation, work communication, education tutoring, life service, etc., the multiple scenarios include, for example, family conversation, workplace meeting, customer service interaction, social casual conversation, etc., and the diversified text lengths include, for example, short sentence question and answer, long text story, explanatory content, etc. Through the rich data distribution, the adaptation ability of the model to the voice generation demand in different scenarios is enhanced, the robustness and generalization of the model training are improved, and the functional bias of the model caused by single data is avoided.

[0075] In an implementation manner, the text token of the text training sample is obtained, and based on the text token, a voice token containing semantic features and voice acoustic features expressed by the text token is constructed, to obtain pre-training data for converting the text token into the voice token, and the pre-training data specifically includes the following steps:

[0076] In step S110, single voice training initial data is obtained, and the single voice training initial data includes voice format dialogue pure text data and long text data set, wherein the voice format dialogue pure text data includes a filtered general natural dialogue text data set and a reconstructed dialogue text data set, and the reconstructed dialogue text data set is obtained by fine-tuning data of pure text instructions based on a preset prompt word through a general large model.

[0077] In step S120, special characters in reply round text content of the voice format dialogue pure text data and special characters in the long text data set are filtered based on a preset rule, and symbol division method is used for segmentation, to obtain preprocessed single voice training initial data.

[0078] In step S130, token conversion is performed on the single voice training initial data, to obtain the text token, and semantic features expressed by the text token are extracted.

[0079] In step S140, voice synthesis is performed on the single voice training initial data, to generate matched single voice speech samples, and voice acoustic features are extracted from the single voice speech samples.

[0080] In step S150, the text token, the semantic features expressed by the text token, and the voice acoustic features are associated, to obtain the voice token.

[0081] In this embodiment, first, single voice training initial data is constructed for single voice training. Figure 2The construction process of the pre-training data is shown. First, the pre-training initial data is divided into four categories, and four categories of natural dialogue text data sets are selected as single voice training initial data. The first is the MuTual data set, which is a two-person multi-round dialogue data set based on the high school English listening question. In order to simulate the open field dialogue more naturally, MuTual further converts the additional questions in the listening question into replies in the dialogue, wherein the turn of the model reply is used for speech synthesis pre-training. The second is the UltraChat data set after length cleaning, which contains 300K multi-round dialogue data, wherein the turn of the model output is used for speech synthesis pre-training data, so that the subsequent training data is aligned with the text-to-speech data. The third is the TinyStories English data set, which is a synthesized data set of short stories generated by large models GPT-3.5 and GPT-4, and the TinyStories-zh Chinese story data set generated by the large model qwen, which only contains long text content, i.e. long text data set, all text data is used for speech synthesis training. The fourth is the large model GPT-4 enhanced data set, which includes a data set simulating daily dialogue generated directly using the large model GPT-4. This part of data is constructed by adjusting the prompt word and calling the interface of the large model GPT-4 to construct natural dialogue, which covers multiple aspects of daily dialogue. Therefore, based on the specific needs of the generated data set, the prompt word with clear dialogue generation logic and quality constraints is constructed. Specifically, the prompt word mainly contains four types of key information. The first type of information is the specific definition of the task, instructing the large model to generate multi-round dialogue that meets the realistic scenario, requiring reasonable content, situational logic, and concise and meaningful statements from participants; the second type of information is the limitation condition of the dialogue type and topic, providing greetings, casual conversation, information exchange, discussion, emotional support, decision-making, and other dialogue types, as well as weather, work, family, health, travel, and technology, and other daily topics, and the prompt word will randomly select one or two dialogue types and a specific topic as the basis for generation; the third type of information is the generation rule constraint, which clearly states that the dialogue only contains two unnamed participants, the total number of turns is even and the number of statements from both parties is equal, and repetition or misuse of simple agreement words is prohibited; the fourth type of information is the output format specification, which requires the large model to return the result in JSON structure, including dialogue type, dialogue topic, and dialogue content three fields, and the dialogue content field needs to be labeled with speaker and corresponding statement one by one. Through the combination of the above prompt word elements, the scene adaptability and data specification of the generated dialogue can be controlled, providing a text basis for subsequent speech synthesis pre-training. In practical applications, Chinese or English prompt words can be selected according to the language preference of the training data, and the specific expression details of the prompt word can be adjusted according to the scene emphasis.

[0082] In addition, some public general fine-tuning datasets used by existing text large language models are used for the construction of speech dialogue data, including filtered moss, openhermes, alpaca, and GPTeacher English fine-tuning datasets, and part of the Chinese instruction fine-tuning dataset of belle. Among them, moss is an open source Chinese multi-turn dialogue dataset with more than 1.1 million data, openhermes is an English instruction fine-tuning dataset with more than 1 million data, alpaca is a cleaned English instruction pair data with more than 50,000 data, GPTeacher is a modular English dataset generated by GPT-4, and belle is a Chinese large-scale instruction dataset with more than 3.5 million data. The input and output of this part of data respectively indicate the data reconstruction of the large model interface, the purpose of which is to construct text data with high quality and in line with the mode of daily dialogue. The prompt words used in the corresponding constructed data aim to optimize the text-speech adaptability, and the purpose is to rewrite the original speech interaction instruction data into a speech format text suitable for the training of speech input-output large models. The prompt words mainly include three types of key restrictions. The first type of restriction makes the generated data colloquial, that is, makes the text close to the mode of real human speech. Specifically, filler words can be appropriately inserted to enhance naturalness, but excessive use of redundant expressions should be avoided; the second type of restriction is the synthesis compatibility constraint, which clearly states that the text must not contain content that cannot be processed by the text-to-speech model. For numerical information, it needs to be presented in English word spelling form rather than Arabic numerals to ensure the feasibility of subsequent speech synthesis; the third type of restriction is the brevity specification, which requires the rewritten text to be concise and avoid lengthy expressions, which is in line with the characteristics of concise communication in daily dialogue. Through the above prompt word elements, the original fine-tuning data is reconstructed in a targeted manner, which can improve the colloquial degree and speech synthesis adaptability of the text data, making it more in line with the actual needs of the speech interaction scene, and providing text basis that meets the quality standards for subsequent speech token construction and model pre-training. Similarly, in practical applications, Chinese or English prompt words can be selected according to the language preference of the training data, and the specific expression details of the prompt words can be adjusted according to the scene.

[0083] The number of turns in which the model replies to the text data set is used for speech synthesis, and the configuration of such text-speech pairs is input into the model for pre-training.

[0084] To further improve the data quality, a multi-layer screening mechanism is specifically designed, that is, the text containing special characters is removed through rule filtering, and for the long text data in UltraChat and TinyStories datasets, symbol division method is used for segmentation. The specific operation is based on fixed length, and priority is given to ensuring that the segmented text ends with punctuation marks. If the text is too long and has no punctuation, it is forced to be segmented and supplemented with punctuation. Finally, based on the segmented data, speech synthesis is realized by calling the Microsoft Edge text-to-speech interface, and three types of single voice models are constructed, including multi-language male voice, female voice adapted to English synthesis, and female voice adapted to Chinese synthesis. The synthesis process sets a retry mechanism, and the failed cases are processed in a loop until the maximum number of attempts is reached or the synthesis is successful, ensuring the integrity and availability of the data. Finally, the characteristic synthesis prompt of such single voice training data adopts a structured parameter configuration form, which contains two types of key information. The first type of key information is the language type parameter, which can be selected from English and Chinese, and is used to match the language scene of the training data. The second type of key information is the speaker parameter, which provides three types of preset voice options, including multi-language male voice, female voice adapted to English synthesis, and female voice adapted to Chinese synthesis, which is used to control the timbre characteristics of the speech sample. Through this standardized prompt format, the speech synthesis configuration can be defined, and the generated speech sample is aligned with the text data and the preset acoustic feature requirements, providing a standardized speech material basis for the construction of subsequent speech tokens.

[0085] In an implementation manner, the text token of the text training sample is obtained, and based on the text token, a speech token containing semantic features and speech acoustic features expressed by the text token is constructed, to obtain pre-training data for converting the text token into the speech token, and the method further includes the following steps:

[0086] Step S160, obtaining multi-voice training initial data, the multi-voice training initial data including dialogue text, corresponding multi-voice speech data, and corresponding speech acoustic feature annotation;

[0087] Step S170, performing data cleaning on the multi-voice training initial data, and performing token conversion to obtain the text token and extract semantic features expressed by the text token;

[0088] Step S180, associating the text token, the semantic features expressed by the text token, and the speech acoustic feature annotation to obtain the speech token.

[0089] In the embodiment, as Figure 2Part of the multi-voice training of the multi-voice emotional life dataset, that is, the multi-voice training, selects some open source and high-quality datasets and applies them to the voice synthesis pre-training process of the model. One of them is the Vccm dataset, which is a voice synthesis dataset optimized based on the TextrolSpeech dataset. Among them, TextrolSpeech contains 330 hours of voice data and 236203 style description texts. On this basis, the Vccm dataset optimizes the pitch distribution, label boundary, and dataset division, and reselects the test set. Its construction is based on the LibriTTS and TextrolSpeech emotional data sets. Five attribute labels of gender, volume, speed, pitch, and emotion are labeled for each voice sample. Through multi-dimensional data labeling and structured processing, a high-quality text-to-speech dataset with fine-grained acoustic features is formed. Among them, LibriTTS is a male and female English voice dataset with more than 500 hours of coverage and more than 2,000 speakers. The second is the Speechcraft dataset, which is a large-scale emotional rich bilingual voice dataset. It generates natural language descriptions with an automatic speech annotation system, covering more than 2 million audio segments, and is equipped with two text prompt versions for labeling. The first text prompt version is a voice description without voice transcription content and focus on voice features, and the second text prompt version is a voice instruction containing voice transcription text and instructional information. It provides rich training resources for bilingual speech understanding, emotion recognition, and cross-modal research, and is suitable for developing voice processing models closer to real scenarios.

[0090] The dataset in the multi-voice training data part does not need to be constructed, but it needs to be cleaned for the text content. Since the dataset still contains many special characters, it will have a great impact on the voice-to-text paradigm, and the format of some datasets is all uppercase and uses words to represent punctuation marks. It is necessary to reconstruct the content according to the rules to finally obtain the corresponding text content that meets the speaking mode.

[0091] Finally, as shown in Figure 2 , by constructing the pre-training initial data containing single-voice training data and multi-voice training data, the pre-training synthesis data is finally obtained.

[0092] Step S200, based on the multi-modal input sample and the pre-training data, construct fine-tuning training data for voice understanding and dialogue generation, wherein the fine-tuning training data contains voice semantic understanding labeling and dialogue generation logic labeling.

[0093] In this embodiment, in order to enhance the understanding ability of the model to the user voice, a batch of accurate and efficient fine-tuning data needs to be constructed for training the speech understanding and standardized generation ability of the multi-modal large model. The data to be constructed should enable the model to have strong audio understanding ability while effectively utilizing context information for reasoning and outputting text content matching the corresponding audio reply. This part of data involves speech understanding alignment and integration of various capabilities of the multi-modal large model, and requires multiple types of data.

[0094] In an implementation manner, the fine-tuning training data includes speech understanding capability alignment data, multi-round dialogue fine-tuning data of voice interaction, text-to-speech generation polishing fine-tuning data, and pre-training capability preservation data set, the fine-tuning training data for speech understanding and dialogue generation is constructed based on the multi-modal input sample and the pre-training data, wherein the fine-tuning training data contains speech semantic understanding annotation and dialogue generation logic annotation, and specifically includes the following steps:

[0095] In step S210, the speech-to-text sample data is obtained, the speech semantic annotation is performed on the speech-to-text sample data, and the speech understanding capability alignment data is constructed based on the speech-to-text sample data and the corresponding speech semantic annotation, wherein the speech understanding capability alignment data is used for learning the speech understanding capability of the model.

[0096] In step S220, the context association logic of the multi-round dialogue data is obtained based on the multi-round dialogue data in the pre-training data, and the fine-tuning data of the multi-round dialogue of voice interaction is constructed based on the multi-round dialogue data and the corresponding context association logic, wherein the multi-round dialogue of voice interaction includes dialogue generation logic annotation, and the fine-tuning data of the multi-round dialogue of voice interaction is used for learning the context association logic of the model.

[0097] In step S230, the spoken language voice output sample data is obtained, and the text-to-speech generation polishing fine-tuning data is constructed, wherein the text-to-speech generation polishing fine-tuning data is used for learning the spoken language of the model.

[0098] In step S240, the multi-modal input sample is obtained, and the pre-training capability preservation data set is constructed, wherein the multi-modal input sample at least includes visual input sample, text input sample and voice input sample, and the pre-training capability preservation data set is used for learning the multi-modal recognition capability of the model.

[0099] In this embodiment, Figure 3The construction process of fine-tuning data is demonstrated. The instruction fine-tuning initial data includes five parts, including pre-training stories, multi-voice data, speech-to-text data, pre-training instruction fine-tuning data, joint capability fine-tuning data, and text, multi-modal capability maintenance data. After processing these five parts, the instruction fine-tuning synthetic data is constructed. After the instruction training of the fine-tuning data set, the model can not only retain the original visual multi-modal dialogue capability to a large extent, but also achieve good speech recognition, speech understanding and speech paradigm generation capability under the condition of less fine-tuning data. Through data organization and training, the model can support high-quality long speech dialogue and realize controllable emotional tone dialogue generation.

[0100] Specifically, the above-mentioned five parts of data can be optimized and summarized into four parts of data, including speech understanding capability alignment data, speech interaction multi-round dialogue fine-tuning data, text speech generation polishing fine-tuning data, and pre-training capability maintenance data set, which are used to construct fine-tuning training data.

[0101] The first part is the speech understanding capability alignment data. This part of data contains the model's recognition of speech content related data, so that the model can effectively understand the speech information and transcribe these contents into text form for output. These data sets are mainly selected from CommonVoice data set and GigaSpeech data set. CommonVoice data set is an open source speech data set of Mozilla, which collects more than 100,000 hours of speech. It mainly uses English data, including audio, text transcription, etc., which is used for open source ASR (Automatic Speech Recognition) training and language diversity research. GigaSpeech data set is a large-scale speech data set of Microsoft, which covers more than 1,000 hours of English speech, covering conference, dialogue and other scenes, with fine annotation, which is used for complex scene ASR model training. The prompt word corresponding to this part of data is "convert speech content to text content" instruction in Chinese and English.

[0102] The second part is the fine-tuning data of multi-round dialogue of speech interaction. This part of data contains the multi-round dialogue data collected and constructed in the pre-training stage. Except for the TinyStories data set, all other data can be used in this stage. This data set will be converted into diversified data through additional data enhancement process. Mutual data set and data set constructed and polished by large model GPT interface. Because the text content of these data sets is more in line with the way of spoken language expression, in order to make the model adapt to diversified input and ensure that only the speech output content is limited to spoken language without changing the original text output form of the model, only this part of data set is used to construct speech output data, and text output data is not constructed. Therefore, these data sets will be constructed from two aspects.

[0103] The first type of data is text input audio output type data, which will generate corresponding voice replies through user text instructions. The corresponding prompt word is "generate voice replies with specified voice characteristics according to text content" in Chinese and English.

[0104] The second type of data is voice input voice output type data. The model aligns voice and text information through a voice understanding encoder, effectively fusing the context content of voice features, and then obtains the corresponding voice reply. The corresponding prompt word is "listen to voice content and generate corresponding voice reply with specified voice characteristics" in Chinese and English.

[0105] The third part is text voice generation polishing fine-tuning data. This part of data is only processed into text input voice output form, which is used to make the model's voice output corresponding text more colloquial, and also to enhance the universality and accuracy of voice colloquial output content, to ensure that the model maintains the output paradigm when outputting voice, and does not generate special characters and long and difficult sentences that do not conform to voice conversation, so as to ensure the accuracy of output segmentation and the robustness of the final voice synthesis module. The organization form of this part of data is consistent with the format and prompt word of the second part of data. The data comes from the VoiceAssistant dataset constructed by the mini-omni framework, which contains more than 400,000 pieces of cleaned daily conversation data.

[0106] The fourth part is a pre-training ability maintenance dataset. This part of data is used to maintain the ability related to the pre-training model. This part of input integrates pre-training data in various sub-datasets to ensure the generation result. The specific sub-datasets are as follows.

[0107] The first sub-dataset is a text-to-speech ability maintenance dataset. This part of data is used to maintain the text-to-speech ability of the model. It is mainly extracted from the pre-training data and converted into a dialogue form. The corresponding prompt word is "generate corresponding voice form according to text content and specified voice or speech parameters" in Chinese and English.

[0108] The second sub-dataset is a voice repetition dataset, which is used to recognize voice content and output it in voice to enhance the most intuitive ability of the model in understanding and generation. As above, the dataset is selected from the pre-training dataset, and the prompt word is "repeat voice content and generate voice of the same content according to specified voice or speech parameters" in Chinese and English.

[0109] The third sub-data is a text capability and visual capability maintenance data set, and the function of the data set is to ensure that the model retains good multi-modal visual capability and text dialogue capability, and the data set includes multiple text instruction data sets and image-text instruction question and answer data sets.

[0110] The fourth sub-data set is a multi-modal emotion recognition data set, and the mainly used data set is an emova-sft-speech-231k data set. The data set is used for training the emotion generation capability of the model, so that the model can efficiently generate voice content with emotion. The corresponding prompt word is the Chinese and English "listen to the voice content, and generate the corresponding voice reply with the specified emotion, timbre and pitch" instruction.

[0111] In step S300, a base model is constructed and pre-trained based on the pre-training data, wherein the base model is used to convert an input text token into a voice token.

[0112] In the embodiment, the pre-training is divided into multiple stages. In the first stage, single timbre data is used to activate the model's bilingual multi-language configuration generation capability to achieve realistic single timbre voice generation. At this time, the model has a relatively perfect text-to-speech capability, but the understanding capability of the self-recurrence module can further process more types of audio data. Therefore, in the second stage, female timbre data and multiple characteristic timbre data generated based on text instruction descriptions are further introduced, including discrete natural language descriptions for emotion (such as emotion: …, pitch: …) and natural language timbre description data. This not only enables the model to have characteristic voice generation capability, but also further enhances the robustness of single timbre synthesis through different timbres.

[0113] The training data organization mode is as shown in the upper half of FIG. 1. Figure 4 As shown in the upper half, it is divided into three parts by special symbols: the first part is a text prompt word, which describes the characteristics of the voice content; the second part is a voice prefix, which contains the timbre information (voice token) of the previous sentence, used to ensure the consistency of the timbre; and the third part is the text to be transcribed, which is the main context information. The model will convert it from text symbols to voice symbols through self-recurrence generation.

[0114] In one implementation mode, the base model is constructed and pre-trained based on the pre-training data, and specifically includes the following steps:

[0115] In step S310, the text token in the pre-training data, the voice token in the foregoing, and the preset prompt word are input into a voice self-recurrence module of the pre-constructed base model to obtain a predicted voice token, wherein the voice self-recurrence module includes 24 Transformer blocks.

[0116] Step S320, output the predicted speech token to a 4096-dimensional prediction head to obtain predicted synthesized speech.

[0117] Step S330, iteratively calculate the loss function, and back-propagate to optimize the parameters of the base model.

[0118] In the first phase of training, the specific goal is to train a text token generation speech token model with basic text-to-speech capability. Specifically, the word embedding of the text is obtained using the vocabulary mapping of the large model Qwen, which contains sufficient text information. After the projection layer, the text information is input into the speech autoregressive module initialized by the text model. The input of the module includes the text content to be converted into speech, the prompt word content indicating the speech conversion style, and the speech token input side by side with the prompt word, which is used to ensure the consistency of the subsequent voice timbre. After the input data is processed by the 24-layer Transformer block of the autoregressive module, the model will predict the subsequent speech token according to the previous text information and output it to the 4096-dimensional prediction head. Since the speech codec Wavtokenizer used can encode and decode more complex acoustic information, combined with the autoregressive module with stronger understanding ability, the model can effectively understand and generate various paralinguistic information of speech, thereby realizing more controllable speech generation.

[0119] In one implementation, when the base model generates speech replies, it generates synthesized speech of any length through block synthesis, specifically including the following steps:

[0120] Step S340, divide the input text corresponding to the speech to be generated into several text blocks with complete semantics by taking the period as the only truncation position;

[0121] Step S350, input the text token corresponding to the first text block into the pre-trained base model in the order of the text blocks, and generate the speech token corresponding to the first text block and the synthesized speech segment through the speech autoregressive module;

[0122] Step S360, input the speech token corresponding to the text block as speech information prompts, and input the text token of the next text block into the base model to generate the corresponding speech token and synthesized speech segment;

[0123] Step S370, generate the synthesized speech segment corresponding to all text blocks in the order of the text blocks, and splice all synthesized speech segments to obtain synthesized speech of any length.

[0124] In this embodiment, to solve the problem of limited speech length generation existing in open source large language model frameworks such as LLaMA-Omni and Freeze-Omni, a method of dividing text into text blocks of similar length for speech synthesis is adopted, and the previous speech information is used to prompt the subsequent generation to ensure the coherence of the speech. Specifically, only the period is used as the truncation position of the text block. This way enables the model to generate speech replies of any length without being limited by the model architecture, while ensuring the accuracy and robustness of long speech at any position.

[0125] Step S400, based on the pre-trained base model, a multi-modal speech interaction large model is constructed, and the multi-modal speech interaction large model is trained using the fine-tuning training data, wherein the multi-modal speech interaction large model is used to regulate the speech acoustic features for generating speech answers based on multi-modal input, and output the speech answers.

[0126] In this embodiment, speech understanding is a core component of the speech dialogue model, and the completeness of its function directly affects the depth of the model's analysis of user intent. Current mainstream models usually focus on extracting and compiling semantic information when processing speech input, but pay relatively little attention to prosodic features such as intonation and timbre. Based on the QwenAudio architecture of the audio language large model, a unified audio information processing framework is constructed by enhancing the representation ability of the encoder, so that the model can capture semantic and prosodic features at the same time, and realize multi-dimensional perception of speech input.

[0127] Specifically, the encoder of Whisper-large-v3 is used as the speech encoder to construct the speech understanding alignment module of the multi-modal large model. Experiments show that by modifying the decoder part of Whisper to Qformer and using a mapping strategy of 400 tokens per minute, the model can efficiently compress speech information while retaining rich paralinguistic features such as timbre, intonation changes, and emotional cues. These features, after being mapped to the text semantic space by the projection layer, can be effectively utilized by the model, significantly improving the multi-modal reasoning ability. The training process is as follows: Figure 3 The lower part represents.

[0128] In one implementation, the multi-modal speech interaction large model is constructed based on the pre-trained base model, and the multi-modal speech interaction large model is trained using the fine-tuning training data, specifically including the following steps:

[0129] Step S410, using the Qformer architecture to modify the Whisper decoder to construct a multi-modal speech interaction large model, the structure of the multi-modal speech interaction large model is expressed as:

[0130]

[0131]

[0132]

[0133]

[0134]

[0135] wherein, is a fixed-length query vector, is a vector, is a hidden layer output, is an audio, is a query query, is the i-th hidden layer output component constituting the fixed-length query vector , is the top output of the pre-trained audio encoder Whisper encoder, is a multi-head self-attention layer, is a multi-head cross-attention layer, is a layer normalization layer, is a multi-layer perceptron, is the corresponding output content of the Whisper encoder, is the corresponding output content of , is the corresponding output content of , is a multi-head self-attention calculation on the fixed-length query vector , represents the total number of initial vectors , represents the output of the cross-attention module, which is used to extract the main content of the input audio, is a hidden layer output after multi-layer perceptron processing on the cross-attention module output , after four layers of the same operation, a learnable linear layer is applied to project the final output to the representation space of the large language model;

[0136] Step S420, using the fine-tuning training data, training the multi-modal voice interaction large model.

[0137] In this embodiment, in this architecture, the Whisper encoder encodes every 30 seconds of audio into audio tags, and the output of the cross-attention module is responsible for extracting the core content of the input audio. After passing through all the layers of the Whisper decoder, the final output is projected to the representation space of the multi-modal large language model through a learnable linear layer.

[0138] Compared with other speech understanding alignment methods, the architecture scheme has three advantages. First, the design of the Qformer architecture is advanced. Second, the pre-trained Whisper architecture is used to initialize the encoder, which can efficiently extract audio information. Finally, the pre-trained Whisper decoder is used to initialize the Qformer, which can significantly improve the speed and accuracy of the model's speech understanding ability alignment.

[0139] In the design of model output, the normalization of speech corresponding text is realized by constructing a specific data set and introducing special markers. Specifically, a fine-tuning data set is constructed to activate the model's speech normalization output ability in different scenarios. At the same time, multiple special tokens are introduced to separate different types of output information. <sosp>As a voice start mark, the content after it is voice control information. The monophonic voice control information includes specific language and speaker. In actual data construction, three kinds of human voices are mainly used for diversified control, including multi-language male voice, female voice adapted to English synthesis, and female voice adapted to Chinese synthesis. The multi-voice and multi-tone control information uses the description after polishing by a large model. <eop>The end of the marked characterised voice description, followed by the actual text content to be converted to speech. <eosp>As an end marker of voice generation, the content outside the marker does not participate in voice conversion.

[0140] After training by the above fine-tuning data set, the model has four core capabilities, including basic multi-modal text-image interaction capability, end-to-end communication capability of voice input and output, mixed communication capability of flexible switching between voice and text, and emotion recognition and specific timbre accurate control generation capability.

[0141] Therefore, the innovative text-voice token alignment mechanism provided by the embodiment realizes efficient processing of text-to-speech synthesis training by constructing a dynamic mapping algorithm and an adaptive coding framework. This method breaks through the length and timbre limitations of traditional synthesis technology, not only accurately processes text input of any length to generate semantically coherent and natural prosodic voice sequences, but also supports diversified timbre, emotion and style characteristic voice synthesis. On the basis of ensuring high fidelity of single timbre synthesis, by introducing a conditional generation model and a cross-modal attention mechanism, fine control of voice acoustic features is realized, effectively improving the flexibility and expressiveness of characteristic audio generation.

[0142] At the same time, in view of the resource bottleneck problem of multi-modal large models in voice interaction capability migration, the embodiment provides a low-resource-consumption fast alignment method and a matching data set. By deeply mining the semantic association between multi-modal data, combining contrastive learning and transfer learning technology, a high-quality fine-tuning data set containing text, voice and emotion labels is constructed. This scheme can quickly activate the voice understanding and generation potential of multi-modal large models, and on the premise of greatly reducing computing resources and training time, it gives the model voice dialogue capability with emotion perception and expression, significantly improving the adaptability and versatility of the model in multi-modal interaction scenarios.

[0143] In addition, with more abundant and comprehensive data sets and more sufficient resources, the understanding and generation of the model can be further expanded, expanding the single voice audio modality to environmental sound, music and medical sound waves and other more extensive aspects, so that the model can have more extensive and comprehensive audio capabilities.

[0144] On the basis of the above-mentioned embodiments, there are also ineffective variants. Specifically, the two training stages are combined into a single-stage training strategy, but this way has double bottlenecks of resource utilization efficiency and model adaptability. This way directly discards the parameters and feature representations accumulated by the first stage of training, fails to effectively reuse the training results of the early stage, and will lead to repeated consumption of computing resources and significant reduction of training efficiency. In addition, when the base model is updated and iterated, due to the lack of modular training structure, the entire model needs to be trained from scratch, greatly increasing the cost of model migration and optimization, and seriously restricting the scalability and rapid adaptability of the model.

[0145] Another ineffective transformation scheme, specifically using a network-level full end-to-end alignment training mode, although intended to realize direct mapping and collaborative optimization of multi-modal information, has uncertain performance improvement. This scheme trains the model as a whole in the final alignment stage, which may cause conflicts in gradient updates and imbalance in parameter adjustments, making it difficult for the model to accurately capture key semantic associations between modalities during alignment. Compared to the direct mapping method based on token generation, this global training method may weaken the alignment ability of the model due to excessive coupling, making it difficult to fully leverage the advantages of each modality, and thus affecting the performance and generalization ability of the model in the speech dialogue task.

[0146] In summary, under the technical scheme of the above embodiments, the present application focuses on the performance optimization and multi-modal adaptation of the speech dialogue system, and achieves breakthroughs in multiple dimensions through innovative technical solutions. In the aspect of long text output, the present application proposes an efficient block synthesis strategy, effectively overcoming the problems of information loss and lack of coherence in traditional speech dialogue models when processing long texts, and enabling the generation of long text content with clear logic and complete semantics. In the field of audio synthesis, the present application constructs a fine-grained emotion and acoustic parameter control model. Based on deep neural networks, this model can accurately analyze users' diverse needs for emotion, timbre, pitch, and speech rate, and through advanced audio synthesis technology, it can map these parameters to high-quality speech output, achieving personalized and immersive voice interaction experience. In the aspect of multi-modal model adaptation, the present application constructs a set of efficient fine-tuning data, enabling multi-modal large models based on different frameworks to quickly integrate into the speech dialogue architecture proposed by the present application, ensuring efficient adaptation while maximizing the preservation of the core capabilities of the original model in text and multi-modal dialogue tasks, significantly improving the versatility and expandability of the system.

[0147] As shown in Figure 5 The embodiment of the present application provides a multi-modal speech interaction large model training system based on voice acoustic feature regulation, which comprises a pre-training data construction module 10, a fine-tuning training data construction module 20, a basic model construction and pre-training module 30, and a multi-modal speech interaction large model construction and training module 40.

[0148] Specifically, the pre-training data construction module 10 is configured to acquire text tokens of a text training sample, and construct speech tokens containing semantic features and speech acoustic features expressed by the text tokens based on the text tokens, to obtain pre-training data for converting the text tokens into the speech tokens; the fine-tuning training data construction module 20 is configured to construct fine-tuning training data for speech understanding and dialogue generation based on the pre-training data and a multi-modal input sample, wherein the fine-tuning training data contains speech semantic understanding labels and dialogue generation logical labels; the base model construction and pre-training module 30 is configured to construct a base model and pre-train the base model based on the pre-training data, wherein the base model is used to convert input text tokens into speech tokens; and the multi-modal speech interaction large model construction and training module 40 is configured to construct a multi-modal speech interaction large model based on the pre-trained base model, and train the multi-modal speech interaction large model using the fine-tuning training data, wherein the multi-modal speech interaction large model is used to regulate speech acoustic features for generating a speech answer based on a multi-modal input, and output the speech answer.

[0149] Based on the above-mentioned embodiments, the application further provides a terminal device, a principle block diagram of which can be shown in Figure 6 The terminal device includes a processor, a memory, a network interface, a display screen, and a temperature sensor connected through a system bus. The processor of the terminal device is configured to provide computing and control capabilities. The memory of the terminal device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the terminal device is configured to communicate with external terminals through network connection. The computer program is executed by the processor to implement a multi-modal speech interaction large model training method based on speech acoustic feature regulation. The display screen of the terminal device can be a liquid crystal display screen or an electronic ink display screen. The temperature sensor of the terminal device is pre-installed in the terminal device and is used to detect the running temperature of the internal device.

[0150] Those skilled in the art can understand that Figure 6 The principle block diagram shown in the above-mentioned embodiments is only a block diagram of part of the structure related to the application scheme, and does not constitute a limitation on the terminal device to which the application scheme is applied. Specifically, the terminal device can include more or fewer components than those shown in the diagram, or combine certain components, or have a different component arrangement.

[0151] In one embodiment, a terminal device is provided, comprising a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by one or more processors, the one or more programs comprising instructions for:

[0152] obtaining a text token of a text training sample, and constructing a speech token comprising a semantic feature and a speech acoustic feature expressed by the text token based on the text token, to obtain pre-training data for converting the text token into the speech token;

[0153] constructing fine-tuning training data for speech understanding and dialogue generation based on the multi-modal input sample and the pre-training data, wherein the fine-tuning training data comprises speech semantic understanding annotation and dialogue generation logic annotation;

[0154] constructing a base model and pre-training the base model based on the pre-training data, wherein the base model is used to convert an input text token into a speech token;

[0155] constructing a multi-modal speech interaction large model based on the pre-trained base model, and training the multi-modal speech interaction large model using the fine-tuning training data, wherein the multi-modal speech interaction large model is used to regulate a speech acoustic feature for generating a speech answer based on a multi-modal input, and output the speech answer.

[0156] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the computer program can include the processes of the above-mentioned embodiments of each method. Any reference to memory, storage, database or other medium used in each embodiment provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0157] In summary, the application discloses a multi-modal speech interaction large model training method and system based on speech acoustic feature regulation, a terminal device and a medium, relates to the technical field of multi-modal speech interaction, and comprises the following steps: first, obtaining a text token of a text training sample, and based on the text token, constructing a speech token containing semantic features and speech acoustic features expressed by the text token to obtain pre-training data for converting the text token into the speech token. Then, based on multi-modal input samples and the pre-training data, constructing fine-tuning training data for speech understanding and dialogue generation, wherein the fine-tuning training data contains speech semantic understanding annotations and dialogue generation logic annotations. Next, based on the pre-training data, constructing a basic model and pre-training the basic model, wherein the basic model is used to convert an input text token into a speech token. Finally, based on the pre-trained basic model, constructing a multi-modal speech interaction large model, and using the fine-tuning training data to train the multi-modal speech interaction large model, wherein the multi-modal speech interaction large model is used to regulate speech acoustic features for generating a speech answer based on multi-modal input and output the speech answer. The application solves the problems of insufficient speech diversity of existing models, difficulty in long speech generation, and large resource consumption of multi-modal adaptation, realizes acoustic features such as tone and emotion, supports controllable long speech coherent generation, quickly activates the speech interaction capability of the model, while retaining the multi-modal core capability, and improves the speech interaction naturalness and generalization of the multi-modal large model.

[0158] The technical features of the above embodiments can be combined in any manner. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present application.

[0159] The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be noted that for ordinary skilled persons in the art, without departing from the concept of the present application, several modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.< / eosp> < / eop> < / sosp>

Claims

1. A multi-modal speech interaction large model training method based on voice acoustic feature regulation, characterized in that, The method comprises: acquiring text tokens of text training samples, and constructing speech tokens containing semantic features and speech acoustic features expressed by the text tokens based on the text tokens, to obtain pre-training data for converting text tokens into speech tokens; based on the pre-training data, constructing a base model and pre-training the base model, wherein the base model is used to convert input text tokens into speech tokens; based on the pre-trained base model, constructing a multi-modal speech interaction large model, and training the multi-modal speech interaction large model using the fine-tuning training data, wherein the multi-modal speech interaction large model is used to regulate speech acoustic features for generating speech answers based on multi-modal input, and output the speech answers; The method comprises: acquiring text tokens of text training samples, and constructing speech tokens containing semantic features and speech acoustic features expressed by the text tokens based on the text tokens, to obtain pre-training data for converting text tokens into speech tokens; acquiring single-person voice training initial data, wherein the single-person voice training initial data comprises speech format dialogue pure text data and long text data set, wherein the speech format dialogue pure text data comprises filtered general natural dialogue text data set and reconstructed dialogue text data set, wherein the reconstructed dialogue text data set is obtained by reconstructing pure text instruction fine-tuning data based on a preset prompt word through a general large model; based on a preset rule, filtering special characters in reply round text content of the speech format dialogue pure text data and special characters in the long text data set, and performing symbol division to obtain preprocessed single-person voice training initial data; performing token conversion on the single-person voice training initial data to obtain the text tokens, and extracting semantic features expressed by the text tokens; performing speech synthesis on the single-person voice training initial data to generate matched single-person voice samples, and extracting speech acoustic features from the single-person voice samples; associating the text tokens, the semantic features expressed by the text tokens, and the speech acoustic features to obtain the speech tokens.

2. The multi-modal voice interaction large model training method based on voice acoustic feature regulation according to claim 1, characterized in that, The method comprises: acquiring text tokens of text training samples, and constructing speech tokens containing semantic features and speech acoustic features expressed by the text tokens based on the text tokens, to obtain pre-training data for converting text tokens into speech tokens; acquiring multi-person voice training initial data, wherein the multi-person voice training initial data comprises dialogue text, corresponding multi-person voice data, and corresponding speech acoustic feature labels; The multi-voice training initial data is cleaned and tokenized to obtain the text token and extract semantic features expressed by the text token; The text token, the semantic features expressed by the text token, and the voice acoustic feature label are associated to obtain the voice token.

3. The multi-modal speech interaction large model training method based on voice acoustic feature regulation according to claim 2, characterized in that, The fine-tuning training data includes voice understanding ability alignment data, voice interaction multi-round dialogue fine-tuning data, text voice generation polishing fine-tuning data, and pre-training ability preservation data set. The fine-tuning training data is constructed based on the pre-training data and the multi-modal input sample for voice understanding and dialogue generation, wherein the fine-tuning training data contains voice semantic understanding label and dialogue generation logic label, including: Obtain voice-to-text sample data, perform voice semantic labeling on the voice-to-text sample data, and construct the voice understanding ability alignment data based on the voice-to-text sample data and the corresponding voice semantic labeling, wherein the voice understanding ability alignment data is used for learning the voice understanding ability of the model. Based on the multi-round dialogue data in the pre-training data, the context association logic of the multi-round dialogue data is obtained, and the fine-tuning data of the voice interaction multi-round dialogue is constructed based on the multi-round dialogue data and the corresponding context association logic, wherein the voice interaction multi-round dialogue includes dialogue generation logic label, and the fine-tuning data of the voice interaction multi-round dialogue is used for learning the context association logic of the model. Obtain spoken voice output sample data and construct the text voice generation polishing fine-tuning data, wherein the text voice generation polishing fine-tuning data is used for learning the spoken language of the model. Obtain multi-modal input samples and construct the pre-training ability preservation data set, wherein the multi-modal input samples at least include visual input samples, text input samples, and voice input samples, and the pre-training ability preservation data set is used for learning the multi-modal recognition ability of the model.

4. The multi-modal voice interaction large model training method based on voice acoustic feature regulation according to claim 1, characterized in that, The base model is constructed and pre-trained based on the pre-training data, including: Input the text token, the voice token in the above, and the preset prompt word in the pre-constructed voice autoregressive module of the base model based on the pre-training data to obtain the predicted voice token, wherein the voice autoregressive module includes 24 layers of Transformer blocks; Output the predicted voice token to a 4096-dimensional prediction head to obtain a predicted synthesized voice; Iteratively calculate the loss function and back-propagate to optimize the parameters of the base model.

5. The multi-modal speech interaction large model training method based on voice acoustic feature regulation according to claim 4, characterized in that, When the base model generates a voice reply, an arbitrary length of synthesized voice is generated by block synthesis, including: Divide the input text corresponding to the voice to be generated into several text blocks with complete semantics with a period as the only truncation position; According to the order of the text blocks, input the text token corresponding to the first text block into the pre-trained base model, and generate the voice token and synthesized voice segment corresponding to the first text block through the voice autoregressive module; The voice token corresponding to the text block is taken as voice information prompt, and the voice token and the text token of the next text block are input into the base model to generate corresponding voice token and synthesized voice segment; According to the order of the text blocks, the synthesized voice segments corresponding to all the text blocks are generated, and all the synthesized voice segments are spliced to obtain a synthesized voice of any length.

6. The multi-modal speech interaction large model training method based on voice acoustic feature regulation according to claim 1, characterized in that, The pre-trained base model is used to construct a multimodal voice interaction large model, and the multimodal voice interaction large model is trained using the fine-tuning training data, including: The Whisper decoder is modified using the Qformer architecture to construct the multimodal voice interaction large model, and the expression of the structure of the multimodal voice interaction large model is: wherein, is a fixed-length query vector, is a vector, is a hidden layer output, is an audio, is a query query, is the i-th hidden layer output component constituting the fixed-length query vector , is the top layer output of the pre-trained audio encoder Whisper encoder, is a multi-head self-attention layer, is a multi-head cross-attention layer, is a layer normalization layer, is a multi-layer perceptron, is the corresponding output content of the Whisper encoder, is the corresponding output content of , is the corresponding output content of , is the corresponding output content of is a multi-head self-attention calculation on the fixed-length query vector , wherein represents the total number of initial vectors, represents the output of the cross-attention module, which is used to extract the main content of the input audio, is a hidden layer output after multi-layer perceptron processing on the cross-attention module output , and after four layers of the same operation, a learnable linear layer is applied to project the final output to the representation space of the large language model. The multimodal voice interaction large model is trained using the fine-tuning training data.

7. A multi-modal speech interaction large model training system based on voice acoustic feature regulation, characterized in that, The system is applied to implement the steps of the multimodal voice interaction large model training method based on voice acoustic feature regulation according to any one of claims 1-6, and the system comprises: A pre-training data construction module is configured to obtain text tokens of text training samples, and construct voice tokens containing semantic features and voice acoustic features expressed by the text tokens based on the text tokens, to obtain pre-training data for converting text tokens into voice tokens; A fine-tuning training data construction module is configured to construct fine-tuning training data for voice understanding and dialogue generation based on multimodal input samples and the pre-training data, wherein the fine-tuning training data contains voice semantic understanding annotations and dialogue generation logic annotations; A base model construction and pre-training module is configured to construct a base model and pre-train the base model based on the pre-training data, wherein the base model is used to convert input text tokens into voice tokens; A multimodal voice interaction large model construction and training module is configured to construct a multimodal voice interaction large model based on the pre-trained base model, and train the multimodal voice interaction large model using the fine-tuning training data, wherein the multimodal voice interaction large model is used to regulate voice acoustic features for generating voice answers based on multimodal input, and output the voice answers.

8. A terminal device, comprising: The terminal device comprises a memory, a processor, and a multimodal voice interaction large model training program based on voice acoustic feature regulation stored in the memory and executable on the processor. When the processor executes the multimodal voice interaction large model training program based on voice acoustic feature regulation, the steps of the multimodal voice interaction large model training method based on voice acoustic feature regulation according to any one of claims 1-6 are implemented.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a multimodal voice interaction large model training program based on voice acoustic feature regulation, and when the multimodal voice interaction large model training program based on voice acoustic feature regulation is executed by the processor, the steps of the multimodal voice interaction large model training method based on voice acoustic feature regulation according to any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Language model training method of native voice mode

    CN118471202A

  • Multi-modal speech emotion recognition method and system based on pre-training model

    CN119339743A