Large model real-time voice interaction method and device based on autoregression voice synthesis

Through the combination of autoregressive streaming speech decoder and large language model, the contradiction between real-time and naturalness of modular speech models is solved, efficient and natural speech generation is achieved, and the response speed and user experience of the voice interaction system are improved.

CN120340472APending Publication Date: 2025-07-18INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510386689.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing modular voice model is difficult to ensure both real-time and naturalness during the speech generation process. The existing voice decoder design makes the generated voice prone to intermittent, unnatural or high delays.

Method used

The autoregressive streaming speech decoder is used to combine with a large language model, and the speech encoder, speech adapter, large language model and text-voice language model is used to generate speech marker sequences using the autoregressive Transformer structure, and the speech synthesis process is optimized through the causal flow matching model and the gating fusion mechanism.

Benefits of technology

It realizes high real-time and high natural speech synthesis, significantly improving the response speed and user experience of the voice interaction system, and significantly improving the fluency and quality of the generated voice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340472A_ABST
    Figure CN120340472A_ABST
Patent Text Reader

Abstract

The invention provides a large-model real-time voice interaction method and device based on autoregressive voice synthesis, and the method comprises the steps: obtaining a voice instruction marked with a target text response and a target voice response, enabling a voice encoder to code the voice instruction into voice representation, and enabling a voice adapter to carry out the dimension reduction and feature conversion of the original voice representation; the large language model generates a hidden state according to the converted voice representation and samples the hidden state to obtain a text sequence; and processing the text sequence by adopting a text-voice language model based on an autoregression Transform structure, generating a voice marking sequence in a streaming manner, and converting the voice marking sequence into a voice signal through a vocoder. According to the method provided by the invention, the naturalness and fluency of speech synthesis are greatly improved while high real-time performance is ensured. The optimized voice decoding architecture effectively reduces the voice generation delay and improves the response speed of the voice interaction system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer science and technology, large language models, speech large models, and natural speech understanding technology, and particularly relates to a large model real-time speech interaction method, device, electronic device, computer-readable storage medium, and computer program product based on autoregressive speech synthesis. Background Art

[0002] Currently, speech interaction systems based on large language models mainly fall into two technical routes: traditional cascaded systems and end-to-end speech large models. The traditional solution realizes the interaction process by serially connecting three independent modules: Automatic Speech Recognition (ASR), Large Language Model (LLM), and Text-To-Speech (TTS). Its technical maturity is relatively high, but it is limited by the inherent defects of the cascaded system. End-to-end speech large models mainly include two technical routes: native speech large models discretize the original speech signal into a symbol sequence through vector quantization or clustering methods, and use a Transformer model with only a decoder architecture to uniformly model the language model of speech and text; modular speech large models form an end-to-end system that can process continuous speech features by adding a speech encoder and a speech decoder before and after the large language model respectively.

[0003] Currently, speech interaction systems based on large language models still have many drawbacks. First of all, one of the biggest problems faced by traditional cascaded systems is the accumulation of errors. Since the three modules of ASR, LLM, and TTS work independently, the output of each module depends on the result of the previous module. This design makes the system have a low tolerance for errors in the previous module, and errors are easily amplified step by step in the pipeline, resulting in unstable voice quality of the final generated speech. In addition, each link in the processing flow of the cascaded system may increase the latency, and the overall response time is relatively long. Especially in application scenarios with high real-time interaction requirements, it may lead to a decline in the user experience. Due to the independent operation between modules, it is also difficult for cascaded systems to effectively capture paralinguistic information (such as tone, emotion, etc.) in speech input, which makes the generated speech often appear rigid and mechanical, lacking the natural feeling of human interaction.

[0004] In end-to-end speech large models, although native speech models overcome the drawbacks of cascaded systems to some extent through a unified architecture, they still face some key challenges. First, native end-to-end models usually require a large amount of speech data for pre-training to effectively model the relationship between speech and text. Although this provides the model with powerful learning capabilities, the dependence on large-scale speech data brings significant cost issues. The costs of data collection, processing, and storage are not only high, but the quality and diversity of the data also directly affect the performance of the model. Especially in applications in specific domains or minority languages, the scarcity of data may lead to the model being unable to effectively cover the specific needs of these domains. In addition, the training process of native end-to-end models may lead to catastrophic forgetting, that is, the model gradually loses its own text processing ability during the training process of speech data, resulting in a decline in the inference ability of the system.

[0005] In contrast, although modular end-to-end speech large models can reduce the demand for large-scale speech data by separately processing speech encoding and decoding tasks, thereby reducing training costs, the current technical solutions also have some limitations. To achieve low-latency speech interaction, existing technologies have proposed a non-autoregressive generation-based streaming speech decoder. Although it achieves a lower response latency by synchronously generating speech and text, limited by the expressive ability of the non-autoregressive model, the naturalness and fluency of the generated speech are usually poor. In addition, existing technologies combine autoregressive and non-autoregressive architectures for speech generation. Although the naturalness of the speech is improved, it cannot achieve a lower response latency. Summary of the Invention

[0006] The purpose of the present invention is to solve the problem that it is difficult to simultaneously ensure real-time performance and naturalness in the process of speech generation by existing modular speech large models. The existing speech decoder design has deficiencies, resulting in the generated speech being prone to discontinuity, unnaturalness, or high latency. Therefore, the present invention proposes an autoregressive streaming speech decoder and combines it with a large language model to finally achieve real-time and natural end-to-end speech interaction with the large language model.

[0007] Aiming at the deficiencies of the existing technology, as Figure 3 shown, the present invention proposes a large model real-time speech interaction method based on autoregressive speech synthesis, which includes:

[0008] Initial step, obtaining a speech instruction with a labeled target text response and a target speech response, the speech encoder encodes the speech instruction into a speech representation, and the speech adapter performs dimensionality reduction and feature transformation on the original speech representation;

[0009] Inference step, the large language model generates a hidden state based on the transformed speech representation and samples the hidden state to obtain a text sequence;

[0010] Conversion step: Process the text sequence using a text-to-speech language model based on the autoregressive Transformer structure, streamingly generate a sequence of speech tokens, and convert this sequence of speech tokens into a speech signal through a vocoder;

[0011] Training step: Construct a first loss function based on the text sequence and the target text response, construct a second loss function based on the speech signal and the target speech response, combine the first loss function and the second loss function to obtain a final loss function, and use the final loss function to first train the speech adapter and the large language model. After the training is completed, train the text-to-speech language model;

[0012] Interaction step: Input the speech to be interacted into the speech encoder and the trained speech adapter to obtain its speech representation. The trained large language model obtains its text sequence based on the speech representation of the speech to be interacted, and the trained text-to-speech language model generates its speech signal based on the text sequence of the speech to be interacted as the speech interaction result of the speech to be interacted.

[0013] The large model real-time speech interaction method based on autoregressive speech synthesis, wherein the adapter performs dimensionality reduction and feature transformation on the original speech representation

[0014] The speech adapter is used to splice every k consecutive frames along the feature dimension of the original speech representation to obtain a dimensionality-reduced feature, and encode the dimensionality-reduced feature to obtain the transformed speech representation.

[0015] The large model real-time speech interaction method based on autoregressive speech synthesis, wherein the text-to-speech language model adopts an autoregressive Transformer structure and autoregressively generates a sequence of speech tokens Y based on the output of the LLM U , where,

[0016] The text-to-speech language model uses a feed-forward network FFN to map the hidden state:

[0017]

[0018] Generate a text sequence Y from the large language model T Calculate the text embedding:

[0019]

[0020] Then, calculate the gating factor g i :

[0021]

[0022] where σ is the sigmoid activation function, Wg and b g are gating weight parameters; the final input representation C is calculated through gated fusion:

[0023]

[0024] This fused representation C = [c1, …, c N is used as the input to this text-to-speech language model;

[0025] The specific process of streaming the generation of the speech token sequence includes:

[0026]

[0027] Among them, represents the read fused representation, R: W is the ratio of reading to writing; after reading R fused representations, this text-to-speech language model generates W speech tokens, and the Mel spectrogram synthesis is performed on these W speech tokens and converted into a speech signal Y through this vocoder S .

[0028] Such as Figure 4 shown, the present invention also proposes a large model real-time voice interaction device based on autoregressive speech synthesis, which includes:

[0029] An initial module, which obtains the voice instructions of the labeled target text response and the target voice response, the voice encoder encodes the voice instructions into a voice representation, and the voice adapter performs dimensionality reduction and feature transformation on the original voice representation;

[0030] An inference module, the large language model generates a hidden state and samples the hidden state according to the converted voice representation to obtain a text sequence;

[0031] A conversion module, which uses a text-to-speech language model based on the autoregressive Transformer structure to process the text sequence, streamingly generates a speech token sequence, and converts the speech token sequence into a speech signal through a vocoder;

[0032] A training module, which constructs a first loss function according to the text sequence and the target text response, constructs a second loss function according to the speech signal and the target voice response, combines the first loss function and the second loss function to obtain a final loss function, and uses the final loss function to first train the voice adapter and the large language model. After the training is completed, the text-to-speech language model is trained;

[0033] The interaction module inputs the voice to be interacted into the voice encoder and the trained voice adapter to obtain its voice representation. The trained large language model obtains its text sequence based on the voice representation of the voice to be interacted, and the trained text-to-speech language model generates its voice signal as the voice interaction result of the voice to be interacted according to the text sequence of the voice to be interacted.

[0034] The large model real-time voice interaction device based on autoregressive speech synthesis, wherein the adapter performs dimensionality reduction and feature transformation on the original voice representation

[0035] The voice adapter is used to splice every k consecutive frames along the feature dimension of the original voice representation to obtain a dimensionality-reduced feature, and encodes the dimensionality-reduced feature to obtain the transformed voice representation.

[0036] The large model real-time voice interaction device based on autoregressive speech synthesis, wherein the text-to-speech language model adopts an autoregressive Transformer structure and autoregressively generates a voice token sequence Y based on the output of the LLM U , where

[0037] The text-to-speech language model maps the hidden state using a feed-forward network FFN:

[0038]

[0039] Calculate the text embedding from the text sequence Y generated by the large language model T :

[0040]

[0041] Then, calculate the gating factor g i :

[0042]

[0043] where σ is the sigmoid activation function, W g and b g are gating weight parameters; calculate the final input representation C through gating fusion:

[0044]

[0045] The fused representation C = [c1,…,c N is used as the input of the text-to-speech language model;

[0046] The specific process of streaming generation of the voice token sequence includes:

[0047]

[0048] Among them, represents the read fusion representation, and R:W is the ratio of reading to writing; after reading R fusion representations, the text-to-speech language model generates W speech tokens, performs Mel spectrogram synthesis on the W speech tokens, and converts them into a speech signal Y through the vocoder S .

[0049] The present invention also proposes a client for any one of the large model real-time voice interaction devices based on autoregressive speech synthesis.

[0050] The present invention also proposes an electronic device, which includes one of the large model real-time voice interaction devices based on autoregressive speech synthesis. The electronic device is either connected to an information display device, which is used to output and display the voice interaction result according to the display parameters, attributes set by the user or through an artificial intelligence model.

[0051] The present invention also proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the large model real-time voice interaction method based on autoregressive speech synthesis are implemented.

[0052] The present invention also proposes a computer program product, including a computer program, where when the computer program is executed by a processor, the steps of the large model real-time voice interaction method based on autoregressive speech synthesis are implemented.

[0053] As can be seen from the above solutions, the advantages of the present invention are as follows:

[0054] As Figure 1 shown, compared with the prior art, the method of the present invention not only ensures high real-time performance but also greatly improves the naturalness and fluency of speech synthesis. The optimized speech decoding architecture effectively reduces the speech generation delay and improves the response speed of the voice interaction system. Experimental results show that in speech question-answering tasks and speech instruction following tasks, this method is significantly superior to models with the same number of parameters in terms of reply quality, speech quality, response delay, etc., improving the user experience and the practicality of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 is a table graph of the effect comparison between the present invention and other methods on the SpokenQA and Speech Instruction Following benchmarks;

[0056] Figure 2 is the overall framework diagram of the present invention;

[0057] Figure 3 is the flowchart of the method of the present invention;

[0058] Figure 4 This is the device module diagram of the present invention;

[0059] Figure 5 This is the schematic structural diagram of the first electronic device of the present invention;

[0060] Figure 6 This is the schematic structural diagram of the application environment of the first electronic device of the present invention;

[0061] Figure 7 This is the schematic structural diagram of the second electronic device of the present invention.

[0062] Reference numerals:

[0063] A - The first electronic device;

[0064] B - The large - model real - time voice interaction device based on autoregressive speech synthesis;

[0065] C - Data acquisition device;

[0066] D - Information display device;

[0067] 1000 - The second electronic device;

[0068] Ⅰ - Computing unit;

[0069] Ⅱ - ROM;

[0070] Ⅲ - RAM;

[0071] Ⅳ - Bus;

[0072] Ⅴ - Interface;

[0073] Ⅵ - Input unit;

[0074] Ⅶ - Output unit;

[0075] Ⅷ - Storage medium;

[0076] Ⅸ - Communication unit. Detailed implementation manners

[0077] It should be noted that in this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non - exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device.

[0078] Without further limitation, an element qualified by the statement "comprising an..." does not exclude the presence of additional identical elements in a process, method, article, or apparatus that comprises the element.

[0079] The processor described in the present invention is the control center of an electronic device, which can be a single processor or a collective term for multiple processing elements. For example, it can be one or more central processing units (CPUs), or it can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0080] Optionally, the processor can execute various functions of the electronic device by running or executing software programs stored in the memory and by invoking data stored in the memory.

[0081] In a specific implementation, as an embodiment, the processor can include one or more CPUs. Each of these processors can be a single-CPU or a multi-CPU. Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). The electronic device can include: servers, desktop computers, laptop computers, smartphones, tablets, embedded computers, etc., where the embedded computer includes vehicles and robots, etc.

[0082] The memory is used to store the software program for implementing the solution of the present invention and is controlled by the processor for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated here.

[0083] It should be noted that the structure of the electronic device shown in the drawings of the present invention does not constitute a limitation thereto. The actual knowledge structure recognition device may include more or fewer components than those shown in the drawings, or combine certain components, or have different component arrangements.

[0084] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0085] It should also be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be understood specifically with reference to the context before and after.

[0086] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0087] It should also be understood that in various embodiments of the present invention, the magnitudes of the serial numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0088] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0089] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0090] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0091] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0092] In the process of researching end-to-end speech large models, the inventors found that when generating speech, the modular speech large models in the prior art often have difficulty ensuring the naturalness and fluency of the speech while ensuring real-time performance. The core of this problem lies in the insufficient design of the speech decoder in the existing modular solutions, resulting in problems such as discontinuous, unnatural, or high-latency speech when generating speech.

[0093] To address this technical challenge, the inventors conducted in-depth research and proposed a brand-new speech decoding architecture to improve the quality of speech generation while meeting the requirements of high real-time performance. First, an autoregressive streaming speech synthesis architecture was introduced. The autoregressive architecture can better model the language characteristics of speech, making the generated speech more in line with the natural rhythm of speech, and this architecture has good scalability. However, the autoregressive architecture usually needs to receive the complete input before starting to generate speech, which will cause a relatively high latency. To achieve real-time generation, the inventors designed an alternating read-write strategy, enabling the decoder to start speech synthesis immediately after obtaining partial input information, thus significantly reducing latency and improving the system's response speed.

[0094] On this basis, to further optimize the naturalness and fluency of the generated speech, the inventors introduced a Causal Flow Matching model, which can more precisely convert speech discrete symbols into Mel spectrogram features. By introducing a Lookahead layer and a causal Transformer layer, the speech streaming synthesis process becomes smoother, avoiding problems such as speech discontinuity or incoherence, thereby improving the final listening quality.

[0095] After determining the efficient speech decoding architecture, the inventors further explored how to integrate it with a large language model (LLM) to implement an end-to-end speech interaction system based on the large language model. For this purpose, the inventors proposed an adaptive gating mechanism, which can dynamically adjust the continuous representation and discrete text output by the large language model to better adapt to the input requirements of the speech decoder. This mechanism can flexibly adjust the fusion strategy in different scenarios, enabling the speech generation process to consider both the context information of the language model and capture precise semantic information.

[0096] To ensure that the model has good generalization ability and end-to-end training effect, the inventors constructed a training dataset containing 200,000 speech-to-speech multi-turn dialogue data and conducted end-to-end training based on this data. This data covers dialogue scenarios in different contexts and includes a wide range of input voices, ensuring that the model can stably generate smooth and natural speech when facing various inputs. Finally, after multiple rounds of optimization and experimental verification, the proposed speech large model has successfully achieved high-quality speech synthesis while maintaining high real-time performance, significantly improving the user experience of the speech interaction system. Specifically, to achieve the above technical effects, the present invention proposes the following key technical points:

[0097] Key Point 1: Autoregressive streaming speech synthesis architecture. The autoregressive architecture can better model the language characteristics of speech, making the generated speech more in line with the natural rhythm, and at the same time reducing latency through the alternating read-write strategy to achieve real-time speech synthesis.

[0098] Key Point 2: Adaptive Gating Mechanism. Dynamically adjust the continuous representations and discrete texts output by the large language model to better suit the input requirements of the speech decoder, and flexibly adjust the fusion strategy in different scenarios to ensure that speech generation can consider context information and accurately express semantics.

[0099] Key Point 3: Speech-to-Speech Multi-Round Dialogue Dataset. Construct a multi-round dialogue dataset containing 200,000 speech-to-speech dialogues, covering a wide range of contexts and voices, to ensure that the model can stably generate high-quality speech under various input conditions and improve generalization ability.

[0100] To make the above features and effects of the present invention more clearly and understandably described, specific embodiments are hereinafter given and detailed descriptions are made in conjunction with the accompanying drawings of the specification. The specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are only for illustrative purposes. The scope of protection of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.

[0101] As Figure 2 shown, the present invention proposes an end-to-end speech interaction system LLaMA-Omni 2 based on a large language model to achieve high-quality real-time speech dialogue. The system adopts a modular architecture, including a speech encoder, a speech adapter, a large language model, a text-to-speech language model (TTS-LM), a flow matching model, and a vocoder, to achieve a complete modeling from speech input to speech output.

[0102] LLaMA-Omni 2 uses the Qwen2.5 series model as the LLM and adopts the encoder of Whisper as the speech encoder. To achieve streaming speech generation, the present invention adopts an autoregressive streaming generation strategy at the TTS-LM end and combines a flow matching model to synthesize the mel spectrogram block by block. In addition, the present invention uses 200,000 multi-round speech dialogue data for training and adopts a two-stage training method to efficiently improve the speech understanding and generation ability of the model.

[0103] The input of the present invention is a speech command X, and its goal is to generate a corresponding text response Y T and a speech response Y S . The core components of the system are as follows:

[0104] First, the Speech Encoder is used to convert the input speech signal X into a series of speech representations E. The present invention adopts the encoder of Whisper-large-v3 to map the speech input into a high-dimensional feature sequence to capture rich speech information.

[0105] Then, the Speech Adapter is used to reduce the dimension and transform the features of the encoded speech representation E to meet the input requirements of the LLM. Specifically, the adapter includes a downsampling module and a Feed-Forward Network (FFN). The downsampling module concatenates every k consecutive frames along the feature dimension to reduce the sequence length, while the FFN further encodes the downsampled features to enhance the representation ability. The finally obtained speech feature representation E' is input into the LLM for processing.

[0106] After obtaining the adapted speech representation E', the large language model sequentially generates the text response Y. T . The output of the LLM includes both the continuous hidden states H and the text sequence Y sampled based on the hidden states. T . The hidden states H have rich context information, while the text sequence Y T provides the exact language content.

[0107] To convert the text sequence into a speech sequence, the present invention designs a Text-to-Speech Language Model (TTS-LM). The TTS-LM adopts an autoregressive Transformer architecture and autoregressively generates the speech token sequence Y based on the output of the LLM. U , where the meaning of k is a natural number, and N is the set of natural numbers.

[0108] To better fuse the hidden states H and the text sequence Y generated by the LLM, T the present invention introduces a Gate Fusion mechanism before the TTS-LM. Specifically, first, a two-layer Feed-Forward Network (FFN) is used to map the LLM hidden states to obtain the continuous hidden states.

[0109]

[0110] Meanwhile, the text embedding is calculated from the text sequence Y T generated by the LLM.

[0111]

[0112] For the i-th word in the text sequence Y T , Emb is the embedding matrix;

[0113] Then, the gating factor g is calculated i :

[0114]

[0115] where σ is the sigmoid activation function, W g and b g are the gating weight parameters. Finally, the final input representation is calculated through gating fusion:

[0116]

[0117] This fused representation C = [c1, …, c N is used as the input of the TTS-LM, ensuring that the model can utilize the context information of the LLM and accurately align the text semantics. To achieve the streaming generation of speech tokens, the present invention adopts the "Read-R-Write-W" strategy, that is:

[0118]

[0119] where N is the sequence length of the target text and M is the sequence length of the target speech tokens, represents the fused representations that have been read, and R:W is the ratio of reading to writing. After reading R fused representations, the TTS-LM can generate W speech tokens, and then the W speech tokens are input into the flow matching model for Mel spectrogram synthesis. Specifically, the present invention adopts a causal flow matching model, enabling the synthesis of Mel spectrogram blocks after every W speech tokens are generated, thereby achieving streaming synthesis. Finally, the Mel spectrogram is converted into the final speech signal Y S .

[0120] To train the model of the present invention, the inventor constructed a dataset of 200,000 multi-turn voice conversations. During the data construction process, first, the large language model was used to generate 200,000 multi-turn text conversation samples, and further, CosyVoice2-0.5B was used to convert the text conversations into speech. During the speech synthesis process, the timbre diversity of the input speech was ensured, while the timbre of the output speech remained consistent. The training process is divided into two stages: In the first stage, the speech understanding (Speech-to-Text) and text-to-speech conversion (Text-to-Speech) parts are trained independently. When training speech understanding, the speech encoder is fixed, and only the speech adapter and the large language model are trained. In the second stage, the LLM and the speech encoder are fixed, and only the TTS-LM model is optimized to enable it to efficiently adapt to the end-to-end voice interaction task.

[0121] The present invention has achieved leading performance in both the SpokenQA and Speech Instruction Following tasks, and at the same time, has achieved a response delay as low as 600 ms, meeting the requirements of real-time voice interaction.

[0122] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied in the above embodiment.

[0123] As Figure 4 shown, the present invention also proposes a large model real-time voice interaction device based on autoregressive speech synthesis, which includes:

[0124] An initial module that obtains a voice command with a marked target text response and a target voice response. The voice encoder encodes the voice command into a voice representation, and the voice adapter performs dimensionality reduction and feature transformation on the original voice representation.

[0125] An inference module where the large language model generates a hidden state based on the transformed voice representation and samples the hidden state to obtain a text sequence.

[0126] A conversion module that processes the text sequence using a text-to-speech language model based on the autoregressive Transformer structure, streamingly generates a voice token sequence, and converts the voice token sequence into a voice signal through a vocoder.

[0127] A training module that constructs a first loss function based on the text sequence and the target text response, constructs a second loss function based on the voice signal and the target voice response, combines the first loss function and the second loss function to obtain a final loss function, and uses the final loss function to first train the voice adapter and the large language model. After the training is completed, the text-to-speech language model is trained.

[0128] An interaction module that inputs the voice to be interacted into the voice encoder and the trained voice adapter to obtain its voice representation. The trained large language model obtains its text sequence based on the voice representation of the voice to be interacted, and the trained text-to-speech language model generates its voice signal as the voice interaction result of the voice to be interacted according to the text sequence of the voice to be interacted.

[0129] In the large model real-time voice interaction device based on autoregressive speech synthesis, the adapter performs dimensionality reduction and feature transformation on the original voice representation

[0130] The voice adapter is used to splice every k consecutive frames along the feature dimension of the original voice representation to obtain a dimensionality-reduced feature, and encodes the dimensionality-reduced feature to obtain the transformed voice representation.

[0131] The described large - model real - time voice interaction device based on autoregressive speech synthesis, where the text - to - speech language model adopts an autoregressive Transformer structure and autoregressively generates a sequence of speech tokens Y based on the output of the LLM U , where,

[0132] The text - to - speech language model uses a feed - forward network FFN to map the hidden state:

[0133]

[0134] Calculate the text embedding from the text sequence Y generated by the large - language model T :

[0135]

[0136] Then, calculate the gating factor g i :

[0137]

[0138] where σ is the sigmoid activation function, W g and b g are gating weight parameters; calculate the final input representation C through gated fusion:

[0139]

[0140] This fused representation C = [c1,…,c N is used as the input to the text - to - speech language model;

[0141] The specific process of streaming - generating a sequence of speech tokens includes:

[0142]

[0143] where, represents the fused representations that have been read, R:W is the ratio of reading to writing; after reading R fused representations, the text - to - speech language model generates W speech tokens, and these W speech tokens are subjected to Mel - spectrum synthesis and converted into a speech signal Y through the vocoder S .

[0144] The present invention also proposes a client for any of the above - mentioned large - model real - time voice interaction devices based on autoregressive speech synthesis.

[0145] As Figure 5 shown, in another embodiment, the present invention also proposes a first electronic device A, which includes the above - mentioned large - model real - time voice interaction device based on autoregressive speech synthesis.

[0146] As shown Figure 6 in the figure, the first electronic device A can also be connected to the data acquisition device C and the information display device D through a wired or wireless information transmission scheme. The data acquisition device C is used to collect voice commands, and the information display device D is used to display the voice interaction results obtained by the analysis of the present invention.

[0147] Among them, the information display device D can process the data output by the first electronic device A based on the information display mechanism to improve the readability of the data output by the first electronic device A. The information display mechanism can be preset manually. For example, the data output by the first electronic device A is visually displayed, and it can display according to the display parameters and / or attributes set by the user. The display parameters can be, for example, the display data range, and the display attributes can be, for example, the display font, color, whether to scroll and play, etc. Present the key information specified by the user to the user. Users can understand this information more timely without having to access the secondary page or scroll the page, saving the user's operations. Or the information display mechanism can be an artificial intelligence AI display model, which can learn the key information of the user according to the user's previous usage habits, such as viewing duration, click times, editing times, etc., and then automatically present rich and necessary key information to the user.

[0148] The present invention also provides a computer program product, the computer program product includes a computer program, the computer program can be stored on a readable storage medium, and when the computer program is executed by a processor, the computer can execute the real-time voice interaction method based on the autoregressive speech synthesis large model provided by the above methods.

[0149] In another embodiment, the present invention further provides a storage medium VIII for storing a computer program for executing the real-time voice interaction method based on the autoregressive speech synthesis large model. It should be understood that the storage medium in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0150] Figure 7 FIG. shows a schematic block diagram of a second electronic device 1000 that can be used to implement the embodiments of the present invention. The second electronic device 1000 is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The second electronic device 1000 may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein. The second electronic device 1000 may be the same as or different from the first electronic device A.

[0151] The second electronic device 1000 includes a computing unit I, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory II (ROM) or computer programs loaded from a storage medium VIII into a random access memory (RAM) III. In the RAM III, various programs and data required for the operation of the device 1000 can also be stored. The computing unit I, the ROM II, and the RAM III are connected to each other via a bus IV. An input / output (I / O) interface V is also connected to the bus IV.

[0152] Multiple components in the second electronic device 1000 are connected to the I / O interface V, including: an input unit VI, such as a keyboard, a mouse, etc.; an output unit VII, such as various types of displays, speakers, etc.; a storage medium VIII, such as a magnetic disk, an optical disc, etc.; and a communication unit IX, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit IX allows the second electronic device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0153] The computing unit I can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit I include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit I executes the various methods and processes described above, such as method steps S1 - S5. For example, in some embodiments, the method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage medium VIII. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM II and / or the communication unit IX. When the computer program is loaded into the RAM III and executed by the computing unit I, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the computing unit I can be configured to execute the method by any other appropriate means (e.g., by means of firmware).

[0154] Although the embodiments of the present invention have been disclosed as above, they are not limited to only the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present invention. For those skilled in the art, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to specific details and the illustrated and described examples here.

Claims

1. A real-time voice interaction method based on autoregressive speech synthesis, characterized in that Including: Initial step: Obtain the voice command of the labeled target text response and the target voice response. The voice encoder encodes the voice command into a voice representation, and the voice adapter performs dimensionality reduction and feature transformation on the original voice representation; Inference step: The large language model generates a hidden state based on the transformed voice representation and samples the hidden state to obtain a text sequence; Conversion step: Use a text-to-speech language model based on the autoregressive Transformer structure to process the text sequence, streamingly generate a voice token sequence, and convert the voice token sequence into a voice signal through a vocoder; Training step: Construct a first loss function based on the text sequence and the target text response, construct a second loss function based on the voice signal and the target voice response, combine the first loss function and the second loss function to obtain a final loss function, and use the final loss function to first train the voice adapter and the large language model. After the training is completed, train the text-to-speech language model; Interaction step: Input the voice to be interacted into the voice encoder and the trained voice adapter to obtain its voice representation. The trained large language model obtains its text sequence based on the voice representation of the voice to be interacted, and the trained text-to-speech language model generates its voice signal as the voice interaction result of the voice to be interacted according to the text sequence of the voice to be interacted.

2. The real-time voice interaction method based on the autoregressive speech synthesis large model according to claim 1, wherein, The adapter performs dimensionality reduction and feature transformation on the original voice representation The voice adapter is used to splice every k consecutive frames along the feature dimension of the original voice representation to obtain a dimensionality-reduced feature, and encode the dimensionality-reduced feature to obtain the transformed voice representation.

3. The real-time voice interaction method based on autoregressive speech synthesis large model according to claim 1 or 2, characterized in that, The text-to-speech language model adopts an autoregressive Transformer architecture and autoregressively generates a sequence of speech tokens Y based on the output of the LLM U , where The text-to-speech language model maps the hidden state using a feed-forward network FFN: The text sequence Y generated from the large language model T Calculate the text embedding: Then, calculate the gating factor g i : where σ is the sigmoid activation function, W g and b g are the gating weight parameters; the final input representation C is calculated through gating fusion: This fused representation C = [c1, …, c N is used as the input to the text-to-speech language model; The specific process of streamingly generating a voice token sequence includes: Among them, represents the read fusion representation, and R:W is the ratio of reading to writing; after reading R fusion representations, the text-to-speech language model generates W speech tokens, performs Mel-spectrum synthesis on the W speech tokens, and converts them into a speech signal Y through the vocoder S .

4. A real-time voice interaction device based on an autoregressive speech synthesis large model, characterized in that, Including: Initial module: Obtain the voice command of the labeled target text response and the target voice response. The voice encoder encodes the voice command into a voice representation, and the voice adapter performs dimensionality reduction and feature transformation on the original voice representation; Inference module: The large language model generates a hidden state based on the transformed voice representation and samples the hidden state to obtain a text sequence; Conversion module: Use a text-to-speech language model based on the autoregressive Transformer structure to process the text sequence, streamingly generate a voice token sequence, and convert the voice token sequence into a voice signal through a vocoder; Training module: Construct a first loss function based on the text sequence and the target text response, construct a second loss function based on the voice signal and the target voice response, combine the first loss function and the second loss function to obtain a final loss function, and use the final loss function to first train the voice adapter and the large language model. After the training is completed, train the text-to-speech language model; The interaction module inputs the voice to be interacted into the voice encoder and the trained voice adapter to obtain its voice representation. The trained large language model obtains its text sequence based on the voice representation of the voice to be interacted. The trained text-to-speech language model generates its voice signal as the voice interaction result of the voice to be interacted according to the text sequence of the voice to be interacted.

5. The real-time voice interaction device based on the autoregressive speech synthesis large model according to claim 4, characterized in that, The adapter performs dimensionality reduction and feature transformation on the original voice representation The voice adapter is used to splice every k consecutive frames along the feature dimension of the original voice representation to obtain the dimensionality-reduced feature, and encodes the dimensionality-reduced feature to obtain the transformed voice representation.

6. The real-time voice interaction device based on autoregressive speech synthesis large model according to claim 4 or 5, characterized in that The text-to-speech language model adopts an autoregressive Transformer architecture and autoregressively generates a sequence of speech tokens Y based on the output of the LLM U , where The text-to-speech language model uses a feed-forward network FFN to map the hidden state: The text sequence Y generated from the large language model T Calculate text embeddings: Then, calculate the gating factor g i : where σ is the sigmoid activation function, W g and b g are the gating weight parameters; the final input representation C is calculated through gated fusion: This fused representation C = [c1, …, c N is used as the input to the text-to-speech language model; The specific process of streaming the generation of the voice token sequence includes: Among them, represents the read fusion representation, and R:W is the ratio of reading to writing; after reading R fusion representations, the text-to-speech language model generates W speech tokens, performs Mel spectrogram synthesis on the W speech tokens, and converts them into a speech signal Y through the vocoder S .

7. A client for any one of the large model real-time voice interaction devices based on autoregressive speech synthesis described in claims 4-6.

8. An electronic device, characterized in that, It includes a large model real-time voice interaction device based on autoregressive speech synthesis described in claims 4-6. The electronic device is connected to an information display device, and the information display device is used to output and display the voice interaction result according to the display parameters, attributes set by the user, or through an artificial intelligence model.

9. A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the large model real-time voice interaction method based on autoregressive speech synthesis described in any one of claims 1-3.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the large model real-time voice interaction method based on autoregressive speech synthesis described in any one of claims 1-3.

Citation Information

Cited By

  • Phonetic symbol generation method and device based on deep neural network architecture, and storage medium

    CN121565139A