Speech generation method and device based on large language model, equipment and medium

By introducing a hybrid LoRA adapter into a large language model and performing multi-stage parameter fine-tuning, combining pre-trained acoustic models and vocoders, the problem of high computing resource requirements in the prior art is solved, and efficient speech generation and fusion of text and speech modalities are achieved.

CN120148472APending Publication Date: 2025-06-13PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510342562.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing speech generation technology based on large language models requires training from scratch, resulting in a high demand for computing hardware and resources, and reducing the efficiency of speech generation.

Method used

By introducing a hybrid LoRA adapter into a large language model and performing multi-stage parameter fine-tuning, semantic features containing semantic information and pronunciation information are generated, and speech generation is achieved by combining pre-trained acoustic models and vocoders.

Benefits of technology

It effectively improves the speech generation efficiency based on large language models, reduces the requirements for computing hardware and resources, and does not require training from scratch, and realizes the effective fusion of text and speech modalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148472A_ABST
    Figure CN120148472A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a voice generation method and device based on a large language model, equipment and a medium. The original text is input into a large language model which is obtained through multi-stage parameter fine tuning and is provided with a mixed LoRA adapter for text processing, and semantic features containing semantic information and rhythm information are generated; inputting the semantic features into a pre-trained acoustic model for feature conversion, and converting the semantic features into corresponding acoustic features; and inputting the acoustic features into a pre-trained vocoder for decoding processing, and generating a voice waveform corresponding to the original text. The big language model with the mixed LoRA adapter obtained through multi-stage parameter fine tuning is processed, the prior knowledge of the big language model is effectively utilized to realize the fusion of two modes of text and voice, the voice generation efficiency is improved, and the requirement on hardware resources is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a speech generation method, device, equipment and medium based on a large language model. Background Art

[0002] Text-To-Speech (TTS) technology refers to the process of generating speech from text. With the development of artificial intelligence technology, speech synthesis plays an increasingly important role in the field of human-computer dialogue. TTS speech synthesis technology has been widely used in various scenarios such as the medical field and the financial field. For example, in the field of healthcare, medical robots can use speech synthesis technology to convey information, tell patients about their treatment plans or tell doctors about their medical conditions, and hospital voice navigation systems can help patients find relevant facilities and services within the hospital; in the field of fintech business, financial institutions can use TTS technology to generate voice prompts and announcements for notifying customers about account changes, new product releases, etc., or create an automatic customer service system with natural language processing capabilities that can answer customers' questions and provide services such as account information query and product consultation.

[0003] In recent years, large language models (LLMs) have demonstrated excellent performance in various text-based tasks such as question answering, machine translation, and common sense reasoning. The development of large language models with speech generation capabilities is closely related to the progress of TTS speech synthesis technology. Currently, a speech generation system is usually rebuilt based on language modeling tasks similar to large language models, but this method only focuses on the goals of TTS and requires training a speech generation system from scratch, which has high requirements for computing hardware and resource consumption, reducing the speech generation efficiency based on large language models. Summary of the Invention

[0004] In view of the above deficiencies of the prior art, the purpose of the present invention is to provide a speech generation method, device, equipment and medium based on a large language model that can be applied to the medical field, fintech or other related fields, and its main purpose is to improve the speech generation efficiency based on the large language model and reduce the requirements for computing hardware and resources.

[0005] The technical solution of the present invention is as follows:

[0006] The first aspect of the present invention provides a speech generation method based on a large language model, including:

[0007] Obtain the original text of the speech to be generated;

[0008] Input the original text into a large language model with a hybrid LoRA adapter for text processing to generate semantic features containing semantic information and prosody information;

[0009] Input the semantic features into a pre-trained acoustic model for feature transformation, and transform the semantic features into corresponding acoustic features;

[0010] Input the acoustic features into a pre-trained vocoder for decoding processing to generate the speech waveform corresponding to the original text;

[0011] Wherein, the large language model with a hybrid LoRA adapter is obtained through multi-stage parameter fine-tuning.

[0012] The second aspect of the present invention provides a speech generation device based on a large language model, including:

[0013] An acquisition module for acquiring the original text of the speech to be generated;

[0014] A text processing module for inputting the original text into a large language model with a hybrid LoRA adapter for text processing to generate semantic features including semantic information and prosody information;

[0015] An acoustic processing module for inputting the semantic features into a pre-trained acoustic model for feature transformation, and transforming the semantic features into corresponding acoustic features;

[0016] A speech decoding module for inputting the acoustic features into a pre-trained vocoder for decoding processing to generate the speech waveform corresponding to the original text;

[0017] Wherein, the large language model with a hybrid LoRA adapter is obtained through multi-stage parameter fine-tuning.

[0018] The third aspect of the present invention provides a computer device, including at least one processor; and,

[0019] A memory communicatively connected to the at least one processor; wherein,

[0020] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned speech generation method based on a large language model.

[0021] The fourth aspect of the present invention provides a non-volatile computer-readable storage medium, which stores computer-executable instructions, and when the computer-executable instructions are executed by one or more processors, the one or more processors can execute the above-mentioned speech generation method based on a large language model.

[0022] Beneficial effects: The present invention discloses a speech generation method, device, equipment and medium based on a large language model. Compared with the prior art, in the embodiments of the present invention, the original text of the speech to be generated is obtained; the original text is input into a large language model with a hybrid LoRA adapter for text processing to generate semantic features containing semantic information and prosody information; the semantic features are input into a pre-trained acoustic model for feature conversion to convert the semantic features into corresponding acoustic features; the acoustic features are input into a pre-trained vocoder for decoding processing to generate a speech waveform corresponding to the original text; wherein, the large language model with a hybrid LoRA adapter is obtained through multi-stage parameter fine-tuning. The large language model with a hybrid LoRA adapter obtained through multi-stage parameter fine-tuning processes the input text, and after obtaining features that integrate the text and speech modalities, speech generation is realized. The prior knowledge of the large language model in understanding text is effectively utilized to achieve effective integration of the two modalities of text and speech, without the need to train from scratch, effectively improving the speech generation efficiency based on the large language model and reducing the requirements for computing hardware and resources. Description of the Drawings

[0023] In order to more clearly illustrate the solutions in the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 It is a schematic diagram of an application environment for the speech generation method based on a large language model provided by an embodiment of the present invention;

[0025] Figure 2 It is a flowchart of the speech generation method based on a large language model provided by an embodiment of the present invention;

[0026] Figure 3 It is another flowchart of the speech generation method based on a large language model provided by an embodiment of the present invention;

[0027] Figure 4 It is a flowchart of step S202 in the speech generation method based on a large language model provided by an embodiment of the present invention;

[0028] Figure 5 It is a flowchart of step S203 in the speech generation method based on a large language model provided by an embodiment of the present invention;

[0029] Figure 6 It is a schematic diagram of the functional modules of the speech generation device based on a large language model provided by an embodiment of the present invention;

[0030] Figure 7 This is a schematic diagram of the hardware structure of the computer device provided by the embodiments of the present invention. Detailed implementation manners

[0031] To make the objectives, technical solutions and effects of the present invention clearer and more definite, the present invention will be further described in detail below. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. The embodiments of the present invention will be introduced below with reference to the accompanying drawings.

[0032] The speech generation method based on a large language model provided by the embodiments of the present invention can be applied in an application environment such as Figure 1 which includes a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0033] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only for example).

[0034] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0035] The server 105 can be a server that provides various services. For example, it can be a background server (only an example) that supports the content browsed by the user using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device. The server 105 can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services (″Virtual Private Server″, or simply referred to as ″VPS″). The server 105 can also be a server of a distributed system, or a server combined with blockchain.

[0036] It should be noted that the speech generation method based on a large language model provided in the embodiments of the present application can generally be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the speech generation device based on a large language model provided in the embodiments of the present invention can also be set in the first terminal device 101, the second terminal device 102, or the third terminal device 103. Or, the speech generation method based on a large language model provided in the embodiments of the present invention can generally also be executed by the server 105. Correspondingly, the speech generation device based on a large language model provided in the embodiments of the present invention can generally be set in the server 105.

[0037] It should be understood that the numbers of the above terminal devices, networks, and servers are only illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.

[0038] As Figure 2 shown, the speech generation method based on a large language model provided in the embodiments of the present invention specifically includes the following steps:

[0039] S201. Obtain the original text of the speech to be generated.

[0040] In this embodiment, the original text refers to the text content that needs to be converted into speech, which can be text in any language, such as Chinese, English, etc. Specifically, the text of the speech to be generated can be obtained through various methods such as user input, file reading, and web crawling. For example, the user enters a paragraph of text in the software interface, or reads the pre-stored text data from the database, etc. This embodiment does not make any limitations. Through text input from various sources, it provides wide applicability for subsequent speech generation.

[0041] Exemplarily, in the application scenarios in the medical field, for example, when doctors make ward rounds, they can input the description of the patient's condition as the original text through the voice assistant interface, such as "The patient's body temperature is 38.5°C, and the main complaints are headache and fatigue", so that the voice assistant can convert this text into voice and broadcast it to nurses or other medical staff, facilitating medical staff to obtain information by listening to the voice when they are busy. Or in the application scenarios in the financial field, for example, the financial broadcast assistant obtains stock market information as the original text, such as "The overall market rose today, and the technology sector performed strongly", and converts this information into voice for users to listen to when driving or busy, etc.

[0042] S202. Input the original text into a large language model with a hybrid LoRA adapter for text processing to generate semantic features containing semantic information and prosodic information, where the large language model with the hybrid LoRA adapter is obtained through multi-stage parameter fine-tuning.

[0043] In this embodiment, the obtained original text is input into a large language model with a hybrid LoRA adapter for text processing. The large language model with the hybrid LoRA adapter is obtained through multi-stage parameter fine-tuning. The LoRA adapter is a parameter-efficient fine-tuning method that adjusts the model's parameters by adding learnable low-rank matrices to the self-attention layer and the feed-forward layer of the model. It can effectively fine-tune the model without significantly increasing the number of model parameters. Based on the prior knowledge of the large language model in understanding text, by adding a LoRA adapter for model fine-tuning and performing multi-stage parameter fine-tuning, the hybrid LoRA adapter after fine-tuning can achieve an effective fusion of text and speech modalities, add the speech modality to the large language model, and the fine-tuned large language model can simultaneously process text and semantic information, converting the input original text into semantic features containing semantic information and prosodic information. This large language model obtained by the multi-stage fine-tuning method avoids training the entire model from scratch and significantly reduces the demand for computing resources.

[0044] Specifically, the generated semantic features not only contain the semantic information of the text but also incorporate prosodic information. Semantic information refers to the meaning of the text, and prosodic information refers to features such as the rhythm and intonation of speech. For example, the semantic information of the sentence "Hello, nice to meet you" is to express greetings, and the prosodic information may include the rise and fall of intonation. In a medical scenario, for example, when the input text is "Please take me to the emergency room", the semantic features generated by the large language model after processing not only contain the location information of the "emergency room" but also have a clear speech intonation, facilitating the patient to understand; in a financial scenario, when the input text is "Buy 1000 shares of xx company's stock", the semantic features generated by the large language model after processing can clearly express the trading intention and at the same time have a clear speech rhythm, facilitating the trader to execute.

[0045] S203. Input the semantic features into a pre-trained acoustic model for feature conversion, and convert the semantic features into corresponding acoustic features.

[0046] In this embodiment, the semantic features output by the large language model are processed by a pre-trained acoustic model, and the semantics and prosody information in the semantic features are further refined into acoustic features. Acoustic features are high-level representations of audio signals, such as Mel spectrograms, etc. Converting semantic features into acoustic features prepares for subsequent speech waveform generation. For example, the emotion of "happy" in the semantic features may be converted into a high pitch and a fast rhythm in the acoustic features. Through the feature extraction and conversion of the acoustic model, the key information in the semantic features can be retained, improving the efficiency of subsequent processing.

[0047] For example, in a medical scenario, converting the semantic features of rehabilitation instructions such as "Please repeat this sentence" into acoustic features can generate speech that better guides patients in language rehabilitation training; in a financial scenario, after converting the semantic features of information such as "The stock market opened today" into acoustic features, the generated speech is more in line with the professional intonation of financial broadcasts, facilitating users to obtain information.

[0048] S204. Input the acoustic features into a pre-trained vocoder for decoding processing to generate the speech waveform corresponding to the original text.

[0049] In this embodiment, the acoustic features are decoded by a pre-trained vocoder. The specific vocoder can use an autoregressive decoder or a non-autoregressive decoder, etc., such as a GAN-based vocoder (such as HiFi-GAN, etc.). The information in the acoustic features is further refined into a speech waveform, thereby decoding and generating a speech waveform corresponding to the original text, and finally converting the text information into audible speech, realizing the complete conversion of text to speech based on the large language model, enabling efficient and smooth speech generation without having to train the model from scratch.

[0050] In the above embodiments, the present invention discloses a speech generation method based on a large language model. The method includes obtaining the original text of the speech to be generated; inputting the original text into a large language model with a hybrid LoRA adapter for text processing to generate semantic features including semantic information and prosody information; inputting the semantic features into a pre-trained acoustic model for feature conversion to convert the semantic features into corresponding acoustic features; and inputting the acoustic features into a pre-trained vocoder for decoding processing to generate the speech waveform corresponding to the original text. Among them, the large language model with a hybrid LoRA adapter is obtained through multi-stage parameter fine-tuning. The large language model with a hybrid LoRA adapter obtained through multi-stage parameter fine-tuning processes the input text, and after obtaining the features that integrate the text and speech modalities, speech generation is achieved. This effectively utilizes the prior knowledge of the large language model in understanding text to achieve the effective integration of these two modalities of text and speech, without the need to train from scratch, effectively improving the speech generation efficiency based on the large language model and reducing the requirements for computing hardware and resources.

[0051] In one embodiment, as Figure 3 shown, the large language model with a hybrid LoRA adapter is obtained through multi-stage parameter fine-tuning by the following steps:

[0052] S301. Collect a text dataset and a corresponding speech dataset, and divide the text dataset and speech dataset into a first-stage dataset, a second-stage dataset, and a third-stage dataset. Among them, both the first-stage dataset and the third-stage dataset contain text data and speech data, and the second-stage dataset contains text data;

[0053] S302. Perform one-stage fine-tuning on the large language model with an initialized LoRA adapter through the text data and speech data in the first-stage dataset;

[0054] S303. Perform second-stage fine-tuning on the LoRA adapter after one-stage fine-tuning through the text data in the second-stage dataset to obtain a hybrid LoRA adapter;

[0055] S304. Perform end-to-end speech generation training on the hybrid LoRA adapter, acoustic model, and vocoder through the text data and speech data in the third-stage dataset to obtain a large language model with a hybrid LoRA adapter.

[0056] In this embodiment, when performing model fine-tuning training, first collect the training data required for multi-stage fine-tuning, including a text dataset and a corresponding speech dataset. The text dataset contains the text data for training, such as sentences, paragraphs, or dialogue content, and the speech dataset is the speech recording corresponding to the text data, which is used to train the modules related to speech generation. And divide the dataset based on the objectives of different-stage fine-tuning. The first-stage dataset contains text and speech data, which is used to initially train the text-to-speech ability of the large language model. The second-stage dataset only contains text data, which is used to optimize the text processing ability of the large language model and avoid the negative impact of the speech generation task on the text understanding ability. The third-stage dataset contains text and speech data, which is used for end-to-end speech generation training to further optimize the speech generation effect of the entire system. By using different datasets in stages, each part of the speech generation system can be optimized more precisely.

[0057] Based on the datasets of each stage collected and divided, first use the text data and speech data in the first-stage dataset to perform one-stage fine-tuning on the large language model with an initialized LoRA adapter, so as to add the speech modality to the large language model, and make the large language model after one-stage fine-tuning initially have the ability to process text and speech modalities. After the first-stage fine-tuning is completed, the initialized LoRA adapter is overall trained into a TTS-LoRA adapter responsible for text-to-speech synthesis tasks. However, after this fine-tuning process, it may have a negative impact on the text understanding ability of the large language model. Therefore, during the second-stage fine-tuning, use the pure text data in the second-stage dataset to further optimize the LoRA adapter after one-stage fine-tuning through semantic understanding tasks. The adapter after two fine-tuning processes includes two parts: a TTS-LoRA adapter (for speech generation) and a text-LoRA adapter (for text processing), thus obtaining a hybrid LoRA adapter. The large language model with a hybrid LoRA adapter after two-stage fine-tuning can improve the model's understanding ability of text content, further optimize the model's text understanding ability, ensure that the performance of the model when processing text data is not affected by the first-stage training, and ensure that the large language model can perform well in both modalities.

[0058] In the last-stage fine-tuning, the hybrid LoRA adapter, acoustic model, and vocoder are trained as a whole to optimize the entire speech generation process. Use the text data and speech data in the third-stage dataset to perform end-to-end speech generation training on the hybrid LoRA adapter, acoustic model, and vocoder to further optimize the quality of speech generation, make the generated speech more natural and fluent, and obtain a large language model with a hybrid LoRA adapter.

[0059] In one embodiment, step S302 includes:

[0060] Load a pre-trained large language model and add an initialized LoRA adapter to the large language model;

[0061] Freeze the parameters of the acoustic model and the vocoder, and perform a speech generation task on the text data in the first-stage dataset through the large language model with the initialized LoRA adapter, the acoustic model with frozen parameters, and the vocoder to obtain the first-stage predicted speech;

[0062] Fine-tune the parameters of the text embedding layer, the output prediction layer, and the initialized LoRA adapter in the large language model according to the difference between the first-stage predicted speech and the speech data in the first-stage dataset.

[0063] In this embodiment, a pre-trained language model is loaded for the first-stage fine-tuning. The pre-trained large language model is a language model that has been trained with a large amount of text data, such as GPT, BERT, etc., and has general language understanding capabilities. An initialized LoRA adapter is added to the pre-trained large language model to introduce a new task, i.e., the speech generation task, into the large language model. The initialized LoRA adapter is inserted into a specific layer (such as the Transformer layer) of the large language model to enable it to adapt to the speech generation task. During the first-stage fine-tuning, the parameters of the acoustic model and the vocoder are first frozen, and their parameters are kept unchanged, that is, only the LoRA adapter in the large language model is trained, significantly reducing the computational resource requirements during the training process. The text data in the first-stage dataset is input into the large language model to generate semantic features, and then the acoustic model and the vocoder are used to generate the speech waveform to obtain the first-stage predicted speech. Then, the difference between the generated first-stage predicted speech and the real speech data (i.e., the corresponding speech data in the first-stage dataset) is calculated, which is measured by using a loss function (such as mean square error or cross entropy, etc.). According to the loss value, the parameters of the text embedding layer, the output prediction layer, and the LoRA adapter in the large language model are adjusted by backpropagation to reduce the gap between the predicted speech and the real speech. Through the first-stage fine-tuning, the speech modality is added to the large language model, so that the large language model with the LoRA adapter can better convert the text into semantic features containing semantic and prosodic information, and initially adapt to the speech generation task in the first stage, laying a foundation for the subsequent multi-stage training, thereby generating results closer to the real speech.

[0064] In one embodiment, step S303 includes:

[0065] Freeze the parameters of the text embedding layer and the output prediction layer in the large language model;

[0066] Perform a semantic understanding task on the text data in the second-stage dataset based on the LoRA adapter fine-tuned in the first stage, the text embedding layer with frozen parameters, and the output prediction layer to obtain a predicted semantic representation;

[0067] Fine-tune the parameters of the LoRA adapter fine-tuned in the first stage according to the difference between the predicted semantic representation and the text data in the second-stage dataset to obtain a hybrid LoRA adapter.

[0068] In this embodiment, after the goal of the first-stage fine-tuning d is completed, it may have a negative impact on the text understanding ability of the large language model. Therefore, during the second-stage fine-tuning, the parameters of the text embedding layer and the output prediction layer in the large language model are frozen. The frozen text embedding layer and output prediction layer are used, combined with the LoRA adapter fine-tuned in the first stage, to process the text data in the second-stage dataset to generate corresponding predicted semantic representations. By freezing the trained text embedding layer and output prediction layer in the first stage and only fine-tuning the LoRA adapter, the LoRA adapter is fine-tuned for the semantic understanding task based on the text dataset, the difference between the predicted semantic representation and the real text data is calculated, and the parameters of the LoRA adapter are updated according to the result of the difference evaluation such as cross-entropy loss to reduce the gap between the predicted semantic representation and the real text, thereby further optimizing the text understanding ability of the model. After the second-stage fine-tuning, the LoRA adapter is divided into two parts, where TTS-LoRA is responsible for the text-to-speech synthesis task to ensure that the model can generate high-quality speech waveforms, while text-LoRA is responsible for the text understanding task to ensure the performance of the model when processing text data, thus forming a hybrid LoRA adapter, making the large language model more stable and accurate when processing text and speech tasks.

[0069] In one embodiment, step S304 includes:

[0070] Freeze all parameters in the large language model except the hybrid LoRA adapter;

[0071] Perform a voice generation task on the text data in the third-stage dataset through the large language model with frozen parameters, the acoustic model to be trained, and the vocoder to obtain a three-stage predicted voice;

[0072] Adjust the parameters of the hybrid LoRA adapter, the acoustic model to be trained, and the vocoder according to the difference between the three-stage predicted voice and the voice data in the third-stage dataset to obtain a large language model with a hybrid LoRA adapter and a trained acoustic model and vocoder.

[0073] In this embodiment, during the fine-tuning of the last stage, all parameters in the large language model except for the hybrid LoRA adapter are frozen. That is, in the subsequent fine-tuning process, only the parameters of the hybrid LoRA adapter will be updated, while the other parts remain unchanged. This focuses on optimizing the hybrid LoRA adapter to better adapt to the speech generation task while maintaining the large language model's ability in text processing. The text data in the third-stage dataset is input into the large language model with frozen parameters (including the hybrid LoRA adapter), the acoustic model to be trained, and the vocoder, and the speech generation task is executed to obtain the three-stage predicted speech, thereby evaluating the performance of the entire system on the third-stage dataset and ensuring the collaborative working effect among the large language model, the acoustic model, and the vocoder. By comparing the differences between the generated three-stage predicted speech and the real speech data in the third-stage dataset, the parameters of the hybrid LoRA adapter, the acoustic model, and the vocoder are adjusted to reduce the gap between the predicted speech and the real speech, obtaining a large language model with a hybrid LoRA adapter and a trained acoustic model and vocoder. Through partial fine-tuning training in three stages instead of additional pre-training or full fine-tuning, the effective fusion of the text and speech modalities is achieved, enabling the efficient conversion of text into high-quality speech signals and significantly reducing the demand for computing resources.

[0074] In one embodiment, the large language model at least includes a text embedding layer, a Transformer layer with a hybrid LoRA adapter, and an output prediction layer. As Figure 4 shown, step S202 includes:

[0075] S401. Input the original text into the large language model with a hybrid LoRA adapter, and perform word embedding and position encoding processing on the original text through the text embedding layer to generate a text embedding sequence;

[0076] S402. Extract features of the self-attention mechanism from the text embedding sequence through the Transformer layer with a hybrid LoRA adapter to obtain a semantic and prosody encoding sequence;

[0077] S403. Map the semantic and prosody encoding sequence to a vector space adapted to the acoustic model through the output prediction layer to obtain semantic features containing semantic information and prosody information.

[0078] In this embodiment, the original text is input into the large language model and first processed through the text embedding layer. First, word embedding is performed on the original text, that is, each word or character in the original text is converted into a vector of a fixed dimension. For example, pre-trained word embeddings (such as Word2Vec, GloVe) are used to map words to vectors. Then, position encoding is used to add position information to each word embedding, enabling the model to capture the sequential relationship in the text. For example, sine and cosine functions are used to generate position encoding and added to the word embedding. The processed word embedding and position encoding are combined to form a text embedding sequence, which is used as the input to the subsequent Transformer layer. For example, for the input text "The patient's body temperature is 38.5°C, and the main complaint is headache", after word embedding and position encoding processing, the generated text embedding sequence can retain the semantic information of keywords such as "body temperature" and "headache" and their positional relationship in the sentence. Another example, for the input text "Your credit card due date is the 5th of each month", the generated text embedding sequence can retain the semantic information of keywords such as "credit card" and "due date" and their positional relationship.

[0079] After that, the text embedding sequence is input into the Transformer layer with a hybrid LoRA adapter, and feature extraction is performed through the self-attention mechanism, which can dynamically calculate the weights of each position in the input sequence and capture global dependencies. For example, the model can automatically focus on the relationship between "38.5°C" in "The patient's body temperature is 38.5°C" and "headache". The hybrid LoRA adapter added to the Transformer layer introduces a low-rank matrix to fine-tune the model parameters, enabling it to handle both the semantic information and prosodic information of the text simultaneously. For example, the hybrid LoRA adapter can enhance the model's ability to capture the semantics and intonation in "the main complaint is headache", so that after being processed by the self-attention mechanism and the hybrid LoRA adapter, the generated encoded sequence not only contains the semantic information of the text but also incorporates prosodic information, resulting in a semantic and prosodic encoded sequence, providing richer features for subsequent speech generation. For example, for the input text "The patient's body temperature is 38.5°C, and the main complaint is headache", after being processed by the Transformer layer and the hybrid LoRA adapter, the generated encoded sequence not only contains the semantic information of "body temperature" and "headache", but also incorporates appropriate intonation and rhythm, making the generated speech more natural.

[0080] After that, the semantic and prosodic encoded sequence is input into the output prediction layer, and the output prediction layer converts the semantic and prosodic encoded sequence into a vector form suitable for the input of the acoustic model, that is, maps it to the vector space suitable for the acoustic model. For example, the encoded sequence is mapped to the vector space of the Mel spectrum so that the acoustic model can further process it. Finally, the generated semantic features not only contain the semantic information of the text but also incorporate prosodic information, providing reliable input data for subsequent speech generation.

[0081] In one embodiment, as Figure 5 shown, step S203 includes:

[0082] S501. Input the semantic feature into a pre-trained acoustic model to perform temporal encoding on the semantic feature, obtaining corresponding temporal information;

[0083] S502. Perform context modeling processing on the semantic feature according to the temporal information, obtaining corresponding acoustic features.

[0084] In this embodiment, when the acoustic model converts the semantic feature, it first performs temporal processing on the semantic feature to capture the changes of the semantic feature in the time dimension. For example, it identifies the temporal relationship between "body temperature" and "38.5°C" in "The patient's body temperature is 38.5°C and the main complaint is headache". The output temporal information reflects the dynamic changes of the semantic feature in time, ensuring that the generated acoustic features are coherent in time, conform to the rhythm and intonation of natural language, providing a basis for subsequent context modeling, and making the generated speech more natural.

[0085] After that, the temporal information is used to further process the semantic feature, and the context relationship between the semantic features is captured to generate corresponding acoustic features. For example, the model can identify the semantic association between "body temperature" and "headache" in "The patient's body temperature is 38.5°C and the main complaint is headache", and generate acoustic features through context modeling, which can clearly express the semantic relationship between "body temperature" and "headache", further enhancing the expression ability of the semantic feature, making the generated acoustic features closer to the features of real speech, and making the generated speech more natural.

[0086] It should be noted that there is not necessarily a certain sequence among the above steps. Those of ordinary skill in the art can understand according to the description of the embodiments of the present invention that in different embodiments, the above steps can have different execution sequences, that is, they can be executed in parallel or exchanged, etc.

[0087] Further referring to Figure 6 , as an implementation of the above Figure 2 shown method, the present invention provides an embodiment of a speech generation device based on a large language model. This device embodiment corresponds to the Figure 2 shown method embodiment, and this device can be specifically applied to various electronic devices.

[0088] As Figure 6 shown, the speech generation device 60 based on a large language model described in this embodiment includes:

[0089] An acquisition module 601, configured to acquire the original text of the speech to be generated;

[0090] A text processing module 602 is configured to input the original text into a large language model with a hybrid LoRA adapter for text processing, and generate semantic features including semantic information and prosody information.

[0091] An acoustic processing module 603 is configured to input the semantic features into a pre-trained acoustic model for feature conversion, and convert the semantic features into corresponding acoustic features.

[0092] A speech decoding module 604 is configured to input the acoustic features into a pre-trained vocoder for decoding processing, and generate a speech waveform corresponding to the original text.

[0093] Wherein, the large language model with a hybrid LoRA adapter is obtained through multi-stage parameter fine-tuning.

[0094] The module referred to in the present invention means a series of computer program instruction segments that can complete specific functions, which is more suitable for describing the execution process of speech generation based on a large language model. For the specific implementation manners of each module, please refer to the corresponding method embodiments above, and details are not described herein again.

[0095] In one embodiment, the device 60 further includes:

[0096] An acquisition and division module is configured to acquire a text data set and a corresponding speech data set, and divide the text data set and the speech data set into a first-stage data set, a second-stage data set, and a third-stage data set. Wherein, both the first-stage data set and the third-stage data set include text data and speech data, and the second-stage data set includes text data.

[0097] A first fine-tuning module is configured to perform first-stage fine-tuning on a large language model with an initialized LoRA adapter through the text data and speech data in the first-stage data set.

[0098] A second fine-tuning module is configured to perform second-stage fine-tuning on the LoRA adapter after the first-stage fine-tuning through the text data in the second-stage data set to obtain a hybrid LoRA adapter.

[0099] A third fine-tuning module is configured to perform end-to-end speech generation training on the hybrid LoRA adapter, the acoustic model, and the vocoder through the text data and speech data in the third-stage data set to obtain a large language model with a hybrid LoRA adapter.

[0100] In one embodiment, the first fine-tuning module includes:

[0101] A model loading unit is configured to load a pre-trained large language model and add an initialized LoRA adapter to the large language model.

[0102] A first task execution unit, configured to freeze the parameters of the acoustic model and the vocoder, and perform a speech generation task on the text data in the first-stage dataset through a large language model with an initialized LoRA adapter, the acoustic model with frozen parameters, and the vocoder, to obtain a first-stage predicted speech;

[0103] A first parameter fine-tuning unit, configured to fine-tune the parameters of the text embedding layer, the output prediction layer, and the initialized LoRA adapter in the large language model according to the difference between the first-stage predicted speech and the speech data in the first-stage dataset.

[0104] In one embodiment, the second fine-tuning module includes:

[0105] A first parameter freezing unit, configured to freeze the parameters of the text embedding layer and the output prediction layer in the large language model;

[0106] A second task execution unit, configured to perform a semantic understanding task on the text data in the second-stage dataset based on the LoRA adapter fine-tuned in the first stage, the text embedding layer with frozen parameters, and the output prediction layer, to obtain a predicted semantic representation;

[0107] A second parameter fine-tuning unit, configured to fine-tune the parameters of the LoRA adapter fine-tuned in the first stage according to the difference between the predicted semantic representation and the text data in the second-stage dataset, to obtain a hybrid LoRA adapter.

[0108] In one embodiment, the third fine-tuning module includes:

[0109] A second parameter freezing unit, configured to freeze all parameters in the large language model except the hybrid LoRA adapter;

[0110] A third task execution unit, configured to perform a speech generation task on the text data in the third-stage dataset through the large language model with frozen parameters, the acoustic model to be trained, and the vocoder, to obtain a third-stage predicted speech;

[0111] A third parameter fine-tuning unit, configured to adjust the parameters of the hybrid LoRA adapter, the acoustic model to be trained, and the vocoder according to the difference between the third-stage predicted speech and the speech data in the third-stage dataset, to obtain a large language model with a hybrid LoRA adapter, and a trained acoustic model and vocoder.

[0112] In one embodiment, the large language model includes at least a text embedding layer, a Transformer layer with a hybrid LoRA adapter, and an output prediction layer, and the text processing module 602 includes:

[0113] A text encoding unit for inputting the original text into a large language model with a hybrid LoRA adapter, performing word embedding and position encoding processing on the original text through a text embedding layer, and generating a text embedding sequence;

[0114] A feature extraction unit for performing feature extraction of the self-attention mechanism on the text embedding sequence through the Transformer layer with the hybrid LoRA adapter to obtain a semantic and prosody encoding sequence;

[0115] A feature mapping unit for mapping the semantic and prosody encoding sequence to a vector space adapted to the acoustic model through the output prediction layer to obtain semantic features containing semantic information and prosody information.

[0116] In one embodiment, the acoustic processing module 603 includes:

[0117] A timing encoding unit for inputting the semantic features into a pre-trained acoustic model, performing timing encoding on the semantic features, and obtaining corresponding timing information;

[0118] A modeling conversion unit for performing context modeling processing on the semantic features according to the timing information to obtain corresponding acoustic features.

[0119] In the above embodiment, the present invention discloses a speech generation device based on a large language model, which obtains the original text of the speech to be generated; inputs the original text into a large language model with a hybrid LoRA adapter for text processing to generate semantic features containing semantic information and prosody information; inputs the semantic features into a pre-trained acoustic model for feature conversion to convert the semantic features into corresponding acoustic features; inputs the acoustic features into a pre-trained vocoder for decoding processing to generate a speech waveform corresponding to the original text; wherein, the large language model with a hybrid LoRA adapter is obtained through multi-stage parameter fine-tuning. The large language model with a hybrid LoRA adapter obtained through multi-stage parameter fine-tuning processes the input text, and after obtaining features that integrate text and speech modalities, realizes speech generation, effectively utilizes the prior knowledge of the large language model in understanding text to effectively integrate these two modalities of text and speech, does not require training from scratch, effectively improves the speech generation efficiency based on the large language model, and reduces the requirements for computing hardware and resources.

[0120] Another embodiment of the present invention provides a computer device, as Figure 7 shown, the computer device 70 includes:

[0121] One or more processors 701 and a memory 702,Figure 7 Taking a processor 701 as an example for introduction, the processor 701 and the memory 702 can be connected through a bus or other means. Figure 7 Taking the connection through a bus as an example.

[0122] The processor 701 is used to complete various control logics of the computer device 70. It can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISCMachine), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of these components. Additionally, the processor 701 can also be any conventional processor, microprocessor, or state machine. The processor 701 can also be implemented as a combination of computing devices. For example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP and / or any other such configuration.

[0123] The memory 702, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions corresponding to the speech generation method based on a large language model in the embodiments of the present invention. The processor 701 executes various functional applications and data processing of the computer device 70 by running the non-volatile software programs, instructions, and units stored in the memory 702, that is, implements the speech generation method based on a large language model in the above method embodiments.

[0124] The memory 702 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device 70, etc. In addition, the memory 702 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 702 optionally includes a memory remotely set relative to the processor 701, and these remote memories can be connected to the computer device 70 through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. One or more units are stored in the memory 702 and, when executed by one or more processors 701, perform the steps of the speech generation method based on a large language model in any of the above method embodiments.

[0125] In the above embodiments, the present invention discloses a computer device, which obtains the original text of the speech to be generated; inputs the original text into a large language model with a hybrid LoRA adapter for text processing to generate semantic features containing semantic information and prosody information; inputs the semantic features into a pre-trained acoustic model for feature conversion to convert the semantic features into corresponding acoustic features; inputs the acoustic features into a pre-trained vocoder for decoding processing to generate the speech waveform corresponding to the original text; wherein, the large language model with a hybrid LoRA adapter is obtained through multi-stage parameter fine-tuning. The large language model with a hybrid LoRA adapter obtained through multi-stage parameter fine-tuning processes the input text, and after obtaining the features that integrate the text and speech modalities, realizes speech generation, effectively utilizes the prior knowledge of the large language model in understanding text to effectively integrate the two modalities of text and speech, does not need to be trained from scratch, effectively improves the speech generation efficiency based on the large language model, and reduces the requirements for computing hardware and resources.

[0126] An embodiment of the present invention provides a non-volatile computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are executed by one or more processors, the steps of the speech generation method based on a large language model in any of the above method embodiments are executed.

[0127] In the above embodiments, the present invention discloses a non-volatile computer-readable storage medium, which obtains the original text of the speech to be generated; inputs the original text into a large language model with a hybrid LoRA adapter for text processing to generate semantic features containing semantic information and prosody information; inputs the semantic features into a pre-trained acoustic model for feature conversion to convert the semantic features into corresponding acoustic features; inputs the acoustic features into a pre-trained vocoder for decoding processing to generate the speech waveform corresponding to the original text; wherein, the large language model with a hybrid LoRA adapter is obtained through multi-stage parameter fine-tuning. The large language model with a hybrid LoRA adapter obtained through multi-stage parameter fine-tuning processes the input text, and after obtaining the features that integrate the text and speech modalities, realizes speech generation, effectively utilizes the prior knowledge of the large language model in understanding text to effectively integrate the two modalities of text and speech, does not need to be trained from scratch, effectively improves the speech generation efficiency based on the large language model, and reduces the requirements for computing hardware and resources.

[0128] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0129] The present invention can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0130] In summary, in the method, device, equipment, and medium for speech generation based on a large language model disclosed in the present invention, the method includes: obtaining the original text of the speech to be generated; inputting the original text into a large language model with a hybrid LoRA adapter for text processing to generate semantic features including semantic information and prosody information; inputting the semantic features into a pre-trained acoustic model for feature conversion to convert the semantic features into corresponding acoustic features; inputting the acoustic features into a pre-trained vocoder for decoding processing to generate a speech waveform corresponding to the original text; wherein, the large language model with a hybrid LoRA adapter is obtained through multi-stage parameter fine-tuning. The large language model with a hybrid LoRA adapter obtained through multi-stage parameter fine-tuning processes the input text, and after obtaining features that integrate the text and speech modalities, realizes speech generation. It effectively utilizes the prior knowledge of the large language model in understanding text to effectively integrate the two modalities of text and speech, without having to train from scratch, effectively improving the speech generation efficiency based on the large language model and reducing the requirements for computing hardware and resources.

[0131] Of course, those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing related hardware (such as a processor, a controller, etc.) through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium, and when the computer program is executed, it can include the processes of the above method embodiments. The storage medium can be a memory, a magnetic disk, a floppy disk, a flash memory, an optical memory, etc.

[0132] It should be noted that if there are software tools or components of other companies in the embodiments of the present application, they are only used for illustrative introduction and do not represent actual use. It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A speech generation method based on a large language model, characterized in that: include: Obtain the original text of the speech to be generated; Inputting the original text into a large language model with a hybrid LoRA adapter for text processing to generate semantic features containing semantic information and prosodic information; Inputting the semantic features into a pre-trained acoustic model for feature conversion, converting the semantic features into corresponding acoustic features; Inputting the acoustic features into a pre-trained vocoder for decoding to generate a speech waveform corresponding to the original text; The large language model with the hybrid LoRA adapter is obtained by multi-stage parameter fine-tuning.

2. The speech generation method based on a large language model according to claim 1, characterized in that: The large language model with hybrid LoRA adapter is obtained by multi-stage parameter fine-tuning through the following steps: Collecting a text data set and a corresponding voice data set, dividing the text data set and the voice data set into a first-stage data set, a second-stage data set and a third-stage data set, wherein the first-stage data set and the third-stage data set both contain text data and voice data, and the second-stage data set contains text data; Using the text and speech data in the first phase of the dataset, the large language model with the initialized LoRA adapter is fine-tuned in one phase; Through the text data in the second-stage data set, the LoRA adapter that has undergone the first-stage fine-tuning is fine-tuned in the second stage to obtain a hybrid LoRA adapter; The hybrid LoRA adapter, acoustic model and vocoder are trained for end-to-end speech generation using the text data and speech data in the third-stage data set to obtain a large language model with a hybrid LoRA adapter.

3. The speech generation method based on a large language model according to claim 2, characterized in that: The first-stage fine-tuning of the large language model with the initialized LoRA adapter is performed using the text data and voice data in the first-stage data set, including: Load a pre-trained large language model and add an initialized LoRA adapter to the large language model; Freeze the parameters of the acoustic model and the vocoder, and perform a speech generation task on the text data in the first-stage data set by using a large language model with an initialized LoRA adapter, an acoustic model with frozen parameters, and a vocoder to obtain a first-stage predicted speech; According to the difference between the first-stage predicted speech and the speech data in the first-stage dataset, parameters of the text embedding layer, the output prediction layer and the initialized LoRA adapter in the large language model are fine-tuned.

4. The speech generation method based on a large language model according to claim 2, characterized in that: The method of performing a second-stage fine-tuning on the LoRA adapter that has undergone the first-stage fine-tuning by using the text data in the second-stage data set to obtain a hybrid LoRA adapter includes: Freeze parameters of a text embedding layer and an output prediction layer in the large language model; Perform semantic understanding tasks on the text data in the second-stage dataset based on the fine-tuned LoRA adapter, the text embedding layer with frozen parameters, and the output prediction layer to obtain predicted semantic representations; According to the difference between the predicted semantic representation and the text data in the second-stage data set, the parameters of the LoRA adapter that has been fine-tuned in one stage are fine-tuned to obtain a hybrid LoRA adapter.

5. The speech generation method based on a large language model according to claim 2, characterized in that: The hybrid LoRA adapter, the acoustic model and the vocoder are trained for end-to-end speech generation through the text data and the speech data in the third stage data set to obtain a large language model with a hybrid LoRA adapter, including: Freezing all parameters in the large language model except the hybrid LoRA adapter; Performing a speech generation task on the text data in the third-stage data set by using a large language model with frozen parameters, an acoustic model to be trained, and a vocoder to obtain a three-stage predicted speech; According to the difference between the three-stage predicted speech and the speech data in the third-stage data set, the parameters of the hybrid LoRA adapter, the acoustic model to be trained, and the vocoder are adjusted to obtain a large language model with a hybrid LoRA adapter and a trained acoustic model and vocoder.

6. The speech generation method based on a large language model according to claim 1, characterized in that: The large language model at least comprises a text embedding layer, a Transformer layer with a hybrid LoRA adapter, and an output prediction layer. The original text is input into the large language model with a hybrid LoRA adapter for text processing to generate semantic features containing semantic information and prosodic information, including: Inputting the original text into a large language model with a hybrid LoRA adapter, performing word embedding and position encoding processing on the original text through a text embedding layer to generate a text embedding sequence; Performing feature extraction of the self-attention mechanism on the text embedding sequence through the Transformer layer with the hybrid LoRA adapter to obtain a semantic and prosodic coding sequence; The semantic and prosodic coding sequence is mapped to a vector space adapted to the acoustic model through the output prediction layer to obtain semantic features containing semantic information and prosodic information.

7. The speech generation method based on a large language model according to claim 1, characterized in that: The step of inputting the semantic features into a pre-trained acoustic model for feature conversion, and converting the semantic features into corresponding acoustic features, includes: Inputting the semantic features into a pre-trained acoustic model, performing temporal encoding on the semantic features, and obtaining corresponding temporal information; Context modeling is performed on the semantic features according to the time series information to obtain corresponding acoustic features.

8. A speech generation device based on a large language model, characterized in that: include: An acquisition module, used to acquire the original text of the speech to be generated; A text processing module, used for inputting the original text into a large language model with a hybrid LoRA adapter for text processing, and generating semantic features containing semantic information and prosodic information; An acoustic processing module, used for inputting the semantic features into a pre-trained acoustic model for feature conversion, and converting the semantic features into corresponding acoustic features; A speech decoding module, used for inputting the acoustic features into a pre-trained vocoder for decoding processing to generate a speech waveform corresponding to the original text; The large language model with the hybrid LoRA adapter is obtained by multi-stage parameter fine-tuning.

9. A computer device, characterized in that: comprising at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the speech generation method based on a large language model as described in any one of claims 1-7.

10. A non-volatile computer-readable storage medium, characterized in that: The non-volatile computer-readable storage medium stores computer-executable instructions, which, when executed by one or more processors, enable the one or more processors to execute the speech generation method based on a large language model as described in any one of claims 1-7.

Citation Information

Cited By

  • Speech recognition method and system based on large language model

    CN120783764A