Speech synthesis method and device based on large language model, equipment and storage medium

CN121214905BActive Publication Date: 2026-06-23PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2025-09-18
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing speech synthesis systems based on large language models suffer from problems such as low naturalness of speech, loss of emotional features, and difficulty in processing technical terms in the fields of fintech and healthcare. As a result, the accuracy of speech synthesis is insufficient and it is difficult to meet the needs of high-value and high-sensitivity applications.

Method used

By constructing triplet and quadruplet datasets, a large language model is trained using a speech encoder and projector to achieve cross-modal alignment between speech and text, preserve continuous speech features and enhance the parsing ability of emotions and professional terms. A staged training strategy is adopted to optimize model parameters and ensure the accuracy of speech synthesis.

Benefits of technology

It improves the accuracy and naturalness of speech synthesis, enabling accurate understanding and generation of speech responses that conform to emotions and context in multiple dialects and professional fields, thereby enhancing the professionalism and naturalness of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121214905B_ABST
    Figure CN121214905B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of artificial intelligence, which can be applied to the medical field and the financial technology field, and discloses a speech synthesis method and device based on a large language model, equipment and a storage medium, wherein the method comprises: obtaining initial data, and generating a triple data set by data construction and data preprocessing on the initial data; performing speech coding and feature projection on the triple data set to obtain speech query embedding; training a speech encoder and a projector in a speech synthesis model to obtain a target speech encoder and a target projector; generating a quadruple data set by data construction and data labeling on the initial data; constructing a training sample sequence based on the quadruple data set, and training a large language model in the speech synthesis model to generate a target speech synthesis model; obtaining query information input by a user, and generating a target speech response by speech synthesis processing based on the query information through the target speech synthesis model. The present application improves the accuracy of speech synthesis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology and can be applied to the medical and financial technology fields. In particular, it relates to a speech synthesis method, device, equipment and storage medium based on a large language model. Background Technology

[0002] In recent years, breakthroughs in large-scale language model technology have brought a new paradigm to speech synthesis systems: using large-scale language models to generate discrete speech tokens to drive vocoders to synthesize speech. This speech synthesis method based on large-scale language models has shown great potential in terms of speech naturalness. However, when deploying existing speech synthesis systems based on large-scale language models in high-value, high-sensitivity vertical sectors such as fintech and healthcare, their inherent shortcomings are further amplified, seriously hindering their practical application. In the financial sector, scenarios such as intelligent customer service, compliance dual recording, and telephone risk control urgently require voice capabilities that can perform highly expressive, multi-dialect, and human-like interactions. When processing voice prompts, existing systems typically need to quantify them into discrete tokens, resulting in the significant loss of emotional features such as anxiety, confusion, and dissatisfaction in the customer's voice. In the medical field, the requirements for voice accuracy, emotional dimension, and privacy are extremely high. When doctors use voice to enter medical records, existing systems struggle to perfectly handle complex medical terminology and sentence structures, and misunderstandings may lead to serious medical accidents.

[0003] Existing speech synthesis methods utilize LLM-driven TTS systems to achieve speech synthesis. This approach processes input by quantizing voice prompts into discrete tokens. This method may result in the loss of subtle timbre and emotional features in the original speech, thereby affecting the naturalness and expressiveness of the generated speech and leading to low accuracy in speech synthesis. Summary of the Invention

[0004] The purpose of this application is to provide a speech synthesis method, apparatus, device, and storage medium based on a large language model to improve the accuracy of speech synthesis.

[0005] To address the aforementioned technical problems, embodiments of this application provide a speech synthesis method based on a large language model, comprising:

[0006] Acquire initial data, and perform data construction and data preprocessing on the initial data to generate a triplet dataset, wherein the initial data includes multilingual speech and text pairing data, sentiment tag data, and TTS synthesized speech data;

[0007] The triplet dataset is subjected to speech encoding and feature projection to obtain the speech query embedding;

[0008] The speech encoder and projector in the speech synthesis model are trained based on the speech query embedding and the triple dataset to obtain the target speech encoder and target projector.

[0009] The initial data is processed by data construction and data labeling to generate a quadruple dataset;

[0010] A training sample sequence is constructed based on the quadruple dataset, and the large language model in the speech synthesis model is trained based on the training sample sequence to generate the target speech synthesis model.

[0011] The system obtains the query information input by the user and performs speech synthesis processing based on the query information using the target speech synthesis model to generate a target speech response.

[0012] To address the aforementioned technical problems, embodiments of this application provide a speech synthesis device based on a large language model, comprising:

[0013] The triplet dataset construction module is used to acquire initial data and perform data construction and data preprocessing on the initial data to generate a triplet dataset. The initial data includes multilingual speech and text pairing data, sentiment tag data, and TTS synthesized speech data.

[0014] The speech query embedding module is used to perform speech encoding and feature projection on the triplet dataset to obtain the speech query embedding;

[0015] The first model training module is used to train the speech encoder and projector in the speech synthesis model based on the speech query embedding and the triplet dataset to obtain the target speech encoder and target projector.

[0016] The quadruple dataset construction module is used to construct and label the initial data to generate a quadruple dataset.

[0017] The second model training module is used to construct a training sample sequence based on the quadruple dataset, and to train the large language model in the speech synthesis model based on the training sample sequence to generate the target speech synthesis model.

[0018] The target speech response generation module is used to obtain query information input by the user and perform speech synthesis processing based on the query information through the target speech synthesis model to generate a target speech response.

[0019] To solve the above-mentioned technical problems, one technical solution adopted by the present invention is to provide a computer device, including one or more processors; and a memory for storing one or more programs, such that the one or more processors implement the speech synthesis method based on a large language model as described above.

[0020] To solve the above-mentioned technical problems, one technical solution adopted by the present invention is: a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the speech synthesis method based on a large language model as described above.

[0021] This invention provides a speech synthesis method, apparatus, device, and storage medium based on a large language model. The method includes: acquiring initial data and performing data construction and preprocessing on the initial data to generate a triplet dataset, wherein the initial data includes multilingual speech and text pairing data, sentiment tag data, and TTS synthesized speech data; performing speech encoding and feature projection on the triplet dataset to obtain a speech query embedding; training a speech encoder and projector in a speech synthesis model based on the speech query embedding and the triplet dataset to obtain a target speech encoder and a target projector; performing data construction and data labeling on the initial data to generate a quadruple dataset; constructing a training sample sequence based on the quadruple dataset and training the large language model in the speech synthesis model based on the training sample sequence to generate a target speech synthesis model; acquiring user-input query information and performing speech synthesis processing on the target speech synthesis model based on the query information to generate a target speech response. This invention, through the construction of a multimodal training dataset and a phased model training strategy, achieves continuous feature encoding and cross-modal alignment of speech prompts while preserving the core understanding capabilities of large language models. This solves the problems of speech prompt quantization loss, difficulty in adapting professional domain data, and insufficient semantic reliability in existing technologies, and is conducive to improving the accuracy of speech synthesis. Attached Figure Description

[0022] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram of an application environment for a speech synthesis method based on a large language model according to an embodiment of the present invention;

[0024] Figure 2 This is a flowchart illustrating the implementation of the speech synthesis method based on a large language model provided in this application embodiment;

[0025] Figure 3 yes Figure 2 A flowchart illustrating a specific implementation method of step S1;

[0026] Figure 4 yes Figure 2 A flowchart illustrating a specific implementation method of step S2;

[0027] Figure 5 yes Figure 2 A flowchart illustrating a specific implementation method of step S3;

[0028] Figure 6 yes Figure 2 A flowchart illustrating a specific implementation of step S4;

[0029] Figure 7 yes Figure 2 A flowchart illustrating a specific implementation of step S5;

[0030] Figure 8 yes Figure 2 A flowchart illustrating a specific implementation of step S6;

[0031] Figure 9 This is a schematic diagram of a speech synthesis device based on a large language model provided in an embodiment of this application;

[0032] Figure 10 This is a schematic diagram of the computer device provided in the embodiments of this application. Detailed Implementation

[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0034] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0035] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0036] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0037] It should be noted that the speech synthesis method based on a large language model provided in this application is generally executed by a server, and correspondingly, the speech synthesis device based on a large language model is generally configured in the server.

[0038] The speech synthesis method based on a large language model provided in this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can receive query information from the client and generate a target voice response based on the query information. In this invention, the server sends the target voice response to the client. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will now be described in detail through specific embodiments.

[0039] The speech synthesis method based on a large language model provided in this application can be applied to medical application scenarios, such as human-like medical assistants and chronic disease management scenarios, doctor-patient communication simulation and training scenarios, etc.; this application can also be applied to financial technology application scenarios, such as hyper-human-like intelligent customer service and financial advisor scenarios, intelligent risk control scenarios based on voice emotion analysis, etc.

[0040] Please see Figure 2 , Figure 2 This paper illustrates a specific implementation of a speech synthesis method based on a large language model.

[0041] It should be noted that if substantially the same result is obtained, the method of this invention is not based on... Figure 2 Limited to the order of the processes shown, this method includes the following steps:

[0042] S1: Obtain initial data, and perform data construction and data preprocessing on the initial data to generate a triplet dataset, wherein the initial data includes multilingual speech and text pairing data, sentiment tag data, and TTS synthesized speech data.

[0043] Specifically, speech-text pairs are extracted from an ASR (Automatic Speech Recognition) corpus; a TTS (Text-to-Speech) system is used to synthesize speech from text generated by an LLM (Large Language Model), constructing text-speech pairs; and descriptive tags (such as dialect and emotion) are added to the text. The triplet dataset consists of speech queries, text transcriptions, and subsequent text. Resampling and normalization are used to ensure acoustic feature consistency, while text cleaning and unified encoding ensure semantic coherence.

[0044] Please see Figure 3 , Figure 3 A specific implementation of step S1 is shown below:

[0045] S11: Obtain the initial data and construct it into triplet format data, wherein the triplet format data includes speech query, text transcription, and subsequent text data. S12: Resample the speech query data in the triplet format data and normalize the resampled speech query data to obtain normalized speech data. S13: Perform text cleaning and unified encoding on the text transcription data and subsequent text data in the triplet format data to obtain encoded text transcription data and subsequent text data. S14: Convert the encoded text transcription data and subsequent text data into word identifier sequences using the word segmenter of the large language model. S15: Construct the triplet dataset based on the normalized speech data and the word identifier sequences.

[0046] Specifically, by constructing a triplet data structure containing speech, text, and context, the speech is resampled and normalized during the data preprocessing stage to eliminate the influence of differences in acquisition devices on acoustic features, while preserving the prosodic characteristics of dialect speech. Text data is cleaned and uniformly encoded to effectively handle complex scenarios where financial terminology and dialect vocabulary are mixed; for example, medical terms such as "myocardial infarction" are uniformly encoded with dialect expressions such as "what are you saying?". A word segmentation sequence is generated using a large language model, mapping text and speech features to a unified semantic space, enabling the model to learn the correspondence between emotional features in speech and semantics in text. The resulting triplet dataset contains both continuous acoustic features and semantic information, providing cross-modal aligned training samples for the model. This application effectively solves the problem of emotional feature loss caused by speech prompt quantization, preserving the prosodic and emotional information of the original speech through non-destructive preprocessing. The constructed triplet data structure enhances the cross-modal association between speech and text, enabling the model to accurately understand regional features and emotional expressions in multi-dialect speech. Unified text encoding and word segmentation ensure accurate parsing of technical terms and dialect words, providing high-quality cross-modal alignment data for subsequent model training and significantly improving the adaptability of speech synthesis systems in multi-dialect scenarios.

[0047] Among them, triplet format data refers to structured data containing speech queries, corresponding text transcriptions, and subsequent text. Specifically, it can be achieved by combining speech dialogue records with corresponding manually annotated transcribed text and context text. This structure can establish cross-modal associations between speech and text. Resampling refers to the process of unifying speech with different sampling rates to a standard sampling rate. This can be achieved using linear interpolation or anti-aliasing filters to eliminate the influence of device differences on acoustic features. Normalization refers to amplitude standardization of the speech signal. This can be achieved using maximum peak value normalization or quantile normalization methods to preserve the prosodic features of the original speech. Text cleaning refers to removing noisy characters and redundant information from the text. This can be achieved using regular expression matching or rule-based filtering algorithms to ensure the accuracy of technical terms and dialect vocabulary. Unified encoding refers to converting multilingual text into a unified character set. This can be achieved using the UTF-8 encoding standard to solve the problem of multilingual mixed processing. Token identification sequence refers to segmenting the text into discrete units that can be processed by the language model. This can be achieved using the BPE algorithm or the WordPiece word segmenter to establish semantic alignment between text and speech.

[0048] S2: Perform speech encoding and feature projection on the triplet dataset to obtain the speech query embedding.

[0049] Specifically, the speech query embedding is generated through continuous acoustic embedding vector mapping, uses a convolutional neural network to capture local acoustic patterns, and utilizes linear projection to align the text semantic space.

[0050] Please see Figure 4 , Figure 4 A specific implementation of step S2 is shown below:

[0051] S21: The speech encoder in the speech synthesis model extracts features from the speech data in the triplet dataset to generate a continuous acoustic embedding vector. S22: The projector in the speech synthesis model performs acoustic local acoustic pattern capture and feature dimension mapping based on the continuous acoustic embedding vector to generate the speech query embedding.

[0052] Specifically, the speech encoder performs frame-level feature extraction on the input speech, generating a continuous acoustic embedding vector containing time-frequency characteristics. For example, it uses an 80-dimensional Mel-frequency spectrum as the basic feature and extracts deep acoustic representations through a multi-layer convolutional network. The projector performs local pattern analysis on the continuous acoustic embedding vector, for example, by capturing intonation transitions or emotional stress regions through a sliding window mechanism, and then adjusts the feature dimension to the same 768-dimensional space as the text word embedding through linear transformation. This processing method not only preserves the complete acoustic information of the original speech, but also optimizes the interaction efficiency between speech features and the large language model through feature space alignment. This application effectively solves the problem of loss of paralinguistic information in the process of voice prompt quantification. In financial customer service scenarios, it can accurately capture emotional features such as anxiety and doubt in customer speech, and in medical scenarios, it can completely preserve physiological indicators such as weakness and pain in patient speech. By capturing local acoustic patterns, it improves the recognition ability of unique pronunciation features in dialect speech, such as the accurate extraction of the entering tone value of Cantonese, making the generated response speech more in line with regional language habits.

[0053] The speech encoder is a computational module used to extract continuous acoustic features from the raw speech signal; Whisper-Small can be used as a speech encoder. The projector is a neural network layer used for feature space transformation; it can be implemented using a self-attention mechanism in conjunction with a fully connected layer. Its function is to capture local acoustic patterns in speech and map high-dimensional acoustic features to a dimension that matches the word embedding space of a large language model.

[0054] S3: Train the speech encoder and projector in the speech synthesis model based on the speech query embedding and the triple dataset to obtain the target speech encoder and target projector.

[0055] Specifically, the speech synthesis model includes a speech encoder, a projector, and a large language model. The embodiments of this application are for training the large language model to achieve modal alignment, thereby generating a trained speech encoder and projector. After training, the speech synthesis model can convert any input speech into semantic embeddings that the LLM can "understand," achieving deep alignment between speech and text modalities.

[0056] Please see Figure 5 , Figure 5 A specific implementation of step S3 is shown below:

[0057] S31: Obtain the text transcription lexical and subsequent text lexical corresponding to the text transcription and subsequent text in the triplet dataset. S32: Concatenate the speech query embedding with the embedding vector of the text transcription lexical to generate an initial input sequence. S33: Freeze the large language model in the speech synthesis model to obtain an initial frozen speech synthesis model. Based on the initial input sequence, predict and generate the next lexical using the initial frozen speech synthesis model to obtain the initial lexical. Sample the cross-entropy loss function to calculate a first loss value based on the initial lexical and the subsequent text lexical. S34: Update the parameters of the speech encoder and projector in the initial frozen speech synthesis model using backpropagation based on the first loss value, to iteratively train the speech encoder and projector in the initial frozen speech synthesis model again, generating the target speech encoder and the target projector.

[0058] Specifically, this method optimizes the speech encoding module through a phased parameter update strategy. First, it extracts the word sequences of the text transcription and subsequent text from the triplet data, concatenating the speech query embedding with the text transcription word embedding to form a joint input sequence. After freezing the parameters of the large language model, the speech encoder and projector parameters are updated only through backpropagation using the cross-entropy loss generated by predicting subsequent text words. This training mechanism ensures that the speech encoder maintains a high degree of alignment with the text semantics during optimization, while preventing modification of the original parameters of the large language model. For example, in a medical scenario, when training the speech encoder to recognize pain features in a patient's speech, the frozen large language model can still accurately understand the medical semantics of the medical record text, preventing a decline in diagnostic keyword recognition ability due to parameter updates. This application effectively solves the problem of core capability degradation of the large language model during speech encoder training. In the medical field, this method ensures that when the speech encoder learns patient emotional features, the model maintains accurate parsing ability for professional terminology in the medical record text, avoiding diagnostic errors caused by parameter interference. In a financial scenario, this method enables the model to maintain strict adherence to compliant text when adapting to dialect speech, preventing the generation of speech responses that do not meet regulatory requirements.

[0059] Furthermore, model training can employ a step-by-step training strategy: Step 1: Train the model using both Chinese and English data, establishing a stable speech-semantic mapping as the model converges. Step 2: Incorporate dialect and sentiment category data to optimize the model's generalization ability to diverse and unstructured speech inputs.

[0060] In this context, text transcription lexical units refer to the discrete symbol sequences converted from transcribed text data by a word segmenter. Specifically, a byte-pair-based word segmenter can be used to establish the semantic alignment between speech and text. Subsequent text lexical units refer to the sequence of lexical units corresponding to the subsequent text content associated with the current speech query. This can be generated by extracting context text through a sliding window, providing supervision signals for speech encoder training. Speech query embeddings refer to the continuous vector representation extracted by the speech encoder and projector. This can be implemented using a convolutional neural network and a linear transform layer, used to capture local acoustic patterns in speech. The initial input sequence refers to the joint representation sequence formed by concatenating the speech query embedding and the text transcription lexical unit embedding. This can be achieved through vector concatenation operations, forcing the model to process the joint semantic information of speech and text simultaneously. The cross-entropy loss function measures the difference between the predicted and true lexical unit distributions. Log-likelihood loss can be used to accurately quantify the speech encoder's deviation in semantic alignment.

[0061] S4: Perform data construction and data labeling on the initial data to generate a quadruple dataset.

[0062] Specifically, the quadruple dataset includes text queries, voice queries, text responses, and voice responses. The voice responses are discretized using a neural codec to achieve joint modeling of multimodal dialogue context.

[0063] Please see Figure 6 , Figure 6 A specific implementation of step S4 is shown below:

[0064] S41: Construct the initial data into quadruple-formatted data, wherein the quadruple-formatted data includes text query, voice query, text reply, and voice reply data. S42: Discretize the voice reply data in the quadruple-formatted data into voice reply tokens using a neural codec. S43: Convert the text query data and text reply data in the quadruple-formatted data into text query tokens and text reply tokens using the word segmenter of the large language model. S44: Construct the quadruple dataset based on the voice query data, the text query tokens, and the text reply tokens in the quadruple-formatted data.

[0065] Specifically, in the construction of the quadruple data, the original interaction data is structured into complete units containing both text and speech input and output. Speech response data is processed by a neural codec to generate discrete speech units. This process extracts acoustic features through a multi-layer convolutional network and then uses vector quantization to generate discrete tokens that match the dimensions of the text units. Text queries and responses are converted into sub-word sequences by a word segmenter, which, together with the waveform data of the original speech query, form the training samples. This multimodal mixed data organization allows the model to simultaneously learn cross-modal associations between speech and text, as well as the generation rules of speech tokens, thereby reducing reliance on precisely aligned data during the training phase. This application enables the speech synthesis model to handle mixed-modal input prompts. In dialect interaction scenarios in the financial field, the model can simultaneously parse the acoustic features in customer speech queries and the professional terminology in text queries. In medical scenarios, the system can generate multimodal responses conforming to medical standards based on doctors' mixed input speech instructions and text parameters, significantly improving its adaptability to vertical domain professional terminology and multi-dialect scenarios.

[0066] Among them, the neural encoder-decoder refers to a speech discretization model based on self-supervised learning, which can be implemented using an acoustic encoder based on residual vector quantization. By converting continuous speech waveforms into discrete speech word sequences, it avoids information loss in the traditional speech feature extraction process. The word segmenter refers to the text processing module built into the large language model, which can be implemented using a sub-word segmentation algorithm based on byte pair encoding. By uniformly processing text queries and responses, it achieves the alignment of text and speech word representations in the embedding space.

[0067] S5: Construct a training sample sequence based on the quadruplet dataset, and train the large language model in the speech synthesis model based on the training sample sequence to generate the target speech synthesis model.

[0068] While maintaining the powerful text understanding capabilities of LLM, this application's embodiments fine-tune its top-level parameters, enabling it to directly and autoregressively generate discrete tokens representing target speech based on speech and text queries.

[0069] Please see Figure 7 , Figure 7 A specific implementation of step S5 is shown below:

[0070] S51: The speech query data in the four-tuple format data is converted into a target speech query embedding using the target speech encoder and the target projector. S52: The text query lexical, the target speech query embedding, and the text response lexical are concatenated sequentially to generate the training sample sequence. S53: The bottom N layers of the target speech encoder, the target projector, and the large language model in the speech synthesis model are frozen to obtain the target frozen speech synthesis model. S54: The target frozen speech synthesis model predicts and generates the next lexical based on the training sample sequence to obtain the target lexical, and a second loss value is calculated based on the target lexical and the speech response lexical using the cross-entropy loss function. S55: The parameters of the top K layers of the large language model in the target frozen speech synthesis model are updated based on the second loss value using backpropagation to retrain the target frozen speech synthesis model iteratively, generating the target speech synthesis model.

[0071] Specifically, the voice query data is processed by extracting acoustic features from a pre-trained target speech encoder, which then maps these features into embedding vectors aligned with the text word dimensions via a projector. Text query words are obtained from a word segmenter and sequentially concatenated with the voice embeddings to form a cross-modal input sequence. A freeze operation on the bottom N layers of the large language model ensures the parameter stability of the basic semantic parsing module, preventing the loss of original text understanding capabilities due to training for the speech generation task. During training, the model predicts speech response words based on the concatenated sequence, using a cross-entropy loss function to measure the difference between the predicted words and the actual speech response words. Backpropagation only updates the weight parameters of the top K layers, allowing the higher-level network to focus on learning speech tag generation rules, while the lower-level language understanding module retains its original function. This phased parameter update mechanism achieves decoupling optimization of language understanding and speech generation capabilities in speech synthesis tasks. In the intelligent customer service scenario in the financial field, this application can accurately capture emotional fluctuations in customer speech, generating response speech with appropriate emotional coloring while maintaining accurate understanding of financial terminology. In the context of electronic medical record generation in the medical field, the model, when synthesizing speech reports, can correctly parse medical terminology text while preserving key health indicator features in the patient's speech, avoiding semantic errors caused by model capability degradation. This solution improves the quality of speech synthesis while ensuring that the core language understanding capabilities of the large language model are not compromised through the optimization of parameter update strategies.

[0072] The first part of the text describes a language model's language model. The first part focuses on embedding speech features into a pre-trained speech encoder and projector. Specifically, it uses a convolutional neural network to extract acoustic features and then maps dimensions through a fully connected layer. This embedding preserves the continuous acoustic information of the original speech, avoiding quantization loss. The second part describes the training sample sequence, which is formed by sequentially concatenating text words with speech embeddings. This can be achieved using tensor concatenation. This sequence construction allows the model to process the joint semantics of text and speech simultaneously. The third part describes freezing the bottom N layers, fixing the parameters of the lower-level networks in the large language model and preventing them from participating in training. This can be achieved by setting a gradient calculation masking flag. This operation preserves the model's original language understanding capabilities and prevents performance degradation caused by parameter perturbations. The fourth part describes updating the top K layers, optimizing only the higher-level networks of the large language model. This can be achieved by calculating gradients and updating the weight matrix using backpropagation. This strategy optimizes the abstract feature representations related to speech generation while maintaining basic language capabilities. The fifth part describes the bottom N layers (e.g., the first two-thirds) and the sixth part describes the top K layers (e.g., the last one-third).

[0073] S6: Obtain the query information input by the user, and perform speech synthesis processing based on the query information through the target speech synthesis model to generate the target speech response.

[0074] Please see Figure 8 , Figure 8 A specific implementation of step S6 is shown below:

[0075] S61: Obtain the query information input by the user. S62: If the query information is a voice prompt, convert the voice prompt into a voice query embedding using the target voice encoder and the target projector. S63: If the query information is a text prompt, convert the text prompt into text units. S64: Perform voice unit prediction on the voice query embedding and / or the text units using the target voice synthesis model to generate a voice token sequence. S65: Decode and reconstruct the voice token sequence into a voice waveform to generate the target voice response.

[0076] Specifically, when the user input is a voice prompt, the target speech encoder extracts frame-level features from the input speech, generating a continuous vector containing prosody and timbre features. The projector maps this vector to an embedding representation aligned with the latent space of the large language model. For text input, the word segmenter converts technical terms or dialect words into a sequence of standard word units. The embedding vectors of both modalities are input into the large language model through concatenation or cross-attention mechanisms. During the autoregressive generation process, the model simultaneously models acoustic features and semantic context to predict subsequent speech tags. When the vocoder reconstructs the waveform based on the tag sequence, the continuous acoustic features in the speech query embedding guide the generation of acoustic parameters such as fundamental frequency and energy, ensuring that the emotional expression of the synthesized speech is consistent with the input prompt. This application can accurately reproduce the anxious emotions in the customer's voice in financial dual recording scenarios and generate empathetic response speech. In medical follow-up scenarios, it supports doctors to use a mixture of technical terminology text and patient dialect speech as input to generate guidance speech that conforms to local dialect habits and retains the patient's emotional characteristics, while avoiding errors in the understanding of medical concepts by the large language model during speech generation.

[0077] Among them, speech query embedding refers to the continuous acoustic feature vector extracted by the target speech encoder, which can be implemented using a convolutional neural network or transformer architecture, and is used to preserve the paralinguistic information in the original speech; text lexical refers to the discrete symbols of the text processed by the word segmenter, which can be implemented using a word-based segmentation algorithm, and is used to maintain the semantic integrity of the text; speech tag sequence refers to the discrete acoustic unit sequence predicted by the large language model, which can be implemented using an acoustic codebook index generated by a neural codec, and is used to drive the vocoder to synthesize waveforms.

[0078] In the scenario of highly human-like intelligent customer service and financial advisory: Customers interact with mobile banking apps, smart speakers, and other terminals via voice to inquire about financial products, check bills, or troubleshoot problems. This application can receive customers' voice questions (which may contain expressions of doubt or anxiety), understand their semantics and emotions, and generate voice responses with reassuring, professional, and confident tones, rather than cold and rigid machine broadcasts. This greatly enhances customer experience and trust. In the scenario of human-like medical assistant and chronic disease management, it can conduct post-discharge follow-ups and daily management of patients with chronic diseases (such as diabetes and hypertension). This application can communicate with patients via telephone or smart devices using a caring and encouraging tone of voice to inquire about their condition and remind them to take medication. Patients can describe their physical condition in the most natural way ("I felt a little dizzy today, but I feel better after taking the medicine"). The system can not only understand the text content but also make a preliminary judgment on the patient's condition through voice characteristics (such as weakness or lack of strength) and generate corresponding comfort and suggestions, or trigger alarms to notify doctors.

[0079] In this embodiment, initial data is acquired, and data construction and preprocessing are performed on the initial data to generate a triplet dataset. The initial data includes multilingual speech and text pairing data, sentiment tagging data, and TTS synthesized speech data. Speech encoding and feature projection are performed on the triplet dataset to obtain a speech query embedding. Based on the speech query embedding and the triplet dataset, the speech encoder and projector in the speech synthesis model are trained to obtain a target speech encoder and a target projector. Data construction and labeling are performed on the initial data to generate a quadruple dataset. A training sample sequence is constructed based on the quadruple dataset, and the large language model in the speech synthesis model is trained based on the training sample sequence to generate a target speech synthesis model. User-input query information is acquired, and the target speech synthesis model performs speech synthesis processing based on the query information to generate a target speech response. This invention, through the construction of a multimodal training dataset and a phased model training strategy, achieves continuous feature encoding and cross-modal alignment of speech prompts while preserving the core understanding capabilities of large language models. This solves the problems of speech prompt quantization loss, difficulty in adapting professional domain data, and insufficient semantic reliability in existing technologies, and is conducive to improving the accuracy of speech synthesis.

[0080] This application effectively solves the problem of emotional feature loss during the quantification process of voice prompts. In financial customer service scenarios, it can accurately identify the anxious emotions in dialect speech and generate empathetic responses. In medical scenarios, it can correctly parse the voice input of professional terminology and generate medical record text and voice recordings that conform to medical standards. The system maintains semantic coherence when processing multi-turn dialogues, avoiding the context breakage problem common in traditional methods, and supports seamless bidirectional switching between voice and text, improving the naturalness and professionalism of human-computer interaction.

[0081] Please refer to Figure 9 As a response to the above Figure 2 The implementation of the method shown in this application provides an embodiment of a speech synthesis device based on a large language model, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0082] like Figure 9 As shown, the speech synthesis device based on a large language model in this embodiment includes: a triplet dataset construction module 71, a speech query embedding module 72, a first model training module 73, a quadruple dataset construction module 74, a second model training module 75, and a target speech response generation module 76, wherein:

[0083] The triplet dataset construction module 71 is used to acquire initial data and perform data construction and data preprocessing on the initial data to generate a triplet dataset. The initial data includes multilingual speech and text pairing data, sentiment tag data, and TTS synthesized speech data.

[0084] The speech query embedding module 72 is used to perform speech encoding and feature projection on the triplet dataset to obtain the speech query embedding.

[0085] The first model training module 73 is used to train the speech encoder and projector in the speech synthesis model based on the speech query embedding and the triplet dataset to obtain the target speech encoder and target projector.

[0086] Quadruple dataset construction module 74 is used to construct and label the initial data to generate a quadruple dataset;

[0087] The second model training module 75 is used to construct a training sample sequence based on the quadruple dataset, and to train the large language model in the speech synthesis model based on the training sample sequence to generate the target speech synthesis model.

[0088] The target speech response generation module 76 is used to obtain query information input by the user and perform speech synthesis processing based on the query information through the target speech synthesis model to generate a target speech response.

[0089] Furthermore, the target speech response generation module 76 includes:

[0090] A query information acquisition unit is used to acquire the query information input by the user;

[0091] The first prompt processing unit is configured to convert the voice prompt information into a voice query embedding through the target voice encoder and the target projector if the query information is voice prompt information.

[0092] The second prompt processing unit is used to convert the text prompt information into text words if the query information is text prompt information;

[0093] The prediction unit is used to embed the speech query into the target speech synthesis model and / or predict the speech lexical units of the text to generate a speech tag sequence.

[0094] The reconstruction unit is used to decode and reconstruct the speech marker sequence into a speech waveform to generate the target speech response.

[0095] Furthermore, the triplet dataset construction module 71 includes:

[0096] A data construction unit is used to acquire the initial data and construct the initial data into triplet format data, wherein the triplet format data includes data of voice query, text transcription and subsequent text;

[0097] The resampling unit is used to resample the speech query data in the triplet format data and normalize the resampled speech query data to obtain normalized speech data.

[0098] The text cleaning unit is used to perform text cleaning and unified encoding on the text transcription data and subsequent text data in the triplet format data to obtain the encoded text transcription data and subsequent text data.

[0099] A lexical identifier sequence generation unit is used to convert the encoded text transcription data and subsequent text data into lexical identifier sequences through the word segmenter of the large language model;

[0100] The triplet dataset generation unit is used to construct the triplet dataset based on the normalized speech data and the lexical identifier sequence.

[0101] Furthermore, the voice query embedding module 72 includes:

[0102] The continuous acoustic embedding vector generation unit is used to extract features from the speech data in the triplet dataset through the speech encoder in the speech synthesis model to generate continuous acoustic embedding vectors.

[0103] The speech query embedding generation unit is used to generate the speech query embedding by capturing local acoustic patterns and mapping feature dimensions based on the continuous acoustic embedding vector through the projector in the speech synthesis model.

[0104] Furthermore, the first model training module 73 includes:

[0105] The subsequent text lexical acquisition unit is used to acquire the text transcription lexical and subsequent text lexical corresponding to the text transcription and subsequent text in the triplet dataset;

[0106] The concatenation unit is used to concatenate the speech query embedding with the embedding vector of the text transcription word to generate an initial input sequence;

[0107] The first loss value calculation unit is used to freeze the large language model in the speech synthesis model to obtain an initial frozen speech synthesis model, predict and generate the next word based on the initial input sequence through the initial frozen speech synthesis model to obtain the initial word, and sample the cross-entropy loss function to calculate the first loss value based on the initial word and the subsequent text word.

[0108] The first parameter update unit is used to update the parameters of the speech encoder and projector in the initial frozen speech synthesis model based on the first loss value using backpropagation, so as to retrain the speech encoder and projector in the initial frozen speech synthesis model to generate the target speech encoder and the target projector.

[0109] Furthermore, the quadruple dataset building module 74 includes:

[0110] The quadruple format data construction unit is used to construct the initial data into quadruple format data, wherein the quadruple format data includes text query, voice query, text reply, and voice reply data;

[0111] A discretization unit is used to discretize the speech response data in the quadruplet format data into speech response lexical units using a neural codec.

[0112] The data conversion unit is used to convert the text query data and text response data in the quadruple format data into text query tokens and text response tokens through the word segmenter of the large language model;

[0113] The quadruple dataset generation unit is used to construct the quadruple dataset based on the speech query data, the text query lexical, and the text response lexical in the quadruple format data.

[0114] Furthermore, the second model training module 75 includes:

[0115] A target speech query embedding generation unit is used to convert speech query data in the quadruple format data into a target speech query embedding through the target speech encoder and the target projector.

[0116] The training sample sequence generation unit is used to sequentially concatenate the text query lexical, the target speech query embedding, and the text response lexical to generate the training sample sequence;

[0117] The freezing unit is used to freeze the target speech encoder, the target projector, and the bottom N layers of the large language model in the speech synthesis model to obtain the target frozen speech synthesis model.

[0118] The second loss value calculation unit is used to predict and generate the next word unit based on the training sample sequence through the speech synthesis model after the target is frozen, obtain the target word unit, and calculate the second loss value based on the target word unit and the speech response word unit using the cross-entropy loss function.

[0119] The second parameter update unit is used to update the parameters of the top K layers of the large language model in the target frozen speech synthesis model based on the second loss value using backpropagation, so as to retrain the target frozen speech synthesis model iteratively and generate the target speech synthesis model.

[0120] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 10 , Figure 10 This is a basic structural block diagram of the computer device in this embodiment.

[0121] Computer device 8 includes a memory 81, a processor 82, and a network interface 83 that are interconnected via a system bus. It should be noted that... Figure 10 Only a computer device 8 with three components—memory 81, processor 82, and network interface 83—is shown. It should be understood that implementing all shown components is not required; more or fewer components may be implemented alternatively. Those skilled in the art will understand that this computer device is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), and embedded devices.

[0122] Computer devices can include desktop computers, laptops, handheld computers, and cloud servers. These devices allow for human-computer interaction with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.

[0123] The memory 81 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 81 may be an internal storage unit of the computer device 8, such as the hard disk or memory of the computer device 8. In other embodiments, the memory 81 may also be an external storage device of the computer device 8, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 8. Of course, the memory 81 may include both internal storage units and external storage devices of the computer device 8. In this embodiment, the memory 81 is typically used to store the operating system and various application software installed on the computer device 8, such as the program code of a speech synthesis method based on a large language model. In addition, the memory 81 can also be used to temporarily store various types of data that have been output or will be output.

[0124] In some embodiments, processor 82 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. This processor 82 is typically used to control the overall operation of the computer device 8. In this embodiment, processor 82 is used to run program code stored in memory 81 or process data, for example, to run the program code of the aforementioned large language model-based speech synthesis method to implement various embodiments of the large language model-based speech synthesis method.

[0125] The network interface 83 may include a wireless network interface or a wired network interface, which is typically used to establish a communication connection between the computer device 8 and other electronic devices.

[0126] This application also provides another embodiment, namely, providing a computer-readable storage medium storing a computer program that can be executed by at least one processor to cause the at least one processor to perform the steps of the speech synthesis method based on a large language model as described above.

[0127] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.

[0128] Obviously, the embodiments described above are merely some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the scope of this application. This application can be implemented in many different forms; rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of protection of this application.

Claims

1. A speech synthesis method based on a large language model, characterized in that, include: Acquire initial data, and perform data construction and data preprocessing on the initial data to generate a triplet dataset, wherein the initial data includes multilingual speech and text pairing data, sentiment tag data, and TTS synthesized speech data; The triplet dataset is subjected to speech encoding and feature projection to obtain the speech query embedding; The speech encoder and projector in the speech synthesis model are trained based on the speech query embedding and the triple dataset to obtain the target speech encoder and target projector. The initial data is processed by data construction and data labeling to generate a quadruple dataset; A training sample sequence is constructed based on the quadruple dataset, and the large language model in the speech synthesis model is trained based on the training sample sequence to generate the target speech synthesis model. The system obtains query information input by the user and performs speech synthesis processing based on the query information using the target speech synthesis model to generate a target speech response. The step of obtaining user-input query information and generating a target speech response by performing speech synthesis processing based on the query information using the target speech synthesis model includes: Obtain the query information input by the user; If the query information is a voice prompt, then the voice prompt is converted into a voice query embedding by the target voice encoder and the target projector; If the query information is a text prompt, then the text prompt is converted into text tokens; The speech query is embedded and / or the text lexical units are predicted by the target speech synthesis model to generate a speech tag sequence. The speech marker sequence is decoded and reconstructed into a speech waveform to generate the target speech response; The step of training the speech encoder and projector in the speech synthesis model based on the speech query embedding and the triplet dataset to obtain the target speech encoder and target projector includes: Obtain the text transcription lexical units and subsequent text lexical units corresponding to the text transcription and subsequent text in the triplet dataset; The speech query embedding is concatenated with the embedding vector of the text transcription word to generate an initial input sequence; Freeze the large language model in the speech synthesis model to obtain an initial frozen speech synthesis model. Based on the initial input sequence, predict and generate the next word unit through the initial frozen speech synthesis model to obtain the initial word unit. Then, use the cross-entropy loss function to calculate the first loss value based on the initial word unit and the subsequent text word unit. The parameters of the speech encoder and projector in the initial frozen speech synthesis model are updated based on the first loss value using backpropagation, so as to retrain the speech encoder and projector in the initial frozen speech synthesis model iteratively to generate the target speech encoder and the target projector.

2. The speech synthesis method based on a large language model according to claim 1, characterized in that, The process of acquiring initial data and performing data construction and preprocessing on the initial data to generate a triplet dataset includes: The initial data is obtained and constructed into triplet format data, wherein the triplet format data includes data of voice query, text transcription and subsequent text; The speech query data in the triplet format data is resampled, and the resampled speech query data is normalized to obtain normalized speech data. The text transcription data and subsequent text data in the triplet format data are cleaned and uniformly encoded to obtain the encoded text transcription data and subsequent text data. The encoded text transcription data and subsequent text data are converted into word identifier sequences by the word segmenter of the large language model; The triplet dataset is constructed based on the normalized speech data and the lexical identifier sequence.

3. The speech synthesis method based on a large language model according to claim 1, characterized in that, The step of performing speech encoding and feature projection on the triplet dataset to obtain the speech query embedding includes: The speech encoder in the speech synthesis model extracts features from the speech data in the triplet dataset to generate a continuous acoustic embedding vector. The speech query embedding is generated by the projector in the speech synthesis model, which performs local acoustic pattern capture and feature dimension mapping based on the continuous acoustic embedding vector.

4. The speech synthesis method based on a large language model according to any one of claims 1 to 3, characterized in that, The step of constructing and labeling the initial data to generate a four-tuple dataset includes: The initial data is constructed into a four-tuple format, wherein the four-tuple format data includes text query, voice query, text reply, and voice reply data; The speech response data in the quadruplet format data is discretized into speech response tokens using a neural codec. The word segmenter of the large language model converts the text query data and text response data in the quadruple format data into text query tokens and text response tokens. The quadruple dataset is constructed based on the voice query data, the text query terms, and the text response terms in the quadruple format data.

5. The speech synthesis method based on a large language model according to claim 4, characterized in that, The step of constructing a training sample sequence based on the quadruplet dataset and training the large language model in the speech synthesis model based on the training sample sequence to generate the target speech synthesis model includes: The target speech encoder and the target projector convert the speech query data in the quadruple format data into a target speech query embedding. The text query terms, the target speech query embedding, and the text response terms are concatenated sequentially to generate the training sample sequence. Freeze the target speech encoder, the target projector, and the bottom N layers of the large language model in the speech synthesis model to obtain the target frozen speech synthesis model; After the target is frozen, the speech synthesis model predicts and generates the next word based on the training sample sequence to obtain the target word. The cross-entropy loss function is then used to calculate a second loss value based on the target word and the speech response word. The parameters of the top K layers of the large language model in the target frozen speech synthesis model are updated based on the second loss value using backpropagation, so as to retrain the target frozen speech synthesis model iteratively and generate the target speech synthesis model.

6. A speech synthesis device based on a large language model, characterized in that, include: The triplet dataset construction module is used to acquire initial data and perform data construction and data preprocessing on the initial data to generate a triplet dataset. The initial data includes multilingual speech and text pairing data, sentiment tag data, and TTS synthesized speech data. The speech query embedding module is used to perform speech encoding and feature projection on the triplet dataset to obtain the speech query embedding; The first model training module is used to train the speech encoder and projector in the speech synthesis model based on the speech query embedding and the triplet dataset to obtain the target speech encoder and target projector. The quadruple dataset construction module is used to construct and label the initial data to generate a quadruple dataset. The second model training module is used to construct a training sample sequence based on the quadruple dataset, and to train the large language model in the speech synthesis model based on the training sample sequence to generate the target speech synthesis model. The target speech response generation module is used to obtain query information input by the user and perform speech synthesis processing based on the query information through the target speech synthesis model to generate a target speech response; The target speech response generation module includes: A query information acquisition unit is used to acquire the query information input by the user; The first prompt processing unit is configured to convert the voice prompt information into a voice query embedding through the target voice encoder and the target projector if the query information is voice prompt information. The second prompt processing unit is used to convert the text prompt information into text words if the query information is text prompt information; The prediction unit is used to embed the speech query into the target speech synthesis model and / or predict the speech lexical units of the text to generate a speech tag sequence. The reconstruction unit is used to decode and reconstruct the speech marker sequence into a speech waveform to generate the target speech response; The first model training module includes: The subsequent text lexical acquisition unit is used to acquire the text transcription lexical and subsequent text lexical corresponding to the text transcription and subsequent text in the triplet dataset; The concatenation unit is used to concatenate the speech query embedding with the embedding vector of the text transcription word to generate an initial input sequence; The first loss value calculation unit is used to freeze the large language model in the speech synthesis model to obtain an initial frozen speech synthesis model, predict and generate the next word based on the initial input sequence through the initial frozen speech synthesis model to obtain the initial word, and use the cross-entropy loss function to calculate the first loss value based on the initial word and the subsequent text word. The first parameter update unit is used to update the parameters of the speech encoder and projector in the initial frozen speech synthesis model based on the first loss value using backpropagation, so as to retrain the speech encoder and projector in the initial frozen speech synthesis model to generate the target speech encoder and the target projector.

7. A computer device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the speech synthesis method based on a large language model as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the speech synthesis method based on a large language model as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Voice generation method and device, equipment and medium

    CN120148474A

  • Word-level end-to-end neural speaker diarization with auxnet

    US20250118292A1