Voice generation model acquisition method and voice generation method

By using text word participle and pronunciation word participle to obtain word element samples and train a large language model for speech generation, combined with distillation and quantization processing, the problem of insufficient speech generation efficiency in the prior art is solved, and efficient and real-time speech synthesis is achieved.

CN120071892APending Publication Date: 2025-05-30GUANGZHOU QUYAN NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510277771.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing voice generation technology is not efficient in real-time voice generation, and it is difficult to meet application scenarios with high requirements for real-time performance.

Method used

By using text word segmenter and pronunciation word segmenter to obtain text word segmenter and pronunciation word segmenter, train the initial speech generation large language model, obtain the optimized speech generation large language model, and obtain the target speech generation large language model for speech synthesis through distillation and quantization processing.

Benefits of technology

It significantly improves the real-time efficiency of speech generation, solves the problem of insufficient efficiency in the existing technology, and maintains high-quality speech synthesis effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071892A_ABST
    Figure CN120071892A_ABST
Patent Text Reader

Abstract

According to the voice generation model obtaining method and the voice generation method provided by the invention, the obtained text lexical element sample and the obtained voice lexical element sample are adopted to train the initial voice generation large language model, the optimized voice generation large language model is obtained, and the text lexical element sample is obtained by using the text word segmentation device. The text word segmentation device is used for improving the text lexical element compression ratio and retaining text information, the voice lexical element samples are obtained by using the voice word segmentation device, and the voice word segmentation device is used for improving the voice lexical element compression ratio and retaining voice information; and according to the optimized voice generation large language model, obtaining a target voice generation large language model, and performing voice generation by using the target voice generation large language model. Therefore, by improving the lexical element compression ratio, the data volume is reduced while the key information is reserved, so that the model calculation amount is reduced, the model reasoning speed is increased, and the real-time voice generation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method, apparatus, storage medium, and computer device for speech generation. Background Art

[0002] In recent years, with the development of speech synthesis technology, personalized speech generation has become an important research direction. Personalized speech generation technology can synthesize highly similar speech according to the user's voice characteristics and text content, and plays an important role in application scenarios such as intelligent voice assistants, online education, speech translation, personalized broadcasting, and virtual customer service. Currently, many speech generation systems are based on deep learning models and are trained using large-scale speech and text data to improve the naturalness and personalization of the generated speech. For example, the VALL-E model (Voice LLM-based Audio Language model) is based on the Transformers mechanism and has achieved remarkable results with its high-quality personalized speech synthesis ability.

[0003] However, the current speech generation technology still faces the problem of insufficient efficiency and is difficult to meet application scenarios with high real-time requirements. Since the generation process involves complex deep learning computations, speech generation models often have a long inference time. Taking the VALL-E model as an example, when generating 10 seconds of speech, it usually requires 20 - 30 seconds of processing time, which is much higher than the actual speech playback duration and significantly affects its feasibility in practical applications. Summary of the Invention

[0004] The purpose of this application aims to solve at least one of the above technical defects, especially the technical defect of insufficient efficiency in real-time speech generation in the prior art.

[0005] In a first aspect, this application provides a method for obtaining a speech generation model. The method includes:

[0006] Training an initial large language model for speech generation with the obtained text token samples and speech token samples to obtain an optimized large language model for speech generation. The text token samples are obtained using a text tokenizer, which is used to improve the text token compression rate and retain text information. The speech token samples are obtained using a speech tokenizer, which is used to improve the speech token compression rate and retain speech information;

[0007] Based on the optimized large language model for speech generation, obtain a target large language model for speech generation, which is used for speech synthesis.

[0008] In one of the embodiments, the process of obtaining the text token samples includes:

[0009] Obtain the generated text sample, and use a text tokenizer to discretize the generated text sample to obtain a text token sample. The text tokenizer is the Qwen2 tokenizer.

[0010] In one embodiment, the process of obtaining the speech token sample includes:

[0011] Obtain a reference speech sample, and use a speech tokenizer to discretize the speech per second in the reference speech sample into M tokens to obtain a speech token sample. The speech tokenizer is composed of residual vector quantization and the ConvNeXt architecture, where M is a positive integer and M < 50.

[0012] In one embodiment, the steps of obtaining the target speech generation large language model according to the optimized speech large language model include:

[0013] Based on a preset objective function, distill the optimized speech generation large language model to obtain a distilled speech generation large language model. Among them, the model architectures of the optimized speech generation large language model and the distilled speech generation large language model are the same;

[0014] After quantizing the distilled speech generation large language model, accelerate it through an inference engine to obtain the target speech generation large language model.

[0015] In one embodiment, the preset objective function is:

[0016]

[0017] Among them, represents the target preset function, and represent hyperparameters, represents the supervised cross-entropy loss of the true speech label corresponding to the speech token sample, represents the distilled speech generation large language model, represents the text token sample, represents the speech token sample, represents the true speech label corresponding to the speech token sample, represents the KL divergence distribution difference between the optimized speech generation large language model and the distilled speech generation large language model, represents the optimized speech generation large language model.

[0018] In one embodiment, the steps of accelerating the distilled speech generation large language model through an inference engine after quantization include:

[0019] After quantizing the model weights and activation values of the distilled speech generation large language model into floating-point numbers with a preset number of bits, ONNX and TensorRT are used to accelerate the quantized distilled speech generation large language model.

[0020] In a second aspect, the present application provides a speech generation method, the method comprising:

[0021] Obtain a target generation text and a target reference speech;

[0022] According to the target generation text and the target reference speech, use the target speech generation large language model to obtain a target synthesized speech, where the target speech generation large language model is obtained by using the speech generation model acquisition method described in any one of the above embodiments.

[0023] In a third aspect, the present application provides a speech generation model acquisition device, the device comprising:

[0024] An optimized speech generation large language model acquisition module, configured to train an initial speech generation large language model by using the obtained text token samples and speech token samples to obtain an optimized speech generation large language model, where the text token samples are obtained by using a text tokenizer for improving the text token compression rate and retaining text information, and the speech token samples are obtained by using a speech tokenizer for improving the speech token compression rate and retaining speech information;

[0025] A target speech generation large language model acquisition module, configured to obtain a target speech generation large language model according to the optimized speech large language model, where the target speech generation large language model is used for speech synthesis.

[0026] In a fourth aspect, the present application provides a speech generation device, the device comprising:

[0027] A target reference speech acquisition module, configured to obtain a target generation text and a target reference speech;

[0028] A target synthesized speech acquisition module, configured to obtain a target synthesized speech according to the target generation text and the target reference speech by using the target speech generation large language model, where the target speech generation large language model is obtained by using the speech generation model acquisition device in the above embodiments.

[0029] In a fifth aspect, the present application provides a computer device, comprising: one or more processors, and a memory;

[0030] The memory stores computer-readable instructions, which when executed by one or more processors, execute the steps of the speech generation model acquisition method described in any one of the above embodiments, and / or, the steps of the speech generation method.

[0031] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:

[0032] In the method for obtaining a speech generation model and the speech generation method provided by the present application, first, when training the initial speech generation large language model, text tokens samples are obtained through a text tokenizer first, which can ensure a high compression rate while retaining text information, and speech tokens samples are obtained through a speech tokenizer to achieve efficient expression of speech data, thus greatly improving the token compression rate and reducing the data volume while retaining key information. Since the data input into the model is compressed and the core information is retained, the model can learn the correspondence between text and speech more quickly and accurately during the training process, reducing ineffective calculations, thereby accelerating model convergence, enabling the optimized speech generation large language model to complete training in a shorter time and achieve better performance. For the target speech generation large language model finally used for speech synthesis, in the real-time speech generation scenario, thanks to the efficient training process in the early stage, when facing the input text, it can reason more quickly based on the learned text-to-speech mapping relationship and generate the corresponding speech. Compared with the traditional model, the target speech generation large language model significantly improves the real-time generation efficiency while maintaining high-quality speech synthesis, solving the problem of insufficient efficiency in real-time speech generation in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0034] Figure 1 It is a schematic flowchart of the method for obtaining a speech generation model provided by an embodiment of the present application;

[0035] Figure 2 It is an example diagram of the training process of the initial speech generation large language model provided by an embodiment of the present application;

[0036] Figure 3 It is a schematic flowchart of the speech generation method provided by an embodiment of the present application;

[0037] Figure 4 It is a schematic structural diagram of the speech generation model obtaining device provided by an embodiment of the present application;

[0038] Figure 5 It is a schematic structural diagram of the speech generation device provided by an embodiment of the present application;

[0039] Figure 6Schematic internal structure diagram of the computer device provided by the embodiment of the present application. Detailed implementation manners

[0040] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0041] The present application provides a speech generation method. The following embodiments are described by taking the application of this method to a computer device as an example. It can be understood that the computer device can be various devices with data processing functions, including but not limited to a single server, a server cluster, a personal laptop, a desktop computer, etc. As Figure 1 shown, the present application provides a chat corpus annotation method, and this method includes:

[0042] S101: Train an initial speech generation large language model using the obtained text token samples and speech token samples to obtain an optimized speech generation large language model. The text token samples are obtained using a text tokenizer, and the text tokenizer is used to improve the text token compression rate and retain text information. The speech token samples are obtained using a speech tokenizer, and the speech tokenizer is used to improve the speech token compression rate and retain speech information.

[0043] Among them, the text token samples refer to the basic units extracted from text data, such as words, sub-words, or characters, etc., and are used to train the initial speech generation large language model. The speech token samples refer to the basic units extracted from speech data, such as phonemes, acoustic features, or discrete representations converted by a speech tokenizer, and are used to train the initial speech generation large language model. The speech generation large language model is a model based on a large-scale neural network, which can generate a natural speech output of the target text according to the input speaker voice and target text, and can be trained by learning a large amount of speech and text data. The text tokenizer is used to segment text data to improve the data compression rate and retain the semantic information in the text as much as possible. The speech tokenizer is used to convert continuous speech signals into discrete speech tokens to achieve data compression while trying to retain the information in the speech.

[0044] In this step, first, the text entered by the user is tokenized by a text tokenizer to compress the text into an efficient token sample, which not only helps reduce the computational amount but also retains the key information in the text. Then, the voice sample data is processed by a voice tokenizer to convert the voice data into a voice token sample. This step also aims to improve the data compression rate and maintain the integrity of the voice information. Next, these processed voice and text token samples are input into a neural network to train the initial voice generation large language model. The generated optimized voice generation large language model can efficiently learn the mapping relationship between text and voice, thus being more accurate and natural when generating personalized voices.

[0045] In an example, first, the initial voice generation large language model needs to be loaded and initialized. Then, the text token sample and the voice token sample are received as input data and input into the initial voice generation large language model to generate the corresponding voice output result. Next, the voice generation result is analyzed to adjust the model parameters of the initial voice generation large language model to obtain the optimized voice generation large language model. For example, by calculating the difference between the generated result and the real voice sample, such as phoneme accuracy, voice fluency, or intonation naturalness, to evaluate the model performance; based on the evaluation result, an optimization algorithm can be used to adjust the model parameters, such as the gradient descent algorithm or the adaptive learning rate algorithm, to reduce the difference degree, making the generated result gradually approach the real voice sample until the model output reaches the expected quality, and finally obtaining the optimized voice generation large language model.

[0046] It can be understood that processing the text and voice samples respectively through the text tokenizer and the voice tokenizer to improve the compression rate of the tokens and maintain the information integrity helps achieve better learning effects with a smaller amount of data in model training. Through the training of the initial voice generation large language model, the model can more accurately restore the voice intonation, emotional expression, and voice coherence in the text-to-voice conversion. The finally obtained optimized voice generation large language model can generate more natural and fluent voice outputs in various scenarios. This execution method not only improves the generation quality of the model but also effectively reduces the computational resource consumption of the model, realizing a more efficient text-to-voice conversion.

[0047] S102: Obtain a target voice generation large language model according to the optimized voice generation large language model. The target voice generation large language model is used for speech synthesis.

[0048] In this step, after obtaining the optimized speech generation large language model, further performance improvement can be made to the optimized speech generation large language model to obtain the target speech generation large language model, so as to achieve high-efficiency and high-quality speech synthesis in actual application scenarios. For example, weight pruning can be adopted to reduce the computational complexity of the model, reduce the inference time, and improve the real-time performance of the generated speech; transfer learning technology can also be used to apply the model that has been trained in similar tasks to the new speech generation task, which can not only utilize the knowledge of the pre-trained model, but also avoid training the model from scratch, shorten the time for the model to adapt to the new task, and improve the efficiency of real-time speech generation; data from the target application scenario can also be used for fine-tuning, enabling the model to generate more natural speech according to specific features such as accent, speech rate, and language. Fine-tuning not only improves the speech quality, but also optimizes the model's adaptability to specific datasets, reduces the training time and resource consumption, and thus improves the efficiency of real-time speech generation.

[0049] In the above embodiment, first, when training the initial speech generation large language model, text token samples are first obtained through a text tokenizer to ensure high compression rate while retaining text information, and speech token samples are obtained through a speech tokenizer to achieve efficient expression of speech data, thereby greatly improving the token compression rate and reducing the data volume while retaining key information. Since the data input into the model is compressed and the core information is retained, the model can learn the correspondence between text and speech more quickly and accurately during the training process, reduce ineffective calculations, and thus accelerate model convergence, enabling the optimized speech generation large language model to complete training in a shorter time and achieve better performance. For the target speech generation large language model finally used for speech synthesis, in the real-time speech generation scenario, due to the efficient training process in the early stage, when facing the input text, it can more quickly perform inference based on the learned text-to-speech mapping relationship and generate the corresponding speech. Compared with traditional models, the target speech generation large language model significantly improves the real-time generation efficiency while maintaining high-quality speech synthesis, solving the problem of insufficient efficiency in real-time speech generation in the prior art.

[0050] In one embodiment, the process of obtaining text token samples includes:

[0051] Obtain a generated text sample, and discretize the generated text sample using a text tokenizer to obtain text token samples, where the text tokenizer is the Qwen2 tokenizer.

[0052] Among them, the generated text sample is a text segment generated by the model or extracted from an existing corpus. The Qwen2 tokenizer is an efficient text processing tool designed to improve the compression rate of tokens by optimizing the text tokenization process while retaining as much text information as possible. It uses advanced natural language processing algorithms to reduce unnecessary token splitting during text processing, resulting in fewer tokens while each token can still accurately express the semantics of the original text. Especially in real-time application scenarios such as speech generation, the Qwen2 tokenizer can provide a more compact and information-rich input for subsequent tasks.

[0053] Specifically, first, it is necessary to obtain the generated text sample, which can be read from a corpus, generated by a model, or extracted from external data sources such as user input, API call results, etc. Then, the obtained text sample is input into the Qwen2 tokenizer, which performs lexical analysis and encoding on the text sample, breaking the text into discrete tokens. Among them, the discretization process includes decomposing the text sample into sentence, word, or character-level segments, and mapping each segment to a numerical form, such as a token ID, by looking up a vocabulary or encoding rules. The Qwen2 tokenizer can accurately segment the text into the most appropriate tokens according to the pre-trained vocabulary, minimizing the number of tokens after tokenization, that is, by compressing redundant tokens in the text while retaining important semantic information. Each token is precisely segmented according to its meaning in the context, thus maximizing the retention of the semantic information of the text. Through the Qwen2 tokenizer, what is finally obtained is a set of samples composed of compressed and information-rich tokens. These tokens not only reduce in number, saving storage and processing time, but also can fully convey the meaning of the original text.

[0054] In this embodiment, using the Qwen2 tokenizer to discretize the generated text sample can significantly improve the text compression rate while retaining the key information of the text, thereby improving the speech generation efficiency and generation quality. When processing text, the Qwen2 tokenizer can efficiently convert complex natural language into a more compact and higher information density token representation. By compressing the text length and reducing redundant information, the tokenizer effectively reduces the computational complexity. This compact token sequence not only reduces the processing time of the model but also enhances the parallel processing ability of the model, making the speech generation faster and smoother.

[0055] In one embodiment, the process of obtaining the speech token sample includes:

[0056] Obtain a reference speech sample, and use a speech tokenizer to discretize each second of speech in the reference speech sample into M tokens to obtain a speech token sample. The speech tokenizer is composed of residual vector quantization and the ConvNeXt architecture, where M is a positive integer and M < 50.

[0057] Among them, the reference speech sample refers to the known speech data used for model training or inference. Residual vector quantization is used to compress speech signals or features, which improves the expressive ability of speech signals by reducing the residuals of errors during the quantization process, thereby enhancing the representation effect of speech. The ConvNeXt architecture is a deep learning architecture based on convolutional neural networks, which can efficiently process image and speech data. It uses advanced convolutional methods to extract useful features from speech signals for better speech processing. M represents the number of tokens into which the speech signal is discretized per second. M is a positive integer and M < 50, that is, the speech signal per second is segmented into less than 50 discrete units. For example, it can be segmented into 25 tokens.

[0058] Specifically, the reference speech sample can be extracted from a specified data source, such as a local database or cloud storage, etc. The reference speech sample can be recording segments of various language types, which are used to support the training of multi-lingual speech generation models. Then, the reference speech sample is input into the speech tokenizer, and the speech tokenizer will segment the speech signal into small discrete units based on residual vector quantization and the ConvNeXt architecture. The speech signal per second is decomposed into less than 50 tokens, and these tokens represent different audio segments, phonemes or other speech units in the speech signal, and each token can accurately describe the speech features in the corresponding time period.

[0059] In this embodiment, when the speech signal is segmented into less than 50 tokens, the number of discrete units to be processed per second will be greatly reduced. With the decrease in the number of tokens, the computational complexity will also decrease accordingly, making the speech generation large language model more rapid in processing and response, thus achieving low latency in real-time speech generation. In addition, when processing speech signals, Residual Vector Quantize (RVQ) optimizes the quantization process, retains more key information, and reduces the influence of irrelevant information, not only improving the quality of speech generation but also further compressing the data volume, enabling the speech generation large language model to maintain high-quality speech output while reducing the computational amount. In addition, as an efficient deep learning model, ConvNeXt can extract more detailed and efficient features from speech signals. Combined with residual vector quantization, the ConvNeXt architecture enhances the model's representation ability and understanding depth of speech signals. Through this efficient feature extraction, the speech generation large language model can not only generate speech content more quickly and accurately, but also maintain high-quality output while improving efficiency, further optimizing the real-time generation effect.

[0060] In an example, such as Figure 2As shown, the main function of the text tokenizer is to discretize the target generated text into a series of discrete text tokens, that is, text token samples, while the speech tokenizer extracts the speaker's speech information and discretizes the reference speech into speech tokens, that is, speech token samples. Subsequently, the text tokens and speech tokens are jointly input into the speech generation large model to generate speech . The specific speech generation process can be expressed as:

[0061]

[0062] where are the parameters of the speech generation large model.

[0063] In one embodiment, the steps of obtaining the target speech generation large language model according to the optimized speech large language model include:

[0064] Based on a preset objective function, distill the optimized speech generation large language model to obtain a distilled speech generation large language model, where the model architectures of the optimized speech generation large language model and the distilled speech generation large language model are the same;

[0065] After quantizing the distilled speech generation large language model, accelerate it through an inference engine to obtain the target speech generation large language model.

[0066] Among them, the preset objective function is a mathematical function used to evaluate the model performance during the training process, which defines the optimization objective that is expected to be achieved when training the model, such as minimizing the loss function or maximizing a certain performance metric. The distilled speech generation large language model is a model after distillation processing, specifically referring to transferring the knowledge of the optimized speech generation large language model to a smaller model with the same structure and function, so that its inference speed is faster and the computing requirements are lower.

[0067] Specifically, at the beginning of the distillation process, the optimized speech generation large language model is first deployed as the teacher model. The teacher model is used to generate reference outputs, which usually include intermediate feature representations and final output results. Then, by extracting the outputs or intermediate representations of the teacher model, the differences between the teacher model and the student model to be trained are calculated. These differences may include the probability distribution differences of the prediction results and the similarities of the intermediate layer features. Subsequently, according to the preset objective function, the parameters of the student model are continuously adjusted to gradually approximate the output behavior of the teacher model, thereby completing the knowledge transfer and obtaining the distilled speech generation large language model. It should be noted that the core of the distillation process for the optimized speech generation large language model is to efficiently transfer the knowledge of the teacher model to a smaller student model with the same structure, enabling the student model to maintain high accuracy on the same task while significantly improving the inference efficiency. By simulating the prediction behavior of the teacher model, the student model can not only provide high-quality speech outputs but also retain certain personalized features when generating speech.

[0068] It can be understood that during the distillation process, using the preset objective function for model distillation is to improve the inference efficiency of the model without changing the model architecture. In this way, the distilled speech generation large language model inherits the main capabilities of the optimized speech generation large language model, ensuring the speech generation quality while significantly reducing the computational overhead of the model. The distilled model is smaller in size, occupying less memory and computational resources, so it can respond to requests faster and improve the efficiency of real-time speech synthesis.

[0069] After completing the model distillation, the distilled speech generation large language model can be optimized for deployment. First, the distilled speech generation large language model is quantized, converting the floating-point parameters in the distilled speech generation large language model into fixed-point numbers or numerical representations with lower precision to reduce the storage space and computational resource requirements of the model. During this process, the model weight distribution can be automatically analyzed to select an appropriate quantization strategy, such as static quantization or dynamic quantization, so as to ensure that while reducing the model accuracy loss, the overall efficiency of the model is improved. Then, the inference engine is used to accelerate and optimize the quantized distilled speech generation large language model. The inference engine will perform a series of operations according to the structure of the model and the hardware environment, such as operator fusion, memory optimization, and hardware instruction optimization, to further improve the execution efficiency of the model in the inference stage. After distillation, quantization, and acceleration by the inference engine, the target speech generation large language model will be obtained. By inputting the text and voice to be synthesized into the target speech generation large language model, natural speech outputs corresponding to the text can be generated with the voice.

[0070] It can be understood that quantizing the distilled speech generation large language model and accelerating the inference engine helps to significantly improve the efficiency and deployment flexibility of speech generation. The quantization process can reduce the computational complexity of the model, decrease memory occupancy and computational costs, enabling the model to run on resource-constrained devices. At the same time, the inference engine optimization reduces the inference latency and increases the throughput through in-depth optimizations targeting hardware characteristics, thus enabling real-time speech generation. In this way, it is possible to effectively balance improving the inference speed, reducing resource consumption, and maintaining the speech generation quality.

[0071] In one embodiment, the preset objective function is:

[0072]

[0073] Wherein, represents the target preset function, and represent hyperparameters, represents the supervised cross-entropy loss of the true speech label corresponding to the speech token sample, represents the distilled speech generation large language model, represents the text token sample, represents the speech token sample, represents the true speech label corresponding to the speech token sample, represents the KL divergence distribution difference between the optimized speech generation large language model and the distilled speech generation large language model, represents the optimized speech generation large language model.

[0074] Specifically, this preset objective function is used for the distilled speech generation large language model, combining the hard label supervision loss and the soft label distillation loss , by adjusting the weight coefficients and , to improve the performance of the distilled speech generation large language model , making it have higher computational efficiency while maintaining high performance. Among them, the supervised cross-entropy loss is the loss calculated by the true speech label , used to measure the difference between the output speech result of the distilled speech generation large language model and the true speech label, ensuring that the model has accurate basic speech generation ability. The KL divergence distribution difference measures the distribution difference between the distilled speech generation large language model and the optimized speech generation large language model by comparing their output probability distributions, helping the distilled speech generation large language model Better learn the prediction behavior of the large model. KL divergence can capture more subtle features and prediction patterns in the speech generation process, enhancing the distilled speech generation large language model 's generalization ability and maintaining high performance even in small data volume scenarios.

[0075] In this embodiment, through the distillation training method of this objective function, the distilled speech generation large language model is close to the optimized speech generation large language model in terms of accuracy and generalization ability, while significantly reducing the model parameter scale and improving the inference efficiency. The hard label loss ensures that the generated result is consistent with the real speech label, improving the generation quality; the soft label loss helps the distilled speech generation large language model learn the complex prediction patterns of the optimized speech generation large language model, reduces the training data requirements, accelerates the model convergence, and improves the generation stability. The overall optimization process not only improves the speed of speech generation but also reduces the computational resource consumption.

[0076] In one embodiment, the steps of accelerating the distilled speech generation large language model through an inference engine after quantization include:

[0077] After quantizing the model weights and activation values of the distilled speech generation large language model into floating-point numbers with a preset number of bits, ONNX and TensorRT are used to accelerate the quantized distilled speech generation large language model.

[0078] Among them, the model weights are the parameters in the model used to represent the relationship between the input and output. In a neural network, the weights determine the output of each layer of neurons. The activation values are the output values of neurons in each layer of the neural network, and these output values affect the input of the next layer. ONNX (Open Neural Network Exchange) is an open format for representing deep learning models, enabling the model to be migrated and optimized between different frameworks (such as PyTorch, TensorFlow, etc.). ONNX aims to improve cross-platform compatibility. TensorRT is designed specifically for accelerating deep learning inference tasks, improving the inference efficiency by optimizing the neural network model and utilizing the hardware acceleration of NVIDIA GPUs.

[0079] Specifically, the model weights and activation values are extracted from the already trained distilled speech generation large language model and then converted into floating-point number representations with a preset number of bits, reducing the storage overhead and computational complexity of the model. In one example, during the quantization process, the quantization ratio can be determined by calculating the average activation value of each layer, and then each value is mapped to the 16-bit floating-point number range, thus significantly reducing the memory occupancy and computational overhead of the model without sacrificing accuracy significantly.

[0080] It can be understood that by quantizing the model weights and activation values of the distilled speech generation large language model into floating-point numbers with a preset number of bits, the memory requirements of the model can be effectively reduced and the inference process can be accelerated. Quantization not only reduces the storage space and is applicable to devices with limited memory, but also significantly improves the computing efficiency. Especially in real-time speech generation and embedded applications, it can greatly improve the response speed. Since the floating-point numbers with a preset number of bits provide sufficient precision, it generally does not significantly affect the output quality of the model, thus ensuring the fluency and naturalness of speech generation. While ensuring high performance, it can also reduce the consumption of computing resources and improve the speech generation efficiency.

[0081] After quantization is completed, model acceleration is performed. Specifically, ONNX provides cross-framework compatibility for the model, enabling the model to be migrated between different platforms and further optimized in the TensorRT engine. TensorRT will utilize GPU hardware acceleration and adopt efficient matrix operations and memory management techniques, reducing the consumption of computing resources during model inference and improving the processing speed. Through such an acceleration mechanism, the speech generation large language model can still maintain a fast response time and high efficiency with lower hardware resources.

[0082] In one embodiment, as Figure 3 shown, the present application provides a speech generation method, which includes:

[0083] S201: Obtain a target generation text and a target reference speech;

[0084] S202: According to the target generation text and the target reference speech, use the target speech generation large language model to obtain a target synthesized speech, and the target speech generation large language model is obtained by using the speech generation model acquisition method described in any one of the above embodiments.

[0085] Among them, the target generation text refers to the text content generated according to the input conditions in the speech generation task and serves as the text basis for speech generation. The target reference speech refers to the speech sample associated with the target generation text and is used to guide the speech generation process.

[0086] Specifically, first, the target generation text and the target reference speech are obtained as input data, and then the target speech generation large language model is used to process the target generation text and the target reference speech to obtain a natural speech that generates the target generation text from the target reference speech, that is, the target synthesized speech.

[0087] In this embodiment, the execution manner of inputting the target generation text and the target reference speech into the target speech generation large language model can significantly improve the efficiency and quality of speech synthesis.

[0088] The speech generation device provided by the embodiments of the present application will be described below. The speech generation device described below can be correspondingly referred to the speech generation method described above. As Figure 4 shown, the present application provides a speech generation device, which includes:

[0089] An optimized speech generation large language model acquisition module 301, configured to train an initial speech generation large language model by using the obtained text token samples and speech token samples to obtain an optimized speech generation large language model. The text token samples are obtained by using a text tokenizer, and the text tokenizer is used to improve the text token compression rate and retain text information. The speech token samples are obtained by using a speech tokenizer, and the speech tokenizer is used to improve the speech token compression rate and retain speech information;

[0090] A target speech generation large language model acquisition module 302, configured to obtain a target speech generation large language model according to the optimized speech large language model, and the target speech generation large language model is used for speech synthesis.

[0091] In one embodiment, the optimized speech generation large language model acquisition module 301 includes:

[0092] A text token sample acquisition unit, configured to obtain a generated text sample, and discretize the generated text sample by using a text tokenizer to obtain text token samples. The text tokenizer is a Qwen2 tokenizer.

[0093] In one embodiment, the optimized speech generation large language model acquisition module 301 includes:

[0094] A speech token sample acquisition unit, configured to obtain a reference speech sample, and use a speech tokenizer to discretize each second of speech in the reference speech sample into M token samples to obtain speech token samples. The speech tokenizer is composed of residual vector quantization and a ConvNeXt architecture, and M is a positive integer and M < 50.

[0095] In one embodiment, the target speech generation large language model acquisition module 302 includes:

[0096] A distilled speech generation large language model acquisition unit, configured to distill the optimized speech generation large language model based on a preset objective function to obtain a distilled speech generation large language model, where the model architectures of the optimized speech generation large language model and the distilled speech generation large language model are the same;

[0097] A target speech generation large language model acquisition unit, configured to perform quantization on the distilled speech generation large language model and then accelerate it through an inference engine to obtain a target speech generation large language model.

[0098] In one embodiment, the preset objective function is:

[0099]

[0100] Among them, represents the target preset function, and represent hyperparameters, represents the supervised cross-entropy loss of the true speech label corresponding to the speech token sample, represents the distilled speech generation large language model, represents the text token sample, represents the speech token sample, represents the true speech label corresponding to the speech token sample, represents the KL divergence distribution difference between the optimized speech generation large language model and the distilled speech generation large language model, represents the optimized speech generation large language model.

[0101] In one embodiment, the target speech generation large language model acquisition unit includes:

[0102] The model deployment optimization subunit is used to quantize the model weights and activation values of the distilled speech generation large language model into floating-point numbers with a preset number of bits, and then use ONNX and TensorRT to accelerate the quantized distilled speech generation large language model.

[0103] In one embodiment, as Figure 5 shown, the present application provides a speech generation device, and the device includes:

[0104] The target reference speech acquisition module 401 is used to acquire the target generated text and the target reference speech;

[0105] The target synthesized speech acquisition module 402 is used to obtain the target synthesized speech according to the target generated text and the target reference speech by using the target speech generation large language model, and the target speech generation large language model is obtained by using the speech generation model acquisition device in the above-mentioned embodiment.

[0106] In one embodiment, the present application further provides a storage medium, in which computer-readable instructions are stored, and when the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the speech generation model acquisition method described in any one of the above-mentioned embodiments, and / or the steps of the speech generation method.

[0107] In one embodiment, the present application further provides a computer device. Computer-readable instructions are stored in the computer device. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the voice generation model acquisition method described in any one of the above embodiments, and / or the steps of the voice generation method.

[0108] Schematically, as Figure 6 shown, Figure 6 is an internal structural diagram of a computer device provided by an embodiment of the present application. The computer device 500 can be provided as a server. Referring to Figure 6 , the computer device 500 includes a processing component 502, which further includes one or more processors, and memory resources represented by a memory 501 for storing instructions executable by the processing component 502, such as application programs. The application programs stored in the memory 501 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 502 is configured to execute instructions to perform the voice generation model acquisition method of any of the above embodiments, and / or the steps of the voice generation method.

[0109] The computer device 500 may further include a power supply component 503 configured to perform power management of the computer device 500, a wired or wireless network interface 504 configured to connect the computer device 500 to a network, and an input / output (I / O) interface 505. The computer device 500 can operate based on an operating system stored in the memory 501, such as Windows Server TM, Mac OS XTM, Unix TM, Linux TM, Free BSDTM or the like.

[0110] Those skilled in the art can understand that Figure 6 the structure shown in

[0111] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising said element. In this text, "a", "an", "the", "this" and "its" may also include the plural form, unless the context clearly indicates otherwise. "Plural" means at least two, such as 2, 3, 5 or 8, etc. "And / or" includes any and all combinations of the related listed items.

[0112] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0113] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for acquiring a speech generation model, characterized in that: The method comprises: The initial speech generation large language model is trained using the acquired text word unit samples and speech word unit samples to obtain an optimized speech generation large language model, wherein the text word unit samples are acquired using a text word segmenter, and the text word segmenter is used to improve the text word unit compression rate and retain text information; and the speech word unit samples are acquired using a speech word segmenter, and the speech word segmenter is used to improve the speech word unit compression rate and retain speech information; According to the optimized speech large language model, a target speech generation large language model is obtained, and the target speech generation large language model is used for speech synthesis.

2. The method for acquiring a speech generation model according to claim 1, characterized in that: The process of obtaining the text word unit sample includes: A generated text sample is obtained, and the generated text sample is discretized by using the text word segmenter to obtain the text word unit sample, wherein the text word segmenter is a Qwen2 word segmenter.

3. The method for acquiring a speech generation model according to claim 1, characterized in that: The process of acquiring the speech word unit sample includes: A reference speech sample is obtained, and the speech segmenter is used to discretize the speech in the reference speech sample into M words per second to obtain the speech word sample, wherein the speech segmenter is composed of residual vector quantization and ConvNeXt architecture, and M is a positive integer and M<50.

4. The method for acquiring a speech generation model according to claim 1, characterized in that: The step of obtaining a target speech generation large language model according to the optimized speech large language model comprises: Based on a preset objective function, distilling the optimized speech generation large language model to obtain a distilled speech generation large language model, wherein the model architecture of the optimized speech generation large language model is the same as the model architecture of the distilled speech generation large language model; After the distilled speech generation large language model is quantized, it is accelerated through an inference engine to obtain the target speech generation large language model.

5. The method for acquiring a speech generation model according to claim 4, characterized in that: The preset objective function is: in, represents the target preset function, and represents the hyperparameter, represents the supervised cross entropy loss of the real speech label corresponding to the speech word sample, represents the distilled speech generation large language model, represents the text word sample, represents the speech word sample, represents the real speech label corresponding to the speech word sample, represents the difference in KL divergence distribution between the optimized speech generation large language model and the distilled speech generation large language model, Represents the optimized speech generation large language model.

6. The method for acquiring a speech generation model according to claim 4, characterized in that: The step of quantizing the distilled speech to generate a large language model and then accelerating it through an inference engine comprises: After the model weights and activation values ​​of the distilled speech generation large language model are quantized into floating-point numbers with a preset number of digits, ONNX and TensorRT are used to accelerate the quantized distilled speech generation large language model.

7. A speech generation method, characterized in that: The method comprises: Obtain target generated text and target reference speech; According to the target generated text and the target reference speech, a large language model is generated using the target speech to obtain a target synthesized speech, wherein the large language model generated by the target speech is obtained by using the speech generation model acquisition method as described in any one of claims 1 to 6.

8. A speech generation model acquisition device, characterized in that: The device comprises: An optimized speech generation large language model acquisition module is used to train the initial speech generation large language model using the acquired text word unit samples and speech word unit samples to obtain an optimized speech generation large language model, wherein the text word unit samples are acquired using a text word segmenter, and the text word segmenter is used to improve the text word unit compression rate and retain text information, and the speech word unit samples are acquired using a speech word segmenter, and the speech word segmenter is used to improve the speech word unit compression rate and retain speech information; The target speech generation large language model acquisition module is used to obtain the target speech generation large language model according to the optimized speech large language model, and the target speech generation large language model is used for speech synthesis.

9. A speech generating device, characterized in that: The device comprises: A target reference speech acquisition module is used to acquire target generated text and target reference speech; The target synthesized speech acquisition module is used to generate a large language model based on the target text and the target reference speech, and obtain the target synthesized speech. The target speech generation large language model is obtained by using the speech generation model acquisition device as described in claim 8.

10. A computer device, characterized in that: include: one or more processors, and memory; The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the speech generation model acquisition method according to any one of claims 1 to 6 and / or the steps of the speech generation method according to claim 7 are performed.