Voice generation method and device based on audio prompt, equipment and medium

By splicing the features of the target text and reference audio and inputting the trained speech generation model, the problem of inefficient speech generation in the prior art is solved, and fast and accurate speech synthesis is achieved, and the generation efficiency is improved.

CN119964547APending Publication Date: 2025-05-09PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510243850.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Existing speech generation models are inefficient in training and inference, resulting in increased device burden and slower model generation.

Method used

By obtaining the target text and reference audio of the speech to be generated, a pre-trained text feature extractor is used to perform multi-level feature extraction, and the audio prompt features are spliced ​​with text features, and a pre-trained speech generation model is input to generate the target speech. This speech generation model is obtained by training the preset stream model for speech mask generation.

Benefits of technology

It improves the efficiency of voice generation based on audio prompts, realizes rapid and accurate output of synthesized speech, reduces complex operations between text and speech, and reduces the amount of computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964547A_ABST
    Figure CN119964547A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a voice generation method and device based on audio prompt, equipment and a medium. Performing multi-level feature extraction on the target text through a pre-trained text feature extractor to obtain multi-level text features; generating a corresponding audio prompt feature according to the reference audio, and splicing the multi-level text feature with the audio prompt feature to obtain a spliced input feature; and inputting the splicing input feature into a pre-trained voice generation model, and generating a target voice corresponding to the target text, the voice generation model being obtained by performing voice mask generation training on a preset stream model. The text and the voice are subjected to feature splicing and then input into the model obtained based on voice mask generation training for voice generation, additional complex operation between the text and the voice is not needed, and the voice generation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, equipment and medium for generating speech based on audio prompts. Background Art

[0002] At present, the research on speech generation models has made great progress. It can simulate and generate the speaker's voice for any given text through a few seconds of audio prompts, providing a variety of high-quality synthetic voices for voice assistants, audio books and entertainment, and has broad application prospects in various fields such as the medical field and the financial field. For example, in the field of medical health, medical institutions can use intelligent voice assistants to help patients obtain medical information, make appointments, consult doctors and other services, and convey the corresponding service information to patients in the form of voice, or the hospital's voice navigation system can also provide real-time voice guidance to facilitate patients' actions within the hospital; for example, in the field of financial technology business, the intelligent voice robot of financial institutions can automatically make outbound calls through voice generation technology based on audio prompts, complete personalized recommendations and collection, etc., to improve work efficiency.

[0003] The current mainstream speech generation models basically have the following modules, such as duration module, text encoder or phoneme alignment module, etc. These modules add a certain burden to the training of the overall model and continuously increase the requirements for equipment, which not only greatly limits the training efficiency, but also limits the speed of model inference generation. Summary of the invention

[0004] In view of the above-mentioned deficiencies in the prior art, the purpose of the present invention is to provide a speech generation method, device, equipment and medium based on audio prompts that can be applied to the medical field, financial technology or other related fields. Its main purpose is to improve the efficiency of speech generation based on audio prompts and to achieve fast and accurate output of synthesized speech.

[0005] The technical solution of the present invention is as follows:

[0006] A first aspect of the present invention provides a method for generating speech based on audio prompts, comprising:

[0007] Obtain target text and reference audio for speech to be generated;

[0008] Performing multi-level feature extraction on the target text by a pre-trained text feature extractor to obtain multi-level text features;

[0009] Generate a corresponding audio prompt feature according to the reference audio, and concatenate the multi-level text feature with the audio prompt feature to obtain a concatenated input feature;

[0010] The concatenated input features are input into a pre-trained speech generation model to generate a target speech corresponding to the target text, wherein the speech generation model is obtained by training a preset stream model for speech mask generation.

[0011] A second aspect of the present invention provides a speech generation device based on audio prompts, comprising:

[0012] An acquisition module is used to acquire the target text and reference audio of the speech to be generated;

[0013] A text processing module, used for performing multi-level feature extraction on the target text through a pre-trained text feature extractor to obtain multi-level text features;

[0014] A feature splicing module, used to generate corresponding audio prompt features according to the reference audio, and splice the multi-level text features with the audio prompt features to obtain spliced ​​input features;

[0015] The speech generation module is used to input the concatenated input features into a pre-trained speech generation model to generate a target speech corresponding to the target text, wherein the speech generation model is obtained by training a preset stream model for speech mask generation.

[0016] A third aspect of the present invention provides a computer device, comprising at least one processor; and

[0017] a memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above-mentioned voice generation method based on audio prompts.

[0019] A fourth aspect of the present invention provides a non-volatile computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by one or more processors, the one or more processors can execute the above-mentioned audio prompt-based speech generation method.

[0020] Beneficial effects: The present invention discloses a method, device, equipment and medium for speech generation based on audio prompts. Compared with the prior art, the embodiments of the present invention obtain the target text and reference audio of the speech to be generated; perform multi-level feature extraction on the target text through a pre-trained text feature extractor to obtain multi-level text features; generate corresponding audio prompt features according to the reference audio, and splice the multi-level text features with the audio prompt features to obtain spliced ​​input features; input the spliced ​​input features into a pre-trained speech generation model to generate a target speech corresponding to the target text, and the speech generation model is obtained by training a preset stream model for speech mask generation. By splicing the input text with the reference audio as the prompt information for speech generation, there is no need to perform additional complex operations between text and speech. Under the guidance of the prompt information, the corresponding target speech can be obtained by mapping the spectrum distribution through the speech generation model, which effectively improves the speech generation efficiency based on audio prompts and achieves fast and accurate output of synthesized speech. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the scheme in the present invention, a brief introduction is given below to the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0022] Figure 1 A schematic diagram of an application environment of the method for generating speech based on audio prompts provided by an embodiment of the present invention;

[0023] Figure 2 A flow chart of a method for generating speech based on audio prompts provided by an embodiment of the present invention;

[0024] Figure 3 A flow chart of step S203 in the method for generating speech based on audio prompts provided in an embodiment of the present invention;

[0025] Figure 4 A flow chart of step S204 in the method for generating speech based on audio prompts provided in an embodiment of the present invention;

[0026] Figure 5 Another flow chart of the method for generating speech based on audio prompts provided by an embodiment of the present invention;

[0027] Figure 6 A schematic diagram of functional modules of a speech generation device based on audio prompts provided in an embodiment of the present invention;

[0028] Figure 7A schematic diagram of the hardware structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical solution and effect of the present invention clearer and more specific, the present invention is further described in detail below. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. The embodiments of the present invention are described below in conjunction with the accompanying drawings.

[0030] The speech generation method based on audio prompts provided in the embodiment of the present invention can be applied in the following aspects: Figure 1 In the application environment of FIG. 1 , it includes a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104 and a server 105. The network 104 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0031] The user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as examples).

[0032] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0033] The server 105 may be a server that provides various services, such as a background server that provides support for the content browsed by the user using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (for example only). The background server may analyze and process the received data such as the user request, and feed back the processing results (such as web pages, information, or data obtained or generated according to the user request) to the terminal device. The server 105 may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server 105 may also be a server for a distributed system, or a server combined with a blockchain.

[0034] It should be noted that the speech generation method based on audio prompts provided in the embodiment of the present application can generally be executed by the first terminal device 101, the second terminal device 102 or the third terminal device 103. Accordingly, the speech generation device based on audio prompts provided in the embodiment of the present invention can also be set in the first terminal device 101, the second terminal device 102 or the third terminal device 103. Alternatively, the speech generation method based on audio prompts provided in the embodiment of the present invention can also generally be executed by the server 105. Accordingly, the speech generation device based on audio prompts provided in the embodiment of the present invention can generally be set in the server 105.

[0035] It should be understood that the numbers of the above terminal devices, networks and servers are only illustrative. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0036] like Figure 2 As shown, the speech generation method based on audio prompts provided in the embodiment of the present invention specifically includes the following steps:

[0037] S201: Obtain target text and reference audio of speech to be generated.

[0038] In this embodiment, the target text is the text content that needs to be converted into speech, and can be any text, such as medical reports, financial statements, user instructions, etc. Specifically, the target text can be obtained by manual input, such as the user inputs the text content that needs to be converted into speech through the interface; or read from a text file, such as reading a text file (such as TXT, JSON format) from a local or server; or generated in real time, such as combining natural language processing (NLP) technology to dynamically generate text according to instructions such as user queries or system feedback, etc.

[0039] Reference audio is used to provide features such as voice style, intonation, and speaking speed, so as to guide the style of generated voice and make it more suitable for specific scenarios. Specifically, there are various ways to obtain reference audio. It can be uploaded by users, such as users uploading specific reference voice samples to define voice style; or obtained from a preset audio library, such as selecting standard voice samples from a predefined audio library, such as a doctor's professional explanation or the standard voice of a financial customer service representative; or it can be obtained through real-time recording, such as recording the user's voice in real time through a microphone, so as to guide the generation of voice similar to the user's voice, etc. Obtaining the target text and reference audio through the above-mentioned methods provides clear input and voice feature guidance for subsequent voice generation, ensuring that the generated voice not only accurately conveys the text content, but also adapts to the needs of different users for generated voice characteristics such as timbre and style.

[0040] For example, in the application scenario of the medical field, the target text can be the patient's medical record description, such as "the patient's temperature is 38.5℃, further examination is required", and the reference audio is an audio clip captured from the doctor's professional explanation voice (such as medical science lectures, etc.), so that the generated voice will have the doctor's professional tone, helping patients better understand medical advice. In the application scenario of the financial field, the target text can be the customer service reply content, such as "your account balance is 10,000 yuan", and the reference audio is the standard voice of the financial customer service recorded in advance, so that the generated voice will have the professionalism and intimacy of the customer service, improving the user experience.

[0041] In one embodiment, after step S201, the method further includes:

[0042] Determine a target sampling rate, and convert the sampling rate of the reference audio to the target sampling rate through a resampling process;

[0043] The reference audio after the sampling rate is adjusted is subjected to pre-emphasis processing and normalization processing to obtain a pre-processed audio.

[0044] In this embodiment, after the reference audio is obtained, data preprocessing is also performed on it to facilitate subsequent feature extraction and splicing. Specifically, the target sampling rate is first determined. The sampling rate refers to the number of times the audio signal is sampled per unit time. The target sampling rate can be the sampling rate of different application scenarios or the sampling rate consistent with the data in the training phase, etc. The sampling rate of the reference audio is converted to the target sampling rate through resampling processing. Specifically, it can be resampled on the time axis through interpolation algorithms such as linear interpolation and spline interpolation, so as to ensure that the reference audio is compatible with the subsequent speech generation module and avoid audio distortion caused by sampling rate mismatch.

[0045] Furthermore, in order to improve the clarity of speech, the reference audio after adjusting the sampling rate is further subjected to pre-emphasis processing and normalization processing, wherein the pre-emphasis processing is to improve the clarity and intelligibility of speech by enhancing high-frequency signals, for example, the high-frequency signals can be enhanced by a pre-emphasis filter. The normalization processing is to adjust the amplitude of the audio signal to a standard range (such as [-1,1] or [0,1]) to eliminate the amplitude difference between different audios, standardize the audio signal, and obtain pre-processed audio that is convenient for subsequent processing.

[0046] For example, in a medical scenario, when the reference audio is the doctor's explanation voice, after the sampling rate of the explanation voice is converted to the target sampling rate, it is further processed by pre-emphasis and normalization to make the explanation voice clearer, and the high-frequency information (such as keywords such as "body temperature" and "check") is more prominent, providing reliable prompt information for subsequent voice generation. In a financial scenario, when the reference audio is the standard voice of the customer service, after the sampling rate of the standard voice is converted to the target sampling rate, it is further processed by pre-emphasis and normalization to enhance the high-frequency signal, make the voice clearer, and improve the subsequent voice generation effect.

[0047] S202: Perform multi-level feature extraction on the target text using a pre-trained text feature extractor to obtain multi-level text features.

[0048] In this embodiment, the acquired target text is input into a pre-trained text feature extractor for multi-level feature extraction. The text data is converted into structured features through text feature extraction so that the computer can understand and process it. The multi-level feature extraction can further capture text information from different levels including word level, sentence level, chapter level, etc., such as the semantics, grammar and contextual information of the text. Specifically, the text features can be processed in a fine-grained manner through the ConveXt V2 module. ConvNeXt V2 is an improved convolutional neural network structure that can capture multi-level text features so that the text can be better aligned with the speech. The target text is input into the trained ConvNeXt V2 module, and the input text is represented as a multi-level feature to facilitate the generation model to understand the text information at different levels, so that the generated speech is more natural in tone, rhythm and rhythm. Alternatively, in other embodiments, multi-level feature extraction can be achieved through other text feature extractors, such as the RoBERTa model that captures cross-sentence and cross-paragraph dependencies in the text through a multi-layer self-attention mechanism, thereby extracting multi-level contextual features, etc. This embodiment is not limited to this.

[0049] S203: Generate corresponding audio prompt features according to the reference audio, and concatenate the multi-level text features with the audio prompt features to obtain concatenated input features.

[0050] In this embodiment, the corresponding audio prompt features are generated after feature extraction based on the reference audio, thereby providing prompt information such as the style, intonation, timbre, and speech speed of the voice. Specifically, this can be achieved by extracting the Mel frequency cepstral coefficients (MFCC) of the audio or other audio feature parameters. The generated audio prompt features are spliced ​​with the multi-level text features to form a complete input feature, which is used to guide the speech generation model to generate a synthetic speech that meets the specific voice characteristics. The specific feature splicing method can be direct splicing or weighted fusion, etc., wherein direct splicing is to splice the multi-level text features and the audio prompt features by dimension to form a complete input vector, and weighted fusion is to weight different features according to the application scenario to enhance the influence of specific features. By combining the multi-level text features and the audio prompt features to guide the subsequent speech generation process, the generated speech not only conforms to the content of the target text, but also inherits the speech characteristics of the reference audio in terms of timbre, intonation, etc., and improves the naturalness and consistency of the speech.

[0051] S204, inputting the concatenated input features into a pre-trained speech generation model to generate a target speech corresponding to the target text, wherein the speech generation model is obtained by performing speech mask generation training on a preset stream model.

[0052] In this embodiment, the spliced ​​input features are input into a speech generation model that has been trained for speech mask generation. The speech generation model is based on a flow model, such as a Rectified Flow model, a conditional flow matching model, etc. During the training process, the text data and the masked speech data are used to learn the process of conversion and mapping between two distribution samples, namely, the masked speech sample and the real speech sample, so as to achieve the conversion of the distribution of the two speech data to achieve the speech synthesis effect, so that the trained speech generation model can convert a noise mapping into the corresponding target speech by splicing the text and voice prompts in the input features. The content of the target speech is consistent with the content of the target text, and the target speech has speech characteristics similar to the reference audio.

[0053] For example, in a medical scenario, the target text and the doctor's reference audio can be used to generate voice prompts for medical equipment or voice guidance for telemedicine, etc., to help patients better understand medical information; or in a financial scenario, the target text and the customer service's reference audio can be used to generate voice prompts for financial account operations or remote operation guidance, etc., to help customers quickly understand account information or operation procedures. In the above speech generation process, there is no need to perform additional complex operations between text and speech. Under the guidance of the spliced ​​input features, the corresponding target speech can be obtained by mapping the spectral distribution through the speech generation model, which greatly reduces the amount of calculation and effectively improves the efficiency of speech generation based on audio prompts.

[0054] In the above embodiment, the present invention discloses a method for speech generation based on audio prompts, which includes obtaining a target text and a reference audio of the speech to be generated; performing multi-level feature extraction on the target text through a pre-trained text feature extractor to obtain multi-level text features; generating corresponding audio prompt features according to the reference audio, and splicing the multi-level text features with the audio prompt features to obtain spliced ​​input features; inputting the spliced ​​input features into a pre-trained speech generation model to generate a target speech corresponding to the target text, wherein the speech generation model is obtained by training a preset stream model for speech mask generation. By splicing the input text with the reference audio as the prompt information for speech generation, there is no need to perform additional complex operations between text and speech. Under the guidance of the prompt information, the corresponding target speech can be obtained by mapping the spectrum distribution through the speech generation model, which effectively improves the efficiency of speech generation based on audio prompts and achieves fast and accurate output of synthesized speech.

[0055] In one embodiment, Figure 3 As shown, step S203 includes:

[0056] S301, converting the reference audio into a reference Mel spectrum, and adding a randomly initialized noise spectrum to the reference Mel spectrum to obtain a corresponding audio prompt feature;

[0057] S302: Fill the multi-level text feature according to the length of the audio prompt feature, and concatenate the filled text feature with the audio prompt feature to obtain the concatenated input feature.

[0058] In this embodiment, when generating spliced ​​input features based on reference audio and multi-level text features, the reference audio is first converted into a signal, and the reference audio is converted into a corresponding reference Mel spectrum by means of a Mel filter bank or the like. Mel spectrum is a spectrum representation method that converts the frequency of an audio signal into a Mel scale. It is based on the characteristics of the human auditory system and converts the sound signal into a frequency domain representation through a series of mathematical transformations, thereby capturing the speech features well and facilitating further analysis and processing. Then, a randomly initialized noise spectrum is spliced ​​into the Mel spectrum. Since the speech generation model is obtained after the speech mask generation training of the convection model, it can realize the mapping conversion of the spectrum distribution of an initial noise to the target spectrum distribution. A randomly initialized noise spectrum is spliced ​​at the end of the reference Mel spectrum to obtain the corresponding audio prompt feature so as to realize the subsequent efficient speech generation.

[0059] Furthermore, since the length of the multi-level text feature and the length of the audio prompt feature after adding the noise spectrum may be inconsistent, in order to enable them to be spliced, the multi-level text feature needs to be padded. The specific filling method can be to use zero padding or other preset padding codes to fill it so that its length is consistent with the audio prompt feature and then splice it to obtain a complete input feature. The input features after padding and splicing can better integrate the information of the text and the reference audio, and the noise spectrum in the spliced ​​input features is the part that the speech generation model needs to predict the speech, so that the distribution mapping capability of the flow model is utilized, and the noise spectrum can be converted into the target spectrum without additional alignment operations on the text and speech during the speech generation process, thereby generating the target speech more naturally and efficiently, and improving the speech synthesis efficiency.

[0060] In one embodiment, Figure 4 As shown, step S204 includes:

[0061] S401, inputting the concatenated input features into a pre-trained speech generation model, converting the noise spectrum into a speech spectrum distribution under the guidance of the generation of the reference Mel spectrum and the multi-level text features, and generating a corresponding predicted Mel spectrum;

[0062] S402: Decode the predicted Mel-spectrogram through a vocoder to generate a target speech corresponding to the target text.

[0063] In this embodiment, the spliced ​​input features are input into the pre-trained speech generation model, and the spliced ​​input features include the reference Mel spectrum, the multi-level text features with the same length after padding, and the noise spectrum to be predicted. The combination of these features provides the speech generation model with rich context information to achieve speech generation guidance. The speech generation model learns the mapping relationship from the noise distribution to the target speech spectrum distribution, i.e., the masked part, through the training of speech mask generation in the training phase, thereby realizing the text-to-speech generation process of the masked part. Therefore, in the inference generation phase, the speech generation model can use the relationship between text and speech learned in the training phase to convert the noise spectrum into the speech spectrum distribution under the guidance of the generation of the reference Mel spectrum and the multi-level text features, and convert the noise spectrum into a predicted Mel spectrum that is consistent with the target text content and the speech characteristics are consistent with the reference Mel spectrum. No additional text semantic alignment module is required, and the generation processing time is saved while generating natural flow speech.

[0064] After the predicted Mel spectrum is output, the predicted Mel spectrum is decoded by the vocoder. The vocoder is a tool for converting the Mel spectrum into a time-domain audio signal, which can be implemented by vocoders such as WaveNet and HiFi-GAN. The predicted Mel spectrum is decoded into a time-domain audio signal by the vocoder to restore the natural and clear target speech. The target speech content generated is consistent with the target text, and the speech characteristics are consistent with the reference audio, realizing an efficient text-to-speech generation process based on audio prompts, meeting the needs of different scenarios such as medical and financial.

[0065] In one embodiment, Figure 5 As shown, the process of obtaining a speech generation model after performing speech mask generation training on a preset stream model includes:

[0066] S501, collecting model training data, where the model training data includes sample text and sample Mel spectrum corresponding to the sample text;

[0067] S502, performing multi-level feature extraction on the sample text by a pre-trained text feature extractor, and aligning the feature extraction result with the sample Mel spectrum to obtain a sample multi-level text feature;

[0068] S503, randomly masking the sample Mel spectrum according to a preset masking strategy to obtain a Mel spectrum to be predicted including the masked part;

[0069] S504, splicing the multi-level text features of the sample and the Mel spectrum to be predicted and inputting them into a preset stream model, performing speech prediction on the mask part, and obtaining a masked reconstructed spectrum;

[0070] S505: Adjust parameters of the stream model based on the difference between the masked reconstructed spectrum and the real spectrum corresponding to the masked part until a trained speech generation model is obtained when a preset convergence condition is met.

[0071] In this embodiment, during the training phase, model training data is collected to train a pre-built streaming model, wherein the model training data includes sample text and a sample Mel spectrum corresponding to the sample text. For example, the sample text may be a sentence or paragraph extracted from a corpus, and the sample Mel spectrum is obtained by speech conversion corresponding to the sample text, thereby providing a reliable data basis for subsequent model training. The pre-built flow model is composed of multiple transformer modules. Generally, the mainstream transformer module includes a layer normalization, a multi-head self-attention module, a layer normalization, and a fully connected layer. The transformer module in the flow model constructed in this embodiment adds a downsampling module in front of the multi-head self-attention module, that is, a layer normalization, a downsampling layer, a multi-head self-attention module, a layer normalization, and a fully connected layer. After the training data is passed in, it passes through these five modules in sequence. Similarly, it is also necessary to perform a residual operation after passing through the first three modules and the last two modules respectively. Compared with the mainstream transformer module, the transformer module used in the speech generation model in this embodiment can reduce the dimension of the matrix for matrix multiplication in the multi-head self-attention module by adding a downsampling layer in front of the multi-head self-attention module, thereby significantly reducing the amount of calculation.

[0072] After collecting the model training data, similar to the inference process, in the training stage, the sample text is first subjected to multi-level feature extraction through the text feature extractor, and the feature extraction result is aligned with the sample Mel spectrum to obtain the sample multi-level text feature. Specifically, the sample text can also be processed through the ConveXt V2 module, and the length of the converted text feature is ensured to be the same as the length of the speech. For the speech data, i.e., the sample Mel spectrum, some parts of the sample Mel spectrum are randomly selected and masked (for example, replaced with zero or a specific value) according to the preset masking strategy. The purpose of random masking is to let the model learn how to predict these masked parts from the context, so as to obtain the predicted Mel spectrum containing the masked part, that is, the length of the predicted Mel spectrum is the same as the original sample Mel spectrum, but it contains the masked part and the non-masked part, where the non-masked part corresponds one-to-one with the sample text at the corresponding position, and the masked part requires the model to predict based on the sample text at the corresponding position, so as to learn the relationship between text and speech.

[0073] The multi-level text features and the Mel spectrum to be predicted with the masked part are spliced ​​together as the input of the stream model constructed above. The stream model predicts the real spectrum of the masked part by learning the mapping relationship from text features to acoustic features, that is, the speech prediction is performed on the masked part to obtain the corresponding masked reconstructed spectrum, which is then compared with the real spectrum (that is, the sample Mel spectrum corresponding to the masked part). Based on the difference between the masked reconstructed spectrum and the real spectrum (such as the mean square error), the loss function is minimized through back propagation and adjustment of model parameters. When the loss function reaches the preset threshold or the training reaches the specified number of iterations, the model training is completed, and a trained speech generation module is obtained.

[0074] In this embodiment, the flow model can sample the Rectified Flow module, which is similar to the conditional flow matching model in that both construct an ordinary differential equation , when t=0, X t It obeys Gaussian distribution, that is, X0 is the noise spectrum, and when t=1, X t is the expected data distribution, that is, X1 is the predicted spectrum corresponding to the masked part. Assuming that the information obtained by text-speech concatenation is μ, the vector field corresponding to the path where data flows from X(t=0) to X(t=1) in Rectified Flow is t is the time. After the model training is completed, an ordinary differential equation can be obtained. For any t in [0,1], a v can be calculated by the conditional flow matching module t (X t |μ), so that the ordinary differential equation can be numerically solved to achieve the conversion from noise X (t = 0) to target speech X (t = 1).

[0075] The speech generation model obtained by training the convection model in this embodiment no longer requires a complex text-to-speech alignment module, and does not require additional complex operations between text and speech. It only needs simple filling, masking and splicing operations to learn complex mapping relationships from text and acoustic features, and enhance the robustness and generalization ability of the model through mask prediction, which can greatly reduce the amount of calculation and ensure that the model can generate high-quality and natural speech in a variety of scenarios.

[0076] In one embodiment, step S502 includes:

[0077] Performing multi-level feature extraction on the sample text by a pre-trained text feature extractor to obtain initial text features;

[0078] Comparing the length of the initial text feature with the length of the sample Mel spectrum;

[0079] If the length of the initial text feature is smaller than the length of the sample Mel spectrum, the initial text feature is padded with a preset padding code to obtain a sample multi-level text feature, and the degree of the sample multi-level text feature is the same as the length of the sample Mel spectrum.

[0080] In this embodiment, when processing the sample text in the training phase, the sample text is first subjected to multi-level feature extraction by a text feature extractor to obtain the initial text feature, for example, feature extraction is performed on the sample text "The patient's body temperature is 38.5°C, further examination is required." The length of the extracted initial text feature is 10, and the corresponding sample Mel spectrum is obtained by voice conversion of the doctor reading this sentence in a standard tone, and the Mel spectrum length obtained by its length may be 20. Since the length of the initial text feature (10) is less than the length of the Mel spectrum (20), an alignment operation is required to fill the initial text feature with a corresponding number of preset filling codes, such as 10 zeros or other data, so that the length of the filled text feature reaches 20, and the sample multi-level text feature is obtained. By inputting the data into the model for simple alignment and padding processing before speech generation, the length of the filled sample multi-level text feature is consistent with the length of the sample Mel spectrum, which is convenient for subsequent splicing and training, and there is no need to perform additional complex alignment generation operations on the text and speech in the speech generation phase, saving data processing and improving training efficiency and generation efficiency.

[0081] In one embodiment, step S503 includes:

[0082] Performing zero-mean normalization processing on the sample Mel spectrum to obtain a sample normalized spectrum;

[0083] A number of continuous or non-continuous regions are randomly selected in the sample normalized spectrum, and mask processing is performed on the selected regions to obtain a Mel spectrum to be predicted including the masked portion.

[0084] In this embodiment, when the training speech data, i.e., the sample Mel spectrum, is randomly masked during the training phase, the sample Mel spectrum is firstly normalized by zero mean, so that the mean of the data is adjusted to zero while keeping the distribution characteristics of the data unchanged. Specifically, each spectrum feature value is subtracted from its mean, so that the mean of the normalized data is zero, and the sample normalized spectrum is obtained. The zero mean normalization process can ensure that the characteristic distribution of the Mel spectrum is more uniform, thereby improving the model's learning ability for speech features. Afterwards, in the normalized Mel spectrum, several continuous or non-continuous regions are randomly selected, which can be time periods or frequency segments of the spectrum, and the selected regions are masked, that is, the values ​​of these regions are replaced with specific mask values ​​(such as zero or random noise, etc.). By randomly masking certain regions, the model learns how to reconstruct the masked part from the context, so that the model can learn how to reconstruct the missing part from the incomplete data based on text and voice prompts, and help the model better learn the contextual relationship of speech features, thereby generating more natural and fluent speech.

[0085] It should be noted that there is not necessarily a certain order between the above-mentioned steps. A person skilled in the art can understand, based on the description of the embodiments of the present invention, that in different embodiments, the above-mentioned steps may have different execution orders, that is, they may be executed in parallel, may be executed interchangeably, and so on.

[0086] Further references Figure 6 , as a response to the above Figure 2 The present invention provides an embodiment of a speech generation device based on audio prompts, and the device embodiment is similar to Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0087] like Figure 6 As shown, the speech generation device 60 based on audio prompts described in this embodiment includes:

[0088] An acquisition module 601 is used to acquire a target text and a reference audio of a speech to be generated;

[0089] The text processing module 602 is used to perform multi-level feature extraction on the target text through a pre-trained text feature extractor to obtain multi-level text features;

[0090] A feature splicing module 603 is used to generate a corresponding audio prompt feature according to the reference audio, and splice the multi-level text feature with the audio prompt feature to obtain a spliced ​​input feature;

[0091] The speech generation module 604 is used to input the concatenated input features into a pre-trained speech generation model to generate a target speech corresponding to the target text, wherein the speech generation model is obtained by training a preset stream model for speech mask generation. The module referred to in the present invention refers to a series of computer program instruction segments that can perform specific functions, which is more suitable for describing the speech generation execution process based on audio prompts than a program. For the specific implementation of each module, please refer to the corresponding method embodiment above, which will not be repeated here.

[0092] In one embodiment, the feature stitching module 603 includes:

[0093] An audio processing unit, used to convert the reference audio into a reference Mel spectrum, and add a randomly initialized noise spectrum to the reference Mel spectrum to obtain a corresponding audio prompt feature;

[0094] A filling and splicing unit is used to fill the multi-level text feature according to the length of the audio prompt feature, and splice the filled text feature with the audio prompt feature to obtain the spliced ​​input feature.

[0095] In one embodiment, the speech generation module 604 includes:

[0096] A guided generation unit, used for inputting the concatenated input features into a pre-trained speech generation model, converting the noise spectrum into a speech spectrum distribution under the guidance of the generation of the reference Mel spectrum and the multi-level text features, and generating a corresponding predicted Mel spectrum;

[0097] A decoding unit is used to decode the predicted Mel spectrum through a vocoder to generate a target speech corresponding to the target text.

[0098] In one embodiment, the apparatus 60 further comprises:

[0099] A collection module, used for collecting model training data, wherein the model training data includes sample text and sample Mel spectrum corresponding to the sample text;

[0100] A text feature extraction module, used to perform multi-level feature extraction on the sample text through a pre-trained text feature extractor, and align the feature extraction result with the sample Mel spectrum to obtain the sample multi-level text feature;

[0101] A random masking module, used for randomly masking the sample Mel spectrum according to a preset masking strategy to obtain a Mel spectrum to be predicted including a masked part;

[0102] A splicing prediction module is used to splice the multi-level text features of the sample and the Mel spectrum to be predicted and input them into a preset stream model, perform speech prediction on the mask part, and obtain a masked reconstructed spectrum;

[0103] The model parameter adjustment module is used to adjust the parameters of the stream model based on the difference between the masked reconstructed spectrum and the real spectrum corresponding to the masked part, until a trained speech generation model is obtained when a preset convergence condition is met.

[0104] In one embodiment, the text feature extraction module includes:

[0105] A multi-level extraction unit, used for performing multi-level feature extraction on the sample text through a pre-trained text feature extractor to obtain initial text features;

[0106] A length comparison unit, used for comparing the length of the initial text feature with the length of the sample Mel spectrum;

[0107] A text filling unit is used for aligning the initial text feature by filling the initial text feature with a preset filling code to obtain a sample multi-level text feature if the length of the initial text feature is smaller than the length of the sample Mel spectrum, and the degree of the sample multi-level text feature is the same as the length of the sample Mel spectrum.

[0108] In one embodiment, the random mask module includes:

[0109] A normalization unit, used for performing zero-mean normalization processing on the sample Mel spectrum to obtain a sample normalized spectrum;

[0110] The mask unit is used to randomly select a number of continuous or non-continuous areas in the sample normalized spectrum, perform mask processing on the selected areas, and obtain a to-be-predicted Mel spectrum containing the masked part.

[0111] In one embodiment, the apparatus 60 further comprises:

[0112] A conversion module, used to determine a target sampling rate, and convert the sampling rate of the reference audio to the target sampling rate through a resampling process;

[0113] The preprocessing module is used to perform pre-emphasis processing and normalization processing on the reference audio after the sampling rate is adjusted to obtain preprocessed audio.

[0114] In the above embodiment, the present invention discloses a speech generation device based on audio prompts, which obtains the target text and reference audio of the speech to be generated; performs multi-level feature extraction on the target text through a pre-trained text feature extractor to obtain multi-level text features; generates corresponding audio prompt features according to the reference audio, and splices the multi-level text features with the audio prompt features to obtain spliced ​​input features; inputs the spliced ​​input features into a pre-trained speech generation model to generate a target speech corresponding to the target text, and the speech generation model is obtained by training a preset stream model for speech mask generation. By splicing the input text with the reference audio as the prompt information for speech generation, there is no need to perform additional complex operations between text and speech. Under the guidance of the prompt information, the corresponding target speech can be obtained by mapping the spectrum distribution through the speech generation model, which effectively improves the speech generation efficiency based on audio prompts and achieves fast and accurate output of synthesized speech.

[0115] Another embodiment of the present invention provides a computer device, such as Figure 7 As shown, the computer device 70 includes:

[0116] One or more processors 701 and memory 702, Figure 7 A processor 701 is used as an example for description. The processor 701 and the memory 702 may be connected via a bus or other means. Figure 7 The example of connecting through bus is taken in the following.

[0117] Processor 701 is used to complete various control logics of computer device 70, and it can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISC Machine) or other programmable logic device, a discrete gate or transistor logic, a discrete hardware component or any combination of these components. In addition, processor 701 can also be any traditional processor, microprocessor or state machine. Processor 701 can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors combined with a DSP and / or any other such configuration.

[0118] The memory 702 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions corresponding to the method for generating speech based on audio prompts in the embodiment of the present invention. The processor 701 executes various functional applications and data processing of the computer device 70 by running the non-volatile software programs, instructions and units stored in the memory 702, that is, the method for generating speech based on audio prompts in the above method embodiment is implemented.

[0119] The memory 702 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device 70, etc. In addition, the memory 702 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 702 may optionally include a memory remotely arranged relative to the processor 701, and these remote memories may be connected to the computer device 70 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof. One or more units are stored in the memory 702, and when executed by one or more processors 701, the steps of the voice generation method based on audio prompts in any of the above-mentioned method embodiments are executed.

[0120] In the above embodiment, the present invention discloses a computer device, which obtains the target text and reference audio of the speech to be generated; extracts multi-level features of the target text through a pre-trained text feature extractor to obtain multi-level text features; generates corresponding audio prompt features according to the reference audio, and splices the multi-level text features with the audio prompt features to obtain spliced ​​input features; inputs the spliced ​​input features into a pre-trained speech generation model to generate a target speech corresponding to the target text, and the speech generation model is obtained by training a preset stream model for speech mask generation. By splicing the input text with the reference audio as the prompt information for speech generation, there is no need to perform additional complex operations between text and speech. Under the guidance of the prompt information, the corresponding target speech can be obtained by mapping the spectrum distribution through the speech generation model, which effectively improves the efficiency of speech generation based on audio prompts and achieves fast and accurate output of synthesized speech.

[0121] An embodiment of the present invention provides a non-volatile computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by one or more processors, the steps of the speech generation method based on audio prompts in any of the above method embodiments are executed.

[0122] In the above embodiment, the present invention discloses a non-volatile computer-readable storage medium, which obtains a target text and a reference audio of a speech to be generated; extracts multi-level features of the target text through a pre-trained text feature extractor to obtain multi-level text features; generates corresponding audio prompt features according to the reference audio, and splices the multi-level text features with the audio prompt features to obtain spliced ​​input features; inputs the spliced ​​input features into a pre-trained speech generation model to generate a target speech corresponding to the target text, and the speech generation model is obtained by training a preset stream model for speech mask generation. By splicing the input text with the reference audio as prompt information for speech generation, there is no need to perform additional complex operations between text and speech. Under the guidance of the prompt information, the corresponding target speech can be obtained by mapping the spectrum distribution through the speech generation model, which effectively improves the efficiency of speech generation based on audio prompts and achieves fast and accurate output of synthesized speech.

[0123] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present invention.

[0124] The present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0125] In summary, the method, device, equipment and medium for speech generation based on audio prompts disclosed in the present invention include: obtaining the target text and reference audio of the speech to be generated; performing multi-level feature extraction on the target text through a pre-trained text feature extractor to obtain multi-level text features; generating corresponding audio prompt features according to the reference audio, and splicing the multi-level text features with the audio prompt features to obtain spliced ​​input features; inputting the spliced ​​input features into a pre-trained speech generation model to generate a target speech corresponding to the target text, wherein the speech generation model is obtained by training a preset stream model for speech mask generation. By splicing the input text with the reference audio as the prompt information for speech generation, there is no need to perform additional complex operations between text and speech. Under the guidance of the prompt information, the corresponding target speech can be obtained by mapping the spectrum distribution through the speech generation model, which effectively improves the efficiency of speech generation based on audio prompts and achieves fast and accurate output of synthesized speech.

[0126] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium, and the computer program can include the processes of the above-mentioned method embodiments when executed. The storage medium can be a memory, a disk, a floppy disk, a flash memory, an optical storage device, etc.

[0127] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of the present application, they are merely used for illustration and do not represent actual use. It should be understood that the application of the present invention is not limited to the above examples, and that those skilled in the art can make improvements or changes based on the above description, and all such improvements and changes shall fall within the scope of protection of the claims attached to the present invention.

Claims

1. A method for generating speech based on audio prompts, characterized in that: include: Obtain target text and reference audio for speech to be generated; Performing multi-level feature extraction on the target text by a pre-trained text feature extractor to obtain multi-level text features; Generate a corresponding audio prompt feature according to the reference audio, and concatenate the multi-level text feature with the audio prompt feature to obtain a concatenated input feature; The concatenated input features are input into a pre-trained speech generation model to generate a target speech corresponding to the target text, wherein the speech generation model is obtained by training a preset stream model for speech mask generation.

2. The method for generating speech based on audio prompts according to claim 1, characterized in that: The step of generating a corresponding audio prompt feature according to the reference audio, and splicing the multi-level text feature with the audio prompt feature to obtain a spliced ​​input feature includes: Convert the reference audio into a reference Mel spectrum, and add a randomly initialized noise spectrum to the reference Mel spectrum to obtain a corresponding audio prompt feature; The multi-level text feature is padded according to the length of the audio prompt feature, and the padded text feature is spliced ​​with the audio prompt feature to obtain the spliced ​​input feature.

3. The method for generating speech based on audio prompts according to claim 2, characterized in that: The step of inputting the concatenated input features into a pre-trained speech generation model to generate a target speech corresponding to the target text includes: Inputting the concatenated input features into a pre-trained speech generation model, converting the noise spectrum into a speech spectrum distribution under the guidance of the generation of the reference Mel spectrum and the multi-level text features, and generating a corresponding predicted Mel spectrum; The predicted Mel spectrum is decoded by a vocoder to generate a target speech corresponding to the target text.

4. The method for generating speech based on audio prompts according to claim 1, characterized in that: The process of obtaining a speech generation model after performing speech mask generation training on a preset stream model includes: Collecting model training data, wherein the model training data includes sample text and sample Mel spectrum corresponding to the sample text; Performing multi-level feature extraction on the sample text by a pre-trained text feature extractor, and aligning the feature extraction result with the sample Mel spectrum to obtain the sample multi-level text feature; Randomly masking the sample Mel spectrum according to a preset masking strategy to obtain a Mel spectrum to be predicted including the masked part; The multi-level text features of the sample and the Mel spectrum to be predicted are concatenated and input into a preset stream model, speech prediction is performed on the mask part to obtain a masked reconstructed spectrum; The parameters of the stream model are adjusted based on the difference between the masked reconstructed spectrum and the real spectrum corresponding to the masked part until a trained speech generation model is obtained when a preset convergence condition is met.

5. The method for generating speech based on audio prompts according to claim 4, characterized in that: The method of performing multi-level feature extraction on the sample text by using a pre-trained text feature extractor and aligning the feature extraction result with the sample Mel spectrum to obtain the sample multi-level text feature includes: Performing multi-level feature extraction on the sample text by a pre-trained text feature extractor to obtain initial text features; Comparing the length of the initial text feature with the length of the sample Mel spectrum; If the length of the initial text feature is smaller than the length of the sample Mel spectrum, the initial text feature is padded with a preset padding code to obtain a sample multi-level text feature, and the degree of the sample multi-level text feature is the same as the length of the sample Mel spectrum.

6. The method for generating speech based on audio prompts according to claim 4, characterized in that: The step of randomly masking the sample Mel spectrum according to a preset masking strategy to obtain a Mel spectrum to be predicted containing a masked portion includes: Performing zero-mean normalization processing on the sample Mel spectrum to obtain a sample normalized spectrum; A number of continuous or non-continuous regions are randomly selected in the sample normalized spectrum, and mask processing is performed on the selected regions to obtain a Mel spectrum to be predicted including the masked portion.

7. The method for generating speech based on audio prompts according to any one of claims 1 to 3, characterized in that: After obtaining the target text and reference audio of the speech to be generated, the method further includes: Determine a target sampling rate, and convert the sampling rate of the reference audio to the target sampling rate through a resampling process; The reference audio after the sampling rate is adjusted is subjected to pre-emphasis processing and normalization processing to obtain a pre-processed audio.

8. A speech generation device based on audio prompts, characterized in that: include: An acquisition module is used to acquire the target text and reference audio of the speech to be generated; A text processing module, used for performing multi-level feature extraction on the target text through a pre-trained text feature extractor to obtain multi-level text features; A feature splicing module, used to generate corresponding audio prompt features according to the reference audio, and splice the multi-level text features with the audio prompt features to obtain spliced ​​input features; The speech generation module is used to input the concatenated input features into a pre-trained speech generation model to generate a target speech corresponding to the target text, wherein the speech generation model is obtained by training a preset stream model for speech mask generation.

9. A computer device, characterized in that: comprising at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the speech generation method based on audio prompts as described in any one of claims 1-7.

10. A non-volatile computer-readable storage medium, characterized in that: The non-volatile computer-readable storage medium stores computer-executable instructions, which, when executed by one or more processors, enable the one or more processors to execute the audio prompt-based speech generation method described in any one of claims 1-7.

Citation Information

Cited By

  • Data desensitization method and device based on TTS speech synthesis, equipment and medium

    CN120748361A

  • Voice editing method and device and electronic equipment

    CN121506157A