Audio generation method and device
By converting reference audio into reference sound feature vectors and pending text into pending text feature vectors, combined with audio generation model for inference and decoding, the problem of unnatural audio generation in the prior art is solved, and a more natural audio generation effect is achieved.
Patent Information
- Application Number
- CN202510792476.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-08-15
AI Technical Summary
In the prior art, the tone, rhythm and other information of the audio generated by the audio generation method are completely obtained based on reference audio and will not be adjusted according to the text content, resulting in the generated audio being unnatural enough.
By obtaining the audio generation model, the reference audio of the target object is converted into a reference sound feature vector, and the pending text is converted into a pending text feature vector, and the model reasoning is performed based on the pending text feature vector and the reference sound feature vector, the audio word elements matching the pending text are obtained, and they are decoded to generate the target audio that simulates the tone of the target object.
It improves the naturalness of synthetic audio, makes the generated audio closer to the tone of the target object, and enhances the naturalness of the audio.
Smart Images

Figure CN120496495A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing and natural language modeling, and in particular to an audio generation method and device. Background Art
[0002] Sound synthesis is the process of converting digital signals into analog sound signals using electronic devices or computer programs to produce artificial sounds. With the development of artificial intelligence, sound synthesis, especially sound synthesis that simulates the timbre of a specific object, is finding increasingly widespread application.
[0003] In existing technologies, reference audio of the specific object to be imitated is typically concatenated with the text content and fed into an audio generation model. The model then performs a backward prediction based on the reference audio and generates audio corresponding to the text content, using the timbre of the specific object. However, this existing audio generation method derives information such as the timbre and rhythm of the generated audio entirely from the reference audio and does not adjust to the text content, resulting in the generated audio sounding less natural. Summary of the Invention
[0004] The present invention provides an audio generation method and device to improve the naturalness of synthesized audio.
[0005] In a first aspect, an embodiment of the present invention provides an audio generation method, the method comprising:
[0006] Acquire an audio generation model, convert reference audio of a target object into a reference sound feature vector, and convert a text to be processed into a text feature vector to be processed, wherein the reference audio includes reference sound information required for synthesizing the target audio;
[0007] Performing model inference based on at least the feature vector of the text to be processed and the reference sound feature vector through the audio generation model to obtain an audio word that matches the text to be processed;
[0008] The audio generation model is used to decode the audio word that matches the text to be processed to obtain the target audio that simulates the timbre of the target object.
[0009] In a second aspect, an embodiment of the present invention further provides an audio generation device, the device comprising:
[0010] a vector encoding module, configured to obtain an audio generation model, convert reference audio of a target object into a reference sound feature vector, and convert a text to be processed into a text feature vector to be processed, wherein the reference audio includes reference sound information required for synthesizing the target audio;
[0011] A model inference module is configured to generate a model using the audio generation model, based at least on a feature vector of the text to be processed and a reference sound feature vector, to perform model inference to obtain an audio word that matches the text to be processed;
[0012] The target audio generation module is used to decode the audio word that matches the text to be processed through the audio generation model to obtain the target audio that simulates the timbre of the target object.
[0013] The technical solution of the embodiment of the present invention obtains a pre-trained audio generation model, converts the target object reference audio into a reference sound feature vector, and converts the to-be-processed text into a to-be-processed text feature vector. Based on at least the to-be-processed text feature vector and the reference sound feature vector, model inference is performed to obtain audio tokens that match the to-be-processed text, and the audio tokens are decoded to obtain target audio that simulates the timbre of the target object. The technical solution of the embodiment of the present invention obtains a reference sound feature vector from the reference audio, separates the sound of the reference audio from the text, better extracts sound information, and makes the generated audio more natural. The audio generation model proposed in this application can support both zero-shot learning mode and single-shot learning mode.
[0014] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0016] Figure 1 is a flowchart of an audio generation method provided by Example 1 of the present invention;
[0017] Figure 2 This is a flowchart of an audio generation method provided by Embodiment 2 of the present invention;
[0018] Figure 3 This is a schematic structural diagram of an audio generating device provided in Embodiment 3 of the present invention;
[0019] Figure 4 This is a structural diagram of an electronic device provided in Example 4 of the present invention. DETAILED DESCRIPTION
[0020] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0021] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. In the embodiments of the present application, certain software, components, models and other existing solutions in the industry may be mentioned, and they should be considered as exemplary. Their purpose is merely to illustrate the feasibility of the implementation of the technical solution of the present application, but it does not mean that the applicant has or will necessarily use the solution.
[0022] The acquisition, transmission, storage, use, and processing of data in the technical solution of this application comply with the relevant provisions of national laws and regulations.
[0023] Example 1
[0024] Figure 1 A flowchart of an audio generation method is provided for embodiment 1 of the present invention. This embodiment is applicable to the case of synthesizing target audio that simulates the timbre of a target object. The method can be executed by an audio generation device, which can be implemented in the form of hardware and / or software, and the audio generation device can be configured in a server.
[0025] like Figure 1 As shown, the method includes:
[0026] S110 , obtaining an audio generation model, converting the reference audio of the target object into a reference sound feature vector, and converting the text to be processed into a text feature vector to be processed.
[0027] Among them, the audio generation model is obtained by pre-training based on the initial audio generation model using the training audio as the training sample. The structure of the initial model may include a Transformer model, a GPT model, etc.
[0028] Furthermore, the audio generation model includes an encoder, an autoregressive model, and a decoder, wherein the encoder is a learnable speaker encoder. In this embodiment, S110 is performed by the encoder.
[0029] The Learnable Speaker Encoder (LEE) is a neural network model that extracts representative features from input audio. The LEE's parameters are trained jointly with the model. Through continuous learning during model training, the reference sound feature vectors generated from reference audio of the same target object are kept as close as possible in feature space, while reference sound feature vectors of different target objects are kept farther apart. This enables the audio generation model to accurately distinguish between different target objects, resulting in generated audio that more closely matches the timbre of each target object, improving the naturalness of the audio.
[0030] The target object can be the user who initiates the dialogue interaction, or it can be an audiovisual character, a virtual character, etc., and the present invention does not limit this. The reference audio is the voice input into the audio generation model through the voice input module when the user initiates the dialogue, or in response to the user's selection operation of the reference audio of the audiovisual character, or in response to the user's selection operation of the timbre reference audio of the virtual character, it is input into the audio generation model. Among them, the reference audio includes the reference sound information required for synthesizing the target audio. The reference sound information may include language information, timbre information, rhythm information, volume information, and speaking speed information. Correspondingly, the reference sound feature vector is used to represent the reference sound features extracted by the encoder of the audio generation model in the reference audio. The reference sound features may include language, timbre, rhythm, volume, and speaking speed.
[0031] The text to be processed is the text content corresponding to the target audio that the audio generation model needs to generate. The text to be processed can be obtained based on the user's input, or it can be generated based on the content of the conversation initiated by the user through a conversation generation model, etc.
[0032] In this embodiment, the audio generation model converts the reference audio of the target object into a reference sound feature vector. Specifically, the reference audio of the target object can be converted into a digital signal, then the digital signal can be framed, windowed, and pre-emphasized, and finally feature extracted from the digital signal to obtain the reference sound feature vector. Alternatively, the reference sound feature vector can be extracted from the reference audio using a pre-trained model. By converting the reference audio into a reference sound feature vector, the key attributes of the reference audio can be captured, thereby providing data support for subsequent audio generation.
[0033] The audio generation model converts the to-be-processed text into a to-be-processed text feature vector. Specifically, this can be achieved using a method based on a bag of words (BoW) model, a method based on distributed semantics, or a pre-trained natural language model, which is not limited in this embodiment. By converting the to-be-processed text into a to-be-processed text feature vector, the conversion of text content into a computable numerical vector is achieved, thereby providing data support for subsequent audio generation.
[0034] Furthermore, for the target object that requires audio synthesis, when the reference audio is input for the first time, the reference audio is converted into a reference sound feature vector through the encoder of the audio generation model. Before the conversation matching the target object ends, the audio generation model can directly use the converted reference sound feature vector and the new text to be processed to obtain the new text to be processed feature vector for model inference.
[0035] Furthermore, the training step of the audio generation model includes:
[0036] S1. Obtain the initial audio generation model and training audio;
[0037] S2. Convert the training audio into training sound feature vectors through the initial audio generation model, and obtain first sound word units that match each training sound feature vector based on each training sound feature vector;
[0038] S3. Extract the second sound words from the training audio, and train the initial audio generation model according to each first sound word and each second sound word to obtain an audio generation model.
[0039] The structure of the initial audio generation model may also include an encoder, an autoregressive model, and a decoder. The initial parameters of the encoder, autoregressive model, and decoder in the initial audio generation model may adopt initial default parameters or input custom parameters. Model training and optimization are achieved by adjusting the parameters of the encoder and autoregressive model. The training audio is audio corresponding to at least two different target objects, and the training sound feature vector is used to represent the sound features contained in the training audio.
[0040] It should be noted that the dimensions of the training sound feature vectors after conversion of each training audio are the same, so as to ensure that the first sound word unit and the second sound word unit obtained based on each training audio can realize the calculation of loss, and ensure that the dimensions of the training sound feature vectors after conversion of each training audio are the same, which can avoid the destruction of the sound timing feature alignment logic due to inconsistent feature dimensions and ensure the good training effect of the model.
[0041] A sound token is a discretized semantic unit. Through the autoregressive model in the audio generation model, sound prediction can be performed based on the training sound feature vector to obtain a sound token, and the sound token predicted by the autoregressive model is used as the first sound token. At the same time, for the training audio, token extraction is directly performed on it. Specifically, acoustic feature extraction, semantic recognition and discretization are performed on the training audio to finally obtain a sound token, and the sound token extracted from the training audio is used as the second sound token. The "first" and "second" in the "first sound token" and "second sound token" in this embodiment are only used to distinguish the source of the sound token, and are not used to indicate the order, etc.
[0042] In this embodiment, the initial audio generation model converts the training audio into a training sound feature vector in the same manner as the reference audio is converted into a reference sound feature vector in this application. Prediction is performed based on the training sound feature vector to obtain first sound tokens that match each training sound feature vector. Simultaneously, second sound tokens are extracted from the training audio. For each piece of training audio, a loss is calculated based on its corresponding first sound token and second sound token. The initial audio generation model is trained based on the loss calculation results, ultimately obtaining an audio generation model.
[0043] The model training method of this embodiment enables the encoder, especially the learnable speaker encoder, to be trained together with the autoregressive model in the audio generation model. This joint training method enables the encoder and the autoregressive model to process more diverse training audio, such as training audio of different languages, different target objects, and different timbres. At the same time, this joint training method enables the audio generation model to learn a wider range of feature representation capabilities and generalization capabilities from diverse training audio. The audio generation model finally trained can support diversified target audio generation, and the generated target audio can be more natural and the audio generation effect is better.
[0044] S120. Perform model inference based on at least the feature vector of the text to be processed and the reference sound feature vector through the audio generation model to obtain an audio word that matches the text to be processed.
[0045] In this embodiment, at least the feature vector of the text to be processed and the feature vector of the reference sound are required so that the audio generation model can combine the desired timbre with the text content to be output during inference, thereby generating synthesized audio that simulates the timbre of the target object. Step S120 of this embodiment can be performed by an autoregressive model within the audio generation model. Furthermore, the autoregressive model is preferably an AR transformer (Autoregressive Transformer).
[0046] The input of the audio generation model of this embodiment includes at least a feature vector of the text to be processed and a reference sound feature vector. Therefore, the audio generation model of this embodiment supports two model reasoning modes: zero-shot mode and one-shot mode. Among them, the zero-sample mode means that during the model training process, each training audio does not match the reference sound information of the target object, that is, the target object corresponding to each training audio does not include the target object in the model reasoning stage, and the model has not learned the relevant information of the target object. The zero-sample mode requires a higher model generalization ability. The one-sample mode means that during the model training process, there is at least one training audio that matches the reference sound information of the target object, that is, the target object corresponding to each training audio includes the target object in the model reasoning stage, and the model has learned the relevant information of the target object.
[0047] It should be noted that the basis for determining whether to use zero-sample mode or single-sample mode is whether the training audio matches the reference sound information. This embodiment determines the matching relationship between the reference sound information and the training audio, rather than the matching relationship between the reference audio and the training audio. This is because the text content, volume, and other aspects of the reference and training audio may differ, resulting in inaccurate judgment results. Therefore, this embodiment can determine whether to use zero-sample mode or single-sample mode by determining the similarity between the reference sound information and the training sound information in the training audio. This setting can improve the accuracy of the model inference mode judgment, thereby improving the accuracy of the target audio generation.
[0048] Furthermore, S120 may include: if it is determined that none of the training audios matches the reference sound information, performing model inference based on the feature vector of the text to be processed and the reference sound feature vector through the audio generation model.
[0049] In zero-sample mode, the audio generation model converts the target object's reference audio into a reference sound feature vector to obtain the desired timbre of the target audio. Simultaneously, the processed text feature vector is used to obtain the desired text content of the target audio. Combining the processed text feature vector with the reference sound feature vector yields the predicted target audio.
[0050] Specifically, the encoder of the audio generation model converts the reference audio of the target object into a reference sound feature vector. Simultaneously, the encoder converts the text to be processed into a text feature vector. The text feature vector and the reference sound feature vector are used as inputs to the autoregressive model, which performs sound prediction to generate an audio token that mimics the timbre of the target object and contains the text content to be processed, ultimately producing the target audio.
[0051] Furthermore, S120 may include: determining that at least one of the training audios matches the reference sound information, then determining a training sound feature vector corresponding to the training audio that matches the target object, and determining a reference text feature vector, wherein the reference text feature vector is converted from text data extracted from the reference audio; performing model inference based on the training sound feature vector, the text feature vector to be processed, the reference sound feature vector, and the reference text feature vector through the audio generation model.
[0052] In single-sample mode, since each training audio track already includes training audio that matches the target subject's reference sound information during model training, the training sound feature vector corresponding to this matching training audio track can reflect the target subject's timbre. Simultaneously, the reference sound feature vector, obtained by converting the target subject's reference video, can also reflect the target subject's timbre. Therefore, combining the training sound feature vectors with the reference sound feature vectors can enhance timbre simulation and make the generated target audio sound more natural.
[0053] Furthermore, the content and format of the input data for the autoregressive model must be unified in both the zero-shot and single-shot modes. The input for the zero-shot mode includes the to-be-processed text feature vector and the reference sound feature vector. Therefore, in the single-shot mode, in addition to the training sound feature vector, the to-be-processed text feature vector, and the reference sound feature vector, a text feature vector is also required as input for the autoregressive model. This can be accomplished by extracting text data from the reference audio, converting the extracted text data, and then using the reference text feature vector as the input for the autoregressive model.
[0054] Specifically, the encoder of the audio generation model converts the reference audio of the target object into a reference sound feature vector. Simultaneously, the encoder converts the to-be-processed text into a to-be-processed text feature vector, and also converts the text data extracted from the reference audio into a reference text feature vector. The training sound feature vector, the to-be-processed text feature vector, the reference sound feature vector, and the reference text feature vector are used as inputs to the autoregressive model. The autoregressive model performs sound prediction to obtain an audio token that simulates the timbre of the target object and contains the content of the to-be-processed text, ultimately yielding the target audio.
[0055] Furthermore, S120 may include: at least splicing the feature vector of the text to be processed and the reference sound feature vector, and using the spliced vector group as the input of the audio generation model to perform model inference; or, at least establishing a unique mapping between the feature vector of the text to be processed and the reference sound feature vector, and using the mapped vector as the input of the audio generation model to perform model inference.
[0056] To facilitate model training, whether it is a zero-sample mode or a single-sample mode, on the premise that the input data content and format of the autoregressive model are unified, it is also necessary to ensure that the form and order of the input data are unified.
[0057] Specifically, for the zero-sample mode, the feature vector of the text to be processed and the reference sound feature vector can be spliced, either with the feature vector to be processed first and the reference sound feature vector second, or with the reference sound feature vector first and the text feature vector to be processed second, and the vector group after splicing is used as the input of the autoregressive model. Accordingly, when the zero-sample mode uses vector splicing to process the input data of the autoregressive model, the vector splicing method must also be used for the single-sample mode. At the same time, the splicing order must be consistent with the splicing order in the zero-sample mode. For example, when the zero-sample mode uses vector splicing with the feature vector to be processed first and the reference sound feature vector second, the vector splicing order of the single-sample mode can be the text feature vector to be processed, the reference text feature vector, the reference sound feature vector, and the training sound feature vector.
[0058] For the zero-shot mode, the to-be-processed text feature vector and the reference sound feature vector can be mapped and then jointly input into the autoregressive model. Accordingly, when the zero-shot mode uses a mapping approach to process the input data of the autoregressive model, the same mapping approach must be used for the single-shot mode. Specifically, a mapping can be established between the to-be-processed text feature vector and the reference sound feature vector, and a mapping can be established between the reference text feature vector and the training sound feature vector, before both are input into the autoregressive model.
[0059] The technical solution of this embodiment does not directly use the reference audio as the input of the audio generation model, especially the autoregressive model, but separates the reference sound feature vector in the reference audio from the text content, so that the audio generation model, especially the autoregressive model, can only obtain the reference sound feature vector in the reference audio, so that the autoregressive model can better extract the reference sound information, and the generated target audio is more natural and has a better simulation effect on the timbre of the target object.
[0060] S130 , decoding the audio word that matches the text to be processed through the audio generation model to obtain target audio that simulates the timbre of the target object.
[0061] Specifically, step S120 of this embodiment may be executed by a decoder in the audio generation model, and the specific working principle of the decoder is not described in detail in this embodiment.
[0062] This embodiment supports both zero-shot and single-shot model inference. During the model training phase, a learnable speaker encoder and an autoregressive model are jointly trained. Especially for single-shot model inference, the generated audio is more natural and better simulates the timbre of the target subject. Furthermore, because sound feature vectors can contain multivariate information, the model inference method of this embodiment can also achieve cross-lingual audio generation, improving the generalization capability of the audio generation model.
[0063] The technical solution of the embodiment of the present invention pre-trains an audio generation model, converts the target object reference audio into a reference sound feature vector, and converts the to-be-processed text into a to-be-processed text feature vector. Based at least on the to-be-processed text feature vector and the reference sound feature vector, model inference is performed to obtain audio tokens that match the to-be-processed text. The audio tokens are then decoded to obtain target audio that simulates the timbre of the target object. The technical solution of the embodiment of the present invention obtains a reference sound feature vector from the reference audio, separates the sound of the reference audio from the text, and better extracts sound information, making the generated audio more natural.
[0064] Example 2
[0065] Figure 2 This is a flowchart of an audio generation method provided in the second embodiment of the present invention. Based on the above embodiments, this embodiment of the present invention further specifically describes the training process of the audio generation model.
[0066] like Figure 2 As shown, the method includes:
[0067] S210: Obtain an initial audio generation model and training audio.
[0068] S220: Convert the training audio into training sound feature vectors through an initial audio generation model, and obtain first sound word units that match each training sound feature vector based on each training sound feature vector.
[0069] This embodiment does not elaborate on the process of obtaining the first sound token by using an autoregressive model to predict sound based on the training sound feature vector. This embodiment also specifically describes the process of segmenting the long training audio to ensure processing speed and computing resource balance during model training when converting the training audio to the training audio sound feature vector.
[0070] Furthermore, S220 may include:
[0071] S221: If the length of the target training audio is greater than the audio length threshold, dividing the target training audio according to the audio length threshold to obtain at least two training audio segments;
[0072] S222, converting each training audio segment into a segment sound feature vector;
[0073] S223: Determine a fusion vector of the sound feature vectors of each segment, and use the fusion vector as a training sound feature vector that matches the target training audio.
[0074] It is understandable that if the length of the training audio is too long when converting it into a training sound feature vector, it may cause a video memory overflow or insufficient computing resources. Therefore, the audio length threshold can be set according to the computing resources allocated during model training, especially according to the amount of video memory resources, as the upper limit of the length of the training audio or training audio segment processed when performing a sound feature vector conversion. For target training audio whose length is greater than the audio length threshold, it is necessary to divide it to obtain at least two training audio segments, and directly convert the target training audio into a training sound feature vector, instead of converting multiple training audio segments into segment sound feature vectors separately, thereby ensuring the processing speed and computing resource balance of the model training process. Avoid video memory overflow or insufficient computing resources.
[0075] In an optional embodiment, the target training audio whose length is greater than the audio length threshold may be divided equally according to the audio length threshold, wherein the length of the training audio segment after the equal division is less than or equal to the audio length threshold.
[0076] For example, if the audio length threshold is 6 seconds and the length of the target training audio is 15 seconds, the target training audio may be divided into three training audio segments of 5 seconds each.
[0077] In another optional embodiment, target training audio with a length greater than the audio length threshold may be sequentially intercepted according to the audio length threshold to obtain multiple training audio segments, wherein the length of the last training audio segment is less than or equal to the audio length threshold.
[0078] For example, if the audio length threshold is 5 seconds and the length of the target training audio is 13 seconds, the lengths of the training audio segments are 5 seconds, 5 seconds, and 3 seconds, respectively.
[0079] In another optional embodiment, dividing the target training audio to obtain at least two training audio segments may further include: determining a sliding window length and a sliding step size based on the length of the target training audio and the audio length threshold, wherein the sliding window length is less than or equal to the audio length threshold; and performing sliding sampling on the target training audio based on the sliding window length and the sliding step size to obtain each training audio segment, wherein the length of the training audio segment is the same as the sliding window length.
[0080] In this embodiment, for target training audio that cannot be evenly divided, in order to ensure accuracy during vector fusion and to make the lengths of the training audio segments as similar as possible, a sliding window sampling method can be used to divide the training audio segments.
[0081] When performing sliding window sampling, the principles include: the sliding window length is less than or equal to the audio length threshold; after sliding sampling based on the sliding window length and sliding step size, each training audio segment can fully cover the target training audio; to save computing resources, the number of sliding samplings is minimized. This can be converted into the solution process of the objective function: Where n is the number of slides, L is the target training audio length, l is the length of the training audio segment, and w is the slide step size. Solve for l and w when n is minimized, where the constraint is l ≤ s, where s is the audio length threshold.
[0082] For example, if the audio length threshold is 5s and the length of the target training audio is 13s, sliding sampling can be performed with a sliding window length of 4s and a sliding step size of 3s. Finally, the sliding window slides three times to obtain four training audio segments: 0-4s, 3-7s, 6-10s and 9-13s.
[0083] This sliding sampling method can ensure that the length of the training audio segments is the same, so that for target training audio with longer lengths, it is easier to perform vector conversion and vector fusion on each training audio segment, thereby improving the processing speed during model training and ensuring balanced computing power.
[0084] In this embodiment, after the target training audio is divided into multiple training audio segments and converted into segment sound feature vectors, each segment sound feature vector is fused, and the fused vector is used as the training sound feature vector that matches the target training audio. Specifically, vector fusion can be performed using weighted summation, average pooling, linear transformation, tensor fusion, and multi-layer perceptron fusion, etc., which are not limited in this embodiment.
[0085] S230. Extract the second sound words from the training audio, and train the initial audio generation model according to each first sound word and each second sound word to obtain an audio generation model.
[0086] The process of extracting the second sound token is not described in detail in this embodiment. Meanwhile, a specific process of model training based on the first sound word and the second sound word is provided.
[0087] Furthermore, S230 may include:
[0088] S231, calculating the loss between each first sound word and each second sound word according to the loss function, and adjusting the parameters of the initial audio generation model according to the loss function calculation result;
[0089] S232. Repeat the operation of converting the training audio into a training sound feature vector until the model meets the training conditions, thereby obtaining an audio generation model.
[0090] The loss function may be a cross-entropy loss function. For each training audio, the loss may be calculated based on its corresponding first sound word and second sound word using the loss function. Based on the losses corresponding to each training audio, an average loss may be calculated as the loss for this training, and based on the average loss, the parameters of the initial audio generation model may be adjusted. Specifically, when the structure of the audio generation model includes an encoder, an autoregressive model, and a decoder, the parameters of the encoder and the autoregressive model may be adjusted.
[0091] Repeat the above process and perform iterative model training until the training conditions are met. Specifically, the training conditions may include reaching a preset number of iterations, the model performance (such as accuracy, F1 score) not improving after multiple consecutive iterations, the model loss being less than a preset threshold or changing very little, or the parameter update amplitude being less than a threshold. This embodiment does not limit the specific content of the conditions for stopping iterative model training.
[0092] S240: Acquire an audio generation model, convert the reference audio of the target object into a reference sound feature vector, and convert the text to be processed into a text feature vector to be processed.
[0093] S250. Perform model inference based on at least the feature vector of the text to be processed and the reference sound feature vector through the audio generation model to obtain an audio word that matches the text to be processed.
[0094] S260: Decode the audio word that matches the text to be processed through the audio generation model to obtain target audio that simulates the timbre of the target object.
[0095] The technical solution of this embodiment adopts the joint training of a learnable speaker encoder and an autoregressive model during the model training phase, which improves the feature representation and generalization capabilities of the audio generation model, so that the trained model can process the multivariate information in the sound feature vector, thereby generating more natural audio and better simulating the timbre of the target object. At the same time, since the sound feature vector can contain multivariate information, the model reasoning method of this embodiment can also achieve cross-language audio generation. In the model reasoning phase, it supports not only the zero-sample model reasoning mode but also the single-sample model reasoning mode. The model has a wider range of applicability and does not need to train the audio generation model separately for different reasoning modes, saving the model training cost.
[0096] Example 3
[0097] Figure 3 This is a structural diagram of an audio generating device provided by the third embodiment of the present invention. Figure 3 As shown, the device includes:
[0098] The vector encoding module 310 obtains an audio generation model, converts the reference audio of the target object into a reference sound feature vector, and converts the text to be processed into a text feature vector to be processed, wherein the reference audio includes the reference sound information required for synthesizing the target audio;
[0099] A model inference module 320 is configured to generate a model using the audio generation model, based at least on the feature vector of the text to be processed and the reference sound feature vector, to obtain an audio word that matches the text to be processed;
[0100] The target audio generation module 330 is used to decode the audio word that matches the text to be processed through the audio generation model to obtain the target audio that simulates the timbre of the target object.
[0101] The technical solution of the embodiment of the present invention pre-trains an audio generation model, converts the target object reference audio into a reference sound feature vector, and converts the to-be-processed text into a to-be-processed text feature vector. Based at least on the to-be-processed text feature vector and the reference sound feature vector, model inference is performed to obtain audio tokens that match the to-be-processed text. The audio tokens are then decoded to obtain target audio that simulates the timbre of the target object. The technical solution of the embodiment of the present invention obtains a reference sound feature vector from the reference audio, separates the sound of the reference audio from the text, and better extracts sound information, making the generated audio more natural.
[0102] Based on the above embodiment, optionally, the device further includes:
[0103] A training audio determination module is used to obtain an initial audio generation model and training audio;
[0104] A first sound word unit determination module is configured to convert the training audio into training sound feature vectors using an initial audio generation model, and obtain, based on each training sound feature vector, a first sound word unit that matches each training sound feature vector;
[0105] The audio generation model training module is used to extract the second sound word from the training audio, and train the initial audio generation model according to each first sound word and each second sound word to obtain the audio generation model.
[0106] Based on the above embodiment, optionally, the model reasoning module 320 includes:
[0107] The first model inference unit is used to perform model inference based on the feature vector of the text to be processed and the reference sound feature vector through the audio generation model if it is determined that the training audio does not match the reference sound information.
[0108] Based on the above embodiment, optionally, the model reasoning module 320 includes:
[0109] a training sound feature vector determining unit configured to, if it is determined that at least one piece of the training audio matches the reference sound information, determine a training sound feature vector corresponding to the training audio that matches the target object, and determine a reference text feature vector, the reference text feature vector being converted from text data extracted from the reference audio;
[0110] The second model inference unit is used to perform model inference based on the training sound feature vector, the to-be-processed text feature vector, the reference sound feature vector and the reference text feature vector through the audio generation model.
[0111] Based on the above embodiment, optionally, the first sound word unit determination module includes:
[0112] a target training audio segmentation unit, configured to segment the target training audio according to the audio length threshold to obtain at least two training audio segments if the length of the target training audio is greater than the audio length threshold;
[0113] A sound feature vector conversion unit, configured to convert each training audio segment into a segment sound feature vector;
[0114] The feature vector fusion unit is used to determine a fusion vector of the sound feature vectors of each segment, and use the fusion vector as a training sound feature vector that matches the target training audio.
[0115] Based on the above embodiment, optionally, the target training audio division unit is specifically configured to:
[0116] Determining a sliding window length and a sliding step size according to the length of the target training audio and the audio length threshold, wherein the sliding window length is less than or equal to the audio length threshold;
[0117] According to the sliding window length and the sliding step size, the target training audio is slidingly sampled to obtain each training audio segment, and the length of the training audio segment is the same as the sliding window length.
[0118] Based on the above embodiment, optionally, the audio generation model training module includes:
[0119] a model parameter adjustment unit, configured to calculate the loss between each first sound word and each second sound word according to a loss function, and adjust the parameters of the initial audio generation model according to the loss function calculation result;
[0120] The training condition judgment unit is used to repeatedly perform the operation of converting the training audio into a training sound feature vector until the model meets the training conditions, thereby obtaining an audio generation model.
[0121] Based on the above embodiment, optionally, the audio generation model includes an encoder, an autoregressive model, and a decoder, and the encoder is a learnable speaker encoder;
[0122] Model parameter adjustment unit, specifically used for:
[0123] According to the loss function calculation results, the parameters of the encoder and autoregressive model are adjusted.
[0124] Based on the above embodiment, optionally, the model reasoning module 320 includes:
[0125] A feature vector concatenation unit, configured to concatenate at least a feature vector of the text to be processed and a feature vector of the reference sound, and use the concatenated vector group as an input of the audio generation model to perform model inference;
[0126] The feature vector mapping unit is used to establish at least a unique mapping between the feature vector of the text to be processed and the reference sound feature vector, and use the mapped vector as the input of the audio generation model to perform model inference.
[0127] The audio generation device provided in the embodiment of the present invention can execute the audio generation method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0128] Example 4
[0129] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0130] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0131] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0132] The processor 11 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the audio generation method.
[0133] In some embodiments, the audio generation method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the audio generation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the audio generation method in any other suitable manner (e.g., by means of firmware).
[0134] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0135] Computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable audio generating device, such that when executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0136] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0137] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0138] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0139] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0140] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0141] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. An audio generation method, characterized in that: include: Acquire an audio generation model, convert reference audio of a target object into a reference sound feature vector, and convert a text to be processed into a text feature vector to be processed, wherein the reference audio includes reference sound information required for synthesizing the target audio; Performing model inference based on at least the feature vector of the text to be processed and the reference sound feature vector through the audio generation model to obtain an audio word that matches the text to be processed; The audio generation model is used to decode the audio word that matches the text to be processed to obtain the target audio that simulates the timbre of the target object.
2. The method according to claim 1, characterized in that The training step of the audio generation model includes: Get the initial audio generation model and training audio; The training audio is converted into a training sound feature vector by an initial audio generation model, and based on each training sound feature vector, a first sound word unit matching each training sound feature vector is obtained; Second sound words are extracted from the training audio, and an initial audio generation model is trained based on each first sound word and each second sound word to obtain an audio generation model.
3. The method according to claim 2, characterized in that The audio generation model is used to perform model reasoning based on at least the feature vector of the text to be processed and the feature vector of the reference sound, including: If it is determined that the training audio does not match the reference sound information, model inference is performed based on the feature vector of the text to be processed and the reference sound feature vector through the audio generation model.
4. The method according to claim 2, characterized in that The audio generation model is used to perform model reasoning based on at least the feature vector of the text to be processed and the feature vector of the reference sound, including: If it is determined that at least one of the training audios matches the reference audio information, determining a training audio feature vector corresponding to the training audio that matches the target object, and determining a reference text feature vector, the reference text feature vector being converted from text data extracted from the reference audio; The audio generation model is used to perform model inference based on the training sound feature vector, the to-be-processed text feature vector, the reference sound feature vector, and the reference text feature vector.
5. The method according to claim 2, characterized in that Converting the training audio into a training sound feature vector includes: If the length of the target training audio is greater than the audio length threshold, dividing the target training audio according to the audio length threshold to obtain at least two training audio segments; Convert each training audio segment into a segment sound feature vector; A fusion vector of the sound feature vectors of each segment is determined, and the fusion vector is used as a training sound feature vector that matches the target training audio.
6. The method according to claim 5, characterized in that Divide the target training audio according to the audio length threshold to obtain at least two training audio segments, including: Determining a sliding window length and a sliding step size according to the length of the target training audio and the audio length threshold, wherein the sliding window length is less than or equal to the audio length threshold; According to the sliding window length and the sliding step size, the target training audio is slidingly sampled to obtain each training audio segment, and the length of the training audio segment is the same as the sliding window length.
7. The method according to claim 2, characterized in that The initial audio generation model is trained according to each first sound word and each second sound word to obtain an audio generation model, including: Calculating the loss between each first sound word and each second sound word according to the loss function, and adjusting the parameters of the initial audio generation model according to the loss function calculation result; The operation of converting the training audio into a training sound feature vector is repeatedly performed until the model meets the training conditions, thereby obtaining an audio generation model.
8. The method according to claim 7, characterized in that The audio generation model includes an encoder, an autoregressive model and a decoder, wherein the encoder is a learnable speaker encoder; According to the loss function calculation results, the parameters of the initial audio generation model are adjusted, including: According to the loss function calculation results, the parameters of the encoder and autoregressive model are adjusted.
9. The method according to claim 1, characterized in that Model inference is performed based on at least the feature vector of the text to be processed and the reference sound feature vector, including: At least concatenate the feature vector of the text to be processed and the feature vector of the reference sound, and use the concatenated vector group as input of the audio generation model to perform model inference; Alternatively, at least a unique mapping is established between the feature vector of the text to be processed and the feature vector of the reference sound, and the mapped vector is used as the input of the audio generation model to perform model inference.
10. An audio generating device, characterized in that: include: a vector encoding module, configured to obtain an audio generation model, convert reference audio of a target object into a reference sound feature vector, and convert a text to be processed into a text feature vector to be processed, wherein the reference audio includes reference sound information required for synthesizing the target audio; A model inference module is configured to generate a model using the audio generation model, based at least on a feature vector of the text to be processed and a reference sound feature vector, to perform model inference to obtain an audio word that matches the text to be processed; The target audio generation module is used to decode the audio word that matches the text to be processed through the audio generation model to obtain the target audio that simulates the timbre of the target object.
Citation Information
Patent Citations
Audio generation method and device, computer equipment and storage medium
CN118841008A
Speech synthesis method and device, electronic equipment and storage medium
CN119181349A
Zero sample speech synthesis method and device based on autoregressive large language model
CN119380696A
Text-based speech synthesis method and device, equipment and storage medium
CN119400150A
Speech synthesis method, device, equipment and computer medium
CN119541451A
Cited By
Data processing method, data processing device, electronic equipment and storage medium
CN121096313A