Audio generation method based on preference optimization
Through the two-stage training framework, the integration of audio and text features is optimized, and the problems of insufficient and deviation in generation quality in traditional audio generation are solved, achieving efficient generation of high-quality audio.
Patent Information
- Application Number
- CN202510571665.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional audio generation technology has insufficient generation quality, lack of human preference alignment, low data utilization efficiency and model architecture limitations, resulting in the generated audio being fuzzy, semantic deviation and mechanical stiffness.
Using a two-stage training framework, firstly extract audio and text features through audio VAE and pre-trained models, stitch the candidate audio, and optimize human preferences through similarity comparison and reinforcement learning, and finally generate audio in line with human preferences.
It enables efficient generation of high-quality audio, which can complete traditional methods for hours to days in minutes, reduce labor costs, generate complex scene audio and prevent catastrophic forgetting.
Smart Images

Figure CN120452413A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and audio generation technology, and in particular to an audio generation method based on preference optimization. Background Art
[0002] Traditional audio content creation relies heavily on professional recording, sound effects, and post-processing, which is time-consuming and costly. While existing text-to-audio (TTA) generation technology can automatically generate audio from text descriptions, it still has significant drawbacks:
[0003] 1. Insufficient generation quality: The audio generated by existing models often suffers from fuzzy sound quality and semantic deviation. For example, when generating "explosion sounds", the generated audio lacks low-frequency impact or spatial sense.
[0004] 2. Lack of alignment with human preferences: Traditional methods rely on objective optimization indicators such as mean square error (MSE), ignoring human subjective aesthetics (such as emotional expression and sound layering), resulting in mechanical and rigid generation results.
[0005] 3. Low data utilization efficiency: The quality of audio-text alignment in public datasets varies greatly, and there is no targeted screening of preferred data, which limits the upper limit of model performance.
[0006] 4. Model architecture limitations: Most models use single-modal encoding (text or audio only), which does not fully integrate multimodal features and makes it difficult to generate complex scene sounds. Therefore, we propose an audio generation method based on preference optimization to solve this problem. Summary of the Invention
[0007] The purpose of the present invention is to provide an audio generation method based on preference optimization to solve the problems raised in the above background technology.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] An audio generation method based on preference optimization comprises the following steps:
[0010] S1. Input audio: Use audio VAE to convert any audio into audio features;
[0011] S2. Input text description: Use pre-trained model to extract text features;
[0012] S3, feature concatenation: concatenate the audio features and text features, input them into the large model, and train them to generate the audio large model trained in the first stage;
[0013] S4, candidate audio generation: Input the text description of the music category, and generate N audios through the audio model trained in the first stage;
[0014] S5. Similarity comparison: perform a similarity comparison between the N generated audios and the original text description, and select the most similar one as the final audio;
[0015] S6. Model iteration: The audio model trained in the first phase is fed into the selected final video as the data for this iteration. This fine-tunes the audio model to generate the second phase trained audio model, allowing the model to generate results that are more in line with human preferences.
[0016] S7, Audio Generation: Input the text description of the music category into the audio model trained in the second stage to generate audio.
[0017] Preferably, in S1, the audio VAE adopts the Stable Audio Open model, whose encoder is a 12-layer convolutional neural network, the decoder is a 10-layer deconvolutional network, and the latent variable dimension is 512.
[0018] Preferably, in S1, audio of any length is resampled to 44.1kHz mono and divided into 30-second segments. If the segment is less than 30 seconds, zero is added, and if the segment exceeds 30 seconds, the middle part is cut off. The audio features extracted by the VAE encoder are continuous vectors.
[0019] Preferably, in S2, the pre-trained models include T5, Clip and FLAN-T5. Clip is used when the input text is described as a short text, and T5 is used when the input text is described as a long text. When multiple pre-trained models are used at the same time, weighted splicing is used, and the weights are determined by grid search of the validation set.
[0020] Preferably, in S5, the similarity comparison is performed manually, and the audio with the best quality that best matches the text description is manually screened out based on all generated audios.
[0021] Preferably, similarity comparison is implemented using an algorithm and CLAP pre-trained large model is used for screening to reduce labor costs and time.
[0022] Preferably, in S3, the audio features and text features are concatenated along the channel dimension to obtain a joint feature vector, the joint feature vector is mapped to the DiT input dimension through a linear projection layer, and a learnable position encoding is added.
[0023] Preferably, in S3, the Flow Matching loss function is used during training, the optimizer is AdamW, and the audio clips are uniformly 30 seconds.
[0024] Preferably, in S6, 2000 texts are randomly selected from the candidate pool in each iteration, and 2000 positive samples are generated and screened, while unselected low-scoring audios of the same text are retained as negative samples.
[0025] Preferably, in S6, during the iteration process, the parameters of the first 16 layers of DiT are frozen, and only the last 8 layers and the projection layer are fine-tuned to prevent catastrophic forgetting.
[0026] Preferably, in S7, the frame number control module of the VAE decoder supports the generation of 10 to 60 seconds of audio, and dynamic range compression and noise threshold are applied to the generated audio to improve the listening quality.
[0027] The beneficial effects of the present invention are:
[0028] In the present invention, the audio generation method based on preference optimization adopts a two-stage training framework. In the pre-training stage, the present invention uses large-scale public data to learn the basic audio generation capabilities, and in the fine-tuning stage, reinforcement learning is used to directly optimize human preference indicators.
[0029] In the present invention, the described audio generation method based on preference optimization can accurately capture cross-modal semantic associations by fusing audio VAE features with multi-encoder text features. In addition, the deep attention mechanism of the DiT architecture supports long-range dependency modeling and can generate audio with complex temporal structures.
[0030] The present invention describes a method for audio generation based on preference optimization. Traditional audio production requires professional sound engineers to complete the task of several hours to several days. This method can generate high-quality candidate results in just minutes. Through CLAP automated screening and reinforcement learning fine-tuning, labor costs are reduced.
[0031] In the present invention, the audio generation method based on preference optimization can generate a full range of audio including music, sound effects, and ambient sounds, and is particularly good at complex event combination scenes;
[0032] In the present invention, the audio generation method based on preference optimization freezes the parameters of the first 16 layers of DiT and only fine-tunes the deep network and projection layer. The present invention effectively prevents catastrophic forgetting problems during the reinforcement learning stage. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 This is a flowchart of an audio generation method based on preference optimization proposed by the present invention. DETAILED DESCRIPTION
[0034] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0035] Reference Figure 1 , an audio generation method based on preference optimization, comprising the following steps:
[0036] S1. Input audio: Use audio VAE to convert any audio into audio features;
[0037] S2. Input text description: Use pre-trained model to extract text features;
[0038] S3, feature concatenation: concatenate the audio features and text features, input them into the large model, and train them to generate the audio large model trained in the first stage;
[0039] S4, candidate audio generation: Input the text description of the music category, and generate N audios through the audio model trained in the first stage;
[0040] S5. Similarity comparison: perform a similarity comparison between the N generated audios and the original text description, and select the most similar one as the final audio;
[0041] S6. Model iteration: The audio model trained in the first phase is fed into the selected final video as the data for this iteration. This fine-tunes the audio model to generate the second phase trained audio model, allowing the model to generate results that are more in line with human preferences.
[0042] S7, Audio Generation: Input the text description of the music category into the audio model trained in the second stage to generate audio.
[0043] In this embodiment, in S1, the audio VAE adopts the Stable Audio Open model, whose encoder is a 12-layer convolutional neural network, the decoder is a 10-layer deconvolutional network, and the latent variable dimension is 512.
[0044] In this embodiment, in S1, audio of any length is resampled to 44.1kHz mono and divided into 30-second segments. If the segment is less than 30 seconds, zero is added, and if the segment exceeds 30 seconds, the middle part is cut off. The audio features extracted by the VAE encoder are continuous vectors.
[0045] In this embodiment, in S2, the pre-trained models include T5, Clip and FLAN-T5. Clip is used when the input text description is a short text, and T5 is used when the input text description is a long text. When multiple pre-trained models are used at the same time, weighted splicing is used, and the weights are determined by grid search of the validation set.
[0046] In this embodiment, in S5, the similarity comparison is manually implemented, and the audio with the best quality that best matches the text description is manually screened out based on all the generated audios.
[0047] In this embodiment, similarity comparison is implemented using an algorithm, and CLAP pre-trained large model is used for screening to reduce labor costs and time.
[0048] In S3, the audio features and text features are concatenated along the channel dimension to obtain a joint feature vector, which is then mapped to the DiT input dimension through a linear projection layer, and a learnable positional encoding is added.
[0049] In this embodiment, in S3, the Flow Matching loss function is used during training, the optimizer is AdamW, and the audio clips are uniformly 30 seconds.
[0050] In this embodiment, in S6, 2000 texts are randomly selected from the candidate pool in each iteration, and 2000 positive samples are generated and screened, while unselected low-scoring audios of the same text are retained as negative samples.
[0051] In this embodiment, in S6, during the iteration process, the parameters of the first 16 layers of DiT are frozen, and only the last 8 layers and the projection layer are fine-tuned to prevent catastrophic forgetting.
[0052] In this embodiment, in S7, the frame number control module of the VAE decoder supports the generation of 10 to 60 seconds of audio, and dynamic range compression and noise threshold are applied to the generated audio to improve the listening quality.
[0053] In this embodiment, through a two-stage training framework, the present invention uses large-scale public data to learn basic audio generation capabilities in the pre-training stage, and directly optimizes human preference indicators through reinforcement learning in the fine-tuning stage.
[0054] The above is a detailed introduction to the audio generation method based on preference optimization provided by the present invention. Specific embodiments are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. An audio generation method based on preference optimization, characterized in that: The steps include: S1. Input audio: Use audio VAE to convert any audio into audio features; S2. Input text description: Use pre-trained model to extract text features; S3, feature concatenation: concatenate the audio features and text features, input them into the large model, and train them to generate the audio large model trained in the first stage; S4, candidate audio generation: Input the text description of the music category, and generate N audios through the audio model trained in the first stage; S5. Similarity comparison: perform a similarity comparison between the N generated audios and the original text description, and select the most similar one as the final audio; S6. Model iteration: The audio model trained in the first phase is fed into the selected final video as the data for this iteration. This fine-tunes the audio model to generate the second phase trained audio model, allowing the model to generate results that are more in line with human preferences. S7, Audio Generation: Input the text description of the music category into the audio model trained in the second stage to generate audio.
2. The audio generation method based on preference optimization according to claim 1, characterized in that: In S1, the audio VAE adopts the Stable Audio Open model, whose encoder is a 12-layer convolutional neural network, the decoder is a 10-layer deconvolutional network, and the latent variable dimension is 512.
3. The audio generation method based on preference optimization according to claim 1, characterized in that: In S1, audio of any length is resampled to 44.1kHz mono and divided into 30-second segments. If the segment is less than 30 seconds, zero is added, and if the segment exceeds 30 seconds, the middle part is cut off. The audio features extracted by the VAE encoder are continuous vectors.
4. The audio generation method based on preference optimization according to claim 1, characterized in that: In S2, the pre-trained models include T5, Clip and FLAN-T5. Clip is used when the input text is described as a short text, and T5 is used when the input text is described as a long text. When multiple pre-trained models are used at the same time, weighted splicing is used, and the weights are determined by grid search on the validation set.
5. The audio generation method based on preference optimization according to claim 1, characterized in that: In S5, the similarity comparison is manually implemented, and the audio with the best quality that best matches the text description is manually screened out based on all the generated audios.
6. The audio generation method based on preference optimization according to claim 1, characterized in that: Similarity comparison is implemented using algorithms and CLAP pre-trained large models are used for screening to reduce labor costs and time; In S3, the audio features and text features are concatenated along the channel dimension to obtain a joint feature vector, which is then mapped to the DiT input dimension through a linear projection layer, and a learnable positional encoding is added.
7. The audio generation method based on preference optimization according to claim 1, characterized in that: In S3, the Flow Matching loss function is used during training, the optimizer is AdamW, and the audio clips are uniformly 30 seconds.
8. The audio generation method based on preference optimization according to claim 1, characterized in that: In S6, 2000 texts are randomly selected from the candidate pool in each iteration, and 2000 positive samples are generated and screened, while unselected low-scoring audios of the same text are retained as negative samples.
9. The audio generation method based on preference optimization according to claim 1, characterized in that: In S6, during the iteration process, the parameters of the first 16 layers of DiT are frozen, and only the last 8 layers and the projection layer are fine-tuned to prevent catastrophic forgetting.
10. The audio generation method based on preference optimization according to claim 1, characterized in that: In the S7, the frame number control module of the VAE decoder supports the generation of 10 to 60 seconds of audio, and dynamic range compression and noise threshold are applied to the generated audio to improve the listening quality.