Method for generating audio from high-quality video
By using semantic pre-training models and video understanding pre-training models to process videos in video generation audio technology, and using Seq2Seq model and vocoder to generate audio, the problem of poor time-related information guidance in the prior art is solved, and a higher quality time alignment effect is achieved.
Patent Information
- Application Number
- CN202510098110.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-13
AI Technical Summary
In the existing video generation audio technology, the guidance effect of time-related information is poor, resulting in low time alignment and quality of the generated audio and video.
The pre-trained model based on semantic pre-trained model and video understanding are used to process videos to obtain semantic information and video understanding features. These features are input into the Seq2Seq model for processing, output the vocal prediction of audio frames, and enhance the guidance ability of time-related information through discretization and embedding vectors, and finally generate audio through the vocoder.
Improve the prediction accuracy and guidance capability of time-related information in video generated audio, thereby improving the time alignment effect and quality of generated audio and video.
Smart Images

Figure CN119988671A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video sound effect generation, and in particular to a method for generating audio from high-quality video. Background Art
[0002] Common video-generated audio technologies are usually divided into several modules: 1) Video prediction module for information related to the time of voice onset; 2) Module for extracting semantic information from the video; 3) Guide the audio representation generation module based on semantic information and information related to the time of voice onset; 4) Restore the audio representation to the audio module through the vocoder. Common methods for information related to the time of voice onset are: ① The voice label of the video frame: represented by 0, 1, 0 for no voice, 1 for voice, this method cannot represent the intensity, and cannot guide the voice intensity in the generated audio well. ② The energy or root mean square (RMS) of the video frame, represented by a decimal of 0 to 1, the size of this value can represent the intensity, and to a certain extent has a guiding effect on the intensity of the generated audio. In the module for guiding the audio representation generation with relevant time information, the relevant time information will be copied from a 1-dimensional value to the same 256-dimensional value, which is used to guide the input of the generation module. From this point of view, the amount of input information for time-related guidance is actually very small, so it is impossible to achieve a good guiding effect.
[0003] Therefore, it is necessary to provide a high-quality video-to-audio generation method to improve the time alignment effect and quality of the generated audio and video. Summary of the invention
[0004] The object of the present invention is to provide a method for generating audio from high-quality video, so as to improve the time alignment effect and quality of the generated audio and video.
[0005] In order to solve the problems existing in the prior art, the present invention provides a method for generating audio from high-quality video, comprising the following steps:
[0006] S1: Process the video based on the semantic pre-training model to obtain semantic information; process the video based on the video understanding pre-training model to obtain video understanding features;
[0007] S2: Obtain video frames of fixed length based on video understanding features;
[0008] S3: A fixed-length video frame is input into the Seq2Seq model, and the Seq2Seq model outputs the utterance prediction of the audio frame, where the utterance prediction of the audio frame is the RMS value;
[0009] S4: Discretize the RMS value into 64 discrete values as follows:
[0010] d(r) = math.floor(64*(ln(1+63|r|) / ln(64))), d(r) is 64 discrete values, r is the RMS value;
[0011] The discretized RMS value corresponds to a 256-dimensional embedding vector;
[0012] S5: Based on semantic information and 256-dimensional embedding vector training, it guides the audio representation generation module;
[0013] S6: Based on the audio representation generation module, the vocoder is used to restore and generate audio.
[0014] Optionally, in the high-quality video-to-audio generation method, the Seq2Seq model is processed as follows:
[0015] S31: Input fixed-length video frames into the Seq2Seq model;
[0016] S32: Process the length of the video frame sequence and construct a mask mark to distinguish the real video frame from the filling part;
[0017] S33: During training and inference, mask marking is used to ensure that the Seq2Seq model only focuses on the real video frames and ignores the padding parts;
[0018] S34: The Seq2Seq model outputs a fixed length of predicted audio frame. By multiplying the fixed length of the predicted audio frame with the mask marker, the padding part is filtered out to generate an audio frame utterance prediction that is consistent with the ratio of the actual video frame number.
[0019] Optionally, in the method for generating audio from high-quality video, the video frame sequence length is processed as follows:
[0020] For input video frame sequences that are insufficient in length, padding is used to fill them up to a fixed length. For input sequences that are too long, they are truncated to a fixed length.
[0021] Optionally, in the method for generating audio from high-quality video, the mask mark is constructed as follows: the mask value of the real video frame is 1, and the mask value of the filled part is 0.
[0022] Optionally, in the high-quality video audio generation method, a 256-dimensional embedding vector is updated during the training process.
[0023] Compared with the prior art, the present invention has the following advantages:
[0024] The present invention can improve the prediction accuracy and guidance capability of time-related information in video-generated audio, thereby improving the time alignment effect and quality of generated audio and video. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 A flow chart of a method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The specific implementation of the present invention will be described in more detail below in conjunction with the schematic diagram. The advantages and features of the present invention will become clearer based on the following description. It should be noted that the drawings are all in a very simplified form and are not in exact proportions, and are only used to facilitate and clearly assist in explaining the purpose of the embodiments of the present invention.
[0027] Hereinafter, if the method described herein includes a series of steps, the order in which the steps are presented herein is not necessarily the only order in which the steps may be performed, and some of the steps described may be omitted and / or some other steps not described herein may be added to the method.
[0028] Common video-generated audio technologies are usually divided into several modules: 1) Video prediction module for information related to the time of voice onset; 2) Module for extracting semantic information from the video; 3) Guide the audio representation generation module based on semantic information and information related to the time of voice onset; 4) Restore the audio representation to the audio module through the vocoder. Common methods for information related to the time of voice onset are: ① The voice label of the video frame: represented by 0, 1, 0 for no voice, 1 for voice, this method cannot represent the intensity, and cannot guide the voice intensity in the generated audio well. ② The energy or root mean square (RMS) of the video frame, represented by a decimal of 0 to 1, the size of this value can represent the intensity, and to a certain extent has a guiding effect on the intensity of the generated audio. In the module for guiding the audio representation generation with relevant time information, the relevant time information will be copied from a 1-dimensional value to the same 256-dimensional value, which is used to guide the input of the generation module. From this point of view, the amount of input information for time-related guidance is actually very small, so it is impossible to achieve a good guiding effect.
[0029] In order to solve the problems existing in the prior art, the present invention provides a method for generating audio from high-quality video, such as Figure 1 As shown, the following steps are included:
[0030] S1: Process the video based on the semantic pre-training model to obtain semantic information; process the video based on the video understanding pre-training model to obtain video understanding features;
[0031] S2: Each video frame is a picture. Usually, picture features are extracted, such as RGB features, which record the color, shape and texture of objects in the picture to represent static features; Flow features, which record the movement of objects and pixels to represent dynamic features. These features can better represent the local information of the video, but cannot represent the global information of the video. It is necessary to use a video understanding model to extract the global information of the entire video. Combining the local information of the video with the global information of the video can better predict the sound of the video.
[0032] Therefore, video frames of fixed length are obtained based on the video understanding features;
[0033] S3: Since the temporal granularity of video frames is usually large, that is, the number of video frames per second is much less than the number of audio frames. For example, the number of video frames in 1 second is 30, while the number of audio frames in 1 second is 100. In order to map video frames to audio frames, the current model structure usually uses two methods to expand the features of video frames: interpolation or replication. However, these two methods often result in each predicted video frame not being perfectly aligned to the audio frame.
[0034] Therefore, the present invention inputs a video frame of a fixed length into a Seq2Seq model, and the Seq2Seq model outputs a voicing prediction of an audio frame, where the voicing prediction of the audio frame is an RMS value (root mean square value);
[0035] Furthermore, the Seq2Seq model is processed as follows:
[0036] S31: Input fixed-length video frames into the Seq2Seq model;
[0037] S32: Process the length of the video frame sequence and construct a mask mark to distinguish the real video frame from the filling part;
[0038] Specifically, the length of the video frame sequence is processed as follows: for an insufficient input video frame sequence, pad it to a fixed length through padding, and for an overlong input, truncate it to a fixed length. The mask mark is constructed as follows: the mask value of the real video frame is 1, and the mask value of the padded part is 0.
[0039] S33: During training and inference, mask marking is used to ensure that the Seq2Seq model only focuses on the real video frames and ignores the padding parts;
[0040] S34: The Seq2Seq model outputs a fixed length of predicted audio frame. By multiplying the fixed length of the predicted audio frame with the mask marker, the padding part is filtered out to generate an audio frame utterance prediction that is consistent with the ratio of the actual video frame number.
[0041] In one embodiment,
[0042] (1) Input feature extraction: The input of the Seq2Seq model is fixed to the features of 300 video frames (10s video).
[0043] (2) Sequence length processing. For input video frame sequences that are not long enough, pad them to 300. For input sequences that are too long, cut off the first 300 frames. At the same time, construct a mask tag: the mask value of the real video frame is 1, and the padding part is 0.
[0044] (3) During training and inference, masks are used to ensure that the model only focuses on the real frames and ignores the filled parts.
[0045] (4) Model output processing. The output audio frame level occurrence of the model (i.e., the predicted discretized RMS value) is fixed to 1000 frames. By multiplying the predicted audio frame with the mask and filtering out the padding part, an audio frame sound prediction that is consistent with the ratio of the actual video frame number is generated.
[0046] The Seq2Seq model solves the time alignment problem caused by traditional interpolation or replication methods by inputting video frame features and predicting the utterance of audio frames. By padding or truncating video frames of different lengths and combining mask operations, it ensures that the output of the model can accurately generate audio utterances in proportion to the number of video frames.
[0047] S4: Discretize the RMS value into 64 discrete values as follows:
[0048] d(r) = math.floor(64*(ln(1+63|r|) / ln(64))), d(r) is 64 discrete values, r is the RMS value, math.floor is a function used to round down floating point numbers;
[0049] The discretized RMS value corresponds to a 256-dimensional embedding vector;
[0050] The discretized RMS values are thus mapped to a higher-dimensional embedding space for better interaction with the model.
[0051] S5: Based on semantic information and 256-dimensional embedding vector training, the audio representation generation module is guided; by discretizing the RMS value and mapping it to the corresponding 256-dimensional embedding, the redundancy caused by direct copying can be avoided, while retaining the effective expression of RMS on the trend of audio sound changes. The present invention not only reduces unnecessary calculations, but also improves the guided generation effect of the model, so that the model can better utilize the RMS curve to generate more natural audio.
[0052] S6: Based on the audio representation generation module, the vocoder is used to restore and generate audio.
[0053] Preferably, the 256-dimensional embedding vector is updated during the training process.
[0054] In summary, compared with the prior art, the present invention has the following advantages:
[0055] The present invention can improve the prediction accuracy and guidance capability of time-related information in video-generated audio, thereby improving the time alignment effect and quality of generated audio and video.
[0056] The above is only a preferred embodiment of the present invention and does not limit the present invention in any way. Any technician in the relevant technical field, without departing from the scope of the technical solution of the present invention, makes any form of equivalent replacement or modification to the technical solution and technical content disclosed in the present invention, which does not depart from the content of the technical solution of the present invention and still falls within the protection scope of the present invention.
Claims
1. A method for generating audio from high-quality video, characterized in that: The following steps are involved: S1: Process the video based on the semantic pre-training model to obtain semantic information; process the video based on the video understanding pre-training model to obtain video understanding features; S2: Obtain video frames of fixed length based on video understanding features; S3: A fixed-length video frame is input into the Seq2Seq model, and the Seq2Seq model outputs the utterance prediction of the audio frame, where the utterance prediction of the audio frame is the RMS value; S4: Discretize the RMS value into 64 discrete values as follows: d(r) = math.floor(64*(ln(1+63|r|) / ln(64))), d(r) is 64 discrete values, r is the RMS value; The discretized RMS value corresponds to a 256-dimensional embedding vector; S5: Based on semantic information and 256-dimensional embedding vector training, it guides the audio representation generation module; S6: Based on the audio representation generation module, the vocoder is used to restore and generate audio.
2. The method for generating audio from high-quality video according to claim 1, wherein: The Seq2Seq model is processed as follows: S31: Input fixed-length video frames into the Seq2Seq model; S32: Process the length of the video frame sequence and construct a mask mark to distinguish the real video frame from the filling part; S33: During training and inference, mask marking is used to ensure that the Seq2Seq model only focuses on the real video frames and ignores the padding parts; S34: The Seq2Seq model outputs a fixed length of predicted audio frame. By multiplying the fixed length of the predicted audio frame with the mask marker, the padding part is filtered out to generate an audio frame utterance prediction that is consistent with the ratio of the actual video frame number.
3. The high-quality video-to-audio generation method according to claim 2, wherein: The way to process the video frame sequence length is as follows: For input video frame sequences that are insufficient in length, padding is used to fill them up to a fixed length. For input sequences that are too long, they are truncated to a fixed length.
4. The method for generating audio from high-quality video according to claim 2, wherein: The mask mark is constructed as follows: the mask value of the real video frame is 1, and the mask value of the filled part is 0.
5. The method for generating audio from high-quality video according to claim 1, wherein: The 256-dimensional embedding vector is updated during the training process.