Method for generating audio from high-quality video

By using semantic pre-training models and video understanding pre-training models to process videos in video generation audio technology, and using Seq2Seq model and vocoder to generate audio, the problem of poor time-related information guidance in the prior art is solved, and a higher quality time alignment effect is achieved.

CN119988671APending Publication Date: 2025-05-13GIANT MOBILE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510098110.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the existing video generation audio technology, the guidance effect of time-related information is poor, resulting in low time alignment and quality of the generated audio and video.

Method used

The pre-trained model based on semantic pre-trained model and video understanding are used to process videos to obtain semantic information and video understanding features. These features are input into the Seq2Seq model for processing, output the vocal prediction of audio frames, and enhance the guidance ability of time-related information through discretization and embedding vectors, and finally generate audio through the vocoder.

Benefits of technology

Improve the prediction accuracy and guidance capability of time-related information in video generated audio, thereby improving the time alignment effect and quality of generated audio and video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988671A_ABST
    Figure CN119988671A_ABST
Patent Text Reader

Abstract

The invention relates to a method for generating an audio from a high-quality video, and the method comprises the following steps: S1, processing the video based on a semantic pre-training model, and obtaining semantic information; processing the video based on the video understanding pre-training model to obtain video understanding features; s2, acquiring a video frame with a fixed length according to the video understanding features; s3, the video frame with the fixed length is input into the Seq2Seq model, the Seq2Seq model outputs sound production prediction of the audio frame, and the sound production prediction of the audio frame is an RMS value; s4, the RMS value is discretized into 64 discrete numerical values, the mode is as follows: d (r) = math.floor (64 * (ln (1 + 63r) / ln (64))), d (r) is the 64 discrete numerical values, and r is the RMS value; the discretized RMS value corresponds to an embedding vector of a 256 dimension, and the RMS value corresponds to the embedding vector of the 256 dimension; s5, guiding an audio representation generation module on the basis of semantic information and 256-dimensional embedding vector training; and S6, based on an audio representation generation module, adopting a vocoder for reduction, and generating audio. According to the invention, the time alignment effect and quality of the generated audio and video can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video sound effect generation, and in particular to a method for generating audio from high-quality video. Background Art

[0002] Common video-generated audio technologies are usually divided into several modules: 1) Video prediction module for information related to the time of voice onset; 2) Module for extracting semantic information from the video; 3) Guide the audio representation generation module based on semantic information and information related to the time of voice onset; 4) Restore the audio representation to the audio module through the vocoder. Common methods for information related to the time of voice onset are: ① The voice label of the video frame: represented by 0, 1, 0 for no voice, 1 for voice, this method cannot represent the intensity, and cannot guide the voice intensity in the generated audio well. ② The energy or root mean square (RMS) of the video frame, represented by a decimal of 0 to 1, the size of this value can represent the intensity, and to a certain extent has a guiding effect on the intensity of the generated audio. In the module for guiding the audio representation generation with relevant time information, the relevant time information will be copied from a 1-dimensional value to the same 256-dimensional value, which is used to guide the input of the generation module. From this point of view, the amount of input information for time-related guidance is actually very small, so it is impossible to achieve a good guiding effect.

[0003] Therefore, it is necessary to provide a high-quality video-to-audio generation method to improve the time alignment effect and quality of the generated audio and video. Summary of the invention

[0004] The object of the present invention is to provide a method for generating audio from high-quality video, so as to improve the time alignment effect and quality of the generated audio and video.

[0005] In order to solve the problems existing in the prior art, the present invention provides a method for generating audio from high-quality video, comprising the following steps:

[0006] S1: Process the video based on the semantic pre-training model to obtain semantic information; process the video based on the video understanding pre-training model to obtain video understanding features;

[0007] S2: Obtain video frames of fixed length based on video understanding features;

[0008] S3: A fixed-length video frame is input into the Seq2Seq model, and the Seq2Seq model outputs the utterance prediction of the audio frame, where the utterance prediction of the audio frame is the RMS value;

[0009] S4: Discretize the RMS value into 64 discrete values ​​as follows:

[0010] d(r) = math.floor(64*(ln(1+63|r|) / ln(64))), d(r) is 64 discrete values, r is the RMS value;

[0011] The discretized RMS value corresponds to a 256-dimensional embedding vector;

[0012] S5: Based on semantic information and 256-dimensional embedding vector training, it guides the audio representation generation module;

[0013] S6: Based on the audio representation generation module, the vocoder is used to restore and generate audio.

[0014] Optionally, in the high-quality video-to-audio generation method, the Seq2Seq model is processed as follows:

[0015] S31: Input fixed-length video frames into the Seq2Seq model;

[0016] S32: Process the length of the video frame sequence and construct a mask mark to distinguish the real video frame from the filling part;

[0017] S33: During training and inference, mask marking is used to ensure that the Seq2Seq model only focuses on the real video frames and ignores the padding parts;

[0018] S34: The Seq2Seq model outputs a fixed length of predicted audio frame. By multiplying the fixed length of the predicted audio frame with the mask marker, the padding part is filtered out to generate an audio frame utterance prediction that is consistent with the ratio of the actual video frame number.

[0019] Optionally, in the method for generating audio from high-quality video, the video frame sequence length is processed as follows:

[0020] For input video frame sequences that are insufficient in length, padding is used to fill them up to a fixed length. For input sequences that are too long, they are truncated to a fixed length.

[0021] Optionally, in the method for generating audio from high-quality video, the mask mark is constructed as follows: the mask value of the real video frame is 1, and the mask value of the filled part is 0.

[0022] Optionally, in the high-quality video audio generation method, a 256-dimensional embedding vector is updated during the training process.

[0023] Compared with the prior art, the present invention has the following advantages:

[0024] The present invention can improve the prediction accuracy and guidance capability of time-related information in video-generated audio, thereby improving the time alignment effect and quality of generated audio and video. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 A flow chart of a method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0026] The specific implementation of the present invention will be described in more detail below in conjunction with the schematic diagram. The advantages and features of the present invention will become clearer based on the following description. It should be noted that the drawings are all in a very simplified form and are not in exact proportions, and are only used to facilitate and clearly assist in explaining the purpose of the embodiments of the present invention.

[0027] Hereinafter, if the method described herein includes a series of steps, the order in which the steps are presented herein is not necessarily the only order in which the steps may be performed, and some of the steps described may be omitted and / or some other steps not described herein may be added to the method.

[0028] Common video-generated audio technologies are usually divided into several modules: 1) Video prediction module for information related to the time of voice onset; 2) Module for extracting semantic information from the video; 3) Guide the audio representation generation module based on semantic information and information related to the time of voice onset; 4) Restore the audio representation to the audio module through the vocoder. Common methods for information related to the time of voice onset are: ① The voice label of the video frame: represented by 0, 1, 0 for no voice, 1 for voice, this method cannot represent the intensity, and cannot guide the voice intensity in the generated audio well. ② The energy or root mean square (RMS) of the video frame, represented by a decimal of 0 to 1, the size of this value can represent the intensity, and to a certain extent has a guiding effect on the intensity of the generated audio. In the module for guiding the audio representation generation with relevant time information, the relevant time information will be copied from a 1-dimensional value to the same 256-dimensional value, which is used to guide the input of the generation module. From this point of view, the amount of input information for time-related guidance is actually very small, so it is impossible to achieve a good guiding effect.

[0029] In order to solve the problems existing in the prior art, the present invention provides a method for generating audio from high-quality video, such as Figure 1 As shown, the following steps are included:

[0030] S1: Process the video based on the semantic pre-training model to obtain semantic information; process the video based on the video understanding pre-training model to obtain video understanding features;

[0031] S2: Each video frame is a picture. Usually, picture features are extracted, such as RGB features, which record the color, shape and texture of objects in the picture to represent static features; Flow features, which record the movement of objects and pixels to represent dynamic features. These features can better represent the local information of the video, but cannot represent the global information of the video. It is necessary to use a video understanding model to extract the global information of the entire video. Combining the local information of the video with the global information of the video can better predict the sound of the video.

[0032] Therefore, video frames of fixed length are obtained based on the video understanding features;

[0033] S3: Since the temporal granularity of video frames is usually large, that is, the number of video frames per second is much less than the number of audio frames. For example, the number of video frames in 1 second is 30, while the number of audio frames in 1 second is 100. In order to map video frames to audio frames, the current model structure usually uses two methods to expand the features of video frames: interpolation or replication. However, these two methods often result in each predicted video frame not being perfectly aligned to the audio frame.

[0034] Therefore, the present invention inputs a video frame of a fixed length into a Seq2Seq model, and the Seq2Seq model outputs a voicing prediction of an audio frame, where the voicing prediction of the audio frame is an RMS value (root mean square value);

[0035] Furthermore, the Seq2Seq model is processed as follows:

[0036] S31: Input fixed-length video frames into the Seq2Seq model;

[0037] S32: Process the length of the video frame sequence and construct a mask mark to distinguish the real video frame from the filling part;

[0038] Specifically, the length of the video frame sequence is processed as follows: for an insufficient input video frame sequence, pad it to a fixed length through padding, and for an overlong input, truncate it to a fixed length. The mask mark is constructed as follows: the mask value of the real video frame is 1, and the mask value of the padded part is 0.

[0039] S33: During training and inference, mask marking is used to ensure that the Seq2Seq model only focuses on the real video frames and ignores the padding parts;

[0040] S34: The Seq2Seq model outputs a fixed length of predicted audio frame. By multiplying the fixed length of the predicted audio frame with the mask marker, the padding part is filtered out to generate an audio frame utterance prediction that is consistent with the ratio of the actual video frame number.

[0041] In one embodiment,

[0042] (1) Input feature extraction: The input of the Seq2Seq model is fixed to the features of 300 video frames (10s video).

[0043] (2) Sequence length processing. For input video frame sequences that are not long enough, pad them to 300. For input sequences that are too long, cut off the first 300 frames. At the same time, construct a mask tag: the mask value of the real video frame is 1, and the padding part is 0.

[0044] (3) During training and inference, masks are used to ensure that the model only focuses on the real frames and ignores the filled parts.

[0045] (4) Model output processing. The output audio frame level occurrence of the model (i.e., the predicted discretized RMS value) is fixed to 1000 frames. By multiplying the predicted audio frame with the mask and filtering out the padding part, an audio frame sound prediction that is consistent with the ratio of the actual video frame number is generated.

[0046] The Seq2Seq model solves the time alignment problem caused by traditional interpolation or replication methods by inputting video frame features and predicting the utterance of audio frames. By padding or truncating video frames of different lengths and combining mask operations, it ensures that the output of the model can accurately generate audio utterances in proportion to the number of video frames.

[0047] S4: Discretize the RMS value into 64 discrete values ​​as follows:

[0048] d(r) = math.floor(64*(ln(1+63|r|) / ln(64))), d(r) is 64 discrete values, r is the RMS value, math.floor is a function used to round down floating point numbers;

[0049] The discretized RMS value corresponds to a 256-dimensional embedding vector;

[0050] The discretized RMS values ​​are thus mapped to a higher-dimensional embedding space for better interaction with the model.

[0051] S5: Based on semantic information and 256-dimensional embedding vector training, the audio representation generation module is guided; by discretizing the RMS value and mapping it to the corresponding 256-dimensional embedding, the redundancy caused by direct copying can be avoided, while retaining the effective expression of RMS on the trend of audio sound changes. The present invention not only reduces unnecessary calculations, but also improves the guided generation effect of the model, so that the model can better utilize the RMS curve to generate more natural audio.

[0052] S6: Based on the audio representation generation module, the vocoder is used to restore and generate audio.

[0053] Preferably, the 256-dimensional embedding vector is updated during the training process.

[0054] In summary, compared with the prior art, the present invention has the following advantages:

[0055] The present invention can improve the prediction accuracy and guidance capability of time-related information in video-generated audio, thereby improving the time alignment effect and quality of generated audio and video.

[0056] The above is only a preferred embodiment of the present invention and does not limit the present invention in any way. Any technician in the relevant technical field, without departing from the scope of the technical solution of the present invention, makes any form of equivalent replacement or modification to the technical solution and technical content disclosed in the present invention, which does not depart from the content of the technical solution of the present invention and still falls within the protection scope of the present invention.

Claims

1. A method for generating audio from high-quality video, characterized in that: The following steps are involved: S1: Process the video based on the semantic pre-training model to obtain semantic information; process the video based on the video understanding pre-training model to obtain video understanding features; S2: Obtain video frames of fixed length based on video understanding features; S3: A fixed-length video frame is input into the Seq2Seq model, and the Seq2Seq model outputs the utterance prediction of the audio frame, where the utterance prediction of the audio frame is the RMS value; S4: Discretize the RMS value into 64 discrete values ​​as follows: d(r) = math.floor(64*(ln(1+63|r|) / ln(64))), d(r) is 64 discrete values, r is the RMS value; The discretized RMS value corresponds to a 256-dimensional embedding vector; S5: Based on semantic information and 256-dimensional embedding vector training, it guides the audio representation generation module; S6: Based on the audio representation generation module, the vocoder is used to restore and generate audio.

2. The method for generating audio from high-quality video according to claim 1, wherein: The Seq2Seq model is processed as follows: S31: Input fixed-length video frames into the Seq2Seq model; S32: Process the length of the video frame sequence and construct a mask mark to distinguish the real video frame from the filling part; S33: During training and inference, mask marking is used to ensure that the Seq2Seq model only focuses on the real video frames and ignores the padding parts; S34: The Seq2Seq model outputs a fixed length of predicted audio frame. By multiplying the fixed length of the predicted audio frame with the mask marker, the padding part is filtered out to generate an audio frame utterance prediction that is consistent with the ratio of the actual video frame number.

3. The high-quality video-to-audio generation method according to claim 2, wherein: The way to process the video frame sequence length is as follows: For input video frame sequences that are insufficient in length, padding is used to fill them up to a fixed length. For input sequences that are too long, they are truncated to a fixed length.

4. The method for generating audio from high-quality video according to claim 2, wherein: The mask mark is constructed as follows: the mask value of the real video frame is 1, and the mask value of the filled part is 0.

5. The method for generating audio from high-quality video according to claim 1, wherein: The 256-dimensional embedding vector is updated during the training process.