A controllable video soundtrack generation method based on cross-modal alignment mechanism
By performing shot and object segmentation on video content, extracting multiple features and generating music using a non-autoregressive network, the problems of lack of controllability and fine-grained modeling in existing methods are solved, and accurate matching and personalized creation of music and video content are achieved.
Patent Information
- Application Number
- CN202510999352.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-07-21
AI Technical Summary
Existing automated video soundtrack generation methods lack controllability and fine-grained structure modeling capabilities, making it difficult to achieve fine-grained control over specific video image areas and musical styles, resulting in inaccurate matching between the generated music and the video content.
By dividing the video content into shot sequences and performing object segmentation, semantic embedding vectors, area features, starting position features, color features and motion vectors are extracted, and a non-autoregressive network is used to generate music feature sequences. The deep consistency of music and video is achieved through a cross-modal alignment mechanism, and a featureless guidance strategy is introduced for user controllable adjustment.
It achieves refined control over video content, enabling precise matching with music in both time and space, enhancing the flexibility of personalized creation and the accuracy of generated music.
Smart Images

Figure CN120510556B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of music generation, and in particular relates to a controllable video soundtrack generation method based on a cross-modal alignment mechanism. Background Art
[0002] With the rapid development of digital media technology, film and television have become an important vehicle for modern cultural communication. As an integral component of video content, music plays a key role in evoking emotion, reinforcing narrative, and creating an immersive audience experience. However, traditional video soundtrack production processes rely heavily on manual creation, resulting in high production costs and long production cycles. It also struggles to adapt to the rapidly iterating demands of short videos and personalized content. In recent years, deep learning-based automatic music generation technology has provided a new solution to this dilemma. By analyzing video content and generating corresponding background music, this approach not only offers significant advantages in efficiency and scale, but also addresses copyright compliance issues to a certain extent.
[0003] Patent application publication number CN112685592A discloses a method and apparatus for generating music for sports videos, relating to the fields of video processing and cloud computing. The specific implementation comprises: obtaining an action rhythm node sequence corresponding to a sports video; searching for at least one audio rhythm node sequence matching the action rhythm node sequence within one or more audio rhythm node sequences corresponding to an audio set, wherein the audio set includes one or more audio units; and searching an index representing the correspondence between audio units and audio rhythm node sequences for the audio unit corresponding to the at least one audio rhythm node sequence, as the audio unit for the music for the sports video. This patent application uses action rhythm nodes and audio rhythm nodes to intelligently and automatically generate music for sports videos, effectively improving the accuracy of the music. While this patent application generates music based on global features (such as the overall video content), which improves the quality of the generated music, it lacks fine-grained control over specific video elements.
[0004] Current approaches to automated video soundtrack generation still suffer from two key limitations:
[0005] 1. Lack of controllability: Existing methods often employ end-to-end "black box" generation mechanisms, which can only generate music that aligns with the overall semantics of the video. This makes it difficult to actively direct the model to focus on specific areas of the video (such as people, actions, and colors) or customize the musical style (such as mood and rhythm), limiting the scope for personalized creation.
[0006] 2. Lack of fine-grained structural modeling capabilities: Some methods attempt to introduce explicit visual features (such as motion and color) as generation conditions, but the granularity is usually coarse and the structural expression is simple. They lack joint modeling of the video in the temporal dimension (such as shot switching and dynamic rhythm) and the spatial dimension (such as picture composition and subject position). As a result, the generated music cannot accurately match the video content in terms of emotional changes and rhythmic dynamics.
[0007] Therefore, there is an urgent need for a video soundtrack generation method that combines high generation quality and refined control capabilities. Summary of the Invention
[0008] The present invention provides a controllable video soundtrack generation method based on a cross-modal alignment mechanism. This method enables users to flexibly adjust the multimodal features in the generation process, achieving deep consistency between music and video in structure, emotion, and rhythm, to meet the increasingly diverse creative needs of practical application scenarios.
[0009] The present invention provides a controllable video soundtrack generation method based on a cross-modal alignment mechanism, comprising:
[0010] The video content is divided into shot sequences, and the first frame image in each shot is segmented to obtain a set of object area images of each first frame;
[0011] Constructing a training model, the training model includes a feature encoding and conversion module and a music generation module, the feature encoding and conversion module is used to extract features from each object area picture to obtain a semantic embedding vector, area features, starting position features, color features and motion vectors, respectively encode each shot, starting position feature and color feature to obtain multiple shot codes, starting position codes and color codes, fuse the semantic embedding vector of each object area picture with the shot code of the corresponding shot and multiply by the corresponding area feature to obtain a first fused feature, fuse the first fused feature, starting position code and color code to obtain a second fused feature, fuse the second fused feature and the motion vector through a block matrix to obtain a dynamic video feature, and pass the dynamic video feature through a non-autoregressive network to obtain a predicted music feature sequence; the music generation module is used to decode the multiple predicted music feature sequences through a transformer to obtain a predicted music embedding vector, and perform music decoding on the predicted music embedding vector to obtain a predicted audio token;
[0012] Constructing a loss function, wherein the loss function includes a music prediction loss function and a music feature loss function, wherein the music prediction loss function is constructed by using a cross loss between the predicted audio token and the real audio code value, and the music feature loss function is constructed by using an L1 loss based on the real music features and the predicted music features;
[0013] Based on the object area picture set, the training model is trained by the loss function to obtain a video soundtrack generation model. When applied, the video is input into the video soundtrack generation model to obtain corresponding music.
[0014] Preferably, feature extraction is performed on each object region image to obtain a semantic embedding vector, area feature, starting position feature, color feature, and motion vector, including:
[0015] The pre-trained ViT model is used to extract region-level semantic representations for each object region image, and then the semantic embedding vector is obtained through the self-attention mechanism;
[0016] The ratio of the number of pixels in the object area to the total number of pixels in the corresponding first frame image is used as the area feature;
[0017] Normalize the position data of the object area in the first frame image to obtain the starting position feature;
[0018] The 256-dimensional color histograms of the three RGB channels of the object area are spliced to obtain the color features;
[0019] The difference between the position feature of the object area in the current frame image and the position feature of the first frame image is used as the motion vector.
[0020] Preferably, encoding each shot, starting position feature, and color feature separately to obtain multiple shot codes, starting position codes, and color codes includes:
[0021] Fourier frequency method is used to encode each shot to obtain multiple shot codes;
[0022] The starting position feature is encoded by the sine and cosine functions, wherein the first half dimension encodes the horizontal coordinate of the starting position feature and the second half dimension encodes the vertical coordinate of the starting position feature, so as to form two-dimensional data, thereby obtaining the starting position code;
[0023] Color coding is achieved by encoding each pixel based on the color intensity of each pixel in the object area.
[0024] Preferably, the starting position feature is encoded to obtain a starting position code, wherein the j-th object area is in the shot i, and the values of the dim, dim+1 dimensions are for: ,in, and are the horizontal and vertical coordinates of the starting position feature of the j-th object area in the lens i, is the scaling factor, is the total dimension.
[0025] Preferably, the color features are encoded to obtain color codes of the red, green and blue channels, and the color codes of the red, green and blue channels are spliced to obtain the final color code, wherein the encoding value of the k-th pixel in the j-th object area in the i-th shot under the color channel C is for: ,in, is the total dimension, C is the color channel, is the number of pixels in the current region, in color channel C, whose color value is equal to k.
[0026] Preferably, the second fusion feature and the motion vector are fused through a block matrix to obtain a dynamic video feature, including:
[0027] The motion features are mapped to angle offsets in the form of trigonometric functions, and the starting position encoding is linearly transformed based on the angle offset to obtain the dynamic video features of each object area at time step t.
[0028] Preferably, before the dynamic video features are input into the non-autoregressive network, they are first dimensionally adjusted and aligned through multi-layer linear transformations, and then the contextual information of the dynamic video features in the time dimension is captured through a multi-head self-attention mechanism.
[0029] Preferably, the dynamic video features are sequentially passed through a non-autoregressive network to obtain a predicted music feature sequence, and the predicted music feature sequence is upsampled through a one-dimensional transposed convolution to obtain a final predicted music feature sequence, wherein the predicted music feature sequence includes pitch, loudness, chroma and timbre feature sequences.
[0030] Preferably, multiple predicted music feature sequences are decoded through a transformer to obtain a predicted music embedding vector, including:
[0031] Multiple predicted music feature sequences are sequentially subjected to feature concatenation and one-dimensional convolution. A multi-head self-attention mechanism is performed based on the result of the one-dimensional convolution and the corresponding real music embedding vector to obtain the predicted music embedding vector. The result of the one-dimensional convolution is used as the query, and the real music embedding vector is used as the key and value.
[0032] Preferably, a featureless guidance strategy is introduced after obtaining the semantic embedding vector, area feature, starting position feature, color feature and motion vector, and after obtaining the predicted music feature sequence, respectively. The featureless guidance strategy is used to enable the user to control the weights of the semantic embedding vector, area feature, starting position feature, color feature, motion vector and predicted music feature sequence.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] The present invention deconstructs the video content to obtain the area, starting position, color and motion vector features of the object area with clear semantics and composition attributes, and integrates the above features through specific encoding, so that the user can actively guide the model to focus on specific picture areas in the video (such as characters, actions, colors) or realize customized adjustment of music style (such as emotions, rhythm), which can realize personalized creation. It can also jointly model the time dimension (such as shot switching, dynamic rhythm) and spatial dimension (such as picture composition, subject position), so that the music can accurately match the video content in terms of emotional changes and rhythm dynamics. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 A controllable video soundtrack generation method based on a cross-modal alignment mechanism is provided in a specific embodiment of the present invention. DETAILED DESCRIPTION
[0036] This paper proposes a controllable video soundtrack generation method based on spatial-temporal deconstruction and cross-modal alignment. The method aims to achieve high-quality, multi-style, and controllable music generation driven by video content, meeting users' personalized soundtrack needs in diverse scenarios. This method incorporates shot segmentation and intra-frame image segmentation in the temporal domain to extract object units with clear semantic and compositional attributes (such as area, position, color, and motion). These are encoded into structured control features, which are then converted into musical features such as pitch, loudness, and timbre through a non-autoregressive neural network and an attention mechanism. Ultimately, background music is generated that is semantically consistent and emotionally compatible with the input video content.
[0037] The present invention provides a controllable video soundtrack generation method based on a cross-modal alignment mechanism, such as Figure 1 Shown, including:
[0038] S1. Video content structure: The video content is divided into shot sequences, and the first frame image in each shot is segmented to obtain a set of object area images for each first frame. The specific steps are as follows:
[0039] In the time dimension, the specific embodiment of the present invention uses the pre-trained shot segmentation model TransNet V2 to detect the shot boundaries of the video frame sequence, identify the shot segments with relatively consistent semantics and rhythm, and represent the original video as a shot sequence. , where V is the input video, S N For the Nth shot.
[0040] In terms of spatial dimension, the specific embodiment of the present invention selects the first frame image of each shot, uses the Segment Anything Model (SAM) to perform intra-frame object segmentation, and extracts image regions with semantic meaning.},in, The jth segmented image region in the i-th frame (e.g., an object, a person, a background block) is retained in the present invention to avoid over-segmentation leading to sparse control information. n The large object areas are combined into the complementary areas, and the rest are expressed as: , thus obtaining a set of object area images. This ensures that all pixels are covered and the area division has control efficiency.
[0041] Furthermore, in this embodiment of the present invention, considering the relatively stable semantics and composition of objects within a shot, the system performs spatial segmentation only on the first frame of each shot to reduce redundant computations. The motion features in subsequent frames are extracted using a motion trajectory module. The shot sequence and image region set output by this module serve as the basis for constructing a video feature extraction and control unit.
[0042] S2. Construct a training model, which includes a feature encoding and conversion module and a music generation module. The feature encoding and conversion module is used to extract structured features with music control potential from the video image area and encode them into a sequence representation for generating guidance. The feature encoding and conversion module encodes and converts around five types of video features: semantics, area, starting position, motion and color.
[0043] The feature encoding and conversion module provided in a specific embodiment of the present invention is used to extract features from each object area image to obtain a semantic embedding vector, area features, starting position features, color features and motion vectors, encode each shot, starting position feature and color feature respectively to obtain multiple shot codes, starting position codes and color codes, fuse the semantic embedding vector of each object area image with the shot code of the corresponding shot and multiply them by the corresponding area feature to obtain a first fused feature, fuse the first fused feature, the starting position code and the color code to obtain a second fused feature, fuse the second fused feature and the motion vector through a block matrix to obtain a dynamic video feature, and pass the dynamic video feature through a non-autoregressive network to obtain a predicted music feature sequence.
[0044] Specifically, a specific embodiment of the present invention extracts video features, wherein the step of extracting a semantic embedding vector includes:
[0045] The pre-trained ViT model is used to extract region-level semantic representations. The system divides the image into several patches and calculates the "cls token" vector of each region as the semantic embedding through the self-attention mechanism. The formula is: ,in, represents the semantic embedding vector of the jth segmented image region in the i-th frame, (Query) is the query vector, similar to the request, (Key): key vector, similar to index, (Value) is a value vector, which represents the value table from which the value is taken. To represent the similarity (dot product) between Query and Key, : scaling factor, It is the dimension of the key, which is used to prevent the dot product value from being too large, resulting in the softmax gradient being too small. Q, K, and V are three sets of vectors obtained by transforming the image patch features of the object area, which are used to extract the semantic relationship between the area and other context areas through the self-attention mechanism.
[0046] The area feature provided by the specific embodiment of the present invention is the relative size of the object area in the first frame image, which is used to measure its visual weight. The calculation method is the ratio of the number of pixels in the object area to the total number of pixels in the corresponding first frame image: , Refers to the relative area (normalized) of the jth segmented region in the i-th frame image - used to calculate area features, Refers to the area The number of non-zero pixels in the region is called the "pixel area" of the region. Refers to the height (Height) multiplied by the width (Width) of the first frame image, that is, the total number of pixels in the entire image.
[0047] The starting position feature provided in the specific embodiment of the present invention is used to describe the layout center of the image area in the composition. The present invention normalizes the coordinates of the area center to In the range, the normalized starting position coordinates of the jth segmentation area in the i-th shot are , contains two components (horizontal position, vertical position), ,in, Refers to the area The horizontal coordinate of the center point in the image; Refers to the vertical axis.
[0048] The specific embodiment of the present invention reflects the regional hue and atmosphere through color features. The present invention splices the 256-dimensional color histograms of the RGB three channels of the object area to obtain the color features, wherein the color histogram statistics of the j-th segmented object area in the i-th shot are performed, and the 256-dimensional color distribution vector is extracted to obtain the color features. for: The meaning of this formula is to perform color histogram statistics on the j-th segmented area in the i-th shot, and count the number of pixels with each color value (color level) in the image (or image area), that is, to count the frequency of occurrence of each color value.
[0049] The motion vector provided by the specific embodiment of the present invention is the dynamic displacement of the object in the time dimension, and the positioning is the difference between the target position in the current frame and its initial frame position, where the displacement vector of the jth segmentation area in the i-th shot at time t is for: - ,in, and They represent the position of the object area in the current frame and the first frame respectively. The features are used to characterize the dynamic behavior of the visual object.
[0050] This embodiment of the present invention encodes each shot to generate a shot code. The goal is to distinguish different shots and reflect their relative temporal relationships, thereby effectively utilizing the temporal information of the video. This embodiment of the present invention uses a Fourier frequency method, similar to the position encoding in the Transformer. This method facilitates the model to learn an attention mechanism based on relative position. The specific calculation is as follows: , ,in, represents the encoding value of the even dimension in the i-th shot, represents the encoding value of the odd dimension in the i-th shot, The number index of the current dimension, , N is a natural number, Represents the total dimension of the encoding vector, for example 128 dimensions, and the semantic embedding vector The dimensions are consistent. In the specific embodiment of the present invention, each lens i will generate a dimension of Vector , used to represent the relative order of shots on the time axis, and the sine-cosine function is used to generate periodic position relationship encoding. For a fixed offset a, the encoding value can be transformed from get.
[0051] The specific embodiment of the present invention uses the area feature as a weight to obtain the regional fusion feature representation of semantic, area, and shot features, that is, the first fusion feature, where the first fusion feature of the jth object area in the i-th shot is for: , Represents the area feature of the jth object area in the i-th shot, that is, the visual weight; is the semantic embedding vector; The lens code for lens i.
[0052] The position code provided in the specific embodiment of the present invention is similar to the shot code, but since it is two-dimensional data (horizontally , vertical coordinates), so the distinction is made as follows: the first half of the dimension of the encoding vector is used Coordinate encoding, the second half dimension is used for Coordinate encoding; the encoding method uses sine and cosine functions to maintain the periodicity and relative computability of the position; to control the distribution range of each dimension information, a dimension-related scaling factor is introduced , respectively used for and Direction, the specific formula is as follows: ,in, and are the horizontal and vertical coordinates of the starting position feature of the j-th object area in the lens i, is the scaling factor, is the total dimension, (If coordinates) or (If coordinates), this encoding method can be directly combined with motion features in the subsequent process By combining linear transformations, the model can perceive the correlation between the starting position and the motion trend, further enhancing the object's space-time modeling capabilities. The present invention places the sin value in the dim dimension and the corresponding cos value in the dim+1 dimension. This pair of values can uniquely determine a specific angle (position) without confusion.
[0053] The color coding provided by the specific embodiment of the present invention is different from the coding method related to time sequence (such as lens position or starting position). The color information itself does not have obvious order. In order to preserve the distribution of the color histogram in the image, the specific embodiment of the present invention divides each pixel into Considered as a dimension , so that each pixel is encoded separately. It can reflect the color intensity, and the specific embodiment of the present invention retains this feature in the encoding.
[0054] In addition, in order to emphasize the differences between pixels, a specific embodiment of the present invention introduces a pixel-related scaling factor, as shown in the following formula, so as to strengthen the pixel differences and the correlation between colors while preserving the color distribution, where the encoding value of the k-th pixel in the j-th object area in the i-th shot under the color channel C is for: ,in, is the total dimension, k is the pixel index, C is the color channel, such as red (R), green (G), blue (B), is the number of pixels in the color channel C in the current image area whose color value is equal to k. For example, there are 4 pixels in the object area, each of which contains three color channels: red (R), green (G), and blue (B). The specific values are as follows: the first pixel: R = 100, G = 150, B = 200; the second pixel: R = 100, G = 150, B = 200; the third pixel: R = 80, G = 120, B = 180; the fourth pixel: R = 100, G = 130, B = 190. Therefore, in the red channel, the pixel value 100 appears 3 times and the pixel value 80 appears once. Therefore, for the red channel, when k=100, , when k=80, .
[0055] The specific embodiment of the present invention further adds the above feature results and uses them as the static representation feature of the object , the time factor will be further introduced in the feature transformation process. Feature fusion is expressed as: ,in, The static features of the object after final fusion, Color coding of the three channels of red, green and blue, Splicing operation.
[0056] After completing the static visual encoding described above, the system in this embodiment of the present invention constructs a multidimensional embedding representation for each object region, including semantics, starting position, and color distribution. These encodings not only reflect the compositional information and perceptual attributes of the object in the image, but also provide a well-structured input foundation for subsequent temporal modeling.
[0057] However, static coding only reflects the state of an object in a single shot, and it is difficult to capture the dynamic changes of the shot over time. For this reason, a specific embodiment of the present invention introduces a feature transformation and alignment mechanism, which linearly transforms the static features through the motion trajectory of the object, thereby obtaining dynamic video features with temporal continuity, which are further used for the generation and matching of music features.
[0058] The specific embodiment of the present invention utilizes motion characteristics The static features are converted into temporal features, which are then input into a non-autoregressive neural network to generate a continuous sequence of music features. Since the initial position encoding is linearly additivity, the specific embodiment of the present invention can use motion to model features in the time dimension: since the starting position encoding is essentially expressing the center position of the region in the form of a trigonometric function, in the position encoding, the specific embodiment of the present invention encodes the column vector of the starting position as follows: ,in, is the starting position angle, derived from the position code The sine and cosine values in . The angular offset caused by the motion is , then the result after transformation is It can be obtained by the following linear transformation: , is the position offset caused by motion, A1 is the motion vector The two-dimensional rotation matrix obtained by mapping, A2 is the sine cosine form extracted from the starting position encoding, and the specific embodiment of the present invention converts the motion feature into a trigonometric function form. Mapping is done as an angular offset, thereby dynamically updating the original position code. Based on this mechanism, the specific embodiment of the present invention further uses block matrix multiplication to combine the motion information of each object with its static features, where: The 2*2 sub-block comes from the motion vector The trigonometric transformation of The 2*1 sub-block comes from the starting position code of the object The two constitute the multiplication term on the right side of the formula, which is an expression of spatial rotation transformation. Finally, the specific embodiment of the present invention obtains the dynamic video features of each object at time step t. After time series modeling, the dynamic video feature vector of the jth image area in the i-th shot at time step t is , Indicates that the motion vector With static features Fusion is performed via block matrix multiplication.
[0059] After obtaining preliminary dynamic video features, this embodiment of the present invention uses multi-layer linear transformations to resize and align these features. A feedforward transformer (FFT) architecture is then introduced, using a multi-head self-attention mechanism to capture the temporal context of these features. This attention mechanism dynamically reallocates weights, effectively selecting time-dependent features that are useful for generation.
[0060] Subsequently, a specific embodiment of the present invention uses a music feature predictor based on a Transposed-Conv1D network to predict music features. Because the attention module has already extracted temporal information, this specific embodiment of the present invention uses a non-autoregressive network to generate the music feature sequence in parallel. Compared to the autoregressive model generated step by step, the non-autoregressive structure significantly improves prediction speed.
[0061] The dropout layer provided in this embodiment of the present invention, after applying the attention mechanism to video features, prevents the model from relying solely on certain local features, which can lead to poor generalization. By randomly dropping some neurons (that is, temporarily setting their outputs to 0), the model is forced to learn to utilize different feature combinations during training, improving its adaptability to new data.
[0062] In the music generation phase, the one-dimensional transposed convolution provided by this embodiment of the present invention uses the previously extracted feature sequences as low-dimensional sequences after compression (e.g., through a Conv1D or linear layer). Using Transposed Conv1D, these sequences can be restored to longer time frames, gradually generating a complete time series of music features (such as pitch, loudness, chroma, and timbre).
[0063] This embodiment of the present invention uses pitch, loudness, chroma, and timbre (represented by the spectral centroid) as core musical features and serves as the generation target. During the training phase, the system calculates the L1 loss between the predicted and true features, and then achieves feature alignment (FA) through backpropagation. The loss function can be expressed as: ,in, The real music features at the mth time step, is the predicted music feature sequence of the mth time step, where is the loss of music features, indicating the error of the model in generating music features, and t is the total number of time steps.
[0064] The music generation module provided in the specific embodiment of the present invention is intended to generate a matching music clip based on the multi-dimensional video control features output by the preceding module, taking into account semantic consistency, emotional alignment and user controllability.
[0065] The music generation module provided in a specific embodiment of the present invention is used to decode multiple predicted music feature sequences through a transformer to obtain a predicted music embedding vector, and then perform music decoding on the predicted music embedding vector to obtain a predicted audio token. The specific process is as follows:
[0066] The specific embodiment of the present invention concatenates all the predicted music features as the input of the music generation module, thereby avoiding cross-modal loss. In this module, the specific embodiment of the present invention first uses a one-dimensional convolutional network to compress the music features, because this structure has the advantages of efficient calculation and strong robustness when processing one-dimensional time series data. Subsequently, a multi-head attention mechanism is introduced to guide music generation. The structure of this module is similar to that of a Transformer using only a decoder, in which the query vector (Query) comes from the real music embedding. (And right shift operation), that is, refer to the previously generated music clips. The multi-head attention mechanism in time is specifically expressed as: That is, through the attention mechanism, the model can capture the relationship between contextual music features, thereby enhancing the global dependency and coherence of the generated music. Since the music features are only spliced and compressed, they themselves remain unchanged, so Fm' is still used here to represent the predicted music feature sequence, where : The original music embedding vector (real music sequence), but shifted right (Shifted Right), as the attention query, : The original music embedding vector (real music sequence), but right-shifted (ShiftedRight), as the attention query, Predicted music features from the video, used as the key and value of attention, : The music embedding prediction value output by the model serves as the representation required for the final music generation.
[0067] In order to measure the difference between the generated result and the true target, the specific embodiment of the present invention uses each codebook Average cross-entropy loss to compare predicted music embeddings The difference between 𝑀 and the real music token 𝑀. Its loss function is defined as follows: ,in, is the average cross entropy loss (Loss) of music prediction, Indicates the number of codebooks, indicating how many independent coding spaces are used in the quantization layer, T indicates the time in seconds, indicating the temporal length of the music, Indicates the frame rate / sampling rate, Indicates real token In the The code value in the codebook (the value is 0 or 1), The output of the model The probability of the predicted audio token at the position (corresponding to the softmax probability of a token at that position)
[0068] After completing cross-temporal attention modeling, this embodiment of the present invention calls a music decoder to convert the predicted music embedding vector into the final audio format. This embodiment of the present invention utilizes a pre-trained MusicGen decoder with its parameters frozen. MusicGen internally employs the Encodec decoding framework, which uses Residual Vector Quantization (RVQ) to compress the audio stream and generate a discrete token representation. This approach not only effectively reduces spatial complexity but also, thanks to extensive training on large-scale datasets, significantly improves the diversity and expressiveness of the generated music.
[0069] After completing the decomposition and alignment of the aforementioned spatiotemporal features, the system has achieved cross-modal generation from video to music. Throughout the generation process, the model fully utilizes all features of both the video and the music. However, in real-world applications, users often do not need to pay attention to all features. They may wish to emphasize certain features, ignore others, or assign different importance weights to them.
[0070] To meet such controllability requirements, it is desirable to introduce some simulated control data during the training phase. However, the manual effort of real-time guided training is high, so the specific embodiment of the present invention designs a feature-free guidance (FFG) strategy to simulate user control of features. The specific approach is: with a certain probability , randomly changing the combined weights of different features, or even directly discarding some features. This mechanism supports controlling the final generated music features from both video and music dimensions. Its training expression is: , The feature perturbation model used in the final training (i.e., the prediction behavior of the entire model under this batch), Represents the inference output of the original model (e.g. a Transformer network) given the features. Represents the original video feature vector; represents the original music feature vector, Indicates the sampling probability of whether to use "feature perturbation" (for example, if it is set to 0.3, 30% of the samples will use perturbation features, and the rest will use original features). , are the perturbation weights of video and music features respectively (vectors, not scalars).
[0071] The specific embodiment of the present invention provides and The feature weights are randomly set during the training process. The reason is that video features are relatively discrete and redundant, so they can be controlled in an on-off manner. However, music features are continuous and combinable, so they need to meet the normalization constraint: These constraints help improve the stability and generalization of the training process. Through this mechanism, users can control the model to focus on specific parts of the video or adjust the weight of features in the generated music, thereby achieving personalized music generation output and enhancing user controllability and satisfaction.
[0072] In summary, this module realizes high-quality music generation driven by video semantics and structural features, and ensures the accuracy, stability and adaptability of the generated content through multiple alignment mechanisms and controllable strategies.
[0073] Specific embodiments of the present invention propose a model that balances control and high-quality output to address the controllability challenge in video-to-music generation. Traditional approaches to generating music from videos often rely on global features (such as the overall video content) to generate music. While this improves the quality of generated music, it lacks fine-grained control over specific video elements. The generation process fails to clearly distinguish between different video components and details, limiting the flexibility users need to customize their music.
[0074] This paper proposes a new model that not only focuses on high-quality generation but also addresses the lack of controllability found in traditional methods. Specifically, by decomposing a video into controllable elements in both time and space, this approach enhances the user's fine-grained control over the music generation process. This allows users to precisely adjust characteristics such as emotion and rhythm to their needs, rather than relying on simple global features. This approach allows users greater creative freedom in music generation, avoiding the limitations of traditional methods that lack precise control.
[0075] Specific embodiments of the present invention propose a feature-free spatiotemporal decomposition and alignment method that introduces a simulation mechanism during the training phase to narrow the control distribution gap between training and inference. Unlike traditional methods, this method introduces a feature-free guidance strategy that randomly drops, emphasizes, or weights some controllable features during training (i.e., Random / Controllable Drop) to simulate real-world scenarios where users have incomplete control, ambiguous expressions, or diverse preferences, thereby narrowing the control distribution gap between training and inference. This strategy does not rely on manual labeling, significantly reducing training costs and improving the model's controllability and generalization capabilities in practical applications.
Claims
1. A controllable video soundtrack generation method based on a cross-modal alignment mechanism, characterized in that: include: The video content is divided into shot sequences, and the first frame image in each shot is segmented to obtain a set of object area images of each first frame; Constructing a training model, the training model includes a feature encoding and conversion module and a music generation module, the feature encoding and conversion module is used to extract features from each object area picture to obtain a semantic embedding vector, area features, starting position features, color features and motion vectors, respectively encode each shot, starting position feature and color feature to obtain multiple shot codes, starting position codes and color codes, fuse the semantic embedding vector of each object area picture with the shot code of the corresponding shot and multiply by the corresponding area feature to obtain a first fused feature, fuse the first fused feature, starting position code and color code to obtain a second fused feature, fuse the second fused feature and the motion vector through a block matrix to obtain a dynamic video feature, and pass the dynamic video feature through a non-autoregressive network to obtain a predicted music feature sequence; the music generation module is used to decode the multiple predicted music feature sequences through a transformer to obtain a predicted music embedding vector, and perform music decoding on the predicted music embedding vector to obtain a predicted audio token; Constructing a loss function, wherein the loss function includes a music prediction loss function and a music feature loss function, wherein the music prediction loss function is constructed by using a cross loss between the predicted audio token and the real audio code value, and the music feature loss function is constructed by using an L1 loss based on the real music features and the predicted music features; Based on the object area picture set, the training model is trained by the loss function to obtain a video soundtrack generation model. When applied, the video is input into the video soundtrack generation model to obtain corresponding music.
2. The method for generating controllable video soundtrack based on a cross-modal alignment mechanism according to claim 1, characterized in that: Feature extraction is performed on each object region image to obtain semantic embedding vectors, area features, starting position features, color features, and motion vectors, including: The pre-trained ViT model is used to extract region-level semantic representations for each object region image, and then the semantic embedding vector is obtained through the self-attention mechanism; The ratio of the number of pixels in the object area to the total number of pixels in the corresponding first frame image is used as the area feature; Normalize the position data of the object area in the first frame image to obtain the starting position feature; The 256-dimensional color histograms of the three RGB channels of the object area are spliced to obtain the color features; The difference between the position feature of the object area in the current frame image and the position feature of the first frame image is used as the motion vector.
3. The method for generating controllable video soundtrack based on a cross-modal alignment mechanism according to claim 1 or 2, characterized in that: Each shot, starting position feature, and color feature are encoded separately to obtain multiple shot codes, starting position codes, and color codes, including: Fourier frequency method is used to encode each shot to obtain multiple shot codes; The starting position feature is encoded by the sine and cosine functions, wherein the first half dimension encodes the horizontal coordinate of the starting position feature and the second half dimension encodes the vertical coordinate of the starting position feature, so as to form two-dimensional data, thereby obtaining the starting position code; Color coding is achieved by encoding each pixel based on the color intensity of each pixel in the object area.
4. The method for generating controllable video soundtrack based on a cross-modal alignment mechanism according to claim 3, characterized in that: The starting position feature is encoded to obtain the starting position encoding, where the j-th object area is in the shot i, and the values of the dim, dim+1 dimensions are for: ,in, and are the horizontal and vertical coordinates of the starting position feature of the j-th object area in the lens i, is the scaling factor, is the total dimension.
5. The method for generating controllable video soundtrack based on a cross-modal alignment mechanism according to claim 3, characterized in that: The color features are encoded to obtain the color codes of the red, green and blue channels, and the color codes of the red, green and blue channels are spliced to obtain the final color code, where the encoding value of the k-th pixel in the j-th object area in the i-th shot under the color channel C is for: ,in, is the total dimension, C is the color channel, is the number of pixels in the current region, in color channel C, whose color value is equal to k.
6. The method for generating controllable video soundtrack based on a cross-modal alignment mechanism according to claim 1, characterized in that: The second fusion feature and the motion vector are fused through a block matrix to obtain dynamic video features, including: The motion features are mapped to angle offsets in the form of trigonometric functions, and the starting position encoding is linearly transformed based on the angle offset to obtain the dynamic video features of each object area at time step t.
7. The method for generating controllable video soundtrack based on a cross-modal alignment mechanism according to claim 1, characterized in that: Before inputting the dynamic video features into the non-autoregressive network, the dimensions are adjusted and aligned through multi-layer linear transformations, and then the contextual information of the dynamic video features in the temporal dimension is captured through a multi-head self-attention mechanism.
8. The method for generating controllable video soundtrack based on a cross-modal alignment mechanism according to claim 1, characterized in that: The dynamic video features are sequentially passed through a non-autoregressive network to obtain a predicted music feature sequence, and the predicted music feature sequence is upsampled through a one-dimensional transposed convolution to obtain a final predicted music feature sequence, wherein the predicted music feature sequence includes pitch, loudness, chroma and timbre feature sequences.
9. The method for generating controllable video soundtrack based on a cross-modal alignment mechanism according to claim 1, characterized in that: Multiple predicted music feature sequences are decoded through the transformer to obtain predicted music embedding vectors, including: Multiple predicted music feature sequences are sequentially subjected to feature concatenation and one-dimensional convolution. A multi-head self-attention mechanism is performed based on the result of the one-dimensional convolution and the corresponding real music embedding vector to obtain the predicted music embedding vector. The result of the one-dimensional convolution is used as the query, and the real music embedding vector is used as the key and value.
10. The method for generating controllable video soundtrack based on a cross-modal alignment mechanism according to claim 1, characterized in that: After obtaining the semantic embedding vector, area feature, starting position feature, color feature and motion vector, and after obtaining the predicted music feature sequence, a featureless guidance strategy is introduced respectively. The featureless guidance strategy is used to enable the user to control the weights of the semantic embedding vector, area feature, starting position feature, color feature, motion vector and predicted music feature sequence.
Citation Information
Patent Citations
Method and device for generating sports video games
CN112685592A
Cross-modal music generation method and device oriented to film and television duplication
CN118551074A
Intelligent MV generation method, system and device based on AIGC and medium
CN119788886A