Potential diffusion model-based cinema video score generation and style control method
By adopting a multi-level controllable generation method based on a potential diffusion model in the video to music generation task, combined with local and global control, the problem of low matching between music and movie themes in the existing technology is solved, efficient and controllable film soundtrack generation is achieved, and the generation performance is improved through evaluation indicators.
Patent Information
- Application Number
- CN202510328247.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art is difficult to achieve efficient and multi-conditional control of music generation in video to music generation tasks, especially in movie soundtracks, where the theme narrative and emotional tone of the music and the movie are low and there is a lack of effective evaluation indicators.
A multi-level controllable video-to-music generation model is adopted based on the latent diffusion model, combining local control (melody and dynamic control of music) and global control (semantic features, emotional features and aesthetic features of movie videos). The model generates flexible and style-controllable movie soundtracks through feature fusion and style control modules.
The efficient generation of the movie soundtrack is achieved, and the music can accurately reflect the emotional tension and narrative intention of the movie, enhance the drama and consistency with the movie content, and improve the model generation performance through evaluation indicators (originality and recognizability).
Smart Images

Figure CN120199208A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large model computing, and specifically to a method for generating and style controlling the background music of movie videos based on the latent diffusion model. Background Art
[0002] Diffusion models have achieved state-of-the-art sample quality in various tasks, such as image generation, image restoration, and video generation. For audio or music synthesis, diffusion models for Mel-spectrogram generation and waveform generation have been studied. However, a major problem with diffusion models is that the iterative generation process in high-dimensional data spaces leads to slow training and inference speeds. Among them, the latent diffusion model has gradually become a common way to solve this problem. It uses the diffusion model in a smaller latent space, which has promoted the transfer of image generation from Gan to Diffusion. For music generation, audio waveforms have redundant information, which increases the modeling complexity and reduces the inference speed. Therefore, Yang et al. used the latent diffusion model to generate embeddings compressed from Mel spectrograms. It is also the efficient and high-quality performance of the latent diffusion model that has gradually made the latent diffusion model a choice for music generation.
[0003] However, since the control of music by videos is often implicit and the latent diffusion model does not have the function of multi-condition control, it is often infeasible to directly transfer the above method to the video-to-music generation task.
[0004] Film scoring plays a key role in enriching the emotional landscape and theme exploration. It is an artistic expression that conveys emotions, atmosphere, and narrative depth to the audience through sound. Automating the film scoring production process through artificial intelligence research represents a major advancement in the film scoring production field in terms of cost efficiency and innovation, which is the combination of artificial intelligence and artistic creation. Currently, music generation has become a promising field in generative modeling, and cross-modal music generation, as a challenging task, has gradually come into people's view in recent years.
[0005] Currently, for the film scoring task, how to match the generated music with the theme narrative and emotional tone of the film remains a challenge because the content of movie videos is more complex and diverse, and previous visual features are not feasible. At the same time, for the field of controllable video music generation, there is currently no good paradigm. Finally, for emerging tasks such as video-generated music, many evaluation metrics are also lacking, and how to objectively evaluate the generation performance of the model well is also a problem. Summary of the Invention
[0006] To solve the above technical problems, the present invention provides a method for generating and controlling the style of film video soundtracks based on a latent diffusion model. The present invention simultaneously utilizes local control (the melody and dynamics control of music) and global control (the semantic, emotional, and aesthetic features of film video segments), enabling the generation of flexible and style-controllable film soundtracks.
[0007] The present invention is implemented through the following technical solutions:
[0008] A method for generating and controlling the style of film video soundtracks based on a latent diffusion model, comprising the following steps:
[0009] Step 1: Using the latent diffusion model, construct a basic video-to-music generation model;
[0010] Step 2: Extract semantic, aesthetic, and emotional features from the film video and perform feature fusion;
[0011] Step 3: Add a style control module to the basic video-to-music generation model;
[0012] Fix the weights of the basic video-to-music generation model, copy the structure and weights of the downsampling and intermediate sampling of the model, and add zero convolutional layers before and after as the style control module;
[0013] In the training stage, use the melody and dynamics features of the music as local control to enter the style control module, use the fused features of the semantic, aesthetic, and emotional features obtained in Step 2 as global control to send to the cross-attention layer, and send the compressed representation of the music mel spectrogram to the pre-trained basic video-to-music generation model.
[0014] In the above technical solution, in Step 1, the video-to-music generation model uses a probabilistic generative model to estimate the true conditional data distribution q(z0|c film ), where the model distribution is expressed as p θ (z0|c film ), where, represents a sample in the latent space compressed from the music mel spectrogram X∈R T×F , where r represents the compression level, C represents the channels of the latent representation, and T and F respectively represent the time and frequency of the mel spectrogram X.
[0015] In the above technical solution, in Step 1, the basic video-to-music generation model is a pre-trained model obtained by fine-tuning the AudioLDM model using the HIMV-200k dataset, and this operation migrates the latent diffusion model from the text-music space to the video-music space.
[0016] In the above technical solution, in step 2, for extracting semantic features: use the image encoder of the CLIP model to extract frame features, and then average the extracted frame features to capture the video semantics.
[0017] In the above technical solution, in step 2, for extracting aesthetic features: use a pre-trained aesthetic feature extraction model to obtain the aesthetic features of the movie video. The aesthetic feature extraction model includes three parts: a visual attribute analysis module, a theme understanding module, and a two-level aesthetic reasoning module; it includes the following steps 2.21 - step 2.23:
[0018] Step 2.21: Obtain the visual attributes of the movie video through the visual attribute analysis module;
[0019] The visual attribute analysis module is used to learn visual attributes to obtain aesthetics. It uses the ResNet50 architecture, deletes the fully connected layer, and allows sharing of feature extraction across branches; and uses six multi-layer perceptrons (MLPs) to further map the shared features to six visual attributes, which are: interesting content of the image, object emphasis, vivid color, depth of field, color harmony, and good lighting;
[0020] Step 2.22: Obtain the theme feature vector of the movie video through the theme understanding module;
[0021] The theme understanding module first uses the ResNet-50 backbone to predict the theme category of the image; then, uses a multi-layer perceptron MLP with PReLU activation function to map the input image to the predicted theme category, where the last fully connected layer generates the theme features representing the theme category; finally, performs a softmax non-linear operation to generate the predicted theme probability;
[0022] Step 2.23: Input the visual attribute information obtained in step 2.21 and the theme feature vector information obtained in step 2.22 into the two-level aesthetic reasoning module to obtain the aesthetic features;
[0023] First, take the theme feature as the central node, and all attribute nodes are connected to the central node of the theme feature; then, obtain the node features through a graph convolutional network (GCN), and these node features serve as theme-aware visual attribute nodes;
[0024] Then extract the general aesthetic features of the image; take the general aesthetic features as the central node, and connect all theme-aware visual attribute nodes to this central node; by using the attribute aesthetic graph, update the node features, integrate the theme-aware visual attributes and the general aesthetic features, so as to obtain the final aesthetic features;
[0025] Finally, append a fully connected layer to map the updated node features to the overall aesthetic quality score, and then embed the overall aesthetic quality score to obtain the aesthetic features for guiding music generation.
[0026] In the above technical solution, in step 2, for extracting emotional features: use a pre-trained weakly supervised video emotion detection and prediction model (WECL) to predict the emotional features of the video.
[0027] In the above technical solution, let f (l) (x (m,l-1) +C style ,m,c film ) represent the l-th block of the style control module, where m is the diffusion time step, x (m,l-1) contains the latent noise features after l-1 blocks, c film is the global control condition, and C style is the local control condition. Then, the data flow of the style control module is as follows:
[0028]
[0029] Among them, Z out is the zero convolution layer, and f l is initialized from the l-th encoder block of the pre-trained base video-to-music generation model;
[0030] For the global control c film , the Q (Query), K (Key), and V (Value) in the cross-attention layer are expressed as:
[0031] Q = W q (z t ), K = W k (c film ), V = W v (c film ); where z t represents the input noise features, and W q , W k and W v are projection matrices;
[0032] For the local control C style , first use two convolutional blocks and a newly added zero convolution layer to align the melody and dynamic features of the music with the resolution of the input noise features; then, integrate them with the normalized input noise features through the shortcuts in the following sequence:
[0033] C = norm(z t ) + Z in (conv mel (c mel )) + Zin (conv dyn (c dyn )); where norm(z t ) represents the normalized input noise feature; z in is the new zero convolution layer; c mel represents the melody; c dyn represents the dynamics.
[0034] In the above technical solution, for the melody c mel , a window size of 260 and a hop size of 160 are used to adjust the chromagram, and the energy from F frequency bins is condensed into 12 pitch classes. Then, the most dominant pitch class is selected for enhancement by maximizing the average at each time step.
[0035] In the above technical solution, for the dynamics c dyn , the loudness is obtained by summing the energy of the frequency bins in each time frame of the linear spectrogram and converting the value to decibels.
[0036] In the above technical solution, the method for evaluating the originality and recognizability of the generated film score is as follows:
[0037] For the evaluation of originality, its calculation formula is:
[0038]
[0039] where n c represents the number of music categories participating in the evaluation, represents the features extracted from the pre-trained model for the mel spectrogram , and N represents the number of samples extracted for each category, which is used to calculate the originality score of the model;
[0040] For the evaluation of recognizability, a one-shot classification model designed for music styles is used to measure the classification accuracy of each generated sample, and the average accuracy of the entire test set is used as the recognizability metric.
[0041] The advantages and beneficial effects of the present invention are:
[0042] The multi-level controllable generation system proposed by the present invention, including local control (such as melody, dynamics) and global control (visual features, emotional features, aesthetic features), injects richer expressiveness into the generation of film scores. This control system enables the generated music to accurately reflect the emotional tension and narrative intention of the film. For example, specific emotional control parameters can adjust the emotional color of the music, making the music show a low atmosphere in sad scenes and convey a sense of tension in intense fighting scenes. Such expressiveness enhances the drama of the score and makes it highly consistent with the film content.
[0043] By utilizing local control and global control during the generation process, personalized music generation for different movie scenes can be achieved. The existence of local control enables the generation model to finely adjust music details, such as customizing styles in aspects like melody and dynamics, so as to adapt to different visual dynamics and picture dynamics. In addition, the present invention uses the compressed representation of the Mel spectrogram to transform high-dimensional music data into a low-dimensional feature space, which significantly reduces the computational complexity during the generation process. Compared with directly generating in the high-dimensional music waveform, the Mel spectrogram and the compressed representation can accelerate the training and generation speed of the model. This framework not only affects the mood of music through emotion feature control parameters, but also enables music to convey specific narrative intentions through the input of visual features. For example, at the climax of a movie, by enhancing the interaction between visual features (such as camera angles, color saturation) and emotion control, the generated music can further amplify the narrative tension at intense and emotionally heightened moments. Such an interactive generation mechanism makes music not only a simple background sound effect, but also an integral part of movie narration, helping the audience to better understand the plot and the emotional development of the characters in the movie. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a flowchart of the present invention.
[0045] For those of ordinary skill in the art, other related drawings can be obtained based on the above drawings without creative efforts. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solution of the present invention will be further described below in conjunction with specific embodiments.
[0047] The present invention provides a method for movie score generation and style control based on a latent diffusion model, which allows for the simultaneous utilization of local control (melody and dynamics control of music) and global control (semantic features, emotion features, and aesthetic features of movie video clips), enabling the generation of flexible and style-controllable movie scores. In this task, given the global control c film and the local control C style , the goal is to learn the conditional generation model p θ (z0|c film , C style ) in the compressed representation z of the music Mel spectrogram X.
[0048] Specifically, the method for movie score generation and style control based on a latent diffusion model proposed by the present invention includes the following steps:
[0049] Step 1: Using the Latent Diffusion Model (LDM), a basic video-to-music generation model was pre-trained.
[0050] Use a probabilistic generative model to estimate the true conditional data distribution q(z0|c film ), where the model distribution is denoted as p θ (z0|c film ), where, denotes a sample in the latent space compressed from the music mel spectrogram X ∈ R T×F , where r represents the compression level, C represents the channels of the latent representation, and T and F represent the time and frequency of the mel spectrogram X respectively. In this embodiment, the basic video-to-music generation model is a pre-trained model obtained by training the AudioLDM model using the HIMV-200k dataset. This operation migrates the latent diffusion model from the text-music space (i.e., text-to-music generation) to the video-music space (i.e., video-to-music generation).
[0051] Step 2: Extract semantic features, aesthetic features, and emotional features from the movie video, and perform feature fusion on the semantic features, aesthetic features, and emotional features.
[0052] Step 2.1: Extract semantic features.
[0053] Use the image encoder of the CLIP model (Contrastive Language-Image Pretraining) to export semantic features, and average the extracted frame features to capture video semantics. CLIP is a pre-trained model for language-image tasks, consisting of a text encoder f text (·) and an image encoder f image (·). On a large-scale image-text dataset, this algorithm is trained using the method of contrastive learning to enhance the similarity between images and texts. Therefore, CLIP can generate semantic priors not only for images but also for videos, which can be used to facilitate music generation.
[0054] To generate a video semantic prior, first use f image (·) to extract 512-dimensional features from video frames at a frame rate of 10 frames per second; then, average the feature sequence along the time dimension to obtain the visual semantic feature c s ∈ R 1×512 , c s contains different semantic representations in the video, such as scenes, people, objects, etc., which can guide music generation.
[0055] Step 2.2: Extract aesthetic features.
[0056] Use a pre-trained aesthetic feature extraction model to obtain the aesthetic features in the movie video. The aesthetic feature extraction model includes three key parts: a visual attribute analysis module, a theme understanding module, and a two-level aesthetic reasoning module.
[0057] Step 2.21: Obtain the visual attributes of the movie video through the visual attribute analysis module.
[0058] The visual attribute analysis module is used to learn visual attributes to obtain aesthetic feeling. It uses the ResNet50 architecture, deletes the fully connected layer, and allows sharing of feature extraction across branches; and also adopts six multi-layer perceptrons (MLPs) to further map the shared features to six visual attributes (interesting content of the image, object emphasis, vivid color, depth of field, color harmony, and good lighting). The specific processing process is as follows:
[0059] First, extract 10 frames of images from the movie clip at a frame rate of 1 frame per second, denoted as a i (i = 1, 2,..., 10). For each input image a i , obtain the hidden features from the shared feature extraction network F a (·) as follows: As shown below:
[0060]
[0061] Then, use six multi-layer perceptrons (MLPs) to construct six attribute branches to further map the hidden features of each image to the visual attributes which are defined as follows:
[0062]
[0063] where m = 1, 2,..., 6 represents 6 MLPs, and represents the six predicted visual attributes. Take the average of the 10 frames to obtain 6 different attributes (the average of different video frames) as the visual attributes of the entire video
[0064] Step 2.22: Obtain the theme feature vector of the movie video through the theme understanding module.
[0065] The theme understanding module, first, it uses the ResNet-50 backbone to predict the theme category of the image; then, adopts a multi-layer perceptron (MLP) with PReLU activation function to map the input image to the predicted theme category, where the last fully connected layer generates the theme features representing the theme category; finally, performs a softmax non-linear operation to generate the predicted theme probability.
[0066] Input 10 frames of images a of the movie video i (i = 1, 2, ..., 10) into the topic understanding module to extract the topic features of each frame of the image Defined as:
[0067]
[0068] where F t represents using the ResNet-50 backbone to predict the topic category of the image; MLP represents the above-mentioned multi-layer perceptron with PReLU activation function.
[0069] Then, average the topic features of these 10 frames to obtain the topic feature vector of the movie video.
[0070] Step 2.23: Input the visual attribute information obtained in Step 2.21 and the topic feature vector information obtained in Step 2.22 into the two-level aesthetic reasoning module to obtain aesthetic features.
[0071] First, consider the relationship between the image topic and visual attributes. Taking the topic feature as the central node, all attribute nodes are connected to the central node of the topic feature; then, obtain the node features through the graph convolutional network (GCN), and this node feature serves as the topic-aware visual attribute node.
[0072] Then extract the general aesthetic features of the image. Here, use the Swin Transformer model to extract general aesthetic features. Swin Transformer is a vision model based on Transformer, which can efficiently extract image features through hierarchical feature mapping and window shifting mechanism. In the aesthetic feature extraction task, Swin Transformer can capture the global semantic information and local detail features of the image.
[0073] Then, based on the relationship between the topic-aware visual attributes and general aesthetics, taking the general aesthetic features as the central node, connect all topic-aware visual attribute nodes to this central node; by using the attribute-aesthetic graph, update the node features, integrate the topic-aware visual attributes and general aesthetic features, so as to obtain the final aesthetic features.
[0074] Finally, append a fully connected layer to map the updated node features to the overall aesthetic quality score, and then embed the overall aesthetic quality score to obtain the aesthetic feature c for guiding music generation a ∈R 1×512 .
[0075] Step 2.3: Extract emotional features.
[0076] Use the pre-trained weakly supervised video emotion detection and prediction model (WECL) to predict the emotion features in the video. The WECL model comes from "Weakly Supervised Video Emotion Detection and Prediction via Cross-Modal Temporal Erasing Network". Embed the one-hot classification label E into the emotion feature c through linear projection e ∈R 1×512 as follows:
[0077] c e = Embed(E).
[0078] Step 2.4: Perform feature fusion on the obtained semantic features, aesthetic features, and emotion features.
[0079] For the obtained semantic feature c s , aesthetic feature c a , and emotion feature c e , if using typical feature vector concatenation, the curse of dimensionality and lack of interaction between features will be faced. Therefore, preferably, the present invention adopts a simplified feature fusion block called lightweight attention feature fusion (LAFF). In this method, the combine harvester combines the semantic features, emotion features, and aesthetic features of the movie video, and uses convex combination in a specific LAFF block, where the learning of the fusion weights is to optimize the generation of movie music.
[0080] Step 3: Add a style control module to the base video-to-music generation model obtained in Step 1.
[0081] Fix the weights of the base video-to-music generation model, copy the structure and weights of the downsampling and intermediate sampling of the model, and add zero convolutional layers before and after as the style control module, which together with the pre-trained video-to-music model constitutes the ControlNet architecture.
[0082] In the training stage, take the melody and dynamics of the music as local control and enter the style control module, take the fusion features of the semantic features, aesthetic features, and emotion features obtained in Step 2 as global control and send them to the cross-attention layer, and send the compressed representation of the music mel spectrogram to the pre-trained base video-to-music generation model.
[0083] Let f (l) (x (m,l-1) + C style , m, c film ) represent the l-th block of the style control module, where m is the diffusion time step, x(m,l-1) The noise latent feature after l-1 blocks, c film is the global control, C style is the local control. Then the data flow of the style control module is as follows:
[0084]
[0085] where, Z out is the zero convolution layer, and f l is initialized from the l-th encoder block of the pre-trained base video-to-music generation model.
[0086] For the global control c film , cross-attention is used to perceive the fused feature for the modeling of the global control model. The Q (Query), K (Key), and V (Value) in the cross-attention layer are expressed as:
[0087] Q = W q (z t ), K = W k (c film ), V = W v (c film );
[0088] where, z t represents the input noise feature, W q , W k and W v are projection matrices.
[0089] For the local control C style , first, two convolutional blocks and a newly added zero convolution layer are used to align the resolution of the melody c mel and the dynamics c dyn features of the music with the input noise feature; then, they are integrated with the normalized input noise feature through the shortcuts in the following sequence:
[0090] C = norm(z t ) + Z in (conv mel (c mel )) + Z in (conv dyn (c dyn ));
[0091] where, norm(z t ) represents the normalized input noise feature; z in is the new zero convolution layer; conv mel represents the convolutional block for processing the music melody feature; conv dyn represents the convolutional block for processing the music dynamics feature.
[0092] For melody c mel , the chromagram is adjusted using a window size of 260 and a hop size of 160, the energy from F frequency bins is condensed into 12 pitch classes, and then it is enhanced by selecting the dominant pitch class by average maximization at each time step.
[0093] For dynamics c dyn , the loudness is obtained by summing the energy of the frequency bins in each time frame of the linear spectrogram and converting the value to decibels, which is closely related to human perception of loudness.
[0094] Furthermore, to reduce the rapid fluctuations caused by note onsets or strikes and to make the dynamics control consistent with the perceived musical intensity, the values are smoothed by using a second-order context window (i.e., Savitzky-Golay filter).
[0095] Step 4: Evaluate the performance of the model in terms of originality and recognizability
[0096] The originality of a film score is a key metric, requiring the creation of works that are unique compared to previous background music. This principle emphasizes the importance of innovation in the composition process, ensuring that each score provides a unique auditory experience and enhances the narrative and emotional depth of the film.
[0097] The recognizability of a film score refers to its ability to be discerned by the audience as embodying a specific musical style, thus helping to identify the style lineage or intent of the score. This attribute is crucial when evaluating the effectiveness of the score in evoking the intended emotional response or thematic associations.
[0098] For the evaluation of originality in the present invention, the calculation formula is:
[0099]
[0100] where n c represents the number of music categories (by music composers) participating in the evaluation, represents the features extracted from the pre-trained model for the mel spectrogram , and N represents the number of samples drawn in each category, used to calculate the originality score of the model.
[0101] For the evaluation of recognizability in the present invention, a one-shot classification model designed for music styles is used to measure the classification accuracy of each generated sample, and the average accuracy of the entire test set is used as the recognizability metric.
[0102] The above has made an exemplary description of the present invention. It should be noted that without departing from the core of the present invention, any simple deformation, modification, or equivalent substitution that can be made by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.
Claims
1. A method for generating music and controlling style of movie videos based on a latent diffusion model, characterized in that: The following steps are involved: Step 1: Using the latent diffusion model, a basic video-to-music generation model is pre-trained; Step 2: Extract semantic features, aesthetic features, and emotional features from movie videos and perform feature fusion; Step 3: Add a style control module to the basic video-to-music generation model; The weights of the pre-trained video-to-music generation model were fixed, the structure and weights of the down-sampling and mid-sampling of the model were copied, and zero convolution layers were added before and after as a style control module; During the training phase, the melody and dynamic features of the music are taken as local controls and fed into the style control module. The fusion features of the semantic, aesthetic and emotional features obtained in step 2 are taken as global controls and fed into the cross-attention layer. The compressed representation of the music Mel spectrum is then fed into the pre-trained basic video to music generation model.
2. The method for generating music and controlling style of movie videos based on a latent diffusion model according to claim 1 is characterized in that: In step 1, the probability generation model is used to estimate the true conditional data distribution q(z0|c film ), where the model distribution is represented by p θ (z0|c film ),in, Represented by the music Mel spectrogram X∈R T×F A sample in the compressed latent space, where r represents the compression level, C represents the channel of the latent representation, and T and F represent the time and frequency of the mel-spectrogram X, respectively.
3. The method for generating music and controlling style of movie videos based on a latent diffusion model according to claim 1 is characterized in that: In step 1, the basic video-to-music generation model is a pre-trained model obtained by fine-tuning the AudioLDM model using the HIMV-200k dataset. This operation transfers the latent diffusion model from the text-music space to the video-music space.
4. The method for generating music and controlling style of movie videos based on a latent diffusion model according to claim 1, characterized in that: In step 2, for extracting semantic features: the image encoder of the CLIP model is used to propose frame features, and the extracted frame features are averaged to capture the video semantics.
5. The method for generating music and controlling style of movie videos based on a latent diffusion model according to claim 1 is characterized in that: In step 2, for extracting aesthetic features: Aesthetic features in movie videos are obtained using a pre-trained aesthetic feature extraction model, which consists of three parts: a visual attribute analysis module, a topic understanding module, and a two-level aesthetic reasoning module. The following steps include steps 2.21 to 2.23: Step 2.21: Obtain visual attributes of the movie video through a visual attribute analysis module; The visual attribute analysis module is used to learn visual attributes to obtain aesthetics. It uses the ResNet50 architecture and removes the fully connected layer to allow shared feature extraction across branches. It also uses six multi-layer perceptrons (MLPs) to further map the shared features to six visual attributes: interesting content of the image, object emphasis, vivid colors, depth of field, color coordination, and good lighting. Step 2.22: Obtain the theme feature vector of the movie video through the theme understanding module; The topic understanding module first uses the ResNet-50 backbone to predict the topic category of the image; then, a multi-layer perceptron MLP with PReLU activation function is used to map the input image to the predicted topic category, where the last fully connected layer generates topic features representing the topic category; finally, a softmax nonlinear operation is performed to generate the predicted topic probability; Step 2.23: Input the visual attribute information obtained in step 2.21 and the theme feature vector information obtained in step 2.22 into the two-level aesthetic reasoning module to obtain aesthetic features; First, the topic feature is used as the central node, and all attribute nodes are connected to the central node of the topic feature. Then, the node feature is obtained through the graph convolutional network (GCN), and the node feature is used as the topic-aware visual attribute node. Then the general aesthetic features of the image are extracted; The general aesthetic feature is used as the central node, and all the theme-aware visual attribute nodes are connected to the central node; by utilizing the attribute aesthetic graph, the node features are updated, and the theme-aware visual attributes and the general aesthetic feature are integrated to obtain the final aesthetic feature; Finally, a fully connected layer is attached to map the updated node features to the overall aesthetic quality score, which is then embedded into the overall aesthetic quality score to obtain the aesthetic features that guide music generation.
6. The method for generating music and controlling style of movie videos based on a latent diffusion model according to claim 1, characterized in that: In step 2, for extracting emotional features: use the pre-trained weakly supervised video emotion detection and prediction model to predict the emotional features in the video.
7. The method for generating music and controlling style of movie videos based on a latent diffusion model according to claim 1, characterized in that: Let f (l) (x (m,l-1) +C style ,m,c film ) represents the lth block of the style control module, where m is the diffusion time step, x (m,l-1) Contains the noise latent features after l-1 blocks, c film is the global control, C style If it is local control, the data flow of the style control module is as follows: Among them, Z out is a zero convolutional layer, and f l Initialize the lth encoder block from the pre-trained base video to music generation model; For global control c film , Q(Query), K(Key) and V(Value) in the cross attention layer are expressed as: Q=W q (z t ),K=W k (c film ),V=W v (c film ), where z t represents the input noise characteristics, W q , W k and W v is the projection matrix; For local control C style , we first align the melodic and dynamic features of the music with the resolution of the input noise features using two convolutional blocks and a newly added zero-convolutional layer; then, we integrate them with the normalized input noise features via a shortcut in the following sequence: C = norm(z t )+Z in (conv mel (c mel ))+Z in (conv dyn (c dyn )); where norm(z t ) represents the normalized input noise characteristic; z in is the new zero convolutional layer; c mel Indicates melody; c dyn Indicates dynamics.
8. The method for generating music and controlling style of movie videos based on a latent diffusion model according to claim 1 is characterized in that: For melody c mel ,The chromaticity map is adjusted using a window size of 260 and a hop size of 160, the energy from F frequency bins is condensed into 12 pitch classes, and then,enhancement is performed by selecting the most dominant pitch class at each time step by maximizing the average.
9. The method for generating music and controlling style of movie videos based on a latent diffusion model according to claim 1, characterized in that: For dynamic c dyn , loudness is obtained by summing the energy of the frequency bins in each time frame of the linear spectrogram and converting the values to decibels.
10. The method for generating music and controlling style of movie videos based on a latent diffusion model according to claim 1, characterized in that: The originality and recognizability of the generated movie soundtracks are evaluated as follows: The calculation formula for the evaluation of originality is: Among them, n c It represents the number of music categories involved in the evaluation. Represents the mel-spectrogram extracted from the pre-trained model The features of , N represents the number of samples extracted in each category, which is used to calculate the originality score of the model; For the evaluation of recognizability, a one-shot classification model designed for the music style is used, and the classification accuracy of each generated sample is measured. The average accuracy over the entire test set is taken as the recognizability measure.
Citation Information
Cited By
Controllable video score generation method based on cross-modal alignment mechanism
CN120510556A
Film and television video full-automatic production system based on digital actors
CN120812371A
Movie music sound effect unified generation method based on large model
CN121438775A