A multi-modal driven mv intelligent synthesis and creation method and system
Patent Information
- Application Number
- CN202611026540.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-10
- Publication Date
- 2026-09-25
Smart Images

Figure CN122824955A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a multimodal driven intelligent MV synthesis and creation method and system. Background Technology
[0002] With the rapid development of the digital music and short video industries, the demand for music visualization content, especially music videos (MVs), continues to rise. Traditional MV production relies on professional teams to complete multiple stages such as shot design, editing, color grading, and special effects compositing. This approach has limitations such as long production cycles, high labor costs, and high creative barriers, making it difficult to meet the market demand for rapid visualization of massive amounts of music content. In recent years, generative AI technology has provided a new path for automated MV creation. However, existing automatic audio and video generation solutions generally suffer from insufficient accuracy in audio-visual matching, weak emotional expression, and limited timing control capabilities. They cannot simultaneously ensure accurate rhythm and nuanced emotional visualization, resulting in a significant gap between the artistic expression of the generated content and professional production standards, thus hindering the effective application of AI-powered intelligent creation in the MV field. Summary of the Invention
[0003] In order to ensure that the generated music video is both accurate in rhythm and subtle in emotional visualization, this application provides a multimodal driven intelligent synthesis and creation method and system for music videos.
[0004] Firstly, this application provides a multimodal-driven intelligent MV synthesis and creation method, employing the following technical solution: A multimodal-driven intelligent MV synthesis and creation method, comprising: The input audio is preprocessed and features are extracted to obtain a shared feature map; The shared feature map is subjected to rhythm structure analysis to output rhythm structured label blocks; and the shared feature map is subjected to sentiment curve analysis to output sentiment label blocks; the rhythm structured label blocks and the sentiment label blocks share the same time coordinate system; Based on the rhythm structured tag blocks, rhythm-gated shot sequence planning is performed, and a shot planning tensor is output. The shot planning tensor defines the start and end time positions of each shot. Furthermore, based on the emotion tag blocks, visual style parameters for emotion mapping are generated, and a visual style parameter sequence is output frame by frame. The visual style parameter sequence defines the visual rendering parameters when rendering each frame. The shot planning tensor and the visual style parameter sequence are encoded into rhythm condition vector and style condition vector, respectively, and then fused into a joint condition vector. This vector is then injected into the cross-attention layer of the diffusion model for dual-condition joint frame-level rendering, generating a video frame sequence aligned with the audio duration shot by shot.
[0005] By employing the aforementioned technical solution, and through dual-channel processing of audio—rhythmic structure analysis and emotional curve analysis—and ensuring that both share the same time coordinate system, strict alignment of rhythmic segmentation and emotional rendering in the time dimension is guaranteed. This means that the generated music video will not exhibit the common disjointedness of "the scene transition is on time but the emotion is off" or "the emotional atmosphere is right but the shot transition is out of sync," because the shot planning tensor defines "where to cut," and the visual style parameter sequence defines "what each frame looks like." Both are locked onto by the same set of audio feature maps on the time axis, fundamentally eliminating temporal misalignment. Furthermore, when the diffusion model generates images frame by frame through denoising, the cross-attention layer simultaneously receives conditional signals from both rhythm and style. This means that when generating each frame, the model must simultaneously satisfy the requirements of "which shot segment the current moment belongs to" and "what visual emotion should be presented at the current moment." Compared to using rhythm or style conditions alone, this dual-condition injection approach results in more natural transitions at shot boundaries, higher emotional saturation in emotionally charged sections, and frame-level precision in the overall visual narrative and musical flow. Furthermore, from audio input to video frame sequence output, there's no need for manual editing points, transition template definitions, or segment-by-segment filter parameter adjustments; the entire pipeline flows seamlessly based on shared feature maps. Rhythm structured tags and emotional tags serve as intermediate representations, preserving interpretability while providing high-quality semantic input for subsequent shot planning and style parameter generation. This ensures that the final music video maintains creative diversity while remaining firmly anchored to the original song's musical logic.
[0006] Optionally, the step of performing rhythmic structure parsing on the shared feature map and outputting rhythmic structured tag blocks specifically includes: The shared feature map is extracted into a high-dimensional rhythm feature sequence through a rhythm feature extraction operation based on a temporal convolutional network. The high-dimensional rhythm feature sequence is input into four parallel fully connected prediction heads, which output frame-level beat probability, frame-level bar boundary probability, frame-level accent intensity, and frame-level paragraph label, respectively; the values of the paragraph label correspond to the intro, verse, chorus, interlude, or outro. The frame-level beat probability is subjected to non-maximum suppression, and frame indices greater than a first threshold are selected and converted into a beat timestamp sequence; the frame-level bar boundary probability is selected with frame indices greater than a second threshold and converted into a bar boundary timestamp sequence; the frame-level accent intensity is segmented according to the beat timestamp sequence and the average is calculated to obtain the discrete accent intensity value at each beat point; and a frame-level segment label sequence is output. The rhythmic structured tag block is formed by combining the beat timestamp sequence, the measure boundary timestamp sequence, the discrete accent intensity value, and the paragraph tag sequence.
[0007] By employing the aforementioned technical solution, continuous frame-level predictions are transformed into discrete and usable descriptions of musical events. The beat timestamp sequence provides the rhythmic skeleton, the measure boundary timestamp sequence marks the breath points of musical phrases, and the discrete accent intensity values quantify the dynamic level of each beat. In addition, the frame-level segment tag sequence annotates the macroscopic structure of the music. This combination of four dimensions allows the rhythmic structure tag block to retain both frame-level precision and event-level semantic clarity. In actual music video generation, the shot planning module can directly determine cut points based on beat timestamps, the intensity and speed of cuts based on accent intensity, whether to insert transition effects based on measure boundaries, and the visual language style of the current segment based on segment tags. The entire decision-making chain is entirely driven by the structure of the music itself, requiring no manual intervention. This ensures both a high degree of automation in generation and a deep fit between the visual output and the musical structure.
[0008] Optionally, the step of parsing the shared feature map for sentiment curves and outputting sentiment tag blocks specifically includes: The shared feature map is extracted into a sentiment feature sequence through a sentiment feature extraction operation based on a temporal convolutional network; The emotional feature sequence is input into three parallel fully connected prediction heads, which output frame-level valence trajectory, frame-level arousal trajectory, and frame-level dominance trajectory, respectively. Temporal smoothing and interpolation upsampling are applied to the valence trajectory, the arousal trajectory, and the dominance trajectory, respectively, to obtain a continuous emotion trajectory aligned with the rendering frame rate; The first and second time derivatives of the continuous emotional trajectory are calculated respectively, and emotional extreme points, emotional turning points and emotional steady-state segments are extracted according to preset detection rules and marked as emotional key frames; The continuous emotional trajectory and the emotional keyframes are combined to form the emotional tag block.
[0009] By employing the aforementioned technical solution, emotional extreme points mark the moments when emotions reach their peak or trough, emotional turning points mark the instants when emotions undergo a directional change, and emotional steady-state segments identify relatively stable emotional intervals. These three types of keyframes collectively form the skeleton of the musical emotional narrative. For the subsequent visual style parameter generation module, emotional keyframes provide clear trigger signals. At emotional extreme points, peak adjustments to color saturation or maximization of effect intensity can be triggered; at emotional turning points, gradual color transitions or changes in composition can be triggered; and in emotional steady-state segments, stable visual style output can be maintained to avoid unnecessary parameter fluctuations. The entire emotional tagging block contains both a continuous emotional trajectory frame by frame and discrete emotional keyframes, enabling the generation of visual style parameter sequences to not only match the subtle changes in emotion but also to produce impactful visual responses at key emotional points. The resulting music video expresses both a macro-level narrative rhythm and micro-level tension in its emotional expression.
[0010] Optionally, the step of planning the shot sequence with rhythm gating based on the rhythm structured tag block specifically includes: The rhythm structured tag block is converted into a vector sequence, which contains three channels: the distance to the nearest beat point, the accent intensity of the beat point to which the current frame belongs, and the paragraph tag channel of the current frame. The vector sequence is input into a sequence prediction operation based on a transformer, and a two-dimensional tensor is output as the shot planning tensor; the first dimension of the two-dimensional tensor is the number of shots, and the second dimension includes the shot start frame index and the shot end frame index. Calculate the rhythm gating loss function to force the shot end time to be constrained to the vicinity of the beat point where the accent intensity is higher than the threshold during the training phase; Based on the changes in valence, arousal, and accent intensity between two adjacent shots, perform a transition type selection operation and output the transition type; Based on the shot planning tensor and the transition type, a complete shot sequence plan is generated.
[0011] By employing the aforementioned technical solutions, when there are drastic changes in valence and arousal between adjacent shots, strong transitions such as hard cuts or flash whites can be chosen to enhance emotional shifts; when emotional changes are gradual, soft transitions such as dissolves or fade-in / fade-out can be chosen to maintain visual continuity; and changes in accent intensity further refine the strength of transitions—strong transitions accompany rising accents, and weak transitions accompany falling accents. The final generated complete shot sequence plan not only precisely aligns with the music rhythm in the time dimension, but also maintains synchronization with the emotional fluctuations of the music in terms of transition language, thus achieving unity in the visual narrative of the music video across three levels: macro structure, beat precision, and emotional expression.
[0012] Optionally, the step of generating visual style parameters for emotion mapping based on the emotion tag blocks specifically includes: The effect value, arousal value, and dominance value corresponding to the current frame are combined into an emotional three-dimensional vector. Through a mapping operation composed of fully connected layers, the emotional three-dimensional vector is mapped into visual style parameters. The visual style parameters include hue shift, color saturation, illumination intensity, depth parameters, and motion blur intensity parameters. For each discrete time point under the rendering frame rate, the corresponding emotional 3D vector is input into the mapping operation, and the visual style parameters of the current frame are output. Based on the user-adjustable emotion mapping sensitivity parameter, the output of the mapping operation and the preset default parameter vector are weighted and fused to obtain the final style parameter; A variation upper limit constraint is imposed on each component of the final style parameter between adjacent frames to prevent visual flicker.
[0013] By adopting the above technical solution, the emotion mapping and visual style parameter generation mechanism constructs a complete closed loop from quantitative emotion input to precise, controllable, and smooth visual parameter output. While ensuring the accuracy and subtlety of emotion expression, it also takes into account the flexibility of user interaction and the stability of image rendering. It can transform abstract emotion labels into intuitive and perceptible changes in image style in a natural, smooth, and adjustable manner, thereby enhancing the immersiveness, expressiveness, and user experience of emotional visual rendering.
[0014] Optionally, the steps of encoding the shot planning tensor and the visual style parameter sequence into rhythm condition vectors and style condition vectors respectively, fusing them into a joint condition vector, and then injecting it into the cross-attention layer of the diffusion model for dual-condition joint frame-level rendering specifically include: Each shot information in the shot planning tensor is encoded as a rhythm condition vector; the rhythm condition vector includes the sum of beat position embedding, accent intensity embedding, paragraph type embedding, and temporal position encoding; The visual style parameters of each frame are converted into a style condition vector through a fully connected layer; The rhythm condition vector and the style condition vector at the same time point are concatenated and then linearly projected to obtain the joint condition vector; In the cross-attention block of the diffusion model U-shaped network, the joint conditional vector is injected as an additional key value sequence into the attention calculation, so that the diffusion model can simultaneously receive shot switching timing guidance and frame-by-frame visual style guidance during the denoising process.
[0015] By adopting the above technical solutions, the dual-condition joint rendering mechanism constructs a diffusion generation framework under the dual constraints of temporal structure and visual attributes. While ensuring the logical coherence and natural rhythm of the video sequence shots, it achieves high-precision and controllable generation of frame-by-frame visual style, improves the temporal rationality, stylistic expressiveness and picture integration of AI-generated videos, and can output high-quality video frame sequences with regular rhythm, unified style and high fit to the preset plan, effectively enhancing the overall viewing experience and controllability of the generated videos.
[0016] Optionally, the rhythm gating loss function is calculated as follows: Calculate the numerical distance between the end time of each shot and the nearest high-emphasis beat, and then multiply the average value by the normalization factor after taking the average value for the number of shots.
[0017] Optionally, in the step of performing dual-condition joint frame-level rendering, a temporal smoothing constraint loss is introduced when training the diffusion model. The temporal smoothing constraint loss is constructed as follows: When calculating the denoising loss of the diffusion model, penalty weights are assigned to the denoising prediction differences between adjacent frames, wherein the penalty weight of adjacent frames within the same shot is greater than the penalty weight of adjacent frames across shots.
[0018] Optionally, the steps following dual-condition joint frame-level rendering may also include: For the reasoning stage, it is generated frame by frame in sequence according to the shot; For each shot, first load the rhythm condition vector of the shot itself as a shared condition for all frames of the shot, then obtain the frame-level style condition vector frame by frame and inject it into the diffusion model for denoising and generation. Perform a transition blending operation on the first frame of the current shot and the last frame of the previous shot, according to the transition type. Once all shots are generated, the complete video frame sequence is stitched together and output.
[0019] By adopting the above technical solution, while fully inheriting the advantages of dual-condition joint rendering in rhythm control and style regulation, it achieves a collaborative design that combines shared rhythm conditions for each shot, injection of style constraints frame by frame, targeted transition mixing, and modular splicing output. This approach also takes into account the operational efficiency, resource consumption, control precision, and visual coherence of long video generation. It effectively solves the pain points of low efficiency, easy temporal deviation, and abrupt shot transitions when generating long videos using diffusion models. It can efficiently and stably output high-quality video content with regular rhythm, delicate style, natural transitions, and complete coherence.
[0020] Secondly, this application provides a multimodal driven intelligent MV synthesis and creation system, which adopts the following technical solution: A multimodal driven intelligent MV synthesis and creation system, comprising: The feature extraction module is used for preprocessing and feature extraction of the input audio to obtain a shared feature map; The rhythm parsing module is used to parse the rhythm structure of the shared feature map and output rhythm structured label blocks. The sentiment analysis module is used to analyze the sentiment curve of the shared feature map and output sentiment tag blocks; the rhythm structured tag blocks and the sentiment tag blocks share the same time coordinate system; The shot planning module is used to plan shot sequences with rhythm gating based on the rhythm structured tag block and output a shot planning tensor; the shot planning tensor defines the start and end time positions of each shot; The style generation module is used to generate visual style parameters for emotion mapping based on the emotion tag block, and outputs the visual style parameter sequence frame by frame; the visual style parameter sequence defines the visual rendering parameters when rendering each frame. The video rendering module is used to encode the shot planning tensor and the visual style parameter sequence into rhythm condition vector and style condition vector respectively, and then fuse them into a joint condition vector. After that, it is injected into the cross attention layer of the diffusion model to perform dual-condition joint frame-level rendering and generate a video frame sequence aligned with the audio duration shot by shot.
[0021] In summary, this application includes at least one of the following beneficial technical effects: It achieves precise frame-level audio-visual coordination in both rhythmic structure and emotional expression. By sharing feature maps, it synchronously completes rhythm and emotion analysis and uses a unified time coordinate system, fundamentally eliminating the problem of temporal misalignment between shot transitions and emotional rendering. At the same time, it integrates the two types of conditions and injects them into the cross-attention layer of the diffusion model, so that the generation of the picture is subject to the dual constraints of shot timing planning and frame-by-frame visual style, improving the performance of MV picture and music in terms of rhythm fit and emotional matching, and achieving a deep unity of visual narrative and musical logic.
[0022] It has built an end-to-end automated, efficient and stable intelligent creation pipeline. The entire process does not require manual configuration of editing points, transition rules and filter parameters. It can automatically complete the entire process from feature analysis, shot planning, style mapping to video rendering based on the input audio. At the same time, it combines designs such as shot-by-shot sequential reasoning, shared rhythm conditions, transition mixing and time sequence smoothing constraints. While ensuring natural transitions and delicate and controllable style, it reduces the consumption of computing resources and the risk of time sequence drift in the generation of long videos, and takes into account creation efficiency, product quality and operation stability. Attached Figure Description
[0023] Figure 1 This is a first flowchart of an embodiment of the method of this application; Figure 2This is a second flowchart of an embodiment of the method of this application; Figure 3 This is a third flowchart of an embodiment of the method of this application; Figure 4 This is the fourth flowchart of an embodiment of the method of this application; Figure 5 This is the fifth flowchart of an embodiment of the method of this application; Figure 6 This is the sixth flowchart of an embodiment of the method of this application. Detailed Implementation
[0024] To make the purpose, technical solution, and advantages of this application clearer, the following description is provided in conjunction with the appendix. Figures 1-6 The present application will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the application.
[0025] The first embodiment of this application discloses a multimodal-driven intelligent MV synthesis and creation method. (Refer to...) Figure 1 The intelligent MV synthesis and creation method includes S110-S140: S110, preprocess and extract features from the input audio to obtain a shared feature map; S120 performs rhythm structure analysis on the shared feature map and outputs rhythm structured label blocks; and performs sentiment curve analysis on the shared feature map and outputs sentiment label blocks; the rhythm structured label blocks and sentiment label blocks share the same time coordinate system; S130 performs rhythm-gated shot sequence planning based on rhythm-structured tag blocks and outputs a shot planning tensor. The shot planning tensor defines the start and end time positions of each shot. It also generates visual style parameters for emotion mapping based on emotion tag blocks and outputs a visual style parameter sequence frame by frame. The visual style parameter sequence defines the visual rendering parameters when rendering each frame. S140 encodes the shot planning tensor and visual style parameter sequence into rhythm condition vector and style condition vector respectively, and then merges them into a joint condition vector. This vector is then injected into the cross-attention layer of the diffusion model to perform dual-condition joint frame-level rendering, generating a video frame sequence aligned with the audio duration shot by shot.
[0026] Specifically, for step S110, the input audio of any format can be uniformly resampled to 16kHz, mono, 16-bit linear PCM format. The signal-to-noise ratio of the high-frequency components of the audio is improved by using a pre-emphasis filter with a coefficient of 0.97. Then, the Hamming window function is used for frame processing. The frame length is set to 25 milliseconds, which corresponds to 400 sampling points, and the frame shift is set to 10 milliseconds, which corresponds to 160 sampling points. A 2048-point fast Fourier transform is performed on each frame to obtain the linear amplitude spectrum. Then, the linear spectrum is mapped to the Mel spectrum by a triangular filter bank with 80 Mel scale distributions. After logarithmic compression, a log-Mel feature map with dimensions [number of frames, 80] is obtained as a shared feature map. The first dimension is the time frame dimension and the second dimension is the Mel frequency band dimension, which is consistent with the input format of the subsequent temporal convolutional network. This shared feature map is also used as the input of the two extraction branches of rhythm features and emotion features, avoiding the repeated calculation of low-level features. All time dimensions are uniformly measured in seconds with the audio start point as the zero point.
[0027] Reference Figure 2 In S120, the steps of parsing the rhythmic structure of the shared feature map and outputting rhythmic structured label blocks specifically include S210-S240: S210 extracts the shared feature map into a high-dimensional rhythmic feature sequence through rhythmic feature extraction operation based on temporal convolutional network; S220 inputs the high-dimensional rhythm feature sequence into four parallel fully connected prediction heads, which output frame-level beat probability, frame-level bar boundary probability, frame-level accent intensity, and frame-level paragraph label respectively; the values of the paragraph label correspond to the intro, verse, chorus, interlude, or outro. S230: Perform non-maximum suppression on the frame-level beat probability and select frame indices greater than the first threshold, converting them into a beat timestamp sequence; select frame indices greater than the second threshold for the frame-level bar boundary probability and convert them into a bar boundary timestamp sequence; divide the frame-level accent intensity into segments according to the beat timestamp sequence and calculate the average to obtain the discrete accent intensity value at each beat point; and output the frame-level segment label sequence. S240 combines the beat timestamp sequence, measure boundary timestamp sequence, discrete accent intensity value, and paragraph label sequence into a rhythmic structured label block.
[0028] Specifically, for step S120, a stacked causal dilated temporal convolutional network can be used. The network consists of six stacked one-dimensional causal convolutional layers, with dilation coefficients increasing progressively from 1 to 32. The kernel size of each convolutional layer is set to 3, and the number of channels gradually increases from 80 dimensions in the input to 256 dimensions. After each convolutional layer, a weight normalization layer, a ReLU activation function, and a Dropout layer with a probability of 0.2 are connected. The structure of the causal convolution ensures that the feature calculation of each frame only depends on the information of the current frame and historical frames, without introducing future data. The dilated convolution expands the receptive field layer by layer. According to the receptive field calculation formula, it can eventually cover approximately 1.27 seconds, or 127 frames, of audio context, fully capturing the periodic patterns of beats and accents, and outputting a high-dimensional rhythmic feature sequence with dimensions of [number of frames, 256]. In another implementation, the receptive field can also be expanded to more than 6 seconds by increasing the number of network layers to 8 to 10 or adjusting the dilation coefficient sequence, so as to better capture the long-term musical structure patterns at the section level, such as the intro to the chorus.
[0029] Each of the four parallel fully connected prediction heads employs a two-layer fully connected structure. The hidden layer dimension is uniformly 128, and it uses the ReLU activation function. The output layer adapts different activation functions and dimensions according to the respective tasks: frame-level beat probability and frame-level measure boundary probability are both single outputs, normalized to the 0-1 range using the Sigmoid activation function, representing the probability that the current frame is a beat point and a measure boundary, respectively; frame-level accent intensity is also a single output, normalized to the 0-1 range using Sigmoid, with higher values indicating stronger audio impact in the frame; frame-level segment labeling is a 5-class classification task, and the output layer... The algorithm has 5 dimensions and outputs probability distributions for five categories—intro, verse, chorus, interlude, and outro—after Softmax activation. Four prediction heads share the high-dimensional rhythmic feature sequence of the input. During training, a multi-task joint loss optimization is employed. Binary cross-entropy loss is used for the beat and measure boundary tasks, mean squared error loss for the accent intensity task, and multi-class cross-entropy loss for the paragraph classification task. The total loss is obtained by weighting the losses in a ratio of 0.3:0.2:0.2:0.3, ensuring a balanced learning across multiple tasks. The weights of each task's loss can be dynamically adjusted based on the specific dataset distribution and task preferences.
[0030] The frame-level beat probability is first subjected to non-maximum suppression through a sliding window with a window length of 3 frames, retaining only the frame with the highest probability value in each window as a candidate beat point. Then, a first threshold (default is 0.5) is set to filter out the frame indices with a probability greater than the first threshold. The frame index is multiplied by a frame shift of 10 milliseconds to convert it into a beat timestamp sequence in seconds. The frame-level section boundary probability is set with a second threshold (default is 0.6), filtering out the frame indices that meet the conditions and converting them into timestamps. At the same time, it is verified by the beat timestamp sequence, retaining only the boundary candidates that fall on the beat point and eliminating invalid boundary predictions. The frame-level accent intensity is divided into segments based on the time interval between two adjacent beat points. The accent intensity values of all frames in each segment are summed and divided by the number of frames in the segment to obtain the discrete accent intensity value corresponding to each beat point, with the value range being 0 to 1. The paragraph label sequence is obtained by taking the category with the highest probability in the Softmax output of each frame and mapping it to the corresponding paragraph type label, finally obtaining a frame-level paragraph label sequence aligned with the original feature frames.
[0031] Finally, the obtained beat timestamp sequence, measure boundary timestamp sequence, discrete accent intensity value and paragraph label sequence are mapped onto the same time axis to construct a structured tensor format. Each beat point is associated with the corresponding accent intensity value and the paragraph label to which it belongs, each measure boundary is bound to the corresponding beat number, and the paragraph label sequence labels the paragraph attributes frame by frame. The whole forms a multi-level rhythmic structured label block containing beat layer, measure layer and paragraph layer. Its time base is completely aligned with the original audio and serves as the input basis for the subsequent shot planning module.
[0032] Reference Figure 3 In S120, the steps of parsing the sentiment curve of the shared feature map and outputting the sentiment label block specifically include S310-S350: S310 extracts the shared feature map into a sentiment feature sequence through a sentiment feature extraction operation based on a temporal convolutional network; S320 inputs the sentiment feature sequence into three parallel fully connected prediction heads, which output the frame-level valence trajectory, frame-level arousal trajectory, and frame-level dominance trajectory, respectively. S330 applies time smoothing and interpolation upsampling to the valence trajectory, arousal trajectory, and dominance trajectory respectively to obtain a continuous emotion trajectory aligned with the rendering frame rate; S340: Calculate the first and second time derivatives of the continuous emotional trajectory, and extract emotional extreme points, emotional turning points and emotional steady-state segments according to the preset detection rules, and mark them as emotional keyframes; S350, based on continuous emotional trajectories and emotional keyframes, is combined into emotional tag blocks.
[0033] Specifically, the emotional feature processing branch is shifted to a temporal convolutional network. The emotional feature extraction operation is also based on a temporal convolutional network. It adopts a 5-layer causal dilated temporal convolutional network with the same structure as the rhythm branch but with independent parameters. The dilation coefficients are 1, 2, 4, 8, and 16 respectively, the kernel size is 3, and the number of channels gradually increases from 80 dimensions to 128 dimensions. The input is also the shared feature map output by S110. This branch focuses on capturing features related to emotional expression in audio, such as timbre, melody, and harmonic changes. The output dimension is an emotional feature sequence of [number of frames, 128]. The design of sharing the underlying features with the rhythm branch reduces the amount of computation and ensures that the temporal dimensions of the two branches are naturally aligned.
[0034] Subsequently, three parallel fully connected prediction heads were used, each containing two fully connected layers with a hidden layer dimension of 64 and a ReLU activation function. The output layer was a single neuron with linear output, corresponding to the three sentiment dimensions of valence, arousal, and dominance. During the training phase, the output was normalized to the interval [-1, 1]. The valence dimension corresponds to the positive or negative tendency of the emotion, with positive values representing positive emotions and negative values representing negative emotions. The arousal dimension corresponds to the intensity of the emotion, with positive values representing excitement and negative values representing calmness. The dominance dimension corresponds to the strength of the sense of control over the emotion, with positive values representing a strong sense of control and negative values representing a sense of weakness. All three prediction heads were trained under supervised conditions using mean squared error loss. The mapping relationship between sentiment dimensions and audio features was fitted based on the publicly available VAD sentiment audio annotation dataset.
[0035] First, one-dimensional Gaussian smoothing is performed on the three frame-level trajectories of valence, arousal, and dominance. The smoothing kernel size and standard deviation can be set according to the audio frame rate (e.g., the smoothing kernel size is set to 5 frames and the standard deviation is 1.5) to eliminate high-frequency noise and jitter in frame-level prediction. Then, a cubic spline interpolation algorithm is used to resample and align the 100fps feature frame sequence to the target rendering frame rate. Taking the common 30fps as an example, each frame corresponds to a time interval of approximately 33.33 milliseconds. The resampling process strictly follows the timestamp for numerical mapping to ensure that each sampling point of the resampled continuous emotional trajectory corresponds one-to-one with the time point of the rendering frame. Finally, three continuous emotional trajectories with the same length as the total number of rendering frames are obtained.
[0036] Based on this, emotion keyframe extraction is performed. First, the first and second time derivatives of three continuous emotion trajectories are calculated by dividing the difference between adjacent frames by the frame interval. The first derivative reflects the rate of change of emotion, and the second derivative reflects the acceleration of the change of emotion. The detection rule for emotion extrema is as follows: the point where the first derivative of the continuous emotion trajectory changes from positive to negative is the emotion maximum point, and the point where it changes from negative to positive is the emotion minimum point. In addition, the difference between the emotion value of the extrema point and the mean emotion value of the preceding and following frames (e.g., 5 frames) is greater than the first emotion difference threshold (e.g., 0.1), thus eliminating false detections of small fluctuations. The detection rule for emotion turning points is the point where the second derivative of the continuous emotion trajectory crosses zero, and the corresponding position... When the absolute value of the first derivative is greater than the second emotional difference threshold (e.g., 0.05 per frame), it represents a node where the rate of emotional change changes changes. The detection rule for the steady-state segment is that the emotional value fluctuation range is less than 0.08 for more than 30 consecutive frames, and the absolute value of the first derivative is consistently lower than 0.02 per frame. The start and end frames of the steady-state segment are marked as keyframes. The above frame number threshold can be adaptively scaled according to the total audio duration and rendering frame rate. For example, 30 frames can be replaced with the larger value between 10% of the total number of frames and 30 frames to adapt to audio segments of different lengths. Finally, all extreme points, turning points, and steady-state segment endpoints are summarized and deduplicated to obtain a complete set of emotional keyframes, which serve as anchor points for subsequent visual style changes.
[0037] After completing the emotional feature processing, the continuous emotional trajectory data frame by frame and the frame index and type information of the emotional keyframes are integrated into a unified structured tag block. This tag block shares the same time coordinate system as the rhythm structured tag block, with the audio start point as the time zero point and seconds as the unit of measurement. This ensures that the rhythm information and emotional information can be accurately matched according to the timestamp in subsequent processing. At the same time, the emotional tag block containing continuous values and structured nodes can provide a constraint basis for the emotional dimension for visual style mapping and shot planning.
[0038] Reference Figure 4 In S130, the steps for planning the shot sequence with rhythm gating based on rhythm structured tag blocks specifically include S410-S450: S410 converts rhythm structured tag blocks into vector sequences, which contain three channels: the distance to the nearest beat point, the accent intensity of the beat point to which the current frame belongs, and the paragraph tag channel of the current frame. S420 takes a vector sequence as input to a sequence prediction operation based on a transformer and outputs a two-dimensional tensor as a shot planning tensor. The first dimension of the two-dimensional tensor is the number of shots, and the second dimension includes the shot start frame index and the shot end frame index. S430, calculates the rhythm gating loss function to force the shot end time to be constrained to the vicinity of the beat point where the accent intensity is higher than the threshold during the training phase; S440 performs a transition type selection operation based on the changes in valence, arousal, and accent intensity between two adjacent shots, and outputs the transition type. S450 generates a complete shot sequence plan based on the shot planning tensor and transition type.
[0039] The rhythm gating loss function is calculated as follows: Calculate the numerical distance between the end time of each shot and the nearest high-emphasis beat, and then multiply the average value by the normalization factor after taking the average value for the number of shots.
[0040] Specifically, for step S130, during the execution of shot sequence planning, the continuous emotional trajectory in the emotional tag block is simultaneously acquired; the 100fps beat timestamp, discrete accent intensity, and paragraph tags are aligned and mapped to the rendering frame rate (e.g., 30fps) according to the timestamp; for each frame at the rendering frame rate, the feature values of three channels are calculated and concatenated into a vector: the first channel is the distance to the nearest beat point, calculating the absolute value of the time difference between the current frame time and the nearest beat points before and after, and then using half of the average beat interval as a normalization factor to map the value to the interval between 0 and 1; the second channel is the accent intensity of the beat point to which the current frame belongs, that is, matching the beat interval where the current frame is located, taking the discrete accent intensity value of the starting beat point of the interval, with a value range of 0 to 1; the third channel is the paragraph tag of the current frame, using 5-dimensional one-hot encoding to correspond to five paragraph categories; finally, the vector dimension of each frame is 7-dimensional, and the entire sequence forms a vector sequence with dimension [total number of rendering frames, 7], which serves as the input to the subsequent transformer model.
[0041] The sequence prediction operation is implemented using a transformer model with an encoder-decoder structure. The encoder consists of four stacked multi-head attention layers with four attention heads and a hidden layer dimension of 128. After inputting the above 7-dimensional vector sequence, it outputs a global context encoding sequence. The decoder uses an autoregressive method to predict shot by shot, outputting the start frame index and end frame index of a shot at a time. The final output is a two-dimensional tensor of shape [number of shots, 2] as the shot planning tensor. The two values of the second dimension are the start frame index and end frame index of the shot, respectively. The end frame index of the previous shot is constrained to be equal to the start frame index of the next shot minus one, ensuring that the shot sequence is continuous and without gaps in time. The transformer's self-attention mechanism can capture the structural patterns of long-distance music segments and generate a basic shot segmentation scheme that conforms to the music narrative logic. This two-dimensional tensor only contains the core start and end time boundaries. The transition and segment information will be integrated through post-processing logic to generate a complete shot sequence plan.
[0042] The rhythm-gated loss function is used to constrain the end position of shots during the training phase. First, a preset accent intensity threshold of 0.7 is set, defining all discrete accent intensity values greater than 0.7 as high-accent beat points. If no high-accent beat point exists in the current audio segment, the beat point with the highest accent intensity is used as a substitute anchor point; if no beat point still exists, the distance error for that shot is set to zero. For each predicted shot, the timestamp corresponding to its end frame is taken, and the absolute value of the time difference between this timestamp and the nearest high-accent beat point that meets the conditions is calculated to obtain the distance error of a single shot. The distance errors of all shots are then... The average distance error is obtained by summing the results and dividing by the total number of shots. Then, it is multiplied by a normalization factor, which is the reciprocal of the average beat interval, to normalize the loss value to the order of 0 to 1. Finally, the rhythm gating loss and the sequence prediction loss of the transformer itself are weighted and summed at a ratio of 1:5, and this sum is used as the total loss for backpropagation training. The weight coefficients mentioned above can be adjusted according to the scale of the dataset and the requirements of the task, or they can be automatically determined by searching through the performance of the validation set. By penalizing the loss, the model is forced to tend to place the shot switching points on the beats with strong accents, thereby improving the fit between the shot switching and the music rhythm.
[0043] Based on this, a transition type selection operation is performed. First, the number of predicted shots is determined. If the number of shots is less than 2, the transition type selection operation is skipped, and no transition is performed by default, directly outputting a single-shot sequence. If the number of shots is greater than or equal to 2, the continuous emotional trajectory in the emotional tag block is obtained, and the feature change at the junction of two adjacent shots is calculated: valence change is the absolute value of the difference between the valence value of the first frame of the subsequent shot and the last frame of the preceding shot; arousal change is similarly calculated using the absolute value of the difference; accent intensity change is the absolute value of the difference between the accent intensity of the first beat of the subsequent shot and the last beat of the preceding shot. Simultaneously, it is determined whether the segment tags corresponding to the two shots have changed. Candidate transition types include four types: hard cut, fade in / out, dissolve, and push-pull transitions. The selection rule follows a priority order from high to low and depends on a preset feature change threshold: when the accent intensity changes... When the change in valence is greater than the stress abruptness threshold (e.g., 0.6) and the change in arousal is greater than the arousal abruptness threshold (e.g., 0.5), a hard cut is selected to adapt to scenes with a strong rhythmic impact. When the paragraph label changes and the change in stress intensity is in the medium to high range (e.g., 0.3 to 0.6), a push-pull transition is selected to adapt to scenes with paragraph progression. When the change in valence is greater than the valence gradation threshold (e.g., 0.4) and the change in arousal is lower than the arousal leveling threshold (e.g., 0.3), a dissolve is selected to adapt to scenes with a smooth emotional transition. For other smooth transition scenes, fade-in and fade-out are selected. In addition, if there is an emotional extreme point or emotional turning point near the shot transition and the change in stress intensity is greater than 0.4, the transition intensity is increased first, and a hard cut or push-pull transition is selected to strengthen the dual impact of emotion and rhythm. If multiple conditions are triggered at the same time, the highest priority is selected, and the transition type of each shot transition is finally output.
[0044] It should be noted that the step of calculating the rhythm gating loss function in S430 is only performed during the model training phase to optimize the parameters. During the inference phase, only S410, S420, S440, and S450 are performed to output the planning results.
[0045] After completing the transition selection, the start and end frame information of each shot in the shot planning tensor is integrated with the transition type at each connection point to form a structured shot sequence. Each shot unit includes shot number, start frame index, end frame index, duration, segment type, in-point transition type, and out-point transition type. At the same time, it is verified that the start and end frames of all shots are continuous and do not overlap, and that the total number of frames is completely consistent with the total number of rendered frames corresponding to the audio duration, forming a complete shot timing planning file, which serves as the timing basis for subsequent frame generation and transition processing.
[0046] Reference Figure 5 In S130, the step of generating visual style parameters for emotion mapping based on the emotion tag block specifically includes S510-S540: S510 combines the effect value, arousal value, and dominance value corresponding to the current frame into a three-dimensional emotion vector. Through a mapping operation consisting of fully connected layers, the three-dimensional emotion vector is mapped into visual style parameters. The visual style parameters include hue shift, color saturation, illumination intensity, depth parameters, and motion blur intensity parameters. S520, for each discrete time point under the rendering frame rate, inputs the corresponding three-dimensional emotion vector into the mapping operation and outputs the visual style parameters of the current frame; S530, based on the user-adjustable emotion mapping sensitivity parameter, performs a weighted fusion of the output of the mapping operation and the preset default parameter vector to obtain the final style parameter; S540 imposes a variation upper limit constraint on each component of the final style parameter between adjacent frames to prevent visual flicker.
[0047] Specifically, in the visual style mapping stage, the mapping operation from sentiment vector to visual style parameters is implemented by a multi-layer fully connected network. First, the valence, arousal, and dominance values of the current frame are concatenated into a three-dimensional sentiment vector as input. In a preferred embodiment, a three-layer structure is adopted: the first fully connected layer has a 3-dimensional input and a 32-dimensional output, using the ReLU activation function; the second layer has a 32-dimensional input and a 16-dimensional output, also using the ReLU activation function; and the third layer has a 16-dimensional input and a 5-dimensional output, corresponding to the five visual style parameters. The dimensions of each layer can be adjusted according to the fitting requirements. The hue shift parameter uses the tanh activation function, with an output range of [-π, π] radians. The offset angle corresponds to the color wheel; the four parameters of color saturation, light intensity, depth of field, and motion blur intensity all use the Sigmoid activation function, and the output is normalized to the range of [0, 1]. They can then be mapped to the actual rendering parameter range. For example, saturation is mapped to 0.5 to 1.5 times the base saturation, light intensity is mapped to 0.8 to 1.2 times the base light, depth of field is mapped to aperture values from f / 1.8 to f / 16, and motion blur intensity is mapped to shutter angles from 0 to 180 degrees. This mapping network is trained under supervision based on manually labeled sentiment-style pairing data to establish a non-linear mapping relationship from sentiment dimension to visual style.
[0048] Then, iterates through each discrete time point under the rendering frame rate, extracts the valence, arousal, and dominance values in the continuous emotional trajectory corresponding to that time point, forms a three-dimensional emotional vector, and inputs it into the fully connected mapping network mentioned above. After forward propagation, it calculates the original output values of the five visual style parameters corresponding to that frame, and completes the calculation of all frames in sequence to obtain the initial frame-by-frame visual style parameter sequence that corresponds one-to-one with the rendering frame, ensuring that the time of style change and emotional change is completely synchronized.
[0049] First, set a user-adjustable emotion mapping sensitivity parameter, with a value range of 0 to 1. 0 represents using the default parameters completely, with emotion not affecting the visual style, while 1 represents using the style parameters of the mapping output completely. The preset default parameter vector is a neutral style parameter with a hue shift of 0 radians and normalized values of saturation, light intensity, depth of field, and motion blur of 0.5. The fusion method is a weighted summation of components, that is, each component of the final style parameter is equal to the sensitivity parameter multiplied by the corresponding component of the mapping output, plus 1 minus the corresponding component of the sensitivity parameter multiplied by the default parameter. This mechanism allows users to flexibly adjust the degree of influence of emotion on visual style according to creative needs.
[0050] Next, to avoid visual flickering caused by abrupt changes in style parameters between adjacent frames, adaptive upper limit constraints on the five components of the final style parameters are applied in conjunction with the emotional keyframes. The default maximum change threshold for a single frame is set according to the physical magnitude of each parameter, for example: hue shift 0.05π radians, saturation 0.02, illumination intensity 0.015, depth of field parameter 0.02, and motion blur intensity 0.03. During processing, the system traverses frame by frame from the first frame. If the current frame is marked as an emotional extreme point or emotional turning point, the maximum change threshold of each component in that frame is multiplied by the set amplification. A coefficient (e.g., 1.5 times) allows for more significant visual style jumps at key emotional changes to match emotional outbursts. If the current frame is in a steady emotional state or a non-key frame, the default threshold is strictly enforced. If the absolute value of the difference between a component in the current frame and the previous frame exceeds the currently effective threshold, the component in the current frame is corrected to the component in the previous frame plus the threshold multiplied by the sign value of the change direction. Through adaptive frame-by-frame limiting based on emotional key frames, the changes in style parameters in non-key areas are smooth and natural, while allowing reasonable style jumps at key nodes. The final output is a smoothed sequence of frame-by-frame visual style parameters.
[0051] Reference Figure 6 S140, the steps of encoding the shot planning tensor and visual style parameter sequence into rhythm conditional vectors and style conditional vectors respectively, fusing them into a joint conditional vector, and then injecting it into the cross-attention layer of the diffusion model for dual-conditional joint frame-level rendering specifically include S610-S640: S610 encodes each shot information in the shot planning tensor into a rhythm condition vector; the rhythm condition vector includes the sum of beat position embedding, accent intensity embedding, paragraph type embedding, and temporal position encoding. S620 converts the visual style parameters of each frame into a style condition vector through a fully connected layer; S630: After concatenating the rhythm condition vector and style condition vector at the same time point, a joint condition vector is obtained by linear projection. In S640, the joint conditional vector is injected as an additional key-value sequence into the attention calculation in the cross-attention block of the diffusion model U-shaped network, so that the diffusion model can simultaneously receive shot switching timing guidance and frame-by-frame visual style guidance during the denoising process.
[0052] In the step of performing dual-condition joint frame-level rendering, a temporal smoothing constraint loss is introduced when training the diffusion model. The temporal smoothing constraint loss is constructed as follows: When calculating the denoising loss of the diffusion model, penalty weights are assigned to the denoising prediction differences between adjacent frames, wherein the penalty weight of adjacent frames within the same shot is greater than the penalty weight of adjacent frames across shots.
[0053] Specifically, in step S140, the diffusion model rendering stage begins. For each shot in the shot planning tensor, four types of rhythm information are extracted, encoded, and summed to obtain a rhythm condition vector: Beat position embedding calculates the relative position of the shot's starting frame within the current measure, normalizes it to the 0-1 range, and generates a 16-dimensional embedding vector through sinusoidal position encoding; accent intensity embedding takes the average accent intensity value of all beat points within the shot and maps it to a 16-dimensional embedding vector through a fully connected layer; segment type embedding maps the segment label to which the shot belongs to a 16-dimensional vector through a learnable embedding table; temporal position encoding uses sinusoidal position encoding to generate a 16-dimensional position code based on the shot's sequence number in the entire video; finally, the four vectors of the same dimension are summed element-wise to obtain the rhythm condition vector, with one vector corresponding to each shot, serving as the rhythm guidance condition shared by all frames within that shot.
[0054] The 5-dimensional visual style parameters of each frame are then input into a fully connected layer. This fully connected layer maps the low-dimensional style parameters to a high-dimensional feature space with the same dimension as the rhythm condition vector (e.g., the output dimension is 64), thus obtaining the style condition vector. The parameters of this fully connected layer are trained together with the diffusion model to ensure that the style condition vector is adapted to the feature space of the diffusion model.
[0055] For the rhythm condition vector and style condition vector at the same time point, they are first concatenated in the channel dimension. Then, through a biased linear projection layer, the concatenated vector is mapped to a target dimension (such as 768 dimensions) that is consistent with the key dimension of the cross-attention module of the diffusion model. This is the joint condition vector. Each rendering frame corresponds to a joint condition vector, which contains both shot rhythm structure information and frame-by-frame visual style information, serving as the conditional guidance input for the denoising process of the diffusion model.
[0056] A latent space diffusion model using the classic U-Net architecture is adopted. This U-Net includes four downsampling stages, one bottleneck layer, and four upsampling stages, with a base channel count of 320. In the 2nd, 3rd, and 4th downsampling stages and their corresponding upsampling stages, the channel count doubles to 640, 1280, and 1280 respectively. In the attention modules of the 2nd and 3rd downsampling and upsampling stages, a cross-attention layer is added. The joint conditional vector serves as the key-value sequence for the cross-attention, querying the image feature map from the current layer of the U-Net. Through attention calculation, the image features focus on the corresponding rhythm and style conditions during denoising, achieving dual-condition joint guidance. Injecting conditional vectors into intermediate layers balances computational cost and the strength of conditional guidance. During training, the original diffusion denoising loss (i.e., the mean square error between predicted and actual noise) is used. The temporal smoothing constraint loss is calculated as follows: For consecutive adjacent frames, the L1 distance between the noise maps predicted by the diffusion model is calculated. For adjacent frames within the same shot, this distance is multiplied by a weighting factor of 2.0; for adjacent frames across shots, it is multiplied by a weighting factor of 0.5. The weighted distances of all adjacent frame pairs are summed and then normalized by the feature dimension. The mean value is taken to obtain the temporal smoothing constraint loss. The denoising loss and the normalized temporal smoothing constraint loss are added together with a weighting of 7:3 to obtain the total training loss. These weighting factors can also be dynamically adjusted based on the error convergence during the training phase. This loss design ensures high continuity of the image content within the same shot, avoiding inter-frame jitter, while allowing for greater image changes across shots, aligning with the visual logic of shot transitions. The trained diffusion model can generate video frames frame-by-frame that are aligned with the audio rhythm and emotional style, and the total duration is perfectly aligned with the audio.
[0057] The steps following dual-condition joint frame-level rendering also include: During the inference phase, the video is generated frame by frame in sequence. For each shot, the rhythm condition vector of the shot itself is loaded first as a shared condition for all frames of the shot. Then, the frame-level style condition vector is obtained frame by frame and injected into the diffusion model for denoising generation. For the first frame of the current shot and the last frame of the previous shot, a transition blending operation is performed according to the transition type. After all shots are generated, the complete video frame sequence is spliced and output.
[0058] Specifically, the inference process is carried out sequentially according to the shot number from smallest to largest. When processing a single shot, the rhythm condition vector corresponding to that shot is loaded first. This vector remains unchanged in all frames of the entire shot as a shared condition. Then, each frame is processed sequentially from the start frame to the end frame of the shot. The style condition vector corresponding to the current frame is extracted and concatenated with the shared rhythm condition vector to obtain a joint condition vector. The joint condition vector is injected into the cross-attention layer of the diffusion model. The DDIM sampler is used to perform a set number of denoising sampling steps (e.g., 20 steps), starting from pure Gaussian noise and gradually denoising to generate the image of the frame. After all frames of a shot have been generated, the next shot is processed. The batch processing method according to shot can reuse the calculation results of the rhythm condition, improve inference efficiency, and at the same time ensure the consistency of the rhythm condition within the same shot.
[0059] After generating a single-shot frame, at the junction of two adjacent shots—that is, between the last frame of the previous shot and the first frame of the current shot—inter-frame blending is performed according to the preset transition type. To ensure strict alignment between the total video duration and audio, the transition processing uses the edge frames of the preceding and following shots as an overlap time window, without increasing the total number of frames: Hard cuts require no additional processing; simply stitch the two frames together. Fade-in / fade-out transitions use the last 5 frames of the previous shot and the first 5 frames of the current shot as an overlap window; the transparency of the previous shot linearly decreases from 1 to 0, while the transparency of the current shot linearly increases from 0 to 1. Dissolve transitions also use 5 frames before and after as an overlap window; the transparency of the two frames gradually changes linearly, with the transparency of the previous frame decreasing from 1 to... Simultaneously, the next frame increases from 0 to 1; push-pull transitions gradually enlarge and fade out the last frame of the preceding shot, and gradually shrink and fade in the first frame of the following shot, achieving a push-pull visual effect in conjunction with the center displacement of the image. The transition window also occupies 5 frames before and after; in practical applications, the number of overlapping window frames can be adaptively adjusted according to the shot length and transition style. For example, the window length can be set to the smaller value between half the shot length and 10 frames to avoid the transition area completely covering short shots; all transition effects are generated based on the pixel mixing of the keyframes before and after the overlapping window, ensuring a smooth and natural transition process. Since the transition frames occupy the original shot planning time, the total number of frames is completely consistent with the total number of rendering frames corresponding to the audio duration.
[0060] After all the frames of the shots have been generated and the transition frames at all the shot connections have been processed, all frame data are stitched together sequentially from the first frame to the last frame in chronological order. The total number of frames is verified to be consistent with the calculated value of audio duration multiplied by the rendering frame rate, ensuring that the audio and video durations are perfectly aligned. Finally, a continuous sequence of image frames is output, which can be further encoded into common video formats such as MP4. The frame rate of the frame sequence is consistent with the preset rendering frame rate. The content of the picture changes shots according to the rhythm and structure of the music, and the visual style is adjusted according to the emotional changes, realizing end-to-end audio-driven intelligent MV synthesis and creation.
[0061] Based on the above method embodiments, the second embodiment of this application discloses a multimodal-driven intelligent music video synthesis and creation system. The multimodal-driven intelligent music video synthesis and creation system of this application embodiment can implement any of the above-described multimodal-driven intelligent music video synthesis and creation methods, and the specific working process of each module in the multimodal-driven intelligent music video synthesis and creation system can be referred to the corresponding process in the above method embodiments.
[0062] For ease of understanding, an example is provided below: A multimodal driven intelligent MV synthesis and creation system includes: The feature extraction module is used for preprocessing and feature extraction of the input audio to obtain a shared feature map; The rhythm parsing module is used to parse the rhythm structure of the shared feature map and output rhythm structured label blocks; The sentiment analysis module is used to analyze the sentiment curve of the shared feature map and output sentiment label blocks; the rhythm structured label blocks and sentiment label blocks share the same time coordinate system; The shot planning module is used for rhythm-gated shot sequence planning based on rhythm-structured tag blocks, and outputs a shot planning tensor; the shot planning tensor defines the start and end time positions of each shot; The style generation module is used to generate visual style parameters based on emotion tag blocks for emotion mapping, and outputs the visual style parameter sequence frame by frame; the visual style parameter sequence defines the visual rendering parameters when rendering each frame. The video rendering module encodes the shot planning tensor and visual style parameter sequence into rhythm condition vectors and style condition vectors, respectively, and then merges them into a joint condition vector. This vector is then injected into the cross-attention layer of the diffusion model to perform dual-condition joint frame-level rendering, generating a video frame sequence aligned with the audio duration shot by shot.
[0063] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.
Claims
1. A multimodal-driven intelligent MV synthesis and creation method, characterized in that, include: The input audio is preprocessed and features are extracted to obtain a shared feature map; The shared feature map is parsed for rhythmic structure, and rhythmic structured label blocks are output. Furthermore, the shared feature map is analyzed for sentiment curves, and sentiment tag blocks are output. The rhythm structured tag block and the emotion tag block share the same time coordinate system; Based on the rhythm-structured tag block, perform rhythm-gated shot sequence planning and output a shot planning tensor. The shot planning tensor defines the start and end time positions of each shot; and based on the emotion tag block, visual style parameters for emotion mapping are generated, and a sequence of visual style parameters is output frame by frame; the sequence of visual style parameters defines the visual rendering parameters when rendering each frame. The shot planning tensor and the visual style parameter sequence are encoded into rhythm condition vector and style condition vector, respectively, and then fused into a joint condition vector. This vector is then injected into the cross-attention layer of the diffusion model for dual-condition joint frame-level rendering, generating a video frame sequence aligned with the audio duration shot by shot.
2. The multimodal driven intelligent MV synthesis and creation method according to claim 1, characterized in that, The steps of parsing the rhythmic structure of the shared feature map and outputting rhythmic structured tag blocks specifically include: The shared feature map is extracted into a high-dimensional rhythm feature sequence through a rhythm feature extraction operation based on a temporal convolutional network. The high-dimensional rhythm feature sequence is input into four parallel fully connected prediction heads, which output frame-level beat probability, frame-level bar boundary probability, frame-level accent intensity, and frame-level paragraph label, respectively; the values of the paragraph label correspond to the intro, verse, chorus, interlude, or outro. The frame-level beat probability is subjected to non-maximum suppression, and frame indices greater than a first threshold are selected and converted into a beat timestamp sequence; the frame-level bar boundary probability is selected with frame indices greater than a second threshold and converted into a bar boundary timestamp sequence; the frame-level accent intensity is segmented according to the beat timestamp sequence and the average is calculated to obtain the discrete accent intensity value at each beat point; and a frame-level segment label sequence is output. The rhythmic structured tag block is formed by combining the beat timestamp sequence, the measure boundary timestamp sequence, the discrete accent intensity value, and the paragraph tag sequence.
3. The multimodal driven intelligent MV synthesis and creation method according to claim 2, characterized in that, The specific steps for parsing the shared feature map for sentiment curves and outputting sentiment label blocks include: The shared feature map is extracted into a sentiment feature sequence through a sentiment feature extraction operation based on a temporal convolutional network; The emotional feature sequence is input into three parallel fully connected prediction heads, which output frame-level valence trajectory, frame-level arousal trajectory, and frame-level dominance trajectory, respectively. Temporal smoothing and interpolation upsampling are applied to the valence trajectory, the arousal trajectory, and the dominance trajectory, respectively, to obtain a continuous emotion trajectory aligned with the rendering frame rate; The first and second time derivatives of the continuous emotional trajectory are calculated respectively, and emotional extreme points, emotional turning points and emotional steady-state segments are extracted according to preset detection rules and marked as emotional key frames; The continuous emotional trajectory and the emotional keyframes are combined to form the emotional tag block.
4. The multimodal driven intelligent MV synthesis and creation method according to claim 3, characterized in that, The specific steps for planning a rhythm-gated shot sequence based on the rhythm-structured tag blocks include: The rhythm structured tag block is converted into a vector sequence, which contains three channels: the distance to the nearest beat point, the accent intensity of the beat point to which the current frame belongs, and the paragraph tag channel of the current frame. The vector sequence is input into a sequence prediction operation based on a transformer, and a two-dimensional tensor is output as the shot planning tensor; the first dimension of the two-dimensional tensor is the number of shots, and the second dimension includes the shot start frame index and the shot end frame index. Calculate the rhythm gating loss function to force the shot end time to be constrained to the vicinity of the beat point where the accent intensity is higher than the threshold during the training phase; Based on the changes in valence, arousal, and accent intensity between two adjacent shots, perform a transition type selection operation and output the transition type; Based on the shot planning tensor and the transition type, a complete shot sequence plan is generated.
5. The multimodal driven intelligent MV synthesis and creation method according to claim 3, characterized in that, The steps for generating visual style parameters for emotion mapping based on the emotion tag blocks specifically include: The effect value, arousal value, and dominance value corresponding to the current frame are combined into an emotional three-dimensional vector. Through a mapping operation composed of fully connected layers, the emotional three-dimensional vector is mapped into visual style parameters. The visual style parameters include hue shift, color saturation, illumination intensity, depth parameters, and motion blur intensity parameters. For each discrete time point under the rendering frame rate, the corresponding emotional 3D vector is input into the mapping operation, and the visual style parameters of the current frame are output. Based on the user-adjustable emotion mapping sensitivity parameter, the output of the mapping operation and the preset default parameter vector are weighted and fused to obtain the final style parameter; A variation upper limit constraint is imposed on each component of the final style parameter between adjacent frames to prevent visual flicker.
6. The multimodal driven intelligent MV synthesis and creation method according to claim 1, characterized in that, The steps of encoding the shot planning tensor and the visual style parameter sequence into rhythm conditional vectors and style conditional vectors respectively, fusing them into a joint conditional vector, and then injecting it into the cross-attention layer of the diffusion model for dual-conditional joint frame-level rendering specifically include: Each shot information in the shot planning tensor is encoded as a rhythm condition vector; the rhythm condition vector includes the sum of beat position embedding, accent intensity embedding, paragraph type embedding, and temporal position encoding; The visual style parameters of each frame are converted into a style condition vector through a fully connected layer; The rhythm condition vector and the style condition vector at the same time point are concatenated and then linearly projected to obtain the joint condition vector; In the cross-attention block of the diffusion model U-shaped network, the joint conditional vector is injected as an additional key value sequence into the attention calculation, so that the diffusion model can simultaneously receive shot switching timing guidance and frame-by-frame visual style guidance during the denoising process.
7. The multimodal driven intelligent MV synthesis and creation method according to claim 4, characterized in that, The rhythm gating loss function is calculated as follows: Calculate the numerical distance between the end time of each shot and the nearest high-emphasis beat, and then multiply the average value by the normalization factor after taking the average value for the number of shots.
8. The multimodal driven intelligent MV synthesis and creation method according to claim 6, characterized in that, In the step of performing dual-condition joint frame-level rendering, a temporal smoothing constraint loss is introduced when training the diffusion model. The temporal smoothing constraint loss is constructed as follows: When calculating the denoising loss of the diffusion model, penalty weights are assigned to the denoising prediction differences between adjacent frames, wherein the penalty weight of adjacent frames within the same shot is greater than the penalty weight of adjacent frames across shots.
9. A multimodal driven intelligent MV synthesis and creation method according to claim 4, characterized in that, The steps following dual-condition joint frame-level rendering also include: For the reasoning stage, it is generated frame by frame in sequence according to the shot; For each shot, first load the rhythm condition vector of the shot itself as a shared condition for all frames of the shot, then obtain the frame-level style condition vector frame by frame and inject it into the diffusion model for denoising and generation. Perform a transition blending operation on the first frame of the current shot and the last frame of the previous shot, according to the transition type. Once all shots are generated, the complete video frame sequence is stitched together and output.
10. A multimodal driven intelligent MV synthesis and creation system, characterized in that, Performing the multimodal-driven intelligent MV synthesis and creation method as described in any one of claims 1 to 9 includes: The feature extraction module is used for preprocessing and feature extraction of the input audio to obtain a shared feature map; The rhythm parsing module is used to parse the rhythm structure of the shared feature map and output rhythm structured label blocks. The sentiment analysis module is used to analyze the sentiment curve of the shared feature map and output sentiment tag blocks; the rhythm structured tag blocks and the sentiment tag blocks share the same time coordinate system; The shot planning module is used to plan shot sequences with rhythm gating based on the rhythm structured tag block and output a shot planning tensor; the shot planning tensor defines the start and end time positions of each shot; The style generation module is used to generate visual style parameters for emotion mapping based on the emotion tag block, and outputs the visual style parameter sequence frame by frame; the visual style parameter sequence defines the visual rendering parameters when rendering each frame. The video rendering module is used to encode the shot planning tensor and the visual style parameter sequence into rhythm condition vector and style condition vector respectively, and then fuse them into a joint condition vector. After that, it is injected into the cross attention layer of the diffusion model to perform dual-condition joint frame-level rendering and generate a video frame sequence aligned with the audio duration shot by shot.