Video background music generation method and system, intelligent terminal and storage medium
By determining and deduplicating candidate keyframes in video background music generation, the redundant calculation problem caused by fixed period sampling is solved, and the processing efficiency is improved.
Patent Information
- Application Number
- CN202510084766.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, video frame sampling is performed using fixed periods, resulting in video frames with highly repetitive information being collected, resulting in redundant calculations and reducing the processing efficiency of the background music generation process.
By obtaining the pending video, multiple candidate keyframes are determined, and deduplication is performed according to the structural similarity between the candidate keyframes, the target keyframe is obtained, and then the background music matching the video is generated based on the target keyframe.
By removing the candidate keyframes of repeated information, redundant calculations are avoided, and the processing efficiency of the background music generation process is improved.
Smart Images

Figure CN120050444A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly relates to a method, system, intelligent terminal, and storage medium for generating background music for videos. Background Art
[0002] Background music is an important part of a video. To avoid copyright issues caused by using existing music as the background music for a video, in some scenarios, it is necessary to generate background music for the video.
[0003] In the prior art, background music can be generated based on video features extracted from a video through a pre-set model. When extracting the video features corresponding to the video, multiple video frames are usually obtained by sampling at a fixed period, and video feature extraction is performed based on the obtained video frames. The problem with the prior art is that when sampling video frames at a fixed period, the actual characteristics of the video are not considered, and it is easy to collect video frames with highly repeated information, resulting in redundant calculations, which is not conducive to improving the processing efficiency of the background music generation process. Therefore, the related technology still needs to be improved and developed. Summary of the Invention
[0004] The main purpose of this application is to provide a method, system, intelligent terminal, and storage medium for generating background music for videos, aiming to solve the technical problem in the related technology that when sampling video frames at a fixed period to extract video features based on the sampled video frames, it is easy to collect video frames with highly repeated information, resulting in redundant calculations, which is not conducive to improving the processing efficiency of the background music generation process.
[0005] To achieve the above object, in the first aspect of this application, a method for generating background music for videos is provided. The method for generating background music for videos includes: obtaining a video to be processed; determining a plurality of candidate key frames from the video to be processed; removing duplicates from the candidate key frames according to the structural similarity between the candidate key frames to obtain target key frames; and generating background music matching the video to be processed according to the target key frames.
[0006] Optionally, determining a plurality of candidate key frames from the video to be processed includes: using the first video frame of the video to be processed as a candidate key frame; for each video frame in the video to be processed except the first video frame, using the candidate key frame that is before the video frame and closest to the video frame as the comparison key frame corresponding to the video frame, obtaining comparison result data between the video frame and the comparison key frame, and if the comparison result data meets a preset threshold condition, using the video frame as a candidate key frame; where the comparison result data includes at least one of inter-frame difference data, color distribution comparison data, and sparse optical flow detection data.
[0007] Optionally, the above-mentioned deduplication process for the above-mentioned candidate key frames according to the structural similarity between the above-mentioned candidate key frames to obtain target key frames includes: sorting the above-mentioned candidate key frames according to the timestamps corresponding to the above-mentioned candidate key frames and obtaining a key frame list; using the first candidate key frame in the above-mentioned key frame list as the target key frame; for each candidate key frame in the above-mentioned key frame list except the first candidate key frame, using the target key frame that is before the above-mentioned candidate key frame and closest to the above-mentioned candidate key frame as the target comparison frame, calculating the structural similarity between the above-mentioned candidate key frame and the above-mentioned target comparison frame, and if the above-mentioned structural similarity is not greater than a preset structural similarity threshold, using the above-mentioned candidate key frame as the target key frame.
[0008] Optionally, the above-mentioned generation of background music matching the above-mentioned video to be processed according to the above-mentioned target key frames includes: extracting features from the above-mentioned target key frames to obtain key frame feature vectors, where the above-mentioned key frame feature vectors are used to represent the semantic features and color features corresponding to the above-mentioned target key frames; generating background music matching the above-mentioned video to be processed according to the above-mentioned key frame feature vectors.
[0009] Optionally, the above-mentioned semantic features include semantic information features and emotional information features;
[0010] The above-mentioned extraction of features from the above-mentioned target key frames to obtain key frame feature vectors includes: encoding the above-mentioned target key frames through a text-to-image joint embedding pre-trained model to obtain scene semantic vectors; extracting hue vectors from the above-mentioned target key frames, analyzing the brightness distribution of the above-mentioned target key frames to obtain brightness vectors, and concatenating the above-mentioned hue vectors and the above-mentioned brightness vectors to obtain color vectors; performing fusion processing on the above-mentioned scene semantic vectors and the above-mentioned color vectors to obtain fusion vectors; concatenating the above-mentioned scene semantic vectors, the above-mentioned fusion vectors, and the above-mentioned color vectors to obtain the above-mentioned key frame feature vectors.
[0011] Optionally, the above-mentioned generation of background music matching the above-mentioned video to be processed according to the above-mentioned key frame feature vectors includes: generating a chord sequence matching the above-mentioned video to be processed according to the above-mentioned key frame feature vectors; generating background music matching the above-mentioned video to be processed according to the above-mentioned key frame feature vectors and the above-mentioned chord sequence.
[0012] Optionally, the above-mentioned generation of background music matching the above-mentioned video to be processed according to the above-mentioned key frame feature vectors and the above-mentioned chord sequence includes: generating a melody sequence according to the above-mentioned key frame feature vectors and the above-mentioned chord sequence; generating an accompaniment sequence according to the above-mentioned chord sequence and the above-mentioned melody sequence; obtaining background music matching the above-mentioned video to be processed according to the above-mentioned chord sequence, the above-mentioned melody sequence, and the above-mentioned accompaniment sequence.
[0013] In a second aspect of the present application, a video background music generation system is provided. The video background music generation system includes: a video acquisition module for acquiring a video to be processed; a key frame selection module for determining a plurality of candidate key frames from the video to be processed; a key frame deduplication module for deduplicating the candidate key frames according to the structural similarity between the candidate key frames to obtain target key frames; and a music generation module for generating background music matching the video to be processed according to the target key frames.
[0014] In a third aspect of the present application, an intelligent terminal is provided. The intelligent terminal includes a memory, a processor, and a video background music generation program stored on the memory and executable on the processor. When the video background music generation program is executed by the processor, the steps of any one of the above video background music generation methods are implemented.
[0015] In a fourth aspect of the present application, a computer-readable storage medium is provided. A video background music generation program is stored on the computer-readable storage medium. When the video background music generation program is executed by a processor, the steps of any one of the above video background music generation methods are implemented.
[0016] As can be seen from the above, in the solution of the present application, a video to be processed is acquired; a plurality of candidate key frames are determined from the video to be processed; the candidate key frames are deduplicated according to the structural similarity between the candidate key frames to obtain target key frames; and background music matching the video to be processed is generated according to the target key frames.
[0017] Compared with the prior art, in the solution of the present application, a plurality of candidate key frames in the video to be processed are first determined. For the candidate key frames, further deduplication processing is performed based on the structural similarity between the candidate key frames, so as to remove the video frames with repeated information and obtain target key frames, and subsequent background music generation process is carried out according to the target key frames. In this way, the candidate key frames with repeated information can be removed based on the deduplication processing, avoiding redundant calculations, and being beneficial to improving the processing efficiency of the background music generation process. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a flowchart of a video background music generation method provided by an embodiment of the present application;
[0020] Figure 2 is a schematic diagram of a training sample processing flow provided by an embodiment of the present application;
[0021] Figure 3 is a schematic diagram of a target key frame extraction flow provided by an embodiment of the present application;
[0022] Figure 4 is a corresponding relationship diagram between video key frames and music bars provided by an embodiment of the present application;
[0023] Figure 5 is a schematic diagram of a key frame feature vector extraction flow provided by an embodiment of the present application;
[0024] Figure 6 is a schematic diagram of a chord progression type statistical result provided by an embodiment of the present application;
[0025] Figure 7 is a schematic diagram of a chord tonality frequency statistical result provided by an embodiment of the present application;
[0026] Figure 8 is a schematic diagram of an auditory V-A emotion space coordinate system provided by an embodiment of the present application;
[0027] Figure 9 is a schematic diagram of the composition modules of a video background music generation system provided by an embodiment of the present application;
[0028] Figure 10 is a schematic diagram of the internal structure principle of an intelligent terminal provided by an embodiment of the present application. Detailed implementation manners
[0029] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0030] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, but the present application can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present application, so the present application is not limited by the specific embodiments disclosed below.
[0031] Early music generation systems were mostly based on rule algorithms or random processes (such as Markov chains), generating chords, melodies, etc. according to preset rules. The biggest drawback of this generation method is that the generated music lacks diversity and creativity. In one application scenario, music generation can be based on deep learning technology. For example, recurrent neural networks or other deep neural networks can be used to generate music. However, the core concern of this solution is music generation itself, rather than the correspondence between music and video content, resulting in the generated music being difficult to effectively align with the given video content in terms of emotional matching, rhythm synchronization, atmosphere, etc. In another application scenario, a background music generation scheme for free-form video content can be adopted. For example, a controllable music generator based on Transformer (CMT, Controllable Music Transformer) can be used, or a symbolic music-video dataset (SymMV, Symbolic Music-Video Dataset) can be used. SymMV affects music generation by extracting semantic information in the video and the color distribution of video key frames. However, the generated music lacks consistency with the color distribution in the video key frames in terms of emotion.
[0032] At the same time, in the current generation of video background music based on free-form video content, video features are usually extracted by sampling multiple consecutive video frames at fixed intervals, such as extracting semantic information and the average value of RGB or optical flow changes. However, this fixed-interval sampling method has high computational complexity and low efficiency. Especially when extracting semantic features, the fixed interval results in highly repetitive semantics of the sampled frames, blurred images, insignificant subjects, and unprominent emotions. Directly extracting the features of these continuously repeated frames will generate a large amount of redundant calculations and reduce the generation efficiency. In addition, it is difficult to set the sampling interval, and an unreasonable interval may lead to the loss of important semantic or emotional information.
[0033] To solve at least one of the above-mentioned multiple technical problems, in the solution of this application, a video to be processed is first obtained, multiple candidate key frames are extracted, and duplicate removal is performed on them according to structural similarity to obtain target key frames. Subsequently, background music matching the video is generated based on the target key frames. Compared with the prior art, this solution avoids redundant calculations by removing candidate key frames with repeated information, thereby improving the processing efficiency of background music generation.
[0034] As Figure 1 shown, an embodiment of this application provides a method for generating video background music. Specifically, the above method includes the following steps:
[0035] Step S100, obtain a video to be processed.
[0036] The video to be processed above is a video for which background music needs to be generated. Based on the video background music generation method provided in the embodiments of the present application, background music that matches the content and emotion of the above-mentioned video to be processed is generated for the above-mentioned video to be processed.
[0037] Step S200: Determine a plurality of candidate key frames from the above-mentioned video to be processed.
[0038] It should be noted that the method for determining candidate key frames from the video to be processed can be set and adjusted according to actual needs. In one application scenario, video frame sampling can be performed using a preset sampling period; in another application scenario, some video frames can be randomly selected from the video to be processed as candidate key frames. In the embodiments of the present application, to obtain a better music generation effect, candidate key frames are determined from the video to be processed according to the similarity between video frames.
[0039] Specifically, determining a plurality of candidate key frames from the above-mentioned video to be processed includes: using the first video frame of the above-mentioned video to be processed as a candidate key frame; for each video frame in the above-mentioned video to be processed except the first video frame, using the candidate key frame that is before the above-mentioned video frame and closest to the above-mentioned video frame as the comparison key frame corresponding to the above-mentioned video frame, obtaining comparison result data between the above-mentioned video frame and the above-mentioned comparison key frame, and if the above-mentioned comparison result data meets a preset threshold condition, then using the above-mentioned video frame as a candidate key frame; wherein, the above-mentioned comparison result data includes at least one of inter-frame difference data, color distribution comparison data, and sparse optical flow detection data.
[0040] Specifically, if the above-mentioned comparison result data is greater than a preset threshold, it meets the preset threshold condition. It should be noted that if the comparison result data includes one of inter-frame difference data, color distribution comparison data, and sparse optical flow detection data, then this type of data being greater than the corresponding threshold is sufficient. If it includes multiple of them, then each type of data needs to be greater than the corresponding threshold to be considered as meeting the preset threshold condition. In the embodiments of the present application, the above-mentioned comparison result data includes three types of data: inter-frame difference data, color distribution comparison data, and sparse optical flow detection data, and corresponding thresholds are set for each of the three types of data, and the corresponding thresholds can be set and adjusted according to actual needs, and no specific limitation is made here.
[0041] Step S300: Perform duplicate removal processing on the above-mentioned candidate key frames according to the structural similarity between the candidate key frames to obtain target key frames.
[0042] In the embodiments of the present application, the structural similarity (SSIM, Structural Similarity Index Measure) between candidate key frames is calculated. For multiple candidate key frames with a structural similarity greater than a preset structural similarity threshold (i.e., a group of relatively similar candidate key frames), only one of the candidate key frames is selected as the target key frame to achieve deduplication and avoid redundant calculations. It should be noted that in the embodiments of the present application, for a group of similar candidate key frames, the candidate key frame with the earliest timestamp is retained as the target key frame.
[0043] Specifically, the above-mentioned deduplication process of the candidate key frames based on the structural similarity between the candidate key frames to obtain the target key frame includes: sorting the candidate key frames according to the timestamps corresponding to the candidate key frames to obtain a key frame list; taking the first candidate key frame in the key frame list as the target key frame; for each candidate key frame in the key frame list except the first candidate key frame, taking the target key frame that is before the candidate key frame and closest to the candidate key frame as the target comparison frame, calculating the structural similarity between the candidate key frame and the target comparison frame, and if the structural similarity is not greater than the preset structural similarity threshold, taking the candidate key frame as the target key frame.
[0044] In this way, in the embodiments of the present application, when performing structural similarity judgment, for a candidate key frame, only the SSIM calculation and comparison are performed with the previous retained frame. By comparing only adjacent frames, the secondary deduplication is completed in linear time, which not only ensures the sufficient deletion of highly repeated key frames but also avoids the computational overhead caused by pairwise comparison. It should be noted that the specific value of the above-mentioned structural similarity threshold can be set and adjusted according to actual needs, and no specific limitation is made here.
[0045] Step S400, generate background music that matches the to-be-processed video according to the above-mentioned target key frame.
[0046] In an application scenario, based on the deduplicated target key frames, background music that matches the to-be-processed video in content and / or emotion can be generated through a music generation model or a video background music generation framework for free-form video content.
[0047] In the embodiments of the present application, the above-mentioned generation of background music that matches the to-be-processed video according to the above-mentioned target key frame includes: extracting features from the above-mentioned target key frame to obtain a key frame feature vector, where the key frame feature vector is used to represent the semantic feature and color feature corresponding to the above-mentioned target key frame; generating background music that matches the to-be-processed video according to the above-mentioned key frame feature vector.
[0048] Specifically, the above semantic features include semantic information features and emotional information features; the above feature extraction of the above target key frame to obtain a key frame feature vector includes: encoding the above target key frame through a text-to-image joint embedding pre-trained model to obtain a scene semantic vector; extracting the hue of the above target key frame to obtain a hue vector, analyzing the brightness distribution of the above target key frame to obtain a brightness vector, and concatenating the above hue vector and the above brightness vector to obtain a color vector; performing fusion processing on the above scene semantic vector and the above color vector to obtain a fusion vector; concatenating the above scene semantic vector, the above fusion vector, and the above color vector to obtain the above key frame feature vector.
[0049] It should be noted that the colors of the picture represent the emotional tendency to a certain extent. Therefore, the color features can also represent the corresponding emotional information. Similarly, the semantics of the video can also represent some emotional information.
[0050] In some embodiments of the present application, the above vector processing process is fused with a pre-constructed valence-arousal emotion model (V-A), so as to improve the representation ability of the obtained key frame feature vector for emotional information.
[0051] Specifically, in the embodiments of the present application, the above-mentioned target key frames are encoded by a Contrastive Language-Image Pre-Training (CLIP) model to obtain scene semantic vectors. The scene semantic vectors are processed by a trained image embedding layer to obtain image embedding vectors, the color vectors are processed by a trained color embedding layer to obtain color embedding vectors, and a trained intermediate attention layer (such as a Transformer) is used to fuse the color embedding vectors and the image embedding vectors to obtain fused vectors. The scene semantic vectors, the fused vectors, and the color vectors are vector-concatenated to obtain the above-mentioned key frame feature vectors. Among them, the CLIP model used is pre-frozen and does not need to be trained. During the training process, the key frame feature vectors obtained during training are input into a fully connected layer and the corresponding V-A coordinates are predicted, and the predicted V-A coordinates are compared with the corresponding label data to update the parameters corresponding to the image embedding layer, the color embedding layer, and the intermediate attention layer, so that the V-A coordinates predicted by the key frame feature vectors obtained based on the adjusted parameters are closer to the real label data, and the representation ability of the corresponding key frame feature vectors for emotional information is improved. In this way, in the task of generating background music, based on the updated image embedding layer, color embedding layer, and intermediate attention layer after training, the final key frame feature vectors that can represent semantic information features, emotional information features, and color features can be obtained. Then, based on the key frame feature vectors, background music that matches the video to be processed in terms of semantics, emotion, and color is generated.
[0052] In an application scenario, background music can be generated based on a pre-trained music generation model, such as CMT can be used. In the embodiments of the present application, background music is generated based on a pre-constructed video background music generation framework, and the music generation process is decoupled into three parts.
[0053] Specifically, generating background music that matches the above-mentioned video to be processed according to the above-mentioned key frame feature vectors includes: generating a chord sequence that matches the above-mentioned video to be processed according to the above-mentioned key frame feature vectors; generating background music that matches the above-mentioned video to be processed according to the above-mentioned key frame feature vectors and the above-mentioned chord sequence.
[0054] Further, generating background music that matches the above-mentioned video to be processed according to the above-mentioned key frame feature vectors and the above-mentioned chord sequence includes: generating a melody sequence according to the above-mentioned key frame feature vectors and the above-mentioned chord sequence; generating an accompaniment sequence according to the above-mentioned chord sequence and the above-mentioned melody sequence; obtaining background music that matches the above-mentioned video to be processed according to the above-mentioned chord sequence, the above-mentioned melody sequence, and the above-mentioned accompaniment sequence.
[0055] In this way, a modular music generation scheme is adopted to achieve fine-grained conditional control of different music structures, improving controllability and flexibility. First, based on chord generation, the emotional and structural foundation is provided for the generation of melodies and accompaniments. Second, using chord tokens as control signals, a melody is generated through a Linear Transformer with linear attention, enhancing efficiency and long-range dependence modeling capabilities. Finally, using chord and melody tokens as conditional signals, a Polyfussion diffusion model is used to generate the accompaniment. During the accompaniment generation process, the Music VAE encoder encodes the tokens, and the conditional signals are dynamically mapped to intermediate features through the cross-attention layer in the U-Net, ensuring a high degree of consistency in emotion and structure between the accompaniment, chords, and melody. The combination of modular generation and conditional control realizes the efficient decomposition and precise regulation of the music generation process, ensuring the generation of high-quality and controllable background music.
[0056] As can be seen from the above, in the video background music generation method provided by the embodiment of the present application, multiple candidate key frames in the video to be processed are first determined. For the candidate key frames, further duplicate removal processing is performed based on the structural similarity between the candidate key frames, so as to remove the video frames with duplicate information and obtain the target key frames, and subsequent background music generation processes are carried out according to the target key frames. In this way, the candidate key frames with duplicate information can be removed based on the duplicate removal processing, avoiding redundant calculations and being beneficial to improving the processing efficiency of the background music generation process.
[0057] In this embodiment, the video background music generation method is described in detail based on specific application scenarios, and the entire processing flow is illustrated in combination with the training and generation processes. The data processing and training processes for background music generation are basically the same, and they can refer to each other. The specific steps sequentially include input processing, key frame metadata design, feature extraction, generation framework, chord-emotion-color mapping modeling, training joint embedding space, and retrieval accuracy index design, comprehensively elaborating the specific implementation of the method.
[0058] Figure 2 It is a schematic diagram of a training sample processing flow provided by the embodiment of the present application. In this embodiment, video segments with background music of 30 - 120 seconds are used as training samples, sourced from music videos and classic movies. The audio is extracted through FFmpeg, denoised and mono-channeled (16 kHz) using Audacity. High-quality piano tracks are screened and separated using Spleeter. The piano tracks are converted into MIDI symbolic music using the Onset and Frames algorithm and saved as.mid files for subsequent modeling.
[0059] Since the generation task requires aligning the timing of the video with the background music, it is necessary to perform bar division on the background music. Bar division splits the continuous music stream into organized bars according to the beats per minute (BPM) and time signature, ensuring compliance with music theory rules and consistency with the overall rhythm and dynamics. A bar typically contains a fixed number of beats, such as 4 / 4, 3 / 4, or 6 / 8 time, and the duration of each beat is determined by the time signature. For example, in 4 / 4 time, there are 4 beats per bar, and each beat is a quarter note. The beats within a bar are divided into strong beats and weak beats, enhancing the rhythm and expressiveness of the music. BPM represents the speed of the music, and together with the time signature, it determines the duration of each bar. This division method ensures an orderly music structure, facilitating emotional expression and timing alignment for the generation task. For example, for music in 4 / 4 time with 120 BPM, each bar lasts for 2 seconds (0.5 seconds per beat). After determining the time signature and BPM value of the background music, the following steps can be taken for bar division: For standard 4 / 4 time, 120 BPM symbolic music, the continuous MIDI note sequence is divided into one bar every 2 seconds (i.e., every 4 beats), ensuring that the notes and rhythm patterns within each bar conform to the 4 / 4 time signature. Use the PrettyMIDI library to extract the time signature and global rhythm information (BPM) from the background music MIDI file and determine the position of the strong beats in each bar. After obtaining the above time signature information, BPM, and downbeat positions, it is possible to locate each music bar. Based on the time signature and BPM, the exact duration of each bar can be calculated using the following formula (1):
[0060]
[0061] In the embodiments of this application, the downbeat is used as the starting beat of each bar, and the boundaries are divided according to the duration of each bar. The start and end times of the bar are mapped to the video timeline to generate video segments corresponding to the bar duration, achieving the alignment of video content with music bars. This alignment provides a basis for key frame extraction, rhythm synchronization, and semantic and emotional matching, ensuring the temporal coordination of visual content and background music.
[0062] Furthermore, based on video output processing, representative semantic and emotional key frames are extracted from the original video to prepare for subsequent semantic, emotional, and color feature extraction. When actually generating the background music, the input is the video without background music, while the video with background music is used in the training phase. To improve the processing efficiency, first, FFmpeg is used to compress the video resolution to 360p and reduce the frame rate to 30fps, and then the video is segmented according to musical measures. For each segment, a hierarchical deduplication algorithm is adopted to retain the most representative key frames based on visual, color, and semantic similarities. This ensures that while reducing storage and computational overhead, key details are maintained, supporting subsequent emotion analysis and music generation tasks.
[0063] Figure 3 is a schematic diagram of a target key frame extraction process provided by an embodiment of the present application. As Figure 3 shown, the extraction of target key frames is generally divided into two stages. The first stage is the extraction of candidate key frames, and the second stage is to determine the target key frames by removing duplicates based on similarity. It should be noted that for the process of extracting target key frames, the corresponding processing methods are the same during the training process and the video background music generation process.
[0064] As Figure 3 shown, in the embodiment of the present application, in the candidate key frame extraction stage, first, the video is read through cv2.VideoCapture, the first frame is set as the initial key frame, and its color histogram is calculated as a reference. Subsequently, for each subsequent frame, frame difference is performed with the nearest previous candidate key frame. If the number of non-zero pixels is lower than the set threshold frame_diff_pixels_threshold, then this frame is skipped. Otherwise, the HSV histograms of the current frame and the previous candidate key frame are calculated and compared through the Bhattacharyya distance. If the distance is less than color_threshold, it is considered that the color distributions are similar, and this frame is skipped. For the frames determined by color, Lucas-Kanade sparse optical flow is used to calculate the motion saliency. If the average displacement of the corner points exceeds flow_threshold, then the current frame is set as a new key frame and the reference state is updated. After all frames are processed, the extracted candidate key frames are saved and a key frame list (frame_index, frame) is generated. Among them, for any current frame, its corresponding previous candidate key frame is the candidate key frame with a timestamp before the current frame and the timestamp closest to the current frame.
[0065] In the similarity deduplication stage, first, the key frames are sorted according to their timestamps, and then the structural similarity index (SSIM) is calculated with the previous retained target key frame in turn. When the SSIM value of two adjacent frames is greater than or equal to the set threshold (such as 0.5), they are regarded as duplicates, and only the frame with the earlier timestamp is retained as the target key frame. This process can choose to delete redundant frames or only record prompts, and complete deduplication in linear time by the method of "only comparing adjacent frames", effectively reducing the computational overhead. The overall algorithm combines inter-frame difference, color histogram and optical flow detection to identify significantly changed frames, and then performs deduplication through SSIM, and finally outputs a concise and representative key frame sequence. This method has high efficiency and scalability, adapts to videos with different resolutions and content features, and ensures good performance in most video processing applications through the use of optimized functions of OpenCV and linear complexity design.
[0066] Figure 4 It is a correspondence diagram between video key frames and music bars provided by an embodiment of the present application. It should be noted that Figure 4 the key frames shown in Figure 4 are the target key frames obtained after extraction. As Figure 4 shown, for training samples, the input video and background music are processed to extract the music bars and the target key frames in the corresponding time periods to ensure the synchronization of music and video on the time axis. During the video key frame extraction process, the target key frame file name retains the timestamp (milliseconds), and each key frame is the earliest one in a series of consecutive similar frames.
[0067] Furthermore, after obtaining the target key frames, their semantic and emotional features are extracted, which belongs to the task of visual emotion analysis (VAE, Visual Emotion Analysis). Traditional methods use image encoders such as deep convolutional neural networks or ViT to classify pictures into fixed emotion categories, but this method ignores the complexity and continuity of emotions. This embodiment adopts the V-A (valence-arousal) emotion model to model the emotional content of key frames. The V-A model describes emotions through two continuous axes: valence (from happy to sad) and arousal (from excited to calm), which can depict and quantify emotional states more delicately and is more applicable than traditional discrete emotion categories, thus improving the accuracy of emotional representation.
[0068] Figure 5 It is a schematic diagram of a key frame feature vector extraction process provided by an embodiment of the present application. Among them, Figure 5 The image processed in is the target key frame determined based on the above processing process. Specifically, a CLIP image encoder is used to extract the scene semantic vector to capture the high-level semantic features of the image (such as scene type, object category). At the same time, low-level visual features (such as brightness, dominant color) are supplemented to more accurately reflect the emotional clues. The high-level features reflect "what" and "what is being done", and the low-level features provide color and atmosphere information. By fusing high-level and low-level features, a more comprehensive emotional feature representation is formed, understanding both semantic content and perceiving the visual atmosphere.
[0069] Specifically, on each video segment, the dominant color of all target key frames involved is extracted based on k-means clustering in the Lab color space (L, a, b) (pixels that are too dark are filtered out before extraction). For the brightness information in the target key frame, the distribution of brightness in the picture is obtained by performing a histogram statistics on the V channel in the HSV space. When calculating the brightness distribution, the continuous brightness values (discrete pixel values from 0 to 255) are divided into 32 blocks by interval (i.e., bins = 32), and the number of pixels in a certain brightness range is counted to obtain the brightness distribution. The brightness distribution of each key frame is characterized and returned by a one-dimensional histogram vector (BrightnessHistogram) with a shape of 1*32.
[0070] In the embodiment of the present application, 5 dominant colors are extracted from each key frame based on k-means clustering (k = 5, and the specific value can be set and adjusted according to actual needs). These 5 dominant colors have different proportions. For example, the preset weight proportions are ratio 1 、ratio 2 、ratio 3 、ratio 4 and ratio 5 (the weight proportion can be set and adjusted according to actual needs). A dominant color center is obtained by weighted averaging the 5 dominant colors based on the weight proportion, and then it is concatenated with the brightness histogram Figure 1 dimensional vector to obtain the final key frame color feature, as shown in the following formulas (2) and (3):
[0071]
[0072] ColorFeature = Concat(WeightedCenter, BrightnessHistogram)(3);
[0073] Among them, (L i, a i , b i ) represents the i-th dominant color, WeightedCenter represents the vector corresponding to the center of the dominant color, and BrightnessHistogram represents the brightness histogram Figure 1 dimensional vector, Concat represents the vector concatenation operation, and ColorFeature represents the color vector, that is, the final key-frame color feature obtained by concatenation.
[0074] In this embodiment, the CLIP model (ViT-B / 32) is used to encode the key-frame images to generate 512-dimensional scene semantic vectors, and the relationships among semantics, colors, and emotions are established by combining color features such as brightness and hue. 7500 images and their normalized V-A coordinates in the OASIS dataset are used as emotion labels. Scene semantics and color features are extracted from each image, and the two are fused through a cross-attention mechanism. Specifically, the color features are mapped to queries, the semantic features are mapped to keys / values, the attention-weighted semantic features are calculated and concatenated with the color features to form a key-frame fusion feature vector. Subsequently, the dimension is reduced through a fully connected layer and a regressor is used to predict the V-A coordinates, and the model parameters are updated based on the prediction loss, thereby generating a key-frame feature vector that fuses semantic, color, and emotion information. As Figure 5 shown, the image is encoded by CLIP to generate a 512-dimensional scene semantic vector, and is further processed through a linear layer to remain 512-dimensional. Since the video background music generation task is different from the image-text matching objective of CLIP, there may be distribution differences when directly using CLIP features. Adding a linear layer enables the model to adapt to these tasks, and the attention mechanism is optimized by fine-tuning the CLIP features, improving the training convergence and performance. In multi-head attention, the linear projection provides a learnable transformation for each attention head, enhancing the expressive ability and the depth of feature interaction, thereby optimizing the feature adaptation ability.
[0075] In Figure 5 , both the key vector and the value vector come from the scene semantic vector of CLIP and obtain the embedding vector through a shared linear projection; the query vector comes from the color features of the key frame and is generated through a 512-dimensional linear color embedding layer. The weights are calculated through dot-product attention to efficiently fuse semantic and color features. In the multi-head attention layer (8 heads) of the Transformer, the output is projected to 256 dimensions through a fully connected layer and concatenated with the CLIP features through a skip connection, finally forming a 768-dimensional fusion key-frame feature vector for predicting the V-A coordinates.
[0076] In the prediction stage, the 768-dimensional features first pass through a fully connected layer to be reduced to 512 dimensions, and then a regressor is used to predict the final V-A coordinate values, so as to achieve the accurate mapping of visual features and emotions that integrate semantic and color features. This design can greatly improve the interaction depth of semantic and color information, ensuring feature adaptability and final task performance.
[0077] It should be further noted that during the training process, the image embedding layer, color embedding layer, and intermediate attention layer have been trained so that the finally output key frame feature vectors can represent the corresponding emotions. Therefore, in the actual background music generation task, based on Figure 5 the processing flow shown, it is only necessary to process until the key frame feature vectors are obtained by splicing, and there is no need to perform the subsequent process of predicting V-A coordinates.
[0078] In the embodiments of the present application, the processing of the background music in the training samples during the training process is also specifically described. Specifically, the Skyline algorithm is used to separate the melody and accompaniment notes of the background music. First, the background music is divided into measures (such as BPM 120, time signature 4 / 4, 2 seconds per measure), and the note duration is standardized with 16th notes as the unit. The melody, as the core of the music, combines features such as pitch, dynamics, and duration to extract the most representative notes, and the remaining notes are used as accompaniment to support the main melody. The specific implementation uses pretty_midi to parse MIDI files, the Skyline algorithm to identify melody notes, and extract information such as pitch, start time, duration, and dynamics to complete the separation and extraction of melody and accompaniment. Secondly, identifying the chord progression in the background music is the key. A chord consists of at least three notes arranged vertically, including the root note, type, and tonality, which determine the musical tone. Major chords (such as C major) convey bright emotions, while minor chords (such as A minor) appear melancholy. The chord progression is an ordered sequence of chords that reflects the style, tonality, and emotional changes of the music and drives the emotional development. Accurately identifying the chord progression depends on determining each chord type and analyzing the relationship between chords. In terms of modeling, a sequence model (such as Chord REMI-Based LinearTransformer) is used to tokenize the chords and learn the chord transition probabilities and change patterns. This method captures the relationship between chords and extracts the complete chord sequence features. The specific implementation uses the chorder and midi21 libraries to automatically detect and extract chords and their progression features from MIDI files, ensuring the capture of chord flow and emotional changes in the time dimension and providing support for music generation and understanding tasks.
[0079] Furthermore, REMI (REvamped MIDI-derived events) is used to encode note events, converting melodies, accompaniments, and chord notes into discrete event sequences, which facilitates machine learning models to understand and generate music. REMI preserves the temporal and structural information of music while simplifying note representation. Each event contains a type (such as Note-On, Note-Off, Chord), a time point (in beats), pitch, velocity, duration, and chord information (root note and quality). In this embodiment, a type field is added to distinguish between melody note (melody_note_on), accompaniment note (accompaniment_note_on), and chord events. The time granularity per measure is quantized to 16th notes to ensure compliance with the REMI encoding specification. This encoding method standardizes music data and supports efficient music generation and analysis.
[0080] For example, for a musical measure with a 4 / 4 time signature, the total duration of each measure is 2s, each measure has 4 beats, each beat is a quarter note, and the duration of a quarter note is four times that of a 16th note. Therefore, a standard time step step is specified as 0.125s, which is the duration of a 16th note. For notes with a duration less than 0.125s, rounding is used for approximation. The quantized duration of each note is as shown in the following formula (4):
[0081] quantize_time = round(time / step)*step (4);
[0082] where time is the actual starting moment of the current note, and step = 0.125s. Under REMI encoding, chord events are regarded as instantaneous events, that is, only the starting time is retained, without duration, so as to represent the specific chord changes at each time step and the chord combination information in each chord progression in a finer granularity. Melody and accompaniment events retain duration information in addition to the starting time.
[0083] Furthermore, chord-emotion alignment is achieved. Table 1 is a mapping table of chord types and emotions provided by the embodiment of the present application. The emotion space is selected as: excited, fearful, tense, sad, and relaxed.
[0084] Table 1
[0085]
[0086] As shown in Table 1, when performing chord-to-emotion mapping, the chord-to-emotion mapping aligned by music experts can be used as a benchmark. This approach is simple and straightforward, but the emotion space is small and cannot fully accommodate emotional feelings. Although directly using static chord-to-emotion mapping is convenient, it ignores individual and cultural differences, context and dynamic characteristics, complex emotional dimensions, and the lack of data-driven verification. Therefore, in this embodiment, an auditory experiment is designed, combined with data-driven statistical analysis and machine learning, to construct a more flexible, detailed, and verifiable chord-to-emotion mapping model, and combine the color emotion in the video key frames with the music emotion. At the same time, a color-to-emotion experiment based on scene association is carried out. In a unified emotional space, map chords and colors to emotions, and establish the association between chords and colors through emotions as a bridge. The experiment indirectly establishes the association between chords and colors by collecting the emotional feedback of the subjects on chord sequences and color images and mapping the two to the same emotional space. Specifically, the experiment constructs an emotion space through the V-A emotion model, uses two psychological dimensions of valence and arousal to form a two-dimensional coordinate system, and describes the emotion characteristics more carefully. Valence measures the degree of pleasure of emotional experience, with one end representing positive pleasure and the other end representing negative unhappiness; arousal measures the intensity and activity of emotions, with one end being high arousal (excited, tense) and the other end being low arousal (calm, relaxed). The position of emotions in the coordinate system, such as "high valence and high arousal" is happy and excited, and "low valence and low arousal" is depressed and tired, thus flexibly depicting the diverse emotional experiences triggered by music and colors.
[0087] Further, prepare the chord progression samples. Most of the music works in the video-background music dataset are in the style of pop music. Therefore, in this embodiment of the present application, 500 pop music pieces are extracted from each of SymMV and MuViSync, for a total of 1000 pop music pieces, and their MIDI-formatted files are obtained. The chord roots and tonality (major or minor) involved in each piece of music are extracted using chorder, and the chord progression of each MIDI file is analyzed using music21.
[0088] Figure 6 It is a schematic diagram of the statistical results of a chord progression type provided by an embodiment of the present application. Figure 7 It is a schematic diagram of the statistical results of the chord tonality frequency provided by an embodiment of the present application. As Figure 6 and Figure 7As shown, in terms of the frequency of chord occurrences, the C chord is the most common, exceeding 30,000 times, indicating the dominant position of C major in pop music and playing an important role as the tonic chord in the chord progression. Following closely are the Amin, F, and G chords, which also have high frequencies, showing that the I-IV-V or related chord progressions are commonly used in pop music. The main chord types are major chords (such as C, F, G) and minor chords (such as Amin, Emin, Dmin), which are the mainstream of modern pop music. At the same time, seventh chords (such as Amin7, G7, E7) are also relatively common, used to enrich chord colors and tonal variations. Diminished chords (such as Ddim) and suspended chords (such as G-sus4) are less common because of their special sound and limited usage scenarios. Tonally, C major and A minor are the most prominent, conforming to the common major-minor key conversion in pop music. Sharp and flat chords such as D#, G#, and Fmin are used for key modulation and expressing special emotional changes.
[0089] After completing the frequency statistics of the chord sequence, the key of each music work is analyzed. The key defines the overall context of the music, distinguishing between "diatonic" and "non-diatonic" notes and chords, and helping to understand the function of chords in the structure. For example, the G chord is the dominant chord (V) in C major, while it is the tonic chord (I) in G major. Key analysis improves the accuracy of chord analysis and supports the generation of melodies and accompaniments with consistent keys.
[0090] To improve the reliability of key detection, three classic algorithms (Krumhansl-Schmuckler, Temperley-Kostka-Payne, and Bellman-Budge) are adopted, and the results are integrated through a voting mechanism. These three algorithms are based on pitch perception, probability models, and harmonic stability respectively, and are suitable for different music types. The integrated voting method overcomes the limitations of a single algorithm and significantly improves the accuracy and robustness of key detection, thus providing a solid foundation for chord progression classification and emotional expression.
[0091] Major keys usually convey positive emotions, while minor keys express sadness. In the dataset, major keys (such as G, D, C major) account for a relatively high proportion, indicating that most music has bright emotions; minor keys (such as D, A, E minor) are less, reflecting less sad or mysterious emotions. The key distribution shows a long tail, with a few keys occurring frequently and most keys being rare. This analysis guides the priority selection of high-frequency keys when generating music to conform to mainstream preferences.
[0092] Based on chord type and key analysis, the key is normalized to major and minor keys, and the chords are classified into tonic chords, dominant chords, and subdominant chords according to their functions. Through frequency statistics, representative chords are extracted: there are 12 major chords, including tonic chords (C, G, F, A), dominant chords (G7, D7, E7), and subdominant chords (F, Dm, A, Bb); there are 9 minor chords, including tonic chords (Am, Dm, Em), dominant chords (E7, B7, G7), and subdominant chords (Dm, Gm, Cm). Tonic chords provide emotional stability, dominant chords create tension and stimulate emotions, and subdominant chords play a transitional role to ensure natural and coherent emotional expression.
[0093] Based on the above-extracted representative chords, 100 music segments of 6 - 10 seconds (4 / 4 time signature, 3 - 5 measures) are sampled from 1000 pop songs, mainly selected from the chorus part to concentrate on the emotional theme, and at the same time including the bridge, prelude, and main melody connection parts to reflect emotional transitions. These segments contain common chord combinations, such as major tonic chords (C, G), dominant chords (G7, D7), subdominant chords (F, Dm), and minor tonic chords (Am, Dm), dominant chords (E7, B7), subdominant chords (Gm, Cm). Chord progressions such as I-IV-V-I, G→A→F→G7 in major keys, i-iv-V-i and Dm→Gm→A→Dm in minor keys, etc., convey different emotions such as stable and pleasant, high tension, sad and tense, and calm and melancholy respectively. These chords are distributed in the complete chord progression structure to ensure that each segment can effectively express the corresponding emotional characteristics and support the training of the emotion prediction model.
[0094] For each sampled music segment containing a complete chord progression, the following annotations are added to form metadata information: Chord type: the chord categories included and their arrangement order. Chord function: the proportion of tonic chords, dominant chords, and subdominant chords. Key transition: whether the segment contains a transition from major to minor or from minor to major. Rhythm pattern: whether the chord changes follow the standard time signature (e.g., under 4 / 4, the chord distribution per measure is 1 - 2) or other distribution forms. Emotional characteristics: obtain the emotional distribution of the segment through the V-A coordinate system.
[0095] Figure 8 It is a schematic diagram of the auditory V-A emotion space coordinate system provided by an embodiment of the present application. In this embodiment, 500 interviewees are invited through an online mini-program to listen to 3 music segments and perform emotion annotation using the same V-A two-dimensional model as the visual emotion space. Some sampled real labels are such as Figure 8As shown. In terms of emotion distribution, the V-A two-dimensional model consistent with the visual emotion space is continued to be adopted, and valence and arousal coordinates are used to describe and classify emotions. Valence ranges from extremely negative (such as sadness) to extremely positive (such as pleasure), and arousal ranges from low activation (such as calm) to high activation (such as excitement). The positions of different emotions on the two-dimensional plane reflect their characteristics. For example, excitement is located in the upper right corner (high valence, high arousal), sadness is located in the lower left corner (low valence, low arousal), and fear is located in the upper left corner (low valence, high arousal). This design simplifies complex emotions into a structured framework, facilitating modeling and analysis. When tagging respondents, emotional trend lines are added to the initial V-A coordinate system to reflect the transitions between emotion quadrants, and the arrow direction shows the dynamic changes of emotions. For example, the transition from fear to anger is manifested by accelerating the rhythm and increasing the volume, while the transition from melancholy to sadness is expressed by slowing down the rhythm and lowering the pitch.
[0096] Through the above experiments, 1500 (music segment, V-A coordinate) samples were collected, and each segment contains 3 - 5 bars, 6 - 10 seconds, and 8 - 12 chords. Through key transposition and rhythm perturbation for data augmentation, the sample size increased to 3000 while maintaining consistent emotional attributes. The music segments were split by chord progressions to obtain 7500 (chord progression, V-A coordinate) samples, and each progression lasts for 2 - 5 seconds. To align chord features with the V-A emotion space, a regression model based on Bi-LSTM and MLP was trained. Chord events are encoded using REMI, which includes type, time, pitch, velocity, duration, and chord information, and a complete representation is formed through the concatenation of embedding vectors, thereby effectively characterizing the emotional features of chords.
[0097] In the sequence modeling stage, a bidirectional LSTM (Bi-LSTM) network is used to encode the chord progressions. The Bi-LSTM captures the forward dependencies from the beginning to the end of the sequence through the forward LSTM, and at the same time captures the backward dependencies from the end to the beginning of the sequence through the backward LSTM. The hidden state at each time step is concatenated between the forward and backward directions to generate a global feature representation that combines the front and back information. To generate a fixed-length embedding representation of the chord progression, the hidden state of the last time step of the Bi-LSTM network is selected as the result of the pooling operation, obtaining a 512-dimensional global feature representation of the chord progression to ensure feature integrity and sequence consistency.
[0098] The 512-dimensional feature vector output by the Bi-LSTM is input into the MLP to predict the Valence and Arousal coordinates. The MLP contains hidden layers with 256, 128, and 64 neurons, uses ReLU activation, and incorporates 20% Dropout and L2 regularization, finally outputting two emotional dimensions. Through the joint design of the Bi-LSTM and the MLP, the model effectively captures the emotional dynamics in the chord progression, optimizes the chord features to match the experimental data, and lays a foundation for the joint emotional embedding space and background music generation.
[0099] When training the video-background music joint embedding space, the semantic and emotional features of the key frames are fused. A 256-dimensional visual semantic vector is generated through the attention mechanism using CLIP and color features, and the 512-dimensional CLIP feature is concatenated to form a 768-dimensional vector for V-A coordinate prediction. At the same time, the 512-dimensional chord feature output by the Bi-LSTM ensures the consistency of visual and chord emotions, thus supporting a high degree of emotional and semantic matching between the background music and the video content.
[0100] In video background music generation, the effective fusion of visual and auditory features is crucial. The joint representation not only improves the control accuracy in the generation stage but also provides a quantitative basis for the evaluation of emotional consistency. This scheme proposes a fine-grained modeling method, which conducts joint embedding training of audition and vision for each chord progression and its corresponding key frame, surpassing the traditional extraction of the whole music features that rely on pre-trained models. By constructing a shared joint embedding space, the visual features of the video key frames are aligned with the features of the chord progression, achieving efficient emotional matching between the background music and the video, and providing a solid foundation for the generation task and emotional alignment evaluation. In terms of time division, a chord progression does not strictly correspond to a measure, but may span multiple measures. After ensuring structural integrity and temporal coherence, a 3-5-measure segment can be divided into 2-3 chord progressions, each about 3 seconds. The key frames within the time range of each chord progression are positive samples, and the key frames of different progressions are negative samples. For example, within [Ti, Ti+3) of Chord_i, the key frames frame_j, frame_j+1, and frame_j+2 form three positive sample pairs with Chord_i.
[0101] In terms of feature extraction, the video key frames are obtained by concatenating the CLIP features fine-tuned by the predicted V-A emotion coordinates task that fuses scene semantic features and color features with the 256-dimensional output of the Transformer based on cross-attention feature fusion, resulting in a 768-dimensional key frame feature encoding output. The chord progression features are formed by concatenating the 512-dimensional chord progression features extracted by the Bi-LSTM network used to align chords with the emotion coordinates based on V-A coordinates and the 256-dimensional features output by the cross-attention fusion Transformer, forming a 768-dimensional chord progression feature representation that fuses scene semantics and scene emotions on the basis of chord features.
[0102] Through the contrastive learning objective function InfoNCE Loss, the 768-dimensional visual features of video key frames and the 768-dimensional chord features of chord progressions are jointly trained. In the joint embedding space, the key frame and chord progression feature vectors within the same time range are made closer, while the feature vectors in different time ranges are made farther apart.
[0103] The contrastive learning between video key frames and chord progressions needs to consider multiple pairs of diverse sample relationships: Suppose there are a total of M chord progressions {Chord_1, …, Chord_M} in the background music of a video, and each chord progression corresponds to several video key frames. Then the InfoNCE loss is calculated based on the following formula (5):
[0104]
[0105] where, represents the feature vector of the j-th key frame, represents the embedding vector of the i-th chord progression, S i represents the set of video key frames within the time range of the i-th chord progression. τ is a temperature parameter used to control the smoothness of the contrastive distribution and can be set and adjusted according to actual needs. M represents the total number of all chord progressions included in the background music of the entire video. In addition, in the loss function of contrastive learning, the base of the logarithm does not affect the final result of the optimization objective because the change in the logarithm base only causes a scaling of a constant factor, and the optimization of model parameters is not affected by this scaling. Considering the simplicity of the natural logarithm and its conformity with many mathematical derivation habits, and its consistency with the implementation of most optimization frameworks, the natural logarithm is used in this embodiment, that is, the base of log is e.
[0106] represents calculating the cosine similarity between two feature vectors, which can be represented in the form of sim(x, y). Then there is the cosine similarity calculation method shown in the following formula (6):
[0107]
[0108] Among them, x and y represent a pair of parameters used for calculation, without any other meaning.
[0109] In this way, by traversing the chord progression time period, it is ensured that the embedded vectors of the video key frames and the corresponding chord progressions are close within the time range, and are far from each other between different time periods. On the basis of time range alignment, an explicit constraint on the V-A coordinates is added to ensure the emotional consistency between the video key frames and the chord progressions within the time range. As shown in the following formula (7):
[0110]
[0111] Among them, represents the V-A coordinate deviation between the vision (video key frame) and the chord progression in the entire video, and is also the deviation between the visual emotion and the chord emotion. represents the V-A coordinates of the j-th video key frame, represents the V-A coordinates of the chord progression in the i-th segment, S i represents the set of video key frames within the time range of the i-th chord progression. This constraint traverses each chord progression Chord_i and all video key frames frame_j within the time range, and penalizes the deviation of their V-A coordinates to ensure that the emotions are aligned in the same emotional space.
[0112] Combining the above two losses, the final loss is determined based on the following formula (8):
[0113]
[0114] Among them, λ represents the weight parameter, which is used to balance the semantic contrast learning and the emotional alignment goal, so as to achieve the dual optimization of the embedding space in the feature and emotional dimensions, and its value can be set and adjusted according to actual needs. represents the final combined loss function. Combining the above two loss functions, the final loss function is composed of: one is the InfoNCE loss between the chord progression and the video key frames aligned in the time range, which is used to optimize the similarity of the feature vectors between the video key frames corresponding to the chord progression; the other is the V-A coordinate penalty term based on the same emotion model, which is used to explicitly constrain the alignment degree between the video key frames and the chord progression in the emotional space. Specifically, the InfoNCE loss traverses each chord progression and all video key frames within its corresponding time range, pulls closer the embedded vectors of the positive sample pairs (that is, the key frames and chord progressions within the same time range), and pushes away the negative sample pairs (that is, the key frames and chord progressions between different time ranges). The definition of this loss is: within the time range of a chord progression Chord_i, for all video key frames frame_j belonging to this time period, calculate The similarity score is obtained, and the smoothness of the distribution is ensured through a normalization term, finally generating a contrastive loss term. To further improve the alignment of emotional expressions, a V-A coordinate penalty term is introduced to constrain the distance between the V-A coordinates of video key frames and chord progressions within the same time range using the second norm, penalizing their deviations. This term calculates the sum of squares of the V-A coordinate differences for all positive sample pairs, ensuring that the emotional representations of key frames and chord progressions within each time range are more consistent. In this way, the final combined loss is determined by combining the losses of the two parts, ensuring the time range alignment characteristics of video key frames and chord progressions, while imposing consistency constraints on them at both the semantic and emotional levels, laying a foundation for subsequent generation tasks.
[0115] Based on the above operations, a shared joint embedding space can be trained to map chord progressions and video key frames into the same V-A emotional space. This embedding space ensures that the background music of the same video and the key frames in its corresponding time period are close to each other in the space. Therefore, for any randomly intercepted segment of the video background music and the key frames collected within its time range, their emotional semantics should have a high cosine similarity in the joint embedding space.
[0116] In addition, a retrieval accuracy metric is designed for the V-A emotional space. The specific steps are as follows: First, the visual features of key frames from different videos are mapped into the joint embedding space through a visual encoder; then, the chord features of the background music segment of the target video are mapped into the same space through a chord encoder. Next, the similarity between the chord features and the key frame features is calculated. If the Top-K similar frames contain the key frames of the original video segment, it is considered a successful retrieval. This metric measures the emotional alignment and cross-modal retrieval performance between the background music chord progression and video key frames, and specifically evaluates the retrieval accuracy and emotional semantic alignment through multiple retrieval metrics.
[0117] The first retrieval metric is the precision at K (P@K) of the top K key frames: the proportion of the number of key frames from the original video within the same time range as the given music segment among the top K key frames in the retrieval results. The correctly matched key frames refer to those that, among the top K video key frames with the highest similarity retrieved, contain key frames that belong to the target video and are within the time range of the corresponding music segment. Only these key frames can be regarded as the correct results of the retrieval task, and other key frames from interfering videos or different time segments of the target video are not counted as correctly matched. The specific calculation method is shown in formula (9) below:
[0118]
[0119] During the retrieval process, the candidate video key-frame set contains frames from the target video and interfering videos. The retrieved Top-K key frames may all come from different videos, partially from different time periods of the target video and interfering videos, all from different time periods of the target video, or all from interfering videos without frames from the target video. Only the key frames from the target video and within the time range of the corresponding music segment among the Top-K are considered correct matches, and other key frames from interfering videos or different time periods are not included in the correct results.
[0120] The second retrieval metric is recall: the proportion of all key frames from the target video and within the time range of the corresponding music segment that are correctly retrieved, which measures whether the generated music can fully cover the information of the key frames within the time range of the corresponding background music segment in the target video. Its specific calculation method is shown in the following formula (10):
[0121]
[0122] Specifically, the design definition of recall is the ratio of the number of Top-K key frames retrieved from the candidate video key frames and within the target time range to the total number of actual key frames within that range for a given music segment in the joint embedding space. For example, if K = 10, and 2 out of the 10 retrieved frames are within the target range, while there are 5 frames in total within the target range, then the recall is 2 / 5 = 0.4. This design measures the coverage of the retrieval results for the key frames within the target time range. The key frames to be retrieved come from the original video and interfering videos.
[0123] The third retrieval metric is mean average precision (mAP): The target video is divided into N segments, each containing background music. By extracting the chord progression features of each music segment and retrieving the Top-K video key frames with the highest similarity for each of these N music segments in the joint embedding vector space. Finally, by calculating the average precision (AP) and further taking its mean (mAP), the retrieval performance of the model on all segments is measured.
[0124] Among them, for each music segment i, according to its retrieval results, the corresponding average precision AP is calculated through the following formula (11):
[0125]
[0126] Among them, P(k) is the precision up to the k-th retrieval result, and its calculation method is shown in the following formula (12):
[0127]
[0128] rel(k) is an indicator function. If the k-th retrieval result is a correct match, the value of rel(k) is 1; otherwise, it is 0. M i is the actual number of key frames within the target time range corresponding to music segment i.
[0129] mAP is the mean of the APs of all music segments, and its calculation method is shown in the following formula (13):
[0130]
[0131] where N represents the total number of music segments.
[0132] In the embodiments of this application, the music generation process is decoupled into three parts, and a corresponding video background music generation framework is set up. Specifically, a multi-stage music generation framework is designed to deeply integrate the features (i.e., key frame feature vectors) that fuse visual semantics and colors in the video frames extracted from the video with music generation, so as to achieve higher-quality background music generation. First, the key frame feature vectors that fuse multiple modal features in the video key frames are extracted. The fused features include semantic features (such as the emotion and semantic information extracted by CLIP), color features (the dominant color information based on clustering), etc. These features are finally integrated into a unified feature representation that fuses visual semantics and colors (i.e., the key frame feature vector) through the cross-attention mechanism, serving as the conditional control signal for music generation. Subsequently, the chord generation module (ChordTransformer) generates a chord sequence based on the fused features of the video, improves the modeling efficiency for long sequences through the linear attention mechanism, and at the same time ensures that the chord sequence is consistent with the emotion and rhythm of the video through cross-attention. Then, the melody generation module (Melody Transformer) further generates a melody sequence based on the chord sequence and video features, captures the complex relationship between the chord and video features through cross-attention, and ensures that the melody complements the chord and video emotion. Finally, the accompaniment generation module uses a diffusion model based on Polyffusion to generate the accompaniment part, encodes the chord and melody sequences into control signals through VAE, and dynamically maps them to the U-Net through cross-attention to affect the diffusion process to generate an accompaniment part consistent with the video features. Ultimately, the background music generated by the framework consists of chords, melody, and accompaniment, and can efficiently and accurately generate background music that matches the video content and has consistent emotion. At the same time, by aligning the emotions of video key frames and music events in the joint embedding space, and introducing multi-scale modeling and richer conditional signals, the diversity and adaptability of the generated music are further improved, providing a comprehensive solution for video background music generation.
[0133] To verify the effectiveness of the background music generation method, the embodiments of this application designed an objective and subjective evaluation scheme based on emotional matching degree, rhythm synchronization, and theme fit. The emotional matching degree is evaluated by calculating the cosine similarity between the generated music and the video in the joint embedding space; the rhythm synchronization analyzes the temporal alignment between the music rhythm structure and the video dynamic changes; the theme fit measures the relevance between the music melody and the video content. In addition, the subjective evaluation combines the consistency of emotion, rhythm, and theme, as well as the scores of visual style, narrative driving force, and overall visual impression to comprehensively verify the performance of the generation framework. Through the experimental results on multiple test sets, the retrieval precision (mAP) and recall rate (Recall) are statistically calculated, which respectively reflect the matching ability and coverage of the generated music and video segments in the embedding space. The experimental results show that the proposed framework outperforms the existing baseline methods in all metrics, especially showing significant improvements in emotional matching and rhythm synchronization, verifying the powerful ability of the model to capture the emotional and dynamic features of the video. This indicates that based on the generation model in the embodiments of this application, not only can high-quality music be generated, but also it can be closely combined with the video content at different levels, significantly enhancing the audio-visual experience.
[0134] Thus, the embodiments of this application propose a video background music generation method, which is mainly divided into four parts: video key extraction, cross-modal feature fusion, hierarchical music generation, and generation effect evaluation. In video key extraction, the visual dynamic features of the video are captured by combining frame difference, HSV color histogram, and sparse optical flow, and SSIM is used for deduplication and screening. In terms of cross-modal feature fusion, the CLIP model is used to extract the semantic information of key frames, and the joint features are constructed by combining the dominant color and brightness histogram. The deep fusion of video visual features and emotional features is realized through the cross-attention mechanism. The music generation part adopts a hierarchical modeling framework. First, the chord sequence is generated by Chord Transformer, and the chord generation is guided by the dynamic and semantic features of the video; then Melody Transformer is used to generate music by combining the chord sequence and video features. Finally, the generated chords and music are used as conditional inputs to the accompaniment generation module based on the diffusion model to complete the accompaniment generation. The entire music generation process adopts cross-modal mapping based on the V-A emotion model, and the emotional-driven multi-modal generation is realized by establishing the mapping relationships between chord-emotion and color-emotion. When evaluating the generation effect, some preliminary evaluation indicators such as partial search precision are redesigned based on video key frames - music segments. The similarity between the generated music and the video content is determined by calculating the cosine similarity between the chord progression features in the background music segment and the video key frame image features in the joint embedding space. The generation of background music with target emotion matching and multi-modal feature fusion efficiently realizes the fine control of the generation process and high-quality multi-modal interaction, providing a controllable and highly interpretable solution.
[0135] Specifically, to address the problems of low matching degree, uncontrollable generation, and low understanding efficiency in existing freestyle video background music generation, this solution uses cross-modal technology to fuse video and music information. The specific method includes using a key frame extraction algorithm based on deduplication to select target key frames representing the main content from the video, avoiding semantic redundancy. Each key frame represents a shot event and corresponds to a specific note event. By retaining metadata such as the global position, timestamp, and deduplication count of the key frames, visual dynamic features such as time interval, playback speed, and average optical flow change are calculated, and the dominant color is extracted using K-means clustering as the color feature. This method reduces redundant frames, accurately extracts visual and color features, improves the matching degree and generation efficiency between the video and the background music, and ensures that the generated music is highly consistent with the video content in terms of semantics, emotion, and color. In terms of semantic feature extraction, combining scene and emotion semantics, a large visual question answering model is used to generate natural language descriptions containing scenes and emotions for each key frame. This model is trained based on a large amount of image-text data by designing appropriate prompts and image inputs, and can capture rich scene and emotion information and output it in natural language. Subsequently, the encoder of CLIP is used to encode these descriptions into the text-image joint embedding space as semantic control signals to guide the generation process of chords and melodies in the background music, thus achieving consistent matching between music and video in terms of semantics and emotion. The proposed video background music generation framework uses Transformer to perform hierarchical modeling on chords and melodies. Chord Transformer is responsible for generating the chord event sequence, and Melody Transformer generates the melody event sequence. When performing sequence modeling, the visual dynamic, color, and semantic emotion features extracted from the video are fused and input into the encoder of Chord Transformer. The decoder generates new chords based on the previous chord sequence and the fused features. The generated chord sequence is further fused with the visual dynamic and emotion features and input into the encoder of Melody Transformer as the context for melody generation. The decoder generates a new melody based on the previous melody sequence. The time steps of the chords and melodies are unified, the chords are decomposed into individual notes and merged with the melody notes to avoid repetition. Pretty_midi is used to convert the merged note sequence into a piano-roll, which is used as the input of the Polyffusion pre-model to generate accompanying notes that match the chords and melodies, and the final background music is output. In the alignment evaluation of the background music and video content, this solution continues the idea of SymMV based on retrieval accuracy, extends the CLIP model to the video-music field, and proposes the Video Music CLIP Precision (VMCP) as a new evaluation metric to measure the correspondence between the video and the music. Due to the use of the deduplication key frame extraction algorithm, the video consists of several key frames, semantic compression is performed, and the motion details between consecutive frames are lost.Experiments show that these motion details are not helpful for background music generation and may even interfere with the generation quality. Overly detailed motion modeling can cause the background music rhythm to deviate from the traditional trend. Each video key frame xi is directly encoded using CLIP to obtain zi (a 512-dimensional feature vector), and then sorted according to the timestamp to obtain [z1, z2, z3, ..., zN]. Here, it is assumed that a total of N key frames are intercepted, so that the key frame feature matrix Z (N*512) consistent with the actual story progression time sequence in the video is obtained. The video background music is divided into bars. For background music with a 4 / 4 time signature, each bar lasts for 2s, each quarter note lasts for 0.5 seconds, and each bar has four quarter notes. The key frame is intercepted and divided every 2 music bars (i.e. every 4s). The division follows the principle that as long as the key frame timestamp falls within the bar time range, it is counted into the bar, so that there may be overlapping video key frames in every two bars. The Transformer-based music tagging model is also used to encode the background music of these two measures, and a positive sample pair is formed with the feature vector Zm (m*512) obtained after CLIP encoding of the m video key frames in the corresponding time range, i.e. (music_i, Zm), where i represents the i-th group of music clips (each clip consists of two measures), and m represents the video key frames included in the corresponding clip.
[0136] Then, a joint embedding space of music clips to video keyframes is obtained by training with the InfoNCE contrast loss as the basis for evaluating the retrieval accuracy index. For the newly generated background music, the alignment is evaluated by calculating the vector similarity between the original video keyframe and the corresponding segment of the generated background music based on the cosine similarity in the joint embedding space of music clips to video keyframes.
[0137] In this way, the video content is effectively compressed using a deduplication-based video key frame extraction algorithm, and the music generation modeling method of cross-modal feature fusion technology and a layered Transformer combined with a diffusion model is used to solve the problems of low matching degree, poor interpretability of the generation process, low controllability, and low feature extraction efficiency in the existing video background music generation technology, thereby achieving high matching degree, high controllability, and high efficiency of video background music generation.
[0138] It should be noted that in terms of video feature extraction, advanced models such as ViT or 3D CNN can be adopted to enhance the ability to model the spatio-temporal dynamics of video content. In the chord generation module, emotional constraints can be introduced or sequence models such as BiLSTM can be used to improve the flexibility and emotional consistency of chord generation. In terms of cross-modal feature fusion, the cross-attention mechanism is optimized through contrastive learning and multi-layer fusion to enhance the alignment effect between visual and music features. For accompaniment generation, GAN or Flow models can be used to improve the generation efficiency and detail control ability. The emotional alignment method can be extended to dynamic emotion or the PAD three-dimensional model to enrich the emotional expression of music. In terms of evaluation metrics, personalized evaluation based on user preferences or GAN authenticity scoring can be introduced to optimize the music quality evaluation. The system architecture can support large-scale real-time generation through cloud deployment and implement modular design, flexibly select generation modules, and expand to application scenarios such as multi-language video translation or advertisement sound effect generation.
[0139] Furthermore, by constructing and optimizing the multi-modal dataset, a rich and diverse training basis can be provided. Expand the types of video content and music styles to ensure that the model adapts to diverse needs. Combine audio and video feature alignment modeling and add text descriptions to enhance cross-modal mapping and emotional control. Improve the time resolution, refine chord annotation, and enhance the ability to capture rhythm and time correlation. Increase multi-dimensional emotion model annotation and improve the annotation quality through user feedback. Supplement low-resource or specific domain samples to enhance the performance of the model in few-shot and cross-domain generation. Establish a dynamic data update mechanism, regularly expand the content, open part of the dataset for community use, and obtain more annotations and improvement suggestions through crowdsourcing. Through multi-dimensional optimization of the dataset, the adaptability, emotional expression, and generation quality of the model are improved, providing a solid data foundation for multi-modal video background music generation.
[0140] As Figure 9 shown, corresponding to the above video background music generation method, an embodiment of the present application also provides a video background music generation system, and the above video background music generation system includes:
[0141] A video acquisition module 910, configured to acquire a video to be processed;
[0142] A key frame selection module 920, configured to determine a plurality of candidate key frames from the above video to be processed;
[0143] A key frame deduplication module 930, configured to perform deduplication processing on the above candidate key frames according to the structural similarity between the above candidate key frames to obtain target key frames;
[0144] A music generation module 940, configured to generate background music matching the above video to be processed according to the above target key frames.
[0145] In this way, first determine multiple candidate key frames in the video to be processed. For the candidate key frames, further perform duplicate removal processing based on the structural similarity between the candidate key frames, so as to remove the video frames with duplicate information and obtain the target key frames, and then perform the subsequent background music generation process according to the target key frames. In this way, the candidate key frames with duplicate information can be removed through duplicate removal processing, avoiding redundant calculations and facilitating the improvement of the processing efficiency of the background music generation process.
[0146] It should be noted that the specific structures and implementation manners of the above video background music generation system and its various modules or units can refer to the corresponding descriptions in the above method embodiments, and will not be elaborated here. The division methods of the various modules of the above video background music generation system are not unique and are not specifically limited here.
[0147] Based on the above embodiments, the present application also provides an intelligent terminal, and its principle block diagram can be as Figure 10 shown. The above intelligent terminal includes a processor, a memory, a network interface, and a display screen connected through a system bus. Among them, the processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a video background music generation program. The internal memory provides an environment for the operation of the operating system and the video background music generation program in the non-volatile storage medium. The network interface of the intelligent terminal is used to communicate with an external terminal through a network connection. When the video background music generation program is executed by the processor, it implements the steps of any of the above video background music generation methods. The display screen of the intelligent terminal can be a liquid crystal display screen or an electronic ink display screen.
[0148] Those skilled in the art can understand that Figure 10 the principle block diagram shown in
[0149] merely shows the block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the intelligent terminal to which the solution of the present application is applied. The specific intelligent terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0150] The embodiment of the present application further provides a computer-readable storage medium, on which a video background music generation program is stored. When the video background music generation program is executed by a processor, the steps of any one of the video background music generation methods provided by the embodiment of the present application are implemented.
[0151] It should be understood that the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
Claims
1. A method for generating background music for a video, characterized in that: The method comprises: Get the video to be processed; Determining a plurality of candidate key frames from the video to be processed; According to the structural similarities between the candidate key frames, the candidate key frames are deduplicated to obtain a target key frame; Based on the target key frame, background music matching the video to be processed is generated.
2. The method for generating video background music according to claim 1, characterized in that: The step of determining a plurality of candidate key frames from the video to be processed comprises: Taking the first video frame of the video to be processed as a candidate key frame; For each video frame except the first video frame in the video to be processed, a candidate key frame that is before the video frame and closest to the video frame is used as a comparison key frame corresponding to the video frame, and comparison result data between the video frame and the comparison key frame is obtained. If the comparison result data meets a preset threshold condition, the video frame is used as a candidate key frame; The comparison result data includes at least one of inter-frame difference data, color distribution comparison data and sparse optical flow detection data.
3. The method for generating video background music according to claim 1, characterized in that: The step of performing deduplication processing on the candidate key frames according to the structural similarities between the candidate key frames to obtain the target key frame includes: Sorting the candidate key frames according to the timestamps corresponding to the candidate key frames and obtaining a key frame list; Taking the first candidate key frame in the key frame list as the target key frame; For each candidate key frame except the first candidate key frame in the key frame list, a target key frame that is before the candidate key frame and closest to the candidate key frame is used as a target comparison frame, and the structural similarity between the candidate key frame and the target comparison frame is calculated. If the structural similarity is not greater than a preset structural similarity threshold, the candidate key frame is used as the target key frame.
4. The method for generating video background music according to claim 1, characterized in that: The step of generating background music matching the video to be processed according to the target key frame includes: Performing feature extraction on the target key frame to obtain a key frame feature vector, wherein the key frame feature vector is used to characterize semantic features and color features corresponding to the target key frame; Based on the key frame feature vector, background music matching the video to be processed is generated.
5. The method for generating video background music according to claim 4, characterized in that: The semantic features include semantic information features and emotional information features; The step of extracting features from the target key frame to obtain a key frame feature vector includes: Encoding the target key frame through a text-to-image joint embedding pre-trained model to obtain a scene semantic vector; Performing hue extraction on the target key frame to obtain a hue vector, performing brightness distribution analysis on the target key frame to obtain a brightness vector, and performing vector splicing on the hue vector and the brightness vector to obtain a color vector; Fusing the scene semantic vector and the color vector to obtain a fused vector; The scene semantic vector, the fusion vector and the color vector are concatenated to obtain the key frame feature vector.
6. The method for generating video background music according to claim 4 or 5, characterized in that: The step of generating background music matching the video to be processed according to the key frame feature vector comprises: Generating a chord sequence matching the video to be processed according to the key frame feature vector; Background music matching the video to be processed is generated according to the key frame feature vector and the chord sequence.
7. The method for generating video background music according to claim 6, characterized in that: The step of generating background music matching the video to be processed according to the key frame feature vector and the chord sequence includes: Generate a melody sequence according to the key frame feature vector and the chord sequence; Generate an accompaniment sequence according to the chord sequence and the melody sequence; Background music matching the video to be processed is obtained according to the chord sequence, the melody sequence and the accompaniment sequence.
8. A video background music generation system, characterized in that: The system comprises: A video acquisition module is used to acquire the video to be processed; A key frame selection module, used to determine a plurality of candidate key frames from the video to be processed; A key frame deduplication module, used for deduplicating the candidate key frames according to the structural similarities between the candidate key frames to obtain a target key frame; The music generation module is used to generate background music matching the video to be processed according to the target key frame.
9. An intelligent terminal, characterized in that: The intelligent terminal includes a memory, a processor, and a video background music generation program stored in the memory and executable on the processor. When the video background music generation program is executed by the processor, the steps of the video background music generation method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a video background music generation program, and when the video background music generation program is executed by the processor, the steps of the video background music generation method as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Automatic video editing method and device based on text and shot similarity, and terminal
CN120499445A
Method and system for generating video and dubby music on basis of multi-model collaboration
CN121078293A
Video intelligent dubbling method and device, electronic equipment and computer storage medium
CN122205185A