Method and system for generating video and dubby music on basis of multi-model collaboration

By employing a multi-model collaborative approach, combining image visual features and video keyframe sentiment analysis, background music that highly matches the video content is generated. This solves the problem of low matching degree between emotion and background music in existing technologies, and improves the efficiency and quality of content creation.

CN121078293APending Publication Date: 2025-12-05CHONGQING LINGJIFANG TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511596086.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

In existing technologies, the matching degree between emotion and background music is low during the image-to-video conversion process. It is impossible to refine the background music parameters by combining the specific visual features of the video keyframes, resulting in a disconnect between the generated background music and the emotional expression of the video.

Method used

A multi-model collaborative approach is adopted to acquire image data, extract visual features to generate structured descriptions, extract video keyframes and obtain emotion categories and confidence levels, fuse color emotion data to generate emotion feature vectors for keyframes, summarize the emotion proportions to determine the dominant emotion, query the emotion-music feature mapping library to obtain background music parameters, and adjust the music generation to be synchronized with the video.

Benefits of technology

It improves the matching degree between emotions and background music, and the generated background music is highly integrated with the video content, thereby enhancing the efficiency and quality of content creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121078293A_ABST
    Figure CN121078293A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video processing, in particular to a method and a system for generating a video and a dubby music based on a multi-model collaboration map. The method comprises the steps of obtaining picture data to generate structured description, and generating a current video based on scene logic and an action sequence; extracting a current video key frame, extracting a visual feature of the key frame, obtaining an emotion category and confidence corresponding to each key frame, and fusing the color emotion data to generate an emotion feature vector of the key frame; summarizing sentiment classification results of all the key frames, counting the proportion of each type of sentiment in the video, preliminarily screening a matched music style range matched with the dominant sentiment, and obtaining corresponding matched music parameters according to the determined dominant sentiment; the system comprises a structured description generation module, a video key frame emotion feature extraction module and a score matching generation module. Through the above mode, the matching degree of emotion and incidental music is improved, the bottleneck in the prior art is broken through, and the ever-increasing integrated content creation requirement is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video processing technology, and in particular to a method and system for generating image-generated video and background music based on multi-model collaboration. Background Technology

[0002] With the rapid development of short video platforms, digital marketing, and online education, users' demand for integrated content creation services that combine image-to-video conversion with intelligent background music is growing. Currently, the process of generating videos from images and matching them with corresponding music suffers from low emotional and musical compatibility, severely hindering the efficiency and quality of content creation. On the one hand, most music generation methods can only generate music based on general emotional tags such as "cheerful" and "sad", and cannot refine the music parameters by combining the specific visual features of video keyframes. On the other hand, even if some technologies can extract emotions from videos, they have not established a precise correspondence between emotions and the core parameters of music, resulting in a disconnect between the generated background music and the emotional expression of the video. Therefore, it is essential to propose a method and system for generating images and music from images to improve the matching degree between emotions and background music. Summary of the Invention

[0003] The purpose of this invention is to provide a method and system for image-generated video and background music based on multi-model collaboration, aiming to solve the technical problem of low matching degree between existing emotions and background music.

[0004] To achieve the above objectives, the present invention employs a method for image-generated video and background music based on multi-model collaboration, comprising the following steps: Acquire image data, extract visual features from the images to generate structured descriptions, and generate the current video based on the scene logic and action sequences in the structured descriptions; Extract keyframes from the current video, extract keyframe visual features, and obtain the sentiment category and confidence level corresponding to each keyframe. Combine sentiment category, confidence level, and color sentiment data from visual features to generate sentiment feature vectors for keyframes. The emotion classification results of all keyframes are summarized, the proportion of each emotion in the video is calculated, the dominant emotion of the video is determined based on the proportion, the range of background music styles that match the dominant emotion is initially screened, the preset emotion-music feature mapping library is queried, and the corresponding background music parameters are obtained based on the determined dominant emotion.

[0005] In the steps of acquiring image data, extracting visual features from the images to generate a structured description, and generating the current video based on the scene logic and action sequence in the structured description: Receive images uploaded by users and standardize their format. Visual features are extracted from the standardized images to obtain feature data containing information on image color distribution, object outlines, and scene type. Generate an initial text description of the image content based on the feature data; A pre-defined structured template is used to optimize the initial text and generate a structured description. The structured template includes three dimensions: timeline, element logic, and emotional tone. The current video is generated based on the scene logic and action sequence in the structured description.

[0006] In the steps of extracting keyframes from the current video, extracting keyframe visual features, obtaining the sentiment category and confidence level corresponding to each keyframe, and fusing the sentiment category, confidence level, and color sentiment data from the visual features to generate the sentiment feature vector of the keyframe: Perform frame parsing on the current video to obtain the total number of frames and frame rate information, and mark candidate keyframes based on the pixel difference between two adjacent frames; All candidate keyframes are grouped and processed. The final keyframes are determined based on the video duration. Candidate frames with duplicate content in each group are removed to obtain the final video keyframe sequence. Visual feature data is obtained based on the keyframe sequence; the visual feature data includes keyframe color sentiment, scene sentiment, and object sentiment data.

[0007] Among them, in the steps of parsing the current video to obtain the total number of frames and frame rate information, and marking candidate keyframes based on the pixel difference between two adjacent frames: Set a threshold, compare the difference value with the threshold, and when the difference value exceeds the preset threshold, mark the frame as a candidate keyframe.

[0008] After performing frame parsing on the current video to obtain the total number of frames and frame rate information, and marking candidate keyframes based on the pixel difference between adjacent frames: Perform region identification on each frame of the video and extract the core region in the frame.

[0009] In the step of performing region identification on each frame of the video and extracting the core region in the frame: Set a ratio, compare the rate of change of the region in the current frame with that in the previous frame, compare the rate of change and the ratio, and when the rate of change exceeds the set ratio, mark the frame as a candidate keyframe.

[0010] After the step of obtaining visual feature data based on the keyframe order, whereby the visual feature data includes keyframe color sentiment, scene sentiment, and object sentiment data: The sentiment category and confidence level of each keyframe are obtained, and the sentiment category, confidence level, and color sentiment data in visual features are fused to generate the sentiment feature vector of the keyframe.

[0011] The steps include: summarizing the emotion classification results of all keyframes, calculating the proportion of each emotion in the video, determining the dominant emotion of the video based on the proportion, initially screening the range of background music styles that match the dominant emotion, querying the preset emotion-music feature mapping library, and obtaining the corresponding background music parameters based on the determined dominant emotion. Based on the determined dominant emotion, the corresponding background music parameters are obtained from the mapping library; the background music parameters include music rhythm, key, core instrument type and chord progression.

[0012] The process involves retrieving the corresponding background music parameters from a mapping library based on the determined dominant emotion; these parameters include musical rhythm, key, core instrument type, and chord progression. Adjust the music generation duration to match the video duration, generate initial background music audio that meets the emotional tone and duration requirements based on the parameters, and obtain the timestamp information corresponding to the video keyframes; Align the audio and video timestamps, analyze the rhythmic accents of the initial soundtrack, match the rhythmic accents of the soundtrack with the switching timestamps of the video keyframes, synchronize the soundtrack with the keyframe images, and generate the final soundtrack.

[0013] This invention also provides a system for graph-generated video and background music based on multi-model collaboration, including a structured description generation module, a video keyframe emotion feature extraction module, and a background music matching and generation module; wherein: The structured description generation module is used to acquire image data, extract visual features of the image to generate a structured description, and generate the current video based on the scene logic and action sequence in the structured description. The video keyframe emotion feature extraction module is used to extract the current video keyframe, extract the visual features of the keyframe, and obtain the emotion category and confidence level corresponding to each keyframe. It then merges the emotion category, confidence level, and color emotion data in the visual features to generate the emotion feature vector of the keyframe. The music matching and generation module is used to summarize the emotion classification results of all keyframes, count the proportion of each emotion in the video, determine the dominant emotion of the video based on the proportion, initially filter the range of music styles that match the dominant emotion, query the preset emotion-music feature mapping library, and obtain the corresponding music parameters based on the determined dominant emotion.

[0014] This invention discloses a method and system for image-generated video and background music based on multi-model collaboration. The method comprises a structured description generation module, a video keyframe emotion feature extraction module, and a background music matching generation module, performing the following steps: acquiring image data, extracting visual features from the images to generate a structured description, and generating the current video based on the scene logic and action sequence in the structured description; extracting keyframes from the current video, extracting keyframe visual features, and obtaining the emotion category and confidence level corresponding to each keyframe; fusing emotion category, confidence level, and color emotion data from the visual features to generate an emotion feature vector for the keyframes; summarizing the emotion classification results of all keyframes, statistically analyzing the proportion of each emotion in the video, determining the dominant emotion of the video based on the proportion, initially screening the range of background music styles matching the dominant emotion, querying a preset emotion-music feature mapping library, and obtaining the corresponding background music parameters based on the determined dominant emotion. Through the above methods, the matching degree between emotion and background music is improved, overcoming existing technical bottlenecks and meeting the growing demand for integrated content creation. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a flowchart of the steps of the method for generating image-based video and background music based on multi-model collaboration according to the present invention.

[0017] Figure 2 This is a flowchart of steps S100 of the present invention.

[0018] Figure 3 This is a flowchart of steps S200 of the present invention.

[0019] Figure 4 This is a flowchart of steps S300 of the present invention.

[0020] Figure 5 This is a schematic diagram of the system for generating video and background music based on multi-model collaboration according to the present invention.

[0021] Figure 6 This is a schematic diagram of the electronic device of the present invention.

[0022] 401 - Structured Description Generation Module, 402 - Video Keyframe Emotional Feature Extraction Module, 403 - Music Matching Generation Module. Detailed Implementation

[0023] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.

[0024] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0025] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0026] Please see Figures 1-4 This invention provides a method for generating image-based videos and background music based on multi-model collaboration, comprising the following steps: S100: Acquire image data, extract visual features from the images to generate a structured description, and generate the current video based on the scene logic and action sequence in the structured description.

[0027] In this embodiment, image data is acquired, visual features of the images are extracted to generate a structured description, and the current video is generated based on the scene logic and action sequence in the structured description. The specific process is as follows: S101: Receive images uploaded by users and perform format standardization processing on the images; S102: Extract visual features from the standardized image to obtain feature data containing information on image color distribution, object outlines, and scene type. S103: Generate an initial text description of the image content based on the feature data; S104: Preset structured template, optimize the initial text based on the structured template to generate a structured description; the structured template includes three dimensions: timeline, element logic, and emotional tone; S105: Generate the current video based on the scene logic and action sequence in the structured description.

[0028] In the above process, when a user uploads an image, the image format needs to be standardized to ensure smooth subsequent processing. Using image processing tools, such as the Pillow library in Python, the uploaded image is first opened using the Pillow library's `Image.open()` function to obtain its format information. If the image format does not conform to the system's preset standard format (such as common JPEG, PNG, etc.), the `convert()` method is used to convert it to the target format. For example, if a BMP format image is uploaded, it is converted to an RGB mode image for subsequent processing. Simultaneously, the image's resolution, size, etc., are adjusted to meet the requirements. For standardized images, the color histogram method can be used to extract color distribution features. First, the image is converted from the RGB color space to the HSV color space, and the histograms of the three HSV channels are calculated. The histogram of each channel can be divided into several bins (e.g., 16 bins). By counting the number of pixels in each bin, the color distribution features of the image are obtained. These features can reflect the proportion of different colors in the image. For object contour feature extraction, the image is first Gaussian smoothed to reduce the impact of noise. Next, the gradient intensity and direction are calculated, and non-maximum suppression is used to refine edges and remove false edges. Finally, double threshold detection is used to determine the true edges, obtaining the object's contour information. In terms of deep learning techniques, convolutional neural networks (CNNs), such as the classic VGG16 model, can be used. The standardized image is input into the VGG16 model, where the convolutional layers automatically learn various features in the image, including object texture and shape information. Through multiple convolutional and pooling operations, high-level visual features are extracted. For scene type information extraction, pre-trained scene classification models, such as ResNet-based scene classification models, can be used. When an image is input into this model, it determines the scene type (e.g., indoor, outdoor, natural) based on the scene feature patterns it has learned, and outputs the corresponding scene type label. Generating initial text descriptions is the process of converting extracted visual features into natural language descriptions, primarily achieved using neural network models. Taking a model based on Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTMs) as an example, the extracted visual feature data is first input into the CNN for further feature learning and abstraction. CNNs' powerful feature extraction capabilities can transform image features into more representative feature vectors. These feature vectors are then input into the LSTM. In the LSTM, a series of gating mechanisms (input gate, forget gate, output gate) control the flow and retention of information, gradually generating initial text describing the image content based on the input visual feature vectors. For example, for an image containing blue sky, white clouds, and grass, the model might generate an initial text description such as "The image shows a vast grassland with blue sky and white clouds above." To make the description clearer and more organized, the initial text needs to be optimized based on a preset structured template to generate a structured description. The preset structured template includes three dimensions: timeline, element logic, and emotional tone. In the timeline dimension, if the image contains elements with a chronological order, such as an image of sunrise, the time points or chronological order of each stage will be clearly described, such as "In the early morning, the sun begins to rise from the horizon, gradually illuminating the sky and the earth." From the perspective of element logic, the various elements in the image will be sorted out to clarify the relationships between them. For example, for an image containing people flying kites in a park, it would be described as "In the park, the person holds the kite string, the kite is floating in the sky, and the person and the kite are connected by the kite string," clearly showing the positional relationship and action logic between the person, the kite, and the park. In terms of emotional tone, the emotional tone is determined based on the overall atmosphere and visual characteristics of the image. If it's an image showcasing beautiful scenery, the emotional tone is likely to be positive and joyful, and descriptions will use words like "beautiful" and "breathtaking," such as "This is a breathtakingly beautiful landscape, with green mountains and clear waters complementing each other, making one feel intoxicated." Conversely, if it's an image depicting a disaster scene, the emotional tone is likely to be negative and somber, and descriptions will use words like "tragic" and "heartbreaking." Optimizing these three dimensions makes the initial text more structured and logical, facilitating subsequent video generation. The current video is generated based on the scene logic and action sequence in the structured description. First, the structured description is parsed to extract scene and action sequence information. For example, if the structured description is "On a sunny beach, a boy runs on the sand and then jumps into the sea to swim," the scene is identified as a beach, and the action sequence is the boy running on the sand and jumping into the sea to swim. Then, the video is generated using a video library or a video generation model based on Generative Adversarial Networks (GANs). If a video library is used, relevant video clips are selected from the library based on the scene and action information, such as clips containing beach scenes, clips of the boy running and swimming. Next, these clips are spliced ​​and edited according to the chronological order of the action sequence, adjusting the clip length, transition effects, etc., to make the video play smoothly and naturally. If a GAN-based video generation model is used, the generator produces a corresponding video frame sequence based on the information in the structured description, then combines these frames into a complete video, and a discriminator continuously optimizes the generated video to ensure it conforms to the described scene logic and action sequence, ultimately generating the current video.

[0029] S200: Extract keyframes from the current video, extract keyframe visual features, and obtain the sentiment category and confidence level corresponding to each keyframe. Combine the sentiment category, confidence level, and color sentiment data from the visual features to generate the sentiment feature vector of the keyframe.

[0030] In this embodiment, keyframes of the current video are extracted, their visual features are extracted, and the sentiment category and confidence level corresponding to each keyframe are obtained. The sentiment category, confidence level, and color sentiment data from the visual features are then fused to generate the sentiment feature vector of the keyframe. The specific process is as follows: S201: Perform frame parsing on the current video to obtain the total number of frames and frame rate information. Mark candidate keyframes based on the pixel difference between two adjacent frames. Set a threshold and compare the difference value with the threshold. When the difference value exceeds the preset threshold, mark the frame as a candidate keyframe. S202: Perform region identification on each frame of the video and extract the core region in the frame; set a ratio, compare the region change rate of the current frame with that of the previous frame, compare the change rate and the ratio, and when the change rate exceeds the set ratio, mark the frame as a candidate keyframe. S203: Group all candidate keyframes, determine the final keyframes based on the video duration, remove candidate frames with duplicate content in each group, and obtain the final video keyframe sequence. S204: Obtain visual feature data based on the keyframe sequence; the visual feature data includes keyframe color sentiment, scene sentiment, and object sentiment data. S205: Obtain the sentiment category and confidence level corresponding to each keyframe, and fuse the sentiment category, confidence level, and color sentiment data in visual features to generate the sentiment feature vector of the keyframe.

[0031] In the above process, when parsing the current video frame, the two adjacent frames are first converted into grayscale images, the difference between corresponding pixels in the two grayscale images is calculated, and the difference image is binarized by setting a threshold. When there are non-zero pixels in the binarized image, it indicates that the pixel difference between the current frame and the previous frame is large, which may be an important change in the picture. At this time, the frame is marked as a candidate keyframe. Besides pixel differences, region change rate is also a crucial criterion for labeling candidate keyframes. Region identification can be performed on each frame of a video using object detection algorithms, such as the YOLO (YouOnlyLookOnce) series of deep learning algorithms. Taking YOLOv5 as an example, pre-trained model weights are first loaded, and then the video frame is input into the model. The model outputs information such as the object category and location in the frame, thereby extracting the core region of the frame. Next, a ratio is set to measure the degree of region change. The region change rate between the current frame and the previous frame is calculated; for example, the intersection-union ratio (IoU) of two regions is used to measure the degree of region overlap. If the IoU is less than the set ratio, it indicates a significant regional change between the current and previous frames, and this frame is then added as a candidate keyframe. This labeling method based on region change rate can effectively capture important changes such as object movement and scene transitions in the video, complementing the pixel difference-based method and improving the accuracy and comprehensiveness of candidate keyframe labeling. After obtaining all candidate keyframes, they need to be grouped to determine the final keyframes. The grouping is primarily based on the video duration. Assuming the video duration is T and the frame rate is fps, the video is divided into time intervals, each interval being t (t can be set according to actual needs, such as 0.5 seconds). Candidate keyframes within each time interval are then grouped together. For example, for a 10-second video with a frame rate of 30fps, if the time interval t is set to 0.5 seconds, it can be divided into 20 time intervals, each containing a maximum of 15 frames (0.5 * 30). The candidate keyframes within these frames are then grouped together. After grouping, candidate frames with duplicate content in each group are removed. Content duplication is determined by calculating the similarity between candidate keyframes. The Structural Similarity Index (SSIM) method can be used to calculate image similarity. First, the candidate keyframe images are normalized to have the same size and brightness range. Then, the SSIM value of two candidate keyframe images is calculated. The SSIM value ranges from -1 to 1; the closer the value is to 1, the more similar the two images are. A similarity threshold is set; when the SSIM value of two candidate keyframes exceeds this threshold, they are considered to have duplicate content, one is retained, and the other is removed. In this way, only the most representative candidate keyframes are retained in each group, resulting in the final video keyframe sequence. This keyframe sequence covers the important content changes in the video while avoiding interference from too many duplicate frames, effectively representing the core information of the video and providing a concise and crucial data foundation for subsequent video analysis and processing. Extracting visual feature data from keyframes is a crucial step in gaining a deeper understanding of video content. This feature data includes keyframe color sentiment, scene sentiment, and object sentiment data. For color sentiment feature extraction, the keyframe images are first converted from the common RGB color space to the HSV color space, which is more in line with human perception. Color distribution features are obtained by counting the number of pixels in each bin. Then, using some pre-established color sentiment mapping models, the color distribution features are mapped to sentiment categories. For example, images with more warm tones (such as red and orange) may be mapped to positive and enthusiastic emotions; images with more cool tones (such as blue and green) may be mapped to calm and peaceful emotions. For scene sentiment feature extraction, a pre-trained deep learning model for scene classification, such as a scene classification model based on the ResNet architecture, is used. Keyframe images are input into the model, which analyzes and judges the scene elements in the image, outputting the scene type to which the keyframe belongs, such as indoor scene, outdoor scene, natural scene, urban scene, etc. Then, based on the association between different scene types and emotions, the scene sentiment is determined. For example, natural scenes are often associated with tranquility and comfort; bustling urban scenes may be associated with vitality and excitement. For object sentiment extraction, object detection algorithms (such as Faster R-CNN) are first used to detect objects in keyframes. The Faster R-CNN model generates object category labels and location information. Then, for each detected object, a pre-trained object sentiment classification model is used to determine the emotion it conveys. For example, if a smiling face is detected, the model might classify it as expressing happiness or joy; if a broken object is detected, it might be classified as expressing sadness or loss. Through these methods, rich visual feature data is comprehensively obtained from keyframes, providing strong support for subsequent sentiment analysis and video processing. Generating sentiment feature vectors is the process of fusing various sentiment information from keyframes, which is crucial for a comprehensive understanding of the emotional content of a video. First, a sentiment classification model is used to obtain the sentiment category and confidence score for each keyframe. This model can be a multi-classification model based on a convolutional neural network (CNN), trained on a large amount of sentiment-annotated image data. When a keyframe image is input into this model, it outputs predicted probabilities for multiple sentiment categories (such as happiness, sadness, anger, fear, surprise, disgust, etc.). The category with the highest probability value is the sentiment category of that keyframe, and this probability value serves as the confidence score, representing the reliability of the classification result. Next, the emotion category, confidence score, and color emotion data from visual features are fused to generate the emotion feature vector for the keyframe. Assuming the emotion category is represented using one-hot encoding, each emotion category is represented by a binary vector, with only the corresponding category position set to 1 and the others to 0. For example, for the happiness emotion category, if there are a total of 6 emotion categories, its one-hot encoding might be [1,0,0,0,0,0]. The one-hot encoded emotion category vector, the confidence score (as a scalar), and the color emotion data (such as the previously extracted color emotion feature vector) are concatenated to form a new vector. To fuse different types of data on the same scale, it may be necessary to normalize these data, such as normalizing the confidence score to the range of 0-1 and standardizing the color emotion feature vector to have a mean of 0 and a standard deviation of 1. The final fused vector is the emotion feature vector of the keyframe, which integrates multiple emotional information and can more comprehensively and accurately represent the emotion contained in the keyframe, providing a crucial data foundation for subsequent operations such as adding music to videos based on emotion.

[0032] S300: Summarize the emotion classification results of all keyframes, count the proportion of each emotion in the video, determine the dominant emotion of the video based on the proportion, initially filter the range of background music styles that match the dominant emotion, query the preset emotion-music feature mapping library, and obtain the corresponding background music parameters based on the determined dominant emotion.

[0033] In this implementation, the emotion classification results of all keyframes are summarized, the proportion of each emotion in the video is statistically analyzed, the dominant emotion of the video is determined based on the proportion, a preliminary selection of background music styles matching the dominant emotion is made, a preset emotion-music feature mapping library is queried, and the corresponding background music parameters are obtained based on the determined dominant emotion. The specific process is as follows: S301: Summarize the emotion classification results of all keyframes, count the proportion of each emotion in the video, determine the dominant emotion of the video based on the proportion, and initially screen the range of background music styles that match the dominant emotion. S302: Query the preset emotion-music feature mapping library, and obtain the corresponding background music parameters from the mapping library based on the determined dominant emotion; the background music parameters include music rhythm, mode, core instrument type and chord progression; S303: Adjust the music generation duration to match the video duration, generate initial background music audio that meets the emotional tone and duration requirements based on the parameters, and obtain the timestamp information corresponding to the video keyframes; S304: Align audio and video timestamps, analyze the rhythmic accent positions of the initial soundtrack, match the rhythmic accents of the soundtrack with the switching timestamps of video keyframes, synchronize the soundtrack with the keyframe images, and generate the final soundtrack.

[0034] In the above process, after completing the sentiment analysis of keyframes, it is necessary to summarize the sentiment classification results of all keyframes to determine the overall emotional tone of the video and then select a suitable range of background music styles. Firstly, by statistically analyzing the proportion of each type of emotion in the video, we can intuitively understand the distribution of different emotions within the video. For example, if a video has 40% of keyframes showing happiness, 20% showing sadness, and other emotions making up a smaller proportion, then happiness is likely the dominant emotion in the video. In actual statistical analysis, data structures in programming languages, such as dictionaries in Python, can be used to store the number of keyframes for various emotions.

[0035] Based on the statistically derived emotional proportions, the dominant emotion of the video is determined. This dominant emotion is a crucial basis for selecting a soundtrack style, as the music style needs to match the overall emotional atmosphere of the video to enhance audience emotional resonance. For example, for a video with a dominant emotion of happiness, the initial range of soundtrack styles might include upbeat pop music, light electronic music, and lively folk music; if the dominant emotion is sadness, the range of soundtrack styles might lean towards lyrical classical music, slow jazz, and deep rock. This method of selecting soundtrack styles based on emotional proportions allows for the rapid identification of music types that match the video's emotion from a wide range of musical styles, providing a clear direction for subsequent soundtrack creation. After determining the range of musical styles, it is necessary to query a pre-defined emotion-music feature mapping library to obtain the musical parameters corresponding to the dominant emotion. This mapping library is built upon extensive music data and emotion analysis research, recording the correspondence between different emotions and musical features. For example, for the emotion of happiness, the mapping library might record a faster musical tempo, typically between 120-140 BPM; the key would be predominantly major, as major keys often evoke a positive and cheerful feeling; the core instrument types might include piano, violin, guitar, etc., which can produce bright and lively timbres, enhancing the happy atmosphere; and the chord progressions might be common cheerful chord progressions, such as C-Am-F.

[0036] The system retrieves music parameters corresponding to the emotion of happiness from a mapping library. These parameters provide specific guidance for generating music that matches the emotional tone, ensuring that the generated music matches the dominant emotion of the video in terms of rhythm, mode, instrument selection, and chord progression, thereby enhancing the emotional expression of the video. After obtaining the background music parameters, the music generation duration needs to be adjusted to match the video duration to ensure that the music can play in sync with the video without being too long or too short.

[0037] After adjusting the music duration, an initial background music audio that matches the emotional tone and duration requirements is generated using a music generation model based on the acquired background music parameters. The music generation model can be a deep learning-based model, such as WaveNet. WaveNet, through learning from a large amount of audio data, can generate audio with specific styles and characteristics based on input parameters. Parameters such as rhythm, mode, core instrument type, and chord progression are input into the WaveNet model, and the model generates corresponding audio waveform data according to these parameter requirements, thus obtaining the initial background music audio. At the same time, obtaining the timestamp information corresponding to the video keyframes is also crucial. Timestamp information accurately records the appearance time of each keyframe in the video, providing a time reference for subsequently synchronizing the background music with the keyframe visuals. Timestamp information can be obtained through video editing software or related video processing libraries. For example, when using FFmpeg for video processing, the timestamps of keyframes can be obtained by parsing the video's metadata. FFmpeg stores the video's timing information in metadata in a specific format; by reading and parsing this metadata, the accurate timestamp of each keyframe can be obtained. The key to generating the final soundtrack lies in precisely synchronizing the initial soundtrack with the video keyframes, ensuring that the rhythm and emotional changes of the music perfectly match the key moments in the video, thereby enhancing the overall expressiveness of the video. First, aligning the audio and video timestamps is fundamental to achieving synchronization. This process is akin to calibrating two different time series, ensuring consistency across the time dimension. The initial timestamp alignment is completed by perfectly aligning the start time with the start time of the video. Use audio analysis tools to detect rhythmic accents, such as the Essentia library. Essentia provides a series of audio feature extraction algorithms, including a rhythmic accent detection algorithm. This algorithm can accurately identify the timing of each rhythmic accent in the initial soundtrack. The core synchronization operation involves mapping the rhythmic accents of the background music to the timestamps of keyframe transitions in the video. When a keyframe transition is detected, its timestamp is matched with the most recent rhythmic accent in the background music. For example, in an action video, when a keyframe shows the protagonist performing a difficult move, the timestamp of this keyframe is matched with the time of a strong drumbeat in the background music. This ensures that the music's strong rhythm complements the protagonist's impressive action on screen, enhancing the visual and auditory impact. By mapping each video keyframe to a specific rhythmic accent in this way, perfect synchronization between the music and the keyframe visuals is achieved, ultimately generating a final background music that is highly integrated with the video content and consistent with its emotional expression.

[0038] Corresponding to the aforementioned embodiments of the method for generating image-generated video and background music based on multi-model collaboration, this application also provides embodiments of a system for generating image-generated video and background music based on multi-model collaboration.

[0039] Figure 5 This is a system block diagram illustrating a multi-model collaborative graph-generated video and background music system according to an exemplary embodiment. (Refer to...) Figure 5 The system may include: a structured description generation module 401, a video keyframe emotion feature extraction module 402, and a background music matching generation module 403; wherein: The structured description generation module 401 is used to acquire image data, extract visual features of the image to generate a structured description, and generate the current video based on the scene logic and action sequence in the structured description. The video keyframe emotion feature extraction module 402 is used to extract the current video keyframe, extract the visual features of the keyframe, obtain the emotion category and confidence level corresponding to each keyframe, and fuse the emotion category, confidence level, and color emotion data in the visual features to generate the emotion feature vector of the keyframe. The music matching and generation module 403 is used to summarize the emotion classification results of all keyframes, count the proportion of each emotion in the video, determine the dominant emotion of the video based on the proportion, initially screen the range of music styles that match the dominant emotion, query the preset emotion-music feature mapping library, and obtain the corresponding music parameters based on the determined dominant emotion.

[0040] In this embodiment, the structured description generation module 401 acquires image data, extracts visual features from the images to generate a structured description, and generates the current video based on the scene logic and action sequence in the structured description; the video keyframe emotion feature extraction module 402 extracts keyframes from the current video, extracts keyframe visual features, and obtains the emotion category and confidence level corresponding to each keyframe, and integrates the emotion category, confidence level, and color emotion data in the visual features to generate the emotion feature vector of the keyframe; the music matching generation module 403 summarizes the emotion classification results of all keyframes, counts the proportion of each emotion in the video, determines the dominant emotion of the video based on the proportion, initially filters the range of music styles that match the dominant emotion, queries a preset emotion-music feature mapping library, and obtains the corresponding music parameters based on the determined dominant emotion; through the above methods, the matching degree between emotion and music is improved, thereby breaking through the existing technical bottlenecks and meeting the growing demand for integrated content creation.

[0041] Regarding the system in the above embodiments, the specific ways in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0042] For the system embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0043] Accordingly, this application also provides an electronic device, including: one or more processors; a memory for storing one or more programs; and when the one or more programs are executed by the one or more processors, causing the one or more processors to implement the above-described method for image-generated video and background music based on multi-model collaboration. Figure 6 The diagram shown is a hardware structure diagram of any device with data processing capabilities, which is part of a system for graph-generated video and background music based on multi-model collaboration provided by an embodiment of the present invention. (Except for...) Figure 6In addition to the processor, memory, and network interface shown, any data processing device in the embodiment may also include other hardware depending on the actual function of the data processing device, which will not be described in detail here.

[0044] Accordingly, this application also provides a computer-readable storage medium storing computer instructions thereon, which, when executed by a processor, implement the method for generating image-generated video and background music based on multi-model collaboration as described above. The computer-readable storage medium can be an internal storage unit of any data-processing device as described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units of any data-processing device and external storage devices. The computer-readable storage medium is used to store the computer program and other programs and data required by the data-processing device, and can also be used to temporarily store data that has been output or will be output.

[0045] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0046] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. A method for generating video and music based on multi-model cooperation, characterized in that, The method comprises the following steps: Obtaining picture data, extracting picture visual features to generate a structured description, and generating a current video based on the scene logic and action sequence in the structured description; Extracting key frames of the current video, extracting key frame visual features, and obtaining the emotion category and confidence corresponding to each key frame, and fusing the emotion category, confidence, and color emotion data in the visual features to generate a key frame emotion feature vector; Summarizing the emotion classification results of all key frames, calculating the proportion of each emotion in the video, determining the dominant emotion of the video according to the proportion, preliminarily screening the music style range matching the dominant emotion, querying a preset emotion-music feature mapping library, and obtaining corresponding music parameters according to the determined dominant emotion.

2. The method for generating video and music based on multi-model cooperation according to claim 1, wherein, In the step of obtaining picture data, extracting picture visual features to generate a structured description, and generating a current video based on the scene logic and action sequence in the structured description: Receiving a picture uploaded by a user, and performing format standardization processing on the picture; Extracting visual features from the standardized picture to obtain feature data containing picture color distribution, object contour, and scene type information; Generating an initial text description of the picture content according to the feature data; Presetting a structured template, optimizing the initial text according to the structured template, and generating a structured description; The structured template includes three dimensions of timeline, element logic, and emotion tone; Generating a current video according to the scene logic and action sequence in the structured description. 3.The method of generating video and music based on multi-model cooperation according to claim 1, wherein, In the step of extracting key frames of the current video, extracting key frame visual features, and obtaining the emotion category and confidence corresponding to each key frame, and fusing the emotion category, confidence, and color emotion data in the visual features to generate a key frame emotion feature vector: Frame analysis is performed on the current video to obtain total frame number and frame rate information of the video, and candidate key frames are marked according to pixel difference values of adjacent two frames; Grouping processing is performed on all candidate key frames, final key frames are determined according to the video length, and candidate frames with repeated content in each group are removed to obtain a final key frame sequence of the video; Visual feature data is obtained according to the key frame sequence; The visual feature data includes key frame color emotion, scene emotion, and object emotion data. 4.The method of generating video and music based on multi-model cooperation according to claim 3, wherein, In the step of frame analysis on the current video, obtaining total frame number and frame rate information of the video, and marking candidate key frames according to pixel difference values of adjacent two frames: A threshold value is set, and the difference value is compared with the threshold value. When the difference value exceeds the preset threshold value, the frame is marked as a candidate key frame.

5. The method for generating video and music based on multi-model cooperation according to claim 4, wherein, After the step of frame analysis on the current video, obtaining total frame number and frame rate information of the video, and marking candidate key frames according to pixel difference values of adjacent two frames: Each frame of the video is subjected to region identification, and the core region in the frame is extracted.

6. The method for generating video and music based on multi-model cooperation according to claim 5, wherein, In the step of region identification on each frame of the video, and extracting the core region in the frame: A proportion is set, and the region change rate of the current frame is compared with that of the previous frame. When the change rate exceeds the set proportion, the frame is additionally marked as a candidate key frame.

7. The method for generating video and music based on multi-model cooperation according to claim 3, wherein, After the step of obtaining visual feature data according to the key frame sequence; The visual feature data includes key frame color emotion, scene emotion, and object emotion data. The emotion category and confidence corresponding to each frame of key frame are acquired, the emotion category, the confidence, and color emotion data in the visual feature are fused to generate an emotion feature vector of the key frame. 8.The method of generating video and music based on multi-model cooperation according to claim 1, wherein, In the step of aggregating emotion classification results of all key frames, counting proportions of each emotion in the video, determining a dominant emotion of the video according to the proportions, preliminarily screening a music style range matched with the dominant emotion, querying a preset emotion-music feature mapping library, and acquiring corresponding music parameters according to the determined dominant emotion: According to the determined dominant emotion, corresponding music parameters are acquired from the mapping library; wherein the music parameters include music rhythm, mode, core instrument type, and chord progression. 9.The method of generating video and music based on multi-model cooperation according to claim 8, wherein, After the step of acquiring corresponding music parameters from the mapping library according to the determined dominant emotion; wherein the music parameters include music rhythm, mode, core instrument type, and chord progression: The music generation duration is adjusted to be consistent with the video duration, initial music audio conforming to the emotion key and duration requirements is generated according to the parameters, and timestamp information corresponding to the video key frame is acquired; The audio and video timestamps are aligned, the rhythm accent position of the initial music is analyzed, the rhythm accent of the music is corresponded to the switching timestamp of the video key frame, the music is synchronized with the key frame picture, and final music is generated.

10. A system for generating video and music based on multi-model cooperation, adopting the method for generating video and music based on multi-model cooperation according to claim 1, characterized in that, The system comprises a structured description generation module, a video key frame emotion feature extraction module, and a music matching generation module; wherein: The structured description generation module is configured to acquire picture data, extract picture visual features to generate a structured description, and generate a current video based on scene logic and action sequences in the structured description; The video key frame emotion feature extraction module is configured to extract a current video key frame, extract key frame visual features, acquire emotion categories and confidence corresponding to each frame of key frame, and fuse emotion categories, confidence, and color emotion data in the visual features to generate an emotion feature vector of the key frame; The music matching generation module is configured to aggregate emotion classification results of all key frames, count proportions of each emotion in the video, determine a dominant emotion of the video according to the proportions, preliminarily screen a music style range matched with the dominant emotion, query a preset emotion-music feature mapping library, and acquire corresponding music parameters according to the determined dominant emotion.

Citation Information

Patent Citations

  • Video processing method, electronic equipment and storage medium

    CN110602550A

  • Video generation method and device and terminal equipment

    CN110830845A

  • Video music matching method and device, electronic equipment and computer readable storage medium

    CN110971969A

  • Method and system for generating video and dubby music on basis of multi-model collaboration

    CN118782045A

  • Video processing method and device, related equipment, storage medium and computer program product

    CN119182914A