Music-driven personalized video generation method and device

By acquiring text descriptions and music feature information, adjusting the content script, and generating video clips synchronized with the music, the problem of the disconnect between emotion and visual style in music video generation in existing technologies is solved, achieving efficient and high-quality personalized video generation.

CN121908080APending Publication Date: 2026-04-21TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies struggle to synchronize emotional expression and visual style in music video generation, resulting in low generation efficiency and quality, high computational resource consumption, and a lack of music-driven mechanisms and cross-modal integration capabilities.

Method used

By acquiring text descriptions and target music, extracting music feature information, adjusting the semantic content and timing of the content script, generating a text sequence that matches the rhythm and mood of the music, and using a large language model and a visual generation model to generate video clips synchronized with the music, which are then spliced ​​together to form the target video.

Benefits of technology

It enables users to quickly generate visual content that resonates with emotions and is synchronized with rhythm through music and text, improving generation efficiency and quality, lowering the creative threshold, and supporting multiple style templates and intelligent rhythm synchronization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121908080A_ABST
    Figure CN121908080A_ABST
Patent Text Reader

Abstract

The invention provides a music-driven personalized video generation method and device, and relates to the technical field of artificial intelligence content generation, and the method comprises the steps: obtaining text description and target music, generating a content script containing a plurality of script segments based on the text description, and extracting music feature information of the target music; performing semantic reconstruction and time slice alignment on the content script based on the music feature information to obtain a text sequence matched with music rhythm and emotion; and after generating a plurality of images based on the text sequence, generating a plurality of video clips based on the plurality of images, and splicing the plurality of video clips according to the music rhythm of the target music to obtain a target video. According to the music-driven personalized video generation method and device provided by the invention, a user can quickly generate visual contents with emotion resonance and rhythm synchronization only through music and texts, the generation efficiency and the quality of the generated contents are greatly improved, and coordinated matching of multi-modal contents is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence-generated content technology, and in particular to a music-driven personalized video generation method and apparatus. Background Technology

[0002] With the development of cross-modal artificial intelligence generated content (AIGC) technology, it has made significant progress in the fields of image, video and audio generation.

[0003] In the field of music generation, related technologies have lowered the barriers to music creation and generation, enabling automatic music accompaniment and timbre conversion based on control parameters. Meanwhile, an increasing number of multimodal generation systems are attempting to fuse text, audio, and visual signals to generate video content with rhythm, emotion, and narrative logic.

[0004] However, the technical solutions in the related technologies cannot guarantee the generation efficiency and quality, and the generated content also lacks emotional expression and visual style. Summary of the Invention

[0005] The purpose of this application is to provide a music-driven personalized video generation method and apparatus that enables users to quickly generate visual content with emotional resonance and rhythmic synchronization using only music and text, greatly improving generation efficiency and the quality of generated content.

[0006] This application provides a music-driven personalized video generation method, including: The process involves acquiring a text description and target music, generating a content script containing multiple script fragments based on the text description, and extracting musical feature information from the target music. Based on the musical feature information, the semantic content and timing of the multiple script fragments in the content script are adjusted to obtain a text sequence that matches the musical rhythm and mood. Multiple images are generated based on the text sequence, and multiple video clips are generated based on the images. These video clips are then spliced ​​together according to the musical rhythm of the target music to obtain the target video.

[0007] Optionally, the music feature information includes at least one of the following: rhythmic structure, melody line, beat density, and emotional features; the text description includes at least one of the following: scene, character, action, and time segment.

[0008] Optionally, generating a content script containing multiple script fragments based on the text description includes: extracting element information from the text description using natural language processing; the element information includes at least one of the following: theme, emotion, and time; generating the content script based on the element information; wherein the content script includes multiple script fragments, and each script fragment includes at least one of the following: scene, character, event, and time.

[0009] Optionally, the step of extracting the musical feature information of the target music includes: determining the rhythm nodes and spectral features of the target music through beat detection and spectral feature analysis; and determining the rhythmic structure, melody line, beat density, and emotional features of the target music based on the rhythm nodes and spectral features of the target music.

[0010] Optionally, adjusting the semantic content and occurrence order of multiple script segments in the content script based on the music feature information to obtain a text sequence that matches the music rhythm and emotion includes: semantically reconstructing and aligning time segments of the content script based on the rhythmic structure, melody line, beat density, and emotional features of the target music to obtain an intermediate script; and assigning each script segment in the intermediate script to a corresponding time interval based on the lyrics timestamp and lyrics paragraph structure of the target music to obtain the text sequence.

[0011] Optionally, after generating multiple images based on the text sequence, generating multiple video segments based on the multiple images, and splicing the multiple video segments according to the musical rhythm of the target music to obtain the target video, includes: generating visual cue words corresponding to the text sequence using a large language model, inputting the visual cue words into an image generation model to generate multiple keyframe images; one script segment corresponds to at least two keyframe images; generating multiple video segments using a visual generation model based on each keyframe image and the corresponding script segment; and splicing the multiple video segments according to the musical rhythm of the target music to obtain the target video.

[0012] This application also provides a music-driven personalized video generation device, comprising: The system includes an information acquisition module for acquiring text descriptions and target music; an information extraction module for generating a content script containing multiple script fragments based on the text descriptions, and extracting musical feature information from the target music; a text generation module for adjusting the semantic content and occurrence sequence of multiple script fragments in the content script based on the musical feature information, to obtain a text sequence that matches the music rhythm and mood; and a video generation module for generating multiple images based on the text sequence, generating multiple video clips based on the multiple images, and splicing the multiple video clips according to the musical rhythm of the target music to obtain the target video.

[0013] Optionally, the text generation module is specifically used to extract element information from the text description using natural language analysis; the element information includes at least one of the following: topic, emotion, and time; the text generation module is further used to generate the content script based on the element information; wherein the content script includes: multiple script fragments, and each script fragment includes at least one of the following: scene, character, event, and time.

[0014] Optionally, the information extraction module is specifically used to determine the rhythm nodes and spectral features of the target music through beat detection and spectral feature analysis; the information extraction module is also specifically used to determine the rhythmic structure, melody line, beat density and emotional features of the target music based on the rhythm nodes and spectral features of the target music.

[0015] Optionally, the text generation module is specifically used to perform semantic reconstruction and time segment alignment on the content script based on the rhythm structure, melody line, beat density and emotional features of the target music to obtain an intermediate script; the text generation module is also specifically used to assign each script segment in the intermediate script to a corresponding time interval based on the lyrics timestamp and lyrics paragraph structure of the target music to obtain the text sequence.

[0016] Optionally, the video generation module is specifically used to generate visual cue words corresponding to the text sequence using a large language model, and input the visual cue words into an image generation model to generate multiple keyframe images; one script segment corresponds to at least two keyframe images; the video generation module is further used to generate multiple video segments based on each keyframe image and the corresponding script segment using a visual generation model; the video generation module is further used to splice the multiple video segments according to the musical rhythm of the target music to obtain the target video.

[0017] This application also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the music-driven personalized video generation method as described above.

[0018] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of any of the above-described music-driven personalized video generation methods.

[0019] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the music-driven personalized video generation method as described above.

[0020] The music-driven personalized video generation method and apparatus provided in this application first acquire a text description and target music, and generate a content script containing multiple script fragments based on the text description, and extract music feature information of the target music; then, based on the music feature information, adjust the semantic content and occurrence order of the multiple script fragments in the content script to obtain a text sequence that matches the music rhythm and emotion; finally, after generating multiple images based on the text sequence, generate multiple video segments based on the multiple images, and splice the multiple video segments according to the music rhythm of the target music to obtain the target video. In this way, users can quickly generate visual content with emotional resonance and rhythmic synchronization using only music and text, greatly improving generation efficiency and the quality of generated content. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is one of the flowcharts illustrating the music-driven personalized video generation method provided in this application; Figure 2 This is the second flowchart of the music-driven personalized video generation method provided in this application; Figure 3 This is the third flowchart of the music-driven personalized video generation method provided in this application; Figure 4 This is a schematic diagram of the structure of the music-driven personalized video generation device provided in this application; Figure 5This is a schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] The terms "first," "second," etc., used in this application's specification are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects; for example, a first object can be one or more. Furthermore, in the specification, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects have an "or" relationship.

[0025] Generative Artificial Intelligence (AIGC) refers to techniques based on generative adversarial networks (GANs), large-scale pre-trained models, and other artificial intelligence methods that learn from and recognize existing data to generate relevant content with appropriate generalization capabilities. The core idea of ​​AIGC is to utilize artificial intelligence algorithms to generate content with a certain degree of creativity and quality. Through training models and learning from large amounts of data, AIGC can generate relevant content based on input conditions or guidance. For example, by inputting keywords, descriptions, or samples, AIGC can generate matching articles, images, audio, etc.

[0026] Despite progress in the multimodal integration of AIGC (AIGC), several bottlenecks remain. Existing systems primarily focus on "text-driven visual generation" or "monomodal audio-assisted generation," lacking a dynamic coupling mechanism between music, rhythm, emotion, and visual content. The visual style of generated videos often clashes with the musical rhythm, making it difficult to create a structured, rhythmically synchronized audiovisual experience. Furthermore, high-resolution and high-frame-rate video generation still faces challenges such as high computational resource consumption and pipeline fragmentation (lack of a unified framework from storyboarding to generation), limiting the application of AIGC in music video generation scenarios. These limitations are mainly reflected in the following aspects: 1. Insufficient cross-modal generation integration capabilities: Text, image, and audio modalities are mostly processed sequentially, lacking unified modeling. 2. Lack of music-driven mechanisms: Related technologies mostly rely on text as the primary controller, with music serving only as an auxiliary feature input. 3. Poor video consistency: I2V generation suffers from inconsistent styles between characters and scenes. 4. Low generation efficiency and quality: High-resolution video generation is computationally intensive and involves a fragmented process. 5. Lack of emotional expression and visual style: Visual effects lack a correspondence with the emotional tone of the music.

[0027] To address the aforementioned technical problems in related technologies, embodiments of this application provide a music-driven personalized video generation method, such as... Figure 1 As shown, this method aims to enable users to quickly generate visual content with emotional resonance and rhythmic synchronization through music. The entire process, from music input to video output, is fully automated by AI, lowering the creative threshold and improving generation efficiency. First, users upload a piece of music or select a track recommended by the system through the application interface. Then, optional inputs include keywords, lyrics, or descriptive phrases to indicate the video theme (such as "city night view," "youthful memories," etc.). The system analyzes the rhythm, melody, intensity, and emotional characteristics of the music; generates corresponding visual scripts (scripts describing the visual content of each rhythmic segment of the video); and generates images and video clips through an AIGC model. Finally, the system automatically stitches the generated video clips into a complete video, overlays subtitles, effects, and music, and outputs the final product. Users can directly preview or download the video and share it on social media platforms.

[0028] The music-driven personalized video generation method provided in this application transforms the passive experience of "listening to music" into the active visual experience of "watching music-generated videos" through AI. Users previously only received auditory emotional stimulation through music, but this invention utilizes AIGC technology to generate video content that resonates with the rhythm, melody, and emotion of the music, achieving a transformation "from music evoking emotion to AI-generated visual content." The core value of the system lies in: automatically matching the rhythm of the music with the visual imagery; generating personalized music videos consistent with the music style and unified in style; and allowing ordinary users to generate high-quality content without video editing or modeling experience.

[0029] The music-driven personalized video generation method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0030] like Figure 2 As shown in the embodiment of this application, a music-driven personalized video generation method is provided, which may include the following steps 201 to 203: Step 201: Obtain the text description and target music, generate a content script containing multiple script fragments based on the text description, and extract the music feature information of the target music.

[0031] The musical feature information includes at least one of the following: rhythmic structure, melody line, beat density, and emotional features; the text description includes at least one of the following: scene, character, action, and time segment.

[0032] It should be noted that the above text description can be user-defined input or obtained through a formatted questionnaire. Taking obtaining the text description through a questionnaire as an example, users upload music files and input text descriptions related to the music (which can be lyrics, keywords, or phrases) through an interactive interface. The goal of this stage is to structure the user's text input into machine-understandable semantic units, including scenes, characters, actions, time segments, etc., laying the foundation for subsequent music alignment and generation.

[0033] Specifically, step 201 above, the step of generating a content script containing multiple script fragments based on the text description, may further include the following steps 201a1 and 201a2: Step 201a1: Extract element information from the text description using natural language processing.

[0034] The element information includes at least one of the following: theme, mood, and time.

[0035] Step 201a2: Generate the content script based on the element information.

[0036] The content script includes multiple script fragments, and each script fragment includes at least one of the following: scene, character, event, and time.

[0037] For example, the system can extract the theme, emotion, and time elements through natural language analysis and automatically generate a preliminary "content script (Memory Raw Text / Script)".

[0038] Specifically, step 201 above, the step of extracting the musical feature information of the target music, may further include the following steps 201b1 and 201b2: Step 201b1: Determine the rhythm nodes and spectral characteristics of the target music through beat detection and spectral feature analysis.

[0039] Step 201b2: Based on the rhythmic nodes and spectral characteristics of the target music, determine the rhythmic structure, melody line, beat density, and emotional characteristics of the target music.

[0040] For example, by using beat tracking and spectral feature analysis (such as Mel-Spectrogram and Spectral Centroid), the rhythmic nodes of music can be determined. Then, the rhythmic structure, melody line, beat density, and emotional characteristics of the target music can be further determined.

[0041] Understandably, rhythmic nodes are sequences of time points obtained through beat detection technology; they are the most basic and regular pulses in music. They are the "skeleton" of music, defining its time grid. Rhythmic structure refers to higher-level patterns organized by rhythmic nodes, such as measures and phrases. It typically involves the cycle of strong and weak beats. Melodic line is the flow of a series of continuous notes across pitches, forming the main "melody" of the music. Beat density is the number of musical events (such as notes and drumbeats) occurring per unit of time (e.g., a measure). High density creates a sense of complexity and tension; low density creates a sense of sparseness and relaxation. Emotional characteristics are the emotions evoked by music, such as joy, sadness, excitement, and tranquility. This is the result of the combined effect of all the above elements.

[0042] In summary, rhythm nodes, as the most basic time grid in music, are the skeleton that constructs all higher-level musical elements: they define the rhythmic structure of music through periodic strong and weak patterns, provide a temporal framework and rhythmic basis for the rise and fall of the melody line and phrasing, and determine the density of beats through the change in the number of events per unit time. Ultimately, these elements—including the node's own speed, stability, and its synergistic effect with spectral characteristics (such as timbre brightness)—interweave together to become the core force driving the emotional characteristics of music (such as excitement or tranquility).

[0043] Step 202: Based on the music feature information, adjust the semantic content and occurrence sequence of multiple script segments in the content script to obtain a text sequence that matches the music rhythm and emotion.

[0044] For example, the system performs semantic reorganization and time segment alignment on the content script generated in the previous stage based on the rhythmic structure, melody line, beat density and emotional characteristics of the target music.

[0045] For example, this stage implements a dynamic synchronization mechanism between video content and music rhythm and emotion, which is a key innovation that distinguishes it from traditional AIGC text-driven generation.

[0046] Specifically, step 202 above may also include the following steps 202a1 and 202a2: Step 202a1: Based on the rhythmic structure, melody line, beat density, and emotional characteristics of the target music, perform semantic reconstruction and time segment alignment on the content script to obtain an intermediate script.

[0047] Step 202a2: Based on the lyrics timestamp and lyrics paragraph structure of the target music, assign each script segment in the intermediate script to the corresponding time interval to obtain the text sequence.

[0048] For example, after semantic restructuring and time segment alignment of the script generated in the previous stage, the content script is assigned to the corresponding time interval by combining the lyrics timestamp and paragraph structure, and then "Music-AlignedScript" can be output, which is a text sequence that matches the rhythm and mood of the music.

[0049] Step 203: After generating multiple images based on the text sequence, generate multiple video segments based on the multiple images, and splice the multiple video segments according to the musical rhythm of the target music to obtain the target video.

[0050] For example, after obtaining the above text sequence, a two-stage AIGC generation (Two-Stage AIGC Pipeline) can be performed to obtain the final target video.

[0051] Specifically, step 203 above may also include steps 203a1 to 203a3: Step 203a1: Generate visual cue words corresponding to the text sequence using a large language model, and input the visual cue words into an image generation model to generate multiple keyframe images.

[0052] One script segment corresponds to at least two keyframe images.

[0053] Step 203a2: Based on each keyframe image and the corresponding script fragment, generate multiple video fragments using a visual generation model.

[0054] Step 203a3: Splice the multiple video segments together according to the rhythm of the target music to obtain the target video.

[0055] For example, such as Figure 3As shown, this stage consists of two parts: Text-to-Image (T2I) generation: Detailed visual cues (Prompts) are generated using a Large Language Model (LLM), and keyframe images are generated using an image generation model (such as FLUX.1-dev or Stable Diffusion). Each segment corresponds to several static frames to ensure consistent scene layout and style. Image-to-Video (I2V) generation: Based on the keyframes and their descriptions, dynamic short clips are generated using a video generation model (such as HunyuanVideo or WAN2.2) to capture character movements and scene changes. Through this two-stage generation mechanism, the system maintains visual consistency while achieving motion changes coordinated with the music rhythm, improving the naturalness and coherence of the generated video.

[0056] For example, such as Figure 3 As shown, the system splices the generated video clips according to the music rhythm and performs detailed post-processing: switching shots according to the rhythm; adding lyrics or text descriptions; and matching colors and brightness to ensure visual consistency. The final output is a complete personalized music video file (M). 2 Video contains multimodal information such as visual images, background music, and subtitle text.

[0057] For example, such as Figure 3 As shown, the music-driven personalized video generation method provided in this application includes two parts, a and b. a) End-to-end workflow, divided into four stages: memory text collection (collecting user memories); structuring and music alignment reorganization (organizing narrative content and aligning it with the music structure); a two-stage AIGC pipeline (generating cues, then generating memory images and memory videos); and synthesis and presentation (editing clips with music and subtitles into a complete M2 video). b) Step-level outputs and results: showing the specific intermediate results of the four stages, and how to transform raw memories into the final completed M2 video—including text materials (notes, scripts, subtitles), visual assets (storyboards, images), and video clips.

[0058] The music-driven personalized video generation method provided in this application has the following advantages: Simple Process: Simply upload your music and ideas for one-click generation. Real-time Preview: The system displays keyframes during generation, allowing users to fine-tune style keywords. Multiple Style Templates: Supports realistic, anime, abstract, and mood-inspired lighting styles. Intelligent Rhythm Synchronization: The system automatically switches camera angles based on the shooting points, creating a rhythmic visual experience synchronized with the music. Expandable Application Scenarios: Music video production, personal video generation, short video platform content creation, digital art exhibitions, etc.

[0059] The music-driven personalized video generation method provided in this application first obtains a text description and target music, and generates a content script containing multiple script fragments based on the text description, and extracts the music feature information of the target music; then, based on the music feature information, it adjusts the semantic content and occurrence order of the multiple script fragments in the content script to obtain a text sequence that matches the music rhythm and emotion; finally, it generates multiple images based on the text sequence, generates multiple video clips based on the multiple images, and splices the multiple video clips according to the music rhythm of the target music to obtain the target video. In this way, users can quickly generate visual content with emotional resonance and rhythmic synchronization using only music and text, greatly improving generation efficiency and the quality of generated content.

[0060] It should be noted that the music-driven personalized video generation method provided in this application embodiment can be executed by a music-driven personalized video generation device, or a control module within that device for executing the music-driven personalized video generation method. This application embodiment uses the execution of the music-driven personalized video generation method by a music-driven personalized video generation device as an example to illustrate the music-driven personalized video generation device provided in this application embodiment.

[0061] It should be noted that, in the embodiments of this application, the music-driven personalized video generation methods shown in the accompanying drawings are all illustrated using one accompanying drawing from one of the embodiments of this application as an example. In specific implementation, the music-driven personalized video generation methods shown in the accompanying drawings of the above methods can also be implemented in conjunction with any other accompanying drawings that can be combined as illustrated in the above embodiments, which will not be elaborated here.

[0062] The music-driven personalized video generation apparatus provided in this application will be described below. The music-driven personalized video generation method described below can be referred to in correspondence with the music-driven personalized video generation method described above.

[0063] Figure 4 This is a schematic diagram of the structure of the music-driven personalized video generation device provided in the embodiments of this application, as shown below. Figure 4 As shown, it specifically includes: Information acquisition module 401 is used to acquire text description and target music; information extraction module 402 is used to generate a content script containing multiple script fragments based on the text description, and to extract the music feature information of the target music; text generation module 403 is used to adjust the semantic content and occurrence sequence of multiple script fragments in the content script based on the music feature information to obtain a text sequence that matches the music rhythm and emotion; video generation module 404 is used to generate multiple images based on the text sequence, generate multiple video fragments based on the multiple images, and splice the multiple video fragments according to the music rhythm of the target music to obtain the target video.

[0064] Optionally, the text generation module 403 is specifically used to extract element information from the text description using natural language analysis; the element information includes at least one of the following: topic, emotion, and time; the text generation module 403 is further used to generate the content script based on the element information; wherein, the content script includes: multiple script fragments, and each script fragment includes at least one of the following: scene, character, event, and time.

[0065] Optionally, the information extraction module 402 is specifically used to determine the rhythm nodes and spectral features of the target music through beat detection and spectral feature analysis; the information extraction module 402 is also specifically used to determine the rhythm structure, melody line, beat density and emotional features of the target music based on the rhythm nodes and spectral features of the target music.

[0066] Optionally, the text generation module 403 is specifically used to perform semantic reconstruction and time segment alignment on the content script based on the rhythm structure, melody line, beat density and emotional features of the target music to obtain an intermediate script; the text generation module 403 is also specifically used to assign each script segment in the intermediate script to a corresponding time interval based on the lyrics timestamp and lyrics paragraph structure of the target music to obtain the text sequence.

[0067] Optionally, the video generation module 404 is specifically used to generate visual cue words corresponding to the text sequence using a large language model, and input the visual cue words into an image generation model to generate multiple keyframe images; one script segment corresponds to at least two keyframe images; the video generation module 404 is further used to generate multiple video segments based on each keyframe image and the corresponding script segment using a visual generation model; the video generation module 404 is further used to splice the multiple video segments according to the musical rhythm of the target music to obtain the target video.

[0068] The music-driven personalized video generation device provided in this application first acquires a text description and target music, and generates a content script containing multiple script fragments based on the text description, as well as extracting musical feature information from the target music. Then, based on the musical feature information, it adjusts the semantic content and occurrence order of the multiple script fragments in the content script to obtain a text sequence that matches the music rhythm and emotion. Finally, it generates multiple images based on the text sequence, generates multiple video clips based on the multiple images, and splices the multiple video clips according to the musical rhythm of the target music to obtain the target video. In this way, users can quickly generate visual content with emotional resonance and rhythmic synchronization using only music and text, greatly improving generation efficiency and the quality of generated content.

[0069] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5 As shown, the electronic device may include a processor 510, a communications interface 520, a memory 530, and a communication bus 540. The processor 510, communications interface 520, and memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a music-driven personalized video generation method. This method includes: first, acquiring a text description and target music, generating a content script containing multiple script fragments based on the text description, and extracting musical feature information from the target music; then, adjusting the semantic content and occurrence sequence of the multiple script fragments in the content script based on the musical feature information to obtain a text sequence that matches the music rhythm and emotion; finally, generating multiple images based on the text sequence, generating multiple video clips based on the multiple images, and splicing the multiple video clips according to the musical rhythm of the target music to obtain the target video. In this way, users can quickly generate visual content with emotional resonance and rhythmic synchronization using only music and text, greatly improving generation efficiency and the quality of generated content.

[0070] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0071] On the other hand, this application also provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer can execute the music-driven personalized video generation method provided by the above methods. This method includes: first, acquiring a text description and target music, and generating a content script containing multiple script fragments based on the text description, and extracting music feature information of the target music; then, adjusting the semantic content and occurrence sequence of the multiple script fragments in the content script based on the music feature information to obtain a text sequence that matches the music rhythm and emotion; finally, generating multiple images based on the text sequence, generating multiple video fragments based on the multiple images, and splicing the multiple video fragments according to the music rhythm of the target music to obtain the target video. In this way, users can quickly generate visual content with emotional resonance and rhythmic synchronization using only music and text, greatly improving generation efficiency and the quality of generated content.

[0072] Furthermore, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned music-driven personalized video generation methods. The method includes: first, acquiring a text description and target music, generating a content script containing multiple script fragments based on the text description, and extracting musical feature information from the target music; then, adjusting the semantic content and occurrence sequence of the multiple script fragments in the content script based on the musical feature information to obtain a text sequence matching the music rhythm and emotion; finally, generating multiple images based on the text sequence, generating multiple video segments based on the multiple images, and splicing the multiple video segments according to the musical rhythm of the target music to obtain a target video. This allows users to quickly generate visual content with emotional resonance and rhythmic synchronization using only music and text, greatly improving generation efficiency and the quality of generated content.

[0073] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0074] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A music-driven personalized video generation method, characterized in that, include: Obtain a text description and target music, generate a content script containing multiple script fragments based on the text description, and extract the music feature information of the target music; Based on the music feature information, the content script is semantically reconstructed and time segment aligned to obtain a text sequence that matches the music rhythm and emotion; After generating multiple images based on the text sequence, multiple video segments are generated based on the multiple images, and the multiple video segments are spliced ​​together according to the musical rhythm of the target music to obtain the target video.

2. The method according to claim 1, characterized in that, The musical feature information includes at least one of the following: rhythmic structure, melody line, beat density, and emotional features; the text description includes at least one of the following: scene, character, action, and time segment.

3. The method according to claim 1 or 2, characterized in that, The process of generating a content script containing multiple script fragments based on the text description includes: Natural language processing is used to extract element information from the text description; the element information includes at least one of the following: topic, sentiment, and time; The content script is generated based on the element information; The content script includes multiple script fragments, and each script fragment includes at least one of the following: scene, character, event, and time.

4. The method according to claim 1 or 2, characterized in that, The extraction of musical feature information of the target music includes: The rhythmic nodes and spectral characteristics of the target music are determined through beat detection and spectral feature analysis. Based on the rhythmic nodes and spectral characteristics of the target music, the rhythmic structure, melodic line, beat density, and emotional characteristics of the target music are determined.

5. The method according to claim 4, characterized in that, The step of adjusting the semantic content and occurrence sequence of multiple script fragments in the content script based on the music feature information to obtain a text sequence that matches the music rhythm and emotion includes: Based on the rhythmic structure, melody line, beat density, and emotional characteristics of the target music, the content script is semantically reconstructed and time segment aligned to obtain an intermediate script. Based on the lyrics timestamps and lyrics paragraph structure of the target music, each script segment in the intermediate script is assigned to a corresponding time interval to obtain the text sequence.

6. The method according to claim 1 or 5, characterized in that, After generating multiple images based on the text sequence, multiple video segments are generated based on the multiple images, and the multiple video segments are spliced ​​together according to the musical rhythm of the target music to obtain the target video, including: A large language model is used to generate visual cue words corresponding to the text sequence, and the visual cue words are input into an image generation model to generate multiple keyframe images; one script segment corresponds to at least two keyframe images; Based on each keyframe image and the corresponding script segment, multiple video segments are generated using a visual generation model. The multiple video clips are spliced ​​together according to the musical rhythm of the target music to obtain the target video.

7. A music-driven personalized video generation device, characterized in that, The device includes: The information acquisition module is used to acquire text descriptions and target music. The information extraction module is used to generate a content script containing multiple script fragments based on the text description, and to extract the musical feature information of the target music; The text generation module is used to adjust the semantic content and occurrence sequence of multiple script segments in the content script based on the music feature information, so as to obtain a text sequence that matches the music rhythm and emotion. The video generation module is used to generate multiple images based on the text sequence, generate multiple video segments based on the multiple images, and splice the multiple video segments according to the musical rhythm of the target music to obtain the target video.

8. The apparatus according to claim 7, characterized in that, The text generation module is specifically used to extract element information from the text description using natural language processing; the element information includes at least one of the following: topic, sentiment, and time; The text generation module is further configured to generate the content script based on the element information; The content script includes multiple script fragments, and each script fragment includes at least one of the following: scene, character, event, and time.

9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the steps of the music-driven personalized video generation method as described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the music-driven personalized video generation method as described in any one of claims 1 to 6.