Video synthesis method and device based on large language model, and medium
Through the video synthesis method based on the large language model, the automation and high-quality generation of video production are achieved, which solves the problems of cumbersome and time-consuming traditional video production processes and improves production efficiency and user experience.
Patent Information
- Application Number
- CN202511140396.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-10-10
AI Technical Summary
The traditional video production process is cumbersome, time-consuming, and costly, lacking end-to-end automated integration. Users need to switch between multiple tools or platforms, resulting in high production barriers and difficulty in quickly generating large amounts of video content.
A video synthesis method based on a large language model is used to generate high-quality video content through multimodal feature extraction, digital human model generation, voice and content reference information processing, video material matching, and reinforcement learning fine-tuning.
It realizes the automation of video production, simplifies the process, reduces manual intervention, improves the quality of video synthesis and viewing experience, and solves the problems of logical discontinuity, sound and image misalignment, and rhythm confusion in traditional methods.
Smart Images

Figure CN120769137A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video processing technology, and mainly to a video synthesis method, device and medium based on a large language model. Background Art
[0002] The traditional video production process usually includes multiple interdependent and highly specialized stages, including: 1. Video material preparation: it is necessary to plan, shoot or collect a large amount of original video clips; 2. Copywriting: conceiving the video content script and writing commentary or subtitles; 3. Voice recording: professional dubbing or using speech synthesis technology to generate narration; 4. Video editing and synthesis: manual editing, timeline alignment and final synthesis of materials, audio, text, special effects and other elements; these processes are not only cumbersome, but also require a lot of manpower, time and professional skills in each link, resulting in high threshold, low efficiency and high cost of video production; especially for application scenarios that require rapid generation of large amounts of video content, the limitations of traditional methods are particularly prominent.
[0003] Existing technologies typically focus on optimizing a single process, lacking effective solutions for end-to-end automated integration of key steps such as copywriting, visual asset generation, speech synthesis, and video editing. Users still need to switch between multiple tools or platforms, failing to fundamentally streamline the process and significantly reduce the barriers to entry and time required for production.
[0004] Therefore, there is an urgent need for an automated video synthesis method that can deeply integrate artificial intelligence technology. Summary of the Invention
[0005] In order to solve the above problems in the prior art, the present invention proposes a video synthesis method based on a large language model.
[0006] The technical solutions of the present invention are as follows: In one aspect, the present invention provides a video synthesis method based on a large language model, the method comprising: Obtaining user input information, including portrait images, reference voices, and content reference information; Performing multimodal feature extraction on the input information to obtain multimodal features, including portrait image features, reference voice features, and content reference information features; Input the portrait image features into the digital human model built based on deep learning to generate the initial video; Input the reference speech features and content reference information features into the preset first language model to generate video copy and storyboard information; Matching is performed in the local video material library based on multimodal features to obtain matching materials; The storyboard information and matching materials are input into the second language model to generate a video content timeline; Perform video synthesis based on the video content schedule, initial video, and video copy, and output the final synthesized video.
[0007] Preferably, the portrait image features are input into a digital human model built based on deep learning, and the specific steps are as follows: Normalizing the portrait features specifically involves converting features of different dimensions into a tensor format compatible with the digital human model; Inputting the standardized portrait features into the feature fusion layer of the digital human model to identify the unique identifier in the portrait features; embedding the unique identifier into the digital human basic mesh and outputting a three-dimensional mesh model of the portrait image; Generate a first frame of video image from the three-dimensional mesh model according to a preset initial posture, initial expression and basic lighting; The first frame of video image is combined with a preset video dynamic frame sequence to generate an initial video.
[0008] Preferably, the specific steps of inputting the reference speech features and the content reference information features into the preset first language model are: Preprocessing the reference speech features and content reference information features; Input the pre-processed content reference information into the generative layer of the first language model to generate the initial video copy; Inputting the preprocessed reference speech features and content reference information features into the feature fusion layer of the first language model to generate fused features; Generate storyboard information based on the initial video text and fusion features.
[0009] Preferably, matching is performed in the local video material library based on multimodal features, wherein the matching process is to calculate the similarity between the current multimodal features and the video material in the local video material library; if the current similarity is greater than a preset similarity threshold, it indicates that the match is successful, and the current video material is recorded as the matching material.
[0010] Preferably, the training process of the second largest speech model includes constructing a training data set and a fine-tuning method based on reinforcement learning; the training data set specifically includes: Acquire multiple video templates according to a preset video type ratio, and parse the video templates to obtain a tuple structure; A preset number of video editing experts use a preset professional review mechanism to cross-score the tuple structures and select tuple structures whose average cross-score exceeds a preset screening threshold; A training dataset is constructed based on the video templates corresponding to the filtered tuple structures.
[0011] Preferably, the reinforcement learning-based fine-tuning method specifically fine-tunes the second language model through a reward function, wherein: The reward function is calculated based on the coherence reward, rhythm reward, synchronization reward and aesthetic reward and their corresponding weights; A coherence bonus is calculated based on the text description of each storyboard; Calculate the rhythm reward based on the video beat sequence and shot switching sequence; Synchronization rewards are calculated based on the time difference between subtitles and voice; Aesthetic rewards are calculated based on keyframe images.
[0012] Preferably, the storyboard information and matching materials are input into the second language model, and the specific steps are as follows: Input the matching material into the feature extraction layer of the second largest language model to obtain matching features; The storyboard information and matching features are input into the feature alignment layer of the second language model. The temporal attention mechanism is used to calculate the matching degree between the storyboard timing requirements and the time length of the matching material, and generate temporal correlation features. Generate video content schedule based on temporal correlation features.
[0013] Preferably, the method further comprises performing multi-track rendering and packaging on the final synthesized video using a video coding protocol.
[0014] On the other hand, the present invention further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method of the present invention when executing the program.
[0015] In another aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the method of the present invention when executed by a processor.
[0016] The present invention has the following beneficial effects: 1. The present invention provides a video synthesis method, device, and medium based on a large language model. Through multimodal feature extraction technology, the unstructured information input by the user is converted into a computable feature vector. Based on feature similarity, the method accurately matches the materials in the local material library, avoiding the blindness of manual material screening and reducing the redundant storage and retrieval costs of effective materials. 2. The present invention provides a video synthesis method, device, and medium based on a large language model. The first large language model receives speech and content features and, through natural language generation and semantic understanding capabilities, automatically generates video copy and storyboard information that conforms to the content logic. The second large language model, based on the storyboard information and matching materials, combines the professional video template logic in the training data to generate a frame-accurate video content schedule, replacing the traditional tedious process of manually writing copy, designing storyboards, and planning timelines. 3. The present invention provides a video synthesis method, device, and medium based on a large language model. Based on the principles of reinforcement learning fine-tuning and reward functions, the method improves the quality of video synthesis and the viewing experience. A coherence reward ensures consistent storyboard logic, a rhythm reward matches shot switching with audio tempo, a synchronization reward ensures alignment of subtitles and speech, and an aesthetic reward optimizes the visual effects of keyframes. The reward function constrains the model output, resolving issues such as "logical discontinuity, misalignment of sound and image, and rhythmic confusion" in traditional automated synthesis, thereby enhancing the user's sense of immersion when viewing. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a specific flow chart of an embodiment of the present invention. DETAILED DESCRIPTION
[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0019] It should be understood that the step numbers used herein are only for convenience of description and are not intended to limit the order in which the steps are executed.
[0020] It should be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0021] The terms “include” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0022] The term "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items.
[0023] Example 1: See also Figure 1 The present invention provides a video synthesis method based on a large language model, the method comprising: S1. Obtain user input information, including a portrait image, a reference voice, and content reference information; S2. performing multimodal feature extraction on the input information to obtain multimodal features, including portrait image features, reference voice features, and content reference information features; The portrait image features include facial geometric feature matrices such as three-dimensional coordinates of facial features, contour curve parameters, texture feature matrices such as skin color pixel distribution, pore / wrinkle texture data, and hair feature matrices such as hairstyle contour coordinates, hair density / color values, etc. The reference speech features include speech spectrum features, intonation curve, speech rate and rhythm parameters, emotional tendency features, etc. The content reference information features include core text semantic vectors, keyword weights, information hierarchical structure, domain labels, etc. S3, input the portrait image features into the digital human model built based on deep learning to generate an initial video; S31, standardizing the portrait features, specifically converting features of different dimensions into a tensor format compatible with the digital human model; S32. Input the standardized portrait features into the feature fusion layer of the digital human model to identify unique identifiers in the portrait features; embed the unique identifiers into the digital human base mesh, and output a three-dimensional mesh model of the portrait image; for example, mapping facial textures to cheek and forehead meshes, and mapping hair textures to scalp meshes; S33, generating a first frame of video image from the three-dimensional mesh model according to a preset initial posture, initial expression and basic lighting; In this embodiment, the initial posture is set to "natural frontal gaze": the head pitch angle of the three-dimensional grid is ≤3°, the left and right deflection angles are ≤2°, and the shoulders are kept horizontal: the angle between the shoulder line and the horizontal line is ≤1°; Set the initial expression to "neutral and slightly relaxed": set the facial muscle parameters to naturally close the mouth corners, for example, the distance between the upper and lower lips is ≤1mm, the eyelid opening and closing degree is 80%, and the eyebrows are naturally stretched, for example, the brow peak height is the brow bone baseline + 2mm; Set the basic lighting to front main light + side fill light, with a light intensity ratio of 3:1; S34, combining the first frame of video image with a preset video dynamic frame sequence to generate an initial video; In this embodiment, setting the video dynamic frame sequence specifically includes the following steps: 1. Define the initial video parameters: duration 5 seconds, frame rate 30 frames / second, generate a total of 150 frames, where frame 0 is the first frame of the video image, i.e. the static reference frame, and frames 1-149 are dynamic frames. The time axis is divided into 0.033 second / frame intervals; 2. Set the basic range of motion: Limit the range of motion to ensure naturalness. Set the head rotation to ≤5° (corresponding to an angle change of ≤0.1° between frames), and the vertical pitch to ≤3°. A single blink lasts 0.3 seconds (9 frames) and is triggered every 2 seconds (i.e., around the 60th and 120th frames). 3. Motion smoothing rules: Use Bezier curve interpolation algorithm to control inter-frame motion, ensuring a natural transition from "start - acceleration - deceleration - stop"; Frame-by-frame dynamic generation: Frames 1-59: Generate a "slight head shaking" dynamic: With the reference frame as the center, complete a "3° left-return-3° right-return" cycle every 30 frames. The angle of the head mesh is adjusted through the mesh vertex displacement algorithm, and the neck mesh is simultaneously slightly stretched, with the stretching amplitude ≤ 2%; Frames 60-68: Generate a "natural blinking" dynamic. Through the eyelid mesh contraction algorithm, the upper eyelid is gradually closed from 80% of the reference position to fully closed, and then gradually opened, while synchronously adjusting the muscles around the eyes; Frames 69-119: Continuing the head shaking dynamic, the facial muscle mesh deformation algorithm is used to fine-tune facial expressions, such as slightly raising the corners of the mouth by 0.5mm; Frames 120-128: Trigger the blink action again; Frames 129-149: The head gradually returns to the reference posture; S35, further comprising performing frame-by-frame comparison on the generated initial video; The structural similarity index algorithm is used to detect the difference between adjacent frames to ensure that the change between frames is ≤5%. If it exceeds the threshold, re-interpolation and adjustment are performed; The core feature position of the portrait in each frame is calibrated using a feature point tracking algorithm to ensure that the deviation of the core feature position is ≤ 1 pixel, ensuring that the digital human image is always consistent with the portrait features; S36, further comprising using a lip synchronization algorithm to accurately match the lip shape of the digital human model with the dubbing; S4. Inputting the reference speech features and the content reference information features into a preset first language model to generate video text and storyboard information; In this embodiment, the first large language model is the Qwen3 vertical domain large model fine-tuned by Lora; S41, preprocessing the reference speech features and the content reference information features; In this embodiment, the preprocessing includes normalizing the reference speech features and structurally encoding the content reference information features; The standardization specifically involves converting audio parameters (such as sampling rate, number of channels, and amplitude range) of the reference speech features into a feature format compatible with the first language model, thereby eliminating format differences between different speech inputs. Specifically, if the content reference information is text (such as documents or webpage text), the core semantic features are extracted through text segmentation and entity recognition and converted into vector format; if the content reference information is image or video clip type reference information, visual features (such as key frame features and scene labels) are extracted and mapped into structured text description vectors; S42: Input the pre-processed content reference information into the generation layer of the first language model to generate an initial video copy; the initial video copy includes narrative logic, language style parameters, etc. S43, inputting the preprocessed reference speech features and content reference information features into the feature fusion layer of the first language model to generate fusion features; The attention mechanism is used to calculate the association weights between the preprocessed reference speech features and the content reference information features, and the two feature vectors are weightedly fused to obtain a fusion feature that includes both speech style and content core. S44. Generate storyboard information based on the initial video text and the fusion features; the storyboard information includes scene description, visual element requirements, and temporal relationships; S5. Matching is performed in the local video material library based on the multimodal features to obtain matching materials, wherein the matching process is to calculate the similarity between the current multimodal features and the video materials in the local video material library; if the current similarity is greater than a preset similarity threshold, it indicates that the match is successful, and the current video material is recorded as the matching material; S6. Input the storyboard information and matching materials into the second language model to generate a video content schedule; the video content schedule defines the time points of element appearance, duration, and superposition relationship; In this embodiment, the second largest language model includes GPT-4 Turbo, Qwen-72B, Gemini Pro Vision, etc. S61, the training process of the second largest speech model includes constructing a training data set and a fine-tuning method based on reinforcement learning; S611, the training data set specifically includes: S6111. Acquire multiple video templates according to a preset video type ratio, and parse the video templates to obtain a tuple structure; In this embodiment, the preset video type ratio is 40% for promotional videos, 35% for news broadcasts, and 25% for educational commentary. The tuple structure includes a text description of the storyboard, a storyboard sequence, the type of visual element, the time point of appearance and duration, a time mark of the audio track, the time of appearance of the subtitle text, a shot switching sequence, a background music beat sequence, etc.; S6112: A preset number of video editing experts use a preset professional review mechanism to cross-score the tuple structures, and select tuple structures whose average cross-score exceeds a preset screening threshold; In this embodiment, five experienced video editing experts were selected to independently cross-score the tuple structures, where the score range is 1-10 points; only the tuple structures with a comprehensive average score ≥ 7.0 were retained; The comprehensive average score is expressed as follows: ; Where, Indicates the Rating indicators; Indicates the A video editing expert for the Ratings made based on the rating indicators; Indicates the The index value of the scoring indicator; Indicates the Index value of a video editing expert; represents the number of video editing experts; S6113. Construct a training data set based on the video template corresponding to the filtered tuple structure; S612: The reinforcement learning-based fine-tuning method specifically comprises fine-tuning the second language model using a reward function, wherein: The reward function is calculated based on the coherence reward, rhythm reward, synchronization reward and aesthetic reward and their corresponding weights, and is expressed as follows: ; ; Where, express Momentary status Take action The reward function after ; Indicates a coherence reward; represents the weight of the coherence reward; Indicates rhythm reward; represents the weight of the rhythm reward; Indicates synchronization reward; Represents the weight of synchronization reward; indicates aesthetic rewards; represents the weight of aesthetic reward; The coherence bonus is calculated based on the text description of each storyboard and is expressed as: ; Where, Indicates the length of the storyboard sequence; Indicates the A text description of each storyboard; Indicates the A text description of each storyboard; Indicates the The index value of a storyboard; Represents the BERT text embedding function; The rhythm reward is calculated based on the video beat sequence and shot switching sequence and is expressed as follows: ; Where, represents the dynamic time warping distance algorithm; Represents the background music beat sequence; Represents a shot switching sequence; represents the preset hyperparameter used to control the sensitivity of the rhythm reward function to the dynamic time warping distance; The synchronization reward is calculated based on the time difference between subtitles and voice, and is expressed as follows: ; Where, Indicates the The appearance time of each subtitle; Indicates the The appearance time of each subtitle corresponding to the voice; Indicates the The index value of a subtitle; Indicates the number of subtitles; represents the time tolerance, which is 0.1 in this embodiment; The aesthetic reward is calculated based on the key frame image and is expressed as: ; Where, Indicates that the parameter is A neuroaesthetic scoring model; represents a key frame image; S62: Input the storyboard information and matching materials into the second language model. The specific steps are as follows: S621: Input the matching material into the feature extraction layer of the second language model to obtain matching features, such as the shot type label, picture resolution parameters, audio track features, etc. of the video clip; S622: Input the storyboard information and matching features into the feature alignment layer of the second language model, use the temporal attention mechanism to calculate the matching degree between the storyboard timing requirements and the time length of the matching material, and generate temporal correlation features; S623: Generate a video content schedule based on the temporal correlation features; The video content schedule includes global information, storyboard details, transition effects, auxiliary element windows, etc. The global information includes total duration, frame rate, and time accuracy; The storyboard details include each storyboard's ID, matching material ID and type, start time / end time in the global timeline, and effective content start and end time; The transition effects include the type of each transition, the involved shots, and the start and end times; The auxiliary element window includes recording the start and end time and location / area description of each window by type; In this embodiment, the specific steps of generating a video content schedule are: The global timeline is established based on the storyboard sequence: the theoretical start time of the first storyboard is 0 seconds, and the end time is the preset duration; The theoretical starting time of a storyboard = The theoretical end time of a storyboard is: end time = start time + preset duration; Map the matching material to the theoretical time interval of the corresponding storyboard, and handle the difference in duration: if the matching material duration is less than the preset storyboard duration, it will be played back at a uniform slow speed in proportion while maintaining the smoothness of the picture. The shortfall will be filled with a fade-in / fade-out transition of equal duration at the beginning and end of the material, where the transition effect is a semi-transparent gradient of the first / last frame of the material; if the matching material duration is longer than the preset storyboard duration, the "keyframe matching degree" in the timing association feature will be used to intercept the clip that best fits the storyboard theme; Add transition effects at the junction of adjacent storyboards. The transition duration is fixed at 0.5 seconds, occupying 0.2 seconds from the end of the previous storyboard and 0.3 seconds from the beginning of the next storyboard. Select the special effect type based on the relevance of the storyboard content. For example, if the comprehensive matching score in the storyboard material matching matrix is ≥80, use "fade in and out", 50-80 points use "slide left and right", and <50 points use "zoom in and out". The mapping matrix is the total number of matching materials), the matrix element value represents the The first storyboard and the The comprehensive matching score of the matching materials is used for the time adaptation and content relevance of each storyboard and the material, where Indicates the The index value of the matching material; S624, further comprising adjusting the preliminary schedule in combination with preset reward function-related verification rules; In this embodiment, the verification rules include coherence verification and rhythm verification; The consistency check is to see whether the scene correlation of adjacent materials complies with the storyboard logic; The rhythm check is to check whether the material switching frequency matches the potential audio beat, and to correct timing conflicts, such as material duration exceeding the allocated duration of the storyboard, and element superposition anomalies, such as display level conflicts of multiple materials in the same frame; S7, synthesize the video based on the video content schedule, the initial video, and the video copy, and output the final synthesized video; S8. The method further includes performing multi-track rendering and packaging on the final synthesized video using a video coding protocol.
[0024] Example 2: This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, a video synthesis method based on a large language model as described in any one of Embodiment 1 is implemented.
[0025] Example 3: This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the method for video synthesis based on a large language model as described in any one of the embodiments 1 is implemented.
[0026] In the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can represent: a, b, c, a and b, a and c, b and c or a and b and c, where a, b, c can be single or multiple.
[0027] Those skilled in the art will appreciate that the various units and algorithm steps described in the embodiments disclosed herein can be implemented using a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0028] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0029] In the several embodiments provided in this application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of this application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory; hereinafter referred to as: ROM), random access memory (Random Access Memory; hereinafter referred to as: RAM), magnetic disk or optical disk, and other media that can store program code.
[0030] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention's description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A video synthesis method based on a large language model, characterized in that: The method comprises: Obtaining user input information, including portrait images, reference voices, and content reference information; Performing multimodal feature extraction on the input information to obtain multimodal features, including portrait image features, reference voice features, and content reference information features; Input the portrait image features into the digital human model built based on deep learning to generate the initial video; Input the reference speech features and content reference information features into the preset first language model to generate video copy and storyboard information; Matching is performed in the local video material library based on multimodal features to obtain matching materials; The storyboard information and matching materials are input into the second language model to generate a video content timeline; Perform video synthesis based on the video content schedule, initial video, and video copy, and output the final synthesized video.
2. The video synthesis method based on a large language model according to claim 1, characterized in that: Input the portrait image features into the digital human model built based on deep learning. The specific steps are as follows: Normalizing the portrait features specifically involves converting features of different dimensions into a tensor format compatible with the digital human model; The standardized portrait features are input into the feature fusion layer of the digital human model to identify the unique identifiers in the portrait features; Embedding the unique identifier into a basic mesh of the digital human and outputting a three-dimensional mesh model of the portrait image; Generate a first frame of video image from the three-dimensional mesh model according to a preset initial posture, initial expression and basic lighting; The first frame of video image is combined with a preset video dynamic frame sequence to generate an initial video.
3. The video synthesis method based on a large language model according to claim 1, characterized in that: The specific steps of inputting the reference speech features and the content reference information features into the preset first language model are: Preprocessing the reference speech features and content reference information features; Input the pre-processed content reference information into the generative layer of the first language model to generate the initial video copy; Inputting the preprocessed reference speech features and content reference information features into the feature fusion layer of the first language model to generate fused features; Generate storyboard information based on the initial video text and fusion features.
4. The video synthesis method based on a large language model according to claim 1, characterized in that: Matching is performed in the local video material library based on multimodal features, where the matching process is to calculate the similarity between the current multimodal features and the video materials in the local video material library; if the current similarity is greater than the preset similarity threshold, it means that the match is successful, and the current video material is recorded as the matching material.
5. The video synthesis method based on a large language model according to claim 1, characterized in that: The training process of the second largest speech model includes constructing a training dataset and fine-tuning methods based on reinforcement learning; The training data set specifically includes: Acquire multiple video templates according to a preset video type ratio, and parse the video templates to obtain a tuple structure; A preset number of video editing experts use a preset professional review mechanism to cross-score the tuple structures and select tuple structures whose average cross-score exceeds a preset screening threshold; A training dataset is constructed based on the video templates corresponding to the filtered tuple structures.
6. The video synthesis method based on a large language model according to claim 5, characterized in that: The reinforcement learning-based fine-tuning method specifically fine-tunes the second language model through a reward function, where: The reward function is calculated based on the coherence reward, rhythm reward, synchronization reward and aesthetic reward and their corresponding weights; A coherence bonus is calculated based on the text description of each storyboard; Calculate the rhythm reward based on the video beat sequence and shot switching sequence; Synchronization rewards are calculated based on the time difference between subtitles and voice; Aesthetic rewards are calculated based on keyframe images.
7. The video synthesis method based on a large language model according to claim 1, characterized in that: Input the storyboard information and matching materials into the second language model. The specific steps are as follows: Input the matching material into the feature extraction layer of the second largest language model to obtain matching features; The storyboard information and matching features are input into the feature alignment layer of the second language model. The temporal attention mechanism is used to calculate the matching degree between the storyboard timing requirements and the time length of the matching material, and generate temporal correlation features. Generate video content schedule based on temporal correlation features.
8. The video synthesis method based on a large language model according to claim 1, characterized in that: The method also includes performing multi-track rendering and packaging on the final composite video using a video coding protocol.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Training-free movie-level long video generation method and system
CN122138019A