Video generation control method and computer readable storage medium
Through a video generation control method, storyboard big models and multiple large model technologies are used to generate film and television content with rich emotional logic and consistent characteristics, which solves multiple technical problems of existing AIGC technology in the film and television field, and significantly improves the efficiency and quality of video content production.
Patent Information
- Application Number
- CN202510496118.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing AIGC technology faces the problems of lack of emotional logic and adaptability of lens language in the film and television field of storyboard script generation, poor character face consistency, difficulty in fusion of scene styles, inability to achieve accurate synchronization of emotional dialogue and sound effects, and insufficient controllability of video generation algorithms, resulting in low efficiency and quality of video content production.
Through a video generation control method, a storyboard script is generated using a storyboard model, and a storyboard diagram, voice and video that is consistent with the reference image, audio and video features are generated. The specific steps include: obtaining the video creative text, input it into the storyboard model to generate the storyboard script and associated prompt words, generating corresponding video, audio and image content through different big models based on the storyboard script and prompt words, and dynamically correlating the voice to be synthesized with the character mouth-shaped features in the storyboard video to generate the target video.
It effectively solves the problems of lack of emotional logic, poor consistency of storyboard characteristics, difficulty in audio synchronization and insufficient controllability of video generation in AIGC technology in film and television creation, and improves the efficiency and quality of video content production.
Smart Images

Figure CN120017931A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a video generation control method and a computer-readable storage medium. Background Art
[0002] At present, Artificial Intelligence Generated Content (AIGC) technology has gradually been integrated into the entire process of film and television production, significantly improving creative efficiency.
[0003] However, existing AIGC technology faces multiple technical challenges in the film and television industry: 1. Storyboard script generation lacks emotional logic and lens language adaptation capabilities, and mechanical output leads to the lack of emotional progression structure; 2. Static storyboards have problems such as poor consistency of character faces and difficulty in integrating scene styles; 3. Speech synthesis and sound effect generation technology cannot achieve accurate synchronization of emotional dialogue and sound and picture; 4. The video generation algorithm is not controllable enough, manifested in defects such as lip alignment deviation and low multi-subject fidelity, which restricts the improvement of content quality. The above technical difficulties reduce the efficiency and quality of video content production.
[0004] Therefore, how to overcome the technical difficulties in applying AIGC technology in the film and television field has become the key to improving the efficiency and quality of video content production. Summary of the invention
[0005] In view of the above problems, the present invention provides a video generation control method and a computer-readable storage medium that overcome the above problems or at least partially solve the above problems. The technical solution is as follows:
[0006] A video generation control method, comprising:
[0007] Get creative text for your video;
[0008] Input the video creative text into the storyboard big model to obtain the storyboard script and related prompt words output by the storyboard big model, wherein the related prompt words include storyboard picture prompt words, voice emotion prompt words and storyboard video prompt words;
[0009] Based on the storyboard script, the storyboard prompt words and the reference image, generating a storyboard consistent with the image features of the reference image through an image macro model;
[0010] Based on the storyboard, the speech emotion prompt words and the reference audio, the speech text to be converted in the storyboard is converted into a speech to be synthesized that is consistent with the speech features of the reference audio and contains emotional expression through a speech big model;
[0011] Based on the storyboard and the storyboard video prompt words, the storyboard is generated into a storyboard video through a video macro model;
[0012] The speech to be synthesized is dynamically associated with the mouth shape features of the character in the storyboard video to generate a target video.
[0013] Optionally, after dynamically associating the speech to be synthesized with the mouth shape features of the character in the storyboard video to generate a target video, the method further includes:
[0014] Obtaining a sound effect that matches the sound effect text to be generated in the storyboard;
[0015] Identifying key dynamic time points in the target video;
[0016] The sound effect and the key dynamic time point are dynamically aligned using a self-attention algorithm to obtain the target video after synthesizing the sound effect.
[0017] Optionally, the training process of the storyboard model includes:
[0018] Obtaining a storyboard script dataset, wherein the storyboard script dataset includes a plurality of storyboard script samples obtained by reverse deconstructing at least one video;
[0019] The basic large model is supervised and fine-tuned using a plurality of the storyboard samples in the storyboard data set to obtain a storyboard large model.
[0020] Optionally, the reference image includes a reference character image and a reference scene image, and the step of generating a storyboard consistent with image features of the reference image through an image macro model based on the storyboard script, the storyboard prompt words and the reference image includes:
[0021] The storyboard script, the storyboard prompt words, the reference character image and the reference scene image are input into a large image model, so that the large image model controls the image generation process and the image editing process corresponding to the generation result of each storyboard screen description in the storyboard script based on the reference character image and the reference scene image according to the storyboard prompt words, and outputs a storyboard whose character appearance features are consistent with the reference character image and whose background features are consistent with the reference scene image.
[0022] Optionally, the storyboard prompt words include image generation prompt words and image editing prompt words, and the image generation process includes: extracting the character appearance features of the reference character image and the background features of the reference scene image, and then converting the character appearance features and the background features into first embedded features according to the image generation prompt words, and after processing the first embedded features using a first diffusion model, combining a self-attention visual model and a second diffusion model to fuse and replace the character clothing area based on the reference character image, so as to obtain an image to be adjusted in which the character appearance features are consistent with the reference character image and the background features are consistent with the reference scene image,
[0023] And / or, the image editing process includes: identifying the area to be edited of the image to be adjusted, extracting the image area features of the area to be edited, and then converting the image area features into a second embedded feature based on the image editing prompt word, and after processing the second embedded feature using an image detail adjustment model, generating a storyboard in which the character appearance features are consistent with the reference character image and the background features are consistent with the reference scene image.
[0024] Optionally, based on the storyboard script, the speech emotion prompt words and the reference audio, converting the speech text to be converted in the storyboard script into speech to be synthesized that is consistent with the speech features of the reference audio and contains emotional expression through a speech big model, comprises:
[0025] The storyboard script, the speech emotion prompt words and the reference audio are input into a speech big model, so that the speech big model extracts the text features of the speech text to be converted in the storyboard script and the speech features of the reference audio, the text features and the speech features are input into a self-attention speech model to generate a basic speech, and according to the speech emotion prompt words, the emotional expression of the basic speech is controlled through a third diffusion model to obtain a speech to be synthesized that is consistent with the speech features of the reference audio and contains emotional expression.
[0026] Optionally, the generating the storyboard into a storyboard video by using a video macro model based on the storyboard and the storyboard video prompt words includes:
[0027] The storyboard and the storyboard video prompt words are input into a large video model, so that the large video model extracts the bone sequence and three-dimensional mesh hand features of the preset action sequence in the storyboard according to the storyboard video prompt words, the bone sequence and the three-dimensional mesh hand features are encoded by a posture encoder and then input into a denoising network, and a bone scaling strategy based on three-dimensional bone length estimation is used to dynamically adjust the joint spacing of the bone sequence, so that the bone topological structure of the preset action sequence forms a spatial correspondence with the bone proportion of the character instance in the storyboard, and a storyboard video that maintains motion consistency with the posture of the character in the storyboard is output.
[0028] Optionally, the generating the storyboard into a storyboard video by using a video macro model based on the storyboard and the storyboard video prompt words includes:
[0029] The storyboard and the storyboard video prompt words are input into the video big model, so that the video big model combines the image contents of multiple subjects in the storyboard and the text prompts for multiple subjects in the storyboard video prompt words through the dual cross-attention layer of the diffuse self-attention architecture during the generation of each frame of the storyboard video, thereby generating a storyboard video containing multiple subjects.
[0030] Optionally, dynamically associating the speech to be synthesized with the mouth shape features of the character in the storyboard video to generate a target video includes:
[0031] Extracting the prosodic features of the speech to be synthesized by a speech encoder, wherein the speech encoder is constructed using a multilingual speech recognition model based on a self-attention mechanism;
[0032] Extracting mouth shape features of the character's face from the storyboard video;
[0033] Establishing a cross-modal cross-attention layer in the denoising network, calculating the attention weight matrix of the prosodic features and the character's mouth shape features, and dynamically adjusting the intensity parameters of the speech-driven mouth shape;
[0034] Applying a pre-trained facial mask generator to generate a spatially constrained template, and dynamically masking a non-mouth area in a feature space of the storyboard video;
[0035] The character's mouth shape features adjusted according to the intensity parameter are fused to the mouth area in the feature space of the storyboard video to generate a target video in which the character's mouth shape and voice are synchronized.
[0036] A computer-readable storage medium stores a program, and when the program is executed by a processor, the video generation control method is implemented.
[0037] By means of the above technical scheme, the present invention provides a video generation control method and a computer-readable storage medium, the method comprising: obtaining a video creative text; inputting the video creative text into a storyboard big model, obtaining a storyboard script and associated prompt words output by the storyboard big model, wherein the associated prompt words include storyboard picture prompt words, voice emotion prompt words and storyboard video prompt words; based on the storyboard script, the storyboard picture prompt words and the reference image, generating a storyboard consistent with the image features of the reference image through the image big model; based on the storyboard script, the voice emotion prompt words and the reference audio, converting the voice text to be converted in the storyboard script into a voice to be synthesized that is consistent with the voice features of the reference audio and contains emotional expression through the voice big model; based on the storyboard and the storyboard video prompt words, generating the storyboard into a storyboard video through the video big model; dynamically associating the voice to be synthesized with the character mouth shape features in the storyboard video to generate a target video. The present invention solves the structural defects of storyboard scripts through a large storyboard model, ensures the consistency of storyboard features based on image generation constrained by reference image features, realizes audio and video synchronization by combining emotional speech synthesis with dynamic lip alignment, and improves video fidelity by guiding the multi-agent generation algorithm through storyboard video prompt words, thereby effectively solving the technical problems existing in the existing AIGC technology in film and television creation, such as lack of emotional logic, poor consistency of storyboard features, difficulty in audio synchronization and insufficient controllability of video generation, and improving the efficiency and quality of video content production.
[0038] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented according to the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0040] Figure 1 A schematic diagram showing a flow chart of an implementation of a video generation control method provided by an embodiment of the present invention;
[0041] Figure 2 A schematic flow chart showing a first specific implementation of a video generation control method provided by an embodiment of the present invention;
[0042] Figure 3 A logic block diagram of a sound effect synthesis process provided by an embodiment of the present invention is shown;
[0043] Figure 4A schematic flow chart showing a second specific implementation of the video generation control method provided by an embodiment of the present invention;
[0044] Figure 5 A flowchart of the work flow of the storyboard model provided by an embodiment of the present invention is shown;
[0045] Figure 6 A flowchart of an image generation and editing process provided by an embodiment of the present invention is shown;
[0046] Figure 7 A flowchart of a speech synthesis process provided by an embodiment of the present invention is shown;
[0047] Figure 8 A schematic flow chart showing a third specific implementation of the video generation control method provided by an embodiment of the present invention;
[0048] Fig. 9 A schematic diagram showing the structure of a video generation control device provided by an embodiment of the present invention is shown;
[0049] Fig.10 A schematic structural diagram of an electronic device provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0050] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present invention and to enable the scope of the present invention to be fully communicated to those skilled in the art.
[0051] Currently, AI-generated content (AIGC) technology is widely used in the film and television creation process, including AI-assisted script analysis to help screenwriters explore potential story logic and character relationships. At the same time, in special effects production, AI's ability to generate virtual scenes and characters has been significantly improved, reducing the time and cost of traditional special effects production, making it possible to present alien scenes and fantasy creatures in some science fiction blockbusters more efficiently and realistically.
[0052] However, although AIGC technology has shown great potential in film and television production, there are still many technical problems that need to be solved in the field of video content production:
[0053] 1. Insufficient expression of storyboard content: Currently, AI-generated storyboards are often mechanical, lacking emotional expression and film and television creation logic. Especially in the processing of emotional turning points, it is impossible to build a progressive emotional progression structure that conforms to the laws of film and television creation. In terms of lens language generation, there is a lack of reasonable lens rhythm and spatial coherence, and the style adaptation ability is also insufficient.
[0054] 2. Poor consistency between characters and scenes: In the process of generating static storyboards, it is difficult for AI to ensure the consistency of the characters, objects and styles of the storyboards. In particular, there are obvious technical difficulties in terms of face consistency, background consistency and clothing area fusion.
[0055] 3. Lack of emotional dialogue and intelligent sound effects: Traditional speech synthesis and sound effect generation methods cannot meet the needs of video content production for emotional dialogue and intelligent sound effects. There are difficulties in controllable zero-sample multi-emotion speech synthesis, as well as challenges in accurately synchronizing sound effect generation with video dynamics.
[0056] 4. Low controllability of video generation algorithms: When generating character videos, traditional video generation algorithms often have problems such as poor lip alignment and mismatch between input images and preset action sequence skeletons. In multi-subject video generation, the fidelity of the protagonist and the background diversity are insufficient, and there is a lack of effective technical means to improve the quality and controllability of generated videos.
[0057] 5. Traditional content production is inefficient and costly: In the traditional video content production process, script polishing takes a long time, shooting scene construction is complex, and post-editing and coordination are cumbersome, resulting in high labor and time costs, making it difficult to meet the market's rapid demand for diversified and high-quality content.
[0058] In order to solve the above problems, the present invention proposes a video generation control method, focusing on four core aspects: the expressiveness of storyboard script content, the consistency of storyboard features, emotional dialogue and intelligent sound effects, and controllable video generation algorithm, aiming to improve the efficiency and quality of video content production and reduce production costs.
[0059] like Figure 1 As shown, a flow chart of an implementation of a video generation control method provided by an embodiment of the present invention is provided, and the method may include:
[0060] S100, obtain video creative text.
[0061] The creative text of a video refers to the text containing the script content, creative ideas and related descriptions, which is used to provide guidance and framework for video production. The creative text of a video can involve detailed descriptions of the storyline, characters, scenes, themes, visual style, sound effects and music, narrative structure, dialogue, etc.
[0062] Specifically, the embodiment of the present invention can obtain the creative text of the video uploaded directly by the user, and can also read the creative text of the video input by the user in real time through the online editing tool.
[0063] S110, inputting the creative text of the video into the storyboard big model, obtaining the storyboard script and associated prompt words output by the storyboard big model, wherein the associated prompt words include storyboard picture prompt words, voice emotion prompt words and storyboard video prompt words.
[0064] Among them, the storyboard model is an artificial intelligence model that can convert the input video creative text into a storyboard script after supervised fine-tuning (SFT) of the base model (BaseModel).
[0065] Among them, the storyboard is a detailed script generated based on the creative text of the video, including the scene number, shot number, duration, scene, picture content, scene size, perspective, camera movement and dialogue of each shot. In the storyboard, there is a unified description format for characters, scenes and props. For example: the description of a character can include height, body shape, hairstyle, facial features, personality, skin color and clothing. The description of a scene can include scene classification, geographical location, environmental atmosphere and time. The description of a prop can include attributes, size, color, material and function.
[0066] Among them, the associated prompt word (Prompt) refers to the text or instructions output by the storyboard big model to guide the big model to supplement and describe various aspects of the storyboard script. The associated prompt word is used to guide the visual, sound and emotional expression in the subsequent video content production.
[0067] The storyboard prompts are text or instructions for generating a specific lens image, which help draw or generate the visual effects of each lens. The storyboard prompts usually describe the composition, angle and color of the lens.
[0068] Among them, the voice emotion prompt words are texts or instructions used to identify the emotions and intonations that a character or narrator should express in a certain scene. The voice emotion prompt words guide the adjustment of the voice pitch, speed and emotion to fit the story context.
[0069] Among them, storyboard video prompts are texts or instructions that guide the visual effects and action performance of specific shots or scenes in video production. Storyboard video prompts can involve aspects such as shot switching, special effects use and action design to ensure the smoothness and visual appeal of the video.
[0070] S120, based on the storyboard script, storyboard prompt words and reference image, generate a storyboard consistent with the image features of the reference image through the image macro model.
[0071] Reference images refer to existing images used to guide the creation or generation of new images. Reference images can be existing works of art, photos, or scenes that contain specific visual elements and styles.
[0072] Optionally, the reference image includes a reference character image and a reference scene image.
[0073] Reference character images refer to existing images used to guide character design and performance. Reference character images are used to provide visual references for the character's appearance, style, and details to ensure consistency in the character's storyboard.
[0074] Reference scene images refer to existing images used to guide scene design and layout. Reference scene images usually show a specific environment, geographical location, or atmosphere. Reference scene images are used to provide references for the layout, color, light, and atmosphere of the scene, ensuring that the generated storyboards are consistent with the desired scene effects.
[0075] The image model is an artificial intelligence model based on deep learning technology, which can generate storyboards that meet the requirements through multimodal inputs such as text and images. The image model can simultaneously analyze the style features of the storyboard script, storyboard prompts, and reference images to achieve cross-modal alignment of image and text information.
[0076] Among them, the storyboard is a visual image generated by the image model based on the storyboard script, storyboard prompts and reference images, representing a specific shot in a film or animation. The storyboard usually depicts the main elements of the scene in a simplified way, including characters, backgrounds, actions and camera angles.
[0077] It is understandable that a storyboard script usually includes multiple storyboard screen description generation results, and the image model can generate corresponding storyboards for each storyboard screen description generation result according to the instructions of the corresponding storyboard prompt words.
[0078] For ease of understanding, an example is given here: Assume that the storyboard script includes the storyboard description generation result A (the camera looks down at the park from a high angle, children are playing on the grass, and parents are chatting around) and the storyboard description generation result B (close-up, a little girl is chasing a butterfly, smiling happily). The storyboard prompt for the storyboard description generation result A is "overhead view, green grass, children are playing, bright sunshine, and tree background", and the storyboard prompt for the storyboard description generation result B is "close-up, little girl, happy expression, butterfly flying in the air". The selected reference image is a photo showing a sunny environment, a scene of children playing, and trees and fountains in the background. The storyboard script, the storyboard prompt and the reference image are input into the image macro model. The storyboard generated by the image macro model for the storyboard description generation result A can be an image of a high-angle view of the park, including a scene of children playing on the grass, and a background of trees and fountains. The storyboard generated for the storyboard description generation result B can be a close-up of the little girl chasing a butterfly, showing the happiness and excitement on her face.
[0079] S130, based on the storyboard script, speech emotion prompt words and reference audio, the speech text to be converted in the storyboard script is converted into a speech to be synthesized that is consistent with the speech features of the reference audio and contains emotional expression through a speech big model.
[0080] Reference audio is an existing audio clip used to guide speech synthesis. Reference audio usually shows specific sound characteristics, emotional expression, intonation and speaking speed, and can be natural dialogue, narration or character dubbing.
[0081] The speech big model is an artificial intelligence model based on deep learning technology, which can generate speech that meets the requirements through multimodal inputs such as text and audio. The speech big model can convert text into natural and fluent speech according to speech emotion prompts, and can simulate different speech features and emotional expressions.
[0082] The speech text to be converted is the text content contained in the storyboard, which is usually extracted from the storyboard by speech text recognition technology. The speech text to be converted is the text to be converted into audio by speech synthesis technology. The speech text to be converted usually includes character dialogue, narration or commentary.
[0083] The speech to be synthesized refers to the speech output generated after being processed by the speech model based on the text to be synthesized in the storyboard. The speech to be synthesized is consistent with the reference audio in terms of timbre, speaking speed and emotion.
[0084] For ease of understanding, an example is given here: the speech text to be converted corresponding to the storyboard description generation result of "a little girl (named Xiaoling) and her dog (named Xiaobai) playing Frisbee on the grass" in the storyboard script is "Xiaoling: "Xiaobai, come and catch the Frisbee!" (excited). Xiaobai (responds with a cute dog voice): "Woof woof!" (friendly)", and the corresponding speech emotion prompt words are "Xiaoling's dialogue: excited, happy, and full of expectations. Xiaobai's response: cute, friendly, and excited". The speech features included in the selected reference audio are "the little girl's voice is crisp and bright, with a moderate speaking speed, and the emotional expression is enthusiastic. The dog's voice simulates a lively and cute tone." The storyboard script, the speech emotion prompt words and the reference audio are input into the speech large model, and the generated speech to be synthesized includes the voice generation result of Xiaoling (showing an excited and happy tone, a crisp tone, and a moderate speaking speed) and the voice generation result of Xiaobai (showing a cute and friendly tone, with a slightly playful voice).
[0085] S140, based on the storyboard and the storyboard video prompt words, generate the storyboard into a storyboard video through the video macro model.
[0086] The video model is an artificial intelligence model based on deep learning technology, which can generate video sequences that meet the requirements through multimodal inputs such as text and images. The video model can generate new dynamic video sequences based on the storyboards according to the instructions of the storyboard video prompts, that is, generate dynamic effects on static storyboards.
[0087] Among them, the storyboard video is a dynamic video clip generated based on the storyboard. The storyboard video can show the storyline, character actions and scene changes.
[0088] For ease of understanding, an example is given here: assuming that the storyboard includes "a little boy sleeping in bed with toys scattered around the room" and "a world of dreams, colorful clouds and stars, a little boy and a robot flying", and the corresponding storyboard video prompts are "a quiet night, warm lights" and "dream colors, cheerful melodies, a feeling of flying", the storyboard and storyboard video prompts are input into the large video model, and the generated storyboard videos are "showing a little boy sleeping quietly in bed with soft lights in the room" and "a little boy and a robot flying in a colorful dream with clouds and stars surrounding them".
[0089] S150, dynamically associating the speech to be synthesized with the mouth shape features of the character in the storyboard video to generate a target video.
[0090] Among them, the character's mouth shape features refer to the movement and shape of the character's mouth when pronouncing in the storyboard video. The character's mouth shape features can include the degree of mouth opening and closing and the shape of the lips. The character's mouth shape features are usually combined with the character's expression and action to enhance the emotional expression of the character in the storyboard video.
[0091] Among them, dynamic association refers to synchronizing the speech to be synthesized with the character's mouth shape features during the storyboard video generation process, so that the character's mouth movement is coordinated with the audio content of the speech to be synthesized.
[0092] The invention provides a video generation control method, which comprises: obtaining a video creative text; inputting the video creative text into a storyboard big model, obtaining a storyboard script and associated prompt words output by the storyboard big model, wherein the associated prompt words include storyboard picture prompt words, voice emotion prompt words and storyboard video prompt words; based on the storyboard script, the storyboard picture prompt words and a reference image, generating a storyboard picture that is consistent with the image features of the reference image through an image big model; based on the storyboard script, the voice emotion prompt words and a reference audio, converting the voice text to be converted in the storyboard script into a voice to be synthesized that is consistent with the voice features of the reference audio and contains emotional expression through a voice big model; based on the storyboard picture and the storyboard video prompt words, generating the storyboard picture into a storyboard video through a video big model; dynamically associating the voice to be synthesized with the mouth shape features of a character in the storyboard video, and generating a target video. The present invention solves the structural defects of storyboard scripts through a large storyboard model, ensures the consistency of storyboard features based on image generation constrained by reference image features, realizes audio and video synchronization by combining emotional speech synthesis with dynamic lip alignment, and improves video fidelity by guiding the multi-agent generation algorithm through storyboard video prompt words, thereby effectively solving the technical problems existing in the existing AIGC technology in film and television creation, such as lack of emotional logic, poor consistency of storyboard features, difficulty in audio synchronization and insufficient controllability of video generation, and improving the efficiency and quality of video content production.
[0093] Optional, based on Figure 1 The method shown, such as Figure 2 As shown, a flowchart of a first specific implementation of the video generation control method provided by an embodiment of the present invention is shown. After step S150, the method may further include:
[0094] S200, obtaining a sound effect that matches the sound effect text to be generated in the storyboard script.
[0095] The sound effect text to be generated is the text content contained in the storyboard script for guiding the generation of sound effects, which is usually extracted from the storyboard script through sound effect text recognition technology. The sound effect text to be generated describes the sound effect elements that need to be reflected in the scene, which may include natural sounds, environmental sounds, character dialogues, and sound effects, etc., in order to guide the creation and synthesis of sound effects during the video content production process.
[0096] The embodiment of the present invention can directly search for sound effects that match the sound effect text to be generated in the storyboard script from the sound effect material library, and can also synthesize matching sound effects in real time based on the sound effect text to be generated and the storyboard video.
[0097] S210: Identify key dynamic time points in the target video.
[0098] Among them, key dynamic time points refer to specific time positions in the video sequence that are related to important events, plot turns or emotional peaks, such as the time points when the protagonist of a costume short drama draws his sword to fight and the character sheds tears.
[0099] S220, using a self-attention algorithm to dynamically align the sound effects and key dynamic time points to obtain a target video after synthesizing the sound effects.
[0100] Specifically, the embodiments of the present invention can analyze the dynamic features of the video at different time points and their interrelationships through the self-attention mechanism. For example, in a chase scene, pay attention to the character's running speed, direction changes, and dynamic elements of the surrounding environment. According to the analysis results of key dynamic time points and the self-attention mechanism, the required sound effects are accurately matched, such as adding a sword sound effect when the protagonist draws the sword, and adding soft sad music when the character cries. In addition, parameters such as the playback time and volume of the sound effects can also be adjusted to achieve accurate synchronization of the sound effects and video dynamics, significantly improving the audience's audio-visual experience.
[0101] Figure 3 The figure shows a logical block diagram of the sound effect synthesis process provided by an embodiment of the present invention. The embodiment of the present invention can combine noise, the sound effect text to be generated and the video to be dubbed, convert the sound effect text to be generated into a text semantic feature sequence, and convert the dubbing video into a video feature sequence for subsequent Transformer processing. The dynamic alignment module is used to identify the key dynamic time points in the video, and the self-attention mechanism is used for analysis to ensure the precise alignment of the sound effect and the video dynamics. The multimodal Transformer fuses the text semantic feature sequence and the video feature sequence to extract cross-modal features to achieve comprehensive modeling of information. The unimodal Transformer performs deep modeling on the cross-modal features to extract more representative sound effect features. The variational decoder converts the extracted sound effect features into synthetic sound effects that are precisely synchronized with the video content.
[0102] The embodiments of the present invention improve the expressiveness and naturalness of video sound effects through accurate dynamic alignment of key dynamic time points of sound effects and videos, thereby improving the efficiency and quality of video content production.
[0103] Optional, based on Figure 1 The method shown, such as Figure 4 As shown, a flow chart of a second specific implementation of the video generation control method provided by an embodiment of the present invention is provided. The training process of the storyboard large model may include:
[0104] S400. Obtain a storyboard script dataset, wherein the storyboard script dataset includes a plurality of storyboard script samples obtained by reverse deconstructing at least one video.
[0105] Reverse deconstruction refers to the process of analyzing and disassembling existing video content to extract its structure, elements, and creative intent. The reverse deconstruction process usually involves dividing the video into multiple shots or scenes and analyzing the characters, composition, action, dialogue, sound effects, and other elements of each shot.
[0106] The embodiment of the present invention can pre-use a multimodal content understanding model to parse the video frames in the film and television works and convert them into text descriptions. At the same time, the subtitle text information is extracted from the video frames, and the text content is mapped to the storyboard through the relevant language model, and the script summary and outline are further restored. Finally, the formed storyboard data set contains multiple storyboard samples, which are obtained based on the reverse deconstruction of at least one video.
[0107] S410. Use multiple storyboard samples in the storyboard data set to perform supervised fine-tuning on the basic large model to obtain the storyboard large model.
[0108] Specifically, the embodiments of the present invention can organize the storyboard samples in the storyboard data set, remove invalid marks such as special characters, garbled characters and incorrect pinyin, and mark the type of storyboard, lens movement, role and lens content, etc., to form supervision information. The preprocessed storyboard data set is divided into a training set, a validation set and a test set for model training, validation and testing. The self-developed basic large model is loaded, and it is continuously verified and adjusted through back propagation and parameter update to finally obtain a fine-tuned storyboard large model.
[0109] The storyboard script dataset composed of the storyboard script samples reversely structured out of the embodiment of the present invention performs supervised fine-tuning on the basic large model, which can effectively improve the trained storyboard large model's understanding and generation capabilities of the video narrative structure, thereby improving the generation efficiency and quality of the storyboard scripts, and helping to further improve the efficiency and quality of video content production.
[0110] Figure 5 The flowchart of the work flow of the storyboard model provided by the embodiment of the present invention is shown. The embodiment of the present invention uses multiple storyboard script samples in the storyboard script data set to supervise and fine-tune the basic model to develop a creative storyboard model. The model significantly improves the coherence of the video narrative, ensuring that the generated storyboard script has high-quality performance in terms of plot design, dramatic conflict, key plot points and emotional expression. Through the application of unified construction strategies for roles, scenes and props and dynamic transformation strategies, the shortcomings of AI-generated storyboard scripts in content expression are successfully solved, so that the generated storyboard scripts meet industrial-grade use standards. The constructed storyboard model can output the storyboard script according to the input video creative text including creativity, story or script, and then construct the storyboard screen props and the actions and states of the storyboard characters according to the storyboard script and unified scene construction and character construction, and finally output high-quality storyboard picture prompts, voice emotion prompts and storyboard video prompts. This method not only improves the creation efficiency of storyboard scripts, but also enhances the innovation ability in the film and television production process, and promotes the development of the film and television industry in the direction of intelligence and automation.
[0111] Optional, in the above Figure 1 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present invention, step S120 may specifically include:
[0112] The storyboard script, storyboard prompt words, reference character images and reference scene images are input into the image macro model, so that the image macro model is based on the reference character images and the reference scene images, controls the image generation process and the image editing process corresponding to the generation result of each storyboard screen description in the storyboard script according to the storyboard prompt words, and outputs a storyboard whose character appearance features are consistent with the reference character image and whose background features are consistent with the reference scene image.
[0113] Specifically, the embodiment of the present invention can input the storyboard script, storyboard prompt words, reference character images and reference scene images into the image model. The image model first controls the image generation process according to the storyboard prompt words, and uses the appearance features of the reference character image to ensure that the generated character is consistent with the reference character. At the same time, the background features of the reference scene image are used to ensure that the generated background is consistent with the reference scene. Then, the character clothing is segmented using the self-attention visual model, the reference clothing image is introduced, and the variational autoencoder (VAE) is combined to achieve the fusion and replacement of the clothing area to generate a preliminary image. Finally, the preliminary image is edited to optimize the details and achieve precise control, and finally the storyboard consistent with the reference character and the reference scene is output to ensure the high consistency of the character appearance and background features.
[0114] The embodiment of the present invention uses a storyboard generation process based on reference character images and reference scene images to enable the large image model to maintain a high degree of consistency between the character's appearance and the background features during the storyboard creation process, so that the final output storyboard not only meets the narrative requirements of the storyboard script, but also effectively conveys the story emotions and visual effects required by the user, thereby significantly improving the generation quality and consistency of the storyboard, which in turn helps to further improve the efficiency and quality of video content production.
[0115] Optionally, the storyboard prompt words include image generation prompt words and image editing prompt words.
[0116] Image generation prompts refer to the descriptive text information provided to the image model during the storyboard generation process, which is used to guide the image model to generate images according to specific themes, styles, compositions and scenes. Image generation prompts usually include the appearance characteristics, emotional state, environmental details and lighting effects of the characters, so as to ensure that the generated results are consistent with the user's intentions, thereby achieving accurate restoration of each storyboard in the storyboard script.
[0117] Image editing prompts refer to descriptive text information used in the post-processing stage of storyboard generation to guide the adjustment and optimization of the generated image. Image editing prompts can include specific requirements for color, detail, composition, and style, helping the model to modify and improve to ensure that the final image effect meets user expectations. The image editing process makes the final output storyboard more visually refined and consistent.
[0118] Optionally, the image generation process includes: extracting the character appearance features of the reference character image and the background features of the reference scene image, and then converting the character appearance features and the background features into first embedded features based on the image generation prompt words, and after processing the first embedded features using the first diffusion model, combining the self-attention visual model and the second diffusion model to fuse and replace the character clothing area based on the reference character image to obtain an image to be adjusted in which the character appearance features are consistent with the reference character image and the background features are consistent with the reference scene image.
[0119] Optionally, the image editing process includes: identifying the area to be edited of the image to be adjusted, extracting image area features of the area to be edited, and then converting the image area features into second embedded features based on image editing prompt words, and after processing the second embedded features using an image detail adjustment model, generating a storyboard in which the character appearance features are consistent with the reference character image and the background features are consistent with the reference scene image.
[0120] For ease of understanding, here we combine Figure 6 To explain: Figure 6The flowchart of the image generation and editing process provided by the embodiment of the present invention is shown. In the image generation process, SigLIP is used to encode the reference character image and the reference scene image, and the character appearance features of the reference character image and the background features of the reference scene image are extracted. At the same time, the appearance features of the reference character are used as conditional inputs through the ID encoder, and contrastive learning is used to improve the consistency of the character. The embedder module converts the character appearance features and background features into embedded features based on the image generation prompt words. The character clothing is segmented using the self-attention visual model, and the reference clothing image is introduced. Combined with the variational autoencoder, the fusion and replacement of the clothing area are realized, and the draft image to be adjusted is generated. After the draft is generated, the image editing stage is entered. At this time, the embodiment of the present invention uses the image editing prompt words and the information of the area to be edited to guide the adjustment of details. First, the area to be edited is encoded by SigLIP, and the multi-layer perceptron (Multi-Layer Perceptron, MLP) is used to realize the dimensional alignment of the reference image and the description text encoding. The MM-DiT fine-tuning model receives the base image, mask image, and conditional input. Based on the conditional input, the VAE encoding and denoising network are used to redraw the mask area of the base image. Finally, the edited image ensures that the characters, objects, and style are consistent with the reference image, generating high-quality storyboards. The overall process effectively combines image generation and editing technology to ensure that the generated images meet high standards in character consistency, background consistency, and detail expression, thereby meeting the generation requirements of storyboards.
[0121] Optional, in the above Figure 1 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present invention, step S130 may include:
[0122] The storyboard script, speech emotion prompt words and reference audio are input into the speech big model, so that the speech big model can extract the text features of the speech text to be converted in the storyboard script and the speech features of the reference audio, and the text features and the speech features are input into the self-attention speech model to generate the basic speech, and according to the speech emotion prompt words, the emotional expression of the basic speech is controlled through the third diffusion model to obtain the speech to be synthesized that is consistent with the speech features of the reference audio and contains emotional expression.
[0123] For ease of understanding, here we combine Figure 7 To explain: Figure 7The flowchart of the speech synthesis process provided by the embodiment of the present invention is shown. The embodiment of the present invention can input the storyboard script, speech emotion prompt words and reference audio into the speech model. The speech extractor provided by the speech model converts the input reference audio into a discrete token sequence, thereby extracting potential speech features. These extracted speech features will be input into the self-attention speech model together with the text tokens extracted from the speech text to be converted in the storyboard script by the text processor, and the context information is obtained by the Multi-HeadAttention mechanism to ensure the consistency of the semantics and emotional expression of the text. Then, the diffusion model controls the emotional expression of the generated synthesized speech based on the speech emotion prompt words and specific emotion tags, thereby realizing speech synthesis of multiple emotions without the need for a large amount of emotion annotation data, and finally generating a speech to be synthesized that is consistent with the reference audio in speech features and rich in emotion.
[0124] The embodiment of the present invention can extract and integrate text and voice features by inputting storyboard scripts, voice emotion prompts and reference audio into a large voice model, thereby generating synthetic voice that is consistent with the reference audio and rich in emotion, significantly improving the naturalness and emotional expression ability of voice synthesis, thereby helping to further improve the efficiency and quality of video content production.
[0125] Optional, in the above Figure 1 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present invention, step S140 may include:
[0126] The storyboard and storyboard video prompts are input into the video model, so that the video model can extract the bone sequence and three-dimensional mesh hand features of the preset action sequence in the storyboard according to the storyboard video prompts, encode the bone sequence and three-dimensional mesh hand features through the posture encoder and then input them into the denoising network, and dynamically adjust the joint spacing of the bone sequence by using the bone scaling strategy based on three-dimensional bone length estimation, so that the bone topology structure of the preset action sequence forms a spatial correspondence with the bone proportion of the character instance in the storyboard, and output the storyboard video that maintains motion consistency with the posture of the character in the storyboard.
[0127] Among them, the skeleton sequence of the preset action sequence refers to a sequence of action skeleton structures and action trajectories that are pre-defined according to a specific animation design, which is used to guide the movement performance of the character in the video.
[0128] Among them, the 3D mesh hand features refer to the hand shape and structural features represented by a 3D mesh model, usually including information such as finger joints and shapes, which are used to realize realistic hand movement animation.
[0129] Among them, 3D bone length refers to the physical length of each bone segment that makes up the character's skeleton, which is used to maintain the character's proportions and realism of movement in animation.
[0130] Among them, the bone scaling strategy refers to the strategy of dynamically adjusting the bone length during the character animation generation process to ensure that the character maintains appropriate proportions and structural consistency in different actions.
[0131] Among them, the bone topology structure refers to the connection relationship and arrangement of the character's bone system, which describes the relative position and connection method between each joint and bone.
[0132] Specifically, the video model provided by the embodiment of the present invention can extract the bone sequence (which can be in OpenPose format) and three-dimensional mesh hand features corresponding to the preset action sequence from the storyboard according to the storyboard video prompt words. The bone sequence is used to determine the basic action structure of the character, while the three-dimensional mesh hand features are used to capture the specific shape and movement of the character's hand. The extracted bone sequence and three-dimensional mesh hand features will be input into the posture encoder, which is responsible for converting these features into digital representations, making subsequent processing more efficient. The encoded feature data is sent to the denoising network to correct the noise that may exist in the input data, thereby ensuring that the generated animation performance is more natural and realistic. In order to deal with the problem of mismatch between the proportions of the character skeleton of the storyboard and the preset action sequence skeleton, the video model introduces a bone scaling strategy: according to the estimation of the three-dimensional bone length, the joint spacing of the bone sequence is dynamically adjusted so that the bone topology of the character can match the actual proportion of the character in the storyboard. Through this series of steps, the video model can effectively convert the static storyboard into a dynamic storyboard video, so that the subject in the image can act and speak realistically, improving the realism and expressiveness of the animation.
[0133] The embodiment of the present invention inputs the storyboard and storyboard video prompt words into the large video model, thereby accurately extracting the action features of the character and dynamically adjusting them, thereby generating an animated video that is consistent with the character posture in the storyboard and is natural and smooth, which helps to further improve the efficiency and quality of video content production.
[0134] Optional, in the above Figure 1 On the basis of one or more corresponding embodiments, in another optional embodiment provided by the embodiment of the present invention, step S140 may include:
[0135] The storyboards and storyboard video prompts are input into the large video model, so that the large video model combines the image contents of multiple subjects in the storyboards and the text prompts for multiple subjects in the storyboard video prompts through the dual cross-attention layer of the diffuse self-attention architecture during the generation of each frame of the storyboard video, thereby generating a storyboard video containing multiple subjects.
[0136] Specifically, the video big model provided by the embodiment of the present invention can extract features from the storyboard and the storyboard video prompt words: the storyboard will be converted into an image feature vector to capture important information in the scene. The storyboard video prompt words will be converted into a text feature vector to extract key information in the text. Then, the image-text attention matrix and the text-image attention matrix will be generated in the dual cross-attention layer of the diffuse self-attention architecture. The image-text attention matrix represents the degree of association between certain parts of the image and the text description. The text-image attention matrix represents the degree of influence of each word in the text on the image features. Through these two attention matrices, the video big model can dynamically adjust the focus on image features and text features when generating each frame, so that when generating each frame, the video big model can generate dynamic pictures based on the results of cross-attention and comprehensively consider the information of the image and text.
[0137] The embodiments of the present invention integrate image embedding and subject-level text prompts into the video generation process through a dual cross-attention layer, thereby supporting multi-subject personalized generation. It not only performs well in single / multi-subject video personalized generation, but also significantly improves subject fidelity and background diversity, which helps to further improve the efficiency and quality of video content production.
[0138] Optional, based on Figure 1 The method shown, such as Figure 8 As shown, a flowchart of a third specific implementation of the video generation control method provided by an embodiment of the present invention is shown, and step S150 may include:
[0139] S800. Extracting prosodic features of the speech to be synthesized through a speech encoder, wherein the speech encoder is constructed using a multilingual speech recognition model based on a self-attention mechanism.
[0140] The speech encoder is a model that can process and recognize speech signals in multiple languages, and is intended to convert the speech to be synthesized into a feature representation that is easy to process. Optionally, the speech encoder can be Whisper.
[0141] Among them, prosodic features refer to the characteristics of rhythm, pitch, stress and pause in the speech to be synthesized. Prosodic features play a role in speech fluency, expression of emotions and semantic transmission in speech communication.
[0142] S810, extracting mouth shape features of the character's face from the storyboard video.
[0143] Among them, the character's mouth shape features refer to the shape and movement characteristics of the character's mouth in the storyboard video. The character's mouth shape features are usually closely related to the character's pronunciation, emotional expression, and interactive dynamics.
[0144] Specifically, the embodiment of the present invention can perform face and mouth detection on each video frame in the storyboard video, crop the character's mouth area and perform image preprocessing, and then use a convolutional neural network to extract the character's mouth shape features in the character's mouth area.
[0145] S820. Establish a cross-modal cross-attention layer in the denoising network, calculate the attention weight matrix of the rhythmic features and the character's mouth shape features, and dynamically adjust the intensity parameters of the speech-driven mouth shape.
[0146] Specifically, the embodiment of the present invention can construct a cross-modal cross-attention layer based on the denoising network in advance to process the relationship between the rhythmic features and the character's mouth shape features. The denoised rhythmic features and the character's mouth shape features are used as input, and the attention weights between the rhythmic features and the character's mouth shape features are calculated through the attention mechanism. According to the calculated attention weights, the rhythmic features and the mouth shape features are weighted to provide a more accurate mouth shape animation drive. Based on the attention weight matrix, the intensity parameter of the voice-driven mouth shape is dynamically adjusted. The intensity parameter determines the amplitude and speed of the mouth shape movement to match the changes in the rhythmic features.
[0147] S830: Use a pre-trained facial mask generator to generate a spatial constraint template, and perform dynamic masking on the non-mouth area in the feature space of the storyboard video.
[0148] Specifically, an embodiment of the present invention can select a suitable pre-trained facial mask generator, which is a model based on deep learning and can accurately generate masks of facial features. The storyboard video is input into the pre-trained facial mask generator, and the facial mask generator generates a spatial constraint template to identify the mouth area and the non-mouth area. In the feature space of the storyboard video, the generated spatial constraint template is used to mask the non-mouth area. The purpose of mask processing is to emphasize the dynamic changes of the mouth in the feature extraction and analysis of the video, while ignoring or reducing the attention to other facial areas to avoid interference. According to the changes in the pronunciation, emotional expression and other performances of the character in the video, the application of the mask is dynamically adjusted, that is, in different video frames, the strength and application range of the mask can be adjusted as needed to adapt to the changes in mouth movement.
[0149] S840, fusing the character's mouth shape features adjusted according to the intensity parameter with the mouth area in the feature space of the storyboard video to generate a target video in which the character's mouth shape and voice are synchronized.
[0150] Specifically, the embodiment of the present invention can perform weighted adjustment on the character's mouth shape features according to the value of the intensity parameter, and fuse the adjusted character's mouth shape features to the mouth area in the feature space of the storyboard video through interpolation or other fusion techniques to generate a target video with synchronized audio and video.
[0151] In order to control the character's mouth shape according to audio while ensuring the coordination of movements, the embodiment of the present invention combines a multilingual speech encoder based on a self-attention mechanism with a cross-attention layer and integrates it into a denoising network. A spatial constraint template is generated through a facial mask, which significantly enhances the synchronization effect of the character's mouth shape and voice, thereby helping to further improve the efficiency and quality of video content production.
[0152] Although operations are depicted in a particular order, this should not be understood as requiring that the operations be performed in the particular order shown or in a sequential order.Multitasking and parallel processing may be advantageous under certain circumstances.
[0153] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0154] Corresponding to the above method embodiment, the embodiment of the present invention also provides a video generation control device, whose structure is as follows: Fig. 9 As shown, it may include: a video creative text obtaining unit 10, a storyboard large model application unit 20, an image large model application unit 30, a voice large model application unit 40, a video large model application unit 50 and a target video generating unit 60.
[0155] The video creative text obtaining unit 10 is used to obtain the video creative text.
[0156] The storyboard big model application unit 20 is used to input the video creative text into the storyboard big model, and obtain the storyboard script and related prompt words output by the storyboard big model, wherein the related prompt words include storyboard picture prompt words, voice emotion prompt words and storyboard video prompt words.
[0157] The image large model application unit 30 is used to generate a storyboard having image features consistent with the reference image through the image large model based on the storyboard script, storyboard prompt words and the reference image.
[0158] The speech big model application unit 40 is used to convert the speech text to be converted in the storyboard script into the speech to be synthesized which is consistent with the speech features of the reference audio and contains emotional expression through the speech big model based on the storyboard script, speech emotion prompt words and reference audio.
[0159] The video big model application unit 50 is used to generate the storyboard into a storyboard video through the video big model based on the storyboard and the storyboard video prompt words.
[0160] The target video generation unit 60 is used to dynamically associate the speech to be synthesized with the mouth shape features of the character in the storyboard video to generate a target video.
[0161] Optionally, the video generation control device may further include: a sound effect synthesis unit.
[0162] The sound effect synthesis unit is used to obtain the sound effect that matches the sound effect text to be generated in the storyboard script; identify the key dynamic time points in the target video; use the self-attention algorithm to dynamically align the sound effect and the key dynamic time points to obtain the target video after synthesizing the sound effect.
[0163] Optionally, the video generation control device may also include: a storyboard large model training unit.
[0164] The storyboard big model training unit is used to obtain a storyboard script dataset, wherein the storyboard script dataset includes multiple storyboard script samples obtained by reverse deconstructing at least one video; the basic big model is supervised and fine-tuned using the multiple storyboard script samples in the storyboard script dataset to obtain the storyboard big model.
[0165] Optionally, the reference image includes a reference character image and a reference scene image.
[0166] Optionally, the image large model application unit 30 is specifically used to input the storyboard script, storyboard prompt words, reference character images and reference scene images into the image large model, so that the image large model is based on the reference character images and reference scene images, controls the image generation process and image editing process corresponding to the generation result of each storyboard screen description in the storyboard script according to the storyboard prompt words, and outputs a storyboard in which the character appearance features are consistent with the reference character image and the background features are consistent with the reference scene image.
[0167] Optionally, the storyboard prompt words include image generation prompt words and image editing prompt words.
[0168] Optionally, the image generation process includes: extracting the appearance features of the character from the reference character image and the background features of the reference scene image, and then converting the appearance features of the character and the background features into first embedded features according to the image generation prompt words, and after processing the first embedded features using the first diffusion model, combining the self-attention visual model and the second diffusion model to fuse and replace the character clothing area based on the reference character image, so as to obtain an image to be adjusted whose appearance features are consistent with the reference character image and whose background features are consistent with the reference scene image,
[0169] Optionally, the image editing process includes: identifying the area to be edited of the image to be adjusted, extracting image area features of the area to be edited, and then converting the image area features into second embedded features based on image editing prompt words, and after processing the second embedded features using an image detail adjustment model, generating a storyboard in which the character appearance features are consistent with the reference character image and the background features are consistent with the reference scene image.
[0170] Optionally, the speech big model application unit 40 is specifically used to input the storyboard script, speech emotion prompt words and reference audio into the speech big model, so that the speech big model extracts the text features of the speech text to be converted in the storyboard script and the speech features of the reference audio, inputs the text features and the speech features into the self-attention speech model to generate basic speech, and controls the emotional expression of the basic speech through the third diffusion model according to the speech emotion prompt words, so as to obtain the speech to be synthesized that is consistent with the speech features of the reference audio and contains emotional expression.
[0171] Optionally, the video large model application unit 50 may include: a character skeleton adjustment subunit.
[0172] The character skeleton adjustment subunit is used to input the storyboard and storyboard video prompts into the video model, so that the video model can extract the skeleton sequence and three-dimensional mesh hand features of the preset action sequence in the storyboard according to the storyboard video prompts, encode the skeleton sequence and three-dimensional mesh hand features through the posture encoder and then input them into the denoising network, and dynamically adjust the joint spacing of the skeleton sequence using a skeleton scaling strategy based on three-dimensional bone length estimation, so that the skeleton topology of the preset action sequence forms a spatial correspondence with the skeleton proportions of the character instance in the storyboard, and output a storyboard video that maintains motion consistency with the posture of the character in the storyboard.
[0173] Optionally, the video large model application unit 50 may include: a multi-subject storyboard processing subunit.
[0174] The multi-subject storyboard processing subunit is used to input the storyboard pictures and storyboard video prompt words into the video big model, so that the video big model can combine the image content of multiple subjects in the storyboard pictures and the text prompts for multiple subjects in the storyboard video prompt words through the dual cross-attention layer of the diffuse self-attention architecture during the generation process of each frame of the storyboard video, thereby generating a storyboard video containing multiple subjects.
[0175] Optionally, the target video generation unit 60 is specifically used to extract the prosodic features of the speech to be synthesized through a speech encoder, wherein the speech encoder is constructed using a multilingual speech recognition model based on a self-attention mechanism; extract the character's mouth shape features from the storyboard video; establish a cross-modal cross-attention layer in the denoising network, calculate the attention weight matrix of the prosodic features and the character's mouth shape features, and dynamically adjust the intensity parameters of the speech-driven mouth shape; use a pre-trained facial mask generator to generate a spatial constraint template, and dynamically mask the non-mouth area in the feature space of the storyboard video; fuse the character's mouth shape features adjusted according to the intensity parameters to the mouth area in the feature space of the storyboard video to generate a target video in which the character's mouth shape is synchronized with the speech.
[0176] The invention provides a video generation control device, which is used for: obtaining a video creative text; inputting the video creative text into a storyboard big model to obtain a storyboard script and associated prompt words output by the storyboard big model, wherein the associated prompt words include storyboard picture prompt words, voice emotion prompt words and storyboard video prompt words; based on the storyboard script, the storyboard picture prompt words and a reference image, generating a storyboard picture that is consistent with the image features of the reference image through an image big model; based on the storyboard script, the voice emotion prompt words and a reference audio, converting the voice text to be converted in the storyboard script into a voice to be synthesized that is consistent with the voice features of the reference audio and contains emotional expression through a voice big model; based on the storyboard picture and the storyboard video prompt words, generating the storyboard picture into a storyboard video through a video big model; dynamically associating the voice to be synthesized with the mouth shape features of a character in the storyboard video to generate a target video. The present invention solves the structural defects of storyboard scripts through a large storyboard model, ensures the consistency of storyboard features based on image generation constrained by reference image features, realizes audio and video synchronization by combining emotional speech synthesis with dynamic lip alignment, and improves video fidelity by guiding the multi-agent generation algorithm through storyboard video prompt words, thereby effectively solving the technical problems existing in the existing AIGC technology in film and television creation, such as lack of emotional logic, poor consistency of storyboard features, difficulty in audio synchronization and insufficient controllability of video generation, and improving the efficiency and quality of video content production.
[0177] Regarding the device in the above embodiment, the specific manner in which each unit performs the operation has been described in detail in the embodiment of the method, and will not be elaborated here.
[0178] The video generation control device includes a processor and a memory. The above-mentioned video creative text acquisition unit 10, storyboard large model application unit 20, image large model application unit 30, voice large model application unit 40, video large model application unit 50 and target video generation unit 60 are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to realize corresponding functions.
[0179] The processor includes a kernel, which calls the corresponding program unit from the memory. One or more kernels can be set, and the kernel parameters can be adjusted to solve the structural defects of the storyboard script through the storyboard large model, ensure the consistency of the storyboard features based on the image generation of the reference image feature constraints, combine emotional speech synthesis with dynamic mouth alignment to achieve audio and video synchronization, and guide the multi-subject generation algorithm through the storyboard video prompt words to improve the video fidelity, thereby effectively solving the technical problems of the existing AIGC technology in film and television creation, such as the lack of emotional logic, poor consistency of storyboard features, difficulty in audio synchronization, and insufficient controllability of video generation, and improving the efficiency and quality of video content production.
[0180] An embodiment of the present invention provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the video generation control method is implemented.
[0181] An embodiment of the present invention provides a processor, which is used to run a program, wherein the video generation control method is executed when the program is run.
[0182] like Fig.10 As shown, an embodiment of the present invention provides an electronic device 1000, which includes at least one processor 1001, at least one memory 1002 connected to the processor 1001, and a bus 1003; wherein the processor 1001 and the memory 1002 communicate with each other through the bus 1003; the processor 1001 is used to call the program instructions in the memory 1002 to execute the above-mentioned video generation control method. The electronic device in this article can be a server, a PC, a PAD, a mobile phone, etc.
[0183] The present invention also provides a computer program product, which, when executed on an electronic device, is suitable for executing a program that initializes the steps of the video generation control method.
[0184] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatuses, electronic devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable device to generate a machine, so that the instructions executed by the processor of the computer or other programmable device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0185] In a typical configuration, an electronic device includes one or more processors (CPU), a memory, and a bus. The electronic device may also include an input / output interface, a network interface, and the like.
[0186] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip. The memory is an example of a computer-readable medium.
[0187] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0188] In the description of the present invention, it should be understood that the terms "up", "down", "front", "back", "left" and "right" etc. indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description. They do not indicate or imply that the positions or elements referred to must have specific directions, be constructed and operate in specific directions. Therefore, they should not be understood as limitations of the present invention.
[0189] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. It should also be noted that the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0190] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0191] The above are only embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention should be included in the scope of the present invention.
Claims
1. A video generation control method, characterized in that: include: Get creative text for your video; Input the video creative text into the storyboard big model to obtain the storyboard script and related prompt words output by the storyboard big model, wherein the related prompt words include storyboard picture prompt words, voice emotion prompt words and storyboard video prompt words; Based on the storyboard script, the storyboard prompt words and the reference image, generating a storyboard consistent with the image features of the reference image through an image macro model; Based on the storyboard, the speech emotion prompt words and the reference audio, the speech text to be converted in the storyboard is converted into a speech to be synthesized that is consistent with the speech features of the reference audio and contains emotional expression through a speech big model; Based on the storyboard and the storyboard video prompt words, the storyboard is generated into a storyboard video through a video macro model; The speech to be synthesized is dynamically associated with the mouth shape features of the character in the storyboard video to generate a target video.
2. The method according to claim 1, characterized in that After dynamically associating the speech to be synthesized with the mouth shape features of the character in the storyboard video to generate the target video, the method further includes: Obtaining a sound effect that matches the sound effect text to be generated in the storyboard; Identifying key dynamic time points in the target video; The sound effect and the key dynamic time point are dynamically aligned using a self-attention algorithm to obtain the target video after synthesizing the sound effect.
3. The method according to claim 1, characterized in that The training process of the storyboard model includes: Obtaining a storyboard script dataset, wherein the storyboard script dataset includes a plurality of storyboard script samples obtained by reverse deconstructing at least one video; The basic large model is supervised and fine-tuned using a plurality of the storyboard samples in the storyboard data set to obtain a storyboard large model.
4. The method according to any one of claims 1 to 3, characterized in that The reference image includes a reference character image and a reference scene image, and the step of generating a storyboard consistent with the image features of the reference image through an image macro model based on the storyboard script, the storyboard prompt words and the reference image includes: The storyboard script, the storyboard prompt words, the reference character image and the reference scene image are input into a large image model, so that the large image model controls the image generation process and the image editing process corresponding to the generation result of each storyboard screen description in the storyboard script based on the reference character image and the reference scene image according to the storyboard prompt words, and outputs a storyboard whose character appearance features are consistent with the reference character image and whose background features are consistent with the reference scene image.
5. The method according to claim 4, characterized in that The storyboard prompt words include image generation prompt words and image editing prompt words, and the image generation process includes: extracting the character appearance features of the reference character image and the background features of the reference scene image, and then converting the character appearance features and the background features into first embedded features according to the image generation prompt words, and after processing the first embedded features using a first diffusion model, combining a self-attention visual model and a second diffusion model to fuse and replace the character clothing area based on the reference character image, to obtain an image to be adjusted in which the character appearance features are consistent with the reference character image and the background features are consistent with the reference scene image, And / or, the image editing process includes: identifying the area to be edited of the image to be adjusted, extracting the image area features of the area to be edited, and then converting the image area features into a second embedded feature based on the image editing prompt word, and after processing the second embedded feature using an image detail adjustment model, generating a storyboard in which the character appearance features are consistent with the reference character image and the background features are consistent with the reference scene image.
6. The method according to any one of claims 1 to 3, characterized in that The method of converting the speech text to be converted in the storyboard script into speech to be synthesized that is consistent with the speech features of the reference audio and contains emotional expression based on the storyboard script, the speech emotion prompt words and the reference audio through a speech big model includes: The storyboard script, the speech emotion prompt words and the reference audio are input into a speech big model, so that the speech big model extracts the text features of the speech text to be converted in the storyboard script and the speech features of the reference audio, the text features and the speech features are input into a self-attention speech model to generate a basic speech, and according to the speech emotion prompt words, the emotional expression of the basic speech is controlled through a third diffusion model to obtain a speech to be synthesized that is consistent with the speech features of the reference audio and contains emotional expression.
7. The method according to any one of claims 1 to 3, characterized in that The step of generating the storyboard into a storyboard video based on the storyboard and the storyboard video prompt words through a video macro model includes: The storyboard and the storyboard video prompt words are input into a large video model, so that the large video model extracts the bone sequence and three-dimensional mesh hand features of the preset action sequence in the storyboard according to the storyboard video prompt words, the bone sequence and the three-dimensional mesh hand features are encoded by a posture encoder and then input into a denoising network, and a bone scaling strategy based on three-dimensional bone length estimation is used to dynamically adjust the joint spacing of the bone sequence, so that the bone topological structure of the preset action sequence forms a spatial correspondence with the bone proportion of the character instance in the storyboard, and a storyboard video that maintains motion consistency with the posture of the character in the storyboard is output.
8. The method according to any one of claims 1 to 3, characterized in that The step of generating the storyboard into a storyboard video based on the storyboard and the storyboard video prompt words through a video macro model includes: The storyboard and the storyboard video prompt words are input into the video macro model, so that the video macro model combines the image contents of multiple subjects in the storyboard and the text prompts for multiple subjects in the storyboard video prompt words through the dual cross-attention layer of the diffuse self-attention architecture during the generation of each frame of the storyboard video, thereby generating a storyboard video containing multiple subjects.
9. The method according to any one of claims 1 to 3, characterized in that The step of dynamically associating the speech to be synthesized with the mouth shape features of the character in the storyboard video to generate a target video includes: Extracting the prosodic features of the speech to be synthesized by a speech encoder, wherein the speech encoder is constructed using a multilingual speech recognition model based on a self-attention mechanism; Extracting mouth shape features of the character's face from the storyboard video; Establishing a cross-modal cross-attention layer in the denoising network, calculating the attention weight matrix of the prosodic features and the character's mouth shape features, and dynamically adjusting the intensity parameters of the speech-driven mouth shape; Applying a pre-trained facial mask generator to generate a spatially constrained template, and dynamically masking a non-mouth area in a feature space of the storyboard video; The character's mouth shape features adjusted according to the intensity parameter are fused to the mouth area in the feature space of the storyboard video to generate a target video in which the character's mouth shape and voice are synchronized.
10. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the video generation control method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Animation generation method and system, medium and electronic terminal
CN113744369A
Method and system for automatically generating crime scene video based on AI voice
CN117392289A
Dynamic video generation method and device, electronic equipment and storage medium
CN118612367A
Method and system for automatically generating novel promotion video based on multi-modal large model
CN119094672A
Image generation method, electronic equipment and storage medium
CN119273810A
Cited By
Text-to-video full link generation method and system based on multi-modal large model
CN120512591A
Text-to-video full-link generation method and system based on multimodal large model
CN120512591B
Voice determination method, electronic equipment and storage medium
CN120833776A
SSML text automatic generation method and device fusing sub-mirror hierarchy information
CN121072489A
Music video generation method and device, equipment and storage medium
CN121418302A