A video generation method, apparatus, device and medium

By automatically generating storyboard scripts and audio, and combining video filtering and duration scaling based on storyboard dimensions, the problem of low video generation efficiency in existing technologies has been solved, realizing an efficient and easy-to-use video generation method that improves video quality and audio matching accuracy.

CN119767101BActive Publication Date: 2026-04-03BEIJING DONGCHEZU TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing technologies struggle to simultaneously meet the multi-dimensional requirements of ease of use, efficiency, and video quality during video generation. Especially in the absence of professional shooting and editing capabilities, video production is cumbersome and inefficient, and one-click video generation algorithms cannot adapt to various application scenarios, such as vehicle scenarios, resulting in rough video content.

Method used

By acquiring basic video data, including scene identifiers, target scripts, and multiple first videos, the system automatically generates storyboard text and audio, filters and scales the videos based on storyboard dimensions, and generates target videos by combining storyboard subtitles, thus achieving an automated video generation process.

Benefits of technology

A complete and efficient automated video generation process has been implemented, which improves video generation efficiency, saves manual time, and the generated video quality and audio matching are more in line with actual application scenarios, with natural and smooth video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119767101B_ABST
    Figure CN119767101B_ABST
Patent Text Reader

Abstract

This disclosure relates to a video generation method, apparatus, device, and medium. The method includes: acquiring basic video data, including scene identifiers, a target script, and multiple first videos; generating storyboard text and storyboard audio based on scene information corresponding to the target script and scene identifiers; performing video filtering and duration scaling on the multiple first videos according to the storyboard dimensions based on the storyboard audio to obtain multiple second videos; generating storyboard subtitles based on the storyboard audio and storyboard text; and generating a target video based on the multiple second videos, multiple storyboard audios, and multiple storyboard subtitles. This disclosure implements a complete and efficient automated video generation process with comprehensive functions, high usability, and saves a significant amount of manual time and effort, greatly improving video generation efficiency. By filtering and scaling the storyboard videos, it effectively improves the matching degree between video and audio, thereby improving the video quality of the generated video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of video processing technology, and in particular to a video generation method, apparatus, device, and medium. Background Technology

[0002] Video is a crucial means of showcasing an item, and video quality plays a key role in enhancing the presentation. Currently, producing high-quality videos requires video shooting and editing skills, as well as significant human time and effort, resulting in low efficiency. To address these issues, deep learning technology can be used to optimize certain steps in video generation. However, video generation still struggles to simultaneously meet the requirements of ease of use, efficiency, and video quality, and further improvements are needed. Summary of the Invention

[0003] To address the aforementioned technical problems, this disclosure provides a video generation method, apparatus, device, and medium.

[0004] This disclosure provides a video generation method, the method comprising:

[0005] Acquire basic video data, which includes scene identifiers, target scripts, and multiple first videos;

[0006] Based on the target script and the scene information corresponding to the scene identifier, generate storyboard text and storyboard audio;

[0007] Based on the storyboard audio, the multiple first videos are filtered and their durations are scaled according to the storyboard dimensions to obtain multiple second videos;

[0008] Generate subtitles for the storyboard based on the storyboard audio and the storyboard text;

[0009] The target video is generated based on the plurality of second videos, the plurality of storyboard audios, and the plurality of storyboard subtitles.

[0010] This disclosure also provides a video generation apparatus, the apparatus comprising:

[0011] The acquisition module is used to acquire basic video data, wherein the basic video data includes scene identifiers, target scripts, and multiple first videos;

[0012] The storyboard module is used to generate storyboard scripts and storyboard audio based on the target script and the scene information corresponding to the scene identifier;

[0013] The video editing module is used to perform video filtering and duration scaling on the multiple first videos according to the segmentation dimensions based on the segmentation audio, to obtain multiple second videos;

[0014] The subtitle module is used to generate subtitles for the storyboard based on the storyboard audio and the storyboard text;

[0015] The video generation module is used to generate a target video based on the plurality of second videos, the plurality of storyboard audios, and the plurality of storyboard subtitles.

[0016] This disclosure also provides an electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the video generation method provided in this disclosure.

[0017] This disclosure also provides a computer-readable storage medium storing a computer program for performing the video generation method provided in this disclosure.

[0018] Compared with the prior art, the technical solution provided in this disclosure has the following advantages: The video generation scheme provided in this disclosure obtains basic video data, including scene identifiers, target scripts, and multiple first videos; based on the scene information corresponding to the target scripts and scene identifiers, it generates storyboard text and storyboard audio; based on the storyboard audio, it performs video filtering and duration scaling processing on the multiple first videos according to the storyboard dimension to obtain multiple second videos; it generates storyboard subtitles based on the storyboard audio and storyboard text; and it generates a target video based on the multiple second videos, multiple storyboard audios, and multiple storyboard subtitles. By adopting the above technical solution, storyboard text and audio can be automatically generated based on scene information corresponding to scene identifiers in the basic video data and the target script, thereby generating storyboard subtitles. The video is then filtered and its duration is scaled using the storyboard audio. Finally, the final target video is generated based on the processed video, storyboard audio, and storyboard subtitles. This achieves a complete and efficient automated video generation process with comprehensive functions, high ease of use, and significant savings in manual time and effort, greatly improving video generation efficiency. Furthermore, the text and audio generated based on scene information are more in line with actual application scenarios. The filtering and duration scaling of the storyboard video effectively improves the matching degree between video and audio, thereby improving the video quality of the generated video. Attached Figure Description

[0019] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0020] Figure 1 A flowchart illustrating a video generation method provided in this embodiment of the present disclosure;

[0021] Figure 2 A schematic diagram illustrating a text generation process provided in an embodiment of this disclosure;

[0022] Figure 3 A schematic diagram illustrating an audio generation process provided in an embodiment of this disclosure;

[0023] Figure 4 A flowchart illustrating another video generation method provided in this embodiment of the disclosure;

[0024] Figure 5 A schematic diagram illustrating a video generation process provided in an embodiment of this disclosure;

[0025] Figure 6 This is a schematic diagram of the structure of a video generation device provided in an embodiment of the present disclosure;

[0026] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0027] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0028] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0029] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0031] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0032] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0033] While deep learning tools can be used to accelerate video generation and improve video quality, they suffer from several drawbacks: Without professional video shooting and editing skills, high-quality content cannot be produced; the production process can be cumbersome, requiring significant time and effort, leading to low production efficiency; one-click video generation algorithms are not adaptable to various application scenarios, such as vehicle scenarios; and the rendering quality of these algorithms is often rough, exhibiting poor matching between visuals and audio, stiff text, and unnatural voice-overs. Video generation still struggles to simultaneously meet the demands of ease of use, efficiency, and video quality, requiring further improvement.

[0034] To address the aforementioned problems, this disclosure provides a video generation method, which will be described below with reference to specific embodiments.

[0035] Figure 1 This is a flowchart illustrating a video generation method provided in an embodiment of the present disclosure. The method can be executed by a video generation device, which can be implemented using software and / or hardware, and is generally integrated into an electronic device. Figure 1 As shown, the method includes:

[0036] Step 101: Obtain basic video data, which includes scene identifiers, target scripts, and multiple first videos.

[0037] The basic video data can be the data that a user inputs when generating or creating a video. The specific data included in the basic video data varies depending on the usage scenario. In this embodiment of the disclosure, taking a vehicle scenario as an example, the basic video data can include a scene identifier, a target script, and a first video. The basic video data can also include dubbing parameters, transition parameters, and background music, etc., which are set according to the actual situation.

[0038] Scene identifiers can be information that indicates which specific use case the video is currently being generated for. For example, for a vehicle scene, the scene identifier could be a vehicle identifier, which indicates which vehicle is being generated for the video. The vehicle identifier could include the vehicle model, manufacturer's name, or trademark. Target scripts are texts that provide detailed guidance and structured requirements for the content and style of the target video to be generated. They can plan the video's presentation in detail, including but not limited to video style, timeline, dialogue, narration, scene settings, motion direction, shot transitions, visual effects, background music, and sound effects. Target scripts can include multiple storyboards. A storyboard is the smallest unit in a video for shooting characters, actions, or events, providing detailed textual descriptions and frame compositions for the content, angle, and movement of a shot.

[0039] The first video can be the original video that needs to be edited or processed, or it can be uploaded by the user. There can be multiple first videos.

[0040] Specifically, when a user needs to create a video, they can first input basic video data. The video generation device can then obtain this basic data. For example, after obtaining the scene identifier selected by the user, the video generation device can first display multiple preset scripts. In response to the user's selection, it can determine the target script from among the multiple preset scripts. Then, it can display the target script, including video upload areas corresponding to multiple storyboard scripts. One storyboard script can correspond to one video upload area, and each video upload area displays the video requirements text that the user needs to upload for that storyboard script. The user can upload the first video corresponding to the accurate storyboard script based on the video requirements text. One storyboard script corresponds to multiple first videos, and finally, multiple first videos corresponding to the target script are obtained. By constraining the user to upload the original video of the corresponding type of storyboard script at the storyboard dimension, for example, in a vehicle scene, if the storyboard script explains the interior of the vehicle and requires the user to upload a video related to the vehicle interior, the content of the subsequently generated video footage matches the text corresponding to the script.

[0041] Optionally, after acquiring multiple first videos, the video generation device can also preprocess them. Preprocessing may include at least one of video stabilization, color space processing, and audio processing. Video stabilization can be implemented using a video stabilization algorithm for shaky footage in the first videos; color space processing can be performed on high dynamic range (HDR) format videos in the first videos, converting them to the standard dynamic range (SDR) color gamut; audio processing may involve segmenting the video from the beginning of the audio stream for videos where original sound is selected for preservation, trimming out preceding audio segments, and removing unnecessary "breathing" sections. Preprocessing the videos can improve video quality, which helps improve the quality of subsequent video generation.

[0042] Step 102: Based on the target script and the scene information corresponding to the scene identifier, generate storyboard text and storyboard audio.

[0043] Scene information can be introductory information, attribute information, etc., related to the specific scene corresponding to the scene identifier. For example, for a vehicle scene, the scene information can be vehicle information, which can be all the information related to a vehicle, such as vehicle model, color, usage, displacement, power, chassis number, inspection information, VIN, vehicle model, etc. The storyboard text can be the text content after filling in the scene information for a storyboard. The storyboard text can describe the content, actions, cinematography, character performance, dialogue, and sound effects of a shot. The storyboard audio can be the dubbing audio generated based on the storyboard text. This embodiment of the disclosure takes the automatic generation of the storyboard audio as an example. In actual scenarios, the first video may include user narration audio as storyboard audio. The storyboard, storyboard text, and storyboard audio are divided according to the storyboard dimension and correspond one-to-one, that is, one storyboard text and one storyboard audio correspond to one storyboard text and one storyboard audio.

[0044] After obtaining the target script and scene identifier, the video generation device can retrieve scene information through an interface based on the scene identifier and obtain the target text template corresponding to the target script. The scene information is then filled into the target text template to obtain multiple storyboard texts; alternatively, multiple storyboard texts can be generated using a model. Subsequently, based on the multiple storyboard texts, the model generates corresponding audio for multiple storyboard scenes.

[0045] In some embodiments, generating storyboard text and storyboard audio based on the target script and scene information corresponding to the scene identifier may include: obtaining a prompt word template corresponding to the target script; filling the prompt word template with the scene information corresponding to the scene identifier to obtain text prompt words; inputting the text prompt words into a text generation model and outputting a target text, wherein the target text includes multiple storyboard texts; and generating a target audio with a preset timbre based on the target text using a timbre replication algorithm, wherein the target audio includes multiple storyboard audios.

[0046] The prompt template can be a template that configures the content and position of the prompt words, used to generate prompt words. Prompt words can be text or sentence fragments used to trigger and guide the model to perform a specific task and generate specific output content. Prompt words can describe the task the model should perform to guide the model to generate specific text, images, audio, etc. The prompt word template in this embodiment is used to generate text prompt words. The prompt word template includes multiple placeholders for filling different information, such as scene information. The text prompt words can be prompt words used to guide the text generation model to generate target text. The text generation model can be a model used to generate text for a video; its specific type is not limited, for example, the text generation model can be a deep learning model, a pre-trained model based on large-scale data, etc. The target text can be the text corresponding to the target script, and can include multiple storyboard texts. The target audio can be the audio corresponding to the target video to be generated, and can include multiple storyboard audios.

[0047] Specifically, when generating storyboard scripts and storyboard videos, the video generation device can first obtain the prompt word template corresponding to the target script. This prompt word template is preset, and different scripts have corresponding prompt word templates. Then, scene information can be filled into the corresponding positions in the prompt word template to obtain the script prompt words. The script prompt words are input into the script generation model for reasoning, and the output is the target script including multiple storyboard scripts. Then, a timbre replication algorithm is used to replicate the preset timbre of the speaker based on the preset audio uploaded by the user or to determine the preset timbre based on the user's selection from multiple local timbres. The target audio that matches the preset timbre is generated according to the target script. The preset audio can be audio uploaded by the user, including their own or other people's voices. The timbre replication algorithm can be a technology that replicates the timbre, speech rate, and short sentence characteristics of a specific speaker's voice.

[0048] For example, the copy prompt may include format controls for the target copy, as shown below: "Please output the results according to the above rules and the following format. The returned format is as follows, where "{xxx}" represents a placeholder: Scene 1\{copy}; Scene 2\{copy}...", which is just an example.

[0049] For example, Figure 2This is a schematic diagram of a text generation process provided in an embodiment of the present disclosure, such as... Figure 2 As shown in the figure, the text generation process may include: determining whether the target script includes the corresponding target text template; if so, capturing scene information and filling it into the target text template to generate target text including multiple storyboard texts; otherwise, capturing scene information and filling it into the prompt word template, and calling the text production model to generate target text including multiple storyboard texts.

[0050] For example, Figure 3 This is a schematic diagram of an audio generation process provided in an embodiment of the present disclosure, such as... Figure 3 As shown in the diagram, the audio production process may include: determining whether to replicate the timbre, which can be determined based on the user's selection; if so, providing a preset audio to replicate the preset timbre; otherwise, providing multiple local timbres and determining the preset timbre based on the user's selection; and then producing the target audio according to the preset timbre based on the target text.

[0051] In the above solution, the text and audio are automatically generated based on scene information, which is more in line with actual application scenarios and effectively solves the problem of stiff text. The cloned real human voice is used as the dubbing for the video to enhance the sense of realism.

[0052] Step 103: Based on the storyboard audio, perform video filtering and duration scaling on multiple first videos according to the storyboard dimension to obtain multiple second videos.

[0053] The second video can be a video obtained by filtering and length scaling the first video, and there can be multiple second videos.

[0054] The video generation device first divides multiple first videos according to the segmentation dimension, determines multiple first videos corresponding to each segment audio, and then selects a certain number of first videos with longer durations from the multiple first videos included in the segment based on the audio duration of the segment audio and performs duration scaling processing so that the total duration of the selected first videos is equal to the audio duration of the segment audio. The number of first videos selected during the above video selection can be determined according to the actual situation.

[0055] For example, Figure 4 A flowchart illustrating another video generation method provided in this disclosure embodiment is shown below. Figure 4 As shown, in one feasible implementation, for each scene audio, step 103 above may include the following steps:

[0056] Step 401: Based on the audio duration and duration threshold of the storyboard audio, determine the number of videos to be extracted and the target duration.

[0057] The audio duration can be the specific duration corresponding to the storyboard audio. The duration threshold can be the minimum duration of the second video determined according to business requirements; that is, the duration of the second video must be greater than or equal to the duration threshold to avoid videos that are too short to be usable. For example, the duration threshold can be set to 2 seconds. The number of extracted videos can be the specific number of first videos to be filtered for a storyboard that includes multiple first videos. The number of prompt videos corresponding to different storyboards can be different. The target duration can be the corresponding duration to which each of the multiple first videos included in a storyboard needs to be scaled according to the audio duration of the storyboard audio. The target duration can include multiple scaled durations, and the scaled duration can be the specific duration after scaling a single first video. This scaled duration can be determined based on the audio duration of the storyboard audio.

[0058] The video generation device iterates through the number of videos in multiple first videos corresponding to the storyboard audio, in reverse order, based on the audio duration of the storyboard audio. When a certain value is found as the number of videos to be extracted, the scaling duration of each of the target durations corresponding to the multiple first videos is greater than or equal to the duration threshold. This value is the number of videos to be extracted, and the target duration is also determined.

[0059] In some embodiments, determining the number of extracted videos and the target duration based on the audio duration of the storyboard audio and a duration threshold may include: determining the number of videos in multiple first videos corresponding to the storyboard audio; determining the number of videos as the number to be processed; determining whether the scaling duration of each first video extracted according to the number to be processed and the audio duration is greater than or equal to the duration threshold; if so, determining the number to be processed as the number of extracted videos, and combining the scaling durations of each extracted first video as the target duration; otherwise, determining the number of videos minus one as the new number to be processed and returning to continue the determination until the number to be processed is one.

[0060] The number to be processed can be a calculated value used to determine whether the possible number of extracted videos can meet the duration threshold. The number to be processed starts from the number of videos of multiple first videos corresponding to the current storyboard audio and decreases in reverse order until a certain value is determined to be the number of extracted videos or the number to be processed is 1.

[0061] The video generation device can first divide multiple first videos according to the segmentation dimension. Assuming there are a total of f segments, the audio duration of the i-th segment can be expressed as... Suppose that the i-th shot includes m first videos, that is, the audio of the i-th shot corresponds to m first videos, and the first video can be represented as... The corresponding original duration is expressed as Let the number of videos be m, and let the number of videos to be processed be m. s ,1<<ms If << m, then after filtering and scaling the videos according to the number of videos to be processed and the audio duration, the first video... Scaling duration It can be represented as: Starting from m, the number of videos to be processed is traversed in reverse order. It is determined whether the scaling duration of all first videos corresponding to the number of videos to be processed is greater than or equal to the duration threshold. If a number of videos to be processed is found that satisfies the threshold, then the number of videos to be processed at this time is determined as the number of videos to be extracted, and the scaling duration of each first video corresponding to this number of videos to be processed is determined as the target duration. If the number of videos to be processed is 1, then 1 is directly used as the number of videos to be extracted, and the target duration only includes one scaling duration which is the audio duration.

[0062] Step 402: Extract multiple third videos from the multiple first videos corresponding to the storyboard audio according to the number of extracted videos.

[0063] The third video can be a first video extracted from multiple first videos included in the current storyboard, with the number of third videos being the number of extracted videos.

[0064] The video generation device can extract the first video from multiple first videos corresponding to the storyboard audio according to a preset order, based on the number of extracted videos at the top, to obtain multiple third videos. The preset order can be the specific extraction order of the multiple first videos, i.e., extracting the first videos according to a preset order. The preset order can include the upload order or the duration sorting result. The upload order can be the order in which the user uploaded each first video from the beginning to the end, and the duration sorting result can be the sorting result based on the duration of each first video in descending order.

[0065] Step 403: Perform duration scaling on multiple third videos according to the target duration to obtain multiple second videos corresponding to the storyboard audio.

[0066] Among them, duration scaling can shorten or lengthen the video duration, which can be achieved by adjusting the playback speed or trimming.

[0067] After extracting multiple third videos, the video generation device, since the target duration has been determined in the above embodiment, that is, the scaling duration of each first video has been determined, the scaling duration of each third video is known at this time; the original duration of each third video is adjusted to the corresponding scaling duration through scaling processing, and the corresponding second video can be obtained. The sum of the durations of the multiple second videos obtained after scaling processing of the multiple third videos corresponding to the storyboard audio is equal to the audio duration.

[0068] In some embodiments, multiple third videos are subjected to duration scaling processing according to a target duration to obtain multiple second videos corresponding to the storyboard audio. This includes: for each third video, determining a comparison duration based on the scaling duration of the third video at the target duration and preset parameters, and determining whether the original duration of the third video is less than or equal to the comparison duration. If so, the original duration of the third video is adjusted to the scaling duration by adjusting the playback speed to obtain the corresponding second video; otherwise, a video segment with the scaling duration is extracted from the third video and determined as the corresponding second video.

[0069] The preset parameter can be a parameter multiplied by the scaling duration to determine the comparison duration. This preset parameter is used to control the playback speed so that it is not too slow when adjusting it later. The specific setting depends on the actual situation; for example, the preset parameter can be 1.2. The comparison duration can be compared with the original duration of the video to determine the duration of which scaling example method to use.

[0070] When the video generation device scales multiple third videos, for each third video duration, it multiplies the scaled duration of the third video by a preset parameter to obtain a comparison duration, and determines whether the original duration of the third video is less than or equal to the comparison duration. If so, the original duration of the third video can be adjusted to the scaled duration by adjusting the playback speed to obtain a second video. Here, if the original duration of the third video is greater than the scaled duration, the playback speed is adjusted from 1 to a value greater than 1 to shorten the duration. If the original duration of the third video is less than the scaled duration, the playback speed is adjusted from 1 to a value less than 1 to stretch the duration. If the original duration of the third video is greater than the comparison duration, a video segment with the scaled duration can be extracted from the third video starting from the 1st second to determine the corresponding second video.

[0071] In the above scheme, by first determining how many videos to select for each segment to meet the duration threshold, the scheme avoids videos that are too short from affecting subsequent video generation. After determining the number of selected videos, the scaled duration can be determined. Then, various scaling methods are used to quickly scale the videos according to the scaled duration, which effectively improves the scaling effect. Finally, the duration of the audio and the corresponding video at the segment level is matched, which helps to improve the duration matching effect of subsequent video generation.

[0072] Step 104: Generate storyboard subtitles based on the storyboard audio and storyboard text.

[0073] Storyboard subtitles can be added to a storyboard according to the storyboard text, along with the corresponding timestamp.

[0074] Specifically, after determining the storyboard audio and storyboard text, the video generation device can divide the storyboard text into individual sentences, and then add timestamps for each sentence according to the audio duration of the storyboard audio using a uniform speed strategy, thereby obtaining the storyboard subtitles.

[0075] In some embodiments, generating storyboard subtitles based on storyboard audio and storyboard text may include: dividing the storyboard text into multiple single sentences according to punctuation marks and character count thresholds; determining the timestamp corresponding to each single sentence based on the audio duration of the storyboard audio according to a uniform speed strategy; and combining multiple single sentences and their corresponding timestamps to determine the corresponding storyboard subtitles.

[0076] The character count threshold can be set as the maximum number of characters for each sentence, and can be specifically set according to actual conditions; for example, the character count threshold can be 15. The uniform speed strategy can allocate the time of each sentence according to the number of characters when adding timestamps. The time of each subtitle is fixed; the more characters, the longer the time of a sentence. The timestamp corresponding to a sentence can be the time segment corresponding to the audio duration of that sentence in the storyboard audio, which can include the start and end times. The combined time segments corresponding to the timestamps of multiple sentences constitute the audio duration.

[0077] When generating storyboard subtitles, the video generation device can first divide the storyboard text into multiple sentences according to punctuation marks. Then, it extracts sentences with a character count greater than a character count threshold and continues to divide them from the middle. For all the resulting sentences, it extracts sentences whose segmentation points are English letters or numbers, moves the segmentation points backward, and re-divides them, resulting in the final multiple sentences. By limiting the character count and moving the segmentation points, the accuracy of sentence segmentation is improved. Assuming there are n sentences, the character count of a sentence is represented as s. i It can be known that s f This indicates the total number of characters in the storyboard text, and the audio duration of the storyboard audio is expressed as... The start time in the timestamp of the i-th single sentence can be represented as: The end time can be expressed as Then, multiple single sentences and the timestamp corresponding to each single sentence can be determined as the subtitles for the current scene. The above processing is performed for each scene to obtain multiple subtitles.

[0078] The execution order of steps 103 and 104 above is only an example. They can also be executed in parallel, or step 104 can be executed first and then step 103.

[0079] Step 105: Generate the target video based on multiple second videos, multiple storyboard audios, and multiple storyboard subtitles.

[0080] Specifically, after obtaining multiple storyboard audios, multiple storyboard subtitles, and multiple second videos according to the storyboard dimension, the video generation device can perform video synthesis using a video synthesis algorithm. During the synthesis process, other parameters in the basic video data can also be edited to finally obtain the target video, which can be a video generated based on the basic video data input by the user.

[0081] In some embodiments, generating a target video based on multiple second videos, multiple storyboard audios, and multiple storyboard subtitles may include: writing multiple second videos, multiple storyboard audios, and multiple storyboard subtitles into track editing parameters according to the storyboard dimension, and obtaining the target video using track editing methods according to the track editing parameters.

[0082] Track editing parameters are various parameters used to adjust the materials on the track when editing a video. Track editing parameters can include the position and size, rotation, transparency, color correction, special effects, etc. of video, audio, and subtitles. By adjusting these parameters, various materials generated from the video can be edited, composited, and have special effects processed to achieve the desired visual effect.

[0083] When generating a target video, the video generation device can construct track editing parameters for each scene based on the requirements of track editing parameters, using multiple second videos, multiple scene audios, and multiple scene subtitles. In addition, other parameters such as dubbing parameters, transition parameters, and background music in the basic video data can also be written into the track editing parameters in a prescribed format. Using these track editing parameters, the final target video is obtained by processing the video through a callback mechanism using track editing. By creating the video through track editing, the efficiency of video production is effectively improved.

[0084] The video generation scheme provided in this disclosure includes obtaining basic video data, which includes scene identifiers, target scripts, and multiple first videos; generating storyboard text and storyboard audio based on the scene information corresponding to the target scripts and scene identifiers; performing video filtering and duration scaling on the multiple first videos according to the storyboard dimensions based on the storyboard audio to obtain multiple second videos; generating storyboard subtitles based on the storyboard audio and storyboard text; and generating a target video based on the multiple second videos, multiple storyboard audios, and multiple storyboard subtitles. By adopting the above technical solution, storyboard text and audio can be automatically generated based on scene information corresponding to scene identifiers in the basic video data and the target script, thereby generating storyboard subtitles. The video is then filtered and its duration is scaled using the storyboard audio. Finally, the final target video is generated based on the processed video, storyboard audio, and storyboard subtitles. This achieves a complete and efficient automated video generation process with comprehensive functions, high ease of use, and significant savings in manual time and effort, greatly improving video generation efficiency. Furthermore, the text and audio generated based on scene information are more in line with actual application scenarios. The filtering and duration scaling of the storyboard video effectively improves the matching degree between video and audio, thereby improving the video quality of the generated video.

[0085] In some embodiments, after multiple first videos are filtered and their duration scaled according to the segment dimensions based on the segment audio to obtain multiple second videos, video generation may further include: determining the canvas size of the second video with the highest resolution and the most aspect ratios among the multiple second videos as the target canvas size; and scaling the video size of the multiple second videos according to the target canvas size.

[0086] The target canvas size can be a reference size for video scaling. This means that for multiple second videos, the video size needs to be adaptively scaled to the target canvas size. The target canvas size can include the canvas width and height. The video size of the second video can include the video height and width. Multi-scale scaling can be understood as scaling the dimensions of multiple second videos according to different ratios to avoid compressing or stretching any video and causing video distortion.

[0087] After determining multiple second videos in the above embodiments, the video generation device can extract the second video with the highest resolution and the largest number of aspect ratios among the multiple second videos, determine the canvas size of the second video as the target canvas size, scale the video sizes of the multiple second videos to the target canvas size in multiple proportions, and fill the excess parts with black borders.

[0088] Optionally, scaling the video sizes of multiple second videos according to the target canvas size can include: determining the video aspect ratio of each second video and the canvas aspect ratio of the target canvas size; for each second video, determining whether the video aspect ratio of the second video is greater than the canvas aspect ratio; if so, adjusting the video width of the second video to the canvas width of the target canvas size, and adjusting the video height according to the ratio of the canvas width to the video width; otherwise, adjusting the video height of the second video to the canvas height of the target canvas size, and adjusting the video width according to the ratio of the canvas height to the video height.

[0089] The canvas aspect ratio of the target canvas size can be the ratio of the canvas width to the canvas height in the target canvas size, and the video aspect ratio of the second video can be the ratio of the video width to the video height in the second video.

[0090] When the video generation device scales the dimensions of multiple second videos at multiple ratios, it assumes that the target canvas size is represented as (h c’ w c The canvas scale is represented by R. c =w c / h c The video size of the i-th second video can be represented as (h i’ w i Let i = 1, ..., k, where k represents the number of videos in the multiple second videos. The proportion of the i-th second video can be expressed as R. i =w i / h i The video size of the i-th second video is scaled to the target size, which includes the target width and the target height, and is represented as... The target size is determined based on a comparison between the video aspect ratio and the brush aspect ratio, expressed by the formula:

[0091]

[0092] When the aspect ratio of the second video is greater than the canvas aspect ratio, the target width for scaling the second video is the canvas width. The width ratio of the canvas width to the video width is calculated, and this ratio is multiplied by the video height to determine the target height. When the aspect ratio of the second video is less than or equal to the canvas aspect ratio, the target height of the second video is determined as the canvas height. The height ratio of the canvas height to the video height is calculated, and this ratio is multiplied by the video width to determine the target width. Then, the video width of the second video is adjusted to the target width, and the video height is adjusted to the target height.

[0093] In the above solution, after processing the video according to the segmentation dimension, each video can be scaled up and down in multiple proportions according to the canvas size of the video with the highest resolution and the largest aspect ratio. This allows the video size of each video to be adaptively scaled to the canvas size, which not only unifies the size of each video, but also prevents compression or stretching of any video from causing video distortion when editing horizontal and vertical screens.

[0094] The video generation process of this disclosure embodiment will be further illustrated by a specific example below. For example, Figure 5 This is a schematic diagram of a video generation process provided in an embodiment of the present disclosure, such as... Figure 5 As shown in the diagram, the video generation process can include: Start. Select Video Creation. Input a video list, script number, and other parameters, i.e., input the basic video data. Video Preprocessing. Script Generation, i.e., generating multiple storyboard scripts. Determine if voice-over is needed; if so, generate audio, i.e., generate multiple storyboard audios, and then perform intelligent editing; otherwise, perform video text recognition, add subtitles based on the recognized text, and then perform intelligent editing. Intelligent Editing, i.e., performs the video filtering, duration scaling, storyboard subtitle generation, and video size adjustment as described in the above embodiments. Composite Video, i.e., finally obtain the target video.

[0095] This solution provides a complete, efficient, and intelligent video generation workflow, integrating intelligent editing, high-quality script generation, realistic voice simulation, transitions, subtitles, and special effects. It offers a user-friendly interface, comprehensive functionality, and simple operation, allowing even those with no editing experience to easily create high-quality videos. By constraining users to upload different types of videos based on scene composition, it ensures that the visual content matches the script. When generating scripts, it uses well-designed prompts to address the issue of stiff scripts in large-scale models, resulting in natural transitions. Multiple script templates are provided for users to choose from, and the generated scripts can include descriptions of character design, tone, style, and output format. A voice replication solution is used to clone real human voices for narration, resulting in natural video effects and voiceovers that closely resemble real human voices, enhancing realism.

[0096] Figure 6 This is a schematic diagram of a video generation device provided in an embodiment of the present disclosure. The device can be implemented by software and / or hardware, and is generally integrated into an electronic device. Figure 6 As shown, the device includes:

[0097] The acquisition module 601 is used to acquire basic video data, wherein the basic video data includes scene identifiers, target scripts, and multiple first videos;

[0098] Storyboard module 602 is used to generate storyboard script and storyboard audio based on the target script and the scene information corresponding to the scene identifier;

[0099] The video editing module 603 is used to perform video filtering and duration scaling on the multiple first videos according to the segmentation dimension based on the segmentation audio to obtain multiple second videos;

[0100] Subtitle module 604 is used to generate storyboard subtitles based on the storyboard audio and the storyboard text;

[0101] The video generation module 605 is used to generate a target video based on the plurality of second videos, the plurality of storyboard audios, and the plurality of storyboard subtitles.

[0102] Optionally, the video editing module 603 includes:

[0103] The first unit is used to determine the number of videos to be extracted and the target duration based on the audio duration and duration threshold of the storyboard audio.

[0104] The second unit is used to extract multiple third videos from multiple first videos corresponding to the storyboard audio according to the number of extracted videos;

[0105] The third unit is used to perform duration scaling processing on the plurality of third videos according to the target duration to obtain a plurality of second videos corresponding to the storyboard audio.

[0106] Optionally, the first unit is used for:

[0107] Determine the number of videos in the multiple first videos corresponding to the storyboard audio;

[0108] The number of videos is determined as the number to be processed. It is then determined whether the scaling duration of each first video extracted according to the number to be processed and the audio duration is greater than or equal to the duration threshold. If so, the number to be processed is determined as the number of extracted videos, and the scaling duration of each extracted first video is combined to determine the target duration.

[0109] Otherwise, the value of minus one of the video count is determined as the new number to be processed, and the process continues until the number to be processed is one.

[0110] Optionally, the third unit is used for:

[0111] For each of the third videos, a comparison duration is determined based on the scaling duration corresponding to the target duration and preset parameters. It is then determined whether the original duration of the third video is less than or equal to the comparison duration. If so, the original duration of the third video is adjusted to the scaling duration by adjusting the playback speed to obtain the corresponding second video.

[0112] Otherwise, a video segment of the zoomed duration is extracted from the third video and determined as the corresponding second video.

[0113] Optionally, the subtitle module 604 is used for:

[0114] The storyboard text is divided into multiple single sentences according to punctuation marks and character count thresholds;

[0115] The timestamps corresponding to each sentence are determined according to a uniform speed strategy based on the audio duration of the storyboard audio.

[0116] The multiple single sentences and their corresponding timestamps are combined to determine the corresponding storyboard subtitles.

[0117] Optionally, the device further includes a scaling module, used for: performing video filtering and duration scaling on the plurality of first videos according to the storyboard dimensions based on the storyboard audio, to obtain a plurality of second videos,

[0118] The canvas size of the second video with the highest resolution and the largest number of aspect ratios among the plurality of second videos is determined as the target canvas size;

[0119] The video sizes of the plurality of second videos are scaled proportionally according to the target canvas size.

[0120] Optionally, the scaling module is specifically used to include:

[0121] Determine the video aspect ratio of each of the second videos and the canvas aspect ratio of the target canvas size;

[0122] For each of the second videos, determine whether the video aspect ratio of the second video is greater than the canvas aspect ratio. If so, adjust the video width of the second video to the canvas width of the target canvas size, and adjust the video height according to the ratio of the canvas width to the video width.

[0123] Otherwise, adjust the height of the second video to the height of the target canvas size, and adjust the video width according to the ratio of the canvas height to the video height.

[0124] Optionally, the video generation module 605 is used for:

[0125] The multiple second videos, multiple storyboard audios, and multiple storyboard subtitles are written into the track editing parameters according to the storyboard dimension, and the target video is obtained by using the track editing method according to the track editing parameters.

[0126] The video generation apparatus provided in this disclosure can execute the video generation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.

[0127] This disclosure also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the video generation method provided in any embodiment of this disclosure.

[0128] Figure 7 This is a schematic diagram of an electronic device provided in an embodiment of the present disclosure. See below for details. Figure 7 The diagram illustrates a structural schematic suitable for implementing the electronic device 700 in the embodiments of this disclosure. The electronic device 700 in the embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0129] like Figure 7 As shown, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0130] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0131] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 708, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the video generation method of embodiments of this disclosure.

[0132] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0133] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0134] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0135] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire basic video data, wherein the basic video data includes scene identifiers, a target script, and multiple first videos; generate storyboard text and storyboard audio based on the scene information corresponding to the target script and the scene identifiers; perform video filtering and duration scaling processing on the multiple first videos according to the storyboard dimensions based on the storyboard audio to obtain multiple second videos; generate storyboard subtitles based on the storyboard audio and the storyboard text; and generate a target video based on the multiple second videos, the multiple storyboard audios, and the multiple storyboard subtitles.

[0136] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0137] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0138] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.

[0139] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0140] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0141] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0142] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0143] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0144] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A video generation method, characterized in that, include: Acquire basic video data, which includes scene identifiers, target scripts, and multiple first videos; Based on the target script and the scene information corresponding to the scene identifier, generate storyboard text and storyboard audio; Based on the storyboard audio, the multiple first videos are filtered and their durations are scaled according to the storyboard dimensions to obtain multiple second videos; Generate subtitles for the storyboard based on the storyboard audio and the storyboard text; A target video is generated based on the plurality of second videos, the plurality of storyboard audios, and the plurality of storyboard subtitles; The process of filtering and scaling the duration of the multiple first videos according to the segmentation audio based on the segmentation dimensions yields multiple second videos, including: Determine the number of videos in the multiple first videos corresponding to the storyboard audio; The number of videos is determined as the number to be processed. It is determined whether the scaling duration of each of the first videos extracted according to the number to be processed and the audio duration of the storyboard audio is greater than or equal to the duration threshold. If so, the number to be processed is determined as the number of extracted videos, and the scaling duration of each of the extracted first videos is combined to determine the target duration. Otherwise, the value of the video count minus one is determined as the new number to be processed, and the process continues until the number to be processed is one. Based on the number of extracted videos and the target duration, the multiple first videos are filtered and their durations are scaled to obtain multiple second videos.

2. The method according to claim 1, characterized in that, Based on the number of extracted videos and the target duration, the multiple first videos are filtered and their durations are scaled to obtain multiple second videos, including: According to the number of extracted videos, extract multiple corresponding third videos from the multiple first videos corresponding to the storyboard audio; The multiple third videos are subjected to duration scaling based on the target duration to obtain multiple second videos corresponding to the storyboard audio.

3. The method according to claim 2, characterized in that, Based on the target duration, the plurality of third videos are subjected to duration scaling processing to obtain a plurality of second videos corresponding to the storyboard audio, including: For each of the third videos, a comparison duration is determined based on the scaling duration corresponding to the target duration and preset parameters. It is then determined whether the original duration of the third video is less than or equal to the comparison duration. If so, the original duration of the third video is adjusted to the scaling duration by adjusting the playback speed to obtain the corresponding second video. Otherwise, a video segment of the zoomed duration is extracted from the third video and determined as the corresponding second video.

4. The method according to claim 1, characterized in that, Generating storyboard subtitles based on the storyboard audio and the storyboard text includes: The storyboard text is divided into multiple single sentences according to punctuation marks and character count thresholds; The timestamps corresponding to each sentence are determined according to a uniform speed strategy based on the audio duration of the storyboard audio. The multiple single sentences and their corresponding timestamps are combined to determine the corresponding storyboard subtitles.

5. The method according to claim 1, characterized in that, After performing video filtering and duration scaling on the multiple first videos according to the segmentation dimensions based on the segmentation audio to obtain multiple second videos, the method further includes: The canvas size of the second video with the highest resolution and the largest number of aspect ratios among the plurality of second videos is determined as the target canvas size; The video sizes of the plurality of second videos are scaled proportionally according to the target canvas size.

6. The method according to claim 5, characterized in that, Scaling the video sizes of the plurality of second videos by multiple ratios according to the target canvas size, including: Determine the video aspect ratio of each of the second videos and the canvas aspect ratio of the target canvas size; For each of the second videos, determine whether the video aspect ratio of the second video is greater than the canvas aspect ratio. If so, adjust the video width of the second video to the canvas width of the target canvas size, and adjust the video height according to the ratio of the canvas width to the video width. Otherwise, adjust the height of the second video to the height of the target canvas size, and adjust the video width according to the ratio of the canvas height to the video height.

7. A video generation apparatus, characterized in that, include: The acquisition module is used to acquire basic video data, wherein the basic video data includes scene identifiers, target scripts, and multiple first videos; The storyboard module is used to generate storyboard scripts and storyboard audio based on the target script and the scene information corresponding to the scene identifier; The video editing module is used to perform video filtering and duration scaling on the multiple first videos according to the segmentation dimensions based on the segmentation audio, to obtain multiple second videos; The subtitle module is used to generate subtitles for the storyboard based on the storyboard audio and the storyboard text; The video generation module is used to generate a target video based on the plurality of second videos, the plurality of storyboard audios, and the plurality of storyboard subtitles; The process of filtering and scaling the duration of the multiple first videos according to the segmentation audio based on the segmentation dimensions yields multiple second videos, including: Determine the number of videos in the multiple first videos corresponding to the storyboard audio; The number of videos is determined as the number to be processed. It is determined whether the scaling duration of each of the first videos extracted according to the number to be processed and the audio duration of the storyboard audio is greater than or equal to the duration threshold. If so, the number to be processed is determined as the number of extracted videos, and the scaling duration of each of the extracted first videos is combined to determine the target duration. Otherwise, the value of the video count minus one is determined as the new number to be processed, and the process continues until the number to be processed is one. Based on the number of extracted videos and the target duration, the multiple first videos are filtered and their durations are scaled to obtain multiple second videos.

8. An electronic device, characterized in that, The electronic device includes: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the video generation method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program for executing the video generation method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Video generation method and device, electronic equipment, storage medium and product

    CN118803173A

  • Video generation method, computing device, computer storage medium and computer program product

    CN118972670A