A video generation method, apparatus, device, storage medium, and program product

By generating multiple candidate shot sequences and determining the target shot sequence, the problem of automatic 3D shot generation was solved, enabling diversified display of 3D video content and enhancing the visual experience of the video.

CN119402727BActive Publication Date: 2026-02-10MOFA (SHANGHAI) INFORMATION TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411507749.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-28
Publication Date
2026-02-10
Estimated Expiration
2044-10-28

AI Technical Summary

Technical Problem

Existing technologies cannot automatically generate 3D lenses, resulting in insufficient diversity in the content displayed in automatically generated 3D videos.

Method used

By acquiring dialogue text and shot information from the input content, multiple candidate shot sequences are generated, and the target shot sequence is determined based on the adjacency relationship between shots, ultimately generating the target video and realizing shot switching and diversified display.

Benefits of technology

It improves the diversity of video content display, showcasing different video content from different perspectives and distances, thus enhancing the video display effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119402727B_ABST
    Figure CN119402727B_ABST
Patent Text Reader

Abstract

A video generation method, device, equipment, storage medium and program product are disclosed. Input content is acquired; wherein the input content contains a script text and shot information; based on the script text and the shot information, multiple candidate shot sequences are generated; according to the adjacent relationship between shots in each candidate shot sequence, a target shot sequence is determined from the multiple candidate shot sequences; and based on the target shot sequence, a target video is generated. The video generation method provided in the embodiment of the present application automatically generates multiple candidate shot sequences based on the script text and the shot information, thereby realizing the diversification of the candidate shot sequences, determining the target shot sequence from the multiple candidate shot sequences, and generating the target video based on the target shot sequence. When the target video is played, various shots in the target shot sequence are switched in sequence, the target video displays different video content from different perspectives and distances, thereby improving the diversity of video content display.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video generation technology, and in particular to a video generation method, apparatus, device, storage medium and program product. Background Technology

[0002] The metaverse is a virtual digital world that creates an immersive and highly interactive virtual environment by combining technologies such as virtual reality, augmented reality, and artificial intelligence.

[0003] The metaverse requires the creation of a virtual environment outside the real physical environment through technological means, thus it is highly dependent on virtual reality technology. The basic implementation of virtual reality technology involves using computers and other devices to generate a realistic three-dimensional virtual world with multiple sensory experiences, including vision, touch, and smell. As is well known, vision is the primary way humans perceive the external world; therefore, three-dimensional (3D) content generation technology holds a crucial position in the field of virtual reality technology.

[0004] 3D content generation technology can be further subdivided into different branches such as 3D motion generation technology, 3D face generation technology, and 3D lens generation technology. Currently, it is possible to automatically generate 3D motion and 3D faces, but it is not possible to automatically generate 3D lenses. Therefore, automatically generated 3D videos usually use a single lens to display content, which affects the diversity of video content display. Summary of the Invention

[0005] This invention provides a video generation method, apparatus, device, storage medium, and program product that can generate videos with continuously switching camera angles, thereby improving the diversity of video content display.

[0006] In a first aspect, embodiments of the present invention provide a video generation method, including:

[0007] Obtain input content; wherein, the input content includes dialogue text and camera information;

[0008] Based on the dialogue text and the shot information, multiple candidate shot sequences are generated;

[0009] Based on the adjacency relationship between shots in each candidate shot sequence, a target shot sequence is determined from a plurality of candidate shot sequences;

[0010] Based on the target shot sequence, a target video is generated.

[0011] Secondly, embodiments of the present invention also provide a video generation apparatus, comprising:

[0012] An input content acquisition module is used to acquire input content; wherein, the input content includes dialogue text and camera information;

[0013] A candidate shot sequence generation module is used to generate multiple candidate shot sequences based on the dialogue text and the shot information;

[0014] The target shot sequence determination module determines the target shot sequence from a variety of candidate shot sequences based on the adjacency relationship between shots in each candidate shot sequence;

[0015] The target video generation module is used to generate a target video based on the target shot sequence.

[0016] Thirdly, embodiments of the present invention also provide an electronic device, the electronic device comprising:

[0017] At least one processor; and

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the video generation method described in the embodiments of the present invention.

[0020] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing computer instructions that are used to cause a processor to execute the video generation method described in the embodiments of the present invention.

[0021] Fifthly, embodiments of the present invention also provide a computer program product, including a computer program that, when executed by a processor, implements the video generation method as described in the embodiments of the present invention.

[0022] This invention discloses a video generation method, apparatus, device, storage medium, and program product. The method involves: acquiring input content, which includes dialogue text and shot information; generating multiple candidate shot sequences based on the dialogue text and shot information; determining a target shot sequence from the multiple candidate shot sequences based on the adjacency relationships between shots in each candidate shot sequence; and generating a target video based on the target shot sequence. The video generation method provided by this invention automatically generates multiple candidate shot sequences based on dialogue text and shot information, thereby achieving diversification of candidate shot sequences. Then, a target shot sequence is determined from the multiple candidate shot sequences, and a target video is generated based on the target shot sequence. When playing the target video, as the various shots in the target shot sequence switch sequentially, the target video displays different video content from different perspectives and distances, thereby improving the diversity of video content display. Attached Figure Description

[0023] Figure 1 This is a flowchart of a video generation method according to an embodiment of the present invention;

[0024] Figure 2 This is an example diagram of input content acquired in one embodiment of the present invention;

[0025] Figure 3 This is an example diagram illustrating the determination of a candidate shot sequence in an embodiment of the present invention;

[0026] Figure 4a This is an example diagram of a video frame displayed by a lens according to a content display category in an embodiment of the present invention;

[0027] Figure 4b This is an example diagram of a video frame displayed in a camera lens representing a virtual human category in an embodiment of the present invention;

[0028] Figure 5 This is a flowchart of a video generation method according to an embodiment of the present invention;

[0029] Figure 6 This is a schematic diagram of the structure of a video generation device according to an embodiment of the present invention;

[0030] Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0031] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.

[0032] Figure 1 This is a flowchart of a video generation method provided in an embodiment of the present invention. This embodiment is applicable to automatically generating videos in which content is displayed by sequential switching of multiple lenses. The method can be executed by a video generation device, which can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, such as a mobile terminal, a PC, or a server.

[0033] As explained in the background section regarding metaverse-related technologies, the metaverse is a virtual environment outside the real physical environment. This virtual environment contains virtual humans (also known as "virtual digital humans") with functions such as content display, information interaction, and communication. They can perform roles such as virtual news anchors, virtual teachers, and virtual broadcasters. 3D motion generation and 3D facial generation technologies within 3D content generation are primarily used to generate virtual humans.

[0034] In different scenarios, users can receive information from virtual humans by watching videos, interact with virtual humans by typing text, and communicate with virtual humans through voice and video dialogues.

[0035] The video generation method provided in this embodiment of the invention is mainly used to generate videos containing virtual humans for content display to users.

[0036] In this embodiment, the user inputs at least a piece of text, and after generating audio based on that text, a video of a virtual person explaining the text can be quickly generated. For example, this could be a news broadcast video, a product introduction video, or a science popularization video. Furthermore, to enhance the video's presentation, this embodiment also allows the user to input video materials, such as images or videos, while inputting the aforementioned text. These images or videos can be used as content displayed during the virtual person's explanation in the generated video. Further, for ease of user operation, the aforementioned video materials can be input in the form of static images, animated images (such as GIFs), videos, PPT documents, or PDF documents. It should be understood that each PPT document contains at least one slide, and one slide is equivalent to one image. Text, WordArt, images, videos, and other content can be inserted into the slide. Similarly, each PDF document contains at least one PDF page, and one PDF page is equivalent to one image. The PDF page contains text, images, and other elements. For example, in one application scenario, a user inputs a PPT document and corresponding dialogue text and shot information. The PPT document contains text, images, and videos. This embodiment can generate a video of a virtual human explaining the PPT document with sequentially changing shots based on the input content. The video can show the user the image of the virtual human and video materials from different perspectives and distances.

[0037] It should be noted that each slide in the PPT document has corresponding notes. Therefore, the dialogue text and shot information in the PPT document can be entered either as independent text or as notes in the PPT document. This embodiment does not limit this.

[0038] like Figure 1 As shown, the method specifically includes the following steps:

[0039] S110, Obtain input content.

[0040] The input includes dialogue text and camera information.

[0041] The dialogue text can be directly entered by the user through the client, retrieved from the client's local storage, or downloaded from the server; there is no restriction on the method of obtaining the dialogue text. The dialogue text can contain only one text fragment or multiple text fragments; there is no restriction on this either.

[0042] As explained above, the video generated in this embodiment is a video of a virtual human explaining a script. The script is a text used by the virtual human to explain the content, and it can be represented in a certain document format, such as a TXT document, a WORD document, or a PDF document. The script is used to generate the voice-over for the video, that is, the virtual human's audio. The voice-over will be present throughout the entire video, and its duration is the same as the duration of the video to be generated, and their timelines are also the same.

[0043] Optionally, the input content may also include images, specifically pictures or videos, which are presented in the generated video to assist the virtual human in explaining the content.

[0044] Understandably, in formal settings such as product launches and educational training sessions, verbal explanations alone can be rather dry. They need to be accompanied by images or videos to deepen the audience's understanding and leave a lasting impression. Therefore, this embodiment also allows users to input images or videos as explanatory material, combining them with the virtual human's language and actions to create a high-quality video.

[0045] Furthermore, this embodiment can also convert dialogue text into corresponding dubbing based on a user-selected voice template (e.g., baritone, bass, soprano, etc.). The voice template can be a pre-designed voice available for user selection or generated from user-recorded speech, and the user-selected voice template matches the virtual character's image. In other words, it selects a suitable timbre for the dubbing, ensuring visual and auditory consistency for the virtual character, making the overall character richer and more complete.

[0046] As explained above, this embodiment can generate video dubbing based on dialogue text, and then generate the corresponding video based on the dubbing. Therefore, one possible implementation is to directly use the dubbing instead of the dialogue text as input content, and use the timeline of the dubbing as the timeline of the video to be generated.

[0047] In addition to the dialogue text, the input content in this embodiment also includes camera information. The camera information includes at least one camera category, which may include at least one of content display shots and character display shots.

[0048] As explained above, the video generated in this embodiment is a video of a virtual human explaining a script. The main elements of the video include the image of the virtual human, the virtual environment in which the virtual human is located, and the material being explained by the virtual human. Correspondingly, the shot categories in this embodiment are mainly classified according to the object the shot focuses on. Content display shots refer to shots where the focus is on the video content (i.e., the material being explained by the virtual human). In this type of shot, the video content is focused on and displayed; that is, in the corresponding frame of this type of shot, the virtual human is secondary, or may not even appear in the frame. Full-view shots and panoramic shots are typical content display shots. Character display shots refer to shots where the focus is on the virtual human. In this type of shot, the virtual human is focused on and displayed; that is, in the corresponding frame of this type of shot, the virtual human is the main subject. Medium shots, close-ups, and extreme close-ups are typical character display shots.

[0049] Furthermore, to make the generated video more professional, the shot categories in this embodiment can also include start shots and end shots. Start shots and end shots refer to shots without a clearly defined focus object. That is, the corresponding image of this type of shot moves continuously from top to bottom, from left to right, from far to near, and from near to far, etc., to display the virtual environment in which the virtual person is located in a comprehensive and three-dimensional way, allowing the audience to naturally approach / leave the current virtual environment along with the changes in the camera.

[0050] Understandably, the opening shot is typically placed at the beginning of a video to showcase the overall environment, establish the tone, or highlight key elements of the scene. For example, the opening shot could be a panoramic shot, a close-up shot, or a wide-angle shot. The closing shot is typically placed at the end of a video to summarize the storyline, leave a lasting impression on the viewer, or reinforce the video's theme or message. For example, the closing shot could be a slow pull-back shot.

[0051] In this embodiment, the shot information is used to mark the text content in the dialogue text, thereby further segmenting the text fragments of the dialogue text into multiple shot fragments according to the shot information markings. Each shot fragment corresponds to a shot category. It should be understood that if the dialogue text contains only one text fragment, the shot information markings are used to segment the entire dialogue text into multiple shot fragments; if the dialogue text contains multiple text fragments, the shot information markings are used to segment each text fragment into multiple shot fragments. In other words, the dialogue text, text fragments, and shot fragments constitute a three-layer structure for segmenting the dialogue text.

[0052] Specifically, when users input dialogue text, they can use different camera information to mark the text content in the dialogue text. The marked text content is then segmented from the adjacent text content to form different camera segments.

[0053] It's important to note that for video, each frame corresponds to a virtual camera lens (referred to as a "lens"). The virtual camera lens renders the scene by simulating the shooting effect of a real camera lens, thus generating the image. Therefore, lenses can be used to define the viewpoint, distance, and duration of the displayed content. Different lenses can be defined using different technical parameters. Each lens's parameters include at least one of the following: focal length, camera pose, and duration. Focal length can be understood as the distance from the lens's optical center to the image plane. A larger focal length (e.g., a telephoto lens) results in a smaller viewpoint and a narrower scene range in the image; a smaller focal length (e.g., a wide-angle lens) results in a larger viewpoint and a wider scene range in the image. Camera pose can include the camera's position and orientation, representing the lens's viewing direction. Duration can be understood as the duration of the lens. It should be understood that within the duration of a lens, parameters such as focal length and camera pose can change over time, causing the corresponding image to change accordingly.

[0054] Since each frame in the video needs to correspond to a shot, if the user only marks a part of the text in the dialogue, this embodiment will automatically mark the shot information for the unmarked text.

[0055] In addition, considering the special nature of the beginning and end of a video, this embodiment allows for mandatory setting of the first 5 seconds and the last 5 seconds of the video, in order to prevent users from forgetting to mark the beginning and end shots, or from mistakenly setting the beginning and end of the video to other types of shots due to operational errors.

[0056] For example, Figure 2 This is an example diagram of the input content obtained in this embodiment, such as... Figure 2As shown, the dialogue text contains three text segments: text segment 1, text segment 2, and text segment 3. Users only use content display shots to mark portions of the text. In this embodiment, the beginning shot is automatically marked at the start of text segment 1, the end shot at the end of text segment 3, and character display shots are marked for the remaining text content. This further divides the three text segments into 2 to 3 shot segments, meaning the entire dialogue text is divided into 8 shot segments, each carrying shot information.

[0057] S120 generates multiple candidate shot sequences based on dialogue text and shot information.

[0058] As explained above regarding the generation of dubbing from dialogue text, the timeline of the dubbing is the same as the timeline of the video to be generated. That is, the dubbing for the target video is generated based on the dialogue text, and the timeline of the dubbing is used as the timeline of the target video. The target video here is the video to be generated.

[0059] The dialogue text is segmented into multiple shot fragments. After generating the dubbing based on the dialogue text, the entire dialogue text corresponds to the entire timeline. Each shot fragment then corresponds to a specific segment on the dubbing timeline, meaning there's a one-to-one correspondence between shot fragments and segments on the dubbing timeline. Each shot fragment carries shot information to identify different shot categories. Therefore, a relationship can be established between segments on the timeline and shot categories; that is, segments on the timeline are also associated with the shot categories carried by their corresponding shot fragments.

[0060] In simple terms, this embodiment first obtains the shot segments corresponding to the shot categories in the dialogue text, and then obtains the time periods corresponding to each shot segment on the time axis to establish a relationship between each time period and the shot category, thereby realizing the division of the time axis into different types of shot time periods based on the shot category.

[0061] As explained above, the dialogue text in this embodiment can contain only one text segment or multiple text segments. When the dialogue text contains multiple text segments, the different text segments are pre-separated, and the entire dialogue text corresponds to the entire timeline. Accordingly, the timeline consists of multiple time slices, with each text segment corresponding to a time slice. Dividing each text segment into multiple shot segments also divides each time slice into different types of shot periods.

[0062] To avoid a single time slice containing only one type of shot segment, such as shots only showing characters, this embodiment, after dividing the timeline into different types of shot segments, also detects the type of shot segment contained in each time slice. If a time slice does not contain a shot segment showing content, the beginning of the time slice (e.g., the first 5 seconds) will be replaced with a shot segment showing content, thereby diversifying the types of shot segments contained in that time slice. This ensures that each time slice includes at least one shot segment of the type showing content, preventing the generated target video from being without content display for an extended period.

[0063] To ensure the professionalism of the shots in the generated video, this embodiment pre-designs multiple different shot combinations for various shot categories. These shot combinations are pre-designed by professional directors based on shooting theory and years of shooting experience, and are stored as source material in the databases corresponding to each shot category, thus forming shot libraries for each shot category. Each shot library contains multiple shot combinations, and each shot combination consists of one or more shots. For example, a shot combination in the character display shot library is: medium shot front view - close-up front view - close-up side view.

[0064] After dividing the timeline into different types of shot segments, each segment in the timeline is associated with a shot category, and each shot category corresponds to a shot library, thus linking each segment in the timeline to the shot library.

[0065] Understandably, to determine the shot sequence corresponding to the entire video, one can first select the corresponding shot combination for each shot segment on the timeline, and then concatenate all the shot combinations in sequence to obtain a shot sequence corresponding to the entire video.

[0066] To ensure the quality of the shot sequence, this embodiment selects shot combinations randomly from the corresponding shot library based on the correlation between the shot time period and the shot library, and then concatenates them in sequence to obtain a candidate shot sequence. Repeating this random selection and sequential concatenation multiple times yields multiple candidate shot sequences. For example, using... Figure 3 For example, there are 8 shot segments on the timeline. By randomly selecting a shot combination from the shot library corresponding to each shot segment and stringing them together in sequence, a candidate shot sequence containing 8 shot combinations can be obtained.

[0067] This involves selecting at least one shot combination from a shot library corresponding to different shot categories and binding it to a shot time segment of the corresponding type to generate multiple candidate shot sequences. These candidate shot sequences are then filtered and scored to determine the target shot sequence, which serves as the shot sequence corresponding to the target video.

[0068] It is understandable that when generating multiple candidate shot sequences, the number of candidate shot sequences can be preset, and then the target shot sequence can be determined from the preset number of candidate shot sequences. On the one hand, the diversity of candidate shot sequences can be increased by randomly selecting shot combinations, and on the other hand, the generation efficiency of the target shot sequence can be improved by limiting the number of candidate shot sequences generated.

[0069] It should be noted that when selecting at least one lens combination from the lens library corresponding to different lens categories, it is also necessary to consider whether the duration of the lens segment matches the duration of the lens combination. This can be understood as follows: each lens parameter includes its duration, and the duration of a lens combination is the sum of the durations of all the lenses it contains.

[0070] Correspondingly, the method of selecting at least one lens combination from the lens library corresponding to different lens categories can be: based on the duration of the lens time period, filter out matching lens combinations from the lens library corresponding to the lens category.

[0071] Specifically, for each shot segment, firstly, all shot combinations in the shot library corresponding to the shot category are filtered based on the duration of the shot segment. Shot combinations with the same duration or whose duration exceeds the shot segment by less than 20% are selected as matching shot combinations. Then, a shot combination is randomly selected from the matching shot combinations as the shot combination corresponding to that shot segment.

[0072] Optionally, if the duration of the shot combination is equal to the duration of the shot segment, then the shot combination is directly determined as the shot combination that matches the shot segment; if the duration of the shot combination is longer than the duration of the shot segment, then the duration of the shot combination is proportionally reduced, and the proportionally reduced shot combination is determined as the shot combination that matches the shot segment.

[0073] Specifically, the method for proportionally reducing shot combinations can be as follows: divide the duration of the shot combination by the duration of the shot segment to obtain the ratio of shot combination to shot segment. Reduce the duration of all shots included in the shot combination according to this ratio, that is, divide the duration of each shot by this ratio to obtain the reduced duration. For example, if the duration of the shot segment is T1, and the duration of shot combination b is T2, T2 > T1, and shot combination b includes shot 1 (duration t1), shot 2 (t2), and shot 3 (t3), then calculate the value of T2 / T1, denoted as T. Then divide the durations t1, t2, and t3 of shot 1, shot 2, and shot 3 by T to obtain the reduced duration. Then, the proportionally reduced shot combination b includes shot 1 (t1 / T), shot 2 (t2 / T), and shot 3 (t3 / T).

[0074] Furthermore, this embodiment can pre-set shot switching rules and / or shot concatenation rules to improve the quality of candidate shot sequences. For example, a medium shot can be followed by a close-up shot, but a close-up shot cannot be followed by a close-up shot. Since the selection of shot combinations corresponding to each shot time period is random, concatenating these shot combinations in sequence may result in situations where the last shot of one shot combination cannot be switched / concatenated with the first shot of the next shot combination. To solve this problem, this embodiment can select shot combinations from front to back in a sequential manner when selecting corresponding shot combinations for each shot time period. That is, first select shot combinations for the first shot time period, and then, based on the last shot in the first shot combination, filter the shot combinations in the shot library corresponding to the second shot time period, eliminating shot combinations that do not conform to the preset shot switching rules and / or shot concatenation rules, and selecting from the remaining shot combinations, thereby ensuring the quality of the generated candidate shot sequence.

[0075] Considering that the video generated in this embodiment is a video of a virtual person explaining the script, the virtual person in the video will walk and move during the explanation. Accordingly, the character display shot specifically includes the character follow shot. This type of shot is a special character display shot. In addition to focusing on the virtual person, it can also follow the virtual person when walking or moving to prevent the virtual person from leaving the frame due to walking or moving.

[0076] To link the character's following of the camera with the virtual human's walking or moving behavior on the timeline, this embodiment, after generating the dubbing for the target video based on the dialogue text, further generates an initial video based on the dubbing. The initial video contains footage of the virtual human's actions or movements, and it is a video showcasing content in a single shot.

[0077] Specifically, you can use voiceover as input for the virtual human video generation software, select a virtual human image in the software, and the software will automatically generate an initial video of the virtual human explaining the content. The virtual human may also explain the content while walking around and stopping.

[0078] The initial video is inspected. If a virtual person is detected walking or moving within the video, meaning the initial video contains a segment of movement, the corresponding time period is determined as the movement period on the initial video's timeline. Since the initial and target videos have the same timeline, the corresponding movement period is also determined on the target video's timeline.

[0079] In this embodiment, if a movement period exists in the timeline, a character-following shot from a character display shot or a panoramic shot from a content display shot is bound to that movement period. Then, following the aforementioned steps, shot combinations are bound to other time periods in the timeline besides the movement period, thereby determining the shot sequence corresponding to the entire target video to be generated. It can be understood that both character-following shots and panoramic shots ensure that the virtual person remains in the frame while walking or moving, preventing the virtual person from appearing out of the frame.

[0080] The aforementioned method for determining movement periods requires generating an initial video based on dubbing. However, there is another method for determining movement periods. Specifically, the input content in this embodiment also includes a video script, which includes at least one of the following: a movement marker. The movement marker can mark textual content in the dialogue, used to indicate that the virtual human in the corresponding video segment needs to walk or move.

[0081] As explained above, the video clips in the dialogue text correspond one-to-one with the video segments on the timeline. Therefore, based on the video clips corresponding to the movement markers, the corresponding movement segments can be determined on the timeline. After determining the movement segments on the timeline, these segments are bound to the character-following shots in the character display shots. This ensures that in the generated target video, when the virtual person walks or moves, the camera always follows the virtual person, guaranteeing that the virtual person remains in the preset position on the screen and preventing the virtual person from leaving the frame.

[0082] S130, determine the target shot sequence from multiple candidate shot sequences based on the adjacency relationship between shots in each candidate shot sequence.

[0083] It should be understood that the shot sequence in the final generated target video is unique. Therefore, after generating multiple candidate shot sequences, it is necessary to determine the best shot sequence as the target shot sequence for generating the target video.

[0084] The quality of a shot sequence depends mainly on the switching and stitching between different shots. Therefore, this embodiment determines the target shot sequence based on the adjacency relationship between shots in each candidate shot sequence.

[0085] One possible approach is to evaluate each candidate shot sequence based on preset shot switching rules and / or preset shot concatenation rules to obtain a score for each candidate shot sequence; and then determine the candidate shot sequence with the highest score as the target shot sequence.

[0086] One method for evaluating at least one candidate shot sequence based on preset shot switching rules is as follows: First, the switching methods of each pair of adjacent shot combinations in the candidate shot sequence are scored to obtain multiple shot switching sub-scores. Then, the multiple shot switching sub-scores are weighted and summed to obtain the shot switching score of the candidate shot sequence. Specifically, the method for scoring the switching methods of each pair of adjacent shot combinations is as follows: For adjacent shot combinations, the last shot of the previous shot combination and the first shot of the next shot combination are obtained. The switching methods of the last shot and the first shot are scored according to the preset shot switching rules to obtain the shot switching sub-score of the adjacent shot combination.

[0087] One method for evaluating at least one candidate shot sequence based on preset shot stitching rules is as follows: for each candidate shot sequence, evaluate the candidate shot sequence according to the preset stitching rules to obtain a shot stitching score.

[0088] In this embodiment, after obtaining the shot transition score and the shot sequence score, the shot transition score and the shot sequence score are accumulated to obtain the score of the candidate shot sequence. Finally, the candidate shot sequence with the highest score is determined as the target shot sequence. In this embodiment, the target shot sequence is determined from at least one candidate shot sequence based on preset shot transition rules and / or preset shot sequence rules, making the shot transitions of the generated video more consistent with the visual effects, thereby improving the display effect of the generated video.

[0089] S140, generates target video based on target shot sequence.

[0090] Specifically, after obtaining the target shot sequence, a target video is generated based on the target shot sequence, so that the target video switches shots according to the target shot sequence during playback.

[0091] Alternatively, the initial video generated in the above embodiments can be adjusted based on the target shot sequence to obtain the target video.

[0092] For example, Figure 4a This is an example diagram of a video frame displayed by the content display lens in this embodiment, such as... Figure 4aThe image shown is a video frame from a video of a virtual person broadcasting a weather forecast. The content of this video frame is the main content of the video frame. Figure 4b This is an example image of a video frame displayed in this embodiment using a person's image as a subject. Figure 4b The image shown is another video frame from a video of a virtual person delivering a weather forecast, in which the virtual person is the main subject.

[0093] The technical solution of this embodiment involves acquiring input content, which includes dialogue text and shot information; generating multiple candidate shot sequences based on the dialogue text and shot information; determining a target shot sequence from the multiple candidate shot sequences based on the adjacency relationship between shots in each candidate shot sequence; and generating a target video based on the target shot sequence. The video generation method provided by this embodiment automatically generates multiple candidate shot sequences based on dialogue text and shot information, thereby achieving diversification of candidate shot sequences. Then, a target shot sequence is determined from the multiple candidate shot sequences, and a target video is generated based on the target shot sequence. When playing the target video, as the various shots in the target shot sequence switch sequentially, the target video displays different video content from different perspectives and distances, thereby improving the diversity of video content display.

[0094] Based on the above embodiments, Figure 5 This is a flowchart of a video generation method provided by the present invention, as shown in the figure. The method includes the following steps:

[0095] S501, Get the input content.

[0096] S502, Based on the dialogue text, generate the dubbing for the target video; use the dubbing timeline as the target video timeline.

[0097] S503 divides the timeline into different types of shot segments based on shot category.

[0098] S504, based on the duration of the shot segment, selects matching shot combinations from the shot library corresponding to the shot category and binds them to the corresponding shot segment.

[0099] S505, detect the shot time period type contained in each time slice; in response to a time slice that does not contain a shot time period of content display shot, replace the start time period of the time slice with the shot time period of content display shot.

[0100] S506, Based on the voiceover, generate an initial video; In response to a video segment containing character movement in the initial video, determine the corresponding movement period in the timeline; Bind the movement period to the character follow-up shot in the character display shot.

[0101] S507 generates multiple candidate shot sequences.

[0102] S508: Evaluate each candidate shot sequence based on preset shot switching rules and / or preset shot concatenation rules to obtain a score for each candidate shot sequence; determine the candidate shot sequence with the highest score as the target shot sequence.

[0103] S509 generates target video based on target shot sequence.

[0104] Figure 6 This is a schematic diagram of the structure of a video generation device provided in an embodiment of the present invention, as shown below. Figure 6 As shown, the device includes:

[0105] The input content acquisition module 610 is used to acquire input content; wherein, the input content includes dialogue text and camera information;

[0106] The candidate shot sequence generation module 620 is used to generate multiple candidate shot sequences based on dialogue text and shot information;

[0107] The target shot sequence determination module 630 determines the target shot sequence from multiple candidate shot sequences based on the adjacency relationship between shots in each candidate shot sequence;

[0108] The target video generation module 640 is used to generate target videos based on target shot sequences.

[0109] Optionally, the lens information includes at least one lens category, and the candidate lens sequence generation module 620 is also used for:

[0110] Based on the dialogue text, generate dubbing for the target video;

[0111] Use the timeline of the dubbing as the timeline of the target video;

[0112] Based on shot categories, the timeline is divided into different types of shot segments; among which, shot categories include at least one of content display shots and character display shots;

[0113] At least one combination of shots is selected from the shot library corresponding to different shot categories and bound to the corresponding shot time period to generate multiple candidate shot sequences.

[0114] Optionally, the dialogue text contains multiple segments, and correspondingly, the timeline consists of multiple time slices. The candidate shot sequence generation module 620 is also used for:

[0115] The type of shot segment contained in each time slice is detected;

[0116] In response to a time slice that does not contain a shot segment of content display, the start segment of the time slice is replaced with the shot segment of the content display.

[0117] Optionally, each shot library contains multiple shot combinations, each shot combination consisting of one or more shots. The candidate shot sequence generation module 620 is also used for:

[0118] Based on the duration of the shot, matching shot combinations are selected from the shot library corresponding to the shot category.

[0119] Optionally, it also includes: an initial video generation module, used for:

[0120] Generate an initial video based on the voice-over;

[0121] Optionally, the candidate shot sequence generation module 620 is also used for:

[0122] In response to a video segment in the initial video that contains a moving person, the corresponding movement period is determined in the timeline;

[0123] Link the movement time period to the following shot of the person in the character display shot.

[0124] Optionally, the target shot sequence determination module 630 is also used for:

[0125] Each candidate shot sequence is evaluated based on preset shot switching rules and / or preset shot concatenation rules to obtain a score for each candidate shot sequence;

[0126] The candidate shot sequence with the highest score is selected as the target shot sequence.

[0127] Optionally, the lens parameters may include at least one of focal length, camera pose, and duration.

[0128] The above-described apparatus can execute the methods provided in all the foregoing embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the above methods. Technical details not described in detail in this embodiment can be found in the methods provided in all the foregoing embodiments of the present invention.

[0129] Figure 7A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components, connections and relationships between components, and their functions shown herein are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0130] like Figure 7 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0131] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0132] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as video generation methods.

[0133] In some embodiments, the video generation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the video generation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the video generation method by any other suitable means (e.g., by means of firmware).

[0134] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0135] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0136] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0137] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0138] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0139] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0140] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the video generation method provided in any embodiment of this application.

[0141] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0142] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0143] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A video generation method, characterized in that, include: Obtain input content; wherein, the input content includes dialogue text and shot information; the shot information is used to mark the text content in the dialogue text, so as to divide the text fragments in the dialogue text into multiple shot fragments according to the markings of the shot information, and each shot fragment corresponds to a shot category; Based on the dialogue text and the shot information, multiple candidate shot sequences are generated; wherein each candidate shot sequence includes multiple shot combinations, and each shot combination consists of one or more shots; Based on the adjacency relationship between shots in each candidate shot sequence, a target shot sequence is determined from a plurality of candidate shot sequences; Based on the target shot sequence, generate the target video; The lens is a virtual camera lens, which renders the scene by simulating the shooting effect of a real camera lens; the target video is a video of a virtual person explaining something, and the script is a text of the virtual person explaining the content, used to generate the audio of the virtual person. The step of determining the target shot sequence from a plurality of candidate shot sequences based on the adjacency relationship between shots in each candidate shot sequence includes: The candidate shot sequences are evaluated based on preset shot switching rules and / or preset shot concatenation rules to obtain a score for each candidate shot sequence; The candidate shot sequence with the highest score is selected as the target shot sequence.

2. The method according to claim 1, characterized in that, The shot information includes at least one shot category, and based on the dialogue text and the shot information, multiple candidate shot sequences are generated, including: Based on the dialogue text, generate the dubbing for the target video; Use the timeline of the dubbing as the timeline of the target video; Based on the shot categories, the timeline is divided into different types of shot segments; wherein, the shot categories include at least one of content display shots and character display shots; At least one combination of shots is selected from the shot library corresponding to different shot categories and bound to the corresponding shot time period to generate multiple candidate shot sequences.

3. The method according to claim 2, characterized in that, The dialogue text comprises multiple segments, and correspondingly, the timeline consists of multiple time slices. After dividing the timeline into different types of shot segments based on the shot category, it also includes: The type of shot segment contained in each of the time slices is detected; In response to the time slice not containing a shot segment of the content display shot, the start time segment of the time slice is replaced with the shot segment of the content display shot.

4. The method according to claim 2, characterized in that, Each of the lens libraries contains multiple lens combinations, and selecting at least one lens combination from the lens libraries corresponding to different lens categories includes: Based on the duration of the shot segment, a matching shot combination is selected from the shot library corresponding to the shot category.

5. The method according to claim 4, characterized in that, After generating the dubbing for the target video based on the dialogue text, the method further includes: Based on the dubbing, an initial video is generated; Accordingly, the step of selecting at least one lens combination from a lens library corresponding to different lens categories further includes: In response to a video segment containing a moving person in the initial video, the corresponding movement period is determined in the timeline; The movement time period is linked to the following shot of the character in the character display shot.

6. The method according to claim 1, characterized in that, The parameters of the lens include at least one of focal length, camera pose, and duration.

7. A video generation apparatus, characterized in that, include: An input content acquisition module is used to acquire input content; wherein, the input content includes dialogue text and shot information; the shot information is used to mark the text content in the dialogue text, so as to divide the text fragments in the dialogue text into multiple shot fragments according to the markings of the shot information, and each shot fragment corresponds to a shot category; The candidate shot sequence generation module is used to generate multiple candidate shot sequences based on the dialogue text and the shot information; wherein each candidate shot sequence includes multiple shot combinations, and each shot combination consists of one or more shots; The target shot sequence determination module determines the target shot sequence from a variety of candidate shot sequences based on the adjacency relationship between shots in each candidate shot sequence; A target video generation module is used to generate a target video based on the target shot sequence; The lens is a virtual camera lens, which renders the scene by simulating the shooting effect of a real camera lens; the target video is a video of a virtual person explaining something, and the script is a text of the virtual person explaining the content, used to generate the audio of the virtual person. The target shot sequence determination module is specifically used for: The candidate shot sequences are evaluated based on preset shot switching rules and / or preset shot concatenation rules to obtain a score for each candidate shot sequence; The candidate shot sequence with the highest score is selected as the target shot sequence.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the video generation method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the video generation method according to any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the video generation method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Video generation method, deep learning model training method, device and equipment

    CN116112621A

  • Virtual image animation generation method and device, electronic equipment and storage medium

    CN116168122A

  • Intelligent director method and system

    CN117499685A

  • UGC dialogue shot editing method and device, storage medium and electronic device

    CN118203837A