Video generation method and device, computer equipment and storage medium
By decoding video clips and extracting multimodal features, determining the narrative intent, and selecting appropriate transition effects for splicing, the problem of low efficiency and poor quality in existing video generation technologies is solved, achieving efficient and high-quality personalized video generation.
Patent Information
- Application Number
- CN202511649503.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2026-02-10
Smart Images

Figure CN121509770A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a video generation method, apparatus, computer equipment, storage medium, and computer program product. Background Technology
[0002] With the development of computer and internet technologies and the arrival of the 5G era, the internet has brought great convenience to modern life, and applications supporting personalized video generation are becoming increasingly widespread. For example, users can use tools to create personalized video content. By using preset templates or rules in these tools, user-uploaded video clips can be automatically spliced together to quickly generate personalized video content for further creation.
[0003] However, in current video generation methods, when users want to create personalized videos, such as in the context of video content creation, they can either manually operate the process or use video tools to connect multiple video clips to form a complete personalized work. But for both professional editors and ordinary users, manual processing is extremely time-consuming and labor-intensive. Automated tools, on the other hand, cannot meet the personalized needs of different users, resulting in low efficiency and poor quality in video creation. Therefore, how to improve video generation efficiency while simultaneously enhancing video quality has become a pressing issue. Summary of the Invention
[0004] Therefore, it is necessary to provide a video generation method, apparatus, computer equipment, computer-readable storage medium, and computer program product to address the aforementioned technical problems, which can effectively improve the efficiency of video generation while ensuring the quality of video generation.
[0005] In a first aspect, this application provides a video generation method. The method includes: decoding at least two video segments to obtain decoded data of at least two video segments; extracting features from the decoded data to obtain multimodal features of the at least two video segments; determining a narrative intent corresponding to the at least two video segments based on the multimodal features; determining a transition effect adapted between the at least two video segments based on the narrative intent; and splicing the at least two video segments based on the transition effect to obtain a target video containing the at least two video segments.
[0006] Secondly, this application also provides a video generation apparatus. The apparatus includes: a decoding module for decoding at least two video segments to obtain decoded data of at least two video segments; an extraction module for extracting features from the decoded data to obtain multimodal features of at least two video segments; a determination module for determining the narrative intent corresponding to the at least two video segments based on the multimodal features; and determining a transition effect adapted between the at least two video segments based on the narrative intent; and a processing module for splicing the at least two video segments based on the transition effect to obtain a target video containing the at least two video segments.
[0007] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and the processor, when executing the computer program, performs the following steps: decoding at least two video segments to obtain decoded data for at least two video segments; extracting features from the decoded data to obtain multimodal features for at least two video segments; determining the narrative intent corresponding to at least two video segments based on the multimodal features; determining a transition effect adapted between at least two video segments based on the narrative intent; and splicing the at least two video segments based on the transition effect to obtain a target video containing at least two video segments.
[0008] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps: decoding at least two video segments to obtain decoded data of at least two video segments; extracting features from the decoded data to obtain multimodal features of at least two video segments; determining the narrative intent corresponding to at least two video segments based on the multimodal features; determining a transition effect adapted between at least two video segments based on the narrative intent; and splicing the at least two video segments based on the transition effect to obtain a target video containing at least two video segments.
[0009] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps: decoding at least two video segments to obtain decoded data of at least two video segments; extracting features from the decoded data to obtain multimodal features of at least two video segments; determining the narrative intent corresponding to at least two video segments based on the multimodal features; determining a transition effect adapted between at least two video segments based on the narrative intent; and splicing the at least two video segments based on the transition effect to obtain a target video containing at least two video segments.
[0010] The aforementioned video generation method, apparatus, computer equipment, storage medium, and computer program product obtain decoded data of at least two video segments by decoding at least two video segments; extract features from the decoded data to obtain multimodal features of at least two video segments; determine the narrative intent corresponding to at least two video segments based on the multimodal features; determine the transition effect adapted between at least two video segments based on the narrative intent; and splice the at least two video segments based on the transition effect to obtain a target video containing at least two video segments. Since multimodal features are obtained by extracting features from the decoded data of at least two video segments, the multimodal features of this application already contain rich multidimensional features from at least two video segments. Therefore, the narrative intent corresponding to at least two video segments determined based on multimodal features is more accurate. Consequently, the transition effect dynamically determined based on the narrative intent and adapted to at least two video segments is more in line with the "story logic" and "emotional changes" contained between video segments. This results in the final target video containing at least two video segments, after splicing the at least two video segments based on the transition effect, having a perfect fusion between the transition animation and the original screen content. This ensures visual smoothness and achieves the technical effect of effectively improving video generation efficiency while ensuring video generation quality. Attached Figure Description
[0011] Figure 1 This is a diagram illustrating the application environment of a video generation method in one embodiment;
[0012] Figure 2 This is a flowchart illustrating a video generation method in one embodiment;
[0013] Figure 3 This is a schematic diagram of a cloud video platform that supports user uploads of custom video clips in one embodiment.
[0014] Figure 4 This is a schematic diagram of the overall process of a video generation method provided in one embodiment;
[0015] Figure 5 This is a schematic diagram showing the video production page provided in one embodiment;
[0016] Figure 6 This is a schematic diagram illustrating the overall process of inferring narrative intent in one embodiment;
[0017] Figure 7 This is a schematic diagram of the adaptive video stitching and transition effect generation process in one embodiment;
[0018] Figure 8 This is a schematic diagram of the overall process of a technical solution for intelligent generation of video stitching and transition effects provided in one embodiment;
[0019] Figure 9 This is a structural block diagram of a video generation device in one embodiment;
[0020] Figure 10 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0022] It should be noted that in the following description, the terms "first, second, and third" are used only to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, and third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0023] The video generation method provided in this application embodiment can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be set up independently, integrated into server 104, or placed in the cloud or on other devices. Terminal 102 acquires at least two video clips uploaded by the user, or terminal 102 searches the database for at least two video clips corresponding to the video clip identifiers input by the user. Terminal 102 can send the acquired at least two video clips to the backend server, i.e., server 104, so that server 104 can decode the at least two video clips to obtain decoded data of at least two video clips, and perform feature extraction on the decoded data to obtain multimodal features of at least two video clips. Further, server 104 can determine the narrative intent corresponding to at least two video clips based on the multimodal features, and determine the transition effect adapted between at least two video clips based on the narrative intent, and perform splicing processing on at least two video clips based on the transition effect to obtain a target video containing at least two video clips, and return the target video containing at least two video clips to terminal 102. Terminal 102 can visualize the target video containing at least two video clips.
[0024] Terminal 102 can be a smartphone, tablet, laptop, desktop computer, smart speaker, smart TV, smartwatch, IoT device, or portable wearable device. IoT devices can include smart in-vehicle devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc.
[0025] Server 104 can be an independent physical server or a service node in a blockchain system. The service nodes in the blockchain system form a peer-to-peer (Peer To Peer) network. The Peer To Peer protocol is an application layer protocol that runs on top of the Transmission Control Protocol (TCP).
[0026] In addition, server 104 can also be a server cluster consisting of multiple physical servers, which can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.
[0027] Terminal 102 and server 104 can be connected via Bluetooth, USB (Universal Serial Bus) or network, etc., and this application does not impose any restrictions.
[0028] In one embodiment, such as Figure 2 As shown, a video generation method is provided. This method can be executed by a server or a terminal alone, or by both a server and a terminal together. This method can be applied to... Figure 1 Taking the terminal in the example, the explanation includes the following steps:
[0029] Step 202: Decode at least two video segments to obtain decoded data for at least two video segments.
[0030] In this application, video clips refer to original video segments uploaded by the user or selected by the user for splicing. The video clips in this application can be obtained in the form of video files. Furthermore, the at least two video clips in this application can originate from different shots or scenes, or they can be different segments from the same scene or shot. Even if a video shot under the same shot is divided into multiple segments, these can still be used as video clips to be spliced. For example, in the same scene, a character's transition from "silent contemplation" to "speaking" constitutes two video clips. The terminal can splice these two video clips from the same scene to obtain a target video containing at least two video clips, and simultaneously determine that the two video clips from the same scene represent an emotional transition scene, automatically generating a smooth transition effect. As another example, if a long shot is divided into two segments (e.g., one of walking, one of looking up at the sky), the terminal can still automatically recognize that the two video clips are continuous narrative segments and connect them through a slight fade-in or stable shot transition.
[0031] Decoded data refers to the data obtained after decoding the original video segment (such as a video file). For example, the decoded data in this application includes at least frame sequences, audio streams, and text information. Among them, the text information can be text attached to the original video segment (i.e., the video file), such as subtitles, dialogue, and other text information.
[0032] Specifically, the video generation method provided in this application can be widely applied to scenarios supporting the generation of personalized videos. That is, different users (operation objects) can use devices that interact with the cloud platform (or application). When a user (operation object) wants to log in to the cloud platform, they can trigger a target operation on the cloud platform (or application) (such as triggering a startup operation on the cloud platform) to open the cloud platform application (APP) on their terminal. Upon entering the cloud platform application's (e.g., a cloud video platform) page, the terminal can display the original video clip to be selected on the cloud video platform page opened by the user. For example, ... Figure 3 The image shows a schematic diagram of a cloud video platform that supports user-uploaded custom video clips. The terminal can display a video clip (i.e., the original video clip) representing a specific scene from a video on the playback page of a video opened by the user. Simultaneously, the user can trigger selection or upload operations for different video clips. For example, the user can click on... Figure 3 The control shown, used to initiate "Build Personalized Video Works," will cause the terminal to respond to the user's input, such as... Figure 3 The launch operation for the target function "Build Personalized Video Works" triggered in the playback page shown displays a subpage (such as a video production page) on the playback page for building personalized video works. For example, the terminal can... Figure 3 The playback page shown displays a floating page (i.e., a subpage) on the right, and through the subpage, it obtains at least two video clips (i.e., original video clips) selected or uploaded by the user.
[0033] In some embodiments, in response to a user-triggered video creation operation, the terminal displays a video production page for generating a target video, and obtains at least two video clips input by the user through the video production page.
[0034] In some embodiments, the terminal responds to an information input operation triggered by the user in the information input box on the video production page, and obtains the video files of at least two video segments entered by the user through the information input box; or, the terminal responds to a user's trigger operation on the shortcut menu control (i.e., the shortcut selection item bar) on the video production page, and obtains the video files of at least two video segments selected by the user through the shortcut menu control.
[0035] It is understood that the method provided in this application can be implemented through interaction between the terminal and the backend server of the cloud application (such as a cloud video platform), or through interaction between the frontend and backend of the terminal. That is, the frontend of the terminal is used to display video clips, and the backend of the terminal is equivalent to the backend server of the cloud application, used to perform logical processing such as decoding and splicing on at least two video clips uploaded by the user.
[0036] Furthermore, after acquiring at least two video clips (i.e., the original video clips) selected or uploaded by the user, the terminal can decode the video files of the at least two video clips, that is, decode the video files of each video clip into a processable frame sequence, audio stream, and text information. At the same time, for some long video clips, the terminal can further perform scene boundary detection and shot segmentation on the decoded frame sequence to obtain a grouped sequence, that is, cut the decoded frame sequence (such as a long video clip) into smaller, logically independent segments.
[0037] Furthermore, after the terminal decodes the video files of at least two video segments—that is, after decoding each video segment into processable frame sequences, audio streams, and text information—the terminal can also perform unified format processing on the decoded data, i.e., frame sequences, audio streams, and text information (such as subtitle text attached to the video file), to obtain unified format frame sequences, audio streams, and text information, preparing for subsequent multimodal feature extraction. For example, in the technical solution of this application, the terminal can uniformly convert the three types of information—frame sequences (images), audio streams (audio), and text—into a standardized feature format that can be used for model input and alignment, so as to facilitate subsequent multimodal fusion analysis.
[0038] In some embodiments, the decoded data is first decoded data or second decoded data; the terminal decodes at least two video segments to obtain frame sequences, audio streams and text information of at least two video segments, and the frame sequences, audio streams and text information constitute the first decoded data; or, the terminal decodes at least two video segments to obtain frame sequences, audio streams and text information of at least two video segments, and performs scene boundary detection on the frame sequences to obtain group sequences corresponding to the frame sequences, and the group sequences, audio streams and text information constitute the second decoded data.
[0039] In some embodiments, the decoded data includes at least frame sequences, audio streams, and text information. Before the terminal extracts features from the decoded data to obtain multimodal features corresponding to at least two video segments, the terminal can perform standardization processing on the decoded frame sequences, audio streams, and text information to obtain temporally sequenced frame sequence vectors, audio stream vectors, and text information vectors, in preparation for subsequent multimodal feature extraction.
[0040] Step 204: Extract features from the decoded data to obtain multimodal features of at least two video segments.
[0041] Among them, multimodal features refer to the features of different dimensions extracted from the video segments to be spliced. For example, the multimodal features in this application include, but are not limited to, three types of features: visual features, audio features and text features. Other types of features may also be included, without specific restrictions here.
[0042] Step 206: Based on multimodal features, determine the narrative intent corresponding to at least two video segments.
[0043] Narrative intent refers to the implicit intention between two video segments. For example, in this application, narrative intent can reflect the "story logic" and "emotional changes" inherent at the connection point (splitting point) between two video segments. Alternatively, narrative intent can include discrete tags (such as tension, relaxation, surprise, turning point, flashback, etc.) or continuous vectors to describe the emotion, rhythm, or plot state at the connection point between video segments. This automated and quantitative inference of "narrative intent" is the cornerstone for achieving subsequent adaptive and intelligent transition effects, completely changing the passive and rigid splicing method of traditional approaches and giving machines a "narrative insight" closer to that of human editors. It is understood that narrative intent in this application includes at least emotional tags and continuous narrative vectors.
[0044] Specifically, after decoding at least two video segments and obtaining the decoded data of at least two video segments, the terminal can perform feature extraction on the decoded data to obtain multimodal features of at least two video segments. That is, the terminal can perform feature extraction on the decoded data of each video segment in multiple dimensions. For example, the terminal can perform feature extraction on the frame sequence in the decoded data to obtain the visual features of at least two video segments; at the same time, it can perform feature extraction on the audio stream in the decoded data to obtain the audio features of at least two video segments; and it can perform feature extraction on the text information in the decoded data to obtain the text features of at least two video segments.
[0045] In some embodiments, the terminal can standardize the frame sequence, audio stream, and text information in the decoded data to obtain time-seriesd frame sequence vectors, audio stream vectors, and text information vectors, and extract features from the image content of the frame sequence vectors to obtain first-dimensional visual features of at least two video segments; simultaneously, it can extract features from the lens characteristics of the frame sequence vectors to obtain second-dimensional visual features of at least two video segments, and extract features from the visual emotion of the frame sequence vectors to obtain third-dimensional visual features of at least two video segments, and determine the visual features of at least two video segments based on the first-dimensional visual features, second-dimensional visual features, and third-dimensional visual features.
[0046] In some embodiments, the terminal can standardize the frame sequence, audio stream, and text information in the decoded data to obtain time-sequential frame sequence vectors, audio stream vectors, and text information vectors. It can also extract features from the speech dialogue content of the audio stream vectors to obtain first-dimensional audio features of at least two video segments. Simultaneously, it can extract features from the background sound effects of the audio stream vectors to obtain second-dimensional audio features of at least two video segments, and extract features from the audio emotion of the audio stream vectors to obtain third-dimensional audio features of at least two video segments. Based on the first-dimensional audio features, second-dimensional audio features, and third-dimensional audio features, it can determine the audio features of at least two video segments.
[0047] Furthermore, the terminal extracts features from the decoded data to obtain multimodal features of the two video segments, including at least visual, audio, and text features. The terminal can then fuse these multimodal features to obtain fused features, and perform sequence analysis on the fused features in the time dimension to obtain temporal context information. The terminal can then use a specially trained intent inference model to map the temporal context information to obtain the narrative intent corresponding to at least two video segments.
[0048] In some embodiments, the terminal can fuse multimodal features, namely visual features, audio features and text features, through a fusion model to obtain fused features; wherein, the fusion model is trained based on a multimodal video dataset, and the multimodal data in the multimodal video dataset carries visual labels, audio labels and text feature labels.
[0049] In some embodiments, the narrative intent in this application is at least one of discrete labels or continuous vectors; the temporal context information includes temporal features, and the terminal can map the temporal features through a first intent inference model to obtain discrete labels or continuous vectors; or, it can map the temporal features through a second intent inference model to obtain discrete labels and continuous vectors.
[0050] In some embodiments, the terminal can also use the dual-branch structure in the second intent inference model to map the temporal features separately, thereby obtaining the emotion label and the continuous narrative vector simultaneously; wherein, the dual-branch structure includes a narrative rhythm prediction branch and an emotion classification branch.
[0051] Step 208: Based on the narrative intent, determine the transition effects that are suitable for at least two video segments.
[0052] The transition effect refers to one or more transition effects selected from the transition effect library. The transition effects in this application can include various types of transition effects. For example, visual transition effects include, but are not limited to: fade-in / fade-out, dissolve, wipe, flash, blur, scale, rotate, dissolve, etc. Motion transition effects include, but are not limited to: zoom, pan, oscillation, camera follow, etc. Time-based transition effects include, but are not limited to: fast cut, time delay, fade-out, flashback, etc. It is understood that the transition effect types in this application can also include other custom-defined transition effects.
[0053] Step 210: Based on the transition effect, at least two video segments are spliced together to obtain a target video containing at least two video segments.
[0054] Among them, splicing processing refers to the process of splicing video segments to be spliced. For example, the splicing processing in this application includes content-aware splicing processing, that is, the methods of splicing at least two video segments in this application include: performing high-precision geometric alignment and content fusion processing, performing temporal completion processing or intelligent erasure processing on dynamic targets (to avoid ghosting or tearing), and performing dynamic and fine adjustment processing on various parameters of transition effects, etc.
[0055] A target video refers to a video containing at least two spliced video clips and transition effects. In some cases, the target video in this application can also be called a video work (a personalized creative video). For example, the target video in this application can be a derivative video of a short drama, that is, a target video obtained by splicing video clips from different short dramas according to a personalized creative idea.
[0056] Specifically, such as Figure 4 The diagram shown illustrates the overall process of the video generation method provided in this application. After the terminal determines the narrative intent corresponding to at least two video segments based on multimodal features, as shown... Figure 4 The processing flow shown allows the terminal to search for suitable transition effect types between at least two video segments from a transition effect library based on narrative intent and transition strategies. It then adaptively adjusts the transition effect parameters based on the visual and audio features of the at least two video segments and the transition effect type, resulting in an updated transition effect – thus achieving dynamic optimization of the transition effect parameters. Furthermore, the terminal can perform splicing processing on at least two video segments using content-aware splicing methods and transition effects to obtain a target video containing at least two video segments. For example, the terminal can perform geometric alignment and content fusion processing on at least two video segments to obtain a spliced video, and then render the spliced video based on the updated transition effect to obtain a target video containing transition effects and at least two video segments – a secondary creative video work.
[0057] For example, suppose the terminal infers the narrative intent output by the model based on the narrative intent, such as: "emotional shift from tension to relaxation," "rapid scene change," or "the passage of time." The terminal maintains a rich library of transition effects. Each transition effect in the library (such as fade-in / fade-out, dissolve, wipe, flash, blur, shimmer, etc.) is pre-labeled, associated with a specific narrative intent type, emotional intensity, or scene transition mode. The terminal can automatically and intelligently match and select the most appropriate transition effect type from the transition effect library based on the "narrative intent" output by the model. This adaptive matching of transition effect types based on "narrative intent" used in this application is no longer a random or fixed selection based on manual rules, but rather a matching based on a deep understanding of the "narrative intent" of the video content. That is, the computer device can, like an experienced editor, select the transition effects that best enhance the atmosphere and advance the narrative according to the needs of the story segment. For example, if the terminal detects that the "narrative intent" output by the model is "sudden fright or event outbreak", the terminal will adaptively select a fast "hard cut" or "flash white" transition effect type; while for the narrative intent of "a memory or longing", the terminal may adaptively select a slow "dissolve" or "soft blur" transition effect type.
[0058] Furthermore, after the terminal adaptively determines the transition effect type of the video segments to be spliced, it can dynamically and finely adjust various parameters of that transition effect. These parameters include, but are not limited to, transition duration, transition area, animation direction, blur intensity, color change, and volume gradient curve. The adjustment of these parameters takes into account subtlety and the visual and audio characteristics of the two video segments to be spliced (e.g., Clip A and Clip B). The key to this adaptive dynamic parameter adjustment approach in this application lies in the "dynamic adaptability" of the parameters. For the same "fade-in / fade-out" transition effect, if the narrative intent (I_narrative) is a "gentle scene transition," the duration may be longer and the change smoother; if the narrative intent (I_narrative) is a "rapidly advancing plot segment," the duration will be shorter. For example, regarding the "wipe" transition effect, the terminal can intelligently determine the direction of the wipe (from left to right, from top to bottom, etc.) based on the movement direction of the main elements or the composition of the images to be spliced, ensuring a perfect blend between the transition animation and the screen content, rather than a rigid application. Ultimately, the terminal automatically performs high-precision geometric alignment and content fusion on the two video clips to be spliced (Clip A and Clip B), and renders the final transition effect based on the adjusted parameters, resulting in a target video containing at least two video clips (Clip A and Clip B) and an adaptive transition effect.
[0059] In this embodiment, at least two video segments are decoded to obtain decoded data for at least two video segments; feature extraction is performed on the decoded data to obtain multimodal features for at least two video segments; based on the multimodal features, the narrative intent corresponding to at least two video segments is determined, and based on the narrative intent, a transition effect suitable for the at least two video segments is determined; based on the transition effect, at least two video segments are spliced together to obtain a target video containing at least two video segments. Since multimodal features are obtained by extracting features from the decoded data of at least two video segments, the multimodal features of this application already contain rich multidimensional features from at least two video segments. Therefore, the narrative intent corresponding to at least two video segments determined based on multimodal features is more accurate. Consequently, the transition effect dynamically determined based on the narrative intent and adapted to at least two video segments is more in line with the "story logic" and "emotional changes" contained between video segments. This results in the final target video containing at least two video segments, after splicing the at least two video segments based on the transition effect, having a perfect fusion between the transition animation and the original screen content. This ensures visual smoothness and achieves the technical effect of effectively improving video generation efficiency while ensuring video generation quality.
[0060] In one embodiment, before decoding at least two video segments to obtain decoded data for at least two video segments, the method further includes:
[0061] In response to a triggered video creation action, a video production page for generating the target video is displayed;
[0062] The video creation page retrieves at least two video clips input by the user.
[0063] Among them, the video creation operation refers to the operation used to open the video production page. For example, the video creation operation in this application can be the operation of clicking the "video creation" control.
[0064] The video creation page can be a standalone page or a subpage of the currently displayed video playback page. Subpages can take various forms, such as overlay pages or pop-ups. For example... Figure 5 The image shown is a schematic diagram of the video production page provided in this application. The video production page displayed on the video playback page in this application, used to generate the target video, can be as follows: Figure 5 The overlay page on the right, as shown, displays the video content on the video playback page while simultaneously displaying the content from the subpage, i.e., the video production page.
[0065] Specifically, in the user-opened, such as Figure 5In the video page shown, while video A is playing on the terminal, the user can browse, for example... Figure 5 The interactive elements on the video page shown are such that, assuming the user clicks on them... Figure 5 The control shown, used to initiate "Build Personalized Video Works," will cause the terminal to respond to the user's input, such as... Figure 5 The video page shown triggers a launch operation or video creation operation targeting the desired function (i.e., the "Build Personalized Video Works" function), as in... Figure 5 The target location on the video page (e.g., the right side of the page) displays a video production page (a floating page, i.e., a subpage) for generating the target video. Through this video production page, at least two video clips input by the user are obtained. This allows the method provided in this application to run on any original video playback page and provides a subpage, i.e., the video production page, as the entry point for obtaining the user-inputted video clips to be processed. This provides users with a better interactive experience and a more automated experience of generating personalized video works. In other words, it allows users to remain unaware of the underlying business processing logic, achieving the technical effect of effectively improving video generation efficiency and quality while ensuring a good user experience.
[0066] In one embodiment, the step of obtaining at least two video clips input by the user through the video creation page includes:
[0067] In response to an input action triggered in the input box on the video production page, retrieve the video files containing at least two video clips entered by the user through the input box; or,
[0068] In response to a trigger action on the shortcut menu control in the video production page, retrieve the video files of at least two video clips entered by the user through the shortcut menu control.
[0069] Here, a shortcut menu control refers to an interactive control on a webpage. For example, the shortcut menu control in this application could be as follows: Figure 5 The controls on the video production page shown on the right are for quickly selecting video segments to be spliced, including but not limited to shortcut menu controls such as "Select Quick Item", "Select Local Video", and "Select Custom Segment".
[0070] Specifically, the terminal responds to the user in such... Figure 5 The launch operation triggered on the video page shown targets the "Video Creation" function. Figure 5 After the video editing page (overlay page) on the right side of the video page shown in the image is displayed, the user can... Figure 5In the "Information Input Box" at the bottom of the overlay page on the right (i.e., the video production page) shown in the image, if the user enters a query in natural language format, such as "Query video clip 1 and video clip 2", the terminal will respond to the user's query. Figure 5 The information input operation triggered in the information input box of the subpage shown retrieves at least two video clips entered by the user through the information input box, namely the video file of video clip 1 and the video file of video clip 2; or, the user can also click as shown Figure 5 The "shortcut menu control" at the bottom of the overlay page on the right (i.e., the video file) shown in the image indicates that the terminal responds to the user's selection of options such as... Figure 5 The shortcut menu control shown in the video production page, such as the "Select Custom Clip" trigger operation, retrieves video files containing at least two custom video clips entered by the user through the shortcut menu control.
[0071] Alternatively, users can click on, such as Figure 5 The "voice control" at the bottom of the overlay page on the right (i.e., the video file) shown in the image captures the user's natural language query statement, such as "query video clip 1 and video clip 2," in real time. The terminal then responds to the user's query. Figure 5 The voice control in the subpage shown is triggered to obtain at least two video clips input by the user via the voice control: video clip 1 corresponds to video file 1, and video clip 2 corresponds to video file 2. This allows the method provided in this application to run on any original video playback page and provides a subpage, i.e., a video production page, as the entry point for obtaining the user-input video clips to be processed. This provides users with a better interactive experience and a more automated experience of generating personalized video works. In other words, it allows users to remain unaware of the underlying business processing logic, achieving the technical effect of effectively improving video generation efficiency and quality while ensuring a good user experience.
[0072] In one embodiment, the decoded data is first decoded data or second decoded data; the step of decoding at least two video segments to obtain decoded data for at least two video segments includes:
[0073] Decode at least two video segments to obtain frame sequences, audio streams, and text information for at least two video segments; the frame sequences, audio streams, and text information constitute the first decoded data; or,
[0074] Decode at least two video segments to obtain frame sequences, audio streams, and text information for at least two video segments; perform scene boundary detection on the frame sequences to obtain the corresponding group sequences; the group sequences, audio streams, and text information constitute the second decoded data.
[0075] Specifically, assuming the user-uploaded video clips to be spliced obtained by the terminal are Clip A and Clip B, the terminal can decode at least two video clips, namely Clip A and Clip B, respectively, to obtain the frame sequences, audio streams, and text information of the two video clips, namely Clip A and Clip B. That is, the frame sequence, audio stream, and text information decoded from Clip A, and the frame sequence, audio stream, and text information decoded from Clip B together constitute the first decoded data. Alternatively, the terminal can decode at least two video clips, namely Clip A and Clip B, respectively, to obtain the frame sequences, audio streams, and text information of the two video clips, namely Clip A and Clip B, and further perform scene boundary detection on the decoded frame sequences of Clip A and Clip B to obtain the grouped sequences corresponding to the frame sequences of Clip A and Clip B, and combine the grouped sequences of Clip A and Clip B, audio streams, and text information to form the second decoded data. In this embodiment of the application, the main function of scene boundary detection and shot segmentation is to identify content change points in video segments and automatically divide long videos into logically independent, finer-grained segments (i.e., grouped sequences), which facilitates subsequent narrative analysis and transition effect generation.
[0076] For example, when a complete video switches from a dialogue scene to a landscape shot, the terminal can identify these two different scenes through shot segmentation or scene boundary detection, and then use different strategies when splicing or adding transition effects. If the video clip uploaded by the user is already short or structurally clear, this step can be skipped. That is, the terminal does not need to perform scene boundary detection or shot segmentation on the decoded frame sequence and can directly proceed to the subsequent multimodal feature extraction and narrative analysis processing flow. Therefore, by performing unified preprocessing on different video clips uploaded by the user—including decoding, shot segmentation, and scene boundary detection—more accurate multimodal data can be provided for subsequent multimodal feature extraction. This enables narrative intent inference driven by multimodal feature collaboration, thereby improving the accuracy of narrative intent inference.
[0077] In one embodiment, the decoded data includes at least a frame sequence, an audio stream, and text information; before performing feature extraction on the decoded data to obtain multimodal features corresponding to at least two video segments, the method further includes:
[0078] The frame sequence, audio stream, and text information are standardized to obtain time-series frame sequence vectors, audio stream vectors, and text information vectors.
[0079] The step of extracting features from the decoded data to obtain multimodal features of at least two video segments includes:
[0080] Feature extraction is performed on the frame sequence vectors to obtain the visual features of at least two video segments;
[0081] Feature extraction is performed on the audio stream vector to obtain audio features of at least two video segments;
[0082] Feature extraction is performed on the text information vector to obtain text features for at least two video segments.
[0083] Standardization processing refers to the process of unifying the format of multimodal decoded data from different video segments; it can also be called data standardization. For example, it involves processing the image, audio, and text data contained in the decoded data into a time-series vector representation, which facilitates feature alignment and fusion analysis in subsequent steps.
[0084] Specifically, assuming the user-uploaded video clips to be spliced, Clip A and Clip B, are obtained by the terminal, the terminal can decode at least two video clips, Clip A and Clip B, respectively, to obtain the frame sequences, audio streams, and text information of the two video clips, Clip A and Clip B. The terminal can then perform standardization processing on the decoded data, that is, it can standardize the decoded frame sequences, audio streams, and text information to obtain time-series-based frame sequence vectors, audio stream vectors, and text information vectors. Furthermore, the terminal can perform feature extraction on the standardized frame sequence vectors to obtain the visual features of at least two video clips, and perform feature extraction on the standardized audio stream vectors to obtain the audio features of at least two video clips, and perform feature extraction on the standardized text information vectors to obtain the text features of at least two video clips.
[0085] For example, the terminal can invoke three pre-trained specialized feature extraction models. Specifically, the terminal can use visual feature extraction model A to extract features from the standardized frame sequence vectors to obtain visual features (F_v) for at least two video segments; simultaneously, it can use audio feature extraction model B to extract features from the standardized audio stream vectors to obtain audio features (F_a) for at least two video segments; and it can use text feature extraction model C to extract features from the standardized text information vectors to obtain text features (F_t) for at least two video segments. This allows for the extraction of multi-dimensional, multimodal features, and the deep, interactive fusion of these three modalities—visual, audio, and text features—allowing them to "dialogue" and "verify" with each other, thus inferring a more accurate narrative intent between video segments.
[0086] In one embodiment, the frame sequence vector includes an image matrix vector, the audio stream vector includes an acoustic feature vector, and the text information vector includes a semantic vector; the step of standardizing the frame sequence, audio stream, and text information to obtain temporally sequenced frame sequence vectors, audio stream vectors, and text information vectors includes:
[0087] Extract keyframes from the frame sequence and convert the keyframes into image matrix vectors;
[0088] Based on a preset sampling rate, the audio stream is converted into an acoustic feature vector;
[0089] Convert text information into semantic vectors.
[0090] Specifically, taking two video clips uploaded by the user, Clip A and Clip B, as an example, the terminal can decode at least two video clips, Clip A and Clip B, respectively, obtaining their frame sequences, audio streams, and text information. The terminal can then standardize the decoded data, producing time-series frame sequence vectors, audio stream vectors, and text information vectors. In other words, the terminal automatically converts the image, audio, and text information contained in the decoded data into a standardized feature format suitable for model input and alignment, facilitating subsequent multimodal fusion analysis. Specifically, the terminal can extract keyframes from the frame sequence and convert them into image matrix vectors, such as extracting keyframes and standardizing them into image matrices at a fixed resolution (e.g., 224×224). Furthermore, the terminal can convert the audio stream into an acoustic feature vector based on a preset sampling rate; for example, the terminal can convert the decoded audio stream data into a Mel-frequency spectrum or acoustic feature vector at a standard sampling rate (e.g., 16kHz). Simultaneously, the terminal can convert the text information contained in the decoded data into semantic vectors. For example, it can convert text information into semantic vectors through word segmentation and encoding (such as BERT or Word2Vec). Ultimately, all three types of data contained in the decoded data will be uniformly processed into temporal vector representations, facilitating feature alignment and fusion analysis in subsequent steps to extract more accurate multimodal features. The extracted multimodal features, namely visual features, audio features, and text features, will then be deeply and interactively fused, allowing these three types of information to "dialogue" and "verify" with each other, thereby enabling a more accurate inference of the narrative intent between video segments.
[0091] In one embodiment, the step of extracting features from frame sequence vectors to obtain visual features of at least two video segments includes:
[0092] Feature extraction is performed on the image content of the frame sequence vectors to obtain the first-dimensional visual features of at least two video segments;
[0093] Feature extraction is performed on the lens characteristics of the frame sequence vectors to obtain the second-dimensional visual features of at least two video segments;
[0094] Visual emotion features are extracted from frame sequence vectors to obtain third-dimensional visual features of at least two video segments;
[0095] Based on the first-dimensional visual features, the second-dimensional visual features, and the third-dimensional visual features, determine the visual features of at least two video segments.
[0096] In this application, the first-dimensional visual features, second-dimensional visual features, and third-dimensional visual features are only used to distinguish visual features of different dimensions. For example, the first-dimensional visual features in this application can be visual features that reflect the content of an image, such as visual features related to the image content, such as objects (people, scene props), scene types (indoor, outdoor, street scene), and people's actions (running, jumping, talking) in a video frame.
[0097] The second dimension of visual features can be visual features that reflect the characteristics of a shot, such as analyzing the shot's framing (long shot, close-up), camera movement (push, pull, pan, tilt), composition, color style, and other shot-related visual features.
[0098] The third dimension of visual features can be visual features that reflect visual emotions, such as judging the emotional tendency of the visual based on the light and shadow, color and other factors in the picture (bright, dark, tense, warm, etc.).
[0099] Specifically, taking two video clips, Clip A and Clip B, uploaded by the user for splicing as an example, the terminal decodes the video files of at least two video clips, Clip A and Clip B, respectively, obtaining their frame sequences, audio streams, and text information. The terminal then standardizes the decoded data, obtaining time-sequential frame sequence vectors, audio stream vectors, and text information vectors. Further, the terminal performs feature extraction on the standardized frame sequence vectors. For example, it can extract features from the image content of the frame sequence vectors to obtain the first-dimensional visual features of the at least two video clips; simultaneously, it can extract features from the shot characteristics of the frame sequence vectors to obtain the second-dimensional visual features of the at least two video clips; and it can extract features from the visual emotion of the frame sequence vectors to obtain the third-dimensional visual features of the at least two video clips. Based on the first, second, and third-dimensional visual features, the visual characteristics of the at least two video clips are determined. For example, the terminal can fuse the extracted three-dimensional visual features, that is, concatenate the first-dimensional, second-dimensional, and third-dimensional visual features to obtain the concatenated visual feature F_v. This concatenated visual feature F_v is the visual feature of video clips ClipA and ClipB. Alternatively, the terminal can also fuse the first-dimensional, second-dimensional, and third-dimensional visual features using a fusion model to obtain the fused visual feature F_v. The fused visual feature F_v is also the visual feature of video clips ClipA and ClipB. This allows for the extraction of multi-dimensional visual features and the deep, interactive fusion of these different dimensions, enabling these three dimensions of visual information to "dialogue" and "verify" with each other, thus inferring more accurate visual features between video clips and providing richer visual features for subsequent multimodal feature fusion.
[0100] In one embodiment, the step of extracting features from the audio stream vector to obtain audio features of at least two video segments includes:
[0101] Feature extraction is performed on the speech dialogue content of the audio stream vector to obtain the first-dimensional audio features of at least two video segments;
[0102] Extract background sound effects features from the audio stream vector to obtain second-dimensional audio features from at least two video segments;
[0103] Audio emotion features are extracted from the audio stream vectors to obtain the third-dimensional audio features of at least two video segments;
[0104] Based on the first-dimensional audio features, the second-dimensional audio features, and the third-dimensional audio features, determine the audio features of at least two video segments.
[0105] In this application, the first-dimensional audio features, second-dimensional audio features, and third-dimensional audio features are only used to distinguish audio features of different dimensions. For example, the first-dimensional audio features in this application can be audio features that reflect speech content. For example, if there is dialogue in a video clip, speech recognition can be performed automatically to extract text content.
[0106] The second dimension of audio features can be audio features that reflect sound effects or background music. For example, analyzing the type, rhythm, melody and emotion of music in a video clip, or identifying environmental sound effects and their semantics in a video clip.
[0107] The third dimension of audio features can be audio features used to express audio emotions, such as judging auditory emotional tendencies based on sound characteristics.
[0108] Specifically, let's take two video clips, Clip A and Clip B, uploaded by the user as an example. The terminal decodes the video files of at least two video clips, Clip A and Clip B, respectively, obtaining their frame sequences, audio streams, and text information. Then, the terminal can standardize these decoded components to obtain temporally sequenced frame sequence vectors, audio stream vectors, and text information vectors. Further, the terminal can extract features from the standardized audio stream vectors. For example, it can extract features from the speech dialogue content of the audio stream vectors to obtain the first-dimensional audio features of the at least two video clips; simultaneously, it can extract features from the background sound effects of the audio stream vectors to obtain the second-dimensional audio features of the at least two video clips; and it can extract features from the emotional content of the audio stream vectors to obtain the third-dimensional audio features of the at least two video clips. Based on the first, second, and third-dimensional audio features, the terminal determines the audio features of the at least two video clips. For example, the terminal can fuse the extracted audio features from the three dimensions, that is, concatenate the first, second, and third-dimensional audio features to obtain the concatenated audio feature F_a. This concatenated audio feature F_a is the audio feature of video clips Clip A and Clip B. Alternatively, the terminal can also fuse the first, second, and third-dimensional audio features using a fusion model to obtain the fused audio feature F_a, which is also the audio feature of video clips Clip A and Clip B. This allows for the extraction of multi-dimensional audio features and the deep, interactive fusion of these different dimensions, enabling the audio information from these three dimensions to "dialogue" and "verify" with each other, thus inferring more accurate audio features between video clips and providing richer audio features for subsequent multimodal feature fusion.
[0109] In one embodiment, the multimodal features include visual features, audio features, and text features; the step of determining the narrative intent corresponding to at least two video segments based on the multimodal features includes:
[0110] Visual features, audio features, and text features are fused to obtain fused features;
[0111] Sequence analysis of the fused features in the time dimension is performed to obtain temporal context information;
[0112] By mapping the temporal context information, the narrative intent corresponding to at least two video segments can be obtained.
[0113] Specifically, let's take two video clips uploaded by the user, Clip A and Clip B, as an example for illustration. Figure 6 The diagram illustrates the overall process of inferring narrative intent. The terminal extracts features from the decoded data to obtain multimodal features of the two video clips, Clip A and Clip B, including visual features (F_v), audio features (F_a), and text features (F_t). Then, as shown... Figure 6 As shown, the terminal can fuse visual, audio, and textual features to obtain a fused feature F_context. It then performs sequence analysis on this fused feature F_context over time to obtain temporal context information. For example, the terminal can input the fused features F_context of consecutive frames (or segments) into a Transformer encoder in chronological order, using a self-attention mechanism to calculate the dependencies between each moment, thereby identifying temporal features such as emotional changes and plot twists (i.e., temporal context information). Furthermore, the terminal can invoke a pre-trained intent inference model and map the temporal context information through this model to obtain the narrative intent corresponding to at least two video segments. This allows for more granular and accurate identification of advanced narrative patterns such as "emotions shifting from calm to escalating," "pacing changing from slow to fast," or "plot development shifting from setup to twist" by analyzing continuous visual, audio, and dialogue context information. This ability is fundamental to understanding "narrative intent" because it allows for a clearer understanding of the story's trajectory, providing more granular data support for subsequent, more accurate inferences of narrative intent.
[0114] In one embodiment, the step of fusing visual features, audio features, and text features to obtain fused features includes:
[0115] The fusion model integrates visual features, audio features, and text features to obtain fused features. The fusion model is trained on a multimodal video dataset, which contains visual, audio, and text feature labels.
[0116] Specifically, taking two video clips, Clip A and Clip B, uploaded by the user for splicing as an example, the terminal extracts features from the decoded data to obtain the multimodal features of the two video clips, Clip A and Clip B, including visual features (F_v), audio features (F_a), and text features (F_t). The terminal then uses a fusion model to fuse these visual, audio, and text features to obtain a fused feature F_context. In this application, the fusion model is trained on a multimodal video dataset, which contains visual, audio, and text feature labels. Specifically, during the pre-training of the fusion model, the training samples used in the training phase include a large number of labeled multimodal video datasets containing annotations for visual, audio, and text features, such as sentiment tags and scene transitions. Furthermore, the training method used in this application includes end-to-end training, optimizing the fusion model through interactive learning with multimodal data, enabling it to accurately understand and infer narrative intent in practical applications. Therefore, compared to the simple feature stacking approach in traditional methods, this embodiment strengthens the relationship between different modalities through an attention mechanism, deeply and interactively fusing the extracted visual, audio, and textual features. This is not merely a simple stacking of features, but rather a carefully designed fusion network (e.g., a Transformer structure based on an attention mechanism) that allows these three types of information to "dialogue" and "verify" each other, ensuring that the information is complementary rather than redundant, thereby effectively improving the accuracy of the inferred narrative intent.
[0117] In one embodiment, the temporal context information includes temporal features; the step of performing sequence analysis on the fused features in the time dimension to obtain the temporal context information includes:
[0118] The encoder determines the dependencies between fused features at different time points based on a self-attention mechanism, thus obtaining temporal features.
[0119] Specifically, the terminal can input the fused features F_context of consecutive frames (or segments) into the Transformer encoder in chronological order. The Transformer encoder uses a self-attention mechanism to calculate the dependencies between different moments, thereby identifying temporal features such as emotional changes and plot twists. Therefore, compared to traditional RNN or LSTM models, this application employs a multi-head attention and positional encoding optimized structure, enhancing the modeling ability for long-term dependencies and complex narrative rhythms. This allows for a more accurate understanding of the overall story flow of the video through the processing method provided in this embodiment, ensuring complementary rather than redundant information, and effectively improving the accuracy of the inferred narrative intent.
[0120] In one embodiment, the narrative intent is at least one of discrete labels or continuous vectors; the temporal context information includes temporal features; the step of mapping the temporal context information to obtain the narrative intent corresponding to at least two video segments includes:
[0121] The temporal features are mapped using the first intent inference model to obtain discrete labels or continuous vectors; or...
[0122] The temporal features are mapped using a second intent inference model to obtain discrete labels and continuous vectors.
[0123] In this application, the first intent inference model and the second intent inference model are only used to distinguish different intent inference models. For example, the second intent inference model in this application can be a dual-path model, and the first intent inference model can be a non-dual-path model.
[0124] Specifically, taking two video clips uploaded by the user, Clip A and Clip B, as an example, the terminal extracts features from the decoded data to obtain the multimodal features of the two video clips, including visual features (F_v), audio features (F_a), and text features (F_t). The terminal then fuses these features to obtain a fused feature F_context. This fused feature F_context is then subjected to sequence analysis in the time dimension to obtain temporal context information. The terminal can then call a pre-trained first intent inference model and a second intent inference model. The first intent inference model is used to map the temporal features to obtain discrete labels or continuous vectors; alternatively, the terminal can simultaneously use the second intent inference model to map the temporal features to obtain both discrete labels and continuous vectors. This allows, after obtaining rich temporal contextual information, a specialized intent inference model to map it into specific, actionable "narrative intents." These intents can be discrete labels (e.g., tension, relaxation, surprise, turning point, flashback) or continuous vector representations to describe the emotions, rhythm, or plot states at the connection points of video segments. This representation of "narrative intents" transcends surface-level content recognition, truly understanding the "story logic" and "emotional changes" inherent between video segments. This automated and quantitative inference of "narrative intents" is the cornerstone for achieving subsequent adaptive and intelligent transition effects. It completely changes the passive and rigid splicing method of traditional approaches, giving machines a "narrative insight" closer to that of human editors. This achieves the technical effect of effectively improving the efficiency of target video generation while ensuring video generation quality.
[0125] In one embodiment, the step of mapping temporal features using a second intent inference model to obtain discrete labels and continuous vectors includes:
[0126] By using the dual-branch structure in the second intent inference model, the temporal features are mapped separately to obtain discrete labels and continuous vectors; the discrete labels include sentiment labels, and the continuous vectors include continuous narrative vectors.
[0127] The dual-branch structure includes a narrative rhythm prediction branch and a sentiment classification branch.
[0128] Specifically, taking two video clips uploaded by the user, Clip A and Clip B, as an example, the terminal extracts features from the decoded data to obtain multimodal features of the two video clips, including visual features (F_v), audio features (F_a), and text features (F_t). The terminal then fuses these features to obtain a fused feature F_context. This fused feature F_context is then subjected to sequence analysis in the time dimension to obtain temporal context information. The terminal can then call a pre-trained second intent inference model. This second intent inference model employs a multi-layer Transformer combined with a dual-branch structure of sentiment classification and rhythm prediction, simultaneously outputting discrete labels and continuous vectors. Discrete labels include sentiment labels, and continuous vectors include continuous narrative vectors, enhancing the richness of expression. In other words, the terminal maps the temporal features through the dual-branch structure of the second intent inference model, specifically by mapping the narrative rhythm prediction branch and the sentiment classification branch, thus simultaneously outputting sentiment labels and continuous narrative vectors.
[0129] It is understood that the intent inference model in this application, which employs a dual-branch structure combining a multi-layer Transformer with sentiment classification and rhythm prediction, uses training samples during the training phase including multimodal video datasets (containing visuals, audio, and subtitles) labeled with narrative sentiment, rhythm changes, and plot twists. Furthermore, the model training method in this application can employ multi-task joint training (simultaneously optimizing classification and regression objectives), and enhance the model's generalization ability across different video genres through transfer learning and contrastive learning, thereby effectively improving the accuracy of the sentiment labels and continuous narrative vectors output by the intent inference model.
[0130] In one embodiment, narrative intent includes emotion tags and continuous narrative vectors; based on the narrative intent, the step of determining a transition effect suitable for at least two video segments includes:
[0131] Based on emotion tags and continuous narrative vectors, search the transition effect library for transition effect types that are suitable between at least two video segments;
[0132] Based on the visual features, audio features, and transition effect types of at least two video clips, the transition effect parameters are adjusted to obtain the updated transition effect.
[0133] The transition effect parameters refer to the various parameters corresponding to the transition effect type. For example, the transition effect parameters in this application include, but are not limited to, parameters such as transition duration, transition area, animation direction, blur intensity, color change, and volume gradient curve. The adjustment of these parameters will take into account the subtlety of the transition and the visual, audio, and textual features of the two video clips (Clip A and Clip B) to be spliced.
[0134] Specifically, let's take two video clips, Clip A and Clip B, uploaded by the user as an example. Based on the multimodal features of Clip A and Clip B, the terminal determines the corresponding narrative intent (I_narrative, including emotion tags and continuous narrative vectors) for both clips. Then, based on the emotion tags and continuous narrative vectors in the narrative intent I_narrative, the terminal searches the transition effect library for a suitable transition effect type between at least two video clips. Based on the transition effect type, the visual features of the two video clips, and the audio features, the terminal adaptively adjusts the transition effect parameters to obtain the updated transition effect. In other words, the terminal automatically matches the transition effect type that best expresses the emotion or rhythm based on the narrative intent (such as emotion tags like "turning point," "tension," "warmth," and "memories"), making the spliced target video more natural and narrative-driven.
[0135] In this embodiment, the innovation lies in the "dynamic adaptability" of the transition effect parameters. For example, for the same "fade-in / fade-out" transition effect type, if the narrative intent I_narrative is "gentle scene transition," the duration may be longer and the change smoother; if the narrative intent I_narrative is "rapidly advancing plot segment," the duration will be shorter. Furthermore, for transition effects like "wipe," the direction of the wipe (from left to right, from top to bottom, etc.) can be intelligently determined based on the movement direction of the main elements or the composition of the frames to be spliced, ensuring a perfect blend between the transition animation and the screen content, rather than a rigid application. This results in a more natural, narrative, and artistic video transition effect, thereby enhancing the overall professionalism and viewing experience of the short drama derivative works, providing users with a smoother, more immersive visual experience and a higher quality visual experience when consuming content.
[0136] In one embodiment, the step of splicing at least two video segments to obtain a target video containing at least two video segments, based on a transition effect, includes:
[0137] Perform geometric alignment and content fusion on at least two video clips to obtain a spliced video;
[0138] The spliced video is rendered based on the transition effect after the parameters are updated, resulting in a target video that includes the transition effect and at least two video segments.
[0139] Geometric alignment refers to spatially correcting the keyframes of two video clips to ensure that the position and direction of motion of the main subject in the image remain consistent.
[0140] Content fusion processing refers to the merging of video frame content from two video segments. For example, by combining semantic segmentation and edge smoothing algorithms, intelligent transitions and texture completion are performed in the splicing area of the two video segments to avoid "sponge-in marks" or "ghosting".
[0141] Specifically, such as Figure 7 The diagram illustrates the processing flow for adaptive video splicing and transition effect generation. Taking two video clips, Clip A and Clip B, uploaded by the user as an example, the terminal determines the corresponding narrative intent I_narrative (including emotion tags and continuous narrative vectors) based on the multimodal features of Clip A and Clip B, as shown below. Figure 7The processing flow shown allows the terminal to search for suitable transition effect types between at least two video segments from a transition effect library based on narrative intent and transition strategies. It then adaptively adjusts the transition effect parameters based on the visual and audio features of the at least two video segments and the transition effect type, resulting in an updated transition effect – thus achieving dynamic optimization of the transition effect parameters. Furthermore, the terminal can perform splicing processing on at least two video segments using content-aware splicing methods and transition effects to obtain a target video containing at least two video segments. For example, the terminal can perform geometric alignment and content fusion processing on at least two video segments to obtain a spliced video, and then render the spliced video based on the updated transition effect to obtain a target video containing transition effects and at least two video segments – a secondary creative video work. Therefore, compared with the traditional pixel-by-pixel stitching method, the technical solution of this application adds dynamic adaptability of parameters, semantic-level target recognition and region protection mechanism. It can automatically identify people or key object areas and optimize stitching boundaries, so that the transition animation and the picture content are perfectly integrated, rather than being applied rigidly. This improves the overall professionalism and viewing experience of the video re-creation work, allowing users to obtain a smoother, more immersive visual experience and a higher quality visual experience when consuming content.
[0142] In one embodiment, the step of performing geometric alignment and content fusion processing on at least two video segments to obtain a spliced video includes:
[0143] Based on feature point matching and optical flow estimation, spatial correction is performed on keyframes of at least two video segments to obtain aligned video segments; wherein the position and motion direction of the main subject in each video frame of the aligned video segment are consistent.
[0144] Edge transition and texture completion are performed on the areas to be spliced in the aligned video clips to obtain a spliced video with merged content.
[0145] The area to be spliced refers to the area in two video segments that needs to be spliced together. For example, the area to be spliced (or the splicing area) in this application can be the area (spliced image) formed by splicing the last frame of the previous video segment with the first frame of the next video segment.
[0146] Specifically, taking two video clips, Clip A and Clip B, uploaded by the user as an example, the terminal determines the narrative intent (I_narrative) (including emotion tags and continuous narrative vectors) corresponding to the two video clips based on the multimodal features of Clip A and Clip B. Then, based on the narrative intent I_narrative, the terminal determines a suitable transition effect between the two video clips and performs splicing processing on the two video clips based on a content-aware splicing method and the transition effect, resulting in a target video containing at least two video clips. For example, the terminal can perform spatial correction on the keyframes of the two video clips based on feature point matching and optical flow estimation to obtain aligned video clips. Then, it can perform edge transition and texture completion processing on the splicing area in the aligned video clips to obtain a spliced video with fused content. In this application, the position and direction of motion of the main subject in each video frame of the aligned video clip remain consistent. Therefore, compared with the traditional pixel-by-pixel stitching method, the technical solution of this application adds dynamic adaptability of parameters, semantic-level target recognition and region protection mechanism. It can automatically identify people or key object areas and optimize stitching boundaries, so that the transition animation and the picture content are perfectly integrated, rather than being applied rigidly. This improves the overall professionalism and viewing experience of the video re-creation work, allowing users to obtain a smoother, more immersive visual experience and a higher quality visual experience when consuming content.
[0147] In one embodiment, after splicing at least two video segments based on a transition effect to obtain a target video containing at least two video segments, the method further includes:
[0148] If a moving target appears at the stitching line position in the target video, based on video frame information before or after the appearance of the moving target, temporal completion or erasure processing is performed on the moving target to obtain the completed or erased target video; or...
[0149] Identify the key subjects in the target video and perform area protection processing on the areas where the key subjects are located to keep the areas where the key subjects are located in the target video intact.
[0150] The moving target refers to a moving target in the video frame. For example, the moving target in this application includes, but is not limited to, objects, people, and other targets.
[0151] Key subjects refer to key objects in a video frame, including people, vehicles, and other key objects. For example, the key subject in video A is person A.
[0152] Specifically, taking two video clips, Clip A and Clip B, uploaded by the user as an example, the terminal performs splicing processing on the two video clips, obtaining a target video containing both clips. The terminal can then detect whether a moving target exists at the splicing seam position in the target video. If a moving target is present at the splicing seam position, the terminal can perform temporal completion or erasure processing (i.e., repair processing) on the moving target based on video frame information before or after the appearance of the moving target, thus obtaining the completed or erased target video. Alternatively, the terminal can identify key subjects in the target video and perform region protection processing on the area where the key subjects are located to maintain the integrity of the area containing the key subjects in the target video. In other words, the technical solution of this application introduces a mechanism based on joint modeling of optical flow and target detection, which can simultaneously identify the direction and speed of object movement and dynamically adjust the splicing seam position. Simultaneously, a multi-frame temporal completion algorithm is added to intelligently restore occluded or broken targets using information from preceding and following frames, reducing "ghosting" and "tearing." Furthermore, the technical solution of this application enables a regional protection strategy for key subjects (such as people or vehicles) to avoid cutting or misalignment in the subject area, thereby achieving a more natural dynamic connection, which in turn improves the overall professionalism and viewing experience of the video re-creation work, allowing users to obtain a smoother, more immersive visual experience and a higher quality visual experience when consuming content.
[0153] In one embodiment, after splicing at least two video segments based on a transition effect to obtain a target video containing at least two video segments, the method further includes:
[0154] Identify dynamic elements in frames of the target video;
[0155] Fine-tuning dynamic elements locally yields the target video after local adjustments; or...
[0156] Smooth the seams of the target video to obtain a smoothed target video; or,
[0157] Motion compensation is applied to the seams of the target video to obtain the motion-compensated target video.
[0158] Specifically, taking two video clips uploaded by the user, Clip A and Clip B, as an example, the terminal performs splicing on the two clips to obtain a target video containing both clips. The terminal can then fine-tune and stabilize the seams of the target video to ensure high consistency between the transition and the final image, eliminating minor jitter and guaranteeing a smooth visual experience. For example, the terminal can use an adaptive stabilization algorithm based on image content to automatically identify dynamic elements in the frames of the target video and perform local fine-tuning on these elements to obtain a locally fine-tuned target video, avoiding unnecessary jitter during splicing. Simultaneously, the terminal can use multi-frame information fusion technology to smooth out minor jitter at the seams of the target video, making the transition area smoother and more natural, resulting in a smoothed target video. Alternatively, the terminal can perform motion compensation processing on the seams of the target video to obtain a motion-compensated target video. This means enhancing motion vector estimation and optimizing the smoothness of image transitions through more refined motion compensation, ensuring a more coherent visual effect after splicing, thereby improving the overall professionalism and viewing experience of video re-creation works, allowing users to obtain a smoother, more immersive visual experience and a higher quality visual experience when consuming content.
[0159] In one embodiment, a video generation method is provided. This method can be executed by a server or a terminal alone, or by both a server and a terminal. The method can be applied to... Figure 1 Taking a terminal as an example, the method includes the following steps: decoding at least two video segments to obtain decoded data of at least two video segments; extracting features from the decoded data to obtain multimodal features of at least two video segments; determining the narrative intent corresponding to at least two video segments based on the multimodal features; determining a transition effect adapted between at least two video segments based on the narrative intent; and splicing at least two video segments based on the transition effect to obtain a target video containing at least two video segments.
[0160] In one embodiment, before decoding at least two video segments to obtain decoded data for at least two video segments, the method further includes: displaying a video production page for generating a target video in response to a triggered video creation operation; and obtaining at least two video segments input by the user through the video production page.
[0161] In one embodiment, obtaining at least two video clips input by the user through the video creation page includes: obtaining video files of at least two video clips input by the user through the information input box in response to an information input operation triggered in the information input box of the video creation page; or, obtaining video files of at least two video clips input by the user through the shortcut menu control in response to a trigger operation on the shortcut menu control in the video creation page.
[0162] In one embodiment, the decoded data is either first decoded data or second decoded data; the step of decoding at least two video segments to obtain decoded data for at least two video segments includes: decoding at least two video segments to obtain frame sequences, audio streams, and text information for at least two video segments, wherein the frame sequences, audio streams, and text information constitute the first decoded data; or, decoding at least two video segments to obtain frame sequences, audio streams, and text information for at least two video segments; performing scene boundary detection on the frame sequences to obtain grouping sequences corresponding to the frame sequences, wherein the grouping sequences, audio streams, and text information constitute the second decoded data.
[0163] In one embodiment, the decoded data includes at least a frame sequence, an audio stream, and text information. Before performing feature extraction on the decoded data to obtain multimodal features corresponding to at least two video segments, the method further includes: standardizing the frame sequence, the audio stream, and the text information to obtain temporally sequenced frame sequence vectors, audio stream vectors, and text information vectors. The step of performing feature extraction on the decoded data to obtain multimodal features of at least two video segments includes: performing feature extraction on the frame sequence vector to obtain visual features of at least two video segments; performing feature extraction on the audio stream vector to obtain audio features of at least two video segments; and performing feature extraction on the text information vector to obtain text features of at least two video segments.
[0164] In one embodiment, the frame sequence vector includes an image matrix vector, the audio stream vector includes an acoustic feature vector, and the text information vector includes a semantic vector. The step of standardizing the frame sequence, audio stream, and text information to obtain time-series frame sequence vectors, audio stream vectors, and text information vectors includes: extracting keyframes from the frame sequence and converting the keyframes into image matrix vectors; converting the audio stream into acoustic feature vectors based on a preset sampling rate; and converting the text information into semantic vectors.
[0165] In one embodiment, the step of extracting features from the frame sequence vector to obtain visual features of at least two video segments includes: extracting features from the image content of the frame sequence vector to obtain first-dimensional visual features of at least two video segments; extracting features from the shot characteristics of the frame sequence vector to obtain second-dimensional visual features of at least two video segments; extracting features from the visual emotion of the frame sequence vector to obtain third-dimensional visual features of at least two video segments; and determining the visual features of at least two video segments based on the first-dimensional visual features, the second-dimensional visual features, and the third-dimensional visual features.
[0166] In one embodiment, the step of extracting features from the audio stream vector to obtain audio features of at least two video segments includes: extracting features from the speech dialogue content of the audio stream vector to obtain first-dimensional audio features of at least two video segments; extracting features from the background sound effects of the audio stream vector to obtain second-dimensional audio features of at least two video segments; extracting features from the audio emotion of the audio stream vector to obtain third-dimensional audio features of at least two video segments; and determining the audio features of at least two video segments based on the first-dimensional audio features, the second-dimensional audio features, and the third-dimensional audio features.
[0167] In one embodiment, the multimodal features include visual features, audio features, and text features; determining the narrative intent corresponding to at least two of the video segments based on the multimodal features includes: fusing the visual features, the audio features, and the text features to obtain fused features; performing sequence analysis on the fused features in the time dimension to obtain temporal context information; and mapping the temporal context information to obtain the narrative intent corresponding to at least two of the video segments.
[0168] In one embodiment, fusing the visual features, audio features, and text features to obtain fused features includes: fusing the visual features, audio features, and text features using a fusion model to obtain fused features; wherein the fusion model is trained based on a multimodal video dataset, and the multimodal data in the multimodal video dataset carries visual labels, audio labels, and text feature labels.
[0169] In one embodiment, the temporal context information includes temporal features; the step of performing sequence analysis on the fused features in the time dimension to obtain the temporal context information includes: determining the dependency relationship between the fused features at each time step by an encoder based on a self-attention mechanism to obtain the temporal features.
[0170] In one embodiment, the narrative intent is at least one of discrete labels or continuous vectors; the temporal context information includes temporal features; the mapping process of the temporal context information to obtain the narrative intents corresponding to at least two video segments includes: mapping the temporal features through a first intent inference model to obtain discrete labels or continuous vectors; or, mapping the temporal features through a second intent inference model to obtain discrete labels and continuous vectors.
[0171] In one embodiment, the step of mapping the temporal features through the second intent inference model to obtain discrete labels and continuous vectors includes: mapping the temporal features through a dual-branch structure in the second intent inference model to obtain discrete labels and continuous vectors; the discrete labels include emotion labels, and the continuous vectors include continuous narrative vectors; wherein the dual-branch structure includes a narrative rhythm prediction branch and an emotion classification branch.
[0172] In one embodiment, the narrative intent includes an emotion tag and a continuous narrative vector; determining a transition effect suitable for at least two video segments based on the narrative intent includes: searching a transition effect type suitable for at least two video segments from a transition effect library based on the emotion tag and the continuous narrative vector; adjusting transition effect parameters based on the visual features, audio features, and transition effect type of at least two video segments to obtain a transition effect with updated parameters.
[0173] In one embodiment, the step of splicing at least two video segments based on the transition effect to obtain a target video containing at least two video segments includes: performing geometric alignment and content fusion processing on at least two video segments to obtain a spliced video; and rendering the spliced video based on the transition effect after parameter updates to obtain a target video containing the transition effect and at least two video segments.
[0174] In one embodiment, the step of performing geometric alignment and content fusion processing on at least two video segments to obtain a spliced video includes: performing spatial correction on keyframes of at least two video segments based on feature point matching and optical flow estimation to obtain aligned video segments; wherein the position and motion direction of the main subject in each video frame of the aligned video segment are consistent; and performing edge transition and texture completion processing on the splicing area in the aligned video segment to obtain a spliced video after content fusion.
[0175] In one embodiment, after splicing at least two video segments based on the transition effect to obtain a target video containing at least two video segments, the method further includes: if a moving target appears at the splicing seam position in the target video, performing temporal completion or erasure processing on the moving target based on video frame information before or after the appearance of the moving target in the target video to obtain the target video after completion or erasure processing; or, identifying a key subject in the target video and performing region protection processing on the area where the key subject is located to keep the area where the key subject is located in the target video intact.
[0176] In one embodiment, after splicing at least two video segments based on the transition effect to obtain a target video containing at least two video segments, the method further includes: identifying dynamic elements in the frames of the target video; performing local fine-tuning on the dynamic elements to obtain the locally fine-tuned target video; or, smoothing the seams of the target video to obtain the smoothed target video; or, performing motion compensation on the seams of the target video to obtain the motion-compensated target video.
[0177] In one embodiment, this application also provides an application scenario in which the above-described video generation method is applied. Specifically, the video generation method is applied in this scenario as follows:
[0178] The technical solution provided in this application embodiment concerns how to connect video clips and make the connection look natural and story-like. For example... Figure 5 The playback interface of the secondary-creation video short drama shown is illustrated. This video short drama is created by the platform (backend server) through decoding at least two video segments, obtaining decoded data for at least two video segments, extracting features from the decoded data to obtain multimodal features of at least two video segments, determining the narrative intent corresponding to at least two video segments based on the multimodal features, determining appropriate transition effects between at least two video segments based on the narrative intent, and finally splicing the at least two video segments based on the transition effects. In other words, the method provided in this application allows the backend server to automatically generate appropriate transition effects based on its understanding of the story intent between video segments. Instead of simple fade-in / fade-out or hard cuts, it can choose a transition method that best matches the current video's narrative rhythm and emotion, making the connection between the two segments smooth and natural, and even enhancing the sense of story. Thus, for example, when creating secondary short dramas, the platform system can automatically provide high-quality, narrative-driven splicing effects, greatly improving the efficiency and professionalism of content creation.
[0179] This application involves the following key technical or functional aspects:
[0180] The technical solution in this application concerns how to connect video clips and make the transitions look natural and narrative. Currently, most video splicing techniques have relatively fixed transition effects or require extensive manual adjustments, making it difficult to automatically achieve an atmosphere that matches the video content.
[0181] The purpose of this technical solution is to enable the system to understand video content on its own. Through this diverse information, the system can determine the story and emotion being conveyed between video segments or within a single segment. Then, based on its understanding of this narrative intent, the system automatically generates appropriate transition effects. It no longer simply fades in or out or cuts abruptly, but rather selects a transition method that best matches the current video's narrative rhythm and emotional tone, making the connection between two segments smooth and natural, and even enhancing the sense of story. Thus, for example, when creating derivative short dramas, the system can automatically provide high-quality, narrative-driven splicing effects, greatly improving the efficiency and professionalism of content creation.
[0182] This refers to scenarios involving the secondary creation of video content, where multiple video clips need to be connected to form a complete work, i.e., splicing together different shots or scenes. To make these connections appear natural, transition effects are usually added between the two clips. Traditional technical solutions mainly fall into two categories:
[0183] The first, and most common, method is manual operation. Users need to import video clips into the software, place them on the timeline, and then, based on their own judgment and ideas, select an effect from the software's transition effects library and manually drag and drop it between two clips. The disadvantage is that it consumes a lot of time and effort and requires a high level of skill from the operator.
[0184] The second approach involves tools that offer simple, automated splicing. This method typically uses pre-set templates or rules. After the user imports video clips, the system follows these pre-set rules, such as adding the same fade-in / fade-out effect between all clips, or randomly selecting several transition effects. The problem is its lack of intelligence; it cannot select the most appropriate transition based on the specific content of the video, the context of the scene, or the implied emotional changes. The transition effects are often abrupt and generic, failing to truly match the story or emotion the video aims to convey.
[0185] The disadvantages of traditional technologies include:
[0186] First, it's inefficient and has a high barrier to entry. Whether you're a professional editor or a regular user, manually selecting and adjusting transition effects is extremely time-consuming and labor-intensive. Creating a transition that matches the atmosphere of the video content requires experience and a good aesthetic sense, which is difficult for non-professional users to achieve.
[0187] Secondly, the level of automation is low, and the effects are not intelligent enough. Some existing automated splicing tools at most select transition effects according to fixed rules or randomly. The transition effects they select are often abrupt and cannot truly match the emotional shifts or narrative rhythm.
[0188] Third, there is a lack of understanding of the deeper meaning of the videos. Existing methods are unable to comprehend the story's intent and emotions, and cannot effectively convey the director's intended emotions and narrative flow. This results in spliced videos lacking coherence and professionalism, making them unsuitable for scenarios with high narrative requirements, such as short drama adaptations.
[0189] The problems that the technical solution of this application can solve are:
[0190] Therefore, this solution proposes an intelligent video splicing and narrative system based on short drama re-creation. By analyzing various information in the video (including images, sound, and text), the system can understand the story intent and emotional direction between video segments.
[0191] Specifically, the technical solution of this application no longer simply applies fixed transitions, but instead allows the system to automatically determine and generate the most suitable transition effect for the current content and emotion based on its understanding of the "narrative intent." For example, if the system determines that the video is transitioning from a fast-moving scene to a quiet, contemplative scene, it may choose a slow, dissolving transition; if it is a sudden event, it may choose a fast cut. This improves creative efficiency and the quality of the work.
[0192] On the product side, such as Figure 3 or Figure 5 As shown, the business effects achieved in video platforms or applications include:
[0193] 1. Rapidly expand the short drama fan creation content library: The platform will no longer rely on large-scale manual editing, and will be able to automatically produce a large amount of high-quality short drama fan creation material at an exponential speed. This directly solves the current problem of insufficient supply of fan creation content, allowing users to find fresh, interesting, and professional fan creation content at any time.
[0194] 2. Significantly increases user dwell time and engagement: Rich, high-quality derivative content is more likely to attract users to click and watch, and content connections are formed between different short dramas, increasing the time users spend on the video app. Intelligent transition effects bring a smoother viewing experience, further enhancing user engagement.
[0195] 3. Driving platform traffic and viewership growth: A large amount of high-quality derivative content will become a new traffic entry point. It is expected to continuously bring an additional 6 million to 9 million views per day, effectively revitalizing the long-tail value of the drama IP and bringing new advertising and membership growth opportunities to the platform.
[0196] On the technology side, such as Figure 8 The diagram shown illustrates the overall process of a technical solution for intelligent generation of video splicing and transition effects provided in this application. Figure 8 As shown, the platform system provided in this application can automatically generate effects that conform to the narrative intent by deeply understanding the video content, ultimately realizing the secondary creation and splicing of videos. The entire process can be divided into the following core steps:
[0197] I. Video Input and Preprocessing: The platform system first receives the raw video clips uploaded by the user. These clips will then undergo the following processes:
[0198] Decoding and Basic Analysis: The platform system decodes video files into processable frame sequences and audio streams. Simultaneously, the platform system performs preliminary scene boundary detection and shot segmentation, dividing a large video into smaller, logically independent segments.
[0199] Data standardization: The platform system processes images, audio, and any accompanying text (such as subtitles) in a unified format to prepare for subsequent multimodal feature extraction.
[0200] II. Multimodal Content Understanding and Feature Extraction
[0201] This is a crucial step in understanding the video content for this application. The platform system will comprehensively extract features from multiple dimensions of the video, not just the visuals, including at least:
[0202] Visual feature extraction:
[0203] Image content: Recognize objects (people, scene props), scene type (indoor, outdoor, street scene), and people's actions (running, jumping, talking) in the image.
[0204] Lens characteristics: Analyze the shot type (long shot, close-up), camera movement (push, pull, pan, tilt), composition, color style, etc.
[0205] Visual emotion: Based on the lighting, color and other factors in the image, we can judge the emotional tendency of the visual (bright, dark, tense, warm). The extracted visual feature vector can be represented by F_v.
[0206] Audio feature extraction:
[0207] Voice content: If there is a dialogue, perform speech recognition and extract the text content.
[0208] Background music: Analyze the music's type, rhythm, melody, and emotion.
[0209] Sound effects: Identify environmental sound effects and their semantics.
[0210] Audio emotion: Judging auditory emotional tendencies based on sound characteristics.
[0211] Text feature extraction:
[0212] Subtitles / Dialogue: Extract subtitles and dialogue from the video itself or user input, perform semantic analysis, and obtain character dialogue information and plot clues.
[0213] III. Narrative Intent Inference (The core component of generating adaptive transition effects for "narrative intent" driven by multimodal collaboration)
[0214] like Figure 6 As shown, this part represents the core innovation within this community. The platform system not only identifies what's in the video, but also aims to "understand" what the video wants to express and how the story is unfolding. This process involves the following steps:
[0215] (1) Deep fusion of multimodal features and context modeling:
[0216] Process: The system performs deep, interactive fusion of the three modalities of information extracted in the second step: visual features (F_v), audio features (F_a), and text features (F_t). This is not simply stacking them together, but rather using a carefully designed fusion network (for example, a Transformer structure based on an attention mechanism) to allow these three types of information to "talk" to and "verify" each other.
[0217] Innovation: Traditional fusion might be just a simple feature stitching, but this application emphasizes depth and interactivity. This means that when a visual signal shows a "smiling person," the fusion network will simultaneously refer to "cheerful music" in the audio or "positive dialogue" in the text to more accurately confirm that this is a "joyful" emotion, rather than just an illusion created by facial expression. This interactivity enables the system to build a less ambiguous contextual understanding of video content.
[0218] (2) Analysis of temporal dependence and narrative rhythm:
[0219] Process: Video is constantly changing, so watching just a single moment or short clip is insufficient. Therefore, the system performs sequence analysis on the fused F_context in the time dimension, that is, it uses the temporal modeling module (TransformerEncoder) to capture the development and change patterns of the video content over time.
[0220] Innovation: The innovation here lies in the fact that the system not only understands the content of a single moment, but also the "story flow on the timeline". For example, by analyzing continuous images, sounds and dialogues, the system can identify more advanced narrative patterns such as "emotions moving from calm to climax", "pacing changing from slow to fast", or "plot development from foreshadowing to turning point". This ability is the foundation for understanding "narrative intent" because it allows us to see the trajectory of the story's development.
[0221] (3) Narrative Intent Quantification and Inference Model:
[0222] Process: After obtaining rich temporal context information, the system uses a specialized intent inference model to map it into specific, actionable "narrative intents". These can be discrete labels (such as tension, relaxation, surprise, turning point, flashback, etc.) or continuous vector representations to describe the emotions, rhythm, or plot state at the connection points of video segments.
[0223] Innovation: This is no longer about manually setting rules, but about enabling machines to autonomously learn and extract the "director's intent" from complex video features. Trained on a large amount of data, this model can automatically distinguish feature patterns under different narrative intentions. For example, it can differentiate between "a calm scene transitioning to another calm scene" (requiring a smooth transition) and "a calm scene suddenly interrupted by a terrifying event" (requiring a fast, impactful transition).
[0224] The core innovation of this application's technical solution in "narrative intent inference" lies in its ability to transcend surface-level content recognition and truly understand the "story logic" and "emotional changes" inherent in video clips through multimodal deep interactive fusion combined with powerful temporal modeling capabilities (especially utilizing advanced models such as Transformer). This automated and quantitative inference of "narrative intent" is the cornerstone for achieving subsequent adaptive and intelligent transition effects, completely changing the traditional passive and rigid splicing method and endowing machines with a "narrative insight" closer to that of human editors.
[0225] IV. Adaptive Video Stitching and Transition Effect Generation (Main Implementation of Adaptive Transition Effect Generation Driven by Multimodal Collaboration and “Narrative Intent”)
[0226] like Figure 7 As shown, this step truly brings the "narrative intent" inferred from the previous step to life, transforming it into visually appealing video transition effects that fit the story, and completing the final video stitching. The intelligence and adaptability of this process are key highlights of this application.
[0227] Transition strategy matching and type selection:
[0228] Process: The system receives narrative intent from the "Narrative Intent Inference" module, such as "emotional shift from tension to relaxation," "rapid scene transition," or "time passage." The system maintains a rich library of transition effects. Each transition effect in the library (such as fade-in / fade-out, dissolve, wipe, flash, blur, shimmer, etc.) is pre-labeled, associated with a specific narrative intent type, emotional intensity, or scene transition mode. Based on the narrative intent output by the "Narrative Intent Inference" module, the system intelligently matches and selects the most appropriate transition effect type from this library.
[0229] Innovation: This processing method no longer involves random or fixed choices based on manual rules, but rather on an understanding of the deep "narrative intent" of the video content. Like an experienced editor, the system can select the transitions that best enhance the atmosphere and advance the narrative, based on the needs of the story. For example, if it detects an intent of "sudden fright or an event breaking out," the system will choose a quick "hard cut" or "flash white"; while for an intent of "a memory or longing," it may choose a slow "dissolve" or "soft blur."
[0230] Dynamic optimization of transition parameters:
[0231] Process: After the system determines the type of transition effect, it will further dynamically and finely adjust various parameters of the transition effect. These parameters include, but are not limited to: transition duration, transition area, animation direction, blur intensity, color change, and volume gradient curve. The adjustment of these parameters will take into account the subtlety of the transition, as well as the visual and audio characteristics of the two video clips (Clip A and Clip B).
[0232] Innovation: The innovation here lies in the "dynamic adaptability" of the parameters. For example, with the same "fade-in / fade-out" effect, if the "narrative intent" (I_narrative) indicates a "gentle scene transition," the duration might be longer and the change smoother; if the "narrative intent" indicates a "rapidly advancing plot segment," the duration will be shorter. Another example is the "wipe" effect. The system can intelligently determine the direction of the wipe (from left to right, from top to bottom, etc.) based on the movement direction of the main elements or the composition of the scene before and after, allowing the transition animation to blend perfectly with the screen content, rather than being applied rigidly.
[0233] Content-aware stitching and effect rendering:
[0234] Process: The system performs high-precision geometric alignment and content fusion on the two video clips to be stitched (Clip A and Clip B), and renders the final transition effect based on dynamically determined parameters. This step also includes video stitching technology, but incorporates more content-aware capabilities.
[0235] For example, dynamic target processing: If there is a moving object near the stitching seam, the system will intelligently avoid the stitching seam through target detection and optical flow tracking, or use multi-frame information to perform temporal completion or intelligent removal of the dynamic target to avoid ghosting or tearing.
[0236] Video stabilization and smoothness: Fine-tuning and stabilization of the image at the seams to ensure high consistency of the image before and after the transition, eliminating slight jitter and ensuring visual smoothness.
[0237] Innovation: The innovation here lies in "content-aware" stitching and rendering, which goes beyond pure image pixel processing. The system knows what's in the image, where it can be cut, and where it cannot. It can even perform protective processing or intelligent repair on key areas, making the final stitching effect not only technically flawless but also more visually natural and in line with human perception.
[0238] The core innovation of this application's technical solution in "adaptive video stitching and transition effect generation" lies in using the deep "narrative intent" inferred from the previous stage as the core driving force to achieve intelligent type selection and dynamic parameter optimization of transition effects. Combined with content-aware geometric alignment and fusion technologies, the transitions between the final stitched video segments are no longer rigid technical connections, but rather organic, smooth, and emotionally resonant links with strong storytelling power. This enables the machine to simulate, and even surpass, the artistic insight and refined operational capabilities of human editors in transition processing.
[0239] The beneficial effects of the technical solution provided in this application are as follows:
[0240] After implementation, the technical solution of this application will generate significant technical and commercial value, mainly reflected in the following aspects:
[0241] 1. Significantly Improves Creation Efficiency and Content Supply: The technical solution proposed in this application automates and intelligently generates video splicing and transition effects, greatly shortening the production cycle and workload for content creators (especially short drama derivative teams). It is expected to help video platforms quickly generate and supply more than 300,000 derivative materials, effectively covering both existing and new dramas, and greatly alleviating the problem of insufficient supply of high-quality content on the platform.
[0242] 2. Significantly increase user activity and viewership: Rich, high-quality derivative content can attract more users. According to calculations, this solution is expected to bring an average daily increase of 6 million to 9 million views to the platform, indicating that users have a strong demand for and higher acceptance of intelligently generated high-quality derivative content.
[0243] 3. Enhance content quality and user experience: Through multimodal collaborative narrative intent inference, the video transition effects generated by this solution are more natural, more narrative and artistic, which will improve the overall professionalism and viewing experience of short drama derivative works, allowing users to have a smoother and more immersive experience when consuming content.
[0244] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0245] Based on the same inventive concept, this application also provides a video generation apparatus for implementing the video generation method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more video generation apparatus embodiments provided below can be found in the limitations of the video generation method described above, and will not be repeated here.
[0246] In one embodiment, such as Figure 9 As shown, a video generation apparatus is provided, including: a decoding module 902, an extraction module 904, a determination module 906, and a processing module 908, wherein:
[0247] The decoding module 902 is used to decode at least two video segments to obtain decoded data of at least two video segments.
[0248] The extraction module 904 is used to extract features from the decoded data to obtain multimodal features of at least two video segments.
[0249] The determining module 906 is used to determine the narrative intent corresponding to at least two of the video segments based on the multimodal features; and to determine the transition effect adapted between the at least two video segments based on the narrative intent.
[0250] The processing module 908 is used to splice at least two video segments based on the transition effect to obtain a target video containing at least two video segments.
[0251] In one embodiment, the apparatus further includes: a display module, configured to display a video production page for generating a target video in response to a triggered video creation operation; and an acquisition module, configured to acquire at least two video clips input by the user through the video production page.
[0252] In one embodiment, the acquisition module is further configured to acquire video files of at least two video segments entered by the user through the information input box in response to an information input operation triggered in the information input box of the video production page; or, in response to a trigger operation on the shortcut menu control in the video production page, acquire video files of at least two video segments entered by the user through the shortcut menu control.
[0253] In one embodiment, the decoded data is either first decoded data or second decoded data; the decoding module is further configured to decode at least two video segments to obtain frame sequences, audio streams, and text information of at least two video segments, wherein the frame sequences, audio streams, and text information constitute the first decoded data; or, to decode at least two video segments to obtain frame sequences, audio streams, and text information of at least two video segments; the apparatus further includes: a detection module configured to perform scene boundary detection on the frame sequences to obtain a grouping sequence corresponding to the frame sequences, wherein the grouping sequence, the audio stream, and the text information constitute the second decoded data.
[0254] In one embodiment, the decoded data includes at least a frame sequence, an audio stream, and text information; the processing module is further configured to perform standardization processing on the frame sequence, the audio stream, and the text information to obtain a time-seriesd frame sequence vector, an audio stream vector, and a text information vector; the extraction module is further configured to perform feature extraction on the frame sequence vector to obtain visual features of at least two video segments; perform feature extraction on the audio stream vector to obtain audio features of at least two video segments; and perform feature extraction on the text information vector to obtain text features of at least two video segments.
[0255] In one embodiment, the extraction module is further configured to perform feature extraction on the image content of the frame sequence vector to obtain at least two first-dimensional visual features of the video segments; perform feature extraction on the lens characteristics of the frame sequence vector to obtain at least two second-dimensional visual features of the video segments; and perform feature extraction on the visual emotion of the frame sequence vector to obtain at least two third-dimensional visual features of the video segments; the determination module is further configured to determine the visual features of at least two video segments based on the first-dimensional visual features, the second-dimensional visual features, and the third-dimensional visual features.
[0256] In one embodiment, the extraction module is further configured to extract features from the speech dialogue content of the audio stream vector to obtain first-dimensional audio features of at least two video segments; extract features from the background sound effects of the audio stream vector to obtain second-dimensional audio features of at least two video segments; extract features from the audio emotion of the audio stream vector to obtain third-dimensional audio features of at least two video segments; and the determination module is further configured to determine the audio features of at least two video segments based on the first-dimensional audio features, the second-dimensional audio features, and the third-dimensional audio features.
[0257] In one embodiment, the multimodal features include visual features, audio features, and text features; the device further includes: a fusion module for fusing the visual features, audio features, and text features to obtain fused features; an analysis module for performing sequence analysis on the fused features in the time dimension to obtain temporal context information; and a processing module for mapping the temporal context information to obtain narrative intents corresponding to at least two of the video segments.
[0258] In one embodiment, the fusion module is further configured to fuse the visual features, the audio features, and the text features using a fusion model to obtain fused features; wherein the fusion model is trained based on a multimodal video dataset, and the multimodal data in the multimodal video dataset carries visual labels, audio labels, and text feature labels.
[0259] In one embodiment, the narrative intent is at least one of discrete labels or continuous vectors; the temporal context information includes temporal features; the processing module is further configured to map the temporal features through a first intent inference model to obtain discrete labels or continuous vectors; or, to map the temporal features through a second intent inference model to obtain discrete labels and continuous vectors.
[0260] In one embodiment, the processing module is further configured to map the temporal features using the dual-branch structure in the second intent inference model to obtain discrete labels and continuous vectors; the discrete labels include emotion labels, and the continuous vectors include continuous narrative vectors; wherein the dual-branch structure includes a narrative rhythm prediction branch and an emotion classification branch.
[0261] In one embodiment, the narrative intent includes an emotion tag and a continuous narrative vector; the device further includes: a search module, configured to search for a transition effect type suitable for at least two video segments from a transition effect library based on the emotion tag and the continuous narrative vector; and an adjustment module, configured to adjust the transition effect parameters based on the visual features, audio features, and the transition effect type of at least two video segments to obtain a transition effect with updated parameters.
[0262] In one embodiment, the processing module is further configured to perform geometric alignment and content fusion processing on at least two of the video segments to obtain a spliced video; and to render the spliced video based on the transition effect after parameter updates to obtain a target video containing the transition effect and at least two of the video segments.
[0263] In one embodiment, the processing module is further configured to perform spatial correction on key frames of at least two video segments based on feature point matching and optical flow estimation to obtain aligned video segments; wherein the position and direction of motion of the main subject of each video frame in the aligned video segments are consistent; and to perform edge transition and texture completion processing on the area to be spliced in the aligned video segments to obtain a spliced video after content fusion.
[0264] In one embodiment, the processing module is further configured to, when a moving target appears at the stitching seam position in the target video, perform temporal completion or erasure processing on the moving target based on video frame information before or after the appearance of the moving target in the target video, to obtain the target video after completion or erasure processing; or, determine the key subject in the target video and perform region protection processing on the area where the key subject is located, so that the area where the key subject is located in the target video remains intact.
[0265] In one embodiment, the apparatus further includes: an identification module for identifying dynamic elements in frames of the target video; and a processing module for performing local fine-tuning on the dynamic elements to obtain the locally fine-tuned target video; or, performing smoothing processing on the seams of the target video to obtain the smoothed target video; or, performing motion compensation processing on the seams of the target video to obtain the motion-compensated target video.
[0266] Each module in the aforementioned video generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0267] In one embodiment, a computer device is provided, which may be a terminal or a server. In this embodiment, the computer device is described as a terminal, and its internal structure diagram is as follows. Figure 10 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a video generation method. The display unit of the computer device is used to form a visually visible image. It can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0268] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0269] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0270] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0271] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0272] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0273] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, etc., and are not limited to these.
[0274] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0275] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A video generation method, characterized in that, The method includes: Decode at least two video segments to obtain decoded data for at least two video segments; Feature extraction is performed on the decoded data to obtain multimodal features of at least two of the video segments; Based on the multimodal features, determine the narrative intent corresponding to at least two of the video segments; Based on the narrative intent, determine the transition effects that are suitable for at least two of the video segments; Based on the transition effect, at least two video segments are spliced together to obtain a target video containing at least two video segments.
2. The method according to claim 1, characterized in that, Before decoding at least two video segments to obtain decoded data for at least two video segments, the method further includes: In response to a triggered video creation action, a video production page for generating the target video is displayed; The video creation page obtains at least two video clips input by the user.
3. The method according to claim 2, characterized in that, The step of obtaining at least two video clips input by the user through the video creation page includes: In response to an information input operation triggered in the information input box on the video production page, obtain video files containing at least two video clips entered by the user through the information input box; or, In response to a trigger operation on the shortcut menu control in the video production page, the system acquires video files containing at least two video clips entered by the user through the shortcut menu control.
4. The method according to claim 1, characterized in that, The decoded data is either first decoded data or second decoded data; the process of decoding at least two video segments to obtain decoded data for at least two video segments includes: Decoding at least two video segments yields frame sequences, audio streams, and text information for at least two video segments, wherein the frame sequences, audio streams, and text information constitute the first decoded data; or, Decode at least two video segments to obtain frame sequences, audio streams, and text information for at least two video segments; perform scene boundary detection on the frame sequences to obtain group sequences corresponding to the frame sequences, and the group sequences, the audio streams, and the text information constitute the second decoded data.
5. The method according to claim 1, characterized in that, The decoded data includes at least frame sequences, audio streams, and text information; before performing feature extraction on the decoded data to obtain multimodal features corresponding to at least two video segments, the method further includes: The frame sequence, the audio stream, and the text information are standardized to obtain time-series frame sequence vectors, audio stream vectors, and text information vectors. The step of extracting features from the decoded data to obtain multimodal features of at least two video segments includes: Feature extraction is performed on the frame sequence vector to obtain visual features of at least two of the video segments; Feature extraction is performed on the audio stream vector to obtain audio features of at least two of the video segments; Feature extraction is performed on the text information vector to obtain text features of at least two of the video segments.
6. The method according to claim 5, characterized in that, The step of extracting features from the frame sequence vector to obtain visual features of at least two video segments includes: Feature extraction is performed on the image content of the frame sequence vector to obtain at least two first-dimensional visual features of the video segments; Feature extraction is performed on the lens characteristics of the frame sequence vector to obtain the second-dimensional visual features of at least two of the video segments; Visual emotion features are extracted from the frame sequence vectors to obtain third-dimensional visual features of at least two of the video segments; Based on the first-dimensional visual features, the second-dimensional visual features, and the third-dimensional visual features, visual features of at least two of the video segments are determined.
7. The method according to claim 5, characterized in that, The step of extracting features from the audio stream vector to obtain audio features from at least two of the video segments includes: Feature extraction is performed on the speech dialogue content of the audio stream vector to obtain the first-dimensional audio features of at least two of the video segments; The background sound effects of the audio stream vector are feature extracted to obtain the second-dimensional audio features of at least two of the video segments; The audio emotion features of the audio stream vector are extracted to obtain the third-dimensional audio features of at least two of the video segments; Based on the first-dimensional audio features, the second-dimensional audio features, and the third-dimensional audio features, determine the audio features of at least two of the video segments.
8. The method according to claim 1, characterized in that, The multimodal features include visual features, audio features, and text features; Determining the narrative intent corresponding to at least two of the video segments based on the multimodal features includes: The visual features, audio features, and text features are fused to obtain fused features; The fused features are subjected to sequence analysis in the time dimension to obtain temporal context information; The temporal context information is mapped to obtain the narrative intent corresponding to at least two of the video segments.
9. The method according to claim 8, characterized in that, The process of fusing the visual features, the audio features, and the text features to obtain fused features includes: The visual features, audio features, and text features are fused using a fusion model to obtain fused features. The fusion model is trained on a multimodal video dataset, where the multimodal data carries visual labels, audio labels, and text feature labels.
10. The method according to claim 8, characterized in that, The narrative intent is at least one of discrete labels or continuous vectors; the temporal context information includes temporal features; the mapping process of the temporal context information to obtain the narrative intent corresponding to at least two video segments includes: The temporal features are mapped using a first intent inference model to obtain discrete labels or continuous vectors; or... The temporal features are mapped using a second intent inference model to obtain discrete labels and continuous vectors.
11. The method according to claim 10, characterized in that, The process of mapping the temporal features using the second intent inference model to obtain discrete labels and continuous vectors includes: By using the dual-branch structure in the second intent inference model, the temporal features are mapped to obtain discrete labels and continuous vectors; the discrete labels include emotion labels, and the continuous vectors include continuous narrative vectors. The dual-branch structure includes a narrative rhythm prediction branch and a sentiment classification branch.
12. The method according to claim 1, characterized in that, The narrative intent includes emotion tags and continuous narrative vectors; The determination of transition effects suitable for at least two video segments based on the narrative intent includes: Based on the emotion tag and the continuous narrative vector, search the transition effect library for a transition effect type that is suitable between at least two of the video segments; Based on the visual features, audio features, and transition effect type of at least two video segments, the transition effect parameters are adjusted to obtain the updated transition effect.
13. The method according to claim 12, characterized in that, Based on the transition effect, the process of splicing at least two video segments to obtain a target video containing at least two video segments includes: At least two of the video segments are geometrically aligned and content-fused to obtain a spliced video. The spliced video is rendered based on the transition effect after the parameters are updated, resulting in a target video that includes the transition effect and at least two video segments.
14. The method according to claim 13, characterized in that, The step of performing geometric alignment and content fusion processing on at least two video segments to obtain a spliced video includes: Based on feature point matching and optical flow estimation, spatial correction is performed on key frames of at least two video segments to obtain aligned video segments; wherein the position and direction of motion of the main subject in each video frame of the aligned video segment are consistent. Edge transition and texture completion processing are performed on the areas to be spliced in the aligned video segments to obtain a spliced video with fused content.
15. The method according to claim 1, characterized in that, After splicing at least two video segments based on the transition effect to obtain a target video containing at least two video segments, the method further includes: If a moving target appears at the stitching line position in the target video, based on video frame information before or after the appearance of the moving target, temporal completion or erasure processing is performed on the moving target to obtain the target video after completion or erasure processing; or... The key subjects in the target video are identified, and the areas where the key subjects are located are protected to ensure that the areas where the key subjects are located in the target video remain intact.
16. The method according to claim 1, characterized in that, After splicing at least two video segments based on the transition effect to obtain a target video containing at least two video segments, the method further includes: Identify dynamic elements in frames of the target video; The dynamic elements are locally fine-tuned to obtain the target video after local fine-tuning; or, The seams of the target video are smoothed to obtain the smoothed target video; or, Motion compensation processing is performed on the seams of the target video to obtain the motion-compensated target video.
17. A video generation apparatus, characterized in that, The device includes: A decoding module is used to decode at least two video segments to obtain decoded data of at least two video segments; The extraction module is used to extract features from the decoded data to obtain multimodal features of at least two of the video segments; The determining module is configured to determine the narrative intent corresponding to at least two of the video segments based on the multimodal features; and to determine a transition effect adapted between the at least two video segments based on the narrative intent. The processing module is used to splice at least two video segments based on the transition effect to obtain a target video containing at least two video segments.
18. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 16.
19. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 16.
20. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 16.