A method and system for generating a video driven by user interaction operations
Patent Information
- Application Number
- CN202610938168.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-26
- Publication Date
- 2026-09-22
AI Technical Summary
[0011]有鉴于此,本申请提供一种用户互动操作驱动的视频生成方法和系统,至少在一定程度上解决相关技术中视频单向输出无法实现与用户交互的问题
[0048]相比于现有技术,本申请提供一种用户互动操作驱动的视频生成方法和系统,其有益效果在于:本申请实施例通过执行上述方法,在生成原视频过程中识别出主要情节点、识别故事转折点、情感节点等作为预设节点,在生成的视频中标记该预设节点,视频播放中系统检测到预设节点,生成视频操作控件,给用户提供修改视频对应情节的通道,当用户输入操作指令,系统自动提取与当前视频片段关联的多个视频帧序列,根据用户选择的目标视频帧序列,提取目标视频帧序列对应的大纲文本,生成自动填充大纲文本的文本输入控件,为用户提供便捷修改的通道,根据用户修改后的文本生成修改视频,为用户提供便捷修改故事视频的通道,由于上述过程在关键情节设置互动节点,因此用户可以通过选择、输入等简单操作修改视频内容,无论是悬疑推理、情感故事,还是知识科普,不再是单方面播放视频,而是提供与用户的交互通道,为用户提供沉浸式体验感,并幅降低用户创作门槛。
Smart Images

Figure CN122802741A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video generation technology, and in particular to a video generation method and system driven by user interaction. Background Technology
[0002] Artificial intelligence-generated content (AIGC) technology, especially the development of text-to-video models, has made it possible to automatically generate high-quality film and television videos based on text descriptions. Creators only need to provide a story outline or script, and the model can automatically generate corresponding video footage and complete preliminary editing, which to some extent lowers the barrier to video content creation, improves production efficiency, and provides basic conditions for creation by the general public.
[0003] However, current mainstream AI video generation technologies and traditional film and television content production methods generally suffer from the following significant drawbacks:
[0004] First, the unidirectional and linear nature of content presentation:
[0005] Whether it's traditional film and television dramas or AI short dramas, their core narrative structure is a pre-set, fixed linear process. The content is output unilaterally from the producer to the audience, who can only passively receive the video plot and cannot operate or modify the plot development, character behavior, or story ending in any way.
[0006] Second, there is insufficient user engagement and immersion when watching videos.
[0007] Because of the lack of interactive elements, when viewers are uninterested in the current plot development, believe that the characters' decisions are unreasonable, or expect a different ending, their only option is to stop watching, fast forward, or switch to other content. This "leave if you're not satisfied" model results in a severe lack of deep audience engagement and immersive experience, and also reduces content completion rates and user stickiness.
[0008] Third, it cannot meet diverse and personalized viewing needs.
[0009] In traditional linear narrative models, a single work can only offer a single plot path. However, different viewers often have different preferences and expectations regarding the truth behind "suspenseful mystery," the resolution of "emotional stories," or the in-depth direction of "knowledge dissemination." Current technology cannot provide customized viewing experiences that align with the interests of different viewers within the same content framework.
[0010] In summary, existing video generation and presentation technologies are essentially still a "broadcast-style" one-way output, failing to address the audience's need for active participation and personalized exploration during the viewing process. Summary of the Invention
[0011] In view of this, this application provides a video generation method and system driven by user interaction, which at least to some extent solves the problem that one-way video output cannot achieve user interaction in related technologies.
[0012] A first aspect of this application provides a user-interactive operation-driven video generation method, the method comprising:
[0013] Once the original video is detected to have reached a preset node, video operation controls are generated and displayed.
[0014] In response to the user's input command to the video operation control, acquire multiple video frame sequences associated with the original video at the preset node;
[0015] In response to the selection operation of any video frame sequence among the plurality of video frame sequences, the outline text and associated video frames corresponding to the selected target video frame sequence are retrieved from the database.
[0016] Play the selected target video frame sequence on the display interface and generate a text input control for manipulating the outline text;
[0017] When a user is detected to have entered a text command through the text input control, a pre-set video generation model is run to generate a modified video according to the text command; or when a user is detected to have entered an image input command through the image input control, a pre-set video generation model is run to generate a modified video according to the image information corresponding to the image modification command.
[0018] In one possible implementation, in response to a user's input command to the video operation control, acquiring a sequence of multiple video frames associated with the original video at the preset node includes:
[0019] In response to the user's input command to the video operation control, the system searches the database for a text segment associated with the preset node.
[0020] Using text segments as indexes, the system searches for the corresponding video frame sequences within pre-created structured metadata.
[0021] In one possible implementation, before searching the database for the text paragraph associated with the preset node, the method further includes:
[0022] In response to the user's input command to the video operation control, the original video is displayed in the entity object control corresponding to the preset node;
[0023] In response to the user's confirmation command for the entity object control, generate an object feature vector for the target entity object corresponding to the confirmation command;
[0024] The database is searched for text paragraphs associated with the preset node, including:
[0025] Calculate the matching value between the object feature vector and multiple text paragraphs, and select the target text paragraph with a matching value greater than a preset threshold.
[0026] In one possible implementation, when a user inputs a text command through the text input control, a pre-set video generation model is run to generate a modified video according to the text command, including:
[0027] When a user inputs a text command through the text input control, a pre-set language model is run to detect the conflict locations of the modified text segment; the conflict locations are the logical conflict locations between the modified text segment and the original video outline text.
[0028] The conflict location is displayed, the user's confirmation of the conflict location is detected, and a pre-set video generation model is run to generate a modified video according to the text instructions.
[0029] In one possible implementation, the conflict location is displayed, a user confirmation operation on the conflict location is detected, and a pre-set video generation model is run to generate a modified video according to the text instructions, including:
[0030] Generate intermediate sentence text and display the conflict location and intermediate sentence text;
[0031] Upon detecting the user's confirmation of the conflict location and the middle sentence text, the pre-set video generation model is run to generate a modified video according to the text instructions.
[0032] A second aspect of this application provides a user-interactive operation-driven video generation system, the system comprising:
[0033] The detection module is used to detect when the original video has played to a preset node, and to generate and display video operation controls;
[0034] An interactive response module is used to respond to the user's input command to the video operation control and obtain a sequence of multiple video frames associated with the original video at the preset node;
[0035] The interactive response module is also used to respond to the selection operation of any video frame sequence among the plurality of video frame sequences by retrieving the outline text and associated video frames corresponding to the selected target video frame sequence from the database.
[0036] The control generation module is used to play the selected target video frame sequence on the display interface and generate text input controls for manipulating the outline text.
[0037] The video generation module is used to generate a modified video according to the text command when a user inputs a text command through the text input control.
[0038] In one possible implementation, the interactive response module includes:
[0039] The data lookup submodule is used to respond to the user's input command to the video operation control, search the database for the text segment associated with the preset node, and use the text segment as an index to search for the video frame sequence corresponding to the text segment in the pre-created structured metadata.
[0040] In one possible implementation, the interactive response module includes:
[0041] The display submodule is used to respond to the user's input command to the video operation control and display the original video to the entity object control corresponding to the preset node;
[0042] The vector generation module is used to respond to the user's confirmation command for the entity object control and generate an object feature vector for the target entity object corresponding to the confirmation command.
[0043] The data search submodule is specifically used to calculate the matching value between the object feature vector and multiple text paragraphs, and select the target text paragraph with a matching value greater than a preset threshold.
[0044] In one possible implementation, the video generation module includes:
[0045] The detection submodule is used to detect the conflict positions of the modified text fragment when a user inputs text commands through the text input control; the conflict positions are the logical conflict positions between the modified text fragment and the corresponding outline text of the original video.
[0046] The video generation submodule is used to display the conflict location, detect the user's confirmation operation on the conflict location, and run a pre-set video generation model to generate a modified video according to the text instructions.
[0047] In one possible implementation, the video generation submodule is specifically used to generate intermediate sentence text, display the conflict location and the intermediate sentence text; detect the user's confirmation operation on the conflict location and the intermediate sentence text, and run a pre-set video generation model to generate a modified video according to the text instructions.
[0048] Compared to existing technologies, this application provides a user-interactive operation-driven video generation method and system, the advantages of which are as follows: In the process of generating the original video, the embodiments of this application identify key plot points, story turning points, emotional points, etc., as preset nodes. These preset nodes are marked in the generated video. During video playback, the system detects the preset nodes and generates video operation controls, providing users with a channel to modify the corresponding plot of the video. When the user inputs an operation command, the system automatically extracts multiple video frame sequences associated with the current video segment. Based on the target video frame sequence selected by the user, the system extracts the outline text corresponding to the target video frame sequence and generates a text input control that automatically fills in the outline text, providing users with a convenient modification channel. The modified video is generated based on the user's modified text, providing users with a convenient channel to modify the story video. Because the above process sets interactive nodes in key plot points, users can modify the video content through simple operations such as selection and input. Whether it is suspenseful reasoning, emotional stories, or popular science, it is no longer a one-way video playback, but provides an interactive channel with the user, providing an immersive experience and significantly reducing the user's creative threshold. Attached Figure Description
[0049] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0050] Figure 1 A schematic diagram of a video generation architecture according to an embodiment of the present disclosure is shown;
[0051] Figure 2 This is a flowchart of the steps of the user interaction-driven video generation method proposed in the embodiments of this application;
[0052] Figure 3 This is a schematic diagram of an example video operation control according to this application;
[0053] Figure 4 This is a schematic diagram of an example entity object control of this application;
[0054] Figure 5 This is a schematic diagram of a text input control for outline text in one example of this application;
[0055] Figure 6 This is a functional block diagram of the user interaction-driven video generation system proposed in the embodiments of this application. Detailed Implementation
[0056] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0057] In this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
[0058] Example 1
[0059] Figure 1 A schematic diagram of a video generation architecture according to an embodiment of this disclosure is shown. (See reference...) Figure 1 The video generation architecture 100 may include a server 110, a terminal 120, and a network 130 providing a communication link. The server 110 and the terminal 120 can be connected via a wired or wireless network 130. The server 110 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, security services, and CDN.
[0060] Terminal 120 can be implemented in hardware or software. For example, when terminal 120 is implemented in hardware, it can be various electronic devices with a display screen and support page display, including but not limited to smartphones, tablets, e-book readers, laptops, and desktop computers. When terminal 120 is implemented in software, it can be installed in the electronic devices listed above; it can be implemented as multiple software programs or software modules (e.g., software programs or software modules used to provide distributed services) or as a single software program or software module, without specific limitations.
[0061] It should be noted that the video generation method provided in this embodiment can be executed by terminal 120, server 110, or jointly by terminal 120 and server 110. It should be understood that... Figure 1 The number of terminals, networks, and servers shown is for illustrative purposes only and is not intended to be a limitation. Any number of terminals, networks, and servers can be used depending on implementation needs.
[0062] Figure 2 This is a flowchart of the steps of the user interaction-driven video generation method proposed in the embodiments of this application, as follows: Figure 2 As shown, the steps include:
[0063] S11: The original video is detected to have played to a preset node, and video operation controls are generated and displayed.
[0064] Original videos refer to narrative videos generated using AICG, such as short dramas and situational advertisements. Artificial Intelligence Generated Content (AICG) refers to performing natural language processing on text data to obtain feature vectors, and then using image generation or video generation models to process these feature vectors to generate videos corresponding to the text data. For example, large language models for natural language processing include ChatGPT (OpenAI), Gemini (Google), NovelAI's Xialong, and DeepSeek. Image generation models include Midjourney, Flux, and CanvaAI. Video generation models include Sora (OpenAI) and Veo (Google).
[0065] In this embodiment of the application, during the process of generating video from story text using the AICG tool, a set of video frame sequences is generated for each text segment, establishing a relationship between the text segments and the video frame sequences. This relationship can be established by creating structured metadata and storing the storage path (video_path), video identifier (video_id), and text segment (source_text) of the video frame sequence in the structured metadata.
[0066] Table 1 shows an example of structured metadata established in this application. As shown in Table 1, one element can be used as an index to query other elements. For example, using video_id as the index, source_text, vector, and video_path can be queried.
[0067] Table 1:
[0068] video_id source_text vector video_path vid_001 The protagonist runs in the rainy night [0.12,0.56,...] / videos / gen_001.mp4 vid_002 The two were talking in the coffee shop [0.89,0.23, ...] / videos / gen_002.mp4
[0069] video_id is the encoding of the generated video sequence; an encoding is created sequentially for each generated video sequence.
[0070] Preset nodes are predefined nodes created during the outline creation process. They are built using story script analysis tools such as react-story-tree, drama-analyzer, and intellivng / mcp-server to perform path and diversity analysis on the story, determine the main plot points, identify story turning points and emotional nodes, and mark these identified plot points, turning points, and emotional nodes as preset nodes in the outline. For example, using... <mark>Tags are used to mark key plot points, story turning points, and emotional nodes in the corresponding text. The video generation model parses the text with the tags to generate a video and embeds the corresponding tag information in the generated video metadata. When the video is played, the tag information is identified to detect whether the original video has played to the preset node.
[0071] S12: In response to the user's input command to the video operation control, obtain a sequence of multiple video frames associated with the original video at the preset node.
[0072] Users can input commands to the video control through long presses, drags, etc., including commands to select video frames and to capture specific video segments.
[0073] When the marker information is detected, video operation controls are displayed on the display interface. Video operation controls can refer to interactive components in the display interface used to select video segments, including buttons, drag boxes, file selectors, etc. Figure 3 This is a schematic diagram of an example video operation control according to this application, such as... Figure 3 As shown, the video operation controls include a selection button and a drag-and-drop box. The user clicks the selection button to select the image frame at the current playback time, and then long-presses the trim handle to drag the drag-and-drop box, selecting the start and end points of the video with the image frame at the current playback time as the midpoint, thus obtaining a specific video segment at a preset node. Further, multiple video frame sequences associated with the original video at the preset node are obtained, that is, multiple video frame sequences associated with the specific video segment are obtained.
[0074] This application embodiment also proposes a specific process for performing step S12, the steps of which include:
[0075] S121: In response to the user's input command to the video operation control, search the database for a text segment associated with the preset node.
[0076] S122: Using the text segment as an index, find the video frame sequence corresponding to the text segment in the pre-created structured metadata.
[0077] In response to user input commands to the video operation controls, the system can determine the specific video segment associated with the node, extract the encoding associated with the specific video segment, search for the text segment (source_text) corresponding to the specific video segment in the structured metadata of the database, and then retrieve the text paragraph associated with the text segment in the outline.
[0078] One example of this application employs a semantic search method to retrieve text paragraphs associated with text fragments from the outline. The text fragments and text paragraphs in the outline are then converted into text features. The similarity between the text fragments and multiple other text paragraphs is calculated, and text paragraphs with similarity exceeding a threshold are selected. After obtaining the text paragraphs, the corresponding video IDs are queried again from the structured metadata in the database, and video frames are extracted from the storage path of the structured metadata.
[0079] In this embodiment of the application, the video clips that are plot-related to the specific video clip selected by the user are automatically displayed according to the above method, providing the user with selectable content to continue watching. The user can directly view the video clips of interest according to their own needs without waiting for other clips to play.
[0080] Before executing step S121, the feature vectors of the corresponding objects of the characters and items can be obtained through the following sub-steps, thereby matching the text paragraphs associated with the current character and item.
[0081] S120: In response to the user's input command to the video operation control, display the original video in the entity object control corresponding to the preset node. In response to the user's confirmation command to the entity object control, generate an object feature vector for the target entity object corresponding to the confirmation command.
[0082] The specific execution process of searching for text paragraphs associated with the preset node in the database is as follows: calculate the matching value between the object feature vector and multiple text paragraphs, and select the target text paragraph with a matching value greater than a preset threshold.
[0083] In one example of this application, entity objects include people, items, etc.; Figure 4 This is a schematic diagram of an example entity object control of this application, such as... Figure 4 As shown, received the Figure 3 Clicking the selection control pauses video playback, runs an embedded computer vision model to identify people and objects in the video frame, extracts their coordinate data, generates a selection area for the people and objects based on the coordinate data, and uses JavaScript to bind click and hover events to the selection area to obtain the entity object control. For example, when the user moves the mouse over a person in the frame, a hover event is triggered, the display device highlights the outline of the corresponding person, guiding the user to click on the object. When the user clicks on the highlighted outline area, the system receives a confirmation command from the user for the entity object control.
[0084] One example implementation of this application involves pre-training a semantic embedding model. The information of the target entity object, "Zhang San, male, the eldest disciple of the Tianjian Sect. He is resolute and taciturn, skilled in swordsmanship, especially proficient in the 'Tiangang Sword Technique'," is input into the semantic embedding model to obtain the object's feature vector. Input the text paragraph "Zhang San practices swordsmanship alone under the moon, his heart filled with longing for his master" into the semantic embedding model, and output the text feature vector. Calculate the cosine similarity between the target entity object and different text paragraphs to obtain the matching value S between the target entity object and different text paragraphs, S= .
[0085] This application embodiment presents users with selectable characters in the manner described above. Users select characters of interest according to their needs, and the system recommends video segments associated with those characters, providing links to directly jump to the videos related to those characters. Users do not need to wait for other segments to play before they can view the corresponding video segments and make corresponding modifications.
[0086] S13: In response to the selection operation of any video frame sequence among the plurality of video frame sequences, retrieve the outline text and associated video frames corresponding to the selected target video frame sequence from the database.
[0087] Obtain multiple video frame sequences associated with the original video at the preset node, and perform the following operations on each video frame sequence: synthesize the frame sequences into a stitched image, generate a preview image URL, display the preview image via API, bind the preview image to the address of the source_text corresponding to the video frame sequence, and use... The label is displayed on the real interface. After the user clicks on the preview image, the system extracts the source_text and jumps to the address of the target video frame sequence, and then plays the target video frame sequence.
[0088] After selecting the corresponding target video frame sequence, the user can continue to use the encoding (video_id) of the target video frame sequence as an index to query the text fragment (source_text) corresponding to the video frame sequence in the structured metadata as the outline text.
[0089] In one example, video frame vectors corresponding to all video frames of the original video are pre-calculated, and a vector database is established for all video frame vectors. During the application, S13 is executed to calculate the target image feature vector of the target video frame sequence, query the vector database for video frame vectors that match the target image feature vector, display the associated video frames corresponding to the matched video frame vectors, and generate an image modification interface for modifying the associated video frames.
[0090] The image editing interface includes tools such as move and select. Users can modify related video frames based on the editing interface, such as deleting parts of people or moving the positions of people. The system responds to user operations on related video frames based on the image editing interface, obtains the image information corresponding to the image modification command, and saves the image information.
[0091] S14: Play the selected target video frame sequence on the display interface and generate a text input control for manipulating the outline text.
[0092] Figure 5 This is a schematic diagram of a text input control for outline text in one example of this application, such as... Figure 5 As shown, the text input control is a dialog box. The system queries structured metadata to extract the text fragment (source_text) corresponding to the target video frame sequence, and fills the generated dialog box with the characters of the text fragment (source_text). The user can directly modify the characters in the dialog box. The system runs AICG to generate a modified video from the modified text fragment.
[0093] S15: When a user is detected to have entered a text command through the text input control, a pre-set video generation model is run to generate a modified video according to the text command; or when a user is detected to have entered an image input command through the image input control, a pre-set video generation model is run to generate a modified video according to the image information corresponding to the image modification command.
[0094] Users can input the image information saved by execution S13 into the system based on the generated image input control. The system then generates a modified video based on the video frames modified by the user. The modified video is generated based on the associated video frames, ensuring the consistency of objects such as people.
[0095] In another example, users can also upload images of people captured in a video using an image input control. Simultaneously, by inputting text commands and images of people, the system runs a pre-set video generation model that combines the people's information with the text commands to generate and modify the video.
[0096] This application embodiment, by executing the above method, identifies key plot points, story turning points, and emotional nodes as preset nodes during the original video generation process. These preset nodes are then marked in the generated video. During video playback, the system detects the preset nodes and generates video operation controls, providing users with a channel to modify the corresponding plot points. When a user inputs an operation command, the system automatically extracts multiple video frame sequences associated with the current video segment. Based on the target video frame sequence selected by the user, it uses multimodal interaction to extract the outline text and associated video frames corresponding to the target video frame sequence, generating a text input control that automatically fills in the outline text, providing users with a convenient modification channel. The system generates modified videos based on user-edited text, providing users with a convenient channel to edit story videos. It also offers a channel to modify related video frames. Users can modify related video frames and upload them to the system, which then generates a new video based on the modified frames, ensuring consistency between the generated videos. Furthermore, because the process incorporates interactive nodes at key plot points, users can modify video content through simple operations such as selection and input. Whether it's a suspenseful mystery, an emotional story, or a science popularization piece, the video is no longer a one-way playback but rather provides an interactive channel with the user, offering an immersive experience and significantly lowering the barrier to entry for user creation.
[0097] This application embodiment also provides a specific process for executing S15:
[0098] S151: When a user inputs a text command through the text input control, a pre-set language model is run to detect the conflict position of the modified text segment; the conflict position is the logical conflict position between the modified text segment and the original video outline text.
[0099] S152: Display the conflict location, detect the user's confirmation operation on the conflict location, and run the pre-set video generation model to generate a modified video according to the text instructions.
[0100] One example of this application employs a Missing Logic Detector by Emotion and Action (MLD-EA) to detect conflict locations in modified text fragments. MLD-EA extracts the actions performed by characters in the text fragment, formats these actions into structured tags, and extracts the emotional states performed by characters in the text fragment, formatting these emotional states into structured tags, thus obtaining the original action sequence and the original emotional sequence. A timeline is created, with coordinates representing consecutive time points, which are sentence numbers in the story. For example, time point K represents the position between sentence k and sentence k+1 in the original text. The original action sequence and the original emotional sequence are filled into the timeline according to the corresponding times, establishing a one-to-one correspondence between the original action sequence and the original emotional sequence at the time points, resulting in the action-emotion trajectory.
[0101] Extract the modification action tags and modification sentiment tags from the modified text fragment. Fill the modification action tags and modification sentiment tags into the time trajectory according to the time point. MLD-EA detects the consistency of the action-sentiment trajectory filled in the modification action tags and modification sentiment tags. When a missing logical link is detected, output the time point K value. According to the K value, find the sentence corresponding to the missing logical link in the text fragment to obtain the conflict position.
[0102] The conflict location is displayed to prompt the user for confirmation. If the user finds the conflict location caused by the modified text acceptable, they can enter a confirmation command, and the system will run a pre-set video generation model to generate the modified video according to the text command.
[0103] Through the above process, this application embodiment detects the rationality of the user-modified text paragraphs. When a logical conflict is detected in the original outline of the user-modified text, the conflict position is located by pointing to the K value of the sentence boundary, and the text content at the conflict position is displayed to prompt the user for confirmation. If the user confirms, the pre-set video generation model continues to run to generate the modified video according to the text instructions, avoiding logical confusion caused by directly generating the video according to the user-modified text and improving the user experience.
[0104] To further ensure the logical integrity of the generated video, this application embodiment also proposes to generate intermediate sentence text based on the context after locating the conflict position, reallocate attention weights based on the intermediate sentence text, and use the intermediate sentence text to trigger the generation of a new hidden layer vector so that the two conflicting vectors return to the same logical trajectory.
[0105] Execution S152: Displaying the conflict location, detecting the user's confirmation operation on the conflict location, and running the pre-set video generation model to generate a modified video according to the text instructions includes the following process:
[0106] Generate intermediate sentence text and display the conflict location and intermediate sentence text;
[0107] Upon detecting the user's confirmation of the conflict location and the middle sentence text, the pre-set video generation model is run to generate a modified video according to the text instructions.
[0108] This application embodiment not only detects and locates the conflicting positions between the modified text fragment and the original outline through the above method, prompting the user for confirmation, but also generates the middle sentence text of the conflicting position to supplement the logical conflict position, thus ensuring the rationality of the generated video.
[0109] In one example of this application, after receiving a user's text instruction to modify the outline text, the first sentence of the modified text is "Xiaoming saw the rain outside the window and felt a pang of sadness," and the second sentence is "He picked up an umbrella and walked out of the house." The first and second sentences are formatted to obtain the original action sequence and the original emotion sequence, and an association is established to obtain the action-emotion trajectory {time step 1 (sadness, looking), time step 2 (none, picking up, walking)}. The MLD-EA model detects that the emotion and action changes at time step 1 and time step 2 are inconsistent with common sense, generates a K value of 1, locates the first and second sentences, and generates the middle sentence text "Xiaoming plans to go out for a walk," ensuring the logical rationality of the text paragraphs, thereby ensuring the logical rationality of the generated video clips.
[0110] Based on the user-interactive operation-driven video generation method provided in Embodiment 1 of this application, correspondingly, Embodiment 2 of this application also provides a user-interactive operation-driven video generation system. Figure 6 This is a functional block diagram of the user interaction-driven video generation system proposed in the embodiments of this application, such as... Figure 6 As shown, the system includes:
[0111] The detection module 601 is used to detect when the original video has played to a preset node, and to generate and display video operation controls;
[0112] The interactive response module 602 is used to respond to the user's input command to the video operation control and obtain a sequence of multiple video frames associated with the original video at the preset node;
[0113] The interactive response module 602 is also used to respond to the selection operation of any video frame sequence among the plurality of video frame sequences by retrieving the outline text and associated video frames corresponding to the selected target video frame sequence from the database.
[0114] The control generation module 603 is used to play the selected target video frame sequence on the display interface and generate a text input control for manipulating the outline text.
[0115] The video generation module 604 is used to generate a modified video according to the text command when a user inputs a text command through the text input control.
[0116] In one possible implementation, the interactive response module includes:
[0117] The data lookup submodule is used to respond to the user's input command to the video operation control, search the database for the text segment associated with the preset node, and use the text segment as an index to search for the video frame sequence corresponding to the text segment in the pre-created structured metadata.
[0118] In one possible implementation, the interactive response module includes:
[0119] The display submodule is used to respond to the user's input command to the video operation control and display the original video to the entity object control corresponding to the preset node;
[0120] The vector generation module is used to respond to the user's confirmation command for the entity object control and generate an object feature vector for the target entity object corresponding to the confirmation command.
[0121] The data search submodule is specifically used to calculate the matching value between the object feature vector and multiple text paragraphs, and select the target text paragraph with a matching value greater than a preset threshold.
[0122] In one possible implementation, the video generation module includes:
[0123] The detection submodule is used to detect the conflict positions of the modified text fragment when a user inputs text commands through the text input control; the conflict positions are the logical conflict positions between the modified text fragment and the corresponding outline text of the original video.
[0124] The video generation submodule is used to display the conflict location, detect the user's confirmation operation on the conflict location, and run a pre-set video generation model to generate a modified video according to the text instructions.
[0125] The specific principles and execution processes of each unit in the language model retrieval data processing device disclosed in Embodiment 2 of this application can be found in the corresponding parts of the user interaction operation-driven video generation method disclosed in Embodiment 1 of this application, and will not be repeated here.
[0126] Example 3
[0127] Embodiment 3 of this application provides a server, including: a processor and a memory, the processor and the memory being connected via a communication bus; wherein, the processor is used to call and execute a program stored in the memory; the memory is used to store the program, the program being used to implement the user interaction operation driven video generation processing method provided in Embodiment 1 of this application.
[0128] Example 4
[0129] Embodiment 4 of this application provides a computer-readable storage medium storing computer-executable instructions for executing a user-interactive operation-driven video generation method as provided in Embodiment 1 of this application.
[0130] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computing software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0131] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0132] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.< / mark>
Claims
1. A video generation method driven by user interaction, characterized in that, The method includes: Once the original video is detected to have reached a preset node, video operation controls are generated and displayed. In response to the user's input command to the video operation control, acquire multiple video frame sequences associated with the original video at the preset node; In response to the selection operation of any video frame sequence among the plurality of video frame sequences, the outline text and associated video frames corresponding to the selected target video frame sequence are retrieved from the database. Play the selected target video frame sequence on the display interface and generate a text input control for manipulating the outline text; When a user is detected to have entered a text command through the text input control, a pre-set video generation model is run to generate a modified video according to the text command; or when a user is detected to have entered an image input command through the image input control, a pre-set video generation model is run to generate a modified video according to the image information corresponding to the image modification command.
2. The method according to claim 1, characterized in that, In response to user input commands to the video operation controls, acquire multiple video frame sequences associated with the original video at the preset node, including: In response to the user's input command to the video operation control, the system searches the database for a text segment associated with the preset node. Using text segments as indexes, the system searches for the corresponding video frame sequences within pre-created structured metadata.
3. The method according to claim 2, characterized in that, Before searching the database for the text paragraph associated with the preset node, the method further includes: In response to the user's input command to the video operation control, the original video is displayed in the entity object control corresponding to the preset node; In response to the user's confirmation command for the entity object control, generate an object feature vector for the target entity object corresponding to the confirmation command; The database is searched for text paragraphs associated with the preset node, including: Calculate the matching value between the object feature vector and multiple text paragraphs, and select the target text paragraph with a matching value greater than a preset threshold.
4. The method according to claim 1, characterized in that, When a user is detected to have entered a text command through the text input control, a pre-set video generation model is run to generate a modified video according to the text command, including: When a user inputs a text command through the text input control, a pre-set language model is run to detect the conflict locations of the modified text segment; the conflict locations are the logical conflict locations between the modified text segment and the original video outline text. The conflict location is displayed, the user's confirmation of the conflict location is detected, and a pre-set video generation model is run to generate a modified video according to the text instructions.
5. The method according to claim 4, characterized in that, Display the conflict location, detect the user's confirmation of the conflict location, and run a pre-set video generation model to generate a modified video according to the text instructions, including: Generate intermediate sentence text and display the conflict location and intermediate sentence text; Upon detecting the user's confirmation of the conflict location and the middle sentence text, the pre-set video generation model is run to generate a modified video according to the text instructions.
6. A user-interactive operation-driven video generation system, characterized in that, The system includes: The detection module is used to detect when the original video has played to a preset node, and to generate and display video operation controls; An interactive response module is used to respond to the user's input command to the video operation control and obtain a sequence of multiple video frames associated with the original video at the preset node; The interactive response module is also used to respond to the selection operation of any video frame sequence among the plurality of video frame sequences by retrieving the outline text and associated video frames corresponding to the selected target video frame sequence from the database. The control generation module is used to play the selected target video frame sequence on the display interface and generate text input controls for manipulating the outline text. The video generation module is used to generate a modified video according to the text command when a user inputs a text command through the text input control; or to generate a modified video according to the image information corresponding to the image modification command when a user inputs an image command through the image input control.
7. The user-interactive operation-driven video generation system according to claim 6, characterized in that, The interactive response module includes: The data lookup submodule is used to respond to the user's input command to the video operation control, search the database for the text segment associated with the preset node, and use the text segment as an index to search for the video frame sequence corresponding to the text segment in the pre-created structured metadata.
8. The user-interactive operation-driven video generation system according to claim 7, characterized in that, The interactive response module includes: The display submodule is used to respond to the user's input command to the video operation control and display the original video to the entity object control corresponding to the preset node; The vector generation module is used to respond to the user's confirmation command for the entity object control and generate an object feature vector for the target entity object corresponding to the confirmation command. The data search submodule is specifically used to calculate the matching value between the object feature vector and multiple text paragraphs, and select the target text paragraph with a matching value greater than a preset threshold.
9. The user-interactive operation-driven video generation system according to claim 6, characterized in that, The video generation module includes: The detection submodule is used to detect the conflict positions of the modified text fragment when a user inputs text commands through the text input control; the conflict positions are the logical conflict positions between the modified text fragment and the corresponding outline text of the original video. The video generation submodule is used to display the conflict location, detect the user's confirmation operation on the conflict location, and run a pre-set video generation model to generate a modified video according to the text instructions.
10. The user-interactive operation-driven video generation system according to claim 9, characterized in that, The video generation submodule is specifically used to generate intermediate sentence text, display the conflict position and intermediate sentence text; detect the user's confirmation operation on the conflict position and intermediate sentence text, and run a pre-set video generation model to generate a modified video according to the text instructions.