Video generation method, device, electronic device, storage medium and program product based on large model

By processing video information with a large model to generate interactive orchestration logic and automatically generate interactive videos, the problem of users' lack of convenient interaction when browsing videos is solved, and efficient information interaction and immersive experience are achieved.

CN118870146BActive Publication Date: 2025-09-30BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411188055.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2025-09-30
Estimated Expiration
2044-08-28

AI Technical Summary

Technical Problem

Users lack convenient ways to interact while browsing videos, and video producers find it difficult to quickly generate interactive videos for user interaction. Existing interactive operations are complex and single.

Method used

Use large models to process video-related information, generate interactive choreography logic, and automatically generate interactive videos. By combining the interactive choreography logic with the original video, information interaction can be achieved.

Benefits of technology

It improves the efficiency of interactive video generation, enhances the convenience and immersion of users' information interaction during video browsing, and improves the interactive experience of the target objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118870146B_ABST
    Figure CN118870146B_ABST
Patent Text Reader

Abstract

The present disclosure provides a video generation method, device, electronic device, storage medium and program product based on a large model, which relates to the field of artificial intelligence technology, especially to technical fields such as intelligent dialogue, deep learning, large models, and generative models, and can be applied to application scenarios such as intelligent assistants, Internet e-commerce, intelligent video browsing, video live broadcast, intelligent search, and video search. The method includes: responding to a video production request for an operation interface, obtaining video-related information, the video-related information including the original video; using the large model to process the video-related information to obtain interactive arrangement logic, the interactive arrangement logic represents the arrangement results for the interactive control, and the interactive control is used for information interaction; and generating an interactive video based on the interactive arrangement logic and the original video, the interactive video is used to display the video content of the original video during playback, and to interact with the target object based on the interactive arrangement logic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to technical fields such as intelligent dialogue, deep learning, large models, and generative models, and can be applied to application scenarios such as intelligent assistants, Internet e-commerce, intelligent video browsing, live video, intelligent search, and video search. Background Art

[0002] With the popularization of smart terminal devices such as smartphones and tablet computers, users can conveniently watch video resources through the display screens of smart terminal devices. Summary of the Invention

[0003] The present disclosure provides a large model-based video generation method, device, electronic device, storage medium and program product.

[0004] According to one aspect of the present disclosure, a video generation method based on a large model is provided, comprising: obtaining video-related information in response to a video production request for an operation interface, the video-related information including the original video; processing the video-related information using the large model to obtain interactive choreography logic, the interactive choreography logic representing the choreography results for interactive controls, the interactive controls being used for information interaction; and generating an interactive video based on the interactive choreography logic and the original video, the interactive video being used to display the video content of the original video during playback, and to interact with the target object based on the interactive choreography logic.

[0005] According to another aspect of the present disclosure, a video generation device based on a large model is provided, including: a first acquisition module, used to respond to a video production request for an operation interface and obtain video-related information, the video-related information including the original video; an interactive orchestration logic acquisition module, used to use the large model to process the video-related information to obtain the interactive orchestration logic, the interactive orchestration logic represents the orchestration results for the interactive controls, and the interactive controls are used for information interaction; and a generation module, used to generate an interactive video based on the interactive orchestration logic and the original video, the interactive video is used to display the video content of the original video during playback, and to interact with the target object based on the interactive orchestration logic.

[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned large model-based video generation method.

[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the above-mentioned large model-based video generation method.

[0008] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the above-mentioned large model-based video generation method when executed by a processor.

[0009] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0011] Figure 1 An exemplary system architecture to which a large model-based video generation method and apparatus according to an embodiment of the present disclosure can be applied is shown;

[0012] Figure 2 A flowchart of a method for generating a video based on a large model according to an embodiment of the present disclosure is shown;

[0013] Figure 3 An application scenario diagram of the large model-based video generation method according to an embodiment of the present disclosure is shown;

[0014] Figure 4 A schematic diagram of a video generation method based on a large model according to an embodiment of the present disclosure is shown;

[0015] Figure 5 A diagram showing an application scenario of a large model-based video generation method according to another embodiment of the present disclosure is shown;

[0016] Figure 6 A diagram showing an application scenario of a large model-based video generation method according to another embodiment of the present disclosure is shown;

[0017] Figure 7 A diagram showing an application scenario of a large model-based video generation method according to yet another embodiment of the present disclosure is shown;

[0018] Figure 8 A block diagram of a video generation device based on a large model according to an embodiment of the present disclosure is shown; and

[0019] Figure 9 A block diagram of an electronic device suitable for implementing a large model-based video generation method according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0020] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0021] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.

[0022] The inventors discovered that when users browse video resources, they are willing to express their emotions and opinions based on the video's plot content, style, and other video information to achieve information expression. However, the way for users to express information by interacting with terminal devices during video browsing is relatively simple and the interactive operations are relatively complex. It is also difficult for video producers to quickly create interactive methods for videos for user interaction.

[0023] The embodiments of the present disclosure provide a method, apparatus, electronic device, storage medium, and program product for video generation based on a large model. The method includes: obtaining video-related information in response to a video production request for an operation interface, the video-related information including the original video; processing the video-related information using the large model to obtain interactive arrangement logic, the interactive arrangement logic representing the arrangement results for interactive controls, and the interactive controls being used for information interaction; and generating an interactive video based on the interactive arrangement logic and the original video, the interactive video being used to display the video content of the original video during playback and to interact with a target object based on the interactive arrangement logic.

[0024] According to the embodiments of the present disclosure, by utilizing a large model to process video-related information including the original video, the interactive choreography logic for enabling the target object to conveniently conduct information interaction is automatically generated, and the interactive video is generated through the interactive choreography logic and the original video. The automatic generation of the interactive video can be achieved based on the information processing and prediction capabilities of the large model, so that the interactive video can be produced with fewer interactive operations, thereby improving the generation efficiency of the interactive video, and enabling the target object to conduct information interaction based on the interactive choreography logic while browsing the original video, thereby improving the boundary of the target object's information interaction.

[0025] Figure 1 An exemplary system architecture to which the large model-based video generation method and apparatus according to an embodiment of the present disclosure can be applied is shown.

[0026] It should be noted that Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not imply that the embodiments of the present disclosure may not be applied to other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which the large-scale model-based video generation method and apparatus may be applied may include a terminal device, but the terminal device may implement the large-scale model-based video generation method and apparatus provided by the embodiments of the present disclosure without interacting with a server.

[0027] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0028] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).

[0029] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0030] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports content browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.

[0031] The server can be a cloud server, also known as a cloud computing server or cloud host. It is a host product within the cloud computing service system that addresses the management difficulties and poor business scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or "VPS"). The server can also be a distributed system server or a server integrated with blockchain.

[0032] It should be noted that the video generation method based on the large model provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the video generation device based on the large model provided in the embodiment of the present disclosure can generally be set in the server 105. The video generation method based on the large model provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the video generation device based on the large model provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.

[0033] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0034] Figure 2 A flowchart of a large model-based video generation method according to an embodiment of the present disclosure is shown.

[0035] like Figure 2 As shown, the large model-based video generation method includes operations S210 to S230.

[0036] In operation S210 , video-related information is acquired in response to a video production request for an operation interface.

[0037] In operation S220 , the video-related information is processed using the large model to obtain interactive arrangement logic.

[0038] In operation S230 , an interactive video is generated based on the interactive choreography logic and the original video.

[0039] According to an embodiment of the present disclosure, the operation interface may include an operation interface of a video generation platform for generating interactive videos. The video-related information may include, but is not limited to, the original video. The video-related information may also include other types of information related to the original video, such as images, text, and charts. The original video may include any type of video information, such as a TV series, a movie, a short video, or an animated video. The embodiments of the present disclosure do not limit the specific type of the original video.

[0040] According to an embodiment of the present disclosure, the interactive choreography logic represents the choreography results for interactive controls, which are used for information interaction. For example, interactive controls may be controls or service resources that can receive information input by users and display response information corresponding to the input information. Interactive controls may include elements such as images, texts, and icons that can be rendered on a display screen. The interactive choreography logic may indicate control configuration properties such as the display duration and display mode of the interactive controls. However, it is not limited to this. The interactive choreography logic may also indicate the arrangement order, execution logic relationship, and the like between multiple interactive controls. The interactive controls may indicate the controls to be choreographed, and the interactive choreography logic may be one or more choreographed interactive controls obtained by configuring the interactive controls based on a large model to process video-related information.

[0041] According to embodiments of the present disclosure, interactive videos are used to display the video content of the original video during playback and to interact with target objects based on interactive programming logic. For example, during playback, the interactive video will play the original video content while displaying interactive controls according to the interactive programming logic, allowing the target object to interact with information through the interactive controls while watching the original video.

[0042] According to the embodiments of the present disclosure, a large model may refer to a deep learning model with large-scale model parameters. A large model generally contains hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. The large model may include a large-scale language model (LLM), a GPT (Generative Pre-trained Transformer), a large visual model, a multimodal large model, and the like. The large model involved in the embodiments of the present disclosure may be a general large model, or it may be an expert large model obtained after fine-tuning (Fine Tune) based on demand, and the embodiments of the present disclosure do not limit this. The large model can be used to process information in any modality such as images, text, video, and audio.

[0043] According to embodiments of the present disclosure, using a large model to process video-related information can include using the large model to process the original video, so that the large model can automatically generate and arrange one or more interactive controls based on its understanding of the original video's plot, dialogue, background image, and other video content, thereby achieving automated generation and arrangement of interactive controls. In this way, an interactive video can be generated by combining interactive arrangement logic with the original video.

[0044] According to the embodiments of the present disclosure, the use of a large model to process video-related information may also include the use of a large model to process video description information related to the original video, so that the large model can determine the key video content in the original video by understanding the video description information, thereby obtaining interactive controls suitable for display during the playback of the original video, and realizing the display mode and interaction mode of the interactive controls to enable the target object to conduct information interaction or emotional communication based on the interactive arrangement logic, thereby enhancing the target object's video viewing experience and viewing immersion.

[0045] According to the embodiments of the present disclosure, using a big model to process video-related information can also include using the big model to process the original video and video description information, so that the big model can combine the video content of the original video and the description content of the video description information, thereby improving the big model's ability to understand the content of the original video, and then improving the adaptability of the interaction method represented by the interactive arrangement logic to the video content of the original video, and further improving the adaptability between the information interaction method of the interactive video and the interaction needs of the target object.

[0046] According to an embodiment of the present disclosure, using a large model to process video-related information to obtain interactive choreography logic may include: using the large model to process the original video to obtain role description information of the video characters involved in the original video; and generating virtual character interactive controls corresponding to the video characters based on the role description information.

[0047] According to an embodiment of the present disclosure, processing an original video using a large model to obtain character description information may include processing the video content of the original video using a multimodal large model, and then identifying the image, dialogue, actions, and other video content of each character in the original video based on the multimodal large model's ability to process information in different modalities, and then outputting character description information that can describe the character's identity, personality, spoken language habits, and other character attributes. The character description information may correspond to the video character involved in the original video.

[0048] According to an embodiment of the present disclosure, the interactive choreography logic may include a virtual character interactive control for interacting with a target object based on the expression style of the video character in the original video. The virtual character interactive control may be generated by processing video-related information based on a large model, configuring the properties of the interactive control to be choreographed, and choreographing the execution logic.

[0049] According to an embodiment of the present disclosure, generating a virtual character interaction control corresponding to a video character based on character description information may include retrieving the virtual character interaction control based on the character description information to obtain the virtual character interaction control. However, this is not limited to the above, and may also include processing the character description information based on a large model, so that the large model can generate information that can more accurately interact with the target object in accordance with the expression style of the video character after fully understanding the expression style of the video character through the character description information, so that the virtual character interaction control generates response information that is more suitable for the expression style of the video character based on the information input by the target object, thereby enhancing the target object's participation and interactive immersion in the process of watching the original video.

[0050] According to an embodiment of the present disclosure, the video-related information further includes video description information for the original video. The video description information may be related to the video content in the original video.

[0051] According to an embodiment of the present disclosure, using a large model to process video-related information to obtain interactive choreography logic may also include: using a large model to process video description information to obtain role description information of video characters involved in the original video; and generating virtual character interactive controls corresponding to the video characters based on the role description information.

[0052] According to the embodiments of the present disclosure, using a large model to process video description information can reduce the amount of information processed by the large model, thereby improving the overall generation efficiency of the interactive orchestration logic.

[0053] According to an embodiment of the present disclosure, the video description information may include character introduction information of the video characters in the original video, for example, text, images, and other information that introduces the character attributes such as the experience, identity, and appearance of character A.

[0054] According to an embodiment of the present disclosure, the video description information may also include a video script of the original video, such as a video script generated based on the subtitles of the original video, or a script required by the relevant personnel who shot the original video.

[0055] According to an embodiment of the present disclosure, the video description information may further include a character image related to a video character.

[0056] According to an embodiment of the present disclosure, a virtual character interaction control may include a virtual character image corresponding to the video character. The virtual character image may include a two-dimensional head portrait, a three-dimensional bust portrait, a three-dimensional full-body portrait, and the like. The virtual character image may be obtained by processing video frames related to the video character in the original video using virtual character modeling technology. Alternatively, the virtual character image may be obtained by processing the character image of the video character using virtual character modeling technology.

[0057] According to embodiments of the present disclosure, video description information can be uploaded or input by the creator of the interactive video. Alternatively, video description information can be obtained by processing the original video's audio, subtitles, video frames, and other video content using a large model. Embodiments of the present disclosure do not limit the specific method for obtaining video description information.

[0058] It should be understood that the video characters involved in the original video may be characters that appear in the original video, such as the male protagonist, female protagonist, etc., or may be characters mentioned in the original video's dialogue, background sound, subtitles and other video content, such as the male protagonist's deceased father included in the male protagonist's dialogue.

[0059] According to an embodiment of the present disclosure, generating a virtual character interaction control corresponding to a video character based on the character description information may further include: using a large model to process the original video and style prompt words to obtain the virtual character interaction control.

[0060] According to an embodiment of the present disclosure, style prompt words are used to control the large model to understand the expression style of the video character, and the style prompt words are obtained by performing semantic recognition on the character description information.

[0061] According to embodiments of the present disclosure, a keyword extraction algorithm can be used to process character description information, and then style prompts can be determined based on the extracted keywords related to the video character. Alternatively, character description information can be processed based on a natural language understanding model (e.g., a trained Transformer model), and text related to the video character's attributes can be predicted based on the natural language understanding model's text prediction capabilities to obtain style prompts. Alternatively, a large model can be used to process character description information, eliminating the need for creators to describe the character by editing or uploading style description text, thereby increasing the flexibility and freedom for creators in producing interactive videos.

[0062] According to the embodiments of the present disclosure, a large model is used to process the original video and style prompt words. The large model can be controlled based on the style prompt words to understand the subtitles, audio, video frames and other video content related to the video character in the original video, so that the large model can understand the expression style of the video character in each playback period of the original video, and can mark the virtual character interaction controls based on the style prompt marks corresponding to the playback period, so that the generated virtual character interaction controls can simulate the expression style corresponding to the style prompt marks according to the style prompt marks corresponding to the playback period and interact with the target object, so that the target object can simulate the video character with the current viewing progress according to the virtual character interaction controls during the process of watching the original video, and interact with the target object through information, thereby enhancing the target object's immersion and participation in watching the video, and facilitating the target object to express emotions or express opinions.

[0063] According to an embodiment of the present disclosure, the virtual character interaction control of the interactive choreography logic may include a dialogue interaction window, which is used to obtain interaction information input by the target object and display response information related to the interaction information based on the expression style of the video character involved in the original video.

[0064] According to an embodiment of the present disclosure, a virtual character interaction control can process the interaction information input by a target object based on a style prompt tag. For example, the virtual character interaction control can obtain the interaction information and style prompt tag input by the target object, input the style prompt tag as a prompt word and the interaction information into a large model, obtain response information that matches the emotional expression intention of the interaction information, and display the response information through the virtual character interaction control to achieve information interaction with the target object by simulating the expression style of the video character.

[0065] According to embodiments of the present disclosure, the interactive information input by the target user and the response information generated by the large model can be displayed in a dialogue interaction window. The dialogue interaction window can be suspended over the original video or set at another location on the display screen related to the original video. The embodiments of the present disclosure do not limit the specific configuration of the dialogue interaction window.

[0066] It should be understood that the virtual character interaction control can also process multiple rounds of interaction information based on the large model to generate multiple rounds of response information, thereby realizing multiple rounds of information interaction with the target object.

[0067] Figure 3 An application scenario diagram of the large model-based video generation method according to an embodiment of the present disclosure is shown.

[0068] like Figure 3As shown, this application scenario may include a display screen 300. The interactive video may include the original video played on display screen 300, which may include a first video character 310 and a second video character 320. The interactive video may also include a dialogue interaction window H310 for a virtual character interaction control. The dialogue interaction window H310 may include a first virtual character image F311 corresponding to the first video character 310, and response information D311 output based on the expression style of the first video character 310. Based on the response information D311, the target subject may enter interaction information D321 "Hmm, be bolder" into the dialogue interaction window H310 of the virtual character interaction control. The large model may process the response information D311 and the interaction information D321 based on the style cue tag corresponding to the playback progress of the original video, and obtain a second response information "OK, OK." The dialogue interaction window H310 of the virtual character interaction control may display the second response information "OK, OK," thereby enabling the target subject to interact with information while watching the original video.

[0069] According to the embodiments of the present disclosure, the virtual character interaction controls automatically generated based on the big model can utilize the big model to process the interaction information and response information between the target object and the virtual character image, to understand the interaction intention of the target object and generate response information that matches the intention of the target object, thereby achieving dynamic response to the input of the target object, improving the naturalness and intelligence of the interactive video in the process of information interaction, and thus improving the interactive experience of the target object.

[0070] According to an embodiment of the present disclosure, using a large model to process video-related information to obtain interactive choreography logic may also include: using the large model to process plot prompt words, as well as at least one item of the original video and video description information, to obtain a triggering moment for controlling the display and / or exit of the virtual character interactive controls.

[0071] In one example, a large model can be used to process plot prompt words and video description information to obtain trigger moments.

[0072] In one example, a large model can be used to process plot prompts and original videos to obtain trigger moments.

[0073] In one example, a large model can be used to process plot prompt words, original video, and video description information to obtain trigger moments.

[0074] According to an embodiment of the present disclosure, plot prompts can be obtained by identifying key plot segments in the video description information, and the triggering time is related to the presentation period of the key plot segment in the original video. For example, the triggering time can be the start time of the presentation of the key plot segment, or the time before or after the start of the presentation of the key plot segment.

[0075] According to an embodiment of the present disclosure, plot prompt words can control the big model to understand the video content and / or video description information of the original video, and key plot information such as the playback period and plot content of the key plot segment, so that the big model can be controlled to generate virtual character interaction controls corresponding to the key plot information of the key plot segment, so that the display or exit timing of the virtual character interaction controls is more in line with the target object's emotions and interactive intentions when watching the key plot video of the original video.

[0076] According to an embodiment of the present disclosure, the interactive choreography logic may include the triggering moment of the virtual character interactive control, so as to control the virtual character interactive control to be displayed or exited according to the interactive choreography logic, so that the target object can interact with information in a timely manner according to the plot development of the original video, thereby conveniently expressing emotions or expressing opinions to enhance the video viewing experience.

[0077] According to an embodiment of the present disclosure, the virtual character interaction control performs information interaction with the target object based on the expression style of the video character involved in the key plot segment.

[0078] According to an embodiment of the present disclosure, the video characters involved in the key plot segments may include the characters that appear in the key plot segments of the original video, or may also include the video characters mentioned in the video content such as subtitles, background introduction text, dialogue, video frames, etc. related to the key plot segments. A large model can be used to process at least one of the original video and video description information to determine the video characters involved in the key plot segments, or other types of deep learning algorithms such as image recognition algorithms can be used to process the original video to obtain the video characters involved in the key plot segments. Alternatively, other types of deep learning algorithms such as natural language understanding algorithms can be used to process the video description information to obtain the video characters involved in the key plot segments.

[0079] According to embodiments of the present disclosure, the expression style of a video character involved in a key plot segment may include the video character's emotions, attitude, and tone of voice in the key plot segment. By interacting with the target subject based on the video character's emotions, attitude, and tone of voice in the key plot segment, the target subject can immerse themselves in the key plot while watching the key plot segment of the original video, achieving a deeper sense of participation and immersion.

[0080] Figure 4 A schematic diagram of a large model-based video generation method according to an embodiment of the present disclosure is shown.

[0081] like Figure 4As shown, the large model can obtain the trigger moments corresponding to each key plot segment in the original video by processing the plot prompt words, the original video, and the video description information. The multiple trigger moments can be represented by the first trigger moment T410, the second trigger moment T420, and the third trigger moment T430 in the playback timeline T400 of the original video. The large model obtains the first virtual character interaction control K410, the second virtual character interaction control K420, and the third virtual character interaction control K430 by processing the plot prompt words, the original video, and the character description information. The first virtual character interaction control K410, the second virtual character interaction control K420, and the third virtual character interaction control K430 can respectively have a mapping relationship with the first trigger moment T410, the second trigger moment T420, and the third trigger moment T430, so that based on the first trigger moment T410, the second trigger moment T420, and the third trigger moment T430, the first virtual character interaction control K410, the second virtual character interaction control K420, and the third virtual character interaction control K430 can be respectively controlled to display or exit during the original video playback process.

[0082] The first virtual character interaction control K410 may include a first virtual character image F411 and a first dialogue interaction window H410. The first virtual character interaction control K410 may interact with the target object based on the expression style of the first virtual character image F411 in the first key plot segment corresponding to the first triggering moment T410, such as the emotion.

[0083] The second virtual character interaction control K420 may include a second virtual character image F421 and a second dialogue interaction window H420. The second virtual character interaction control K420 may interact with the target object based on the second virtual character image F421's emotional expression, such as in the second key plot segment, corresponding to the second triggering moment T420. For example, in the second dialogue interaction window H420, based on character A's confusion in the second key plot segment, a second response message D420, "Should I confess my feelings?", may be expressed through voice communication.

[0084] The third virtual character interaction control K430 includes a second virtual character image F421 and a third dialogue interaction window H430. The third virtual character interaction control K430 can interact with the target object based on the second virtual character image F421's emotional expression, such as the third key plot segment, at the third triggering moment T430. For example, based on character A's frustration during the third key plot segment, the third response message D430, "Am I not good enough for her?", can be expressed via voice in the third dialogue interaction window H430.

[0085] According to the embodiments of the present disclosure, while watching the key plot segments of the original video, the target object can communicate and interact with the video character of the current playback progress based on the plot progression of the original video, thereby participating in the plot progression of the original video and improving participation and immersion.

[0086] According to an embodiment of the present disclosure, the key plot segment includes a product display segment displaying the target product, the plot prompt words include product plot prompt words used to control the large model to associate the target product, and the virtual character interaction control is also used to display product link information related to the target product.

[0087] According to an embodiment of the present disclosure, the product link information may include marketing resource information of the target product, or may also include purchase link information for purchasing the target product. The product link information may be displayed based on any form of information such as images, text, etc.

[0088] According to an embodiment of the present disclosure, product plot prompts may include identifiers, names, link resource information, and the like related to the target product. Product plot prompts can be used to control the large model to understand the target product that needs to be recommended in the current product display segment, and associate the product link information of the target product with the product display segment. By processing the product plot prompts and at least one of the original video and video description information through the large model, the triggering moment for displaying the product link information can be obtained, and the target product and product link information that appear in the original video during the display process can be simultaneously displayed to the target object in the product display segment to improve product promotion indicators such as the click-through rate and conversion rate of the product link information.

[0089] According to the embodiments of the present disclosure, the response information contained in the virtual character interaction control may include any form such as text, audio, image, virtual image action, resource connection, etc., so that the virtual character interaction control can support multimodal interaction methods such as text, voice, video, audio, etc., flexibly interact with the target object in various scenarios, and enhance the information interaction experience of the target object.

[0090] According to an embodiment of the present disclosure, the video generation method based on the big model may also include: playing the interactive video in the operation interface based on the interactive arrangement logic; obtaining the interactive information corresponding to the input request in response to the input request for the dialogue interaction window in the interactive video; using the big model to process the interactive information to generate response information corresponding to the expression style of the video character; and displaying the response information in the dialogue interaction window.

[0091] According to an embodiment of the present disclosure, the interactive video played in the operation interface may include the original video played and the virtual character interaction controls displayed according to the interactive arrangement logic. The creator can preview the generated effect of the current interactive video by browsing the interactive video in the operation interface.

[0092] According to the embodiments of the present disclosure, creators can also generate input requests by inputting interactive information into the dialogue interaction window of the interactive video being previewed. They can then simulate the target object's information interaction based on the input interactive information to test the interactive video playback effect, and promptly view the response information through the dialogue interaction window. This allows creators to test the generated interactive video based on interactive information such as the triggering moment of the virtual character interactive control and the information content of the response information fed back by the current operation interface, and can promptly adjust the interactive arrangement logic to improve the interactive video playback effect or information interaction effect.

[0093] According to an embodiment of the present disclosure, the video generation method based on the large model may also include: displaying the virtual character interaction controls in the current interaction choreography logic in the interaction editing box of the operation interface; and updating at least one virtual character interaction control in the current interaction choreography logic in response to an editing request for the interaction editing box.

[0094] According to an embodiment of the present disclosure, the creator can browse the display effects, triggering moments, triggering logic and other properties of the virtual character interaction controls in the interactive editing box, and perform editing operations on the interactive editing box based on the creator's personalized needs, thereby generating an editing request, and using the configuration parameters carried in the editing request to update the properties of the virtual character interaction space, thereby achieving flexible addition, deletion, replacement, and modification of virtual character interaction controls, thereby generating new interactive choreography logic that can meet the creator's needs.

[0095] According to an embodiment of the present disclosure, the video generation method based on the big model may also include: responding to an update request for the interactive editing box of the operation interface, obtaining update requirement information corresponding to the update request; and using the big model to process the update requirement information so as to update the current interactive arrangement logic and obtain a new interactive arrangement logic.

[0096] According to embodiments of the present disclosure, update requests can be generated based on the creator's interactive operations on the user interface. Creators can describe the content that requires modification to the current interactive choreography logic based on their personalized update requirements, thereby implementing update request information. Update request information can include any type of information, such as images and videos. Update request information can also include any content such as updating the avatar image, updating the expression style, and updating the trigger time.

[0097] According to the embodiments of the present disclosure, the large model uses the update requirement information as a new prompt word to process the update requirement information and video-related information, thereby updating the current interactive arrangement logic and obtaining a new interactive arrangement logic. This allows creators to update the interactive arrangement logic by communicating with the operation interface, thereby generating a new interactive video, thereby improving the creator's interactive video production efficiency and the flexibility of modifying interactive videos.

[0098] According to an embodiment of the present disclosure, using a large model to process interactive information and generate response information corresponding to the expression style of a video character may include: performing emotion recognition on the interactive information to obtain an emotion recognition result; using the large model to process the interactive information and emotional prompt words determined based on the emotion recognition result to obtain response information that matches the emotion recognition result.

[0099] According to the embodiments of the present disclosure, emotion recognition can be performed on interactive information based on a deep learning model. For example, interactive information can be processed based on a natural language understanding model to generate emotion recognition results that characterize the target object (or creator) of the input interactive information, such as frustration or happiness. This allows semantic fusion to be performed by processing the semantic themes expressed in the interactive information and the emotional prompt words used to indicate the emotional type through a large model. This allows the generated response information to more accurately match the emotional type of the target object (or creator), thereby enabling interactive feedback on the emotional state of the target object (or creator), more realistically simulating information interaction between video characters and the target object, enabling deep immersive video interaction with the target object (or creator), and achieving emotional resonance with the target object (or creator).

[0100] In one example, the emotion recognition result may be a depressed emotion type, and the corresponding response information may be a response voice with a soothing expression style.

[0101] In one example, the emotion recognition result may be a happy emotion type, and the corresponding response information may be a response voice with an excited expression style.

[0102] According to embodiments of the present disclosure, the timing of an input request matches the presentation of a key plot segment in the original video. For example, a creator may input interactive information while viewing a key plot segment, thereby generating an input request. Response information corresponding to the input request may be displayed during the playback of the key plot segment, promptly addressing the creator's emotional needs.

[0103] It should be understood that the virtual character interaction controls displayed in the operation interface involved in the embodiments of the present disclosure can also be displayed on the display screen of the device of the target object watching the interactive video, so that the target object can achieve an immersive experience by interacting with information for the interactive video.

[0104] According to an embodiment of the present disclosure, using a large model to process interaction information and emotion prompt words determined based on emotion recognition results may include: using a large model to process interaction information, emotion prompt words, and key plot text determined based on key plot segments.

[0105] According to the embodiments of the present disclosure, the use of a large model to process interactive information, emotional prompt words and key plot texts can enable the large model to more accurately understand the progress of the current key plot and the emotional needs of the target object (or creator), and respond more accurately based on the expression style of the video character, so that the response information can more accurately match the emotional needs of the target object (or creator), thereby improving the browsing experience of the interactive video.

[0106] Figure 5 An application scenario diagram of a large model-based video generation method according to another embodiment of the present disclosure is shown.

[0107] like Figure 5 As shown, this application scenario may include an operation interface 500. The interactive editing box within operation interface 500 may include an arrangement logic box 510 and an interactive control box 520. Operation interface 500 may also include a preview box 530. Within arrangement logic box 510, the creator can use the "Global Settings" edit bar to perform interactive operations, enabling global configuration operations such as adding virtual characters, adding original videos, and adjusting the order of interactive controls. The current interactive arrangement logic may display the interactive arrangement logic for each of the three episodes in the original video, categorized by drop-down display boxes for "Episode 1," "Episode 2," and "Episode 3." The creator may select the drop-down display box corresponding to "Episode 3" to view the interactive arrangement logic. In response to the viewing operation, arrangement logic box 510 may display the virtual character interactive controls associated with "Episode 3" in the interactive arrangement logic, representing the video characters "Character A" and "Character B." The virtual character interactive controls may display response information based on the "soothing" expression style. Furthermore, the original video for Episode 3 may be free.

[0108] like Figure 5As shown, the interactive control box 520 can display operation elements such as "Add Role", "Role Association", "Interactive Content", and "Add Control", so that the creator can perform interactive operations on the interactive control box 520 to modify the virtual character interactive control in the current interactive arrangement logic. Text information can also be entered for operation elements such as "Question", "Option 1", "Option 2", and "Add Option" to modify the test content of the virtual character interactive control. The original video can be played in the preview box 530. The original video can include a first video character 531 and a second video character 532. The dialogue interaction window H530 of the virtual character interactive control can generate a first response message "Should I go on stage?" and display option 1 "Well, be bolder" and option 2 "Wait a minute". The creator can click on option 1 "Well, be bolder" to enter interactive information. The large model can process the interactive information "Well, be bolder" and the first response information "Should I go on stage?" and generate a second response message "Okay, okay..."

[0109] Figure 6 An application scenario diagram of a large model-based video generation method according to another embodiment of the present disclosure is shown.

[0110] like Figure 6 As shown, this application scenario may include an operation interface 600. The interactive editing box within operation interface 600 may include an arrangement logic box 610 and a character editing box 620. Operation interface 600 may also include a preview box 630. Within arrangement logic box 610, the creator can use the "Global Settings" edit bar to perform interactive operations, enabling global configuration operations such as adding virtual characters, adding original videos, and adjusting the order of interactive controls. The current interactive arrangement logic may display the interactive arrangement logic for each of the three episodes in the original video, categorized by drop-down display boxes for "Episode 1," "Episode 2," and "Episode 3." The creator may select the drop-down display box corresponding to "Episode 3" to view the interactive arrangement logic. In response to this, arrangement logic box 610 may display the virtual character interactive controls associated with "Episode 3" in the interactive arrangement logic, representing the video characters "Character A" and "Character B." The virtual character interactive controls may display response information based on the "Excited" expression style. Furthermore, the original video for Episode 3 may be free.

[0111] like Figure 6As shown, the character editing box 620 can display operation elements such as "name", "upload photo", "intelligent image generation", "task introduction", "opening remarks", "background file", "sound", and "image", so that the creator can perform interactive operations on the character editing box 620 to modify the virtual character interactive controls in the current interactive arrangement logic. For example, the video character can be modified based on the "name" operation element, the photo of the video character can be uploaded based on the "upload photo" operation element to generate the virtual character image, and the virtual character image can be automatically generated using a large model based on the "intelligent image generation" operation element. The character description information can be uploaded based on the "character introduction" operation element, the first response information can be modified based on the "opening remarks" operation element, the video description information can be uploaded based on the "background file" operation element, and the audio information related to the video character can be uploaded based on the "sound" operation element so that the response information can be pronounced according to the timbre of the video character. The prepared virtual character image can be directly uploaded based on the "image" operation element.

[0112] The preview box 630 can preview the response effect of the virtual character interaction control in the current interactive choreography logic in real time. For example, the virtual character in the virtual character interaction control can display the first response information "Should I go on stage?", and the creator can input the interaction information "Well, be bolder" into the input window 631. The large language model can process the first response information "Should I go on stage?" and the interaction information "Well, be bolder" to generate the second response information "But what if the performance goes wrong" to simulate information interaction with the target object, so as to facilitate the creator to quickly view the information interaction effect of the interactive choreography logic, and facilitate the creator to iteratively adjust the interactive choreography logic through timely feedback information, reduce the time cost of repeated testing, and improve the efficiency of video generation.

[0113] Figure 7 An application scenario diagram of a large model-based video generation method according to another embodiment of the present disclosure is shown.

[0114] like Figure 7As shown, this application scenario may include an operation interface 700. The interactive editing box within operation interface 700 may include an arrangement logic box 710 and a payment editing box 720. Operation interface 700 may also include a preview box 730. The arrangement logic box 710 allows creators to perform interactive operations through the "Global Settings" edit bar, enabling global configuration operations such as adding virtual characters, adding original videos, and adjusting the order of interactive controls. The current interactive arrangement logic may display the interactive arrangement logic for each of the three episodes in the original video, categorized by drop-down display boxes for "Episode 1," "Episode 2," and "Episode 3." Creators may view the drop-down display box corresponding to "Episode 3." In response to this, arrangement logic box 710 may display the interactive arrangement logic, indicating that the virtual character interactive controls associated with "Episode 3" represent "Character A" and "Character B," and that the virtual character interactive controls may display responses based on the "Excited" expression style. Furthermore, the original video for Episode 3 may be a paid video.

[0115] like Figure 7 As shown, the payment editing box 720 can display operation elements such as "trigger moment", "payment type", "pricing setting", "add role", "role association", "select role", and "recommendation method", so that the creator can perform interactive operations on the payment editing box 720 to modify the virtual character interaction control in the current interactive arrangement logic. For example, the trigger moment of the virtual character interaction control can be modified based on the "trigger moment" operation element, the interactive video of episode 3 can be modified to paid browsing or free browsing based on the "payment type" operation element, and the payment amount can be set based on the "pricing setting" operation element. The video role corresponding to the virtual character interaction control is added based on the "add role" operation element, the video role corresponding to the virtual character interaction control is determined based on the "role association" operation element, and the recommended expression style of the product link information of the virtual character interaction control is determined to be "affectionate" based on the "recommendation method" operation element. The preview box 730 can play a video containing the first video character 731 and the second video character 732. For a key plot clip of video character 732, the dialogue interaction window H730 of the avatar interaction control can display a first response message D730, "Please support us," based on the avatar image and expression style of the first video character 731. This first response message D730 can also include a product link for purchasing this episode. If the creator performs a transaction based on the product link, the dialogue interaction window H730 can display a second response message, "Thank you...", based on the avatar image and expression style of the first video character 731. This allows creators to quickly test whether the payment function of the interactive video is working properly.

[0116] According to an embodiment of the present disclosure, the dialogue interaction window may further include virtual character images corresponding to multiple video characters, for example, a first virtual character image corresponding to a first video character, and a second virtual character image corresponding to a second video character.

[0117] According to an embodiment of the present disclosure, using a large model to process interaction information and generate response information corresponding to the expression style of a video character may also include: using a large model to process interaction information and first response information related to the first virtual character image to obtain second response information related to the second virtual character image.

[0118] According to the embodiments of the present disclosure, the interaction information obtained in the current dialogue interaction window and the response information corresponding to each virtual character image can be processed based on the large model, thereby realizing information interaction between virtual character images, and information interaction between multiple virtual character images and the target object.

[0119] According to an embodiment of the present disclosure, the first response information is related to the expression style of the first virtual character, and the second response information is related to the expression style of the second virtual character. This allows the interactive dialogue window to simulate a group chat interaction between the target subject and the first and second video characters, thereby enhancing the target subject's interactive experience and improving immersion.

[0120] Figure 8 A block diagram of a large model-based video generation device according to an embodiment of the present disclosure is shown.

[0121] like Figure 8 As shown, the video generation device 800 based on the large model includes: a first acquisition module 810, an interactive arrangement logic acquisition module 820 and a generation module 830.

[0122] The first acquisition module 810 is configured to acquire video-related information in response to a video production request for an operation interface, where the video-related information includes the original video.

[0123] The interactive arrangement logic obtaining module 820 is used to process the video-related information using the large model to obtain the interactive arrangement logic. The interactive arrangement logic represents the arrangement results for the interactive controls, which are used for information interaction.

[0124] The generation module 830 is used to generate an interactive video based on the interactive arrangement logic and the original video. The interactive video is used to display the video content of the original video during playback and to interact with the target object based on the interactive arrangement logic.

[0125] According to an embodiment of the present disclosure, the interactive choreography logic obtaining module includes: a first processing submodule and a virtual character interactive control obtaining submodule.

[0126] The first processing submodule is used to process the original video using the large model to obtain role description information of the video characters involved in the original video.

[0127] The virtual character interaction control acquisition submodule is used to generate virtual character interaction controls corresponding to the video character based on the character description information, wherein the interaction arrangement logic includes the virtual character interaction controls, and the virtual character interaction controls are used to interact with the target object based on the expression style of the video character in the original video.

[0128] According to an embodiment of the present disclosure, the video-related information further includes video description information for the original video.

[0129] According to an embodiment of the present disclosure, the interactive choreography logic obtaining module includes: a second processing submodule and a virtual character interactive control obtaining submodule.

[0130] The second processing submodule is used to process the video description information using the large model to obtain role description information of the video characters involved in the original video.

[0131] The virtual character interaction control acquisition submodule is used to generate virtual character interaction controls corresponding to the video character based on the character description information, wherein the interaction arrangement logic includes the virtual character interaction controls, and the virtual character interaction controls are used to interact with the target object based on the expression style of the video character in the original video.

[0132] According to an embodiment of the present disclosure, the virtual character interaction control obtaining submodule includes a virtual character interaction control obtaining unit.

[0133] The virtual character interaction control acquisition unit is used to use the large model to process the original video and style prompt words to obtain the virtual character interaction control, wherein the style prompt words are used to control the large model to understand the expression style of the video character, and the style prompt words are obtained by semantic recognition of the character description information.

[0134] According to an embodiment of the present disclosure, the video-related information further includes video description information corresponding to the original video.

[0135] According to an embodiment of the present disclosure, the interactive choreography logic obtaining module includes a triggering moment obtaining submodule.

[0136] The trigger moment acquisition submodule is used to use the large model to process plot prompt words and at least one of the original video and video description information to obtain the trigger moment for controlling the display and / or exit of the virtual character interactive controls, wherein the plot prompt words are obtained by identifying key plot segments in the video description information, and the trigger moment is related to the display period involving the key plot segments in the original video.

[0137] According to an embodiment of the present disclosure, the virtual character interaction control performs information interaction with the target object based on the expression style of the video character involved in the key plot segment.

[0138] According to an embodiment of the present disclosure, the key plot segment includes a product display segment displaying the target product, the plot prompt words include product plot prompt words used to control the large model to associate the target product, and the virtual character interaction control is also used to display product link information related to the target product.

[0139] According to an embodiment of the present disclosure, the video description information includes at least one of the following: role introduction information of a video character in the original video; a video script of the original video; and a role image related to the video character.

[0140] According to an embodiment of the present disclosure, the virtual character interaction control of the interactive choreography logic includes a dialogue interaction window, which is used to obtain interaction information input by the target object and display response information related to the interaction information based on the expression style of the video character involved in the original video.

[0141] According to an embodiment of the present disclosure, the video generation device based on the large model further includes: a playback module, an interaction information acquisition module, a response information generation module and a response information display module.

[0142] The playback module is used to play interactive videos in the operation interface based on the interactive arrangement logic.

[0143] The interaction information acquisition module is used to respond to an input request for the dialogue interaction window in the interactive video and obtain interaction information corresponding to the input request.

[0144] The response information generation module is used to use the large model to process the interaction information and generate response information corresponding to the expression style of the video character.

[0145] The response information display module is used to display the response information in the dialogue interaction window.

[0146] According to an embodiment of the present disclosure, the response information generation module includes: an emotion recognition submodule and a response information acquisition submodule.

[0147] The emotion recognition submodule is used to perform emotion recognition on the interactive information and obtain emotion recognition results.

[0148] The response information acquisition submodule is used to use the large model to process the interaction information and the emotional prompt words determined based on the emotion recognition results to obtain response information that matches the emotion recognition results.

[0149] According to an embodiment of the present disclosure, the request time of the input request matches the presentation time of the key plot segment in the original video.

[0150] According to an embodiment of the present disclosure, the response information obtaining submodule includes a key plot text obtaining unit.

[0151] The key plot text acquisition unit is used to use the large model to process interactive information, emotional prompt words and key plot text determined based on key plot fragments.

[0152] According to an embodiment of the present disclosure, the dialogue interaction window includes virtual character images corresponding to multiple video characters.

[0153] According to an embodiment of the present disclosure, the response information generating module further includes a second response information obtaining submodule.

[0154] The second response information obtaining submodule is used to use the large model to process the interaction information and the first response information related to the first virtual character image to obtain the second response information related to the second virtual character image, wherein the multiple virtual character images include the first virtual character image and the second virtual character image, the first response information is related to the expression style of the first virtual character image, and the second response information is related to the expression style of the second virtual character image.

[0155] According to an embodiment of the present disclosure, the video generation device based on the large model further includes: an update requirement information acquisition module and a first update module.

[0156] The update requirement information acquisition module is used to respond to an update request for the interactive edit box of the operation interface and acquire update requirement information corresponding to the update request.

[0157] The first update module is used to use the large model to process the update requirement information so as to update the current interactive arrangement logic and obtain a new interactive arrangement logic.

[0158] According to an embodiment of the present disclosure, the video generation device based on the large model further includes: a virtual character interactive control display module and a second update module.

[0159] The virtual character interaction control display module is used to display the virtual character interaction controls in the current interaction arrangement logic in the interaction editing box of the operation interface.

[0160] The second updating module is configured to update at least one virtual character interaction control in the current interaction arrangement logic in response to an editing request for the interaction editing box.

[0161] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0162] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a video generation method based on a large model.

[0163] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute a large model-based video generation method.

[0164] According to an embodiment of the present disclosure, a computer program product includes a computer program, which implements a large model-based video generation method when executed by a processor.

[0165] Figure 9 A block diagram of an electronic device suitable for implementing a large model-based video generation method according to an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0166] like Figure 9 As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. Computing unit 901, ROM 902, and RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to bus 904.

[0167] Various components in the device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0168] The computing unit 901 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as the large-model-based video generation method. For example, in some embodiments, the large-model-based video generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the large-model-based video generation method described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the large-model-based video generation method by any other suitable means (e.g., via firmware).

[0169] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0170] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0171] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0172] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0173] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0174] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0175] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0176] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A video generation method based on a large model, comprising: In response to a video production request for an operation interface, obtaining video-related information including an original video and video description information corresponding to the original video; Processing the video-related information using a large model to obtain interactive arrangement logic, wherein the interactive arrangement logic represents an arrangement result for an interactive control, wherein the interactive control is used for information interaction; as well as An interactive video is generated based on the interactive arrangement logic and the original video, wherein the interactive video is used to display the video content of the original video during playback and to perform information interaction with the target object based on the interactive arrangement logic. The interactive arrangement logic is obtained as follows: Obtaining role description information of video characters involved in the original video by processing the original video or the video description information using the large model; and By processing the character description information based on the large model, a virtual character interaction control corresponding to the video character is generated, so that the virtual character interaction control interacts with the target object based on the expression style of the video character in the original video. The virtual character interactive control is included in the interactive arrangement logic, and is obtained by configuring the attributes and executing the logic arrangement of the interactive control to be arranged based on the large model processing the video related information, and includes a virtual character image corresponding to the video character, and The virtual character image is obtained by processing video frames related to the video character or character images related to the video character in the original video based on a virtual character modeling technology.

2. The method according to claim 1, wherein Generating a virtual character interaction control corresponding to the video character includes: The large model is used to process the original video and style prompt words to obtain the virtual character interaction control, wherein the style prompt words are used to control the large model to understand the expression style of the video character, and the style prompt words are obtained by semantically recognizing the character description information.

3. The method according to claim 1, wherein The interactive arrangement logic further includes: The large model is used to process plot prompt words, as well as at least one of the original video and the video description information, to obtain a trigger moment for controlling the display and / or exit of the virtual character interactive control, wherein the plot prompt words are obtained by identifying key plot segments in the video description information, and the trigger moment is related to the display period involving the key plot segments in the original video.

4. The method according to claim 3, wherein: The virtual character interaction control performs information interaction with the target object based on the expression style of the video character involved in the key plot segment.

5. The method according to claim 3 or 4, wherein: The key plot segment includes a product display segment displaying a target product, the plot prompt words include product plot prompt words for controlling the large model to associate with the target product, and the virtual character interaction control is further used to display product link information related to the target product.

6. The method according to claim 1, wherein The video description information includes at least one of the following: Role introduction information of the video characters in the original video; a video script of the original video; A character image associated with the video character.

7. The method according to claim 1, wherein The virtual character interaction control of the interactive choreography logic includes a dialogue interaction window, which is used to obtain the interaction information input by the target object and display response information related to the interaction information based on the expression style of the video character involved in the original video.

8. The method according to claim 7, wherein: The method further comprises: Based on the interactive arrangement logic, playing the interactive video in the operation interface; In response to an input request for the dialogue interaction window of the interactive arrangement logic in the interactive video, obtaining interaction information corresponding to the input request; Processing the interaction information using the large model to generate response information corresponding to the expression style of the video character; and The response information is displayed in the dialogue interaction window.

9. The method according to claim 8, wherein Generating response information corresponding to the expression style of the video character includes: Performing emotion recognition on the interaction information to obtain an emotion recognition result; The interaction information and the emotion prompt words determined based on the emotion recognition result are processed using the large model to obtain response information matching the emotion recognition result.

10. The method according to claim 9, wherein: The request time of the input request matches the presentation time of the key plot segment in the original video; The response information that matches the emotion recognition result includes: The interactive information, the emotional prompt words, and the key plot text determined based on the key plot segments are processed using the large model to obtain the response information.

11. The method according to any one of claims 8 to 10, wherein The dialogue interaction window includes a plurality of virtual character images corresponding to a plurality of video characters; The step of generating response information corresponding to the expression style of the video character includes: The large model is used to process the interaction information and first response information related to the first virtual character image to obtain second response information related to the second virtual character image, wherein the multiple virtual character images include the first virtual character image and the second virtual character image, the first response information is related to the expression style of the first virtual character image, and the second response information is related to the expression style of the second virtual character image.

12. The method according to any one of claims 8 to 10, wherein The method further comprises: In response to an update request for the interactive edit box of the operation interface, obtaining update requirement information corresponding to the update request; and The update requirement information is processed using the large model to update the current interactive arrangement logic and obtain a new interactive arrangement logic.

13. The method according to claim 1, wherein The method further comprises: Displaying the virtual character interaction controls in the current interaction arrangement logic in the interaction editing box of the operation interface; and In response to an edit request for the interactive edit box, at least one virtual character interactive control in the current interactive choreography logic is updated.

14. A video generation device based on a large model, comprising: A first acquisition module is configured to, in response to a video production request for an operation interface, acquire video related information including an original video and video description information corresponding to the original video; An interactive arrangement logic acquisition module is used to process the video-related information using a large model to obtain interactive arrangement logic, wherein the interactive arrangement logic represents an arrangement result for an interactive control, and the interactive control is used for information interaction; as well as A generation module is used to generate an interactive video based on the interactive arrangement logic and the original video, wherein the interactive video is used to display the video content of the original video during playback and to interact with the target object based on the interactive arrangement logic. The interactive arrangement logic obtaining module includes: a processing submodule, configured to obtain role description information of video characters involved in the original video by processing the original video or the video description information using the large model; and A virtual character interaction control acquisition submodule is configured to generate a virtual character interaction control corresponding to the video character by processing the character description information based on the large model, so that the virtual character interaction control interacts with the target object based on the expression style of the video character in the original video. The virtual character interactive control is included in the interactive arrangement logic, and is obtained by configuring the attributes and executing the logic arrangement of the interactive control to be arranged based on the large model processing the video related information, and includes a virtual character image corresponding to the video character, and The virtual character image is obtained by processing video frames related to the video character or character images related to the video character in the original video based on a virtual character modeling technology.

15. The device according to claim 14, wherein The virtual character interaction control acquisition submodule includes: The virtual character interaction control acquisition unit is used to use the large model to process the original video and style prompt words to obtain the virtual character interaction control, wherein the style prompt words are used to control the large model to understand the expression style of the video character, and the style prompt words are obtained by semantically recognizing the character description information.

16. The device according to claim 14, wherein The interactive arrangement logic obtaining module further includes: A trigger moment acquisition submodule is used to use the large model to process plot prompt words, as well as at least one of the original video and the video description information, to obtain a trigger moment for controlling the display and / or exit of the virtual character interactive control, wherein the plot prompt words are obtained by identifying key plot segments in the video description information, and the trigger moment is related to the display period involving the key plot segments in the original video.

17. The device according to claim 16, wherein The virtual character interaction control performs information interaction with the target object based on the expression style of the video character involved in the key plot segment.

18. The device according to claim 16 or 17, wherein The key plot segment includes a product display segment displaying a target product, the plot prompt words include product plot prompt words for controlling the large model to associate with the target product, and the virtual character interaction control is further used to display product link information related to the target product.

19. The device according to claim 14, wherein The video description information includes at least one of the following: Role introduction information of the video characters in the original video; a video script of the original video; A character image associated with the video character.

20. The apparatus according to claim 14, wherein The virtual character interaction control of the interactive choreography logic includes a dialogue interaction window, which is used to obtain the interaction information input by the target object and display response information related to the interaction information based on the expression style of the video character involved in the original video.

21. The device according to claim 20, wherein The device further comprises: A playback module, configured to play the interactive video in the operation interface based on the interactive arrangement logic; an interaction information acquisition module, configured to respond to an input request of the dialogue interaction window of the interaction arrangement logic in the interactive video and acquire interaction information corresponding to the input request; a response information generating module, configured to process the interaction information using the large model and generate response information corresponding to the expression style of the video character; and The response information display module is used to display the response information in the dialogue interaction window.

22. The device according to claim 21, wherein The response information generation module includes: An emotion recognition submodule, configured to perform emotion recognition on the interaction information and obtain an emotion recognition result; The response information acquisition submodule is used to use the large model to process the interaction information and the emotion prompt words determined based on the emotion recognition result to obtain response information matching the emotion recognition result.

23. The device according to claim 22, wherein The request time of the input request matches the presentation time of the key plot segment in the original video; The response information obtaining submodule is further configured to utilize the large model to process the interaction information, the emotional prompt words, and the key plot text determined based on the key plot segments to obtain the response information.

24. The device according to any one of claims 21 to 23, wherein The dialogue interaction window includes a plurality of virtual character images corresponding to a plurality of video characters; The response information generation module includes: The second response information obtaining submodule is used to use the large model to process the interaction information and the first response information related to the first virtual character image to obtain the second response information related to the second virtual character image, wherein the multiple virtual character images include the first virtual character image and the second virtual character image, the first response information is related to the expression style of the first virtual character image, and the second response information is related to the expression style of the second virtual character image.

25. The device according to any one of claims 21 to 23, wherein The device further comprises: an update requirement information acquisition module, configured to, in response to an update request for the interactive edit box of the operation interface, acquire update requirement information corresponding to the update request; and The first updating module is used to process the update requirement information using the large model so as to update the current interactive arrangement logic and obtain a new interactive arrangement logic.

26. The apparatus according to claim 14, wherein The device further comprises: A virtual character interaction control display module, configured to display the virtual character interaction controls in the current interaction arrangement logic in the interaction editing box of the operation interface; and The second updating module is configured to update at least one virtual character interaction control in the current interaction arrangement logic in response to an editing request for the interaction editing box.

27. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 13.

28. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 13.

29. A computer program product comprising a computer program, wherein The computer program implements the method according to any one of claims 1 to 13 when executed by a processor.

Citation Information

Patent Citations

  • Information interaction method and device, electronic equipment and storage medium

    CN116680376A

  • Video-based interaction method and device, equipment, storage medium and program product

    CN116841436A