Video generation method and apparatus, medium, and device

By receiving video description text to generate target text and controlling virtual character actions, virtual scene videos are automatically generated, solving the problem of high complexity in video generation in the prior art and achieving efficient automatic video generation.

WO2025139163A1PCT designated stage expired Publication Date: 2025-07-03BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Application Number
PCT/CN2024/122727
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-26
Filing Date
2024-09-30
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the prior art, the video generation process in a virtual scene is complicated, and users need to manually edit scripts, materials and video processing, resulting in cumbersome operations and inefficient efficiency.

Method used

By receiving the video description text, a target text containing the virtual scene and the virtual character interaction text is generated, a control instruction sequence is generated based on the interactive text, and a virtual character action is recorded through the virtual camera to automatically generate the target video.

Benefits of technology

The video generation process is simplified, the user operation complexity is reduced, the generation efficiency and automation level is improved, and the manual workload is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024122727_03072025_PF_FP_ABST
    Figure CN2024122727_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present invention provide a video generation method and apparatus, a medium, and a device. The method comprises: receiving a video description text, and generating a target text on the basis of the video description text, wherein the target text comprises at least one sub-text, each sub-text corresponds to one virtual scene, and the sub-text comprises a virtual character in the virtual scene and an interactive text corresponding to the virtual character; on the basis of the interactive text in the sub-text, generating a control instruction corresponding to the virtual character in the sub-text, and on the basis of the control instruction corresponding to each sub-text, generating a control instruction sequence corresponding to the target text; and controlling the virtual character to execute a control instruction in the control instruction sequence, and by means of a virtual camera, recording frames obtained by executing the control instruction by the virtual character to obtain a target video corresponding to the video description text.
Need to check novelty before this filing date? Find Prior Art

Description

Video generation method, device, medium and equipment

[0001] This application claims priority to Chinese Patent Application No. 202311814363.4 filed on December 26, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field

[0002] Embodiments of the present disclosure relate to a video generation method, apparatus, medium, and device. Background Art

[0003] Video creation and sharing are increasingly becoming a new way for multiple users to interact. For example, users can create narrative videos using virtual characters and environments within virtual scenes, recreating famous scenes from novels or films, or generating original video plots.

[0004] In related technologies, when generating videos in a virtual scene, users mainly obtain video materials by playing and interacting with virtual characters in the virtual scene and recording videos, and then edit the video materials through video editing to obtain corresponding videos.

[0005] Summary of the Invention

[0006] This summary is provided to briefly introduce concepts that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0007] In a first aspect, the present disclosure provides a video generation method, the method comprising:

[0008] Receive a video description text, and generate a target text based on the video description text, wherein the target text includes at least one subtext, each subtext corresponds to a virtual scene, and the subtext includes a virtual character in the virtual scene and an interactive text corresponding to the virtual character;

[0009] Based on the interactive text in the sub-text, generating control instructions corresponding to the virtual character in the sub-text, and generating a control instruction sequence corresponding to the target text based on the control instructions corresponding to each sub-text;

[0010] The virtual character is controlled to execute the control instructions in the control instruction sequence, and a picture obtained by the virtual character executing the control instructions is recorded through a virtual camera to obtain a target video corresponding to the video description text.

[0011] In a second aspect, the present disclosure provides a video generation device, the device comprising:

[0012] A first generating module is configured to receive a video description text and generate a target text based on the video description text, wherein the target text includes at least one subtext, each subtext corresponds to a virtual scene, and the subtext includes a virtual character in the virtual scene and an interactive text corresponding to the virtual character;

[0013] A second generating module is configured to generate control instructions corresponding to the virtual character in the sub-text based on the interactive text in the sub-text, and to generate a control instruction sequence corresponding to the target text based on the control instructions corresponding to each sub-text;

[0014] The processing module is used to control the virtual character to execute the control instructions in the control instruction sequence, and record the pictures obtained by the virtual character executing the control instructions through a virtual camera to obtain the target video corresponding to the video description text.

[0015] In a third aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, which implements the steps of the method described in the first aspect when executed by a processing device.

[0016] In a fourth aspect, the present disclosure provides an electronic device, comprising:

[0017] a storage device having a computer program stored thereon;

[0018] A processing device is used to execute the computer program in the storage device to implement the steps of the method described in the first aspect.

[0019] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale. In the drawings:

[0021] FIG1 is a flow chart of a video generation method according to an embodiment of the present disclosure;

[0022] FIG2 is a schematic diagram of a process for generating a target text based on the video description text according to an embodiment of the present disclosure;

[0023] FIG3 is a flow chart of a video generation method according to an embodiment of the present disclosure;

[0024] FIG4 is a block diagram of a video generating apparatus according to an embodiment of the present disclosure; and

[0025] FIG5 shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION

[0026] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0027] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0028] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0029] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0030] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0031] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0032] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0033] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0034] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0035] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0036] At the same time, it is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0037] FIG1 is a flowchart of a video generation method according to an embodiment of the present disclosure. As shown in FIG1 , the method may include:

[0038] In step 11, a video description text is received, and a target text is generated based on the video description text. The target text includes at least one subtext, each of which corresponds to a virtual scene, and the subtext includes a virtual character in the virtual scene and an interactive text corresponding to the virtual character.

[0039] Among them, the video description text can be a brief description of the plot input by the user for video generation. For example, "Tom held a party", then in this step, the corresponding target text can be generated based on the video description text, that is, based on the brief description of the plot input by the user, the script is expanded and generated to obtain a target text with detailed scenes and interactive features. As an example, each sub-text can correspond to a virtual scene, and the virtual scene can be a scene in a preset scene library. The virtual scenes corresponding to different sub-texts can be the same, and each sub-text can be a script for representing a storyboard. The virtual character can be a role determined to interact based on the content in the video description text. Then, based on the video description text input by the user, a target text with a detailed feature description that conforms to the video plot can be generated, without the user having to manually edit the character's dialogue or operation and perform scene switching and settings, thus saving manual operations.

[0040] In step 12, based on the interactive text in the sub-text, control instructions corresponding to the virtual character in the sub-text are generated, and based on the control instructions corresponding to each sub-text, a control instruction sequence corresponding to the target text is generated.

[0041] The interactive text may include interactive operations performed by a virtual character, such as dialogue between the virtual character and the interactive text. This step can further determine the control instructions for the virtual character corresponding to each interactive text. When the virtual character is controlled to execute the control instructions, the virtual character can be driven to perform the corresponding interaction. If the target text contains multiple subtexts, the control instructions corresponding to each subtext are concatenated based on the order of the subtexts to obtain a control instruction sequence corresponding to the target text. This control instruction sequence then contains the control instructions for the virtual character in the entire target text.

[0042] In step 13, the virtual character is controlled to execute the control instructions in the control instruction sequence, and the images obtained by the virtual character executing the control instructions are recorded by a virtual camera to obtain a target video corresponding to the video description text.

[0043] Among them, the virtual camera (Virtual Camera) can be a rendering camera in a virtual scene, such as a virtual camera pre-configured in a game scene. Based on the virtual camera, the action scenes of the virtual character in the virtual scene can be displayed, and the action scenes of the virtual character can be further recorded based on the recording function of the virtual camera to obtain a video.

[0044] Controlling the virtual character to execute the control instructions in the control instruction sequence may involve driving the virtual character to complete the control instructions via a computer. A command executor may be configured based on an open API in an application corresponding to the virtual scene. The command executor may be mounted as a global service within the application for execution by the application's program. For example, the refresh rate of the command executor may be the game frame rate to ensure smooth video generation.

[0045] In this step, the virtual character is controlled to execute the control instructions, resulting in a video of the virtual character performing the corresponding actions. Thus, by controlling the virtual character to execute the control instruction sequence corresponding to the target text, the virtual character can perform corresponding interactive operations based on the content of the target text. The video corresponding to the virtual character executing the control instructions is then recorded using a virtual camera, resulting in a target video of the virtual character performing the plot of the target text.

[0046] As an example, the video recorded by the virtual camera can be exported as the target video. As another example, the interactive text can include the dialogue text of the virtual character, and corresponding subtitle information can be generated based on the dialogue text, and the display time of the subtitle information can be determined based on the execution time of the control instruction corresponding to the dialogue text. The recorded video can then be post-processed to add the subtitle information to the recorded video, and the post-processed video can be used as the target video. The way the subtitle information is displayed, such as the font, color, font size, position and other attributes, can be configured based on the actual application scenario, and this disclosure does not limit this.

[0047] Therefore, in the above technical solution, the text of the video description text can be automatically expanded and analyzed based on the content description in the video description text to obtain its corresponding target text. Afterwards, the control instructions of the virtual character can be generated based on the target text, so that the virtual character can be controlled to execute the control instructions so that the virtual character performs the plot according to the interactive content indicated by the interactive text. At the same time, combined with the virtual camera, recording is performed to obtain the target video that drives the virtual character to perform the content corresponding to the video description text. In this way, the complexity of video generation can be effectively reduced, user operations can be simplified, and the technical requirements of the user for the video generation process can be greatly reduced. In addition, there is no need for users to carry out complex processes such as script writing, material operation and recording, video editing and post-processing, etc., which effectively reduces the manual workload. In addition, the automation level and efficiency of video generation can be effectively improved.

[0048] In a possible embodiment, an exemplary implementation of generating the target text according to the video description text may include:

[0049] Based on the video description text and multiple preset virtual scenes, determine the virtual characters in the video description text and at least one outline text, wherein each outline text corresponds to a virtual scene, and the outline text includes the virtual characters in the virtual scene and a description text for representing the interactive plot in the virtual scene.

[0050] Among them, the description text of the interactive plot corresponding to the virtual scene may include interactive operations between multiple virtual characters in the virtual scene, and may also include interactive operations between virtual characters in the virtual scene and objects in the virtual scene. The description text can be a brief outline of the interactive operations performed by the virtual characters, which can be expressed through natural language text.

[0051] Multiple virtual scenes can be pre-set in the scene library to support video generation in a variety of virtual scenes, increasing the diversity of the generated videos. As an example, a video description text and multiple pre-set virtual scenes can be input into a language processing model to output virtual characters participating in the interaction in the video description text, as well as an outline text of the video description text. The language processing model can be implemented based on a large language model (LLM).

[0052] When determining the virtual character and at least one outline text in the video description text, the narrative structure in the outline text can be controlled by adding an example example in the input prompt word prompt. The example can include an example video description text and its corresponding example virtual character and example outline text, so that the language processing model can determine the virtual character and outline text in the input video description text based on the example.

[0053] Afterwards, for each outline text, an interactive text corresponding to the virtual character in the outline text is generated according to the outline text and the historical associated text corresponding to the outline text, wherein the interactive text includes a dialogue text and / or a narration text.

[0054] The order of the generated outline text can be used to represent the plot timeline corresponding to the video description text. The plot of the first outline text can be the historical plot of the later scene text outline. For example, if the generated outline text contains A1-A10, A1 is the plot of the first act, A1 is the historical plot of A2, A1 and A2 are the historical plot of A3, and so on. For A8, A1-A7 can all be its corresponding historical plot.

[0055] In this embodiment, the history-related text may be used to represent text corresponding to the historical plot associated with the current outline text.

[0056] In this step, the interactive text corresponding to the outline text can be generated based on the outline text and its corresponding historical associated text. The interactive text includes the dialogue text of the virtual character, as shown below:

[0057] Tom: Welcome to the party;

[0058] C1: Thank you.

[0059] The interactive text may also include narration text corresponding to the virtual character, for example: Tom and C1 walk to the dining table together.

[0060] As an example, based on the outline text and the historical associated text corresponding to the outline text, generating the interactive text corresponding to the virtual character in the outline text can be done by inputting the outline text and the historical associated text into the LLM model to obtain the interactive text corresponding to the outline text.

[0061] Therefore, in the above technical solution, an outline text and virtual characters can be first generated based on the video description text, thereby obtaining description texts for multiple storyboards, which briefly describe the storyboard plot. Then, based on the outline text and its corresponding historical context, interactive text containing detailed features corresponding to each outline text is iteratively generated. This eliminates the need for users to write scripts, effectively reducing the manual workload of users in the video creation process and streamlining the video generation process.

[0062] In a possible embodiment, the outline text has a sequence identifier, which is used to indicate the display position of the video generated based on the outline text. If the outline text contains A1-A10 as described above, the video can be generated in this order.

[0063] Accordingly, the method further comprises:

[0064] After the interactive text is generated, summary information corresponding to the interactive text is determined and stored.

[0065] After generating the interactive text, a summary can be generated based on commonly used summary generation methods in the art, which is not limited in this disclosure. For example, the summary information corresponding to each interactive text can be stored in a memory pool for easy playback based on this summary information. A schematic diagram of the process for generating target text based on the video description text is shown in Figure 2.

[0066] If the sequence identifier of the outline text indicates that the outline text is the first outline text, the history-related text corresponding to the outline text is empty.

[0067] For the outline text A1, its sequence identifier indicates that the outline text is the first outline text, and there is no other plot before it, so its corresponding historical association text can be a null value.

[0068] If the sequence identifier of the outline text indicates that the outline text is not the first outline text, matching is performed based on the summary information corresponding to the outline text and the stored interactive text.

[0069] Taking outline text A4 as an example, its sequence identifier indicates that the outline text is not the first outline text, and then the outline text can be matched with the summary information corresponding to the stored interactive text, that is, the outline text A4 is matched with A1, A2 and A3 respectively.

[0070] As an example, the summary information corresponding to the stored interactive text and the outline text are matched, and the description text in the outline text and the summary information are matched. As an example, the semantic text similarity corresponding to the description text and the summary information can be calculated, and the summary information in the summary information whose semantic text similarity with the description text exceeds a similarity threshold is used as the matched summary information. The semantic text similarity can be determined by an NLP (Natural Language Processing) model. When determining the semantic text similarity between the summary information and the description text, the summary information and the description text can be separated and divided, and the semantic text similarity is calculated based on the separated and divided sub-texts, and the summary information that exceeds the similarity threshold is used as the matched summary information, so that the summary information or part of the summary information of the interactive text can be matched.

[0071] If the summary information is matched, the interactive text corresponding to the previous outline text of the outline text and the matched summary information are used as the historical associated text corresponding to the outline text.

[0072] As an example, if the description text of A4 is matched with the summary information of the interactive texts corresponding to A1, A2, and A3 respectively, and it is determined that part of the summary information in A1 is matched, then the matched partial summary information in A1 and the interactive text corresponding to the previous outline text A3 of A4 can be used as the historical associated text corresponding to A4.

[0073] If no summary information is matched, the interactive text corresponding to the previous outline text of the outline text is used as the historical associated text corresponding to the outline text.

[0074] As an example, if the description text of A4 is matched with the summary information of the interaction texts corresponding to A1, A2, and A3 respectively, and no summary information is matched, the interaction text corresponding to A3, the previous outline text of A4, can be directly used as the historical associated text corresponding to A4.

[0075] Therefore, through the above technical solution, the interactive text that has been generated can be further generated into its summary information and stored. Then, when the outline text of the subsequent plot generates the interactive text, it can be matched from the stored summary information, so that the historical associated text related to the current outline text can be matched from the historical plot to assist the plot generation of the current outline text, improve the fineness of the generation of the interactive text, and at the same time improve the relevance and consistency between the video generation process and the historical plot, and improve the logical fluency of the video generation.

[0076] In a possible embodiment, the outline text further includes role description information of the virtual character in the virtual scene; accordingly, the method further includes:

[0077] Based on the character description information and preset character clothing components, a target clothing component corresponding to each of the virtual characters is determined, and the virtual characters are rendered based on the target clothing components.

[0078] As an example, for each virtual character, the appearance characteristics of the virtual character can be determined based on the character description information of the virtual character. As an example, the character description information can include the type of the virtual character. For example, if the virtual character C1 can be a company employee, then character clothing components corresponding to different types of virtual characters can be pre-set. For example, for company employees, a suit can be set. For virtual character C1, a suit can be selected from the pre-set character clothing components as the target clothing component corresponding to the virtual character. If multiple suits exist, one set can be randomly selected as the target clothing component, or matching can be performed based on other descriptions of the virtual character's appearance characteristics in the character description information.

[0079] As another example, the character description information may include a personality description and clothing description corresponding to the virtual character. As an example, the selection of character clothing components based on the personality description and clothing description may be implemented based on the LLM model. For example, the task prompts for the LLM model may be completed through a prompt project, so that the LLM model can select target clothing components for different virtual characters. For example, the personality description and clothing description as well as a plurality of preset character clothing components may be input into the LLM model, and the target clothing component may be determined based on the output of the model. Afterwards, the virtual character may be rendered based on the target clothing component, thereby configuring clothing corresponding to the target clothing component for the virtual character.

[0080] Therefore, through the above technical solution, the diversity of virtual characters in virtual scenes can be further improved, the comprehensiveness of the characteristics of virtual characters in the video generation process can be improved, the automatic configuration of virtual characters can be realized, and user operations can be effectively saved.

[0081] In a possible embodiment, the outline text further includes role description information of the virtual character in the virtual scene; accordingly, the method further includes:

[0082] A target sound feature corresponding to each of the virtual characters is determined based on the character description information and a plurality of preset sound features.

[0083] In the virtual scene, gender and age characteristics can be pre-configured for each virtual character, which can be a cartoon image or a human figure. Accordingly, after the virtual character is determined, its character description information can include characteristic information such as gender and age corresponding to the virtual character.

[0084] As an example, speech synthesis models corresponding to multiple timbres can be pre-configured, and the corresponding role attributes can be configured for each speech synthesis model. For example, speech synthesis models corresponding to girls and boys can be set separately, and different speech synthesis models can be set for young people, middle-aged people and the elderly under different genders.

[0085] After the virtual character is determined, a speech synthesis model corresponding to the virtual character can be determined based on the character description information corresponding to the virtual character, and the feature identifier corresponding to the speech synthesis model can be used as the target sound feature.

[0086] To improve the diversity of speech synthesis, multiple different speech synthesis models can be set for the same age group. For example, for virtual character C2, based on its corresponding character description information of female and young, if multiple speech synthesis models are determined, one can be randomly selected from the multiple models as its corresponding target sound feature. For virtual character C3, based on its corresponding character description information of female and young, if multiple speech synthesis models are determined, one can be randomly selected from the multiple models as its corresponding target sound feature, except for the speech synthesis model corresponding to C2. This allows different virtual characters to be synthesized using different sound features to a certain extent.

[0087] generating a dialogue voice corresponding to the virtual character according to the target sound feature and the dialogue text corresponding to the virtual character;

[0088] When controlling the virtual character to execute the control instruction, a dialogue voice corresponding to the control instruction and the virtual character is played.

[0089] As an example, the speech synthesis model can be implemented based on a common speech synthesis method in the art. For each virtual character, the dialogue text corresponding to the virtual character can be input into the speech synthesis model corresponding to the target sound feature, thereby generating the dialogue speech.

[0090] In this embodiment, if the control instruction is used to control the virtual character to speak, the dialogue voice of the virtual character corresponding to the control instruction is played during the execution of the control instruction, so that the virtual camera can record the dialogue voice at the same time when recording, without the need for audio insertion and processing in the later stage of the video, further simplifying the video generation process.

[0091] In a possible embodiment, before generating control instructions corresponding to the virtual character in the subtext based on the interactive text in the subtext, the method may further include:

[0092] Displaying the target text. In this embodiment, the generated target text can be displayed in the display interface so that the user can confirm whether the target text is consistent with the content to be created.

[0093] As an example, the user may confirm the target text through a confirmation operation, and then control instructions may be generated based on the confirmed target text to complete the video generation process.

[0094] As another example, based on their own creative ideas, the user can edit the displayed target script text to modify the target text. For example, the user can edit the target text by clicking the edit control 2, and in response to receiving the user's editing operation, the text obtained by the editing operation is used as the new target text.

[0095] In this embodiment, the text generated after the user's editing can be used as the new target text, and the subsequent control instructions can be generated based on this new target text. This can improve the consistency between the target text used to generate the control instructions and the user's video creation intention, thereby improving the accuracy of the subsequently generated video and enhancing the user's satisfaction with the target video.

[0096] In a possible embodiment, an exemplary implementation of generating a control instruction corresponding to the virtual character in the subtext based on the interactive text in the subtext may include:

[0097] Determine an interactive object in the virtual scene corresponding to the subtext.

[0098] Among them, objects that can interact with virtual characters in different virtual scenes can be pre-set, such as virtual objects such as tables, sofas, and televisions in virtual scenes. Virtual characters can interact with them to perform plot interpretations, such as virtual character A turning on the TV in the virtual scene.

[0099] According to the interactive object and the interactive text in the subtext, a control instruction corresponding to the virtual character in the subtext is generated.

[0100] As an example, control instructions corresponding to the virtual character can be pre-defined, such as speaking, moving, expression, action, object operation, etc., which can be set according to the actual application scenario, and this disclosure does not limit this.

[0101] In this step, the interactive objects and the interactive text in the subtext can be input into the LLM model to generate control instructions corresponding to the virtual characters. As an example, the control instructions can be represented by a two-dimensional list, where each element in the list is used to represent an instruction set at a time step, and the instruction set contains the instructions corresponding to each virtual object at that time step. For example, if the interactive text contains virtual characters C1 and C2, the instruction set at each time step contains the instructions corresponding to C1 and C2 at that time step. The instructions at multiple time steps are spliced ​​to obtain the instruction set.

[0102] Therefore, through the above technical solution, when determining the control instructions corresponding to the virtual character, the interactive objects in the virtual scene can be used to generate the control instructions, so as to interpret the plot based on the preset assets in the virtual scene and ensure the effectiveness of the control instructions of the virtual character.

[0103] In a possible embodiment, an exemplary implementation of generating a control instruction corresponding to the virtual character in the subtext based on the interactive object and the interactive text in the subtext may include:

[0104] The interactive text in the subtext is parsed to obtain the target narration text in the interactive text.

[0105] The target narration text includes at least one of the following:

[0106] a narration text at the beginning of the subtext;

[0107] the narration text at the end of the subtext;

[0108] A continuous narration text within the subtext.

[0109] As an example, in this step, the interactive text can be first parsed to determine the dialogue text and narration text in the interactive text. The text corresponding to the speaking virtual character in the interactive text is the dialogue text, and the remaining text is the narration text. The dialogue text and narration text are separated by delimiters in the interactive text. The dialogue text between two adjacent delimiters is considered as one dialogue text, and the narration text between two adjacent delimiters is considered as one narration text.

[0110] Afterwards, it is determined whether the text at the start position and the end position of the subtext is a narration text, that is, whether the first text and the last text in the interactive text are narration texts. If so, the narration text is determined to be the target narration text.

[0111] Furthermore, for the narration text outside the start position and the end position of the subtext, if there is a continuous narration text, the continuous narration text is used as the target narration text.

[0112] Based on the first input prompt mode, the target narration text and the interactive object are input into a narration processing model to obtain a first control instruction corresponding to the virtual character;

[0113] Based on a second input prompt method, the text in the interactive text except the target narration text and the interactive object are input into a dialogue processing model to obtain a second control instruction corresponding to the virtual character, wherein the first input prompt method and the second input prompt method are different.

[0114] The narration processing model and the dialogue processing model can both be implemented based on the LLM model. By setting different input prompt modes for the narration processing model and the dialogue processing model, respectively, a prompt word prompt is constructed based on the corresponding input prompt mode to constrain the model output and obtain the corresponding control instructions. For example, for the dialogue processing model, the prompt word prompt of its corresponding first input prompt mode can be preset as follows: Determine the control instructions based on the following text, requiring the selection of the virtual character's facial expressions, movements, and orientations to be accurate, rich, and vivid. For example, for the narration processing model, the prompt word prompt of its corresponding second input prompt mode can be preset as follows: Determine the control instructions based on the following text, requiring a certain degree of imagination to reflect the plot in the target narration text based on the interactive objects in the virtual scene, fully express the information in the target narration text, and at the same time not excessively generate a plot that exceeds the content of the target narration text.

[0115] Afterwards, the first control instruction and the second control instruction are spliced ​​according to the order of the texts in the interactive text to obtain the control instruction corresponding to the virtual character in the interactive text.

[0116] Therefore, through the above technical solution, control instructions can be generated for each sub-text, and during the generation process, the dialogue text and narration text of the virtual character can be processed accordingly to improve the accuracy of the obtained control instructions and provide effective support for the subsequent control of the virtual character to perform plot interpretation.

[0117] In a possible embodiment, before controlling the virtual character to execute the control instructions in the control instruction sequence, the method further includes:

[0118] Determining a camera movement mode corresponding to the control instruction;

[0119] The lens parameters of the virtual camera are adjusted based on the lens movement mode corresponding to the control instruction.

[0120] In order to increase the diversity of video recording shots, in this embodiment, multiple camera movement modes can be pre-set, and lens parameters corresponding to each camera movement mode can be configured, so that video generation can be performed based on a variety of shooting methods. Accordingly, before each control instruction is executed, the camera movement mode corresponding to the control instruction can be determined first, so that the lens parameters of the virtual camera can be adjusted based on the lens parameters corresponding to the camera movement mode. Then, the virtual character is controlled to execute the control instruction, and the virtual camera after the lens parameters are adjusted records the image formed by the virtual character executing the control instruction. The camera movement mode can be switched during the video generation process, and the richness of details in the generated video can be improved by simulating the real video recording process, simplifying the operation process of video post-processing.

[0121] In a possible embodiment, an exemplary implementation of determining the camera movement mode corresponding to the control instruction may include:

[0122] According to the instruction type corresponding to the control instruction, the instruction types corresponding to the preset multiple camera movement modes are queried.

[0123] As shown above, control instructions are a variety of instructions pre-set in the instruction executor, which can include the following types of instructions:

[0124] {SayTo,GoTo,GoSit,LookAt,Use,HandHeldAndUse,PlayAnimation,PlayExpression,Stroll,InitRegionIns,Follow,ChangeClothes,PlaySound}.

[0125] The instruction type can be set based on the actual application scenario, and this disclosure does not limit this.

[0126] Camera movement modes can include: single-person full-body, single-person half-body, single-person face profile, multi-person shot, and character follow-from-behind. For each camera movement mode, corresponding command types are pre-set. For example, the command types for single-person full-body can include {SayTo, GoTo, LookAt, Use, PlayAnimation}. The command types for single-person half-body can include {SayTo, LookAt, PlayAnimation}.

[0127] The camera movement mode corresponding to the instruction type corresponding to the control instruction is determined as the candidate camera movement mode corresponding to the control instruction.

[0128] If the instruction type of the control instruction K1 is SayTo, the camera movement mode corresponding to the SayTo type can be used as the candidate camera movement mode of the control instruction K1, that is, a single person's full body and a single person's half body can be used as the candidate camera movement modes corresponding to the control instruction K1.

[0129] Determine a camera movement mode corresponding to the control instruction based on the candidate camera movement modes.

[0130] For example, if there are multiple candidate camera movement methods, one can be randomly selected as the corresponding camera movement method. As another example, each camera movement method can be pre-assigned a corresponding weight. Samples can then be taken based on the weights of the candidate camera movement methods, and the sampled camera movement methods can be determined as the camera movement method corresponding to the control instruction. This can further improve the randomness and diversity of camera movement methods, thereby ensuring a diverse range of perspectives in the generated video, improving the automation of video generation, and enhancing user satisfaction with the generated videos.

[0131] In a possible embodiment, the method may further include:

[0132] Determining a sound effect corresponding to the control instruction according to an instruction type and instruction parameters corresponding to the control instruction;

[0133] When the operation of the virtual character is controlled based on the control instruction, a sound effect corresponding to the control instruction is played.

[0134] In this embodiment, a corresponding sound effect can be configured for a control command. For example, if a sound effect is required when controlling a virtual character to make a certain expression, the command type can be PlayExpression. For example, if the command parameter is a surprised expression, a surprise sound effect is configured; if the command parameter is a laughing expression, a haha ​​sound effect is configured. Another example is if a sound effect is required when controlling a virtual character to perform a certain action, the command type can be PlayAnimation. For example, if the command parameter is a bowing action, a bowing sound effect is configured; if the command parameter is a dancing action, a dancing sound effect is configured. Sound effects can be configured and modified according to specific application scenarios.

[0135] Accordingly, a sound effect query can be performed based on the command type and command parameters of the control command. If a sound effect is found, indicating that the control command has a corresponding sound effect, the sound effect corresponding to the control command will be played when the virtual character's operation is controlled based on the control command. If no sound effect is found, it means that the control command has no corresponding sound effect, and no processing is required. Figure 3 is a flow chart of a video generation method provided according to one embodiment of the present disclosure.

[0136] Therefore, through the above technical solution, corresponding sound effects can be configured for control instructions, so that the generated video can include the interaction of virtual characters while also including features of more dimensions, further improving the content richness of the generated video.

[0137] In a possible embodiment, the instruction type may include the type of play action. In order to increase the diversity of interactive operations of virtual characters in the generated video, in this embodiment, multiple action animations in the action library can be pre-generated based on video-to-motion. For example, the action video of a real person can be pre-recorded, and then the motion sequence of key points of the human skeleton can be obtained based on computer vision recognition technologies such as posture estimation, multi-instance segmentation of the human body, and ReID (Re-identification). The motion sequence of key points of the virtual character with a medium proportion can be estimated based on the forward-backward algorithm, thereby obtaining the action animation corresponding to the virtual character performing the action and storing it. Then, when it is determined that the instruction type of the virtual character is a play action, the corresponding action animation can be queried from the action library for playback, thereby increasing the diversity of actions that can be performed by the virtual character and improving the accuracy and efficiency of the execution of the control instructions.

[0138] Based on the same inventive concept, the present disclosure further provides a video generation device, as shown in FIG4 , wherein the device 10 includes:

[0139] A first generating module 100 is configured to receive a video description text and generate a target text based on the video description text, wherein the target text includes at least one subtext, each subtext corresponds to a virtual scene, and the subtext includes a virtual character in the virtual scene and an interactive text corresponding to the virtual character;

[0140] A second generating module 200 is configured to generate control instructions corresponding to the virtual character in the sub-text based on the interactive text in the sub-text, and to generate a control instruction sequence corresponding to the target text based on the control instructions corresponding to each sub-text;

[0141] The processing module 300 is used to control the virtual character to execute the control instructions in the control instruction sequence, and record the images obtained by the virtual character executing the control instructions through a virtual camera to obtain a target video corresponding to the video description text.

[0142] Optionally, the first generating module includes:

[0143] A first determining submodule is configured to determine, based on the video description text and a plurality of preset virtual scenes, a virtual character in the video description text and at least one outline text, wherein each outline text corresponds to a virtual scene, and the outline text includes a virtual character in the virtual scene and a description text for representing an interactive plot in the virtual scene;

[0144] The first generating submodule is configured to generate, for each outline text, interactive text corresponding to the virtual character in the outline text according to the outline text and the historical associated text corresponding to the outline text, wherein the interactive text includes dialogue text and / or narration text.

[0145] Optionally, the outline text has a sequence identifier; and the apparatus further comprises:

[0146] A first determining module is used to determine and store summary information corresponding to the interactive text after generating the interactive text;

[0147] A second determining module is configured to determine that if the sequence identifier of the outline text indicates that the outline text is the first outline text, then the history associated text corresponding to the outline text is empty;

[0148] The third determination module is used to match the outline text with the summary information corresponding to the stored interactive text if the sequence identifier of the outline text indicates that the outline text is not the first outline text; if the summary information is matched, the interactive text corresponding to the previous outline text of the outline text and the matched summary information are used as the historical associated text corresponding to the outline text; if the summary information is not matched, the interactive text corresponding to the previous outline text of the outline text is used as the historical associated text corresponding to the outline text.

[0149] Optionally, the outline text further includes role description information of the virtual character in the virtual scene; and the device further includes:

[0150] The fourth determining module is configured to determine a target clothing component corresponding to each virtual character based on the character description information and preset character clothing components, and render the virtual character based on the target clothing component.

[0151] Optionally, the outline text further includes role description information of the virtual characters in the virtual scene;

[0152] The device further comprises:

[0153] a fifth determining module, configured to determine a target sound feature corresponding to each of the virtual characters based on the character description information and a plurality of preset sound features;

[0154] a third generating module, configured to generate a dialogue voice corresponding to the virtual character based on the target sound feature and the dialogue text corresponding to the virtual character;

[0155] The first playing module is used to play the dialogue voice corresponding to the control instruction and the virtual character when controlling the virtual character to execute the control instruction.

[0156] Optionally, the device further comprises:

[0157] a display module, configured to display the target text before the second generation module generates control instructions corresponding to the virtual character in the sub-text based on the interactive text in the sub-text;

[0158] The sixth determining module is configured to, in response to receiving an editing operation from the user, use the text obtained by the editing operation as a new target text.

[0159] Optionally, the second generating module includes:

[0160] A second determining submodule is used to determine an interactive object in the virtual scene corresponding to the subtext;

[0161] The second generating submodule is configured to generate control instructions corresponding to the virtual character in the subtext according to the interactive object and the interactive text in the subtext.

[0162] Optionally, the second generating submodule includes:

[0163] a parsing submodule, configured to parse the interactive text in the subtext to obtain a target narration text in the interactive text;

[0164] A first processing submodule is configured to input the target narration text and the interactive object into a narration processing model based on a first input prompt mode, and obtain a first control instruction corresponding to the virtual character;

[0165] a second processing submodule, configured to input the text in the interactive text other than the target narration text and the interactive object into a dialogue processing model based on a second input prompt mode, to obtain a second control instruction corresponding to the virtual character, wherein the first input prompt mode and the second input prompt mode are different;

[0166] a splicing submodule, configured to splice the first control instruction and the second control instruction in the order of the respective texts in the interactive text, to obtain a control instruction corresponding to the virtual character in the interactive text;

[0167] The target narration text includes at least one of the following:

[0168] a narration text at the beginning of the subtext;

[0169] the narration text at the end of the subtext;

[0170] A continuous narration text within the subtext.

[0171] Optionally, the device further comprises:

[0172] A seventh determining module, configured to determine a camera movement mode corresponding to a control instruction before the processing module controls the virtual character to execute a control instruction in the control instruction sequence;

[0173] The adjustment module is used to adjust the lens parameters of the virtual camera based on the lens movement mode corresponding to the control instruction.

[0174] Optionally, the seventh determining module includes:

[0175] A query submodule, configured to query the command types corresponding to the preset multiple camera movement modes according to the command type corresponding to the control command;

[0176] a third determining submodule, configured to determine the camera movement mode corresponding to the instruction type corresponding to the control instruction as a candidate camera movement mode corresponding to the control instruction;

[0177] The fourth determining submodule is used to determine the camera movement mode corresponding to the control instruction based on the candidate camera movement modes.

[0178] Optionally, the device further comprises:

[0179] an eighth determining module, configured to determine a sound effect corresponding to the control instruction according to an instruction type and instruction parameters corresponding to the control instruction;

[0180] The second playing module is used to play the sound effect corresponding to the control instruction when controlling the virtual character to execute the control instruction.

[0181] Reference is now made to FIG5 , which illustrates a schematic diagram of the structure of an electronic device (e.g., a terminal device or server) 600 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device illustrated in FIG5 is merely an example and should not limit the functionality or scope of use of the embodiments of the present disclosure.

[0182] As shown in Figure 5, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0183] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 5 shows the electronic device 600 with various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may be implemented or present instead.

[0184] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0185] It should be noted that the computer-readable medium described above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0186] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0187] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0188] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: receives a video description text, generates a target text based on the video description text, wherein the target text contains at least one sub-text, each sub-text corresponds to a virtual scene, and the sub-text contains a virtual character in the virtual scene and an interactive text corresponding to the virtual character; based on the interactive text in the sub-text, generates a control instruction corresponding to the virtual character in the sub-text, and generates a control instruction sequence corresponding to the target text based on the control instructions corresponding to each sub-text; controls the virtual character to execute the control instructions in the control instruction sequence, and records the picture obtained by the virtual character executing the control instructions through a virtual camera to obtain a target video corresponding to the video description text.

[0189] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0190] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0191] The modules described in the embodiments of the present disclosure may be implemented in software or hardware. In some cases, the name of a module does not necessarily limit the module itself. For example, the first generation module may also be described as a "module that receives video description text and generates target text based on the video description text."

[0192] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0193] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0194] According to one or more embodiments of the present disclosure, Example 1 provides a video generation method, the method including:

[0195] Receive a video description text, and generate a target text based on the video description text, wherein the target text includes at least one subtext, each subtext corresponds to a virtual scene, and the subtext includes a virtual character in the virtual scene and an interactive text corresponding to the virtual character;

[0196] Based on the interactive text in the sub-text, generating control instructions corresponding to the virtual character in the sub-text, and generating a control instruction sequence corresponding to the target text based on the control instructions corresponding to each sub-text;

[0197] The virtual character is controlled to execute the control instructions in the control instruction sequence, and a picture obtained by the virtual character executing the control instructions is recorded through a virtual camera to obtain a target video corresponding to the video description text.

[0198] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1, wherein generating a target text according to the video description text includes:

[0199] Determining, based on the video description text and a plurality of preset virtual scenes, a virtual character in the video description text and at least one outline text, wherein each outline text corresponds to a virtual scene, and the outline text includes a virtual character in the virtual scene and a description text for representing an interactive plot in the virtual scene;

[0200] For each outline text, an interactive text corresponding to the virtual character in the outline text is generated according to the outline text and the historical associated text corresponding to the outline text, wherein the interactive text includes a dialogue text and / or a narration text.

[0201] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 2, wherein the outline text has a sequence identifier; and the method further includes:

[0202] After generating the interactive text, determining and storing summary information corresponding to the interactive text;

[0203] If the sequence identifier of the outline text indicates that the outline text is the first outline text, the history-related text corresponding to the outline text is empty;

[0204] If the sequence identifier of the outline text indicates that the outline text is not the first outline text, matching the outline text with summary information corresponding to the stored interactive text;

[0205] If the summary information is matched, the interactive text corresponding to the previous outline text of the outline text and the matched summary information are used as the historical associated text corresponding to the outline text;

[0206] If no summary information is matched, the interactive text corresponding to the previous outline text of the outline text is used as the historical associated text corresponding to the outline text.

[0207] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 2, wherein the outline text further includes character description information of a virtual character in the virtual scene;

[0208] The method further comprises:

[0209] Based on the character description information and preset character clothing components, a target clothing component corresponding to each of the virtual characters is determined, and the virtual characters are rendered based on the target clothing components.

[0210] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 2, wherein the outline text further includes character description information of a virtual character in the virtual scene;

[0211] The method further comprises:

[0212] Determining a target sound feature corresponding to each virtual character based on the character description information and a plurality of preset sound features;

[0213] generating a dialogue voice corresponding to the virtual character according to the target sound feature and the dialogue text corresponding to the virtual character;

[0214] When controlling the virtual character to execute the control instruction, a dialogue voice corresponding to the control instruction and the virtual character is played.

[0215] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 1, wherein, before generating the control instructions corresponding to the virtual character in the subtext based on the interactive text in the subtext, the method further includes:

[0216] displaying the target text;

[0217] In response to receiving the user's editing operation, the text obtained by the editing operation is used as the new target text.

[0218] According to one or more embodiments of the present disclosure, Example 7 provides the method of Example 1, wherein generating a control instruction corresponding to the virtual character in the subtext based on the interactive text in the subtext includes:

[0219] Determining an interactive object in a virtual scene corresponding to the subtext;

[0220] According to the interactive object and the interactive text in the subtext, a control instruction corresponding to the virtual character in the subtext is generated.

[0221] According to one or more embodiments of the present disclosure, Example 8 provides the method of Example 7, wherein generating a control instruction corresponding to the virtual character in the subtext based on the interactive object and the interactive text in the subtext includes:

[0222] Parsing the interactive text in the subtext to obtain a target narration text in the interactive text;

[0223] Based on the first input prompt mode, the target narration text and the interactive object are input into a narration processing model to obtain a first control instruction corresponding to the virtual character;

[0224] inputting the text in the interactive text other than the target narration text and the interactive object into a dialogue processing model based on a second input prompt mode to obtain a second control instruction corresponding to the virtual character, wherein the first input prompt mode and the second input prompt mode are different;

[0225] splicing the first control instruction and the second control instruction according to the order of the respective texts in the interactive text to obtain a control instruction corresponding to the virtual character in the interactive text;

[0226] The target narration text includes at least one of the following:

[0227] a narration text at the beginning of the subtext;

[0228] the narration text at the end of the subtext;

[0229] A continuous narration text within the subtext.

[0230] According to one or more embodiments of the present disclosure, Example 9 provides the method of Example 1, wherein, before controlling the virtual character to execute the control instruction in the control instruction sequence, the method further includes:

[0231] Determining a camera movement mode corresponding to the control instruction;

[0232] The lens parameters of the virtual camera are adjusted based on the lens movement mode corresponding to the control instruction.

[0233] According to one or more embodiments of the present disclosure, Example 10 provides the method of Example 9, wherein determining the camera movement mode corresponding to the control instruction includes:

[0234] According to the instruction type corresponding to the control instruction, query the instruction types corresponding to the preset multiple camera movement modes;

[0235] Determining the camera movement mode corresponding to the command type corresponding to the control command as a candidate camera movement mode corresponding to the control command;

[0236] Determine a camera movement mode corresponding to the control instruction based on the candidate camera movement modes.

[0237] According to one or more embodiments of the present disclosure, Example 11 provides the method of Example 1, wherein the method further includes:

[0238] Determining a sound effect corresponding to the control instruction according to an instruction type and instruction parameters corresponding to the control instruction;

[0239] When controlling the virtual character to execute the control instruction, a sound effect corresponding to the control instruction is played.

[0240] According to one or more embodiments of the present disclosure, Example 12 provides a video generating device, the device comprising:

[0241] A first generating module is configured to receive a video description text and generate a target text based on the video description text, wherein the target text includes at least one subtext, each subtext corresponds to a virtual scene, and the subtext includes a virtual character in the virtual scene and an interactive text corresponding to the virtual character;

[0242] A second generating module is configured to generate control instructions corresponding to the virtual character in the sub-text based on the interactive text in the sub-text, and to generate a control instruction sequence corresponding to the target text based on the control instructions corresponding to each sub-text;

[0243] The processing module is used to control the virtual character to execute the control instructions in the control instruction sequence, and record the pictures obtained by the virtual character executing the control instructions through a virtual camera to obtain the target video corresponding to the video description text.

[0244] According to one or more embodiments of the present disclosure, Example 13 provides a computer-readable medium having a computer program stored thereon, which implements the steps of the method described in any one of Examples 1-11 when executed by a processing device.

[0245] According to one or more embodiments of the present disclosure, Example 14 provides an electronic device, including:

[0246] a storage device having a computer program stored thereon;

[0247] A processing device is used to execute the computer program in the storage device to implement the steps of the method described in any one of Examples 1-11.

[0248] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in the present disclosure.

[0249] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0250] Although the subject matter has been described using language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. Regarding the apparatus in the above-described embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method and will not be elaborated upon here.

Claims

1. A video generation method, comprising: Receiving a video description text, and generating a target text according to the video description text, wherein the target text includes at least one sub - text, each sub - text corresponds to a virtual scene, and the sub - text includes virtual characters in the virtual scene and interaction texts corresponding to the virtual characters; Generating control instructions corresponding to the virtual characters in the sub - text based on the interaction texts in the sub - text, and generating a control instruction sequence corresponding to the target text based on the control instructions corresponding to each sub - text; Controlling the virtual characters to execute the control instructions in the control instruction sequence, and recording the images obtained by the virtual characters executing the control instructions through a virtual camera to obtain a target video corresponding to the video description text.

2. The video generation method according to claim 1, wherein, The generating the target text according to the video description text includes: Determining virtual characters and at least one outline text in the video description text according to the video description text and a plurality of preset virtual scenes, wherein each outline text corresponds to a virtual scene, and the outline text includes virtual characters in the virtual scene and description texts for representing interaction plots in the virtual scene; For each outline text, generating interaction texts corresponding to the virtual characters in the outline text according to the outline text and the historical associated text corresponding to the outline text, and the interaction texts include dialogue texts and / or narration texts.

3. The video generation method according to claim 2, wherein, The outline text has a sequence identifier; the video generation method further includes: After generating the interaction texts, determining and storing summary information corresponding to the interaction texts; If the sequence identifier of the outline text indicates that the outline text is the first outline text, the historical associated text corresponding to the outline text is empty; If the sequence identifier of the outline text indicates that the outline text is not the first outline text, then matching the outline text with the summary information corresponding to the stored interaction texts; If summary information is matched, using the interaction text corresponding to the previous outline text of the outline text and the matched summary information as the historical associated text corresponding to the outline text; If no summary information is matched, using the interaction text corresponding to the previous outline text of the outline text as the historical associated text corresponding to the outline text.

4. The video generation method according to claim 2 or 3, wherein, The outline text further includes character description information of the virtual characters in the virtual scene; The video generation method further includes: Determining a target clothing component corresponding to each virtual character based on the character description information and preset character clothing components, and rendering the virtual characters based on the target clothing components.

5. The video generation method according to claim 2 or 3, wherein, The outline text further includes character description information of the virtual characters in the virtual scene; The video generation method further includes: Determining a target voice feature corresponding to each virtual character according to the character description information and a variety of preset voice features; Generating dialogue voices corresponding to the virtual characters according to the target voice features and the dialogue texts corresponding to the virtual characters. When controlling the virtual character to execute the control instruction, play the dialogue voice corresponding to the control instruction and the virtual character.

6. The video generation method according to any one of claims 1-5, wherein, Before generating the control instruction corresponding to the virtual character in the sub-text based on the interactive text in the sub-text, the video generation method further includes: Display the target text; In response to receiving the user's editing operation, use the text obtained from the editing operation as the new target text.

7. The video generation method according to any one of claims 1-6, wherein, The generating of the control instruction corresponding to the virtual character in the sub-text based on the interactive text in the sub-text includes: Determine the interaction object in the virtual scene corresponding to the sub-text; Generate the control instruction corresponding to the virtual character in the sub-text according to the interaction object and the interactive text in the sub-text.

8. The video generation method according to claim 7, wherein, The generating of the control instruction corresponding to the virtual character in the sub-text according to the interaction object and the interactive text in the sub-text includes: Parse the interactive text in the sub-text to obtain the target narration text in the interactive text; Based on the first input prompt method, input the target narration text and the interaction object into the narration processing model to obtain the first control instruction corresponding to the virtual character; Based on the second input prompt method, input the text in the interactive text other than the target narration text and the interaction object into the dialogue processing model to obtain the second control instruction corresponding to the virtual character, wherein the first input prompt method and the second input prompt method are different; Splice the first control instruction and the second control instruction in the order of each text in the interactive text to obtain the control instruction corresponding to the virtual character in the interactive text; Wherein, the target narration text includes at least one of the following: The narration text at the start position of the sub-text; The narration text at the end position of the sub-text; The continuous narration text in the sub-text.

9. The video generation method according to any one of claims 1-8, wherein, Before controlling the virtual character to execute the control instruction in the control instruction sequence, the video generation method further includes: Determine the camera movement method corresponding to the control instruction; Adjust the lens parameters of the virtual camera based on the camera movement method corresponding to the control instruction.

10. The video generation method according to claim 9, wherein, The determining of the camera movement method corresponding to the control instruction includes: According to the instruction type corresponding to the control instruction, query the instruction types corresponding to multiple preset camera movement methods; Determine the camera movement method corresponding to the instruction type corresponding to the control instruction queried as the candidate camera movement method corresponding to the control instruction; Determine the camera movement method corresponding to the control instruction based on the candidate camera movement method.

11. The video generation method according to any one of claims 1-10 further includes: Determine the sound effect corresponding to the control instruction according to the instruction type and instruction parameters corresponding to the control instruction; Play the sound effect corresponding to the control instruction when controlling the virtual character to execute the control instruction.

12. A video generation device includes: A first generation module, configured to receive a video description text and generate a target text according to the video description text, wherein the target text includes at least one sub-text, each sub-text corresponding to a virtual scene, and the sub-text includes a virtual character in the virtual scene and interaction text corresponding to the virtual character; A second generation module, configured to generate a control instruction corresponding to the virtual character in the sub-text based on the interaction text in the sub-text, and generate a control instruction sequence corresponding to the target text based on the control instructions corresponding to each sub-text; A processing module, configured to control the virtual character to execute the control instructions in the control instruction sequence, and record the screen obtained by the virtual character executing the control instructions through a virtual camera to obtain a target video corresponding to the video description text.

13. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by a processing device, the steps of the video generation method according to any one of claims 1-11 are implemented.

14. An electronic device, comprising: A storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the video generation method according to any one of claims 1-11.

Citation Information

Patent Citations

  • Video generation method and device, electronic equipment and storage medium

    CN115396595A

  • Video generation method and device, electronic equipment and storage medium

    CN115942039A

  • Method for generating, evaluating and optimizing plot information based on artificial intelligence technology

    CN115970289A

  • Video generation method and device

    CN116684663A

  • Method and system for automatically generating story video in meta universe

    CN117177003A

Cited By

  • Short play video generation method and device, equipment, storage medium and program product

    CN121509781A

  • Short video generation method, device and equipment, storage medium and program product

    CN121509781B