Video generation method and device, live broadcast method and device, intelligent agent and electronic equipment

By obtaining and processing live scripts, combining the voice and lip rendering technology of virtual images, adjusting the movements and expressions of virtual images, the problem of insufficient expression of virtual images live broadcast is solved, and high-quality virtual image live broadcast is achieved.

CN120302121APending Publication Date: 2025-07-11BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510536564.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing virtual image live broadcast technology is difficult to achieve a high degree of unity between virtual image and elements such as explanation content, body movements, expressions, etc. in live broadcasts, and its expressions are poor.

Method used

By obtaining the script text and script tags in the live script, the virtual image is voiced and lip-reflective rendering is performed based on the script text, and the virtual image is adjusted in combination with the script tag to generate a more expressive target video.

Benefits of technology

It has improved the expressiveness of virtual images, improved the quality and interactive effect of virtual image live broadcast content, reduced the cost of live broadcast, and achieved stable live broadcast of 7×24 hours.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302121A_ABST
    Figure CN120302121A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method, a live broadcast method and device, an intelligent agent and electronic equipment, and relates to the technical field of artificial intelligence, in particular to the technical field of deep learning, natural language processing, computer vision, large models and digital live broadcast. According to the specific implementation scheme, a live broadcast script is obtained, wherein the live broadcast script comprises a script text and a plurality of script tags; reasoning and rendering the voice and lip movement of the virtual image based on the script text to obtain an initial video; and based on the plurality of script tags, adjusting at least one of actions, expressions and tones of the virtual image in the initial video to obtain a target video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technologies, particularly to the fields of deep learning, natural language processing, computer vision, large models, and digital live broadcast technologies. Specifically, it relates to a video generation method, a live broadcast method, a device, an intelligent agent, and an electronic device. Background Art

[0002] Virtual avatar live broadcast is a new form of live broadcast based on technologies such as artificial intelligence, computer graphics, and real-time rendering. It uses virtual digital humans to replace or assist real human hosts in real-time content output. Its core is to combine digitally generated virtual avatars with real human actions, voices, interactions, etc. to form a highly anthropomorphic live broadcast experience. Summary of the Invention

[0003] The present disclosure provides a video generation method, a live broadcast method, a device, an intelligent agent, and an electronic device.

[0004] According to one aspect of the present disclosure, there is provided a video generation method, including: obtaining a live broadcast script, where the live broadcast script includes script text and a plurality of script tags; based on the script text, performing inference rendering on the voice and lip movement of a virtual avatar to obtain an initial video; and based on the plurality of script tags, adjusting at least one of the actions, expressions, and tones of the virtual avatar in the initial video to obtain a target video.

[0005] According to another aspect of the present disclosure, there is provided a live broadcast method, including: adding a patch material in an initial live broadcast room to obtain a target live broadcast room; and playing a target video in the target live broadcast room, where the target video is generated by using the method described above.

[0006] According to another aspect of the present disclosure, there is provided a video generation device, including: a first obtaining module, configured to obtain a live broadcast script, where the live broadcast script includes script text and a plurality of script tags; an inference module, configured to perform inference rendering on the voice and lip movement of a virtual avatar based on the script text to obtain an initial video; and an adjustment module, configured to adjust at least one of the actions, expressions, and tones of the virtual avatar in the initial video based on the plurality of script tags to obtain a target video.

[0007] According to another aspect of the present disclosure, there is provided a live broadcast device, including: an adding module, configured to add a patch material in an initial live broadcast room to obtain a target live broadcast room; and a playing module, configured to play a target video in the target live broadcast room, where the target video is generated by using the method described above.

[0008] According to another aspect of the present disclosure, there is provided an agent, including: an input module for receiving input information; a processing module for generating a target video based on the input information received by the above input module by using the method as described above, and using the target video as output information; and an output module for outputting the output information obtained by the above processing module.

[0009] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.

[0010] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method as described above.

[0011] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program which, when executed by a processor, implements the method as described above.

[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0013] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0014] Figure 1 Schematically shows an exemplary system architecture to which the video generation method and apparatus according to an embodiment of the present disclosure can be applied.

[0015] Figure 2 Schematically shows a flowchart of the video generation method according to an embodiment of the present disclosure.

[0016] Figure 3A Schematically shows a schematic diagram of a live script according to an embodiment of the present disclosure.

[0017] Figure 3B Schematically shows a schematic diagram of a live script according to another embodiment of the present disclosure.

[0018] Figure 4 Schematically shows a schematic diagram of the video generation method according to an embodiment of the present disclosure.

[0019] Figure 5Schematically shows a schematic diagram of a video generation method according to an embodiment of the present disclosure.

[0020] Figure 6 Schematically shows a flowchart of a live broadcast method according to an embodiment of the present disclosure.

[0021] Figure 7 Schematically shows a block diagram of a video generation device according to an embodiment of the present disclosure.

[0022] Figure 8 Schematically shows a block diagram of a live broadcast device according to an embodiment of the present disclosure.

[0023] Figure 9 Schematically shows a block diagram of an agent according to an embodiment of the present disclosure.

[0024] Figure 10 Shows a schematic block diagram of an example electronic device that can be used to implement the embodiments of the present disclosure. Detailed implementation manners

[0025] The following makes an explanation of the exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the descriptions of well-known functions and structures are omitted in the following.

[0026] Live broadcast is currently an extremely important part of the commercial and entertainment content fields. However, due to reasons such as high professionalism, high cost, and limited live broadcast duration of traditional live broadcasts by real people, commercial activities carried out based on live broadcasts have a relatively high threshold.

[0027] With the rapid development of artificial intelligence technology, virtual avatar live broadcasts have been widely accepted and used in the market. Virtual avatar live broadcasts can use virtual digital humans to replace or assist real human hosts in real-time content output, achieving 7×24-hour live broadcasts at a relatively low cost, and effectively solving the problems of live broadcast cost, duration, and stability.

[0028] At the same time, along with the expansion of the application scenarios of virtual avatar live broadcasts, users have put forward higher requirements for the expressiveness of virtual avatars. For example, current virtual avatar products are already indistinguishable from real people in terms of intuitive appearance and voice, but they still cannot achieve a high degree of unity of elements such as the content of the explanation, body movements, facial expressions, and material pictures, making the performance of virtual avatars single and dull, and the expressiveness is poor.

[0029] In view of this, embodiments of the present disclosure provide a video generation method, a live broadcast method, a device, an intelligent agent, and an electronic device. Combining the concept of a script in a live broadcast by a real person, and based on script-driven, a high degree of matching between the explanation content of a virtual image and actions, expressions, etc. is achieved, thereby enhancing the expressiveness of the virtual image. Specifically, the video generation method includes: obtaining a live broadcast script, where the live broadcast script includes script text and a plurality of script tags; based on the script text, performing inference rendering on the voice and lip movements of the virtual image to obtain an initial video; and based on the plurality of script tags, adjusting at least one of the actions, expressions, and tones of the virtual image in the initial video to obtain a target video.

[0030] Figure 1 Schematically shows an exemplary system architecture to which the video generation method and device according to embodiments of the present disclosure can be applied.

[0031] It should be noted that Figure 1 The shown is only an example of the system architecture to which embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios. For example, in another embodiment, the exemplary system architecture to which the video generation method and device can be applied may include a terminal device, but the terminal device can implement the video generation method and device provided by embodiments of the present disclosure without interacting with a server.

[0032] As Figure 1 shown, the system architecture 100 according to this embodiment may include a terminal device 101, an intelligent agent 102, a network 103, and a server 104.

[0033] The terminal device 101 may be various electronic devices having a display screen and supporting network communication, including but not limited to smartphones, tablets, laptop portable computers, desktop computers, and the like.

[0034] A large language model may be configured in the intelligent agent 102 to generate answer text for questions input by users.

[0035] The network 103 may include various connection types, such as wired and / or wireless communication links, and the like.

[0036] The server 104 may be a server providing various services. For example, the server 104 may provide computing resource support for the operation of the intelligent agent 102.

[0037] It should be noted that the video generation method provided by embodiments of the present disclosure can generally be executed by the terminal device 101. Correspondingly, the video generation device provided by embodiments of the present disclosure can also be set in the terminal device 101.

[0038] Alternatively, the video generation method provided by the embodiments of the present disclosure can generally also be executed by the server 104. Correspondingly, the video generation device provided by the embodiments of the present disclosure can generally be set in the server 104. The video generation method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 104 and capable of communicating with the terminal device 101 and / or the server 104. Correspondingly, the video generation device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 104 and capable of communicating with the terminal device 101 and / or the server 104.

[0039] It should be understood that Figure 1 the numbers of the terminal devices, networks, agents, and servers in

[0040] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure, and application, etc., of the user's personal information all comply with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and it does not violate public order and good customs.

[0041] In the technical solution of the present disclosure, before obtaining or collecting the user's personal information, the authorization or consent of the user is obtained.

[0042] Figure 2 Schematically shows a flowchart of the video generation method according to the embodiments of the present disclosure.

[0043] As Figure 2 shown, the method 200 may include operations S210 to S230.

[0044] In operation S210, a live script is obtained, and the live script includes a script text and a plurality of script tags.

[0045] In operation S220, based on the script text, the voice and lip movement of the virtual image are inferred and rendered to obtain an initial video.

[0046] In operation S230, based on the plurality of script tags, at least one of the actions, expressions, and tones of the virtual image in the initial video is adjusted to obtain a target video.

[0047] Optionally, the above operations S210 to S230 can be executed by an agent, and at least one neural network model can be configured in the agent. The agent can call the at least one neural network model to implement the methods of the operations S210 to S230.

[0048] The manner in which the agent obtains the live script is not limited herein. For example, the agent can provide a text input window, and the agent can extract the live script from the information input by the user through the text input window. For another example, a generative model can be configured in the agent, and the generative model can output a live script based on the requirement information input by the user to the agent. Optionally, the agent can also provide a modification function for the live script, that is, the agent can display the obtained live script and synchronously modify the obtained live script in response to operations such as text insertion, text deletion, and text modification by the user on the live script.

[0049] The live script can include script text and multiple script tags.

[0050] The script text can include the voice-over text for a single-person live broadcast or the interactive text for a multi-person live broadcast. The script text can be composed of multiple text paragraphs, and the starting position of each text paragraph can indicate the virtual character explaining the paragraph. For example, a text paragraph can be expressed as "Host: Content 1", which means that in the generated video, the text content "Content 1" of the script text is explained by the virtual character "Host". For another example, a text paragraph can be expressed as "Host: Content 1. Assistant: Content 2", which means that in the generated video, the text content "Content 1" of the script text is explained by the virtual character "Host", and the text content "Content 2" of the script text is explained by the virtual character "Assistant".

[0051] The script tags can be embedded in the script text and can be set at any position in the script text. For example, script tags can be set at the starting position of each text paragraph in the script text. For another example, script tags can be set in the middle of the text paragraph. The form of expression of the script tags can be text. When the script tags are expressed in text form, delimiter symbols can be set on both sides of the script tags to distinguish the text content of the script tags from the text content of the script text. The delimiter symbols can be set according to the specific application scenario. For example, the delimiter symbols can be parentheses, double quotes, etc., or other preset symbols such as " / / ", " / *", etc. For example, a text paragraph containing script tags can be expressed as "Host: (Action: A1) (Expression: B1) Content 1, (Action: A2) (Expression: B2) Content 2". If the delimiter symbols are parentheses, it can be determined that the text paragraph contains the text contents "Content 1" and "Content 2", as well as the script tags "Action: A1", "Expression: B1", "Action: A2", and "Expression: B2".

[0052] The virtual character can be generated by the intelligent agent based on the configuration information input by the user, which is not limited herein. Through the Text To Speech (TTS) method, the text content of the script text can be converted into voice output. Combining with the voice configuration of the virtual character, the voice output can be converted into the voice output of the virtual character. Based on the voice output of the virtual character, the lip movements of the virtual character can be inferred and rendered so that the lip movements of the virtual character are synchronized with the voice of the virtual character. The intelligent agent can capture the animation of the virtual character explaining the script text by voice and use the animation of this period as the initial video output.

[0053] The script label can be configured with the scope of the script text it affects. The script label can be used to control the actions, expressions, etc. of the virtual character in the corresponding initial video within the scope of the script text it affects. Based on the multiple script labels, the actions, expressions, tones, etc. of the virtual character in multiple video segments of the initial video can be adjusted so that the virtual character in the obtained target video has stronger expressiveness.

[0054] According to the embodiments of the present disclosure, by taking the live script as the core element of video generation, for the created virtual character, using the script text included in the live script, the virtual character can be driven to explain by voice, and the lip movements of the virtual character can be inferred and rendered to obtain an initial video with the voice and lip movements of the virtual character synchronized. According to the script labels included in the live script, the actions, expressions, tones, etc. of the virtual character in the initial video can be adjusted personalized, so that a more expressive target video can be obtained, which can effectively improve the quality of the live content of the virtual character.

[0055] The following will make a further description of the Figure 2 method shown with reference to the accompanying drawings and in combination with specific embodiments.

[0056] Optionally, the live script can be generated based on the information input by the user. The intelligent agent can obtain the live configuration information and generate a live script based on the live configuration information.

[0057] The intelligent agent can provide an input window for configuration information and an upload window for files. Through this window, the intelligent agent can obtain the live configuration information input by the user.

[0058] In the application scenario with the live broadcast "person - goods - scene" as the core architecture, the live broadcast configuration information may include product information, reference knowledge, persona information, material information, and other input information. The product information may include the basic information of the products in this live broadcast, the order of the products, the main promoted products, etc. The reference knowledge may include the public knowledge on the Internet or the knowledge in various knowledge bases. The persona information may include the number of virtual avatars required for this live broadcast, as well as the persona, image, voice, etc. information of each virtual avatar. The material information may include the picture materials of the products, the video materials of the interaction actions with the products, the front and back patch materials of the live broadcast room, etc. The other input information may include the preset interaction information, etc. After obtaining the live broadcast configuration information, the intelligent agent can also perform structured processing on the live broadcast configuration information so that the intelligent agent can call the live broadcast configuration information at any time when generating the live broadcast script or generating the live broadcast video.

[0059] For example, for the product information, the product information input by the user may only contain the product names of at least one product. The intelligent agent can obtain the product names of at least one product respectively; determine the product detail pages based on the product names of the products; and use the second large - language model to deconstruct the detail pages of at least one product respectively to obtain structured product information.

[0060] The intelligent agent can obtain the product detail pages of the products from the public information based on the product names of the products through the connected knowledge base or through the search engine connected to the Internet. The product detail pages may contain product names, product production information, product pictures, etc.

[0061] The second large - language model may be a multi - modal understanding large - language model. This second large - language model can understand the content of the product detail pages and output the product information in a structured manner. For example, the second large - language model can extract the sub - information corresponding to each item of the structured information from the product detail pages respectively, and the extracted sub - information can be output according to the preset data structure to obtain structured product information.

[0062] Another example, for the persona information, the persona information input by the user may contain the description information of at least one virtual avatar to be generated described in natural language. The intelligent agent can obtain the description information of at least one virtual avatar to be generated respectively; extract the feature information from the description information of the virtual avatar to be generated to obtain the sub - information of the persona of the virtual avatar to be generated; and obtain the persona information based on the sub - information of the persona of at least one virtual avatar to be generated respectively.

[0063] For the description information of each virtual image to be generated, based on items such as the character design, image, and voice that need to be determined, the information corresponding to the items can be extracted from the description information and output in a preset data structure to obtain structured sub-information of the character design.

[0064] For another example, for the material information, the material information input by the user can include material images such as product ingredient diagrams and commodity price diagrams. The intelligent agent can use the configured large language model to perform multimodal understanding on the content expressed by the material images, rename and store the material images based on the understanding results for later use when generating a live script into a video.

[0065] After completing the structured processing of the live broadcast configuration information, the live broadcast configuration information can be used to generate a live script. For example, a first large language model for script generation can be configured in the intelligent agent. By inputting the prompt text generated based on the live broadcast configuration information into the first large language model, a live script output by the first large language model can be obtained.

[0066] The first large language model can be obtained by performing reinforcement learning on the base model using target training samples. For example, based on the target training samples, the base model is reinforced using GRPO (Group Relative Policy Optimization), a within-group relative reward optimization strategy, to obtain the first large language model. The target training samples can include samples labeled with multi-person interaction copywriting.

[0067] Optionally, combined with real-person live broadcast data and data annotation for the real-person live broadcast data, training methods such as SFT (Supervised Fine-Tuning) and data flywheel can be used on the large model to train the base model, enabling the base model to generate live scripts that are adapted to different commodity categories, conform to the host's character design, and contain multi-dimensional elements such as text and behavior. Then, by using reinforcement learning on the base model with training samples labeled with multi-person interaction copywriting, the trained first large language model can output live scripts with multi-person interaction and containing multi-dimensional elements such as the behaviors and expressions of multiple individuals.

[0068] Figure 3A Schematically shows a schematic diagram of a live script according to an embodiment of the present disclosure.

[0069] As Figure 3AAs shown, the live script can be a live script including a single virtual character, and the live script can be divided into three text paragraphs. The first text paragraph contains the script label "Action Label 1" and the script text "Text Content 1", the second text paragraph contains the script label "Expression Label 1" and the script text "Text Content 2", and the third text paragraph contains the script label "Action Label 2" and the script text "Text Content 3". All three text paragraphs can be explained by the virtual character "host".

[0070] Figure 3B A schematic diagram showing a live script according to another embodiment of the present disclosure is schematically illustrated.

[0071] As Figure 3B shown, the live script can be a live script including two virtual characters. The live script can be divided into 3 text paragraphs. Among them, the first text paragraph and the third text paragraph can be explained by the virtual character "host", and the second text paragraph can be explained by the virtual character "assistant host". The first text paragraph can contain the script labels "Action Label A1", "Expression Label B1" and the script text "Text Content C1". The first text paragraph can contain the script labels "Action Label A1", "Expression Label B1" and the script text "Text Content C1". The second text paragraph can contain the script labels "Action Label A2", "Expression Label B2" and the script text "Text Content C2". The third text paragraph can contain the script labels "Action Label A3", "Action Label A4", "Expression Label B3", "Expression Label B4" and the script texts "Text Content C3", "Text Content C4".

[0072] Before generating the video, the user can also configure the virtual characters to be used in the video to be generated.

[0073] Optionally, the agent can provide virtual character templates, and the user can select the required virtual characters from the virtual character templates. The virtual character templates can be approved real person templates, or can be templates of virtual characters generated by large language models and configured by developers, which are not limited here.

[0074] Alternatively, the agent can provide multiple options for the image, persona, voice, etc. of the virtual character respectively. Based on the user's selection operation, the agent can call the large language model to create a virtual character.

[0075] Or, a virtual character can also be created based on the live configuration information input by the user to the agent. For example, based on the character setting information and voice information included in the persona information, virtual character materials are extracted from the material library to combine and obtain an initial virtual character; and based on the image information included in the persona information, the initial virtual character is rendered to obtain a virtual character.

[0076] After determining the virtual avatar and the live script, the virtual avatar can be controlled to perform the live script, and the animation of the virtual avatar performing the live script can be intercepted to obtain the target video.

[0077] The process of controlling the virtual avatar to perform the live script can include two adjustment stages. In the first adjustment stage, based on the script text, the voice and image of the virtual avatar can be synchronized to obtain the initial video. In the second adjustment stage, based on the script tags, the actions, expressions, tones, etc. of the virtual avatar in the video can be adjusted to make the virtual avatar more expressive.

[0078] Optionally, in the first adjustment stage, based on the script text, inferring and rendering the voice and lip movement of the virtual avatar to obtain the initial video can include the following operations:

[0079] Based on the script text, call the voice model to generate audio data; based on the timeline of the audio data and the script text, call the visual model to infer and render the lip movement of the virtual avatar to obtain video data; and embed the audio data into the video data to obtain the initial video.

[0080] The voice model can be any text-to-speech model, which is not limited here.

[0081] After calling the voice model to generate audio data, the timestamps of the pronunciation moments of each word corresponding to the script text in the audio data can be determined, and the timestamps of the pronunciation moments of all words in the script text can form the timeline of the audio data.

[0082] Using the timeline of the audio data, the start time and end time of the virtual avatar explaining each word in the script text can be determined. Based on the start time, the lips of the virtual avatar can be controlled to start moving, and based on the end time, the lips of the virtual avatar can be controlled to stop moving. The lip movement of the virtual avatar controlled between the start time and the end time can represent the lip movement of a person explaining the word. Optionally, the lip movement material corresponding to the word can be obtained, and using the visual model, the lip movement of the virtual avatar can be rendered between the start time and the end time based on the lip movement material to achieve the synchronization between the lip movement and the voice of the virtual avatar and obtain the initial video.

[0083] Optionally, when inferring and rendering the lip movement of each word, the visual model can also be used to adjust the lip movement of the next word based on the lip movement of the previous word to ensure the coherence between the lip movements when the virtual avatar is explaining and improve the expressiveness of the virtual avatar.

[0084] For example, when adjusting the lip movement of the current character, it is possible to determine whether there will be a pause between the current character and the previous character. When it is determined that there is a pause, the virtual image can be inferred and rendered solely based on the lip movement material corresponding to the current character. If it is determined that there is no pause between the current character and the previous character, then when inferring and rendering the virtual image, the starting movement of the lip movement of the current character can be adjusted based on the ending movement of the lip movement corresponding to the previous character. For example, if the ending movement of the lip movement of the previous character is represented as opening the lips, and the current character is pronounced with the lips open, then during the inference and rendering, the movement between the lips closing and opening can be deleted.

[0085] Optionally, it is possible to determine whether there is a pause according to the character design of the virtual image. For example, if the character design information of the virtual image includes a gentle tone, it can be determined that there is a pause between each character. Alternatively, punctuation marks can be added to the script text to punctuate the script text. When there is a punctuation mark between the current character and the previous character, it can be determined that there is a pause between the current character and the previous character.

[0086] Optionally, in the second adjustment stage, based on multiple script tags, at least one of the actions, expressions, and tones of the virtual image in the initial video can be adjusted to obtain a target video, which may include the following operations:

[0087] Based on multiple script tags, the live script is segmented to obtain multiple script segments, and each script segment is associated with at least one script tag; the first video segment corresponding to the script segment in the initial video is determined; based on at least one script tag associated with the script segment, at least one of the actions, expressions, and tones of the virtual image in the first video segment corresponding to the script segment is adjusted to obtain a second video segment; and the multiple second video segments are spliced in the text order of the live script to obtain the target video.

[0088] The script tags can be used as segmentation points to segment the live script to obtain multiple script segments. The obtained script segments can start with one or more script tags and then be followed by a section of script text. Taking Figure 3B the script text shown as an example, its third text paragraph can be segmented into two script segments. The first script segment can include "action tag A3", "expression tag B3", and the script text "text content C3", and the second script segment can include "action tag A4", "expression tag B4", and the script text "text content C4".

[0089] Based on the types of script tags associated with a script segment, corresponding methods can be selected to adjust the virtual character in the first video segment corresponding to the script segment. Classified by type, script tags can be divided into action tags, expression tags, etc. The type of a script tag can be determined by semantic recognition of the text content of the script tag. For example, if the text content of a script tag is "pause slightly and raise the hand slightly", based on semantic analysis of this text content, it can be determined that the type of this script tag is an action tag. Another example, if the text content of a script tag is "slow down the tone and have a sincere look in the eyes", based on semantic analysis of this text content, it can be determined that the type of this script tag is an expression tag. Or, the type of a script tag can be determined according to the keywords included in the text content of the script tag. For example, if the text content of a script tag is "Action: shake the head", based on the keyword "Action" it contains, it can be determined that the type of this script tag is an action tag. Another example, if the text content of a script tag is "Expression: helpless", based on the keyword "Expression" it contains, it can be determined that the type of this script tag is an expression tag.

[0090] In some embodiments, if the script tag associated with a script segment is an action tag, based on at least one script tag associated with the script segment, at least one of the action, expression, and tone of the virtual character in the first video segment corresponding to the script segment is adjusted to obtain a second video segment, which may include the following operations:

[0091] Based on the action materials corresponding to the action tag, a multi-modal model is called to perform frame rendering on the first video segment to adjust the action of the virtual character in the first video segment to obtain a second video segment.

[0092] Optionally, the action of the virtual character can be decomposed into the actions of multiple action control points in the virtual character. By calling the multi-modal model, the action materials can be parsed into control instructions for each of the multiple action control points. Synchronously executing the control instructions for each of the multiple action control points can control the virtual character to perform limb actions according to the action represented by the action materials, so as to adjust the action of the virtual character in the first video segment to obtain a second video segment.

[0093] In some embodiments, if the script tag associated with a script segment is an expression tag, based on at least one script tag associated with the script segment, at least one of the action, expression, and tone of the virtual character in the first video segment corresponding to the script segment is adjusted to obtain a second video segment, which may include the following operations:

[0094] Based on the expression materials corresponding to the expression tags, call a multi-modal model to perform screen rendering on the first video clip to adjust the expression of the virtual character in the first video clip, and obtain the third video clip; and based on the emotion type represented by the expression tags, call a voice model to adjust the tone of the virtual character in the third video clip to obtain the second video clip.

[0095] Similarly, the change in the expression of the virtual character can be represented as the facial actions of the virtual character, and the facial actions of the virtual character can be decomposed into the actions of multiple expression control points in the virtual character. By calling the multi-modal model, the expression materials can be parsed into the control instructions of multiple expression control points respectively. Synchronously executing the control instructions of multiple expression control points respectively can control the virtual character to adjust the facial expression according to the actions represented by the expression materials, so as to adjust the facial expression of the virtual character in the first video clip and obtain the third video clip.

[0096] The change in expression is generally accompanied by the change in the tone of the virtual character. For example, when the expression changes from peaceful to serious, the voice of the virtual character also needs to become serious accordingly, and there should be more stress in the tone at this time.

[0097] The voice model can be called to adjust the tone of the virtual character in the third video clip by adjusting the speaking speed, pronunciation focus, etc. of the virtual character in the third video clip to obtain the second video clip.

[0098] Optionally, when changing the speaking speed of the virtual character in the third video clip, the visual model can also be called to fine-tune the lip movement of the virtual character to match the lip movement with the voice.

[0099] The script tag can also include a material tag. Similar to the action tag and the expression tag, the material tag can be obtained by performing semantic recognition on the text content of the script tag, or can be determined based on the keywords in the text content of the script tag, which will not be elaborated here. The material tag can indicate that when explaining the current position, add the material represented by the text content of the material tag in the video. This material can be obtained from the live broadcast configuration information input by the user.

[0100] In some embodiments, the script tag associated with the script segment can be a material tag, then the material represented by the text content of the material tag can be added to the first video clip corresponding to the script segment to obtain the second video clip.

[0101] For example, a script segment containing a material tag can be expressed as: "(Material: Cold Food) Let's talk about the poem Cold Food now. You may not know the background of the finger..." In this script segment, the script tag can be expressed as "Material: Cold Food", and the picture material corresponding to "Cold Food" can be added to the first video segment corresponding to the script segment to obtain the second video segment. Therefore, when the virtual image explains the script segment, the picture material corresponding to "Cold Food" can be displayed in the video screen, thereby improving the interactive effect of the live video.

[0102] After adjusting each first video segment in sequence to obtain a plurality of second video segments, the plurality of second video segments may be spliced ​​into a target video.

[0103] Optionally, in order to achieve a natural and smooth connection of actions, expressions, tones, etc. and reduce sudden changes, when splicing video clips, a compensating video clip may be added between adjacent second video clips.

[0104] For example, for the i-th second video clip, a compensating video clip for connecting the two can be generated based on the movements, expressions, and tones of the virtual image at the starting position of the i-th second video clip and the movements, expressions, and tones of the virtual image at the ending position of the i-1th second video clip, and the two can be spliced ​​in the order of i-1th second video clip-compensating video clip-i-th second video clip to achieve the splicing of the i-1th second video clip and the i-th second video clip.

[0105] Alternatively, the connecting parts of adjacent second video clips may be adjusted to achieve smooth connection of adjacent second video clips.

[0106] For example, for the i-th second video segment, the i-th second video segment can be inferred and rendered based on at least one script tag associated with the i-th second video segment and at least one script tag associated with the i-1-th second video segment to obtain a fifth video segment, where i is an integer greater than 1; and multiple fifth video segments are spliced ​​according to the text order of the live broadcast script to obtain a target video.

[0107] Specifically, based on at least one script tag associated with the i-1th second video clip, the action, expression and tone of the virtual image at the end of the i-1th second video clip can be inferred, and based on the action, expression and tone of the virtual image at the end of the i-1th second video clip, the action, expression and tone of the virtual image in a subsequent time period can be inferred. The i-th second video clip can be inferred and rendered using the action, expression and tone of the virtual image in the subsequent time period to obtain the fifth video clip.

[0108] Figure 4A schematic diagram showing a video generation method according to an embodiment of the present disclosure is schematically illustrated. Using the method of this embodiment, a target video including a single virtual avatar can be generated.

[0109] As Figure 4 shown, the agent can receive the live broadcast configuration information 401 input by the user and input the live broadcast configuration information 401 as a query text into the first large language model 402 to obtain the output live script 403. The live script 403 may include a script text 4031 and a plurality of script tags 4032.

[0110] Using the script text 4031, the agent can separately call the voice model 404 and the visual model 405 to perform inference rendering on the lip movement and voice of the virtual avatar to obtain the initial video 406.

[0111] The live script 403 can be segmented into multiple script segments based on the plurality of script tags 4032. Similarly, the initial video 406 can be segmented into multiple first video segments 407 based on the segmentation method of the plurality of script segments.

[0112] Each script segment may have one or more associated script tags 4032. Based on the script tag 4032 being an action tag or an expression tag, the actions, expressions, tones, etc. of the virtual avatar in the first video segment 407 can be adjusted to obtain the second video segment 408.

[0113] The multiple second video segments 408 can be spliced, and when splicing, the video segments at the splicing points can be adaptively adjusted to achieve natural and smooth connection of the actions, expressions, tones, etc. of the virtual avatar in different second video segments 408, thereby obtaining the target video 409.

[0114] According to an embodiment of the present disclosure, through the adjustment in the second stage, the actions, expressions, voices, materials, etc. of the virtual avatar in the video can be made consistent with the explanation, thereby effectively improving the expressiveness of the virtual avatar and the quality of the generated video.

[0115] Optionally, there may be multiple narrators in the live script, and correspondingly, the script text of the live script may include multiple virtual avatars. To enhance the interaction effect between the multiple virtual avatars, after adjusting the first video segment based on the script tags associated with the script segments, the actions, expressions, tones, etc. of another virtual avatar can also be adjusted based on the actions, expressions, tones, etc. of one virtual avatar.

[0116] According to an embodiment of the present disclosure, adjusting at least one of the actions, expressions, and tones of the virtual avatar in the initial video based on a plurality of script tags to obtain a target video may include the following operations:

[0117] Based on multiple script tags, the live script is segmented to obtain multiple script segments, and at least one script tag is associated with each script segment; determine the first video segment in the initial video corresponding to the script segment; based on at least one script tag associated with the script segment, adjust at least one of the actions, expressions, and tones of the virtual character in the first video segment corresponding to the script segment to obtain a second video segment; determine the first virtual character associated with the script segment and the second virtual character associated with the adjacent script segment, where the script segment and the adjacent script segment are adjacent according to the text order of the live script; in the case where the first virtual character is different from the second virtual character, based on at least one script tag associated with the script segment and at least one script tag associated with the adjacent script segment, adjust at least one of the actions, expressions, and tones of the second virtual character in the second video segment corresponding to the script segment to obtain a fourth video segment; and splice the multiple fourth video segments according to the text order of the live script to obtain the target video.

[0118] A script segment can correspond to one virtual character among multiple virtual characters. Adjusting at least one of the actions, expressions, and tones of the virtual character in the first video segment corresponding to the script segment based on at least one script tag associated with the script segment to obtain a second video segment can be expressed as adjusting at least one of the actions, expressions, and tones of the one virtual character in the first video segment corresponding to the script segment based on at least one script tag associated with the script segment to obtain a second video segment.

[0119] Optionally, when adjusting the actions, expressions, and tones of one virtual character, especially when adjusting the expressions and tones of the one virtual character, the expressions and tones of other virtual characters in the video segment can be adjusted with a smaller amplitude to add the feedback of the other virtual characters to the actions, expressions, and tones of the one virtual character, so as to achieve the matching of the actions, expressions, and tones among multiple virtual characters.

[0120] When there is a switch in the virtual character targeted between script segments, the virtual character in the current script segment can be adjusted based on the script tags of the adjacent script segments to achieve a more natural multi-person interaction.

[0121] For example, the virtual image associated with the current script segment is virtual image V1, and the virtual image associated with the previous script segment of the current script segment is virtual image V2. When adjusting the virtual image in the second video segment corresponding to the current script segment, the actions, expressions, tones, etc. of virtual image V2 in the second video segment corresponding to the current script segment can also be adjusted slightly according to the script tags corresponding to the previous script segment, so as to achieve the continuity of the actions, expressions, tones, etc. of virtual image V2.

[0122] Figure 5 FIG. schematically shows a schematic diagram of a video generation method according to an embodiment of the present disclosure. Using the method of this embodiment, a target video including multiple virtual images can be generated.

[0123] As Figure 5 shown, the agent can receive the live broadcast configuration information 501 input by the user, and input the live broadcast configuration information 501 as a query text into the first large language model 502 to obtain the output live broadcast script 503. The live broadcast script 503 may include script text 5031 and multiple script tags 5032.

[0124] Using the script text 5031, the agent can respectively call the speech model 504 and the visual model 505 to perform inference rendering on the lip movement and speech of the virtual image to obtain the initial video 506.

[0125] The live broadcast script 503 can be segmented into multiple script segments based on multiple script tags 5032. Similarly, the initial video 506 can be segmented into multiple first video segments 507 based on the segmentation method of the multiple script segments.

[0126] Each script segment may have one or more associated script tags 5032. For each script segment, a first virtual image corresponding to the script segment can be determined, and at least one other virtual image other than the first virtual image can be determined from multiple virtual images. Based on the script tag 5032 being an action tag or an expression tag, the actions, expressions, tones, etc. of the first virtual image in the first video segment 507 can be adjusted, and the actions, expressions, tones, etc. of at least one other virtual image can be adjusted accordingly in a small amplitude to obtain the second video segment 508. The adjustment amplitude of each of the at least one other virtual image can be randomly generated, or can also be determined based on the importance of the other virtual images. For example, the importance of the "host" is greater than the importance of the "assistant host" is greater than the importance of the "field controller", etc., which is not limited herein.

[0127] In the case where the first virtual image associated with the current script segment is different from the second virtual image associated with an adjacent script segment, especially the previous script segment of the current script segment, in the second video segment corresponding to the current script segment, the actions, expressions, tones, etc. of the second virtual image may change suddenly. Therefore, based on at least one script tag 509 associated with the previous script segment, the actions, expressions, tones, etc. of the second virtual image in the second video segment 508 corresponding to the current script segment can be adjusted slightly to achieve smooth connection of the actions, expressions, tones, etc. of the second virtual image in adjacent video segments, so as to obtain the fourth video segment 510.

[0128] Multiple fourth video segments 510 can be spliced, and during splicing, the video segments at the splicing points can be adaptively adjusted to achieve natural and smooth connection of the actions, expressions, tones, etc. of the virtual image in different fourth video segments 510, thereby obtaining the target video 511.

[0129] Figure 6 The flowchart of the live broadcast method according to an embodiment of the present disclosure is schematically shown.

[0130] As Figure 6 shown, the method 600 may include operations S610 to S620.

[0131] In operation S610, a patch material is added to the initial live broadcast room to obtain the target live broadcast room.

[0132] In operation S620, the target video is played in the target live broadcast room.

[0133] The target video may be generated by using the video generation method described above, which will not be elaborated here.

[0134] Optionally, when playing the target video in the target live broadcast room, interaction information from the audience can be obtained. The interaction information may include the audience's comments, bullet screens, etc. Based on the interaction information, a live broadcast script in the interaction scenario can be generated, and using the video generation method described above, an interactive video can be generated based on the live broadcast script in the interaction scenario, and the interactive video can be played in the target live broadcast room to achieve interaction between the virtual image and the audience in the target live broadcast room and improve the expressiveness of the virtual image. Optionally, the interactive video can be played at the end of the playback of a video segment of the target video.

[0135] According to an embodiment of the present disclosure, a virtual image can be used to replace a real person for live broadcast. After generating the target video by using the video generation method described above, the target video can be played cyclically in the target live broadcast room to achieve 7×24-hour live broadcast of the virtual image, which can effectively reduce the live broadcast cost, increase the live broadcast duration, and improve the stability of the live broadcast.

[0136] Figure 7 Schematically shows a block diagram of a video generation device according to an embodiment of the present disclosure.

[0137] As Figure 7 shown, the video generation device 700 may include a first acquisition module 710, an inference module 720, and an adjustment module 730.

[0138] The first acquisition module 710 is configured to acquire a live script, where the live script includes script text and a plurality of script tags.

[0139] The inference module 720 is configured to perform inference rendering on the voice and lip movement of the virtual image based on the script text to obtain an initial video.

[0140] The adjustment module 730 is configured to adjust at least one of the actions, expressions, and tones of the virtual image in the initial video based on the plurality of script tags to obtain a target video.

[0141] According to an embodiment of the present disclosure, the adjustment module 730 includes a first adjustment unit, a second adjustment unit, a third adjustment unit, and a fourth adjustment unit.

[0142] The first adjustment unit is configured to segment the live script based on the plurality of script tags to obtain a plurality of script segments, and each script segment is associated with at least one script tag.

[0143] The second adjustment unit is configured to determine a first video segment corresponding to the script segment in the initial video.

[0144] The third adjustment unit is configured to adjust at least one of the actions, expressions, and tones of the virtual image in the first video segment corresponding to the script segment based on at least one script tag associated with the script segment to obtain a second video segment.

[0145] The fourth adjustment unit is configured to splice the plurality of second video segments in the text order of the live script to obtain a target video.

[0146] According to an embodiment of the present disclosure, at least one script tag includes an action tag.

[0147] According to an embodiment of the present disclosure, the third adjustment unit includes a first adjustment subunit.

[0148] The first adjustment subunit is configured to call a multimodal model to perform frame rendering on the first video segment based on the action material corresponding to the action tag to adjust the actions of the virtual image in the first video segment to obtain a second video segment.

[0149] According to an embodiment of the present disclosure, at least one script tag includes an expression tag.

[0150] According to an embodiment of the present disclosure, the third adjustment unit includes a second adjustment subunit and a third adjustment subunit.

[0151] The second adjustment subunit is configured to call a multimodal model to perform screen rendering on the first video segment based on the expression material corresponding to the expression label, so as to adjust the expression of the virtual image in the first video segment to obtain a third video segment.

[0152] The third adjustment subunit is configured to call a voice model to adjust the tone of the virtual image in the third video segment based on the emotion type represented by the expression label to obtain a second video segment.

[0153] According to an embodiment of the present disclosure, the initial video includes a plurality of virtual images.

[0154] According to an embodiment of the present disclosure, the adjustment module 730 further includes a fifth adjustment unit, a sixth adjustment unit, and a seventh adjustment unit.

[0155] The fifth adjustment unit is configured to determine a first virtual image associated with the script segment and a second virtual image associated with an adjacent script segment, and the script segment and the adjacent script segment are adjacent according to the text order of the live script.

[0156] The sixth adjustment unit is configured to, when the first virtual image is different from the second virtual image, based on at least one script label associated with the script segment and at least one script label associated with the adjacent script segment, adjust at least one of the actions, expressions, and tones of the second virtual image in the second video segment corresponding to the script segment to obtain a fourth video segment.

[0157] The seventh adjustment unit is configured to splice a plurality of fourth video segments according to the text order of the live script to obtain a target video.

[0158] According to an embodiment of the present disclosure, the fourth adjustment unit includes a fourth adjustment subunit and a fifth adjustment subunit.

[0159] The fourth adjustment subunit is configured to perform inference rendering on the i-th second video segment based on at least one script label associated with the i-th second video segment and at least one script label associated with the (i - 1)-th second video segment to obtain a fifth video segment, where i is an integer greater than 1.

[0160] The fifth adjustment subunit is configured to splice a plurality of fifth video segments according to the text order of the live script to obtain a target video.

[0161] According to an embodiment of the present disclosure, the inference module 720 includes a first inference unit, a second inference unit, and a third inference unit.

[0162] The first inference unit is used to call a speech model to generate audio data based on the script text.

[0163] The second inference unit is used to call a visual model to infer and render the lip movement of the virtual avatar based on the timeline of the audio data and the script text, and obtain video data.

[0164] The third inference unit is used to embed the audio data into the video data to obtain an initial video.

[0165] According to an embodiment of the present disclosure, the video generation device 700 further includes a second acquisition module and a generation module.

[0166] The second acquisition module is used to acquire live broadcast configuration information.

[0167] The generation module is used to generate a live broadcast script based on the live broadcast configuration information.

[0168] According to an embodiment of the present disclosure, the generation module includes a generation unit.

[0169] The generation unit is used to input the prompt text generated based on the live broadcast configuration information into the first large language model to obtain a live broadcast script, wherein the first large language model is obtained by performing reinforcement learning on the base model using target training samples, and the target training samples include samples labeled with multi-person interaction copywriting.

[0170] According to an embodiment of the present disclosure, the live broadcast configuration information includes product information.

[0171] According to an embodiment of the present disclosure, the second acquisition module includes a first acquisition unit, a second acquisition unit, and a third acquisition unit.

[0172] The first acquisition unit is used to acquire the product names of at least one product respectively.

[0173] The second acquisition unit is used to determine the product detail page based on the product name of the product.

[0174] The third acquisition unit is used to deconstruct the product detail pages of at least one product respectively using a second large language model to obtain structured product information.

[0175] According to an embodiment of the present disclosure, the live broadcast configuration information includes character setting information.

[0176] According to an embodiment of the present disclosure, the second acquisition module includes a fourth acquisition unit, a fifth acquisition unit, and a sixth acquisition unit.

[0177] The fourth acquisition unit is used to acquire the description information of at least one virtual avatar to be generated respectively.

[0178] A fifth acquisition unit, configured to extract feature information from the description information of the virtual image to be generated, so as to obtain the character sub-information of the virtual image to be generated.

[0179] A sixth acquisition unit, configured to obtain character setting information based on the character sub-information of at least one virtual image to be generated.

[0180] According to an embodiment of the present disclosure, the video generation device 700 further includes a combination module and a rendering module.

[0181] The combination module is configured to extract virtual image materials from the material library based on the character setting information and voice information included in the character setting information, so as to combine and obtain an initial virtual image.

[0182] The rendering module is configured to perform screen rendering on the initial virtual image based on the image information included in the character setting information, so as to obtain a virtual image.

[0183] Figure 8 Schematically shows a block diagram of a live broadcast device according to an embodiment of the present disclosure.

[0184] As Figure 8 shown, the live broadcast device 800 may include an adding module 810 and a playing module 820.

[0185] The adding module is configured to add a patch material to the initial live broadcast room to obtain a target live broadcast room.

[0186] The playing module is configured to play a target video in the target live broadcast room.

[0187] According to an embodiment of the present disclosure, the target video is generated by using the method according to any one of claims 1 to 12.

[0188] The target video may be generated by using the video generation method described above, which will not be elaborated here.

[0189] Figure 9 Schematically shows a block diagram of an agent according to an embodiment of the present disclosure.

[0190] In an embodiment of the present disclosure, inspired by the von Neumann architecture in modern computer theory, as Figure 9 shown, the AI agent 900 may include three core modules: an input module 910, an output module 920, and a processing module 930. The processing module 930 may include a control unit 931, a storage unit 932, and an arithmetic unit 933.

[0191] The input module 910 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the external world (e.g., users or the external environment) and converting it into a format that the AI agent 900 can understand and process. The input module 910 is the primary link for the AI agent 900 to interact with the external world, enabling the AI agent 900 to efficiently and accurately obtain necessary "sensory" information from the external world and respond to this information.

[0192] In the example, the input information received by using the input module 910 can be a live script input by the user or live configuration information.

[0193] In the example, the processing module 930 is the core support for the AI agent 900's ability to handle complex tasks. The processing module 930 can, based on the input information received by the input module 910, execute the video generation method as described above by invoking one or more large language models to obtain a target video and output the target video as output information.

[0194] In the example, the control unit 931 in the processing module 930 will continuously interact with the storage unit 932, the arithmetic unit 933, and / or the output module 920 during operation. However, it should be noted that in the embodiments of the present disclosure, the control unit 931 acts as a single initiator to initiate communication with the storage unit 932, the arithmetic unit 933, and / or the output module 920, and there is no communication coupling between the storage unit 932, the arithmetic unit 933, and the output module 920.

[0195] In the example, the performance of the control unit 931 can be closely related to the large model on which the AI agent 900 is based. To fully utilize the capabilities of the large language model, the internal structure of the control unit 931 can be designed to be highly configurable and extensible to cope with various different types of tasks and requirements in real-world scenarios.

[0196] The storage unit 932 can be responsible for memorizing information such as historical conversations and event streams. The text generated in each round as described above can be included in the storage unit 932.

[0197] In the example, after the AI agent 900 obtains a text generation request, the AI agent 900 can call a large model to execute a task corresponding to the input information and output corresponding text. The corresponding text can be stored in the storage unit 932. The AI agent 900 can retrieve relevant data resources from the storage unit 932 and feedback them to the control unit 931. Then, the control unit 931 can utilize the feedback data resources to generate a target answer. It is also possible to retrieve relevant data resources from the storage unit 932 and feedback them to the control unit 931. Then, the control unit 931 can utilize the returned data resources to generate a target answer. And the target answer is passed to the output module 920.

[0198] The operation unit 933 can be regarded as a predefined tool library. As mentioned above, renderers and display controls can be included in the operation unit 933.

[0199] In the example, when the AI agent 900 needs to render multiple output data, relevant renderers and display controls can be called from the operation unit 933 and feedback them to the control unit 932. Then, the control unit 932 can utilize the feedback renderers and display controls to render the first search result and pass the first search result to the output module 920. It can be understood that although large language models have excellent language understanding and generation capabilities, like humans, the tasks they can solve without any tools are very limited. When the AI agent 900 is given the ability to call tools, it can achieve tasks such as performing mathematical operations with the help of a calculator, completing data analysis with the help of Python, and completing prediction tasks with the help of a certain tool.

[0200] In the example, the output module 920 can output the output information of the target task specification.

[0201] According to the AI agent 900 of the embodiments of the present disclosure, the intelligence level can be simply and effectively improved, and the flexibility and versatility can be enhanced.

[0202] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0203] According to the embodiments of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.

[0204] According to the embodiments of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.

[0205] According to an embodiment of the present disclosure, a computer program product includes a computer program which, when executed by a processor, implements the method described above.

[0206] Figure 10 FIG. shows a schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0207] As Figure 10 shown, the device 1000 includes a computing unit 1001 which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0208] A plurality of components in the device 1000 are connected to the input / output (I / O) interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, an optical disk, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0209] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above, such as a video generation method or a live broadcast method. For example, in some embodiments, the video generation method or the live broadcast method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the video generation method or the live broadcast method described above can be executed. Alternatively, in other embodiments, the computing unit 1001 can be configured to execute the video generation method or the live broadcast method by any other suitable means (e.g., by means of firmware).

[0210] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0211] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0212] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0213] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0214] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0215] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0216] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitations are imposed herein.

[0217] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A video generation method, comprising: Obtaining a live script, where the live script includes script text and multiple script tags; Based on the script text, inferring and rendering the speech and lip movements of the virtual character to obtain an initial video; And Based on the multiple script tags, adjusting at least one of the actions, expressions, and tones of the virtual character in the initial video to obtain a target video.

2. The method according to claim 1, wherein The step of adjusting at least one of the actions, expressions, and tones of the virtual character in the initial video based on the multiple script tags to obtain a target video includes: Based on the multiple script tags, splitting the live script to obtain multiple script segments, and each script segment is associated with at least one script tag; Determining a first video segment corresponding to the script segment in the initial video; Based on at least one script tag associated with the script segment, adjusting at least one of the actions, expressions, and tones of the virtual character in the first video segment corresponding to the script segment to obtain a second video segment; and Splicing multiple second video segments in the text order of the live script to obtain the target video.

3. The method according to claim 2, wherein, The at least one script tag includes an action tag; Among them, the step of adjusting at least one of the actions, expressions, and tones of the virtual character in the first video segment corresponding to the script segment based on at least one script tag associated with the script segment to obtain a second video segment includes: Based on the action material corresponding to the action tag, calling a multi-modal model to perform frame rendering on the first video segment to adjust the action of the virtual character in the first video segment to obtain the second video segment.

4. The method according to claim 2, wherein The at least one script tag includes an expression tag; Among them, the step of adjusting at least one of the actions, expressions, and tones of the virtual character in the first video segment corresponding to the script segment based on at least one script tag associated with the script segment to obtain a second video segment includes: Based on the expression material corresponding to the expression tag, calling a multi-modal model to perform frame rendering on the first video segment to adjust the expression of the virtual character in the first video segment to obtain a third video segment; and Based on the emotion type represented by the expression tag, calling a speech model to adjust the tone of the virtual character in the third video segment to obtain the second video segment.

5. The method according to claim 2, wherein, The initial video includes multiple virtual characters; The method further includes: Determining a first virtual character associated with the script segment and a second virtual character associated with an adjacent script segment, where the script segment and the adjacent script segment are adjacent in the text order of the live script; In the case where the first virtual character is different from the second virtual character, based on at least one script tag associated with the script segment and at least one script tag associated with the adjacent script segment, adjusting at least one of the actions, expressions, and tones of the second virtual character in the second video segment corresponding to the script segment to obtain a fourth video segment; and Splice multiple of the fourth video segments according to the text order of the live script to obtain the target video.

6. The method according to claim 2, wherein The step of splicing multiple of the second video segments according to the text order of the live script to obtain the target video includes: Based on at least one script tag associated with the i-th second video segment and at least one script tag associated with the (i - 1)-th second video segment, perform inference rendering on the i-th second video segment to obtain a fifth video segment, where i is an integer greater than 1; and Splice multiple of the fifth video segments according to the text order of the live script to obtain the target video.

7. The method according to claim 1, wherein The step of performing inference rendering on the voice and lip movement of the virtual character based on the script text to obtain the initial video includes: Based on the script text, call a voice model to generate audio data; Based on the time axis of the audio data and the script text, call a visual model to perform inference rendering on the lip movement of the virtual character to obtain video data; and Embed the audio data into the video data to obtain the initial video.

8. The method according to claim 1, further comprising: Obtain live configuration information; And Generate the live script based on the live configuration information.

9. The method according to claim 8, wherein The step of generating the live script based on the live configuration information includes: Input the prompt text generated based on the live configuration information into a first large language model to obtain the live script, where the first large language model is obtained by performing reinforcement learning on a base model using target training samples, and the target training samples include samples labeled with multi-person interaction copywriting.

10. The method according to claim 8, wherein, The live configuration information includes product information; Wherein, the step of obtaining the live configuration information includes: Obtain the product names of at least one product respectively; Based on the product names of the products, determine the product detail pages; and Use a second large language model to deconstruct the detail pages of the at least one product respectively to obtain the structured product information.

11. The method according to claim 8, wherein, The live configuration information includes character set information; Wherein, the step of obtaining the live configuration information includes: Obtain the description information of at least one virtual character to be generated respectively; Extract feature information from the description information of the virtual character to be generated to obtain the character set sub-information of the virtual character to be generated; and Based on the character set sub-information of at least one virtual character to be generated respectively, obtain the character set information.

12. The method according to claim 11, further comprising: Based on the character setting information and voice information included in the character set information, extract virtual character materials from the material library to combine into an initial virtual character; And Based on the image information included in the character set information, perform screen rendering on the initial virtual character to obtain the virtual character.

13. A live method, comprising: Add a patch material in the initial live room to obtain a target live room; And Play a target video in the target live room, where the target video is generated by using the method according to any one of claims 1 to 12.

14. A video generation device, comprising: A first acquisition module, configured to acquire a live script, where the live script includes script text and a plurality of script tags; An inference module, configured to perform inference rendering on the voice and lip movement of a virtual image based on the script text to obtain an initial video; And An adjustment module, configured to adjust at least one of the actions, expressions, and tones of the virtual image in the initial video based on the plurality of script tags to obtain a target video.

15. A live broadcast method, including: An adding module, configured to add a patch material in an initial live broadcast room to obtain a target live broadcast room; And A playing module, configured to play a target video in the target live broadcast room, where the target video is generated by using the method according to any one of claims 1 to 12.

16. An intelligent agent, including: An input module, configured to receive input information; A processing module, configured to generate a target video by using the method according to any one of claims 1 to 12 based on the input information received by the input module, and use the target video as output information; And An output module, configured to output the output information obtained by the processing module.

17. An electronic device, including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 13.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 13.

19. A computer program product, including a computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1 to 13.