Digital human video generation method and device based on large model, intelligent agent, electronic equipment and storage medium
Through the digital human video generation method based on the big model, the language and visual big model are used to process action description information together to generate high-quality digital human videos, which solves the problem of insufficient motion expression and improves the expressiveness and naturalness of the video.
Patent Information
- Application Number
- CN202510536606.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-11
AI Technical Summary
In scenes such as e-commerce live broadcasts and film and television product production, the generated digital human videos have insufficient action expressiveness, resulting in poor video quality and affecting the user's viewing experience.
Using a big model-based method, by obtaining action description information, the language model is used to generate oral clip text in the target script, and combined with the visual model to process the target script and action video clips to generate target videos, so that the target object can perform oral expressions during the execution of the specified action, ensuring the coordination and adaptation of oral expressions and action expressions.
It improves the expressiveness and naturalness of oral expression of the target object in the target video, improves the video quality, enhances the coordination of action and oral play, and meets the diversity and professional needs of users.
Smart Images

Figure CN120302122A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to technical fields such as deep learning, large models, and computer vision, and can be applied to scenarios such as video live streaming, advertisement production, and e-commerce sales. Background Art
[0002] With the rapid development of Internet technology, users can conveniently browse resource information such as videos through intelligent terminal devices such as smart phones. For example, users can browse live videos through smart phones to understand the detailed information of products. Summary of the Invention
[0003] The present disclosure provides a method, apparatus, intelligent agent, electronic device, and storage medium for generating digital human videos based on large models.
[0004] According to one aspect of the present disclosure, there is provided a method for generating a digital human video based on a large model, including: obtaining requirement information, where the requirement information includes action description information for describing a specified action video segment, and the action video segment represents a specified action of a target object; processing the requirement information using a first large model to obtain a target script, where the target script includes target voiceover segment text that matches the action description information; and processing the target script and the action video segment using a second large model to obtain a target video for displaying the target object performing a voiceover based on the target voiceover segment text during the execution of the specified action.
[0005] According to another aspect of the present disclosure, there is provided a device for generating a digital human video based on a large model, including: an obtaining module for obtaining requirement information, where the requirement information includes action description information for describing a specified action video segment, and the action video segment represents a specified action of a target object; a target script obtaining module for processing the requirement information using a first large model to obtain a target script, where the target script includes target voiceover segment text that matches the action description information; and a target video obtaining module for processing the target script and the action video segment using a second large model to obtain a target video for displaying the target object performing a voiceover based on the target voiceover segment text during the execution of the specified action.
[0006] According to another aspect of the present disclosure, there is provided an intelligent agent of artificial intelligence, including: an input module for receiving input information; a processing module for determining a target task based on the input information received by the input module, determining a first large model and a second large model based on the target task, and executing the method for generating a digital human video based on a large model provided in the embodiments of the present disclosure by calling the first large model and the second large model to obtain output information; and an output module for outputting the output information obtained by the processing module.
[0007] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the large model-based digital human video generation method provided by the embodiments of the present disclosure.
[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the large model-based digital human video generation method provided by the embodiments of the present disclosure.
[0009] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements the large model-based digital human video generation method provided by the embodiments of the present disclosure.
[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0012] Figure 1 Schematically shows an exemplary system architecture to which the content processing method and apparatus according to the embodiments of the present disclosure can be applied;
[0013] Figure 2 Schematically shows a flowchart of the large model-based digital human video generation method according to the embodiments of the present disclosure;
[0014] Figure 3 Schematically shows a schematic diagram of the principle for determining an action video segment according to the embodiments of the present disclosure;
[0015] Figure 4 Schematically shows a schematic diagram of determining action description information according to the embodiments of the present disclosure;
[0016] Figure 5 Schematically shows a schematic diagram of determining a target script according to the embodiments of the present disclosure;
[0017] Figure 6 Schematically shows a schematic diagram of generating a target video according to the embodiments of the present disclosure;
[0018] Figure 7 Schematically shows a schematic diagram of determining the voice-over speech data according to the embodiments of the present disclosure;
[0019] Figure 8 Schematically shows a schematic flowchart of determining a dynamic video segment according to an embodiment of the present disclosure;
[0020] Figure 9 Schematically shows a schematic diagram of a digital human video generation method based on a large model according to an embodiment of the present disclosure;
[0021] Figure 10 Schematically shows a block diagram of a digital human video generation device based on a large model according to an embodiment of the present disclosure;
[0022] Figure 11 Schematically shows a block diagram of the structure of an intelligent agent of artificial intelligence according to an embodiment of the present disclosure; and
[0023] Figure 12 Shows a schematic block diagram of an example electronic device that can be used to implement the digital human video generation method based on a large model according to the embodiments of the present disclosure. Detailed implementation manners
[0024] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0025] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and public order and good customs are not violated.
[0026] The inventors found that digital human videos generated in scenarios such as e-commerce live broadcasts and film and television product production have problems such as insufficient action expressiveness, resulting in poor video quality and affecting the user's video viewing experience.
[0027] Embodiments of the present disclosure provide a method, apparatus, intelligent agent, electronic device, storage medium, and computer program product for generating a digital human video based on a large model. The method for generating a digital human video based on a large model includes: obtaining requirement information, where the requirement information includes action description information for describing a specified action video segment, and the action video segment represents a specified action of a target object; processing the requirement information using a first large model to obtain a target script, where the target script includes target oral broadcast segment text that matches the action description information; and processing the target script and the action video segment using a second large model to obtain a target video for displaying the target digital human performing the specified action and making an oral broadcast based on the target oral broadcast segment text.
[0028] According to the embodiments of the present disclosure, by processing the requirement information including action description information using a first large model, a target oral broadcast segment text that can match the specified action represented by the action description information is generated. Thus, based on the relatively powerful understanding ability and text output ability of the first large model, the semantics of the target oral broadcast segment text in the target script can be adapted to the action intention represented by the specified action. By processing the target script including the target oral broadcast segment text and the action video segment using a second large model to generate a target video for driving the target object, the target object in the target video can make an oral broadcast expression based on the target oral broadcast segment text that is adapted to the specified action during the execution of the specified action, so that the oral broadcast expression and the action expression of the target object in the target video are coordinated and adapted, improving the naturalness of the oral broadcast expression of the target object in the target video, enhancing the expressiveness of the target object, and further improving the quality of the target video.
[0029] Figure 1 An exemplary system architecture to which the method and apparatus for generating a digital human video based on a large model according to the embodiments of the present disclosure can be applied is schematically shown.
[0030] It should be noted that Figure 1 The illustration is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios. For example, in another embodiment, the exemplary system architecture to which the method and apparatus for generating a digital human video based on a large model can be applied may include a terminal device, but the terminal device can implement the method and apparatus for generating a digital human video based on a large model provided by the embodiments of the present disclosure without interacting with a server.
[0031] Such as Figure 1As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0032] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only for example).
[0033] The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0034] The server 105 may be a server providing various services, such as a background management server that supports the content browsed by users using the terminal devices 101, 102, 103 (only for example). The background management server may analyze and process data such as received user requests, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0035] The server 105 may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server 105 may also be a server of a distributed system, or a server combined with a blockchain.
[0036] It should be noted that the digital human video generation method based on a large model provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the digital human video generation device based on a large model provided by the embodiments of the present disclosure can generally be arranged in the server 105. The digital human video generation method based on a large model provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103, and / or the server 105. Correspondingly, the digital human video generation device based on a large model provided by the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103, and / or the server 105.
[0037] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0038] Figure 2 FIG. schematically shows a flowchart of a digital human video generation method based on a large model according to an embodiment of the present disclosure.
[0039] As Figure 2 shown, the digital human video generation method based on a large model includes operations S210 to S230.
[0040] In operation S210, demand information is obtained.
[0041] According to an embodiment of the present disclosure, the demand information includes action description information for describing a specified action video segment, and the action video segment represents a specified action of a target object. For example, the action video segment can represent a target object performing specified actions such as cutting vegetables and stir-frying. For another example, the action video segment can represent a target object performing specified actions such as rotating and displaying a commodity, unpacking, folding, tapping, and squeezing. The embodiments of the present disclosure do not limit the specific type of the specified action in the action video segment.
[0042] According to an embodiment of the present disclosure, the action description information is used to describe any action attributes such as the action type, action position, action amplitude, and action rate of the specified action. In addition, the action description information can also be used to describe action-related objects such as commodities and props related to the specified action. The embodiments of the present disclosure do not limit the specific information type described by the action description information, as long as it is related to the specified action.
[0043] It should be understood that the number of target objects in the action video clip can be one or more, and the action video clip can also represent one or more specified actions of the target object. The embodiments of the present disclosure do not limit the number of target objects in the action video clip or the number of specified actions represented. As long as the action description information can clearly describe the number of specified actions, the specified actions and
[0044] It should be noted that the target object can be a real person, but is not limited thereto. The target object can also be a virtual object such as a virtual digital human. The embodiments of the present disclosure do not limit the specific type of the target object.
[0045] In operation S220, the demand information is processed by the first large model to obtain a target script.
[0046] According to the embodiments of the present disclosure, the target script includes the target voiceover segment text that matches the action description information. The target voiceover segment text matches the action description information, which can be understood as that the text expression mode of the voiceover segment text matches the action expression mode or action intention of the specified action represented by the action description information.
[0047] In one example, the action description information can be the squeezing test action performed on the target commodity, and the target voiceover segment text can be understood as the voiceover content related to the squeezing test action, such as the commodity characteristics, squeezing deformation situation, and commodity deformation recovery shape that need to be explained during the process of the target object performing the squeezing test action on the target commodity.
[0048] In some embodiments, the first large model can be a language large model. The language large model (Large Language Model, LLM) is an artificial intelligence model constructed based on deep learning technology. The language large model usually has a huge number of parameters, and the number of parameters of the language large model can reach billions or even hundreds of billions. The huge number of parameters enables the language large model to capture the subtle features and complex patterns of language, so as to better understand the demand semantics of the demand information, and generate the target voiceover segment text that is more accurately adapted to the action description information based on the action description information. Thus, the target script can include the target voiceover segment text that drives the target object to perform voiceover expression and is adapted to the specified action, realizing that the voiceover expression mode of the target object in the target video is adapted to the action expression mode based on the adapted target voiceover segment text and action description information.
[0049] In operation S230, the target script and the action video clip are processed by the second large model to obtain a target video for displaying the target digital human performing voiceover based on the target voiceover segment text during the execution of the specified action.
[0050] In some embodiments, the second large model can be a visual large model.
[0051] According to an embodiment of the present disclosure, a Vision Large Model (VLM) can be an artificial intelligence model for processing or generating visual data such as images or videos. The Vision Large Model usually contains billions or even hundreds of billions of parameters and can generally be used to process multi-modal data such as text and images. The Vision Large Model can have the ability to generate images or videos. For example, the Vision Large Model can generate video data that matches the demand intention by processing multi-modal data such as text and images.
[0052] According to an embodiment of the present disclosure, the target digital human can be an object that is the same as or similar to the target object. For example, the target digital human and the target object can represent the same real person. Or the target digital human can be a digital human obtained by changing the appearance information of the target object such as the image, color, and clothing of the target object. To facilitate the explanation of the large model-based digital human video generation method provided by the embodiments of the present disclosure, the embodiments of the present disclosure use the "target object" to represent both the target digital human and the target object. The target object in the target video or dynamic video segment involved in the embodiments of the present disclosure can be understood as the target digital human, and the embodiments of the present disclosure will not elaborate here.
[0053] According to an embodiment of the present disclosure, by using the Vision Large Model to process the target script and the action video segment, the Vision Large Model can deeply understand the matching relationship between the text of the target voiceover segment in the target script and the action video segment, and fully understand the semantics of other script contents in the target script except for the text of the target voiceover segment and the action-related information such as the action attributes and action objects of the specified action. As a result, the generated target video can smoothly display that the target object, during the execution of the specified action, conducts a voiceover expression based on the target voiceover segment text in a suitable expression manner, enabling the target object in the target video to enhance the expression effect and expressiveness by naturally performing the suitable specified action during the voiceover expression process, thereby improving the quality of the target video.
[0054] It should be noted that the target video determined according to the large model-based digital human video generation method provided by the embodiments of the present disclosure can be applied to the live e-commerce scenario. For example, a live marketing video of a live streamer can be generated based on the large model-based digital human video generation method provided by the embodiments of the present disclosure. However, it is not limited to this. The large model-based digital human video generation method provided by the embodiments of the present disclosure can also be applied to any application scenarios such as animation production, film and television product production, and metaverse scene construction. The large model-based digital human video generation method provided by the embodiments of the present disclosure does not limit the specific application scenario.
[0055] It should be noted that the acquisition of information involved in any embodiment of the present disclosure, including but not limited to action video clips, requirement information, etc., is obtained under the authorization of the relevant object or institution. Before obtaining the information, the purpose is informed, and necessary encryption or desensitization measures are taken for the acquired information, which complies with the provisions of relevant laws and regulations and does not violate public order and good customs.
[0056] It should be noted that for the convenience of explaining the video generation method provided by the present disclosure, in the embodiments of the present disclosure, the first large model involved is exemplified by a language large model, and the second large model involved is exemplified by a vision large model. The language large model or vision large model involved in the embodiments of the present disclosure is not used to limit the model structure or type of the first large model or the second large model.
[0057] In one example, the action description information can be associated with the action video clip, and the associated action video clip can be determined based on the action description information in the requirement information, so as to use the vision large model to process the target script and the action video clip associated with the action expression information to obtain the target video.
[0058] According to the embodiments of the present disclosure, the requirement information may further include at least one of the object role attribute information of the target object, the target commodity information for the target video, and the target virtual prop information for the target video.
[0059] According to the embodiments of the present disclosure, the object role attribute information of the target object represents information related to the target object's attributes such as the role type, role setting, role personality, speech rate, intonation, etc. of the target object in the target video. By using the language large model to process the requirement information including the object role attribute information, the target script can be matched with the object attributes such as the role type and role setting of the target object, and the expression mode of the target voiceover segment text or other voiceover segment texts in the target script can conform to the object attributes of the target object, so as to improve the authenticity and naturalness of the expression of the target object in the target video and meet the user's needs. At the same time, the language large model can be used to process the requirement information including the object role attribute information to generate a target video in which multiple target objects conduct voiceover explanations, so as to improve the scene diversity of the target video and meet the actual needs of users to watch videos.
[0060] According to the embodiments of the present disclosure, the target commodity information can represent the attribute information of commodities such as shoes and clothes that need to be displayed in the target video. The commodity information can include any type of commodity attributes such as the color, size, specific mark, quantity, etc. of the commodity. The embodiments of the present disclosure do not limit the specific type of the commodity information.
[0061] According to the embodiments of the present disclosure, the target virtual props information for the target video may represent the attribute information of any virtual props element such as the lucky bag price element that needs to appear in the target video. For example, it may represent the type, size, color, display timing, etc. of the virtual props element. The embodiments of the present disclosure do not limit the specific type of the virtual props element represented by the target virtual props information.
[0062] According to the embodiments of the present disclosure, by utilizing a language macro model to process at least one of object attribute information, target product information, and target virtual prop information, the language macro model can be effectively controlled to combine the action attributes such as the action intention and action amplitude of the specified action represented by the action description information to generate script content that is more expressive and adapted to the action attributes, thereby making the expressiveness of the target object in the target video richer, and satisfying the matching of the expression style with the action intention of the specified action, thereby enhancing the diversity and expression ability of the target video.
[0063] In some embodiments, the demand information may also include object attribute information that needs to be modified for the target object's skin color, clothing, etc. The target digital person or target object in the target video may be an image updated according to the object attribute information of the demand information.
[0064] In one embodiment, the demand information may include object attribute information, target product information, and action description information. The object attribute information may include role information, expression information, etc., and the target product information includes information such as components and ingredients of brand A products. The target script obtained by processing the demand information using the language big model can be represented based on the following / / enclosed paragraph.
[0065] / / Host: (tone: calm) (action: pick up two opened boxes of Brand A, put down the box in the left hand, point to the contents of the box in the right hand and explain how to use it) Look, the method of use is very simple (the assistant also said: very simple!), once a month, use for 6 consecutive months, take a break for 6 months, and only need 2 boxes a year. Each time you use it, mix the A and B agents, and then use a roller to micro-prick, it is painless and the recovery period is short. Do not touch water within 12 hours after use, and pay attention to sun protection.
[0066] Host: (tone: calm) (action: pick up the packaged roller on the table to show, while the assistant uses his hands in front of his face to show how to use the roller) This roller is a tool that can be used together to better help nutrients penetrate and absorb. Just like loosening the soil, it allows the nutrients in the seeds to take root and sprout better, and the skin can better absorb nutrients.
[0067] Assistant host: (intonation: calm) (action: bring in a tool kit from off-screen, open it and show it, while the host opens the package of the roller and takes out the roller to show) Babies, everything in the tool kit is provided for you (the host says at the same time: very considerate!), you can use it directly as soon as you get it, which is very convenient. Moreover, our product doesn't hurt and has a short recovery period, and it won't affect your normal life.
[0068] Assistant host: (intonation: enthusiastic) (action: take out a bottle of essence and a bottle of kinetic energy essence from the A brand box above, show them side by side, and then put them back into the box) Look, these two bottles are Agent A and Agent B. When they are combined, they can exert powerful effects. Just like two superheroes joining hands, they can defeat various skin problems and make the skin become white, tender and smooth.
[0069] Host: (intonation: enthusiastic) (expression: happy) All friends, for such a good product and so many benefits, what are you still waiting for? The opportunity is rare and the inventory is limited. Miss this time, and it may take a long time to have such a discount again.
[0070] Assistant host: (intonation: enthusiastic) (action: pick up the KT board showing the comparison photos of using the A brand before and after, introduce the usage effect) (expression: surprised) Babies, take a look at these comparison pictures again. The effect is really immediate (the host says at the same time: this is exaggerated. Before using, there were various skin problems, and after using, the skin is like having a new layer, becoming white, tender and shiny. Do you all want to have such skin? / /
[0071] In this embodiment, in the same paragraph, the content following "Host:" is the target oral broadcast clip text of the host target object with the host role attribute. The content related to "(intonation: calm)" can indicate that the intonation attribute of the target object expressing this target oral broadcast clip text in the target video is calm. "(action: pick up the packaged roller on the table and show it, while the assistant host makes an example action of using the roller in front of the face with hands)" can represent the action description information of the host target object and the assistant host target object respectively. "This roller is the tool for paired use. It can better help the penetration and absorption of nutritional components. Just like loosening the soil of the land, it can make the seed nutritional components take root and germinate better, and make the skin absorb nutrition better." can represent the target oral broadcast clip text of the host target object. "(the assistant host says at the same time: very simple!)" represents the assistant host's echoing target oral broadcast clip text. "(expression: happy)" represents the expression attribute information of the host target object.
[0072] According to the target script provided in this embodiment, the vision large model can accurately capture the object attributes of multiple target objects with different character attributes and the semantics of the specified actions adapted to the target oral broadcast segment text by processing the target script and the action video segment, so that the generated target video can enable the multiple target objects to perform oral broadcasts according to their respective character attribute requirements and display the specified actions according to the choreography logic, so as to improve the expressiveness of the target video.
[0073] In one embodiment, the action description information can be determined based on the action video segment. For example, the action video segment can be post-processed by the vision understanding large model to obtain the action description information.
[0074] In one embodiment, the action video segment is determined based on the following operations: detecting the position change of the key points of the target object in the initial video to obtain the position change detection result; determining the initial action video segment from the initial video according to the position change detection result; detecting the action type of the initial action video segment to obtain the action type for the initial action video segment; and determining the initial action video segment that matches the preset action type as the action video segment.
[0075] According to an embodiment of the present disclosure, the initial video may include video frames depicting the target object performing one or more specified actions. By detecting the position change of the key points of the body parts such as the hands, legs, and torso of the target object in the initial video, and using the position change detection result to accurately determine the start time and end time of the specified action, the initial action video segment representing various action types can be more accurately determined from the initial video.
[0076] According to an embodiment of the present disclosure, by detecting the action type of the initial action video segment, the action video segment that matches the preset action type can be obtained from the initial action video segment, so as to more accurately screen out the action video segment representing the specified action with the preset action type. The initial video is clipped according to the action-related time information such as the start time and end time of the specified action to obtain the action video segment.
[0077] According to an embodiment of the present disclosure, the initial action video segment can be detected based on the target detection algorithm to obtain the initial action type. For example, the initial action video segment can be detected by a target detection model constructed based on the convolutional neural network algorithm. The specific method for determining the initial action type in the embodiments of the present disclosure is not limited.
[0078] In one embodiment, the action type of the specified action may be a high-performance action type. The specified action with a high-performance action type may be an action that accurately expresses object attribute information such as the emotion and role attributes of the target object. Or the specified action with a high-performance action type may also be an action that specifically demonstrates the performance, style, and appearance of the target commodity.
[0079] In one embodiment, the specified action may be a high-performance action, and the high-performance action may be composed of actions of multiple body parts of the target object such as hand actions and body trunk actions. The action video segment related to the high-performance action can be determined from the initial video through the following steps 1.1 to 1.7.
[0080] Step 1.1: Perform key point detection on the live video of real people as the initial video. The key points include the body trunk pose key points and gesture key points in the live video of real people.
[0081] Step 1.2: According to the key point positions of each video frame in the live video of real people. According to the key point positions, the body movement trajectory, left hand movement trajectory, and right hand movement trajectory of the target object can be determined as the position change detection results.
[0082] Step 1.3: Set an observation window with a window duration of 2s. Analyze the movement amplitude within the observation window according to various movement trajectories. When the position change amplitude of the key point positions between adjacent frames is greater than the set threshold, the time stamp of any frame in the adjacent two frames can be determined as the starting moment of the initial action.
[0083] Step 1.4: Move the observation window at a stride of 0.5s, continuously detect the position change amplitude of the key points within the observation window. If the position change amplitude is greater than the set threshold, repeat Step 1.4 until the live video of real people ends or the position change amplitude is less than the set threshold. Then, the time stamp of the video frame where the current observation window slides to is considered the termination of the specified action.
[0084] Step 1.5: Repeat Step 1.3 and Step 1.4 until the video ends to obtain the starting moments and ending moments of multiple initial action video segments.
[0085] Step 1.6: Split the initial video according to the starting moments and ending moments of the initial action video segments to obtain multiple initial action video segments.
[0086] Step 1.7: Use target detection models such as pose recognition models and gesture recognition models to identify the initial action types represented by each of the multiple initial action video segments, and use the preset action type to determine the action video segment representing the specified action from the multiple initial action segments.
[0087] It should be noted that the acquisition of information involved in this embodiment, including but not limited to information such as live video of real people, is obtained under the authorization of relevant objects or institutions. Moreover, the purpose of acquisition is informed before obtaining the information, and necessary encryption or desensitization measures are taken for the acquired information, which complies with the provisions of relevant laws and regulations and does not violate public order and good customs.
[0088] Figure 3 Schematically shows a schematic diagram of the principle for determining action video segments according to an embodiment of the present disclosure.
[0089] As Figure 3 shown, by observing the window, the position change of key points in the initial video 301 is detected, and the position change detection result may include at least one of the left - hand movement trajectory, the right - hand movement trajectory, and the body movement trajectory. By detecting that the position change amplitude of the key points of the left - hand movement trajectory is greater than a preset amplitude threshold, the starting time of the first left - hand movement in the first left - hand movement moment is obtained, and the detection continues from the starting time of the first movement until the position change amplitude of the key points of the left - hand is less than or equal to the preset amplitude threshold, and the ending time of the first left - hand movement in the first left - hand movement moment is obtained. The position change amplitude of the key points of the left - hand in the initial video 301 is repeatedly detected until the nth left - hand movement moment is obtained. Based on a detection method similar to or the same as that for the left - hand movement trajectory, the first right - hand movement moment to the nth right - hand movement moment, and the first body movement moment to the nth body movement moment can be obtained. According to the first body movement moment to the nth body movement moment, body posture recognition is performed on each initial action video segment corresponding to the first body movement moment to the nth body movement moment in the initial video 301, and the body action type is obtained. The body action type may be, for example, "step back". Gesture recognition is performed on each initial action video segment corresponding to the first left - hand movement moment to the nth left - hand movement moment in the initial video 301, and gesture recognition is performed on each initial action video segment corresponding to the first right - hand movement moment to the nth right - hand movement moment in the initial video 301, and the left - hand action type or the right - hand action type of the initial action video is obtained. The right - hand action type may be "gesture ratio 1". The action type may include the body action type, the left - hand action type, and the right - hand action type. By summarizing the action type of each initial action video segment and the action moment corresponding to the initial action video segment, the segmentation sub - task parameter 302 for segmenting the initial video 301 can be obtained. By using the action type in the segmentation sub - task parameter 302 to select the target sub - task parameter matching the specified action, and using the starting time start_time and the ending time end_time of the action in the target sub - task parameter to segment the initial video 301, the action video segment is obtained.
[0090] According to an embodiment of the present disclosure, the action description information may be generated by understanding an action video clip. For example, the action description information may be determined for the following operations: processing the action video clip and the text related to the clip of the action video clip by using a third large model to obtain the action description information.
[0091] In some embodiments, the third large model includes a multimodal large model. It should be noted that, for the convenience of describing the video generation method provided by the present disclosure, the third large model involved in the embodiments of the present disclosure may be described by taking the multimodal large model as an example. The multimodal large model involved in the embodiments of the present disclosure is not used to limit the specific model structure or type of the third large model.
[0092] According to an embodiment of the present disclosure, the text related to the clip may be understood as the text related to the action video clip. The text related to the clip may include the subtitle text of the action video clip, the voice-over text of the action video, the advertisement text displayed in the action video clip, the product name text, the background board text, and so on. The text related to the clip may be recognized from the video frames of the action video clip through text recognition technology, or may also be obtained by performing speech recognition on the speech segment data of the action video clip through speech recognition technology. The embodiments of the present disclosure do not limit the specific manner of recognizing the text related to the clip.
[0093] According to an embodiment of the present disclosure, the multimodal large model may be a large model capable of processing multimodal data such as images and texts. The multimodal large model may understand the action attribute information of a specified action, such as the action intention, action type, action amplitude, and action object of the specified action, by processing the action video clip and the text related to the clip, so that the action description information can more accurately represent the action attribute information, enabling the language large model to more accurately generate the target voice-over clip text that matches the action attribute information according to the more accurate action description information, thereby improving the expressiveness and video richness of the target object in the target video by performing the specified action that matches the voice-over clip text.
[0094] In one embodiment, the action description information is determined based on the following operations: processing the action video clip and the text related to the clip of the action video clip by using a multimodal large model to obtain the action attribute information; and processing the action attribute information by using a language large model to obtain the action description information that matches the semantics of the action intention data.
[0095] According to an embodiment of the present disclosure, the action intention data in the action attribute information is used to represent the explanation intention corresponding to a specified action. For example, the action intention data can represent the intention of the target object to explain the anti-torsion performance of the target commodity, or can also represent the intention of explaining the effect after applying the target commodity such as skin cream. The embodiment of the present disclosure does not limit the type of the explanation intention represented by the action intention data.
[0096] According to an embodiment of the present disclosure, the action description information is used to describe at least one of the following action attribute information: the item information related to the specified action, the virtual prop information related to the specified action, the object role information of the target object performing the specified action, and the action type information of the specified action.
[0097] According to an embodiment of the present disclosure, by using a multimodal large model to generate object attribute information including action intention data, and then using a language large model to process the object attribute information to obtain action description information that can accurately represent action attributes such as the explanation intention for a specified action, it is possible to filter out redundant description information for the specified action in the action video segment and improve the representation accuracy and naturalness of the action description information in representing object attributes such as action intention in a collaborative manner based on multiple large models. Thus, accurate action description information can be used to prompt the language large model to generate a target script with natural expression and logical coherence. Based on the target script as a prompt word for the visual large model, the visual large model is prompted to generate a target video in which the oral expression and action expression of the target object in the target video are adapted to each other, so as to control multiple large models to collaborate to generate a high-quality target video based on accurate action description information.
[0098] In one embodiment, it can be based on Figure 4 and the following embodiments to determine the action description information.
[0099] Figure 4 Schematically shows a schematic diagram for determining the action description information according to an embodiment of the present disclosure.
[0100] As Figure 4As shown, by performing automatic speech recognition (ASR) on the speech segment data of the action video segment, speech subtitles for the action video segment are obtained. By sampling the video frames of the action video segment, the sampled video frames are obtained. Based on optical character recognition (OCR) technology, text recognition is performed on the video frames to obtain the video text that appears in the action video segment. The multimodal large model 410 is used to process the speech subtitles, the sampled video frames, and the video text to obtain action attribute information. The action attribute information includes character attributes, initial action descriptions, product information, prop information, and action intentions. By using the language large model 420 to process the action attribute information, redundant information removal and information correction are performed on the action attribute information to obtain action description information.
[0101] According to an embodiment of the present disclosure, using the first large model to process the requirement information to obtain the target script includes: using the language large model to process the requirement information to obtain the script outline for the target video; performing knowledge retrieval based on the script outline to obtain the script materials for the target voiceover segment text; and using the language large model to process the script materials to obtain the target script.
[0102] According to an embodiment of the present disclosure, the script materials indicate the knowledge that matches the requirement intention represented by the requirement information. The requirement information includes the action description information, and at least one of the object role attribute information of the target object, the target product information for the target video, and the target virtual prop information for the target video. The script outline can represent the planning framework information of the target script to be generated in multiple dimensions such as the script content theme, character attribute setting, script plot structure, and script style positioning. By using the relatively powerful semantic understanding ability and thinking ability of the language large model to generate the script outline, it is possible to perform an overall planning of the target script according to the requirement intention represented by the requirement information. Furthermore, the script materials retrieved based on the script outline can meet the professional knowledge requirements for generating the target script. Thus, the language large model can be used to process the script materials with relatively accurate knowledge materials to generate the target script, so as to improve the professionalism and richness of the subsequent target video, and improve the expression level of professional knowledge in the target video based on the target voiceover text expression and specified action expression that match the target object.
[0103] In one embodiment, using the language large model to process the script materials to obtain the target script includes using the language large model to process the script materials and the script outline to obtain the target script. Thus, the generation of the target script that matches the requirement intention of the requirement information by the language large model can be restricted based on the creative intention framework represented by the script outline, and the creation quality of the target script can be improved.
[0104] According to an embodiment of the present disclosure, performing knowledge retrieval based on a script outline to obtain script materials for the text of a target voiceover segment includes: processing the script outline using a language model to obtain query information; and performing knowledge retrieval based on the query information to obtain script materials.
[0105] In one embodiment, performing knowledge retrieval based on a script outline may include retrieving in a professional knowledge base based on query information such as keywords and semantic vectors in the script outline.
[0106] In one embodiment, performing knowledge retrieval based on a script outline may include processing requirement information and the script outline using a language model to obtain query information, and using the query information to perform retrieval in a professional knowledge base. Thus, based on obtaining the script outline through systematic planning of the creative intent of scriptwriting based on requirement information, relevant knowledge writing materials can be actively collected around the key planning directions of the planned script outline, improving the professionalism, accuracy of the script materials, and the consistency with the overall creative intent of the script, thereby enhancing the logical coherence of the target script and the expression effect of the text of the target voiceover segment in the target script.
[0107] In one embodiment, performing knowledge retrieval based on the query information to obtain script materials includes: performing knowledge retrieval based on the query information to obtain initial script materials; performing semantic relevance detection on the initial script materials and preset requirement conditions to obtain a defect detection result indicating that the initial script materials do not meet the preset requirement conditions; processing the defect detection result and the script outline using a language model to obtain updated query information; and performing knowledge retrieval based on the updated query information to obtain script materials.
[0108] According to an embodiment of the present disclosure, the preset requirement conditions may be conditions for constraining the relevance between the script materials and the creative intent of the target script. The preset requirement conditions may be determined based on requirement information, or may be determined in other ways. For example, the preset requirement conditions may be determined based on information input through interactive operations. The preset requirement conditions are, for example, the number of script materials, the knowledge types represented by the script materials, etc. The embodiments of the present disclosure do not limit the specific setting method of the preset requirement conditions.
[0109] In one embodiment, the initial script material and preset requirement conditions can be processed based on a language large model to obtain a defect result representing the defect type of the initial script material. Thus, the initial script material that has been generated currently can be reflected on and inspected based on the cooperation of the language large model, and the defect result representing the defect type is used as a prompt to control the language large model to adjust the retrieval strategy by processing the defect detection result and the script outline, so that the updated query information can adjust the semantic accuracy of the query information or the implied range of the query information. Thus, the knowledge retrieval process can be continuously optimized iteratively based on the updated query information, and a deep cyclic optimization retrieval link of thinking, retrieving, reflecting, and retrieving again can be realized based on the language large model until the script material that meets the preset requirement conditions is generated, so that the script material can achieve the characteristics of clear content logic, sufficient knowledge information support, and strong executability, thereby improving the creation quality of the target script.
[0110] According to an embodiment of the present disclosure, using a language large model to process script material to obtain a target script may include: in the case where the script outline or requirement information indicates that the target script is a long script exceeding a preset word count threshold or a preset video duration threshold, using the language large model to process the script material to obtain a first script segment. Using the language large model to process the script material and the first script segment to obtain a second script segment. Thus, the language large model can be used to iteratively process the currently generated script segments and script material to obtain updated script segments until the last script segment is generated. The generated multiple script segments are fused to obtain the target script.
[0111] In one embodiment, the script segment may include the target voiceover segment text. Using the language large model to process the script material to obtain a target script includes: using the language large model to process the script material to obtain a first target voiceover segment text; using the language large model to process the script material and the first target voiceover segment text to obtain a second target voiceover segment text; and determining the target script based on the first target voiceover segment text and the second target voiceover segment text.
[0112] By using the language large model to process the script material and the currently generated target voiceover segment script, it is possible to realize script creation based on a multi-step cooperation method based on the language large model, so that the target voiceover segment text output each time can be more accurately matched with the action attribute information of the corresponding specified action, thereby improving the quality of the target script.
[0113] In one embodiment, in the case where the script outline or requirement information indicates that the target script is a short script not exceeding a preset word count threshold or a preset video duration threshold, the language large model can be used to process the script material to directly obtain the target script.
[0114] In one embodiment, a language large model is used to process the script material to obtain the text of the first target voice-over segment, which may include using the language large model to process the script material and the script outline to obtain the text of the meaningful target voice-over segment. Using the language large model to process the script material and the first script segment to obtain the second script segment may include using the language large model to process the script material, the script outline, and the first script segment to obtain the second script segment.
[0115] In one embodiment, the generated target script can be verified and evaluated based on an evaluation module to evaluate whether the currently generated target script meets the preset script conditions according to the script detection results. In the case where the currently generated target script does not meet the preset script conditions, based on the reflection mechanism of the language large model, the language large model can be used to process the script defect type, the script material, and the script outline represented by the script detection results to iteratively optimize the current target script until a target script that meets the preset script conditions is output. Thereby, the quality of the target video can be improved.
[0116] Figure 5 Schematically shows a schematic diagram of determining a target script according to an embodiment of the present disclosure.
[0117] As Figure 5 shown, the requirement information 501 includes the object attribute information of the target object, the target commodity information, the action description information, and the live broadcast room prop information as the target virtual prop information. The requirement information 501 is input into the script material planning module M510, and the script material is output. The script generation module M520 processes the script material and outputs the target script 502.
[0118] Among them, the script generation module M520 includes a content planning component, a query information generation component, a knowledge retrieval component, and a reflection and optimization component. The content planning component performs script content planning on the target script by calling the language large model to process the requirement information 501 to obtain a script outline. The query information generation component obtains query information for knowledge retrieval by calling the language large model to process the script outline and the requirement information 501. The knowledge retrieval component performs a knowledge retrieval operation using the current query information obtained from the query information generation component to obtain the initial script material. The reflection and optimization component calls the language large model to process the initial script material generated by the knowledge retrieval component in each round to implement semantic relevance detection of the initial script material and the preset requirement conditions. And in the case where the detection result is a defect detection result, the defect detection result is transmitted to the query information generation component. So that the query information generation component can generate updated initial script material by calling the language large model to process the script outline, the requirement information 501, and the defect detection result. After iteratively detecting the already generated initial script material, the reflection and optimization component determines the script material that meets the preset requirement conditions.
[0119] The script generation module M520 includes a script generation component and a script evaluation component. The script generation component can generate a short script of the short script type by calling a large language model to process at least one of the script material and the script outline. Or, according to the long script type indicated by the requirement information 501, the script generation component can also generate the text of the first target voiceover segment by calling a large language model to process at least one of the script material and the script outline. By calling a large language model to process at least one of the script material and the script outline, and the text of the first target voiceover segment, the text of the second target voiceover segment is generated until the text of N target voiceover segments is generated, and a long script of the long script type is obtained. Where N is an integer greater than 1. The script evaluation component calls a large language model to process the long script or the short script to perform evaluation operations such as verifying the long script or the short script and checking for logical errors, so as to obtain the target script 502 that meets the evaluation requirement conditions.
[0120] According to an embodiment of the present disclosure, the target script includes a plurality of target voiceover segment texts arranged in sequence; the arrangement positions of the plurality of target voiceover segment texts in the target script can be represented based on the script structure of the target script. A plurality of action video segments can correspond to the plurality of target voiceover segment texts.
[0121] According to an embodiment of the present disclosure, using the second large model to process the target script and the action video segments includes: using a vision large model to process two associated action video segments among the plurality of action video segments to obtain a transition action video segment; and using a vision large model to process the target script, the associated action video segments, and the transition action video segment to obtain the target video.
[0122] In some examples, the transition action video segment indicates the transition action between two different specified actions represented by the two associated action video segments, and the two associated action video segments are determined based on the arrangement positions of the plurality of target voiceover segment texts in the target script. The two associated action video segments can be two action video segments with an action connection relationship among the plurality of action video segments. For example, they can be two action video segments corresponding to two adjacent target voiceover segment texts in the target script.
[0123] In some examples, a vision large model is used to process transition action prompts and two associated action video segments among multiple action video segments to obtain a transition action video segment. The transition action prompts can be used to control the vision large model to understand the specified actions corresponding to the two associated action video segments respectively, so that the transition action video segment can more accurately represent the transition action between the two different specified actions. Thus, by arranging the transition action video segment between the two different associated action video segments, the two different specified actions can be naturally transitioned through the transition action represented by the transition action video segment.
[0124] In some examples, by using a vision large model to process a target script, associated action video segments, and transition action video segments, based on the demand intention matching the demand information represented by the target script, the vision large model can be controlled to more naturally fuse the associated action video segments and the transition action video segments, so that when multiple target action video segments in the target video can be adapted to the expression mode of the target voiceover segment text, the multiple specified actions can be vividly and naturally demonstrated through the target transition action video segment in the target action video segment, improving the display effect of the target video.
[0125] In some examples, using a vision large model to process a target script, associated action video segments, and transition action video segments to obtain a target video includes: using the vision large model to process the object attribute information, associated action video segments, and transition action video segments in the target voiceover segment text to obtain an intermediate video; and driving the lip movements of the target object in the intermediate video based on the voiceover audio data determined according to the target voiceover segment text to obtain the target video.
[0126] The voiceover audio data can be audio data representing the voiceover segment text in the target script. The intermediate video can be a video that can more accurately and naturally display multiple specified actions. By driving the lip movements of the target object in the intermediate video with the voiceover audio data, the lip movements of the target object in the target video can be adapted to the voice information of the voiceover audio data during the process of voiceover expression, further improving the naturalness and expressiveness of the target video.
[0127] It should be noted that any type of algorithm can be used to process the voiceover audio data and the intermediate video to achieve driving the lip movements of the target object in the intermediate video based on the voiceover audio data determined according to the target voiceover segment text to obtain the target video. For example, a diffusion model can be used to process the voiceover audio data and the intermediate video, but it is not limited to this. Other types of algorithms can also be used to process the voiceover audio data and the intermediate video. The embodiments of the present disclosure do not limit the specific manner of driving the lip movements of the target object in the intermediate video.
[0128] In one embodiment, the oral speech data and the intermediate video can be processed based on a vision large model to obtain a target video, thereby improving the adaptability between the lip movements of the target object in the target video and the oral speech data.
[0129] Figure 6 Schematically shows a schematic diagram of generating a target video according to an embodiment of the present disclosure.
[0130] As Figure 6 shown, the target script can include three different script parts, namely, the A product script part, the B product script part, and the C product script part. The action description information in the A product script part can determine that the action video segments used to generate the target video in the action video segment library are A1 and A2. The action description information in the B product script part can determine that the action video segments used to generate the target video are B1 and B3. The action description information in the C product script part can determine that the action video segments used to generate the target video are C1, C2, and C4. According to the positions of the action description information corresponding to the multiple action video segments A1 and A2, B1 and B3, and C1, C2, and C4 in the target script, the arrangement order of the multiple action video segments A1 and A2, B1 and B3, and C1, C2, and C4 in the initial video segment sequence can be determined.
[0131] By using the vision large model to process the context script information and context voice audio information in the target script related to the action video segments A1 and A2, and processing the action video segments A1 and A2 to obtain the transition action video segment A12 of the action video segments A1 and A2 as associated action video segments. By using the vision large model to process the action video segment A1 and the context script information and context voice audio information in the target script related to the action video segment A1, the transition action video segment A11 of the action video segment A1 as an associated action video segment is obtained, and the transition action video segment A11 can transition the start video content of the target video to the specified action represented by the action video segment A1. By using the language large model to process the associated action video segments, multiple transition action video segments A22, B13, C11, C12, and C24 can be obtained. It should be understood that the transition action video segment A22 can represent the transition action between the specified actions of the action video segments A2 and B1, the transition action video segment B13 can represent the transition action between the specified actions of the action video segments B1 and B2, the transition action video segment C11 can represent the transition action between the specified actions of the action video segments B2 and C1, the transition action video segment C12 can represent the transition action between the specified actions of the action video segments C1 and C2, and the transition action video segment C24 can represent the transition action between the specified actions of the action video segments C2 and C4.
[0132] In one example, the visual large model can be used to process the last video frame of the action video clip A1 and the starting video frame of A2 to obtain the transitional action video clip A12.
[0133] By inserting multiple transitional action video clips A11, A12, A22, B13, C11, C12, and C24 into the corresponding positions of the initial video clip sequence, a video clip sequence in which the transitional action video clips and the action video clips are arranged in order is obtained. Based on the voiceover audio data, the video clip sequence is fused to obtain the target video audio sequence. The visual large model is used as an expression driving model, and the visual large model is used to process the object attribute information corresponding to the action video clip or the transitional action video clip in the target script and the action video clip and the transitional action video clip in the target video audio sequence, so as to control the visual large model to control the expression of the target object during the voiceover in the action video clip or the transitional action video clip according to the expression attribute information such as the expression category and the expression position described in the target script, and an intermediate video is obtained. The lip driving model is used to drive the target object in the intermediate video to perform lip movements corresponding to the text in the target script according to the voiceover audio data, and thus the target video is obtained.
[0134] According to an embodiment of the present disclosure, the target video is determined by driving the lip movements of the target object based on the preset voiceover data, and the voiceover data is determined by the following operations: using a language large model to process the target script to obtain prosody features; and performing speech synthesis on the target script based on the prosody features to obtain the voiceover data.
[0135] In some embodiments, the prosody features characterize the voiceover prosody of the text sentences in the target script, and the voiceover prosody may include attribute information related to the voiceover expression rhythm, such as the pause mode, the pause duration, the intonation attribute, the voiceover speech rate, etc. By processing the target script with the language large model, the prosody features adapted to the expression mode of the specified action and the text of the voiceover segment can be output based on the relatively complete semantic of the voiceover segment text in the target script and the relatively accurate description method of the specified action by the action description information. Thus, by performing speech synthesis on the target script based on the prosody features, the obtained voiceover data can control the target object to perform a natural and vivid voiceover expression according to the demand intentions of the specified action, the text of the voiceover segment, and other object attribute information in the demand information. Furthermore, the accuracy and vividness of the expression of the target object in the target video can be improved by fusing the voiceover data and the intermediate video.
[0136] In some embodiments, performing speech synthesis on a target script based on prosodic features to obtain voiceover speech data may include: performing speech synthesis on a text sentence in the target script based on prosodic features to obtain sentence-level audio data; updating the audio time attributes of word-level sub-audio data in the sentence-level audio data based on the word-level time attributes of text words in the target script to obtain voiceover speech data.
[0137] In one embodiment, the sentence-level audio data may represent the voice audio data of a target object's voiceover expression of a text sentence in the target voiceover segment text. The audio time attributes of the word-level sub-audio data in the sentence-level audio data may be updated based on the timestamps of the text words in the text sentence in the target voiceover segment text as the word-level time attributes, so as to align the word-level sub-audio data in the sentence-level audio data with the text words in the target voiceover segment text, so that the voiceover speech data can accurately represent each text word in the target voiceover segment text according to the prosodic expression rhythm of the prosodic feature table, thereby improving the expression accuracy of the voiceover speech data for the target voiceover segment text.
[0138] Figure 7 Schematically shows a schematic diagram of determining voiceover speech data according to an embodiment of the present disclosure.
[0139] As Figure 7 shown, the voiceover speech data generated in this embodiment is used to generate a live video. It should be understood that the live video may be a target video. The live broadcast room server 701 may be a server for publishing live videos. The live broadcast room server 701 sends a speech synthesis request carrying the target script to the communication proxy module 702. The speech synthesis request is used to request the generation of voiceover speech data. The communication proxy module 702 sends the target script to the scheduling module 703. The scheduling module 703 processes the target script by calling the prosody large model 704 to obtain the prosodic features for the target script. The prosody large model 704 may also perform semantic analysis on the target script to clause the target script to obtain the text sentences in the target script, as well as the roles and prosodic features for the text sentences. The prosodic features may represent the voiceover expression rhythms of multiple target objects as different roles respectively in the target script, as well as the alternating expression rhythms between multiple target objects. The scheduling module 703 sends the N text sentences in the target script and the prosodic features of the target script to the audio synthesis module 705, so as to schedule the audio synthesis module 705 to perform speech synthesis on the N text sentences according to the prosodic features based on the text-to-speech (TTS) algorithm to obtain N sentence-level audio data respectively corresponding to the N text sentences. The text-to-speech algorithm may include speech synthesis algorithm models such as WaveNet. In addition, the audio synthesis module 705 may also perform word-level time attribute annotation on the text words in the text sentence to obtain the word-level timestamps of the text words in the text sentence.
[0140] The scheduling module 703 schedules the alignment model 706 to update the audio time attributes of the word-level sub-audio data of the sentence-level audio data by using the word-level timestamps of the text words in the text sentence, so as to obtain speech audio data with word-granularity time attribute alignment, by sending N text sentences and N sentence-level audio data to the alignment model. The alignment model 706 can perform multi-track alignment on the N sentence-level audio data. For example, the audio data can be separated according to the respective roles of multiple target objects, to obtain independent track data of the target objects for each role, and unmuted silent sub-data is filled in the silent parts in the independent track data to achieve precise synchronization of the independent tracks of the multi-role target objects. The scheduling module 703 sends the N sentence-level audio data, the time axis of the N text sentences of the target script, and the word-level timestamps in the text sentence to the communication proxy module 702, and the communication proxy module 702 uploads the complete speech audio data to the storage device of the live room server 701, so that the live room server can perform an asynchronous callback based on the storage access link of the speech audio data. The live room server 701 is enabled to call the complete speech audio data to drive the lip movement of the target object, and generate a live room video as the target video.
[0141] According to an embodiment of the present disclosure, the method for generating a digital human video based on a large model may further include the following operations: in response to a target interaction instruction, using a language large model to process the dynamic video requirement information for the target interaction instruction to obtain a dynamic video segment script; generating a video segment based on the dynamic video script to obtain a dynamic video segment; and inserting the dynamic video segment into the target video.
[0142] In some embodiments, the target interaction instruction may be obtained during the playback of the target video, and the target interaction instruction may be determined based on the user's interaction operation.
[0143] In some embodiments, the target interaction instruction includes at least one of a product order placement instruction, a comment generation instruction, and a like behavior instruction. The dynamic video requirement information corresponding to the target interaction instruction can characterize the interaction requirements of the user during the viewing of the target video. Using a language large model to process the dynamic video requirement information for the target interaction instruction can make the obtained dynamic segment script meet the interaction requirement intention of the user. Furthermore, a dynamic video segment can be obtained by using a vision large model to process the dynamic video segment script, and the dynamic video segment is inserted into the target video during playback to meet the interaction requirements of the user during the viewing of the target video.
[0144] In some embodiments, there is a mapping relationship between the dynamic video demand information and the target interaction instruction. The dynamic video demand information may include dynamic video action description information for dynamic video segments. By using a language large model to process the dynamic video demand information, the dynamic video segment text can drive the vision large model to generate a dynamic video segment in which the target object makes a vivid oral broadcast expression during the execution of the dynamic video action, so as to enhance the immersion degree of the user during the viewing of the dynamic video segment.
[0145] In some embodiments, in response to the target interaction instruction, the dynamic video demand information for the target interaction instruction is processed by using a language large model, and the obtained dynamic video segment script includes: in response to the target interaction instruction, making a dynamic video task decision based on the target interaction instruction to obtain a task decision result; using the language large model to process the task type and the context script content determined from the target script based on the insertion position information to obtain the dynamic video segment script.
[0146] In one embodiment, the task decision result includes the task type related to the dynamic video segment and the insertion position information of the dynamic video segment in the target video. The task decision result can be obtained by processing the target interaction instruction based on a preset decision model. The decision model can be constructed based on machine learning algorithms such as decision trees, or can also be constructed based on other types of algorithms. The embodiments of the present disclosure do not limit this.
[0147] In one embodiment, the context script content determined from the target script based on the insertion position information can represent the context script semantic content at the insertion position in the target script. Thus, by using the language large model to process the task type and the context script content, the language large model can generate a dynamic video segment text that can connect the target script content before and after the insertion position when fully understanding the context plot where the dynamic video segment should be inserted into the target video, so that the generated dynamic video segment can be displayed more naturally, vividly, and semantically logically after being inserted into the insertion position of the target video, so as to enhance the overall expressiveness and interactivity of the target video and improve the video quality.
[0148] In one example, in the dynamic video segment, a specified object can perform an oral broadcast according to the lip movements corresponding to the dynamic video segment text, and at the same time, play the oral broadcast audio data representing the dynamic video segment text to enhance the expressiveness level of the dynamic video segment.
[0149] Figure 8 Schematically shows a flowchart of determining a dynamic video segment according to an embodiment of the present disclosure.
[0150] As Figure 8As shown, the dynamic video segment can be determined based on operations S801 to S806.
[0151] In operation S801, an interaction event is detected. The target interaction instruction is determined by detecting interaction instruction parameters such as the number of interaction instructions and the type of interaction instructions.
[0152] In operation S802, a task decision is made according to the target interaction instruction, and a decision result on whether to trigger the dynamic video segment generation task is obtained, and the task type for the dynamic video segment is determined. In this embodiment, the target interaction instruction can be the comment content including the keyword "shelf life" posted by the user in the live broadcast room. The task type can be a dynamic video generation segment for displaying the shelf life.
[0153] In operation S803, according to the task decision result, the insertion position information of the dynamic video segment is determined from the target video being played. The insertion position indicated by the insertion position information can be a candidate insertion point preset before the target video is played, or the position between two different target action video segments in the target video can also be determined by detecting the playing progress of the current target video as the insertion position.
[0154] In operation S804, a dynamic video segment script is generated. The language large model is used to process the comment content of the target interaction instruction and the context script content determined from the target script based on the insertion position information to obtain dynamic query information. The retrieval enhancement strategy is used to perform knowledge retrieval using the dynamic query information and generate dynamic script materials. The language large model is used to process the dynamic script materials to obtain the dynamic video segment script as the reply content of the target interaction instruction.
[0155] In operation S805, a dynamic video segment is generated. The dynamic script materials are processed by a speech synthesis model to obtain dynamic voice audio data. The vision large model is used to process the dynamic script materials, so that the target object can be driven according to the action, expression and other information indicated by the dynamic script materials to generate a dynamic video segment. In addition, the target action video segment associated with the dynamic video segment can be determined through the insertion position information, and the vision large model is used to process the dynamic video segment and the target action video segment to obtain the transition action video segment between the dynamic video segment and the target action video segment. And the dynamic voice audio data is used to drive the lip movement of the target object in the dynamic video segment and the transition action video segment to obtain the dynamic video segment and the transition video segment for introducing the "shelf life" of the target commodity.
[0156] In operation S806, a dynamic video segment insertion task is performed. A dynamic video segment and a transition video segment for introducing the "shelf life" of the target product are inserted into the insertion position in the target video that is being played, so as to add the dynamic video segment and the transition video segment to the target video, or use the dynamic video segment and the transition video segment to replace part of the video segment content in the target video.
[0157] Figure 9 Schematically shows a schematic diagram of a digital human video generation method based on a large model according to an embodiment of the present disclosure.
[0158] As Figure 9 shown, the initial video is a stored live video. By using the action video segment analysis component 911 to call the multi-modal large model and the language large model to collaboratively process the action video segment in the live video and the video text for the live video, action description information is obtained. The video text may include the speech subtitle text of the live video, and the speech subtitle text is generated by the speech recognition component 912. The action video segment may be obtained by detecting the position change of the key points of the target object in the live video and segmenting the initial video according to the position change detection result.
[0159] The script generation module 920 includes a target script generation component 921 and a dynamic video segment script generation component 922. The target script generation component 921 is used to call the language large model to process the action description information to obtain the target script. The video generation component 931 of the video generation module 930 is used to call the visual large model to process the target script, the action video segment and the voiceover speech data, and obtain a digital human live video in which the target object performs a specified action and simultaneously performs a voiceover expression as the target video. The voiceover speech data may be generated by using the speech synthesis component 932 of the video generation module 930 to process the target script. The digital human live video includes a host object and an assistant object as different target objects, and they cooperate through a two-person echoing interactive dialogue and other methods to perform a voiceover expression. At the same time, the voice of the voiceover expression can be adapted to the specified action, and rich information such as matching expressions and intonations can be displayed.
[0160] During the playback of the digital human live video, the dynamic video clip script generation component 922 can detect the target interaction instruction, generate a dynamic video clip generation task, and determine the insertion position information to be inserted into the digital human live video. By calling the language large model to process the dynamic video requirement information indicated by the target interaction instruction and the context script content corresponding to the insertion position information, a dynamic video clip script is obtained. The dynamic video clip generation component 933 of the video generation module 930 generates a dynamic video clip by calling the vision large model to process the dynamic video clip script, the action video clip corresponding to the dynamic action description information in the dynamic video clip script, and the dynamic voice-over speech data representing the dynamic video clip script. The dynamic voice-over speech data can be determined by the speech synthesis component 932 by calling the speech synthesis model to process the dynamic video clip script. By inserting the dynamic video clip into the insertion position in the digital human live video during the live broadcast, the interaction experience with the user during the live broadcast can be improved.
[0161] Based on the large model-based digital human video generation method provided in the above embodiments, an embodiment of the present disclosure also provides a large model-based digital human video generation device.
[0162] Figure 10 A block diagram of a large model-based digital human video generation device according to an embodiment of the present disclosure is schematically shown.
[0163] As Figure 10 shown, the large model-based digital human video generation device 1000 includes: an acquisition module 1010, a target script acquisition module 1020, and a target video acquisition module 1030.
[0164] The acquisition module 1010 is configured to acquire requirement information, where the requirement information includes action description information for describing a specified action video clip, and the action video clip represents a specified action of a target object.
[0165] The target script acquisition module 1020 is configured to process the requirement information by using a first large model to obtain a target script, where the target script includes a target voice-over segment text that matches the action description information.
[0166] The target video acquisition module 1030 is configured to process the target script and the action video clip by using a second large model to obtain a target video for displaying the target digital human performing a voice-over based on the target voice-over segment text during the execution of the specified action.
[0167] According to an embodiment of the present disclosure, the action description information is determined based on the following operations: using a third large model to process the action video clip and the segment-related text for the action video clip to obtain the action description information.
[0168] According to an embodiment of the present disclosure, the action description information is determined based on the following operations: processing the action video segment and the segment-related text for the action video segment using a third large model to obtain action attribute information, where the action intention data in the action attribute information is used to represent the explanation intention corresponding to the specified action; and processing the action attribute information using a language large model to obtain action description information that matches the semantics of the action intention data.
[0169] In some embodiments, the third large model includes a multimodal large model.
[0170] According to an embodiment of the present disclosure, the action description information is used to describe at least one of the following action attribute information: item information related to the specified action, virtual prop information related to the specified action, object role information of the target object performing the specified action, action type information of the specified action.
[0171] According to an embodiment of the present disclosure, the target script obtaining module 1020 includes: a first processing sub-module, a retrieval sub-module, and a target script obtaining sub-module.
[0172] The first processing sub-module is configured to process the requirement information using a language large model to obtain a script outline for the target video.
[0173] The retrieval sub-module is configured to perform knowledge retrieval based on the script outline to obtain script materials for the target voice-over segment text, and the script materials indicate knowledge that matches the requirement intention represented by the requirement information.
[0174] The target script obtaining sub-module is configured to process the script materials using a language large model to obtain the target script.
[0175] According to an embodiment of the present disclosure, the retrieval sub-module includes: a query information obtaining unit and a retrieval unit.
[0176] The query information obtaining unit is configured to process the script outline using a language large model to obtain query information.
[0177] The retrieval unit is configured to perform knowledge retrieval based on the query information to obtain the script materials.
[0178] According to an embodiment of the present disclosure, the retrieval unit includes: an initial script material obtaining sub-unit, a defect detection result obtaining sub-unit, a query information obtaining sub-unit, and a script material obtaining sub-unit.
[0179] The initial script material obtaining sub-unit is configured to perform knowledge retrieval based on the query information to obtain initial script materials.
[0180] A defect detection result acquisition subunit, configured to perform semantic relevance detection on the initial script material and the preset requirement conditions, and obtain a defect detection result indicating that the initial script material does not meet the preset requirement conditions.
[0181] A query information acquisition subunit, configured to use a large language model to process the defect detection result and the script outline, and obtain updated query information.
[0182] A script material acquisition subunit, configured to perform knowledge retrieval based on the updated query information, and obtain script materials.
[0183] According to an embodiment of the present disclosure, the target script acquisition sub-module includes: a first acquisition unit, a second acquisition unit, and a target script acquisition unit.
[0184] The first acquisition unit is configured to use a large language model to process the script material, and obtain a first target voice-over segment text.
[0185] The second acquisition unit is configured to use a large language model to process the script material and the first target voice-over segment text, and obtain a second target voice-over segment text.
[0186] The target script acquisition unit is configured to determine a target script based on the first target voice-over segment text and the second target voice-over segment text.
[0187] According to an embodiment of the present disclosure, the requirement information further includes at least one of the following: object role attribute information of the target object, target commodity information for the target video, and target virtual prop information for the target video.
[0188] According to an embodiment of the present disclosure, the target script includes a plurality of target voice-over segment texts arranged in sequence; wherein, the target video acquisition module 1030 includes: a transition action video segment acquisition sub-module and a target video acquisition sub-module.
[0189] The transition action video segment acquisition sub-module is configured to use a visual large model to process two associated action video segments among a plurality of action video segments, and obtain a transition action video segment, wherein the transition action video segment indicates a transition action between two different specified actions represented by the two associated action video segments, and the two associated action video segments are determined based on the arrangement positions of the plurality of target voice-over segment texts in the target script.
[0190] The target video acquisition sub-module is configured to use a visual large model to process the target script, the associated action video segments, and the transition action video segments, and obtain a target video.
[0191] According to an embodiment of the present disclosure, the target video acquisition sub-module includes: an intermediate video acquisition unit and a target video acquisition unit.
[0192] An intermediate video acquisition unit, configured to process object attribute information, associated action video segments, and transition action video segments in the text of the target oral broadcast segment by using a vision large model, so as to obtain an intermediate video.
[0193] A target video acquisition unit, configured to drive the lip movements of a target object in the intermediate video based on the oral broadcast speech data determined according to the text of the target oral broadcast segment, so as to obtain a target video.
[0194] According to an embodiment of the present disclosure, the target video is determined by driving the lip movements of the target object based on preset oral broadcast speech data, where the oral broadcast speech data is determined through the following operations: processing a target script by using a language large model to obtain prosody features, where the prosody features represent the oral broadcast expression prosody of text sentences in the target script; and performing speech synthesis on the target script based on the prosody features to obtain the oral broadcast speech data.
[0195] According to an embodiment of the present disclosure, performing speech synthesis on the target script based on the prosody features to obtain the oral broadcast speech data includes: performing speech synthesis on the text sentences in the target script based on the prosody features to obtain sentence-level audio data; and updating the audio time attributes of word-level sub-audio data in the sentence-level audio data based on the word-level time attributes of text words in the target script to obtain the oral broadcast speech data.
[0196] According to an embodiment of the present disclosure, the digital human video generation device 1000 based on a large model further includes: a dynamic video segment script acquisition module, a dynamic video segment generation module, and an insertion module.
[0197] The dynamic video segment script acquisition module is configured to, in response to a target interaction instruction, process dynamic video requirement information for the target interaction instruction by using a language large model to obtain a dynamic video segment script.
[0198] The dynamic video segment generation module is configured to generate video segments based on the dynamic video script to obtain dynamic video segments.
[0199] The insertion module is configured to insert the dynamic video segments into the target video.
[0200] According to an embodiment of the present disclosure, the dynamic video segment script acquisition module includes: a task decision sub-module and a dynamic video segment script acquisition sub-module.
[0201] The task decision sub-module is configured to, in response to a target interaction instruction, perform dynamic video task decision based on the target interaction instruction to obtain a task decision result, where the task decision result includes a task type related to the dynamic video segment and insertion position information of the dynamic video segment in the target video.
[0202] A dynamic video clip script acquisition sub-module, which is used to process a task type and context script content determined from a target script based on insertion position information by using a language large model, so as to obtain a dynamic video clip script.
[0203] According to an embodiment of the present disclosure, the target interaction instruction includes at least one of the following: a commodity order placing instruction, a comment generation instruction, and a like behavior instruction.
[0204] According to an embodiment of the present disclosure, the action video clip is determined based on the following operations: performing position change detection on key points of a target object in an initial video to obtain a position change detection result; determining an initial action video clip from the initial video according to the position change detection result; performing action type detection on the initial action video clip to obtain an action type for the initial action video clip; and determining the initial action video clip that matches a preset action type as the action video clip.
[0205] Figure 11 A structural block diagram of an intelligent agent of artificial intelligence according to an embodiment of the present disclosure is schematically shown.
[0206] In an embodiment of the present disclosure, as Figure 11 shown, the AI intelligent agent 1100 may include an input module 1110, a processing module 1120, and an output module 1130.
[0207] The input module 1110 is used to receive input information;
[0208] The processing module 1120 is used to determine a target task based on the input information received by the input module, determine a language large model and a vision large model based on the target task, and execute the large model-based digital human video generation method provided by the embodiment of the present disclosure by calling the language large model and the vision large model to obtain output information;
[0209] The output module 1130 is used to output the output information obtained by the processing module.
[0210] According to an embodiment of the present disclosure, the input module 1110 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the outside world (such as users or the external environment), and converting it into a format that the AI intelligent agent 1100 can understand and process. The input module 1110 is the primary link for the AI intelligent agent 1100 to interact with the outside world, enabling the AI intelligent agent 1100 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.
[0211] In the example, the input module 1110 may input the demand information, action video clip, etc. described above.
[0212] In the example, the processing module 1120 is the core support for the AI agent 1100 to handle complex tasks. The processing module 1120 can execute the digital human video generation method based on the large model described above.
[0213] In the example, the performance of the processing module 1120 can be closely related to the large model on which the AI agent 1100 is based. To fully utilize the capabilities of the large model, the internal structure of the processing module 1120 can be designed to be highly configurable and extensible to handle various different types of tasks and requirements in real-world scenarios.
[0214] In the example, after the AI agent 1100 obtains the requirement information, the processing module 1120 can use the language large model to process the requirement information to obtain the target script, use the vision large model to process the target script and the action video clips to obtain the target video, and transmit the target video to the output module 1130.
[0215] It can be understood that although the language large model has excellent language understanding and generation capabilities, like humans, the tasks it can solve without any tools are very limited. When the AI agent 1100 is given the ability to call tools, it can perform tasks such as completing mathematical operations with the help of a calculator, performing data analysis with the help of Python, and obtaining weather forecasts with the help of a search engine.
[0216] In the example, the output module 1130 can output the target video described above.
[0217] The AI agent 1100 according to the embodiments of the present disclosure can simply and effectively improve the degree of intelligence, and improve flexibility and versatility.
[0218] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0219] According to the embodiments of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.
[0220] According to the embodiments of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.
[0221] According to the embodiments of the present disclosure, a computer program product includes a computer program, and the computer program implements the method as described above when executed by a processor.
[0222] Figure 12 FIG. 1 shows a schematic block diagram of an exemplary electronic device that can be used to implement embodiments of the large model-based digital human video generation method of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementations of the present disclosure described and / or claimed herein.
[0223] As Figure 12 shown, the device 1200 includes a computing unit 1201 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 12012 into a random access memory (RAM) 1203. In the RAM 1203, various programs and data required for the operation of the device 1200 can also be stored. The computing unit 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0224] A plurality of components in the device 1200 are connected to the I / O interface 1205, including: an input unit 1206, such as a keyboard, a mouse, etc.; an output unit 1207, such as various types of displays, speakers, etc.; a storage unit 12012, such as a magnetic disk, an optical disk, etc.; and a communication unit 1209, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1209 allows the device 1200 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0225] The computing unit 1201 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 executes the various methods and processes described above, such as the digital human video generation method based on a large model. For example, in some embodiments, the digital human video generation method based on a large model can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 12012. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into the RAM 1203 and executed by the computing unit 1201, one or more steps of the digital human video generation method based on a large model described above can be executed. Alternatively, in other embodiments, the computing unit 1201 can be configured to execute the digital human video generation method based on a large model in any other suitable manner (e.g., by means of firmware).
[0226] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0227] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0228] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0229] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0230] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0231] A computer system may include a client and a server. The client and the server are generally far away from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0232] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.
[0233] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for generating a digital human video based on a large model, comprising: Obtaining requirement information, where the requirement information includes action description information for describing a specified action video segment, and the action video segment represents a specified action of a target object; Processing the requirement information by using a first large model to obtain a target script, where the target script includes target oral broadcast segment texts that match the action description information; Processing the target script and the action video segment by using a second large model to obtain a target video for displaying that the target digital human performs an oral broadcast based on the target oral broadcast segment texts during the execution of the specified action.
2. The method according to claim 1, wherein, The action description information is determined based on the following operations: Processing the action video segment and segment-related texts for the action video segment by using a third large model to obtain the action description information.
3. The method according to claim 1 or 2, wherein The action description information is determined based on the following operations: Processing the action video segment and segment-related texts for the action video segment by using a third large model to obtain action attribute information, where the action intention data in the action attribute information is used to represent an explanation intention corresponding to the specified action; and Processing the action attribute information by using a language large model to obtain action description information that matches the semantics of the action intention data, and the third large model includes a multimodal large model.
4. The method according to claim 3, wherein The action description information is used to describe at least one of the following action attribute information: Item information related to the specified action, virtual prop information related to the specified action, object role information of the target object performing the specified action, action type information of the specified action.
5. The method according to claim 1, wherein, The processing the requirement information by using a first large model to obtain a target script includes: Processing the requirement information by using a language large model to obtain a script outline for the target video, and the first large model includes the language large model; Performing knowledge retrieval based on the script outline to obtain script materials for the target oral broadcast segment texts, where the script materials indicate knowledge that matches the requirement intention represented by the requirement information; and Processing the script materials by using the language large model to obtain the target script.
6. The method according to claim 5, wherein, The performing knowledge retrieval based on the script outline to obtain script materials for the target oral broadcast segment texts includes: Processing the script outline by using the language large model to obtain query information; Performing knowledge retrieval based on the query information to obtain the script materials.
7. The method according to claim 6, wherein The performing knowledge retrieval based on the query information to obtain the script materials includes: Performing knowledge retrieval based on the query information to obtain initial script materials; Performing semantic relevance detection on the initial script materials and preset requirement conditions to obtain a defect detection result indicating that the initial script materials do not meet the preset requirement conditions; Processing the defect detection result and the script outline by using the language large model to obtain updated query information; and Performing knowledge retrieval based on the updated query information to obtain the script materials.
8. The method according to claim 5, wherein The processing the script materials by using the language large model to obtain the target script includes: Process the script material using the language large model to obtain the first target voiceover segment text; Process the script material and the first target voiceover segment text using the language large model to obtain the second target voiceover segment text; and Determine the target script based on the first target voiceover segment text and the second target voiceover segment text.
9. The method according to any one of claims 1, 5 to 8, wherein, The requirement information further includes at least one of the following: The object role attribute information of the target object, the target commodity information for the target video, and the target virtual prop information for the target video.
10. The method according to claim 1, wherein, The target script includes a plurality of the target voiceover segment texts arranged in sequence; Among them, the processing of the target script and the action video segments using the second large model includes: Process two associated action video segments among the plurality of action video segments using the vision large model to obtain a transition action video segment, where the transition action video segment indicates the transition action between two different specified actions represented by the two associated action video segments, and the two associated action video segments are determined based on the arrangement positions of the plurality of target voiceover segment texts in the target script, and the second large model includes the vision large model; and Process the target script, the associated action video segments, and the transition action video segment using the vision large model to obtain the target video.
11. The method according to claim 10, wherein, The processing of the target script, the associated action video segments, and the transition action video segment using the vision large model to obtain the target video includes: Process the object attribute information in the target voiceover segment text, the associated action video segments, and the transition action video segment using the vision large model to obtain an intermediate video; and Drive the lip movements of the target object in the intermediate video based on the voiceover speech data determined according to the target voiceover segment text to obtain the target video.
12. The method according to claim 1, wherein The target video is determined based on preset voiceover speech data driving the lip movements of the target object, wherein the voiceover speech data is determined based on the following operations: Process the target script using the first large model to obtain a prosody feature, where the prosody feature characterizes the voiceover expression prosody of the text sentences in the target script; and Perform speech synthesis on the target script based on the prosody feature to obtain the voiceover speech data.
13. The method according to claim 13, wherein, The performing speech synthesis on the target script based on the prosody feature to obtain the voiceover speech data includes: Perform speech synthesis on the text sentences in the target script based on the prosody feature to obtain sentence-level audio data; Update the audio time attribute of the word-level sub-audio data in the sentence-level audio data based on the word-level time attribute of the text words in the target script to obtain the voiceover speech data.
14. The method according to claim 1, wherein The method further includes: In response to a target interaction instruction, process the dynamic video requirement information for the target interaction instruction using the language large model to obtain a dynamic video segment script; Generate video segments based on the dynamic video script to obtain dynamic video segments; and Insert the dynamic video segments into the target video.
15. The method according to claim 14, wherein, In response to the target interaction instruction, processing the dynamic video requirement information for the target interaction instruction using a language large model to obtain a dynamic video segment script, including: In response to the target interaction instruction, making a dynamic video task decision based on the target interaction instruction to obtain a task decision result, where the task decision result includes a task type related to the dynamic video segment and insertion position information of the dynamic video segment in the target video; Processing the task type and the context script content determined from the target script based on the insertion position information using the language large model to obtain the dynamic video segment script.
16. The method according to claim 1 or 2, wherein, The action video segment is determined based on the following operations: Performing position change detection on key points of the target object in the initial video to obtain a position change detection result; Determining an initial action video segment from the initial video based on the position change detection result; Performing action type detection on the initial action video segment to obtain an action type for the initial action video segment; and Determining the initial action video segment that matches the preset action type as the action video segment.
17. A digital human video generation device based on a large model, including: An acquisition module, configured to acquire requirement information, where the requirement information includes action description information for describing a specified action video segment, and the action video segment represents a specified action of a target object; A target script acquisition module, configured to process the requirement information using a first large model to obtain a target script, where the target script includes a target oral broadcast segment text that matches the action description information; A target video acquisition module, configured to process the target script and the action video segment using a second large model to obtain a target video for displaying the target digital human performing the specified action and making an oral broadcast based on the target oral broadcast segment text.
18. An intelligent agent of artificial intelligence, including: An input module, configured to receive input information; A processing module, configured to determine a target task based on the input information received by the input module, determine a first large model and a second large model based on the target task, and execute the method according to any one of claims 1 to 16 by calling the first large model and the second large model to obtain output information; An output module, configured to output the output information obtained by the processing module.
19. An electronic device, including: At least one processor; And A memory communicatively connected to the at least one processor; where The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 16.
20. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 16.
21. A computer program product, including a computer program, where the computer program implements the method according to any one of claims 1 to 16 when executed by a processor.
Citation Information
Cited By
Digital human high-quality video generation method and system based on data driving
CN120689477A
Large model-based audio and video data generation method, training method and intelligent agent
CN121306089A
Audio and video data generation methods, training methods, and intelligent agents based on large models
CN121306089B
Resource recommendation video generation method and device, electronic equipment and storage medium
CN121908077A
Digital human audio and video processing method and device, storage medium and program product
CN122205198A