Method and apparatus for generating digital human video based on large model, agent, electronic device, storage medium and computer program product
Large-scale models are used to generate digital human videos with coordinated and natural expressions, addressing the lack of movement expression in existing videos to enhance their quality and user experience.
Patent Information
- Application Number
- JP2025145788
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-04-25
- Filing Date
- 2025-09-03
- Publication Date
- 2025-12-16
AI Technical Summary
Digital human videos in e-commerce live broadcasts and movie viewing product production often lack movement expression, leading to poor video quality and a subpar user experience.
A method and system utilizing large-scale models, including a linguistic model to generate target broadcast segment text and a visual model to process the scenario and action video segments, ensuring coordinated and natural broadcast expressions that match the specified actions.
Improves the expressiveness and quality of digital human videos by adapting broadcast expressions to match action intentions, enhancing the naturalness and authenticity of the target object's movements and actions.
Smart Images

Figure 2025183260000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to technical fields such as deep learning, large-scale models, and computer vision, and can be applied to scenes such as live video, advertisement creation, and e-commerce. [Background technology]
[0002] With the rapid development of Internet technology, users can conveniently view resource information such as videos through smart terminal devices such as smartphones, etc. For example, users can view live videos through their smartphones to learn detailed product information. Summary of the Invention [Means for solving the problem]
[0003] The present disclosure provides a method, device, agent, electronic device, and storage medium for generating digital human video based on a large-scale model.
[0004] According to one aspect of the present disclosure, a method for generating digital human video based on a large-scale model is provided, including: obtaining demand information including action description information for describing a specified action video segment, where the action video segment represents a specified action of a target object; processing the demand information using a first large-scale model to obtain a target scenario including target broadcast segment text matching the action description information; and processing the target scenario and the action video segment using a second large-scale model to obtain a target video for displaying a target digital human broadcasting based on the target broadcast segment text in the process of performing the specified action.
[0005] According to another aspect of the present disclosure, a digital human video generation apparatus based on a large-scale model is provided, including: an acquisition module for acquiring demand information including action description information for describing a specified action video segment, wherein the action video segment represents a specified action of a target object; a target scenario acquisition module for processing the demand information using a first large-scale model to obtain a target scenario including target broadcast segment text matching the action description information; and a target video acquisition module for processing the target scenario and the action video segment using a second large-scale model to obtain a target video for displaying a target digital human broadcasting based on the target broadcast segment text in the process of performing the specified action.
[0006] According to another aspect of the present disclosure, an artificial intelligence agent is provided, including an input module for receiving input information, a processing module for identifying a target task based on the input information received by the input module, identifying a first large-scale model and a second large-scale model based on the target task, and obtaining output information by invoking the first large-scale model and the second large-scale model to perform the large-scale model-based digital human video generation method provided in the embodiments of the present disclosure, and an output module for outputting the output information obtained by the processing module.
[0007] According to another aspect of the present disclosure, there is provided an electronic device including at least one processor and a memory communicatively coupled to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor can perform a large scale model based digital human video generation method provided in an embodiment of the present disclosure.
[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to cause the computer to perform a large scale model-based digital human video generation method provided in an embodiment of the present disclosure.
[0009] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program for, when executed by a processor, implementing the large-scale model-based digital human video generation method provided in the embodiments of the present disclosure.
[0010] It should be noted that the contents described in this section are not intended to identify key points or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will be easily understood from the following description.
[0011] The drawings are intended to provide a better understanding of the invention and are not intended to limit the invention. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 illustrates a schematic diagram of an exemplary system architecture to which the content processing method and apparatus according to an embodiment of the present disclosure can be applied. [Figure 2] FIG. 2 schematically illustrates a flowchart of a method for generating digital human video based on a large-scale model according to an embodiment of the present disclosure. [Figure 3] FIG. 3 illustrates a principle diagram of identifying motion video segments according to an embodiment of the present disclosure. [Figure 4] FIG. 4 illustrates a schematic diagram for specifying behavioral description information according to an embodiment of the present disclosure. [Figure 5] FIG. 5 illustrates a schematic diagram of identifying a target scenario according to an embodiment of the present disclosure. [Figure 6] FIG. 6 illustrates a schematic diagram of generating a target video according to an embodiment of the present disclosure. [Figure 7] FIG. 7 illustrates a schematic diagram of identifying broadcast audio data according to an embodiment of the present disclosure. [Figure 8] FIG. 8 illustrates a flow chart diagram for identifying dynamic video segments according to an embodiment of the present disclosure. [Figure 9] FIG. 9 illustrates a schematic diagram of a large-scale model-based digital human video generation method according to an embodiment of the present disclosure. [Figure 10] FIG. 10 schematically illustrates a block diagram of a large-scale model-based digital human video generation device according to an embodiment of the present disclosure. [Figure 11] FIG. 11 is a schematic block diagram illustrating an artificial intelligence agent according to an embodiment of the present disclosure. [Figure 12] FIG. 12 shows a schematic block diagram of an exemplary electronic device for implementing the large-scale model-based digital human video generation method of an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0013]
[0023] Exemplary embodiments of the present disclosure will be described below with reference to the drawings. However, to facilitate understanding, various details of the embodiments of the present disclosure are included for illustrative purposes only. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of known functions and structures will be omitted in the following description.
[0014] In the technical solution disclosed herein, the acquisition, storage, application, etc. of relevant user personal information shall all comply with the provisions of relevant laws and regulations, adopt necessary confidentiality measures, and not violate public order and morals.
[0015] The inventor discovered that digital human videos generated in scenes such as e-commerce live broadcasts and movie viewing product production often lack movement expression, resulting in poor video quality and affecting users' video viewing experience.
[0016] Embodiments of the present disclosure provide a method, device, agent, electronic device, storage medium, and computer program product for generating digital human video based on a large-scale model, which includes: obtaining demand information including action description information for describing action video segments representing specified actions of a specified target object; processing the demand information using a first large-scale model to obtain a target scenario including target broadcast segment text matching the action description information; and processing the target scenario and the action video segments using a second large-scale model to obtain a target video for displaying a target digital human broadcasting based on the target broadcast segment text in the process of performing the specified action.
[0017] According to an embodiment of the present disclosure, a first large-scale model is used to process demand information including action description information, thereby generating target broadcast segment text that can match the specified action represented by the action description information, thereby adapting the semantics of the target broadcast segment text in a target scenario to the action intention represented by the specified action based on the high comprehension ability and text output ability of the first large-scale model. A second large-scale model is used to process a target scenario including the target broadcast segment text and action video segments to drive the target object and generate a target video, thereby enabling broadcast expressions to be performed based on the target broadcast segment text that adapts to the specified action during the process of the target object in the target video performing the specified action, and the broadcast expression and action expression of the target object in the target video are adapted in coordination, thereby improving the naturalness of the broadcast expression of the target object in the target video, improving the expressiveness of the target object, and further improving the quality of the target video.
[0018] FIG. 1 schematically illustrates an exemplary system architecture to which the large-scale model-based digital human video generation method and apparatus according to an embodiment of the present disclosure can be applied.
[0019] 1 merely illustrates an example of a system architecture to which the embodiments of the present disclosure can be applied, so as to allow those skilled in the art to understand the technical contents of the present disclosure, and does not mean that the embodiments of the present disclosure cannot be applied to other devices, systems, environments, or scenes. For example, in another embodiment, an exemplary system architecture to which the large-scale model-based digital human video generation method and apparatus can be applied may include a terminal device, but the terminal device does not need to interact with a server to realize the large-scale model-based digital human video generation method and apparatus according to the embodiments of the present disclosure.
[0020] 1, system architecture 100 according to this embodiment includes terminal devices 101, 102, 103, a network 104, and a server 105. Network 104 provides a medium for a communication link between terminal devices 101, 102, 103 and server 105. Network 104 may include various connection types, such as wired and / or wireless communication links.
[0021] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to send and receive messages, etc. Various communication client applications may be installed on terminal devices 101, 102, and 103, such as (by way of example only) knowledge browsing applications, web browser applications, search applications, instant messaging tools, mailbox clients, and / or social platform software.
[0022] The terminal devices 101, 102, and 103 may be various electronic devices that have a display and support web page browsing, including, but not limited to, smartphones, tablet computers, laptop computers, and desktop computers.
[0023] Server 105 may be a server that provides various services, such as a background management server (by way of example only) that provides support for content viewed by users using terminal devices 101, 102, and 103. The background management server performs processing such as analysis on data such as received user requests, and can feed back the processing results (e.g., web pages, information, or data obtained or generated based on the user requests) to the terminal devices.
[0024] The server 105 may be a cloud computing server or cloud host, which is a host product in a cloud computing service system, to solve the drawbacks of high management difficulty and poor service scalability that exist in conventional physical hosts and VPS services (abbreviated as "Virtual Private Server" or "VPS"). The server 105 may be a server in a distributed system or a server connected to a blockchain.
[0025] It should be noted that the large-scale model-based digital human video generation method according to the embodiments of the present disclosure may be generally performed by server 105. Accordingly, the large-scale model-based digital human video generation apparatus according to the embodiments of the present disclosure may be generally provided in server 105. The large-scale model-based digital human video generation method according to the embodiments of the present disclosure may be performed by a server or server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Accordingly, the large-scale model-based digital human video generation apparatus according to the embodiments of the present disclosure may be provided in a server or server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.
[0026] It should be understood that the number of terminal devices, networks, and servers in Figure 1 is merely an example, and any number of terminal devices, networks, and servers may be included as needed.
[0027] FIG. 2 schematically illustrates a flowchart of a method for generating digital human video based on a large-scale model according to an embodiment of the present disclosure.
[0028] As shown in FIG. 2, the large-scale model-based digital human video generation method includes operations S210 to S230.
[0029] In operation S210, demand information is acquired.
[0030] According to an embodiment of the present disclosure, the demand information includes action description information for describing a specified action video segment, where the action video segment represents a specified action of a target object. For example, the action video segment may represent the target object performing a specified action such as cutting, frying, etc. Also, for example, the action video segment may represent the target object performing a specified action on a product such as rotating, opening, folding, tapping, pressing, etc. The embodiment of the present disclosure does not limit the specific type of the specified action in the action video segment.
[0031] According to an embodiment of the present disclosure, the action description information is used to describe any action attribute of the specified action, such as the action type, action position, action amplitude, action speed, etc. The action description information may also be used to describe an action-related object, such as a product or item, related to the specified action. The embodiment of the present disclosure does not limit the specific information type described by the action description information, as long as it is related to the specified action.
[0032] It should be understood that the number of target objects in an action video segment may be one or more, and the action video segment may show one or more specified actions of the target object, and the embodiments of the present disclosure do not limit the number of target objects in an action video segment and the number of specified actions shown, as long as the action description information can clearly describe the number of specified actions.
[0033] It should be noted that the target object may be, but is not limited to, a real person, and the target object may also be a virtual object such as a virtual digital human, and the embodiments of the present disclosure do not limit the specific type of target object.
[0034] In operation S220, the first large-scale model is utilized to process the demand information to obtain a target scenario.
[0035] According to an embodiment of the present disclosure, the target scenario includes a target broadcast segment text that matches the action description information. The target broadcast segment text matching the action description information can be understood as the text expression manner of the broadcast segment text matching the action expression manner or action intention of the specified action expressed by the action description information.
[0036] In one example, the action description information may represent a pressure test operation to be performed on a target product, and the target broadcast segment text may be understood as broadcast content related to the pressure test operation, such as product characteristics, pressure deformation status, and product deformation recovery shape that need to be explained in the process of the target object performing the pressure test operation on the target product.
[0037] In some embodiments, the first large-scale model may be a language large-scale model. A language large-scale model (LLM) is an artificial intelligence model built based on deep learning technology. A language large-scale model typically has a large number of parameters, and the number of parameters for a language large-scale model can reach billions or even hundreds of billions. The large number of parameters allows the language large-scale model to capture subtle features and complex modes of language, better understand the semantics of the demand information, and more accurately generate target broadcast segment text that matches the action description information based on the action description information. This allows the target scenario to include driving the target object to broadcast expression. Then, by matching the target broadcast segment text of the specified action, it is possible to present the broadcast expression manner of the target object in the target video that matches the action expression manner based on the matched target broadcast segment text and the action description information.
[0038] In operation S230, the second large-scale model is used to process the target scenario and the action video segment to obtain a target video for displaying the target digital human broadcasting based on the target broadcast segment text in the process of performing the specified action.
[0039] In some embodiments, the second large scale model may be a visual large scale model.
[0040] According to an embodiment of the present disclosure, a Vision Large Model (VLM) may be an artificial intelligence model for processing or generating visual data such as images or videos. The Vision Large Model generally includes billions or even hundreds of billions of parameters and may be typically used to process multimodal data such as text, images, etc. The Vision Large Model may have image or video generation capabilities. For example, the Vision Large Model may process multimodal data such as text, images, etc. to generate video data that matches a demand intent.
[0041] According to embodiments of the present disclosure, the target digital human may be the same or similar object as the target object; for example, the target digital human and the target object may represent the same real person. Alternatively, the target digital human may be a digital human obtained by modifying the object appearance information, such as the image, color, or clothing, of the target object. To facilitate the description of the large-scale model-based digital human video generation method provided in embodiments of the present disclosure, the embodiments of the present disclosure refer to both the target digital human and the target object as the "target object." The target object of the target video or dynamic video segment according to embodiments of the present disclosure may be understood as the target digital human, and a description of the embodiments of the present disclosure will be omitted here.
[0042] According to an embodiment of the present disclosure, by using a visual large-scale model to process a target scenario and action video segments, the visual large-scale model can deeply understand the matching relationship between the target broadcast segment text and the action video segments in the target scenario, and can fully understand the semantics of other scenario content other than the target broadcast segment text in the target scenario and action-related information such as action attributes of the specified action and action objects. The generated target video broadcasts the target broadcast segment text based on a suitable expression method during the process of the target object performing the specified action, and the target object in the target video naturally performs the suitable specified action during the broadcast expression process, thereby improving the expression effect and expressive power and improving the quality of the target video.
[0043] The target video identified by the large-scale model-based digital human video generation method according to the embodiment of the present disclosure may be applied to live e-commerce scenes. For example, a live marketing video for a streamer may be generated based on the large-scale model-based digital human video generation method according to the embodiment of the present disclosure. However, the large-scale model-based digital human video generation method according to the embodiment of the present disclosure is not limited thereto. The large-scale model-based digital human video generation method according to the embodiment of the present disclosure may also be applied to any application scenario, such as video production, video product production, and meta-universe scene construction. The large-scale model-based digital human video generation method according to the embodiment of the present disclosure is not limited to a specific application scenario.
[0044] It should be noted that the acquisition of information in any embodiment of the present disclosure includes, but is not limited to, information such as motion video segments, demand information, etc., all of which are acquired with the permission of the relevant object or organization, a notice of purpose is given before acquiring the information, necessary encryption or anti-allergy measures are adopted for the acquired information, it complies with the provisions of relevant laws and regulations, and is not contrary to public order and morals.
[0045] For ease of explanation of the video generation method according to the present disclosure, the first large-scale model according to the embodiment of the present disclosure is exemplified as a linguistic large-scale model, and the second large-scale model according to the embodiment of the present disclosure is exemplified as a visual large-scale model. The linguistic large-scale model or the visual large-scale model according to the embodiment of the present disclosure does not limit the model structure or type of the first large-scale model or the second large-scale model.
[0046] In one example, the action description information may be associated with an action video segment, and by identifying the associated action video segment based on the action description information in the demand information, the visual large-scale model can be used to process the action video segment associated with the target scenario and the action expression information to obtain the target video.
[0047] According to an embodiment of the present disclosure, the demand information may further include at least one of object character attribute information of the target object, target product information used in the target video, and target virtual item information used in the target video.
[0048] According to an embodiment of the present disclosure, the object character attribute information of a target object represents information about the target object attributes, such as the character type, character setting, character personality, speaking speed, and tone of the target object in the target video. By using a language large-scale model to process the demand information including the object character attribute information, the target scenario can be matched with the object attributes, such as the character type and character setting, of the target object, and the expression manner of the target broadcast segment text or other broadcast segment text in the target scenario can be adapted to the object attributes of the target object, thereby improving the authenticity and naturalness of the target object expression in the target video and meeting user demands. At the same time, by processing the demand information including the object character attribute information using the language large-scale model, a target video broadcast by multiple target objects can be generated, thereby improving the scene diversity of the target video and meeting the actual demands of users when watching videos.
[0049] According to an embodiment of the present disclosure, the target product information can indicate attribute information of products such as shoes, clothes, etc. to be displayed in the target video, and the product information can include any type of product attribute such as product color, size, specific mark, number, etc., and the embodiment of the present disclosure does not limit the specific type of product information.
[0050] According to an embodiment of the present disclosure, the target virtual item information used in the target video may indicate attribute information of any virtual item element, such as a lucky bag meta selling price element, that is required to appear in the target video, such as the type, size, color, and display timing of the virtual item element. The embodiment of the present disclosure does not limit the specific type of virtual item element represented by the target virtual item information.
[0051] According to an embodiment of the present disclosure, by using a language large-scale model to process at least one of object attribute information, target product information, and target virtual item information, the language large-scale model can be effectively controlled to combine with action attributes such as the action intention and action range of the specified action represented by the action description information, thereby generating scenario content with richer expressiveness and matching the action attributes, thereby making the expressiveness of the target object in the target video richer, satisfying the requirement that the expression style matches the action intention of the specified action, and improving the diversity and expressiveness of the target video.
[0052] In some embodiments, the demand information may further include object attribute information that requires modifying object attributes of the target object, such as skin color, clothing, etc. The target digital human or target object in the target video may be an updated image according to the object attribute information in the demand information.
[0053] In one embodiment, the demand information may include object attribute information, target product information, and action description information. Here, the object attribute information may include character information, facial expression information, etc., and the target product information includes information on the components and ingredients of brand A products. The target scenario obtained by processing the demand information using a language large-scale model can be expressed based on the following paragraphs enclosed in " / / ".
[0054] / / Streamer: (Tone: Calm) (Movement: Holding two open boxes of brand A, putting down the box in his left hand, pointing to the contents in the box in his right hand and explaining how to use it) Look at this. It's easy to use (the streamers both say: It's easy!). Use it once a month for six consecutive months, with a six-month break, and you only need two boxes a year. Each time you use it, mix agent A and agent B together. It will sting a little with the roller, but it's not painful and the recovery period is short. Avoid contact with water within 12 hours of use. Just be careful to use sunscreen.
[0055] Streamer: (Tone: Calm) (Action: Pick up and display the wrapped roller on the table, while the streamer demonstrates the roller's movements in front of the face) This roller is a tool used in combination with other tools, allowing nutrients to penetrate and be absorbed better. It softens the soil, allowing seeds to root better and the skin to better absorb nutrients.
[0056] Coast Streamer: (Tone: Calm) (Action: Bring in a tool bag from off-screen and open it. At the same time, the streamer opens the roll packaging and takes out the roll.) The tool bag is now complete (the streamer says very gently). You can use it directly, which is very convenient. This product is painless, has a short repair period, and will not affect your normal life.
[0057] Coast Streamer: (Tone: Excited) (Movement: Take out one bottle of concentrate and one bottle of mesotherapy cream from the aforementioned Brand A box, display them side by side, then put them back in the box) These two are Agent A and Agent B, and when combined, they have a powerful effect. Like two superheroes working together, they can solve various skin problems and make skin whiter and smoother.
[0058] Streamer: (Tone: Excited) (Expression: Happy) Everyone, with such great products and so many benefits, what are you waiting for? Opportunities like this are rare and stock is limited, so if you miss this one, you may have to wait a long time for another discount like this.
[0059] Coast Streamers: (Tone: Excited) (Action: Holding up a KT board showing comparison photos of before and after use of Brand A, and introducing the effects of use) (Expression: Surprise) Everyone, look at these comparison photos of use, the effects are really remarkable (Streamers say at the same time: This is amazing. Before use, the skin had various problems, but after use, the skin has changed by one layer, becoming whiter, softer, and more lustrous. Wouldn't you like skin like this too? / / In this example, in the same paragraph, the content after "Streamer:" may indicate the target broadcast segment text of a streamer target object with a streamer character attribute, and the content related to "(tone: calm)" may indicate that the tone attribute of the target object expressing the target broadcast segment text in the target video is calm. "(Action: Pick up and display a wrapped roller on the table, while the coast streamer demonstrates an exemplary motion of the roller in front of the face)" may represent the respective action description information of the streamer target object and the coast streamer target object. "This roller is a tool used in combination, which can better penetrate and absorb nutrients. It can soften the soil, promote seed nutrient rooting, and help skin better absorb nutrients" may indicate the target broadcast segment text of the streamer target object. "(The coast streamer simultaneously says: It's easy!)" may indicate the target broadcast segment text of the coast streamer's response. "(Expression: Enjoyable)" indicates the expression attribute information of the streamer target object.
[0060] According to the target scenario provided by this embodiment, the visual large-scale model can process the target scenario and the action video segment to more accurately capture the object attributes of each of the target objects with different character attributes and the semantics of the specified actions that match the target broadcast segment text, and the generated target video can broadcast expressions of the multiple target objects according to the needs of their respective character attributes and display them according to the organization logic in accordance with the specified actions, so as to improve the expressiveness of the target video.
[0061] In one embodiment, the action description information may be determined based on the action video segments, for example, by post-processing the action video segments with a visual language large scale model to obtain the action description information.
[0062] In one embodiment, the action video segment is identified based on the following operations: performing position change detection on key points of the target object in the initial video to obtain a position change detection result; identifying an initial action video segment from the initial video for the position change detection result; performing action type detection on the initial action video segment to obtain an action type used in the initial action video segment; and identifying the initial action video segment that matches the predetermined action type as the action video segment.
[0063] According to an embodiment of the present disclosure, the initial video may include video frames in which the target object performs one or more designated actions, and position change detection is performed on key points of the target object's body parts, such as hands, feet, and torso, in the initial video. The position change detection results can be used to more accurately identify the start and end times of the designated actions, thereby more accurately identifying initial action video segments representing various action types from the initial video.
[0064] According to an embodiment of the present disclosure, by performing action type detection on the initial action video segment, action video segments matching a predetermined action type can be identified from the initial action video segment, and action video segments representing specified actions having the predetermined action type can be more accurately selected. The initial video is clipped according to action-related time information such as the start time and end time of the specified action to obtain action video segments.
[0065] According to an embodiment of the present disclosure, an initial action video segment can be detected based on a target detection algorithm to obtain an initial action type. For example, the initial action video segment can be detected based on a target detection model constructed by a convolutional neural network algorithm. The embodiment of the present disclosure does not limit the specific manner of identifying the initial action type.
[0066] In one embodiment, the action type of the designated action may be a high-expression action type. A designated action having a high-expression action type may be an action that accurately expresses object attribute information such as the mood or character attributes of a target object. Alternatively, a designated action having a high-expression action type may be an action that provides a directional display of the performance, style, or appearance of a target product.
[0067] In one embodiment, the specified action may be a high-expression action, and the high-expression action may be configured based on actions of multiple body parts of the target object, such as hand actions, torso actions, etc. By following steps 1.1 to 1.7 below, action video segments related to the high-expression action can be identified from the initial video.
[0068] Step 1.1: Perform keypoint detection using a live video of a real person as the initial video, and the keypoints include body torso pose keypoints and gesture keypoints in the live video of the real person.
[0069] Step 1.2: Based on the keypoint locations of each frame in the live video of a real person, the body movement trajectory, left hand movement trajectory, and right hand movement trajectory of the target object can be identified as position change detection results.
[0070] Step 1.3: Set an observation window with a window time of 2 seconds, analyze the movement width based on each type of movement trajectory within the observation window, and if the position change width in the adjacent frame of the keypoint position is greater than the set threshold, the timestamp of one of the two adjacent frames can be identified as the movement start time of the initial movement.
[0071] Step 1.4: Move the observation window at a step of 0.5 seconds and continue to detect the position change range of key points within the observation window. If the position change range is greater than the set threshold, repeat step 1.4. When the live video of the actual person ends or the position change range becomes smaller than the set threshold, the timestamp of the video frame in which the current observation window is located is considered to be the end of the specified operation.
[0072] Step 1.5: Repeat step 1.3 and step 1.4 until the video ends, obtaining the start and end times of multiple initial motion video segments.
[0073] Step 1.6: Divide the initial video into multiple initial motion video segments based on the start time and end time of the initial motion video segment.
[0074] Step 1.7: Recognize the initial action type represented by each of the multiple initial action video segments using a target detection model such as a posture recognition model or a gesture recognition model, and identify an action video segment representing a specified action from the multiple initial action segments using the predetermined action type.
[0075] In addition, the information acquired in this embodiment includes, but is not limited to, information such as live video of actual people, and all of these are acquired with the permission of the relevant object or organization, the purpose is notified before the information is acquired, necessary encryption or anti-allergy measures are adopted for the acquired information, it complies with the provisions of relevant laws and regulations, and does not violate public order and morals.
[0076] FIG. 3 illustrates a principle diagram of identifying motion video segments according to an embodiment of the present disclosure.
[0077] As shown in Figure 3, position change detection is performed on key points in the initial video 301 using an observation window, and the obtained position change detection result can include at least one of a left hand movement trajectory, a right hand movement trajectory, and a body movement trajectory. By detecting that the position change width of the left hand key point in the left hand movement trajectory is greater than a predetermined width threshold, the start time of the first left hand movement at the first left hand movement time is obtained. By continuing detection from the first movement start time until the position change width of the left hand key point becomes equal to or less than the predetermined width threshold, the end time of the first left hand movement at the first left hand movement time is obtained. The position change width of the left hand key point in the initial video 301 is repeatedly detected until the nth left hand movement time is obtained. Based on a detection method using the same or similar left hand movement trajectory, the nth right hand movement time can be obtained from the first right hand movement time, and the nth body movement time can be obtained from the first body movement time. Based on the first body movement time through the nth body movement time, body posture recognition is performed on each initial movement video segment corresponding to the first body movement time through the nth body movement time in the initial video 301, respectively, to obtain a body movement type. The body movement type may be, for example, "stepping back." Gesture recognition is performed on each initial movement video segment corresponding to the first left hand movement time through the nth left hand movement time in the initial video 301, respectively, and gesture recognition is performed on each initial movement video segment corresponding to the first right hand movement time through the nth right hand movement time in the initial video 301, respectively, to obtain a left hand movement type or a right hand movement type for the initial movement video. The right hand movement type may be "gesture 1." The movement types may include a body movement type, a left hand movement type, and a right hand movement type. By combining the movement type of each initial movement video segment and the movement time corresponding to the initial movement video segment, a division subtask parameter 302 for dividing the initial video 301 can be obtained. The action type in the split subtask parameters 302 is used to select target subtask parameters that match the specified action, and the action video segments are obtained by splitting the initial video 301 using the action start time start_time and action end time end_time in the target subtask parameters.
[0078] According to an embodiment of the present disclosure, the action description information can be generated by understanding the action video segment, for example, by using a third large scale model to process the action video segment and the segment-related text used in the action video segment to obtain the action description information.
[0079] In some embodiments, the third large scale model includes a multimodal large scale model. Note that, for ease of explanation of the video generation method according to the present disclosure, the third large scale model according to the embodiments of the present disclosure may be described using a multimodal large scale model as an example, and the multimodal large scale model according to the embodiments of the present disclosure does not limit the specific model structure or type of the third large scale model.
[0080] According to an embodiment of the present disclosure, segment-related text may be understood as text related to an action video segment, and may include subtitle text of the action video segment, action video broadcast text, advertising text displayed in the action video segment, product name text, background board text, etc. The segment-related text may be obtained by recognizing video frames of the action video segment using text recognition technology, or by recognizing audio segment data of the action video segment using voice recognition technology. The embodiment of the present disclosure does not limit the specific manner of recognizing and obtaining segment-related text.
[0081] According to an embodiment of the present disclosure, the multimodal large-scale model may be a large-scale model capable of processing multimodal data such as images, text, etc., and the multimodal large-scale model can understand the action attribute information of the specified action expressed by the action video segment, such as the action intention, action type, action width, and action object of the specified action, by processing the action video segment and the segment-related text, thereby allowing the action description information to more accurately represent the action attribute information, and the linguistic large-scale model can more accurately generate target broadcast segment text matching the action attribute information based on the more accurate action description information, and further, the target object in the target video performs the specified action matching the broadcast segment text, thereby improving the expressiveness and richness of the video.
[0082] In one embodiment, the action description information is identified based on the operations of using a multimodal large-scale model to process action video segments and segment-related text used in the action video segments to obtain action attribute information, and using a linguistic large-scale model to process the action attribute information to obtain action description information that matches the semantics of the action intention data.
[0083] According to an embodiment of the present disclosure, the action intention data in the action attribute information is used to indicate a lecture intention corresponding to a specified action. For example, the action intention data may represent an intention of the target object to lecture on the twist resistance of a target product, or may represent an intention to demonstrate the effect of applying a target product such as an emollient cream. The embodiment of the present disclosure does not limit the type of lecture intention represented by the action intention data.
[0084] According to an embodiment of the present disclosure, the action description information is used to describe at least one action attribute information of the specified action, the item information related to the specified action, the virtual item information related to the specified action, the object character information of the target object that performs the specified action, and the action type information of the specified action.
[0085] According to an embodiment of the present disclosure, object attribute information including action intention data is generated using a multimodal large-scale model, and the object attribute information is further processed using a linguistic large-scale model to obtain action description information that can accurately represent action attributes such as the lecture intention for a specified action. This eliminates redundant information describing the specified action in the action video segment based on a collaboration scheme of multiple large-scale models, thereby improving the accuracy and naturalness of the expression of object attributes such as action intentions in the action description information. This allows the accurate action description information to be presented to the linguistic large-scale model so as to generate a target scenario with a natural and logically continuous expression. This allows the visual large-scale model to present a presentation word based on the target scenario, so as to adapt the broadcast representation scheme of the target object in the generated target video to the action representation scheme. This realizes the generation of a high-quality target video by controlling the collaboration of multiple large-scale models based on the accurate action description information.
[0086] In one embodiment, the behavioral description information can be identified based on FIG. 4 and the following example.
[0087] FIG. 4 illustrates a schematic diagram for specifying behavioral description information according to an embodiment of the present disclosure.
[0088] As shown in FIG. 4, automatic speech recognition (ASR) is performed on the audio segment data of the action video segment to obtain the audio subtitles used in the action video segment. Video frame sampling is performed on the action video segment to obtain sampled video frames. Text recognition is performed on the video frames based on optical character recognition (OCR) technology to obtain video text in the action video segment. A multimodal large-scale model 410 is used to process the audio subtitles, sampled video frames, and video text to obtain action attribute information. The action attribute information includes character attributes, initial action descriptions, product information, item information, and action intentions. A linguistic large-scale model 420 is used to process the action attribute information to realize redundant information removal and information correction for the action attribute information, and action description information is obtained.
[0089] According to an embodiment of the present disclosure, processing demand information using a first large-scale model to obtain a target scenario includes processing demand information using a language large-scale model to obtain a scenario outline to be used in the target video, performing knowledge search based on the scenario outline to obtain scenario material to be used in the target broadcast segment text, and processing the scenario material using the language large-scale model to obtain the target scenario.
[0090] According to an embodiment of the present disclosure, the scenario material indicates knowledge matching the demand intent expressed by the demand information. The demand information includes at least one of action description information, object character attribute information of the target object, target product information used in the target video, and target virtual item information used in the target video. The scenario outline can indicate planning framework information for the target scenario to be generated in multiple dimensions, such as scenario content theme, character attribute settings, scenario structure, and scenario style orientation. By using the relatively powerful semantic understanding and thinking capabilities of the language large-scale model to generate the scenario outline, it is possible to realize overall planning for the target scenario according to the demand intent expressed by the demand information. Furthermore, the scenario material obtained by searching based on the scenario outline can meet the need for expertise for generating the target scenario. By using the language large-scale model to process scenario material with relatively accurate knowledge material to generate the target scenario, the expertise and richness of the subsequent target video can be improved. The target broadcast text expression and specified action expression matching the target object can be used to improve the expression level of expertise in the target video.
[0091] In one embodiment, using a linguistic large-scale model to process the scenario material and obtain a target scenario includes using a linguistic large-scale model to process the scenario material and the scenario outline and obtain a target scenario, thereby constraining the linguistic large-scale model based on a creative intention framework represented by the scenario outline to generate a target scenario that matches the demand intention of the demand information, thereby improving the creative quality of the target scenario.
[0092] According to an embodiment of the present disclosure, performing a knowledge search based on the scenario outline to obtain scenario material to be used in the target broadcast segment text includes processing the scenario outline using a language large-scale model to obtain query information, and performing a knowledge search based on the query information to obtain the scenario material.
[0093] In one embodiment, performing knowledge retrieval based on the scenario outline may include performing an expert knowledge-based search based on query information such as keywords, semantic vectors, etc. in the scenario outline.
[0094] In one embodiment, performing knowledge search based on the scenario outline may include processing the demand information and the scenario outline based on a linguistic large-scale model to obtain query information, and performing expert knowledge-based search using the query information. This allows obtaining a scenario outline by systematically planning and developing a scenario production intention based on the demand information, and then actively collecting related knowledge along the key planning directions of the planned scenario outline, thereby improving the expertise, accuracy, and consistency of the scenario material with the overall scenario production intention, and further improving the logical consistency of the target scenario and the expressive effect of the target broadcast segment text in the target scenario.
[0095] In one embodiment, performing a knowledge search based on the query information to obtain scenario material includes performing a knowledge search based on the query information to obtain initial scenario material, performing semantic relevance detection between the initial scenario material and specified demand conditions to obtain a defect detection result indicating that the initial scenario material does not satisfy the specified demand conditions, processing the defect detection result and the scenario outline using a language large-scale model to obtain updated query information, and performing a knowledge search based on the updated query information to obtain the scenario material.
[0096] According to an embodiment of the present disclosure, the predetermined demand condition may be a condition for restricting the relationship between the scenario material and the creative intention of the target scenario. The predetermined demand condition may be determined based on demand information or may be determined based on other methods, such as information input through an interaction operation. The predetermined demand condition may be, for example, the number of scenario materials, the knowledge type represented by the scenario material, etc. The embodiment of the present disclosure does not limit the specific method for setting the predetermined demand condition.
[0097] In one embodiment, the initial scenario material and predetermined requirements are processed based on the linguistic large-scale model to obtain defect results indicating the defect types of the initial scenario material. The currently generated initial scenario material is then subjected to Evans and checks based on the collaboration of the linguistic large-scale model, and defect results indicating the defect types are presented. The linguistic large-scale model can then process the defect detection results and the scenario outline to control the search strategy adjustment. The updated query information can adjust the semantic accuracy of the query information or the semantic range of the query information. This allows the knowledge search process to be continuously optimized based on the updated query information. A deep circular optimization search link is realized that realizes thinking, searching, Evans, and re-searching based on the linguistic large-scale model until scenario material that meets the predetermined requirements is generated. This allows the scenario material to have clear content logic, sufficient knowledge information support, and strong feasibility, further improving the creation quality of the target scenario.
[0098] According to an embodiment of the present disclosure, processing the scenario material using a language large-scale model to obtain a target scenario may include processing the scenario material using a language large-scale model to obtain a first scenario segment when the scenario outline or demand information indicates that the target scenario is a long scenario exceeding a predetermined character count threshold or a predetermined video duration threshold. The scenario material and the first scenario segment are processed using the language large-scale model to obtain a second scenario segment. This allows the currently generated scenario segment and the scenario material to be repeatedly processed using the language large-scale model until the last scenario segment is generated, thereby obtaining an updated scenario segment. The generated multiple scenario segments are merged to obtain the target scenario.
[0099] In one embodiment, the scenario segment may include a target broadcast segment text, and processing the scenario material using a language large-scale model to obtain a target scenario includes processing the scenario material using a language large-scale model to obtain a first target broadcast segment text, processing the scenario material and the first target broadcast segment text using a language large-scale model to obtain a second target broadcast segment text, and identifying the target scenario based on the first target broadcast segment text and the second target broadcast segment text.
[0100] By processing the scenario material and the currently generated target broadcast segment scenario using a large-scale language model, scenario creation can be realized using a multi-step collaborative method based on a large-scale language model, and the target broadcast segment text output each time can be accurately matched with the action attribute information of the corresponding specified action, improving the quality of the target scenario.
[0101] In one embodiment, if the scenario outline or demand information indicates that the target scenario is a short scenario that does not exceed a predetermined character count threshold or a predetermined video duration threshold, the language large-scale model can be used to process the scenario material and directly obtain the target scenario.
[0102] In one embodiment, processing the scenario material using a language large-scale model to obtain a first target broadcast segment text may include processing the scenario material and a scenario outline using a language large-scale model to obtain the first target broadcast segment text. Processing the scenario material and a first scenario segment using a language large-scale model to obtain a second scenario segment may include processing the scenario material, the scenario outline, and the first scenario segment using a language large-scale model to obtain the second scenario segment.
[0103] In one embodiment, the generated target scenario may be verified and evaluated based on the evaluation module, and whether the currently generated target scenario satisfies predetermined scenario conditions may be evaluated based on the scenario detection result. If the currently generated target scenario does not satisfy the predetermined scenario conditions, the language large-scale model may be used to iteratively optimize the current target scenario by processing the scenario defect type, scenario material, and scenario outline indicated by the scenario detection result based on the Evans mechanism of the language large-scale model until a target scenario that meets the predetermined scenario conditions is output. This may improve the quality of the target video.
[0104] FIG. 5 illustrates a schematic diagram for identifying a target scenario according to an embodiment of the present disclosure.
[0105] 5, demand information 501 includes object attribute information of the target object, target product information, action description information, and live room item information as target virtual item information. The demand information 501 is input to a scenario material planning module M510, which outputs scenario materials. The scenario materials are processed using a scenario generation module M520, and a target scenario 502 is output.
[0106] Here, the scenario material planning module M510 includes a content planning component, a query information generation component, a knowledge search component, and an Evans optimization component. The content planning component performs scenario content planning for a target scenario by calling a language large-scale model processing demand information 501 to obtain a scenario outline. The query information generation component processes the scenario outline and demand information 501 by calling a language large-scale model to obtain query information used for knowledge search. The knowledge search component performs a knowledge search operation using the current query information obtained from the query information generation component to obtain initial scenario material. The Evans optimization component performs semantic relevance detection between the initial scenario material and specified demand conditions by calling a language large-scale model to process the initial scenario material generated by the knowledge search component in each round. If the detection result indicates the existence of a defect detection result, the defect detection result is sent to the query information generation unit, and the query information generation component calls the language large-scale model to process the scenario outline, demand information 501, and the defect detection result to generate updated initial scenario material. The Evans optimization component iteratively detects the generated initial scenario material and identifies scenario material that meets the specified demand conditions.
[0107] The scenario generation module M520 includes a scenario generation component and a scenario evaluation component. The scenario generation component can generate a short scenario having a short scenario type by calling a language large-scale model to process at least one of the scenario material and the scenario outline. Alternatively, depending on the long scenario type indicated by the demand information 501, the scenario generation component can further generate a first target broadcast segment text by calling a language large-scale model to process at least one of the scenario material and the scenario outline. By calling the language large-scale model to process at least one of the scenario material and the scenario outline and the first target broadcast segment text, second target broadcast segment texts are generated until N target broadcast segment texts are generated, thereby obtaining a long scenario having a long scenario type, where N is an integer greater than 1. The scenario evaluation component can process the long or short scenario by calling a large-scale language model to perform evaluation operations such as verification and logical error checks on the long or short scenario to obtain a target scenario 502 that satisfies the evaluation demand conditions.
[0108] According to an embodiment of the present disclosure, a target scenario includes a plurality of target broadcast segment texts arranged in order, and the arrangement positions of the plurality of target broadcast segment texts in a sub-target scenario can be represented based on the scenario structure of the target scenario. A plurality of action video segments may correspond to a plurality of target broadcast segment texts.
[0109] According to an embodiment of the present disclosure, processing the target scenario and the action video segment using the second large-scale model includes processing two related action video segments of the plurality of action video segments using the visual large-scale model to obtain a transitional action video segment, and processing the target scenario, the related action video segment, and the transitional action video segment using the visual large-scale model to obtain a target video.
[0110] In some embodiments, the transitional action video segment indicates a transitional action between two different specified actions represented by two related action video segments, and the two related action video segments are identified based on the sequence positions of the multiple target broadcast segment texts in the target scenario. The two related action video segments may be two action video segments having an action connection relationship among the multiple action video segments. For example, the two related action video segments may be two action video segments corresponding to two adjacent target broadcast segment texts in the target scenario, respectively.
[0111] In some embodiments, a visual large-scale model is utilized to process the transition action presentation word and two related action video segments among the plurality of action video segments to obtain a transition action video segment. The transition action presentation word controls the visual large-scale model to understand designated actions corresponding to each of the two related action video segments, allowing the transition action video segment to more accurately represent the transition action between the two different designated actions, and by arranging the transition action video segment between the two different related action video segments, the two different designated actions can be naturally transitioned by the transition action represented by the transition action video segment.
[0112] In some embodiments, using a visual large-scale model to process a target scenario, related action video segments, and transition action video segments can control the visual large-scale model to blend the related action video segments and transition action video segments relatively naturally based on the demand intention matching the demand information represented by the target scenario, and when multiple target action video segments in the target video can conform to the expression method of the target broadcast segment text, multiple specified actions can be displayed naturally and dynamically by the target transition action video segments in the target action video segments, thereby improving the display effect of the target video.
[0113] In some embodiments, processing the target scenario, the associated action video segments, and the transition action video segments using a visual large-scale model to obtain a target video includes processing object attribute information, the associated action video segments, and the transition action video segments in the target broadcast segment text using a visual large-scale model to obtain an intermediate video; and driving lip movements of the target object in the intermediate video based on broadcast audio data identified based on the target broadcast segment text to obtain the target video.
[0114] The broadcast audio data may be audio data representing the broadcast segment text in the target scenario. The intermediate video may be a video that can more accurately and naturally display multiple specified actions, and the lip movement of the target object in the intermediate video may be driven by the broadcast audio data, so that the lip movement of the target object in the target video can be matched with the audio information of the broadcast audio data during the broadcast expression of the target object in the target video, thereby further improving the naturalness and expressiveness of the target video.
[0115] It should be noted that the broadcast audio data and the intermediate video may be processed based on any type of algorithm to drive the lip movement of the target object in the intermediate video based on the broadcast audio data identified based on the target broadcast segment text, thereby achieving the target video. For example, the broadcast audio data and the intermediate video may be processed based on a diffusion model, but are not limited thereto. The broadcast audio data and the intermediate video may be processed based on other types of algorithms. The embodiments of the present disclosure do not limit the specific manner of driving the lip movement of the target object in the intermediate video.
[0116] In one embodiment, the broadcast audio data and the intermediate video are processed based on a large-scale visual model to obtain a target video, thereby improving the fit between the lip movements of the target object in the target video and the broadcast audio data.
[0117] FIG. 6 illustrates a schematic diagram of generating a target video according to an embodiment of the present disclosure.
[0118] As shown in FIG. 6, the target scenario may include three different scenario portions: an A product scenario portion, a B product scenario portion, and a C product scenario portion. The action description information in the A product scenario portion may specify that the action video segments for generating the target video in the action video segment library are A1 and A2. The action description information in the B product scenario portion may specify that the action video segments for generating the target video are B1 and B3. The action description information in the C product scenario portion may specify that the action video segments for generating the target video are C1, C2, and C4. Based on the positions in the target scenario of the action description information corresponding to each of the action video segments A1 and A2, B1 and B3, and C1, C2, and C4, the arrangement order of the action video segments A1 and A2, B1 and B3, and C1, C2, and C4 in the initial video segment sequence may be specified.
[0119] By using the visual large-scale model to process context scenario information and context voice audio information related to the action video segments A1 and A2 in the target scenario and processing the action video segments A1 and A2, the action video segments A1 and A2 are obtained as a transition action video segment A12 of the related action video segments. By using the visual large-scale model to process context scenario information and context voice audio information related to the action video segment A1 and the action video segment A1 in the target scenario, the action video segment A1 is obtained as a transition action video segment A11 of the related action video segments, and the transition action video segment A11 can transition the starting video content of the target video to a specified action represented by the action video segment A1. By using the linguistic large-scale model to process the related action video segments, multiple transition action video segments A22, B13, C11, C12, and C24 can be obtained. In addition, the transition action video segment A22 represents a transition action between the specified actions of each of the action video segments A2 and B1, the transition action video segment B13 represents a transition action between the specified actions of each of the action video segments B1 and B2, the transition action video segment C11 represents a transition action between the specified actions of each of the action video segments B2 and C1, the transition action video segment C12 represents a transition action between the specified actions of each of the action video segments C1 and C2, and the transition action video segment C24 represents a transition action between the specified actions of each of the action video segments C2 and C4.
[0120] In one example, the visual large scale model may be utilized to process the last video frame of the motion video segment A1 and the beginning video frame of A2 to obtain the transitional motion video segment A12.
[0121] The transitional action video segments A11, A12, A22, B13, C11, C12, and C24 are inserted into the corresponding positions of the initial video segment sequence to obtain a video segment sequence in which the transitional action video segments and the action video segments are arranged in order. The video segment sequence is fused based on the broadcast audio data to obtain a target video audio sequence. The visual large-scale model is used as an expression-driven model to process object attribute information corresponding to the action video segments or transitional action video segments in the target scenario and the action video segments and transitional action video segments in the target video audio sequence, so that the visual large-scale model controls the facial expressions of the target object to broadcast expressions in the action video segments or transitional action video segments according to the facial expression attribute information, such as the facial expression category and expression position, described in the target scenario, to obtain an intermediate video. The lip shape-driven model drives the target object in the intermediate audio according to the broadcast audio audio data to perform lip movements corresponding to the characters in the target scenario, to obtain the target video.
[0122] According to an embodiment of the present disclosure, a target video is identified by driving lip movements of a target object based on predetermined broadcast audio data, which is identified based on the following operations: processing a target scenario using a language large-scale model to obtain prosodic features; and synthesizing the target scenario based on the prosodic features to obtain broadcast audio data.
[0123] In some embodiments, the prosodic features represent the broadcast expression prosody of the text sentence in the target scenario, and the broadcast expression prosody may include attribute information related to the rhythm of the broadcast expression, such as pause type, pause duration, intonation attributes, and broadcast speech rate. By processing the target scenario using a large-scale language model, prosodic features that match the expression style of the specified action and the broadcast segment text can be output based on the relatively complete semantics of the broadcast segment text in the target scenario and the relatively accurate description style of the specified action in the action description information. Thus, by synthesizing the target scenario based on the prosodic features, the resulting broadcast audio data can control the target object to naturally and dynamically perform broadcast expressions in accordance with the demand intent of the specified action, the broadcast segment text, and other object attribute information in the demand information. Furthermore, by fusing the broadcast audio data with the intermediate video, the accuracy and flexibility of the target object expression in the target video can be improved.
[0124] In some embodiments, speech synthesizing the target scenario based on the prosodic features to obtain broadcast audio data may include speech synthesizing text sentences in the target scenario based on the prosodic features to obtain sentence-level audio data, and updating audio time attributes of word-level sub-audio data in the sentence-level audio data based on word-level time attributes of the text words in the target scenario to obtain the broadcast audio data.
[0125] In one embodiment, the sentence-level audio data may represent speech audio data in which the target object broadcasts a text sentence in the target broadcast segment text. By updating the audio time attributes of the word-level sub-audio data in the sentence-level audio data using the timestamps of the text words in the text sentence in the target broadcast segment text as word-level time attributes, alignment between the word-level sub-audio data in the sentence-level audio data and the text words in the target broadcast segment text can be achieved, so that the broadcast speech data can accurately represent each text word in the target broadcast segment text according to the broadcast expression prosody represented in the prosodic feature table, thereby improving the expression accuracy of the broadcast speech data with respect to the target broadcast segment text.
[0126] FIG. 7 is a schematic diagram illustrating a method for identifying broadcast audio data according to an embodiment of the present disclosure.
[0127] As shown in FIG. 7 , it should be understood that the broadcast audio data generated in this embodiment can be used to generate live video, and the live video can be the target video. The live room server 701 can be a server for distributing live video, and the live room server 701 sends a speech synthesis request including a target scenario to the communication proxy module 702. The speech synthesis request is used to request the generation of broadcast audio data. The communication proxy module 702 sends the target scenario to the scheduling module 703, which then calls the prosodic large-scale model 704 to process the target scenario and obtain prosodic features used in the target scenario. The prosodic large-scale model 704 can also perform semantic analysis on the target scenario to segment the target scenario and obtain text sentences in the target scenario as well as characters and prosodic features used in the text sentences. The prosodic features can represent the broadcast expression rhythm of each of the multiple target objects in the target scenario as different characters, and the alternating expression rhythm between the multiple target objects. The scheduling module 703 sends the N text sentences in the target scenario and the prosodic features of the target scenario to the audio synthesis module 705, thereby scheduling the audio synthesis module 705 to synthesize the N texts according to the prosodic features based on a text-to-speech (TTS) algorithm to obtain N sentence-level audio data corresponding to each of the N texts. The text-to-speech algorithm may include a speech synthesis algorithm model such as WaveNet. The audio synthesis module 705 may also perform word-level time attribute labeling on the text words in the text sentences to obtain word-level timestamps for the text words in the text sentences.
[0128] The scheduling module 703 sends the N text sentences and the N sentence-level audio data to the alignment model, scheduling the alignment model 706 to update the audio time attributes of the word-level sub-audio data of the sentence-level audio data using the word-level timestamps of the text words in the text sentences to obtain voice audio data with aligned word-granularity time attributes. The alignment model 706 can perform multi-track alignment on the N sentence-level audio data, for example, by separating the audio data according to each character of multiple target objects to obtain independent track data for each character's target object, and filling the mute parts of the independent track data with unvoiced mute sub-data to achieve accurate synchronization of the independent tracks of the multi-character target objects. The scheduling module 703 sends the N sentence-level audio data, the time axes of the N text sentences of the target scenario, and the word-level timestamps of the text sentences to the communication proxy module 702. The communication proxy module 702 uploads the complete voice audio data to the storage device of the live room server 701, allowing the live room server 701 to perform asynchronous callbacks based on the storage access link for the voice audio data. The live room server 701 can call the complete voice audio data to drive the lip movements of the target object to generate the live room video as the target video.
[0129] According to an embodiment of the present disclosure, the digital human video generation method based on a large-scale model may further include: in response to a target interaction command, using a linguistic large-scale model to process dynamic video demand information used in the target interaction command to obtain a dynamic video segment scenario; performing video segment generation based on the dynamic video scenario to obtain a dynamic video segment; and inserting the dynamic video segment into the target video.
[0130] In some implementations, the target interaction instructions may be obtained during playback of the target video or may be identified based on a user's interaction operations.
[0131] In some embodiments, the target interaction instruction includes at least one of a product ordering instruction, a comment creation instruction, and a like behavior instruction. The dynamic video demand information corresponding to the target interaction instruction can represent the interaction demand of the user in the process of watching the target video. By using a linguistic large-scale model to process the dynamic video demand information used in the target interaction instruction, the obtained dynamic segment scenario can satisfy the user's interaction demand intention. Furthermore, by using a visual large-scale model to process the dynamic video segment scenario, a dynamic video segment can be obtained, and the dynamic video segment can be inserted into the playing target video to satisfy the user's interaction demand while watching the target video.
[0132] In some embodiments, the dynamic video demand information may have a mapping relationship with the target interaction command, and the dynamic video demand information may include dynamic video action description information used in the dynamic video segment. By using the linguistic large-scale model to process the dynamic video demand information, the dynamic video segment text can drive the visual large-scale model to generate a dynamic video segment that can be broadcast live during the process of the target object performing the dynamic video action, thereby improving the immersion of the user during the process of watching the dynamic video segment.
[0133] In some embodiments, in response to a target interaction instruction, utilizing a language large-scale model to process dynamic video demand information used in the target interaction instruction and obtain a dynamic video segment scenario includes: in response to the target interaction instruction, performing a dynamic video task determination based on the target interaction instruction and obtaining a task determination result; and using a language large-scale model to process context scenario content identified from the target scenario based on task type and insertion position information and obtain a dynamic video segment scenario.
[0134] In one embodiment, the task determination result includes a task type associated with the dynamic video segment and insertion position information of the dynamic video segment in the target video, and the target interaction instruction is processed based on a predetermined decision model to obtain the task determination result. The decision model may be built based on a machine learning algorithm such as a decision tree, or may be built based on other types of algorithms, and the embodiments of the present disclosure are not limited thereto.
[0135] In one embodiment, the context scenario content identified from the target scenario based on the insertion position information may indicate the context scenario semantic content of the insertion position in the target scenario. Thus, by using the large-scale language model to process the task type and the context scenario content, the large-scale language model can generate dynamic video segment text that can connect the target scenario content before and after the insertion position when it fully understands the context scenario in which the dynamic video segment is to be inserted into the target video. Thus, after the generated dynamic video segment is inserted into the insertion position of the target video, it can display the dynamic video segment and the target video after the insertion position in a relatively natural, vivid, and semantically coherent manner, thereby improving the expressiveness and interactivity of the entire target video and improving video quality.
[0136] In one example, the dynamic video segment may display broadcast audio data representing the dynamic video segment text being played simultaneously while a designated object broadcasts according to lip movements corresponding to the dynamic video segment text, in order to enhance the expressiveness level of the dynamic video segment.
[0137] FIG. 8 schematically illustrates a flowchart for identifying dynamic video segments according to an embodiment of the present disclosure.
[0138] As shown in FIG. 8, a dynamic video segment can be identified based on operations S801 to S806.
[0139] In operation S801, an interaction event is detected. A target interaction instruction is identified by detecting interaction instruction parameters such as the instruction number, instruction type, etc. of the interaction instruction.
[0140] In operation S802, a task determination is performed based on the target interaction instruction to determine whether to trigger a dynamic video segment generation task, and a task type to be used for the dynamic video segment is identified. In this embodiment, the target interaction instruction may be a comment content containing the keyword "expiration date" posted by a user in a live room. The task type may be a dynamic video generation segment for displaying the expiration date.
[0141] In operation S803, based on the task determination result, identify insertion position information of a dynamic video segment from the target video being played. The insertion position indicated by the insertion position information may be a predetermined candidate insertion point before the playback of the target video, or the insertion position may be identified as a position between two different target motion video segments in the target video by detecting the playback process of the current target video.
[0142] In operation S804, a dynamic video segment scenario is generated. The comment content of the target interaction command and the context scenario content identified from the target scenario based on the insertion position information are processed using a linguistic large-scale model to obtain dynamic query information. Knowledge search is performed using the dynamic query information based on a search enhancement strategy to generate dynamic scenario material. The dynamic scenario material is processed using the linguistic large-scale model to obtain a dynamic video segment scenario as a reply content of the target interaction command.
[0143] In operation S805, a dynamic video segment is generated. The dynamic scenario material is processed using a voice synthesis model to obtain dynamic voice and audio data. By processing the dynamic scenario material using a visual large-scale model, a target object can be driven based on information such as actions and facial expressions indicated by the dynamic scenario material to generate a dynamic video segment. In addition, a target action video segment related to the dynamic video segment can be identified using insertion position information, and the visual large-scale model can be used to process the dynamic video segment and the target action video segment to obtain a transition action video segment between the dynamic video segment and the target action video segment. Then, the dynamic voice and audio data can be used to drive the lip action of the target object in the dynamic video segment and the transition action video segment to obtain a dynamic video segment and a transition video segment for introducing the "best before" date of the target product.
[0144] In operation S806, a dynamic video segment insertion task is executed, in which a dynamic video segment and a transition video segment for introducing the "best before date" of the target product are inserted into an insertion position in the target video being played, thereby realizing adding a dynamic video segment and a transition video segment to the target video or replacing some video segment content in the target video with the dynamic video segment and the transition video segment.
[0145] FIG. 9 illustrates a schematic diagram of a large-scale model-based digital human video generation method according to an embodiment of the present disclosure.
[0146] 9, the initial video is a stored live video, and an action video segment analysis component 911 is used to invoke and cooperate with a multimodal large-scale model and a linguistic large-scale model to process action video segments in the live video and video text used in the live video to obtain action description information. Here, the video text may include audio subtitle text of the live video, which is generated by a speech recognition component 912. The action video segments may be obtained by performing position change detection on key points of a target object in the live video and dividing the initial video into video segments based on the position change detection results.
[0147] The scenario generation module 920 includes a target scenario generation component 921 and a dynamic video segment scenario generation component 922. The target scenario generation component 921 is used to call a linguistic large-scale model to process action description information to obtain a target scenario. The target video generation component 931 of the video generation module 930 calls a visual large-scale model to process the target scenario, action video segments, and broadcast audio data to obtain a digital human live video as the target video, which performs a specified action based on a target object and makes a broadcast expression. Here, the broadcast audio data may be generated by processing the target scenario using the voice synthesis component 932 of the video generation module 930. The digital human live broadcast video includes a streamer object and a co-streamer object as different target objects, and the two people cooperate to make a broadcast expression through methods such as backchannel interaction. The voice and specified action of the broadcast expression can be adapted to display rich information such as matching facial expressions and tone of voice.
[0148] During digital human live video playback, the dynamic video segment scenario generation component 922 can identify insertion position information to be inserted into the digital human live video by detecting a target interaction command, i.e., a dynamic video segment generation task. The language large-scale model is invoked to process the dynamic video demand information indicated by the target interaction command and the context scenario content corresponding to the insertion position information to obtain a dynamic video segment scenario. The dynamic video segment generation component 933 of the video generation module 930 invokes the visual large-scale model to process the dynamic video segment scenario, action video segments corresponding to the dynamic action description information in the dynamic video segment scenario, and dynamic broadcast audio data representing the dynamic video segment scenario to generate a dynamic video segment. Here, the dynamic broadcast audio data may be identified by invoking a speech synthesis model based on the speech synthesis component 932 to process the dynamic video segment scenario. Inserting a dynamic video segment at an insertion position in the live digital human live video can improve the interaction experience with the live user.
[0149] Based on the large-scale model-based digital human video generation method according to the above embodiment, an embodiment of the present disclosure further provides a large-scale model-based digital human video generation device.
[0150] FIG. 10 illustrates a block diagram of a large-scale model-based digital human video generation device according to an embodiment of the present disclosure.
[0151] As shown in FIG. 10, a large-scale model-based digital human video generation device 1000 includes an acquisition module 1010, a target scenario acquisition module 1020, and a target video acquisition module 1030.
[0152] The acquiring module 1010 is used to acquire demand information, where the demand information includes action description information for describing a specified action video segment, where the action video segment represents a specified action of a target object.
[0153] The target scenario acquisition module 1020 is used to process the demand information using the first large-scale model to obtain a target scenario including target broadcast segment text that matches the action description information.
[0154] The target video acquisition module 1030 is used to process the target scenario and the action video segment using the second large-scale model, and obtain the target video for broadcasting based on the target broadcast segment text while the target digital human is performing the specified action.
[0155] According to an embodiment of the present disclosure, the action description information is identified based on an operation of using a third large scale model to process the action video segments and segment-related text used in the action video segments to obtain the action description information.
[0156] According to an embodiment of the present disclosure, the action description information is identified based on the following operations: using a third large-scale model to process the action video segments and segment-related texts used in the action video segments to obtain action attribute information, where the action intention data in the action attribute information is for expressing the lecture intention corresponding to the specified action; and using a language large-scale model to process the action attribute information to obtain action description information that matches the semantics of the action intention data.
[0157] In some embodiments, the third large scale model includes a multimodal large scale model.
[0158] According to an embodiment of the present disclosure, the action description information is used to describe at least one action attribute information of item information related to the specified action, virtual item information related to the specified action, object character information of the target object that performs the specified action, and action type information of the specified action.
[0159] According to an embodiment of the present disclosure, the target scenario acquisition module 1020 includes a first processing sub-module, a search sub-module, and a target scenario acquisition sub-module.
[0160] The first processing sub-module is used to process the demand information using a linguistic large-scale model to obtain a scenario outline to be used for the target video.
[0161] The search sub-module is used to perform a knowledge search based on the scenario outline to obtain scenario materials to be used in the target broadcast segment text, and the scenario materials indicate knowledge that matches the demand intent expressed by the demand information.
[0162] The target scenario acquisition sub-module is used to process scenario materials using a linguistic large-scale model to obtain a target scenario.
[0163] According to an embodiment of the present disclosure, the search sub-module includes a query information obtaining unit and a search unit.
[0164] The query information obtaining unit is used to process the scenario outline using a linguistic large-scale model to obtain query information.
[0165] The search unit is used to perform knowledge search based on the query information to obtain scenario materials.
[0166] According to an embodiment of the present disclosure, the search unit includes an initial scenario material obtaining subunit, a defect detection result obtaining subunit, a query information obtaining subunit, and a scenario material obtaining subunit.
[0167] The initial scenario material acquisition subunit is used to perform knowledge search based on the query information to obtain the initial scenario material.
[0168] The defect detection result acquisition subunit is used to detect the semantic association between the initial scenario material and the predetermined demand condition, and to obtain a defect detection result indicating that the initial scenario material does not satisfy the predetermined demand condition.
[0169] The query information obtaining subunit is used to process the fault detection results and the scenario outline using the linguistic large-scale model to obtain updated query information.
[0170] The scenario material acquisition subunit is used to perform knowledge search based on the updated query information to obtain scenario materials.
[0171] According to an embodiment of the present disclosure, the target scenario acquisition sub-module includes a first acquisition unit, a second acquisition unit, and a target scenario acquisition unit.
[0172] The first acquisition unit is used to process the scenario material using a linguistic large-scale model to obtain a first target broadcast segment text.
[0173] The second acquisition unit is used to process the scenario material and the first target broadcast segment text using the linguistic large-scale model to obtain a second target broadcast segment text.
[0174] The target scenario obtaining unit is used to identify a target scenario based on the first target broadcast segment text and the second target broadcast segment text.
[0175] According to an embodiment of the present disclosure, the demand information further includes at least one of object character attribute information of the target object, target product information used in the target video, and target virtual item information used in the target video.
[0176] According to an embodiment of the present disclosure, the target scenario includes multiple target broadcast segment texts arranged in sequence, where the target video acquisition module 1030 includes a transition action video segment acquisition sub-module and a target video acquisition sub-module.
[0177] The transition action video segment acquisition submodule is used to process two related action video segments among the multiple action video segments using a visual large-scale model to obtain a transition action video segment, where the transition action video segment indicates a transition action between two different specified actions represented by the two related action video segments, and the two related action video segments are identified based on the alignment positions in the target scenario of the multiple target broadcast segment texts.
[0178] The target video acquisition sub-module is used to utilize a visual large-scale model to process the target scenario, the relevant action video segments and the transition action video segments to obtain the target video.
[0179] According to an embodiment of the present disclosure, the target video acquisition sub-module includes an intermediate video acquisition unit and a target video acquisition unit.
[0180] The intermediate video acquisition unit is used to process the object attribute information, the associated action video segments and the transition action video segments in the target broadcast segment text using the visual large scale model to obtain the intermediate video.
[0181] The target video acquisition unit is used to drive lip movements of the target object in the intermediate video based on the broadcast audio data identified based on the target broadcast segment text to obtain the target video.
[0182] According to an embodiment of the present disclosure, the target video is identified by driving lip movements of the target object based on predetermined broadcast audio data that is identified based on the operations of processing the target scenario using a language large-scale model to obtain prosodic features that represent the prosody of the broadcast expression of text in the target scenario, and synthesizing the target scenario based on the prosodic features to obtain broadcast audio data.
[0183] According to an embodiment of the present disclosure, speech synthesizing the target scenario based on prosodic features to obtain broadcast speech data includes speech synthesizing text sentences in the target scenario based on prosodic features to obtain sentence-level audio data, and updating audio time attributes of word-level sub-audio data in the sentence-level audio data based on word-level time attributes of text words in the target scenario to obtain broadcast speech data.
[0184] According to an embodiment of the present disclosure, the large-scale model-based digital human video generation device 1000 further includes a dynamic video segment scenario acquisition module, a dynamic video segment generation module, and an insertion module.
[0185] The dynamic video segment scenario acquisition module is used to process the dynamic video demand information used in the target interaction command using the language large-scale model in response to the target interaction command to obtain a dynamic video segment scenario.
[0186] The dynamic video segment generation module is used to generate a video segment based on a dynamic video scenario and obtain a dynamic video segment.
[0187] The insertion module is used to insert the dynamic video segment into the target video.
[0188] According to an embodiment of the present disclosure, the dynamic video segment scenario acquisition module includes a task determination sub-module and a dynamic video segment scenario acquisition sub-module.
[0189] The task determination submodule is used to, in response to the target interaction instruction, perform dynamic video task determination based on the target interaction instruction, and obtain a task determination result including a task type associated with the dynamic video segment and insertion position information of the dynamic video segment in the target video.
[0190] The dynamic video segment scenario acquisition sub-module is used to process the identified context scenario content from the target scenario based on the task type and insertion position information using the language large-scale model to obtain a dynamic video segment scenario.
[0191] According to an embodiment of the present disclosure, the target interaction instruction includes at least one of a product ordering instruction, a comment creation instruction, and a Like action instruction.
[0192] According to an embodiment of the present disclosure, an action video segment is identified based on the following operations: performing position change detection on key points of a target object in an initial video to obtain a position change detection result; identifying an initial action video segment from the initial video based on the position change detection result; performing action type detection on the initial action video segment to obtain the action type used in the initial action video segment; and identifying an initial action video segment that matches the specified action type as an action video segment.
[0193] FIG. 11 is a schematic block diagram illustrating an artificial intelligence agent according to an embodiment of the present disclosure.
[0194] In an embodiment of the present disclosure, as shown in FIG. 11, an AI agent 1100 may include an input module 1110, a processing module 1120, and an output module 1130.
[0195] The input module 1110 is used to receive input information; The processing module 1120 is used to identify a target task based on the input information received by the input module, identify a linguistic large-scale model and a visual large-scale model based on the target task, and obtain output information by calling the linguistic large-scale model and the visual large-scale model to perform the large-scale model-based digital human video generation method according to the embodiment of the present disclosure; The output module 1130 is used to output the output information obtained by the processing module.
[0196] According to an embodiment of the present disclosure, input module 1110 is responsible for receiving or sensing information, such as queries, requests, commands, signals, or data, from the outside (e.g., a user or the external environment) and converting it into a format that can be understood and processed by AI agent 1100. Input module 1110 is the first element for AI agent 1100 to interact with the outside world, allowing AI agent 1100 to efficiently and accurately obtain and respond to necessary "sensory" information from the outside world.
[0197] In an example, the input module 1110 can input demand information, motion video segments, etc., as described above.
[0198] In the example, the processing module 1120 is the core support for processing complex task capabilities of the AI agent 1100. The processing module 1120 can implement the large-scale model-based digital human video generation method described above.
[0199] In an example, the performance of processing module 1120 may be related to the large-scale model on which AI agent 1100 is based. To fully utilize the capabilities of large models, the internal structure of processing module 1120 may be designed to be highly configurable and scalable to accommodate a variety of different types of tasks and demands in real-world scenes.
[0200] In the example, after the AI agent 1100 obtains demand information, the processing module 1120 can use a linguistic large-scale model to process the demand information to obtain a target scenario, and use a visual large-scale model to process the target scenario and action video segments to obtain a target video, and transmit the target video to the output module 1130.
[0201] As can be seen, large-scale language models have excellent language understanding and generation capabilities, but like humans, the tasks they can solve without using any tools are very limited. If the AI agent 1100 is given the ability to call tools, it can achieve tasks such as completing mathematical operations using a computer, completing data analysis using Python, and completing weather forecasts using a search engine.
[0202] In an example, the output module 1130 may output the target video described above.
[0203] The AI agent 1100 according to the embodiment of the present disclosure can simply and effectively improve its intelligence, flexibility, and versatility.
[0204] According to embodiments of the present disclosure, the present disclosure further provides an electronic device, a computer-readable storage medium, and a computer program product.
[0205] According to an embodiment of the present disclosure, an electronic device includes at least one processor and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method as described above.
[0206] According to an embodiment of the present disclosure, there is provided a non-transitory computer-readable storage medium having stored thereon computer instructions for causing a computer to perform the above-described method.
[0207] According to an embodiment of the present disclosure, a computer program product includes a computer program that, when executed by a processor, implements the method as described above.
[0208] 12 shows a schematic block diagram of an exemplary electronic device for implementing a large-scale model-based digital human video generation method according to an embodiment of the present disclosure. The electronic device is intended to represent various types of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The electronic device may also represent various types of mobile devices, such as personal digital assistants, mobile phones, smartphones, wearable devices, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are merely exemplary and do not limit the implementation of the present disclosure as described and / or claimed herein.
[0209] 12, the device 1200 includes a computing unit 1201, which may perform various appropriate operations and processes based on a computer program stored in a read-only memory (ROM) 1202 or loaded from a storage unit 1208 into a random access memory (RAM) 1203. The RAM 1203 may further store various programs and data necessary for the operation of the device 1200. The computing unit 1201, the ROM 1202, and the RAM 1203 are interconnected via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0210] The components of device 1200 are connected to an I / O interface 1205, which includes an input unit 1206 such as a keyboard, a mouse, etc., an output unit 1207 such as various types of displays, speakers, etc., a storage unit 1208 such as a magnetic disk, an optical disk, etc., and a communication unit 1209 such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1209 enables device 1200 to exchange information and data with other devices via a computer network such as the Internet and / or various electrical networks.
[0211] The computing unit 1201 may be various general-purpose and / or specialized processing modules having processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, computing units running various machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 performs each of the methods and processes described above, such as the large-scale model-based digital human video generation method. For example, in some embodiments, the large-scale model-based digital human video generation method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1208. In some embodiments, some or all of the computer program may be loaded and / or installed into the device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into the RAM 1203 and executed by the computing unit 1201, it may perform one or more steps of the large-scale model-based digital human video generation method described above. Alternatively, in other embodiments, the computing unit 1201 may be configured to perform the large-scale model-based digital human video generation method in any other suitable manner (eg, via firmware).
[0212] Various embodiments of the systems and techniques described herein above may be implemented in digital electronic circuitry systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may be embodied in one or more computer programs that can be executed and / or interpreted by a programmable system that includes at least one programmable processor, which may be a special purpose or general purpose programmable processor, and that can receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0213] Program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, so that when the program code is executed by the processor or controller, the functions and operations specified in the flowcharts and / or block diagrams are performed. The program code may be executed entirely on a device, partially on a device, partially on a device as a separate software package, and partially on a remote device, or entirely on a remote device or server.
[0214] In the context of this disclosure, a machine-readable medium may be a tangible medium, and may contain or store a program for use in or in connection with an instruction execution system, device, or electronic device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or electronic device, or any suitable combination of the above. More specific examples of machine-readable storage media include an electrical connection of one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0215] To provide for user interaction, a computer may implement the systems and techniques described herein and include a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor) for displaying information to a user, and a keyboard and pointing device (e.g., a mouse or trackball) through which a user can provide input to the computer. Other types of devices may also provide for user interaction; for example, the feedback provided to the user may be any form of sensing feedback (e.g., visual feedback, auditory feedback, or tactile feedback) and may receive input from the user in any form (including voice input, speech input, or tactile input).
[0216] The systems and techniques described herein can be implemented in a computing system including background components (e.g., a data server), or a computing system including middleware components (e.g., an application server), or a computing system including front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system including any combination of such background, middleware, or front-end components. The components of the system can be connected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include, by way of example, a local area network (LAN), a wide area network (WAN), and the Internet.
[0217] The computer system may include a client and a server. The client and server are generally remote and typically interact via a communication network. The relationship between the client and the server is created by computer programs running on the corresponding computers and having a client-server relationship. The server may be a cloud server, a server in a distributed system, or a server in a blockchain combination.
[0218] It should be understood that various types of flows shown above may be used, and operations may be rearranged, added, or deleted. For example, the operations described in this disclosure may be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this specification is not limited thereto.
[0219] The above specific embodiments do not limit the scope of protection of the present disclosure. Those skilled in the art should understand that various modifications, combinations, subcombinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present disclosure should be included within the scope of protection of the present disclosure.
Claims
1. obtaining demand information including action description information for describing a specified action video segment, the action video segment representing a specified action of a target object; processing the demand information using a first large scale model to obtain a target scenario including target broadcast segment text that matches the behavior description information; and processing the target scenario and the action video segment using a second large-scale model to obtain a target video for displaying a target digital human broadcasting based on the target broadcast segment text in the process of performing the specified action. A method for generating digital human videos based on a large-scale model.
2. The behavior description information is The action description information is identified based on an operation of processing the action video segments and segment-related texts used in the action video segments using a third large-scale model to obtain the action description information.
2. The method of claim 1 .
3. The behavior description information is Using a third large-scale model, the action video segment and the segment-related text used in the action video segment are processed to obtain action attribute information, and the action intention data in the action attribute information is for expressing a lecture intention corresponding to the specified action; using a linguistic large-scale model to process the action attribute information and obtain action description information that matches the semantics of the action intention data, wherein the third large-scale model is identified based on an operation that includes a multimodal large-scale model; 3. The method according to claim 1 or 2.
4. The behavior description information is and for describing at least one action attribute information of item information related to the specified action, virtual item information related to the specified action, object character information of a target object that performs the specified action, and action type information of the specified action.
4. The method of claim 3.
5. utilizing a first large-scale model to process the demand information to obtain a target scenario; processing the demand information using a linguistic large-scale model to obtain a scenario outline to be used for the target video, the first large-scale model including the linguistic large-scale model; performing a knowledge search based on the scenario outline to obtain scenario materials to be used in the target broadcast segment text, the scenario materials indicating knowledge matching the demand intent represented by the demand information; and processing the scenario material using the language large-scale model to obtain the target scenario.
2. The method of claim 1 .
6. performing a knowledge search based on the scenario outline to obtain scenario material to be used in the target broadcast segment text; processing the scenario outline using the linguistic large-scale model to obtain query information; performing a knowledge search based on the query information to obtain the scenario material; 6. The method of claim 5.
7. performing a knowledge search based on the query information to obtain the scenario material, performing a knowledge search based on the query information to obtain initial scenario materials; performing semantic relevance detection between the initial scenario material and a predetermined requirement, and obtaining a defect detection result indicating that the initial scenario material does not satisfy the predetermined requirement; processing the defect detection results and the scenario outline using the linguistic large-scale model to obtain updated query information; performing a knowledge search based on the updated query information to obtain the scenario material; 7. The method of claim 6.
8. Processing the scenario material using the language large-scale model to obtain the target scenario, processing the scenario material using the language large-scale model to obtain a first target broadcast segment text; processing the scenario material and the first target broadcast segment text using the linguistic large-scale model to obtain a second target broadcast segment text; identifying the target scenario based on the first target broadcast segment text and the second target broadcast segment text.
6. The method of claim 5.
9. The demand information is Further, the target video includes at least one of object character attribute information of the target object, target product information used in the target video, and target virtual item information used in the target video; 9. The method according to any one of claims 1, 5 to 8.
10. the target scenario includes a plurality of the target broadcast segment texts arranged in order; processing the target scenario and the action video segments utilizing a second large scale model; utilizing a visual large-scale model to process two related action video segments of the plurality of action video segments to obtain a transitional action video segment, the transitional action video segment showing a transitional action between two different specified actions represented by the two related action video segments, the two related action video segments being identified based on alignment positions in the target scenario of the plurality of target broadcast segment texts, and the second large-scale model including the visual large-scale model; utilizing the visual large-scale model to process the target scenario, the relevant action video segments, and the transition action video segments to obtain the target video.
2. The method of claim 1 .
11. processing the target scenario, the relevant action video segments, and the transition action video segments using the visual large-scale model to obtain the target video; utilizing the visual large-scale model to process object attribute information in the target broadcast segment text, the associated action video segments, and the transition action video segments to obtain intermediate videos; and driving lip movements of a target object in the intermediate video based on broadcast audio data identified based on the target broadcast segment text to obtain the target video.
11. The method of claim 10.
12. the target video is identified by driving lip movements of the target object based on predetermined broadcast audio data; The broadcast audio data is processing the target scenario using the first large-scale model to obtain prosodic features representing broadcast expressive prosody of text sentences in the target scenario; the broadcast speech data is identified based on an operation of performing speech synthesis on the target scenario based on the prosodic features to obtain the broadcast speech data.
2. The method of claim 1 .
13. performing speech synthesis for the target scenario based on the prosodic features to obtain the broadcast speech data, performing speech synthesis on the text sentences in the target scenario based on the prosodic features to obtain sentence-level audio data; updating audio time attributes of word-level sub-audio data in the sentence-level audio data based on word-level time attributes of text words in the target scenario to obtain the broadcast audio data; 13. The method of claim 12.
14. In response to a target interaction command, utilizing a linguistic large-scale model to process dynamic video demand information used in the target interaction command to obtain a dynamic video segment scenario; performing video segment generation based on the dynamic video scenario to obtain dynamic video segments; inserting the dynamic video segment into the target video.
2. The method of claim 1 .
15. In response to the target interaction command, processing dynamic video demand information used in the target interaction command using a language large-scale model to obtain a dynamic video segment scenario, In response to the target interaction instruction, performing dynamic video task determination based on the target interaction instruction to obtain a task determination result, the task determination result including a task type associated with the dynamic video segment and insertion position information of the dynamic video segment in the target video; utilizing the language large-scale model to process the context scenario content identified from the target scenario based on the task type and the insertion position information to obtain the dynamic video segment scenario; 15. The method of claim 14.
16. The motion video segment comprises: Performing position change detection on key points of the target object in the initial video to obtain a position change detection result; identifying an initial action video segment from the initial video for the position change detection result; performing action type detection on the initial action video segment to obtain an action type used in the initial action video segment; and identifying an initial action video segment that matches a predetermined action type as the action video segment.
3. The method according to claim 1 or 2.
17. an acquisition module for acquiring demand information including action description information for describing a specified action video segment, the action video segment representing a specified action of a target object; a target scenario acquisition module for processing the demand information using a first large-scale model to obtain a target scenario including target broadcast segment text that matches the behavior description information; a target video acquisition module for processing the target scenario and the action video segment using a second large-scale model to obtain a target video for displaying a broadcast based on the target broadcast segment text in the process of the target digital human performing the specified action; A digital human video generation device based on a large-scale model.
18. an input module for receiving input information; a processing module for identifying a target task based on the input information received by the input module, identifying a first large-scale model and a second large-scale model based on the target task, and obtaining output information by calling the first large-scale model and the second large-scale model to perform the method according to any one of claims 1, 2, 5-8, and 10-15; an output module for outputting the output information obtained by the processing module, An artificial intelligence agent characterized by
19. at least one processor; a memory communicatively coupled to the at least one processor; the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor can perform the method of any one of claims 1, 2, 5-8, and 10-15; An electronic device characterized by:
20. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: The computer instructions are used to cause the computer to carry out the method according to any one of claims 1, 2, 5-8, and 10-15. A non-transitory computer-readable storage medium having computer instructions stored thereon.
21. A computer program for implementing the method according to any one of claims 1, 2, 5 to 8 and 10 to 15 when executed by a processor.
1. A computer program product comprising:
Citation Information
Patent Citations
Method for generating digital human, model training method, device, equipment and medium
CN115082602A
Speech creation device, speech creation method, and speech creation program
JP2025064907A
Image generation system, image generation method, and image generation program
JP7578209B1