Video live streaming method and apparatus based on digital human technology

By generating scripts to control digital live broadcasts, using structured text and style features training models, the problem of insufficient intelligence of digital live broadcasts is solved, and more anthropomorphic and flexible live broadcast effects are achieved.

WO2025148415A1PCT designated stage expired Publication Date: 2025-07-17HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD

Patent Information

Application Number
PCT/CN2024/121996
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-30
Filing Date
2024-09-27
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

The lack of intelligence in digital live broadcasts leads to excessive dependence on manual control, affecting the quality of live broadcasts.

Method used

By obtaining the information input by the user, the digital human model is configured for live video broadcasting, and the model is trained using structured text, picture style features and sound style features to reduce manual control dependence.

Benefits of technology

It improves the degree of anthropomorphism and flexibility of digital human live broadcasts, reduces the dependence on manual control, and enhances the intelligence of digital humans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024121996_17072025_PF_FP_ABST
    Figure CN2024121996_17072025_PF_FP_ABST
Patent Text Reader

Abstract

A video live streaming method and apparatus based on digital human technology. The method comprises: acquiring first information inputted by a user, configuring a first digital human model to generate a script on the basis of the first information, and configuring a first digital human to perform video live streaming on the basis of the script, wherein the first information is information for display during live streaming, the script comprises information of at least one paragraph of text, and information of each paragraph of text among the information of the at least one paragraph of text is associated with a performance style feature. Therefore, during video live streaming, the first digital human has a performance style learned on the basis of the first digital human model, thus improving the degree of personification of digital humans, reducing the dependence of digital humans on manual control.
Need to check novelty before this filing date? Find Prior Art

Description

A video live broadcast method and device based on digital human technology

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on January 13, 2024, with application number 202410050578.1 and application name “A method and device for interactive communication based on digital human technology”, the entire contents of which are incorporated by reference into this application; this application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on April 30, 2024, with application number 202410543832.1 and application name “A method and device for live video broadcast based on digital human technology”, the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The present application relates to the field of digital human technology, and in particular to a video live broadcast method and device based on digital human technology. Background Art

[0004] With the continuous development of computer technology, the internet live streaming industry has become a major new industry. Internet live streaming refers to the use of various types of video recording equipment to capture the host's image and audio, and push them to each user's terminal device via a server in the form of a video stream.

[0005] Based on this, digital humans (or metahumans), designed based on real people, are gradually being applied to the internet live streaming industry due to their realistic appearance, high reusability, and wide range of application scenarios. Digital humans are virtual avatars with human or human-like appearance and behavior patterns, created through technologies such as modeling, motion capture, and artificial intelligence (AI), and presented on display devices.

[0006] In live broadcast scenarios, digital humans lack intelligence, resulting in poor quality. To make digital humans more human-like and closer to real people, they often require manual control. This results in a reliance on manual control, which is limited by the controller's own live broadcast experience.

[0007] Summary of the Invention

[0008] The embodiments of the present application provide a method and device for live video broadcasting based on digital human technology, which are used to increase the intelligence and flexibility of digital humans and reduce the reliance on manual control during live video broadcasting of digital humans.

[0009] In the first aspect, the present application provides a method for live video broadcasting based on digital human technology. The method can be applied to a cloud management platform, which is used to manage infrastructure. The infrastructure can be a device, or a module of a device (such as a chip), or a system corresponding to the device. The device can be a terminal device or a network device (such as a server). The infrastructure stores a first digital human model of the user, and the first digital human model is associated with a first digital human. Based on this, the method may include the following steps: obtaining first information input by the user, which is information for display during live broadcast; then configuring the first digital human model to generate a script based on the first information, and then configuring the first digital human to broadcast the video according to the script. Wherein, the script includes at least one paragraph of text information, and each paragraph of the at least one paragraph of text information is associated with a performance style feature.

[0010] In this method, the script's textual content can be displayed in segments based on semantics, words, and symbols (commas, periods, etc.). That is, each segment of text can be understood as a piece of text information, and therefore, the script can be understood as consisting of at least one segment of text information. Any segment of text information can be a word, a character, a sentence, a symbol, etc., without specific limitations here. The performance style characteristics associated with each segment of text information can be understood as a performance style tag associated with that segment of text information. During a live video broadcast, the first digital human can be controlled to display the corresponding text information based on the performance style tag, thereby increasing the digital human's intelligence and flexibility, making the digital human in the live video broadcast more vivid and improving the degree of anthropomorphism of the digital human in the live video broadcast. Furthermore, the digital human is completely controlled by the script, thereby reducing reliance on manual control during the live video broadcast.

[0011] In one possible implementation, before configuring the first digital human model to generate a script based on the first information, the first digital human model obtains second information input by the user, which is associated with the live video. The first digital human model then identifies the speech and keyframes in the live video based on the second information, converts the speech in the live video into first text, and converts the first text into structured text based on the structure of the first text. The first text is then converted into structured text based on the structure of the first text. The first digital human model is then learned based on the text features, the image style features, and the sound style features. The keyframes include one or more of the following: keyframes showing the host's actions, keyframes showing the host's expressions, and keyframes showing image changes due to camera changes. The text features of the structured text include one or more of the following: structural features, semantic preference features, and language style features. The image style features of the keyframes include one or more of the following: action features, expression features, and image features.

[0012] In this implementation, structured text represents text information with structural features and semantic preference features. Structural features can be understood as a structural description of multiple objects in the text, such as the structural description including the description order of multiple objects in the text, the composition structure of each object, and the description order of each component in the composition structure of each object. Semantic preference features can be understood as description preferences for each object. Language style features represent language usage preferences in structured text, or in other words, language style features represent the language preferences used by the host in the live video broadcast, such as the host's preference for using dialects, poetry, two-part allegorical sayings, a combination of Chinese and English, and famous quotes. For this reason, structured text can be understood as text content that can reflect the structure of the text.

[0013] Voice style features may include loudness, tone, etc., which can be understood as the voice characteristics of the host in the live video, such as the host's accent, tone when describing the product, and other characteristics.

[0014] Live video can be understood as a training sample. To this end, the first digital human model is used to learn the relevant features of the live video (including text features, picture style features, and sound style features). The first digital human model is also used to output the script, so that the relevant features of the output script are associated with the relevant features of the live video (or the features are similar), thereby achieving the similarity between the features of the live video controlled by the script and the features of the live video. To this end, the features of the digital human in the live video are similar to the features of the host in the live video, thereby improving the degree of anthropomorphism of the digital human video live broadcast.

[0015] In one possible implementation, the text features of the script are learned based on the structural features of the structured text; or, the text features of the script are learned based on the structural features and semantic preference features of the structured text; or, the text features of the script are learned based on the structural features, semantic preference features and language style features of the structured text.

[0016] In this implementation, the script's text features include one or more of the following: structural features and semantic preference features. It can be understood that the script's structural features are learned based on the structural features of the structured text, the script's semantic preference features are learned based on the semantic preference features of the structured text, and the script's language style features are learned based on the language style features of the structured text. In other words, the script's structural features are similar to those of the structured text, the script's semantic preference features are similar to those of the structured text, and the script's language style features are similar to those of the structured text. In other words, the script is similar to the structured text of the live video. Therefore, the digital human in the live video broadcast based on this script will resemble the host in the live video, thereby improving the degree of anthropomorphism of the digital human in the live video broadcast.

[0017] In one possible implementation, the performance style features of the script include picture style features, and the picture style features of the script are learned based on the action features of the key frames; or, the picture style features of the script are learned based on the action features and expression features of the key frames; or, the picture style features of the script are learned based on the action features, expression features and picture features of the key frames; configuring the first digital human to perform live video according to the script includes one or more of the following: the actions of the first digital human during the live video broadcast, the expressions of the first digital human during the live video broadcast, and the display screen of the first digital human during the live video broadcast.

[0018] In this implementation, it can be understood that the script's visual style features are similar to those of the live video. That is, during the live video broadcast, the first digital human's movements are similar to those of the host, the first digital human's expressions are similar to those of the host, and the displayed images are similar to those of the live video, thereby further improving the degree of anthropomorphism of the digital human's live video broadcast.

[0019] In one possible implementation, the performance style features of the script include voice style features, and the voice style features of the script are learned based on the voice style features of the speech; and configuring the first digital human to perform live video according to the script includes the sound of the first digital human during the live video broadcast.

[0020] In this implementation, it can be understood that the voice style characteristics of the script are similar to the voice style characteristics of the live video. When the digital human is controlled to perform live video according to the script, the voice characteristics of the digital human are similar to the voice characteristics of the host in the live video, thereby further improving the degree of anthropomorphism of the digital human live video.

[0021] In one possible implementation, the infrastructure also stores a second digital human model of the user, which is associated with the first digital human. To this end, the infrastructure can obtain the first bullet comment text of the first digital human during a live video broadcast, which triggers an audience interaction event. The second digital human model is then configured to generate a first response text based on the first bullet comment text, and the first digital human is configured to broadcast the video based on the first response text. The first response text is associated with the bullet comment response preference characteristics in the live video.

[0022] In this implementation, the barrage response preference feature can be understood as the host's preference for responding to barrages during a live video, such as whether the host tends to respond to barrages about product selling points or small talk. Based on this, the barrages the digital human responds to during a live video broadcast are similar to those the host responds to, thereby enhancing the digital human's anthropomorphism during the live video broadcast. Optionally, the first response text is also associated with performance style features.

[0023] In a second aspect, the present application provides a video live broadcast device based on digital human technology, the device being used to manage infrastructure, wherein the infrastructure stores a first digital human model of a user, the first digital human model being associated with a first digital human, and the device may include an acquisition module and a processing module. The acquisition module is used to acquire first information input by the user, where the first information is information to be displayed during live broadcast; the processing module is used to configure the first digital human model to generate a script based on the first information, and configure the first digital human to perform a video live broadcast based on the script, wherein the script includes at least one segment of text information, and each segment of the at least one segment of text information is associated with a performance style feature.

[0024] In a possible implementation, the acquisition module is further configured to: before configuring the first digital human model to generate a script based on the first information, acquire second information input by the user, where the second information is associated with the live video;

[0025] The processing module is also used to: identify the voice and key frames in the live video according to the second information, convert the voice in the live video into a first text, and convert the first text into a structured text according to the structure of the first text; extract the text features of the structured text, the picture style features of the key frames, and the sound style features of the voice, and learn the first digital human model based on the text features, the picture style features, and the sound style features. Among them, the key frames include one or more of the following: key frames showing the host's actions, key frames showing the host's expressions, and key frames where the picture changes due to changes in the lens; the text features of the structured text include one or more of the following: structural features, semantic preference features, and language style features; the picture style features of the key frames include one or more of the following: action features, expression features, and picture features;

[0026] In one possible implementation, the text features of the script are learned based on the structural features of the structured text; or, the text features of the script are learned based on the structural features and semantic preference features of the structured text; or, the text features of the script are learned based on the structural features, semantic preference features and language style features of the structured text.

[0027] In one possible implementation, the performance style features of the script include picture style features, and the picture style features of the script are learned based on the action features of the key frames; or, the picture style features of the script are learned based on the action features and expression features of the key frames; or, the picture style features of the script are learned based on the action features, expression features and picture features of the key frames; the processing module is specifically used for one or more of the following: controlling the actions of the first digital human during live video broadcast, controlling the expressions of the first digital human during live video broadcast, and controlling the display screen of the first digital human during live video broadcast.

[0028] In one possible implementation, the performance style features of the script include voice style features, and the voice style features of the script are learned based on the voice style features of the speech; the processing module is specifically used to control the voice of the first digital human during live video broadcast.

[0029] In one possible implementation, the infrastructure also stores a second digital human model of the user, and the second digital human model is associated with the first digital human. The acquisition module is also used to: obtain the first barrage text of the first digital human when performing a live video broadcast, and the first barrage text is triggered by an audience interaction event; the processing module is also used to: configure the second digital human model to generate a first response text based on the first barrage text, and configure the first digital human to perform a live video broadcast based on the first response text, and the first response text is associated with the barrage response preference characteristics in the live video.

[0030] In a third aspect, the present application provides a computer device comprising a processor, a memory, a communication interface, and a bus. The processor, memory, and communication interface are connected via a bus and communicate with each other. The memory is used to store computer execution instructions. When the computer device is running, the processor executes the computer execution instructions in the memory to utilize the hardware resources in the computer device to perform the operating steps of the method described in the first aspect or any possible implementation of the first aspect.

[0031] In a fourth aspect, the present application provides a computer device cluster, comprising multiple computer devices as described in the third aspect, each computer device being used to independently or jointly execute the operating steps of the method described in the first aspect or any possible implementation of the first aspect.

[0032] In a fifth aspect, the present application provides a chip system, which includes at least one chip and a memory, and the at least one chip is used to read and execute the program stored in the memory to implement the operating steps of the method described in the first aspect or any possible implementation method of the first aspect.

[0033] In the sixth aspect, a non-volatile computer-readable storage medium is provided, wherein the non-volatile computer-readable storage medium includes a program, and when the program is run on a device, the device executes the operating steps of the method described in any possible implementation method of the first party or the first aspect.

[0034] In a seventh aspect, a computer program product is provided. When the computer program product is run on a device, the device is caused to perform the operating steps of the method described in the first aspect or any possible implementation of the first aspect.

[0035] Based on the implementations provided in the above aspects, this application can also be further combined to provide more implementations. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] FIG1 is a schematic diagram of an application scenario provided by this application;

[0037] FIG2 is a flow chart of a method for live video streaming provided by the present application;

[0038] FIG3 is a flow chart of another method for live video broadcasting provided by the present application;

[0039] FIG4 is a schematic diagram of a process for learning and obtaining a first digital human model provided by the present application;

[0040] FIG5 is a schematic diagram of the system architecture of a cloud management platform provided by this application;

[0041] FIG6 is a schematic structural diagram of a video live broadcast device based on digital human technology provided by the present application;

[0042] FIG7 is a schematic diagram of the structure of a computing device provided by the present application;

[0043] FIG8 is a schematic diagram of the structure of another computing device provided by the present application;

[0044] FIG9 is a schematic structural diagram of another computing device provided in this application. DETAILED DESCRIPTION

[0045] With the continuous development of computer technology, the internet live streaming industry has become a major new sector. Internet live streaming refers to the use of various video recording devices to capture the host's image and audio, and then push the video stream via a server to each user's terminal device. Based on this, digital humans (Digital Humans / Meta Humans) designed based on real people are gradually being applied to the internet live streaming industry due to their realistic appearance, strong reusability, and wide range of application scenarios. Digital humans are virtual avatars with human or humanoid appearance and behavior patterns, created through technological means such as modeling, motion capture, and AI, and presented on display devices.

[0046] In related digital human live broadcast technologies, the digital human live broadcast lacks intelligence, resulting in poor quality. In order to make the digital human more human-like and closer to a real person in live broadcasts, manual control of the digital human is often required, such as the following:

[0047] Set parameters based on the experience of business personnel, and use large models to generate scripts and dialogues during the live broadcast.

[0048] During the live broadcast, business personnel control the digital human to respond in real time, such as audience comments, the effect of live streaming sales, and temporary field control.

[0049] During the live broadcast, business personnel control the digital human's movements and expressions according to preset requirements.

[0050] In summary, in real-world scenarios, digital human live streaming demonstrates a strong reliance on human operators. The source of script lines, decision-making regarding script arrangement, and real-time control during the performance are all manually implemented. Therefore, digital human live streaming is largely limited by the accumulated experience of human operators.

[0051] Therefore, the present application provides a video live broadcast method based on digital human technology, which is used to perform digital human live broadcast, increase the intelligence and flexibility of digital human live broadcast, make the digital human in the video live broadcast vivid, improve the degree of humanization of the digital human live broadcast, and reduce the dependence on manual control during the digital human video live broadcast.

[0052] The present application is described in detail below with reference to the accompanying drawings.

[0053] First, the application scenarios to which this application is applicable are introduced. As shown in Figure 1, a user can send first information to the server 120 through a client running on the terminal device 110. The first information is information for display during live broadcast. After obtaining the first information, the cloud management platform (or service end) running on the server 120 inputs the first information into the user's first digital human model, and then obtains the script output by the first digital human model. The script includes at least one paragraph of text information, and each paragraph of text information in at least one paragraph of text information is associated with performance style features. The first digital human is then controlled to perform live video according to the script, and the live video is displayed on the terminal device 110. In other words, the terminal device 110 can display a live broadcast room, which includes a live video broadcast according to the script. Optionally, the live video broadcast can be displayed by a client running on the terminal device 110.

[0054] In order to make the purpose, technical solutions and advantages of this application clearer, some terms involved in the embodiments of this application are first explained below.

[0055] (1) Digital Human Technology

[0056] Digital human technology uses modeling, motion capture, artificial intelligence (AI), and other technologies to create virtual avatars with human or humanoid appearance and behavior patterns, and then presents these avatars on display devices. Digital humans are typically controlled by scripts, which can be understood as specific program logic and include the movements, speech, and text displayed by the digital human.

[0057] (2) Structured text

[0058] Structured text is generated by converting the first text in the live video according to its structure. The structure of structured text can be understood as a tree structure, where child nodes represent live broadcast structure words, and each child node is a component of its corresponding parent node. In other words, structured text includes the structural content of each object in the first text as well as the textual content of each object in the first text.

[0059] For example, a live video is a live video of a product promotion. The first text corresponding to the live video includes: the first product and the second product, the origin, price, and color of the first product, and the size, performance, and appearance of the second product. Based on this, the structured content of the structured text includes: root node, child node 1, child node 2, ..., child node 8. The correspondence between the structured content and the text content is as follows:

[0060] The root node represents the live video.

[0061] Child node 1 represents the first product, child node 2 represents the second product, the live broadcast order of child node 1 is before the live broadcast order of child node 2, and the root node is the parent node of child node 1 and child node 2.

[0062] Subnode 3 represents the origin of the first product, subnode 4 represents the price of the first product, and subnode 5 represents the color of the first product, that is, subnode 1 is the parent node of subnodes 3, 4, and 5, and the live broadcast order is subnode 3, subnode 4, and subnode 5; subnode 6 represents the volume of the second product, subnode 7 represents the performance of the second product, and subnode 8 represents the appearance of the second product, that is, subnode 2 is the parent node of subnodes 6, 7, and 8, and the live broadcast order is subnode 6, 7, and 8.

[0063] Based on the above description, we can extract the structural features of the structured text and the semantic preference features of the structured text. Semantic preference features can be understood as description preferences (or description habits) for each object. For example, for the first product, the description preference represents the three selling points of the first product: origin, price, and color.

[0064] (3) Language style characteristics

[0065] The language style feature represents the language usage preference of the structured text. The language style feature may include poetic style, allusion style, musical style, etc. Taking poetic style and allusion style as an example, for the description of the origin of the first product, the structured text prefers to use the poetry corresponding to the origin to describe the origin, and prefers to use the allusions corresponding to the origin to describe the origin.

[0066] (4) Picture style characteristics

[0067] Picture style features represent the visual preferences used for the live video's visual style. Picture style features can include action features, expression features, or visual features. Action features represent the body movements made by the host in the live video, such as tilting the head or raising a hand. Expression features represent the facial expressions made by the host in the live video, such as smiling, laughing, and raising an eyebrow. Visual features represent the live display screen that the host needs to show in the live video.

[0068] For example, the live broadcast display screen may include images after the lens moves or switches. For example, when describing the color of a product, the live broadcast display screen is a close-up of the product. When describing the size of the product, the live broadcast display screen is a close-up comparison of the product and a reference object. The live broadcast display screen may also include images of changes in lens focal length. For example, when describing the appearance of a product, the live broadcast display screen is a picture of the change in lens focal length for the product. It should be understood that the live broadcast display screen may also include images of changes in distance, etc., and the live broadcast display screen is not specifically limited here.

[0069] Optionally, the visual style features are extracted from keyframes in the live video. These keyframes include one or more of the following: keyframes showing the host's actions, keyframes showing the host's expressions, and keyframes showing changes in the image due to camera lens changes. Examples of scenes where camera lens changes cause image changes include: camera movement / switching causing image switching, camera focal length changes causing changes in depth of field, etc. It should be understood that keyframes can be multiple consecutive image frames, and therefore, keyframes can also be understood as video segments that need to be learned.

[0070] (5) Voice style characteristics

[0071] Voice style features represent the voice preferences used by the host in live videos. Voice style features can include tone features, accent features, volume features, etc. For example, the host's accent and the tone of voice when describing a product.

[0072] (6) Danmu response preference characteristics

[0073] The barrage response preference feature indicates the host's preference for responding to barrage comments during a live video. For example, the barrage response preference feature indicates that the host prefers responding to comments about product selling points or chat comments during a live video.

[0074] In one possible implementation, the bullet comment response preference feature is extracted based on the second bullet comment text and the second response text in the live video, and the second response text is obtained from the first text based on the second bullet comment text. The bullet comment response preference feature, the second response text, and the second bullet comment text are used to train the second digital human model.

[0075] (7) The first digital human model

[0076] The first digital human model is learned based on live video as a training sample. The first digital human model is used to output a script so that the features of the output script are similar to those of the training sample.

[0077] Optionally, the training process of the first digital human model may include the following:

[0078] S1: Obtain second information and determine the live video based on the second information. The second information is associated with the live video, for example, the second information is a link, an identifier, or other information of the live video, or the second information is a source file of the live video. Therefore, the live video can be determined based on the second information.

[0079] S2: Identify the voice and key frames in the live video. Key frames include one or more of the following: key frames showing the host's actions, key frames showing the host's expressions, and key frames showing changes in the image due to camera changes.

[0080] The live video can be a live video of a real person (hereinafter referred to as the anchor), and the voice can be understood as the voice of the anchor in the live video. There are multiple key frames, so the key frames can be understood as video clips that need to be learned. For example, the key frames showing the anchor's actions may include image frames of the anchor's body movements in the live video, such as the anchor's head tilting image frame, hand raising image frame, etc.; the key frames showing the anchor's expression may include image frames of the anchor's facial expressions in the live video, such as the anchor's laughing image frame, eyebrow raising image frame, etc.; the key frames of the picture change due to lens change may include the picture after the lens moves or switches, for example, when describing the color of the product, the live display picture is a close-up of the product, and when describing the volume of the product, the live display picture is a close-up of the comparison between the product and the reference object.

[0081] Optionally, key frames that change due to lens changes may also include images that change in lens focal length. For example, when describing the appearance of a product, the live broadcast may display image frames corresponding to the lens focal length change of the product. It should be understood that key frames that change due to lens changes may also include image frames that change in perspective, etc., and are not specifically limited here.

[0082] S3: Convert the voice in the live video into a first text, and convert the first text into a structured text according to the structure of the first text.

[0083] The first text can be understood as the host's voice and text content, and the structure of the first text can be understood as a structural description of multiple objects in the first text. The structural description can include the description order corresponding to multiple objects in the first text, the composition structure of each object, and the description order of the components of each object, etc.

[0084] In one possible implementation, the speech in the live video can be converted into the first text through an automatic speech recognition (ASR) model, where the ASR model includes but is not limited to DeepSpeech, ESPnet, etc., and this application does not limit this.

[0085] After obtaining the first text, the first text is converted into a structured text. Please refer to the above (II) and will not elaborate on it here.

[0086] S4: Extract the text features of the structured text (also known as the text features of the live video), the picture style features of the key frames (also known as the picture style features of the live video), and the sound style features of the voice in the live video (also known as the sound style features of the live video), and obtain a first digital human model based on the text features, picture style features, and sound style features of the live video.

[0087] With reference to the introduction of (ii) and (iii) above, the text features of the live video include one or more of the following: structural features, semantic preference features, and language style features. In one possible implementation, a large language model (LLM) can extract structural features, semantic preference features, language style features, picture style features, and sound style features, and the LLM can learn to obtain a first digital human model based on these features. Among them, LLM refers to a model with a large number of parameters and a complex structure. These models can process a large amount of data and generate high-quality prediction results. During the live broadcast of digital humans, these models can be used to generate scripts and dialogues (i.e., scripts) for live performances, making the content of digital human speech richer. The specific learning process is not limited in this application.

[0088] It can be understood that the first digital human model can also be obtained after learning one or more of these features (including structural features, semantic preference features, language style features, picture style features and sound style features). This application does not limit the feature combination used for learning the first digital human model.

[0089] In one possible implementation, the first digital human model is used to generate first output information based on first input information, such that the characteristics of the first output information are similar to the textual, visual, and audio features of the live video. This first output information is the script used to control the digital human during the live broadcast. This allows the digital human, controlled according to the script, to resemble the host during the live broadcast, making the digital human vivid, increasing its intelligence and flexibility, and enhancing its anthropomorphism. Furthermore, the digital human is controlled by the script, reducing its reliance on manual control.

[0090] (8) Second Digital Human Model

[0091] The second digital human model is learned based on training samples of a second response text and a second bullet comment text that are associated with each other in the live video. The second digital human model is used to generate second output information based on second input information, so that the characteristics of the second output information are similar to the bullet comment response preference characteristics of the second response text. The second input information is the first bullet comment text when the digital human is performing a live video broadcast, and the second output information is the first response text corresponding to the first bullet comment text. In other words, the bullet comment response preference characteristics of the output first response text are similar to the bullet comment response preference characteristics of the second response text, so that the bullet comment responded by the digital human is similar to the bullet comment responded by the host in the live video, thereby improving the digital human's degree of anthropomorphism.

[0092] In one possible implementation, the training process of the second digital human model includes the following:

[0093] Establish an association between the second response text with the barrage response preference characteristics and the second barrage text. For example, if the second barrage text is "Is shipping free?" and the second response text is "Yes." Then, based on the second response text and the second barrage text, a second digital human model is obtained through LLM learning. Optionally, the first response text output by the second digital human model is also associated with the visual and audio style characteristics of the live video. This ensures that the performance style of the digital human when presenting the first response text is similar to the performance style of the host responding to the barrage in the live video, thereby enhancing the digital human's degree of anthropomorphism.

[0094] Based on the above introduction, Figure 2 is a flow chart of a video live broadcast method based on digital human technology provided by this application. In this process, after configuring the first digital human model to generate a script according to the information to be displayed during the live broadcast, the first digital human is configured to perform video live broadcast according to the script.

[0095] As shown in FIG2 , the method may include the following steps:

[0096] Step 210: Acquire first information input by the user, where the first information is information displayed during live broadcast.

[0097] In one possible implementation, the first information is of the same type as the information displayed in the live video used to train the first digital human model (including fruits, electronic products, furniture, etc.). For example, the first information and the information displayed in the live video are both about fruits, or both are about electronic products.

[0098] Step 220: Configure the first digital human model to generate a script based on the first information, where the script includes at least one piece of text information, and each piece of text information in the at least one piece of text information is associated with a performance style feature.

[0099] In one possible implementation, the script includes at least one piece of text information, and the text features of the at least one piece of text information (i.e., the text features of the script) are obtained based on structured text learning of the live video. Each piece of text information in the at least one piece of text information is associated with a performance style feature, which is obtained based on the performance style learning of the live video.

[0100] Each segment of text information can be obtained by segmenting the script text based on semantics, words, symbols (commas, periods, etc.). For example, if the script text includes "This apple is very big and sweet," based on semantics, the following three segments of text information are obtained: "This apple," "very big," and "very sweet." The style features associated with each segment of text information can be understood as the performance style label when displaying the text information. For example, the label corresponding to "this apple" indicates that the display screen is apple-like, the label corresponding to "very big" instructs the digital human to perform actions that reflect the size of the apple, and the label corresponding to "very sweet" instructs the digital human to make expressions that reflect the sweetness of the apple.

[0101] In one possible implementation, the text features of the script include structural features and semantic preference features, wherein the structural features of the script are learned based on the structural features of the structured text, and the semantic preference features of the script are learned based on the semantic preference features of the structured text. Based on the description of the first digital human model mentioned above, it can be understood that the structural features of the script are similar to the structural features of the structured text, and the semantic preference features of the script are similar to the semantic preference features of the structured text. Therefore, the structural features and semantic preference features of the script are similar to the structural features and semantic preference features of the structured text, respectively. For example, the product description structure and product description semantic preference performed in the live video broadcast based on this script are similar to the product description structure and product description semantic preference performed in the live video broadcast, thereby improving the degree of anthropomorphism of the digital human live video broadcast.

[0102] In one possible implementation, the text features also include the script's language style features, which are learned based on the language style features of structured text. Based on the description of the first digital human model, it can be understood that the script's language style features are similar to those of structured text. In other words, the script's language preferences are similar to those of structured text. Therefore, the language preferences used by the digital human controlled by the script are similar to those used by the host in the live video, further enhancing the degree of anthropomorphism of the digital human in the live video.

[0103] In one possible implementation, the performance style features associated with each piece of text information include the screen style features of the script, which are learned based on the screen style features of the live video. The screen style features include one or more of the following: action features, expression features, or screen features. For a specific description, please refer to the introduction of screen style features in the above-mentioned related technology and will not be repeated here. Based on the description of the first digital human model above, it can be understood that the screen style features of the script are similar to the screen style features of the live video. In other words, the digital human can be controlled to perform live video according to the script, including one or more of the following: the digital human's actions, the digital human's expressions, or the live display screen. In this way, during the live video broadcast, the digital human's performance style is similar to that of the host in the live video, and the screen changes during the live video broadcast are close to the screen changes in the live video, further improving the degree of anthropomorphism of the digital human's live video broadcast.

[0104] In one possible implementation, the performance style features associated with each piece of text information also include the script's language style features. These script's language style features are learned based on the sound style features of the live video. For a detailed description, please refer to the introduction to sound style features in the aforementioned related art and will not be repeated here. Based on the description of the first digital human model, it can be understood that the script's language style features are similar to the sound style features of the live video. In other words, the digital human's voice can be controlled based on the script to make its voice style similar to that of the host in the live video, thereby further enhancing the degree of anthropomorphism of the digital human video live broadcast.

[0105] It can be understood that the text features of a script may include one or more of the structural features of the script, the semantic preference features of the script, and the language style features of the script; similarly, the performance style features of the script may include one or more of the visual style features of the script and the language style features of the script.

[0106] Step 230: Configure the first digital human to perform live video broadcast according to the script.

[0107] In the above technical solution, during live video broadcasts, the first digital human can be controlled to display corresponding text information based on the performance style characteristics associated with the script, thereby increasing the digital human's intelligence and flexibility, making the digital human in the live video broadcast vivid and lifelike, and improving the digital human's degree of anthropomorphism in the live video broadcast. Furthermore, the digital human is completely controlled by the script, thereby reducing reliance on manual control during the live video broadcast.

[0108] In one possible implementation, during a live video broadcast, the first digital human also obtains the first bullet-screen text during the live broadcast. Then, upon determining that the first bullet-screen text triggers an audience interaction event, the second digital human model is configured to generate a first response text based on the first bullet-screen text (i.e., the first bullet-screen text is input into the second digital human model, and the second digital human model outputs a second response text), and the first digital human is controlled to display the second response text. For example, during a live video broadcast, the first digital human obtains the first bullet-screen text "Do you support free shipping?", and the second digital human model obtains the first response text "Support", and then controls the digital human to reply "Support" in response to the bullet-screen text.

[0109] Optionally, the second response text may be associated with the language style features and / or sound style features of the script in step 220 , so that the second response text is displayed according to the language style features and / or sound style features of the script.

[0110] Optionally, before controlling the digital human to display the first response text, it can be determined whether the same first bullet comment text has been repeated multiple times. If so, the first response text is discarded. Otherwise, the digital human is controlled to display the first response text. It should be understood that for different products, any two first bullet comment texts are considered different.

[0111] In the above technical solution, during live video broadcasts, the digital human responds to barrage comments in real time through a second digital human model, thereby increasing the digital human's intelligence and flexibility, making the digital human in the live video broadcast more vivid and more human-like. Furthermore, no human intervention is required when responding to barrage comments, thus reducing the digital human's reliance on manual control during live video broadcasts.

[0112] In one possible implementation, the above-mentioned method for live video broadcasting based on digital human technology is implemented with a client running on a terminal device and a server (such as a cloud management platform) running on a server as the execution entities. Referring to Figure 3, the method may include the following steps:

[0113] Step 310: The cloud management platform obtains first information input by the user on the client, where the first information is information displayed during live broadcast.

[0114] Refer to the above step 210 and do not elaborate on it here.

[0115] Step 320: The cloud management platform configures the first digital human model to generate a script based on the first information. The script includes at least one piece of text information, and each piece of text information in the at least one piece of text information is associated with a performance style feature.

[0116] Refer to the above step 220 and do not elaborate on it here.

[0117] Step 330: The cloud management platform configures the first digital human to perform live video broadcast according to the script.

[0118] Refer to the above step 230 and do not elaborate on it here.

[0119] In one possible implementation, a client running on a terminal device and a server running on a server (such as a cloud management platform) are used as the execution entities to implement the method of learning and obtaining the first digital human model. Referring to Figure 4, the method may include the following steps:

[0120] Step 410: The cloud management platform obtains the second information input by the user on the client, where the second information is associated with the live video.

[0121] In this process, the second information is the link, identifier and other information of the live video. The second information is used to determine the source file of the live video, and the live video is used as a training sample for machine learning.

[0122] Step 420: The cloud management platform determines relevant features of the live video, where the relevant features include one or more of the following: text features, picture style features, and sound style features of the live video.

[0123] In this process, the cloud management platform converts the host's speech in the training sample (i.e., live video) into a first text and then converts the first text into structured text. The cloud management platform then obtains video frames such as the host's actions, expressions, and live images in the live video. It also obtains the voice information (such as volume and pitch) of the host describing the product in the live video.

[0124] Then, the cloud management platform extracts the structural features, semantic preference features and language style features of the structured text, extracts the action features, expression features and picture features corresponding to these video frames, and extracts the sound style features of the sound information.

[0125] Step 430: The cloud management platform learns and obtains a first digital human model based on relevant features.

[0126] In this process, the cloud management platform learns and obtains a first digital human model based on these features (including the structural features, semantic preference features, and language style features of the structured text, the motion features, expression features, and image features of the video frame, and the voice style features). Please refer to the description of (VII) above for the specific learning process, which will not be repeated here. Based on this, after receiving the first information input by the user on the client, the cloud management platform executes the method flow of Figure 2 or Figure 3 above.

[0127] To better illustrate the above technical solution, FIG5 provides a schematic diagram of the system architecture of a cloud management platform, which can implement the methods of FIG2 to FIG4. Referring to FIG5, the system includes an identification module 510, a learning module 520, and an imitation module 530. Optionally, the identification module 510 and the learning module 520 can be applied to the server, and the imitation module 530 can be applied to the client; alternatively, the identification module 510, the learning module 520, and the imitation module 530 can all be applied to the server, with the client displaying the live video.

[0128] Recognition module 510 is configured to convert the host's voice in a live broadcast video (hereinafter referred to as the first live broadcast video for ease of description) into a first text, and then convert the first text into structured text. For example, the host's voice may include an introduction to a first product and a second product, including an introduction to the origin, taste, price, and color of the first product and an introduction to the size, performance, and appearance of the second product. This information is then included in the first text.

[0129] Identification module 510 is also configured to obtain video frames, including the host's actions, expressions, and live video footage, from the first live video. For example, when describing the taste of the first product, the host's action is a "thumbs-up," their expression is a "smile," and the live video footage is a "close-up of the host's upper body." Optionally, identification module 510 is also configured to obtain the host's voice information from the first live video. For example, when describing the taste of the first product, the host's voice is loud and high-pitched.

[0130] Learning module 520 is configured to extract the structural features, semantic preference features, and language style features of the structured text, and extract the motion features, expression features, and image features corresponding to the video frames. Optionally, learning module 520 also extracts the voice style features of the sound information. Based on these features (including the structural features, semantic preference features, and language style features of the structured text, the motion features, expression features, and image features of the video frames, and the voice style features), a first digital human model is learned.

[0131] Optionally, the learning module 520 can send the first digital human model to the imitation module 530, or the imitation module 530 can actively request the first model from the learning module 520, or the imitation module 530 obtains the first digital human model according to a preset storage address, wherein the preset storage address is used by the learning module 520 to store the first digital human model.

[0132] [Corrected 14.11.2024 in accordance with Rule 91] Imitation module 530 is used to input first information into the first digital human model, obtain a script output by the first digital human model, and control the first digital human to perform a live video broadcast based on the script (for ease of description, the video broadcast by the first digital human is referred to as the second live video below). For example, the first information includes product information corresponding to the first, second, and third objects. That is, when the first digital human performs a live video broadcast, the second live video introduces the first, second, and third objects respectively. For the first object, the first object's origin, taste, price, and color are described, and for the second and third objects, the second and third objects' size, performance, and appearance are described. The first product in the second live video is of the same type as the first object, such as fruit; the second product is of the same type as the second and third objects, such as electronic devices. Furthermore, during the live broadcast, both the first and second live videos use a poetic style for description. Furthermore, when describing the taste of the first object, the first digital human makes a "thumbs-up" gesture and a "smiling" expression, and the screen displays a "close-up of the first digital human's upper body." Moreover, when the first digital person describes the taste of the first object, the volume is high and the voice is loud.

[0133] In other words, the product description structure and semantic preferences of the first digital human in the second live video are similar to those of the host in the first live video. The language style of the first digital human is similar to that of the host. Furthermore, the movements and expressions of the first digital human are similar to those of the host, and the images displayed in the second live video are similar to those displayed in the first live video. Furthermore, the vocal characteristics of the first digital human when describing the product are similar to those of the host. Therefore, the anthropomorphism of the digital human in the live video is enhanced, and its intelligence and flexibility are increased. Furthermore, the digital human is completely controlled by the script, reducing its reliance on manual control during the live video.

[0134] In one possible implementation, while the first digital person is live streaming a video, the recognition module 510 also obtains the first bullet text in the live stream and then determines whether the first bullet text triggers an audience interaction event. For example, if the first bullet text is "Is free shipping supported?", the first bullet text is determined to be an audience interaction event targeting the product's selling point, i.e., the first bullet text is determined to have triggered an audience interaction event.

[0135] [Corrected 14.11.2024 in accordance with Rule 91] The imitation module 530 then inputs the first bullet comment text into the second digital human model, obtains the second digital human model's output of the first response text, and controls the digital human to display the first response text. For example, while the first digital human is live streaming a video, the cloud management platform obtains the first bullet comment text as "Support or not, free shipping?" The second digital human model obtains the first response text as "Support," and then controls the first digital human to reply "Support" in response to the bullet comment.

[0136] The above text, in conjunction with Figures 2 to 5, describes in detail the video live broadcast method based on digital human technology provided by the present application. The following text, in conjunction with Figures 4 to 7, describes the video live broadcast device and computing equipment provided by the present application for executing the above method.

[0137] Figure 6 is a schematic diagram of the structure of a video live broadcast device based on digital human technology provided by this application. This video live broadcast device can be used to implement the above-mentioned video live broadcast method based on digital human technology, and thus can also achieve the beneficial effects of the above-mentioned method.

[0138] As shown in Figure 6, the video live broadcast device 600 includes an acquisition module 610 and a processing module 620; the acquisition module 610 is used to obtain the first information input by the user, and the first information is the information used for display during the live broadcast; the processing module 620 is used to configure the first digital human model to generate a script based on the first information, and configure the first digital human to perform video live broadcast according to the script, and the script includes at least one piece of text information, and each piece of text information in the at least one piece of text information is associated with performance style characteristics.

[0139] The acquisition module 610 and the processing module 620 can be implemented by software or hardware. For example, the implementation of the acquisition module 610 will be described below using the acquisition module 610 as an example. Similarly, the implementation of the processing module 620 can refer to the implementation of the acquisition module 610.

[0140] As an example of a software functional unit, the acquisition module 610 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the acquisition module 610 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.

[0141] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.

[0142] As an example of a hardware functional unit, the acquisition module 610 may include at least one computing device, such as a server. Alternatively, the acquisition module 610 may be implemented using a central processing unit (CPU), an application-specific integrated circuit (ASIC), or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a data processing unit (DPU), a neural network processing unit (NPU), a system on chip (SoC), an offload card, an accelerator card, or any combination thereof.

[0143] The multiple computing devices included in the acquisition module 610 can be distributed in the same region or in different regions. The multiple computing devices included in the acquisition module 610 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the acquisition module 610 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, DPUs, NPUs, SoCs, offload cards, accelerator cards, and other computing devices.

[0144] It should be noted that, in other embodiments, the acquisition module 610 and the processing module 620 can be used to respectively execute any steps in the video live broadcast method based on digital human technology.

[0145] In a possible implementation, the acquisition module 610 is also used to: obtain the second information input by the user before configuring the first digital human model to generate a script based on the first information, and the second information is associated with the live video; the processing module 620 is also used to: identify the voice and key frames in the live video based on the second information, convert the voice in the live video into a first text, and convert the first text into a structured text based on the structure of the first text; extract the text features of the structured text, the picture style features of the key frames, and the sound style features of the voice, and learn the first digital human model based on the text features, the picture style features, and the sound style features. Among them, the key frames include one or more of the following: key frames showing the host's actions, key frames showing the host's expressions, and key frames where the picture changes due to changes in the lens; the text features of the structured text include one or more of the following: structural features, semantic preference features, and language style features; the picture style features of the key frames include one or more of the following: action features, expression features, and picture features;

[0146] In one possible implementation, the text features of the script are learned based on the structural features of the structured text; or, the text features of the script are learned based on the structural features and semantic preference features of the structured text; or, the text features of the script are learned based on the structural features, semantic preference features and language style features of the structured text.

[0147] In one possible implementation, the performance style features of the script include picture style features, and the picture style features of the script are learned based on the action features of the key frames; or, the picture style features of the script are learned based on the action features and expression features of the key frames; or, the picture style features of the script are learned based on the action features, expression features and picture features of the key frames; the processing module 620 is specifically used for one or more of the following: controlling the actions of the first digital human during live video broadcast, controlling the expressions of the first digital human during live video broadcast, and controlling the display screen of the first digital human during live video broadcast.

[0148] In one possible implementation, the performance style features of the script include voice style features, and the voice style features of the script are learned based on the voice style features of the speech; the processing module 620 is specifically used to control the voice of the first digital human during live video broadcast.

[0149] In one possible implementation, the infrastructure also stores a second digital human model of the user, and the second digital human model is associated with the first digital human. The acquisition module 610 is also used to: obtain the first barrage text of the first digital human when performing a live video broadcast, and the first barrage text is triggered by an audience interaction event; the processing module 620 is also used to: configure the second digital human model to generate a first response text based on the first barrage text, and configure the first digital human to perform a live video broadcast based on the first response text, and the first response text is associated with the barrage response preference characteristics in the live video.

[0150] Figure 7 is a schematic diagram of the structure of a computing device provided by the present application. This computing device can be used to implement the above-mentioned video live broadcast method based on digital human technology, and thus can also achieve the beneficial effects of the above-mentioned method.

[0151] As shown in FIG7 , computing device 100 includes a processor 104 and a communication interface 108. Processor 104 and communication interface 108 are coupled to each other. It will be appreciated that communication interface 108 may be a transceiver or an input / output interface. Optionally, computing device 100 may also include a memory 106 for storing instructions executed by processor 104, input data required by processor 104 to execute instructions, or data generated by processor 104 after executing instructions.

[0152] Optionally, the processor 104, the communication interface 108, and the memory 106 are interconnected via a bus 102. The bus 102 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. Buses may be classified into address buses, data buses, and control buses.

[0153] It can be understood that the processor 104 in the embodiments of the present application may include any one or more computing devices such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP) or a digital signal processor (DSP), an ASIC, an FPGA, a CPLD, an NPU, a SoC, an offload card, an acceleration card, etc.

[0154] The memory 106 may include volatile memory, such as random access memory (RAM). The processor 104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD). In addition, the memory 106 may also be implemented using storage class memory (SCM), phase change memory (PCM), or other types of storage media.

[0155] It is worth noting that the same type of storage medium can be configured in the same computing device to implement the function of memory 106, or two or more types of storage media can be configured to implement the function of memory 106. This application does not limit this.

[0156] Memory 106 stores executable program code, which processor 104 executes to implement the functions of acquisition module 610 and processing module 620, thereby implementing the method for live video broadcasting based on digital human technology. In other words, memory 106 stores instructions for executing the method for live video broadcasting based on digital human technology.

[0157] Alternatively, the memory 106 stores executable code, and the processor 104 executes the executable code to respectively implement the functions of the aforementioned video live broadcast device 600, thereby implementing the video live broadcast method based on digital human technology. In other words, the memory 106 stores instructions for executing the video live broadcast method based on digital human technology.

[0158] [Corrected 14.11.2024 according to Rule 91] The communication interface 108 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to enable communication between the computing device 100 and other devices or a communication network.

[0159] In one possible implementation, the computing device 100 may also include a chip system, which includes a processor and a power supply circuit. The power supply circuit is used to power the processor, and the processor is used to execute the operation steps corresponding to the video live broadcast method based on digital human technology. For the sake of brevity, it is not described here in detail. The processor can be implemented by a GPU, or by a computing device or AI chip such as a DPU, NPU, XPU, SoC, offload card, or accelerator card.

[0160] In one possible implementation, the computing device 100 may include multiple types of processors 104, i.e., the computing device 100 is a heterogeneous device. For example, the computing device 100 may include a CPU and a GPU. At least one of the processors 104 may execute the steps corresponding to the method for live video broadcasting based on digital human technology. For the sake of brevity, this description will not be repeated here.

[0161] As a possible implementation, the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0162] As a possible implementation, as shown in FIG8 , the present application further provides a computing device cluster including at least one computing device 100. The memory 106 in one or more computing devices 100 in the computing device cluster may store the same instructions for executing the video live broadcast method based on digital human technology.

[0163] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the method for live video broadcasting based on digital human technology. In other words, the combination of one or more computing devices 100 can jointly execute the instructions for the method for live video broadcasting based on digital human technology.

[0164] It should be noted that the memory 106 in different computing devices 100 in the computing device cluster can store different instructions, each for executing part of the functions of the live video broadcasting apparatus 600. In other words, the instructions stored in the memory 106 in different computing devices 100 can implement the functions of one or more modules in the acquisition module 610 and the processing module 620.

[0165] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network (WAN) or a local area network (LAN), among others. FIG. 9 illustrates a possible implementation. As shown in FIG. 9 , two computing devices 100A and 100B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 106 in the computing device 100A stores instructions for executing the functions of the acquisition module 610. Simultaneously, the memory 106 in the computing device 100B stores instructions for executing the functions of the processing module 620.

[0166] The connection method between the computing device clusters shown in Figure 9 can be considered to be that the video live broadcast method based on digital human technology provided in this application requires a large amount of storage of questions in the questionnaire and the options associated with each question, so it is considered to hand over the functions implemented by the processing module 620 to the computing device 100B for execution.

[0167] It should be understood that the functions of the computing device 100A shown in FIG9 may also be completed by multiple computing devices 100. Similarly, the functions of the computing device 100B may also be completed by multiple computing devices 100.

[0168] This application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similar to the connection method of the computing device cluster described in Figures 7 and 8. The difference is that the memory 106 of one or more computing devices 100 in this computing device cluster can store the same instructions for executing the video live broadcast method based on digital human technology.

[0169] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the method for live video broadcasting based on digital human technology. In other words, the combination of one or more computing devices 100 can jointly execute instructions for executing the method for live video broadcasting based on digital human technology.

[0170] It should be noted that the memory 106 in different computing devices 100 in the computing device cluster can store different instructions for implementing the functions of the video live broadcast device 600.

[0171] This application also provides a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to execute a video live broadcast method based on digital human technology.

[0172] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium capable of being stored by a computing device, or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, hard disk, or magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute a method for live video streaming based on digital human technology.

[0173] It is understood that the various numbers used in the embodiments of this application are merely for ease of description and are not intended to limit the scope of the embodiments of this application. The order of the sequence numbers of the above-mentioned processes does not necessarily imply a specific order of execution; the order of execution of the processes should be determined by their functions and inherent logic.

Claims

1. A video live streaming method based on digital human technology, characterized in that, The method is applied to a cloud management platform for managing infrastructure in which a user's first digital human model is stored, and the first digital human model is associated with a first digital human. The method includes: Obtaining first information input by the user, where the first information is information for display during a live broadcast; Configuring the first digital human model to generate a script according to the first information, where the script includes at least one piece of text information, and each piece of text information in the at least one piece of text information is associated with a performance style feature; Configuring the first digital human to conduct a video live broadcast according to the script.

2. The method according to claim 1, wherein Before configuring the first digital human model to generate a script according to the first information, the method further includes: Obtaining second information input by the user, where the second information is associated with the live broadcast video; Identifying the voice and key frames in the live broadcast video according to the second information, where the key frames include one or more of the following: key frames showing the actions of the host, key frames showing the expressions of the host, and key frames where the picture changes due to a camera change; Converting the voice in the live broadcast video into a first text, and converting the first text into a structured text according to the structure of the first text; Extracting the text features of the structured text, the picture style features of the key frames, and the sound style features of the voice, where the text features of the structured text include one or more of the following: structural features, semantic preference features, language style features, and the picture style features of the key frames include one or more of the following: action features, expression features, picture features; Learning the first digital human model according to the text features, the picture style features, and the sound style features.

3. The method according to claim 1 or 2, characterized in that, The text features of the script are learned based on the structural features of the structured text; or the text features of the script are learned based on the structural features and semantic preference features of the structured text; or the text features of the script are learned based on the structural features, semantic preference features, and language style features of the structured text.

4. The method according to claim 1 or 2, characterized in that, The performance style features of the script include picture style features, and the picture style features of the script are learned based on the action features of the key frames; or the picture style features of the script are learned based on the action features and expression features of the key frames; or the picture style features of the script are learned based on the action features, expression features, and picture features of the key frames; The configuring the first digital human to conduct a video live broadcast according to the script includes one or more of the following: the actions of the first digital human during the video live broadcast, the expressions of the first digital human during the video live broadcast, and the display picture of the first digital human during the video live broadcast.

5. The method according to claim 1 or 2, characterized in that, The performance style features of the script include sound style features, and the sound style features of the script are learned based on the sound style features of the voice; The configuring the first digital human to conduct a video live broadcast according to the script includes the sound of the first digital human during the video live broadcast.

6. The method according to any one of claims 1-5, characterized in that, The infrastructure also stores a user's second digital human model, and the second digital human model is associated with a first digital human. The method further includes: Obtaining first bullet screen text during a video live broadcast of the first digital human, where the first bullet screen text triggers an audience interaction event; Configuring the second digital human model to generate a first response text based on the first bullet screen text, where the first response text is associated with bullet screen response preference features in the live video; Configuring the first digital human to conduct a video live broadcast according to the first response text.

7. A video live broadcast device based on digital human technology, characterized in that, The device is used to manage the infrastructure, and the infrastructure stores a user's first digital human model, and the first digital human model is associated with a first digital human. The device includes: An obtaining module, configured to obtain first information input by the user, where the first information is information to be displayed during a live broadcast; A processing module, configured to configure the first digital human model to generate a script according to the first information, where the script includes at least one piece of text information, and each piece of text information in the at least one piece of text information is associated with a performance style feature; Configuring the first digital human to conduct a video live broadcast according to the script.

8. The device according to claim 7, characterized in that, The obtaining module is further configured to: Before configuring the first digital human model to generate a script according to the first information, obtain second information input by the user, where the second information is associated with the live video; The processing module is further configured to: Identify the voice and key frames in the live video according to the second information, where the key frames include one or more of the following: key frames showing the actions of the anchor, key frames showing the expressions of the anchor, key frames where the picture changes due to a camera change; Convert the voice in the live video into a first text, and convert the first text into a structured text according to the structure of the first text; Extract the text features of the structured text, the picture style features of the key frames, and the sound style features of the voice, where the text features of the structured text include one or more of the following: structural features, semantic preference features, language style features, and the picture style features of the key frames include one or more of the following: action features, expression features, picture features; Learning the first digital human model according to the text features, the picture style features, and the sound style features.

9. The device according to claim 7 or 8, characterized in that The text features of the script are learned based on the structural features of the structured text; or, the text features of the script are learned based on the structural features and semantic preference features of the structured text; or, the text features of the script are learned based on the structural features, semantic preference features, and language style features of the structured text.

10. The device according to claim 7 or 8, characterized in that The performance style features of the script include picture style features, and the picture style features of the script are learned based on the action features of the key frames; or, the picture style features of the script are learned based on the action features and expression features of the key frames; or, the picture style features of the script are learned based on the action features, expression features, and picture features of the key frames; The processing module is specifically configured to perform one or more of the following: Control the actions of the first digital human during video live streaming, control the expressions of the first digital human during video live streaming, and control the display screen of the first digital human during video live streaming.

11. The device according to claim 7 or 8, characterized in that, The performance style characteristics of the script include voice style characteristics, and the voice style characteristics of the script are learned based on the voice style characteristics of the speech. The processing module is specifically configured to: Control the voice of the first digital human during video live streaming.

12. The device according to any one of claims 7-11, characterized in that, The infrastructure also stores a second digital human model of the user. The second digital human model is associated with the first digital human. The acquisition module is further configured to: Acquire the first bullet screen text of the first digital human during video live streaming, and the first bullet screen text triggers an audience interaction event. The processing module is further configured to: Configure the second digital human model to generate a first response text according to the first bullet screen text, and the first response text is associated with the bullet screen response preference characteristics in the live video. Configure the first digital human to perform video live streaming according to the first response text.

13. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium includes a program, and when the program runs on the device, the device is caused to execute the method according to any one of claims 1-6.

14. A cluster of computing devices, characterized in that, The computing device cluster includes at least one computing device. The at least one computing device includes at least one chip and a memory. The at least one chip is configured to read and execute program instructions stored in the memory to implement the method according to any one of claims 1-6.

15. A program product, characterized in that, When the program product runs on the device, the device is caused to execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Multi-anchor virtual live broadcast method and device

    CN114979682A

  • Virtual character live broadcast method and device, electronic equipment and storage medium

    CN115643467A

  • Live broadcast method and device based on virtual anchor, electronic equipment and medium

    CN117156164A

  • Live video generation method and device based on intelligent digital human model

    CN117319699A

  • Virtual digital person video generation method and device, storage medium, and terminal

    WO2023124933A1

Cited By

  • Real-time interaction method and system for digital human live broadcast, electronic equipment and program product

    CN120812308A

  • Digital human explanation and display control method based on voice recognition driving

    CN121191518A