Video live broadcast method and device based on digital human technology

By using a digital human model to generate scripts based on learned characteristics from real livestreaming data, the method enhances the realism and adaptability of digital livestreaming, reducing human dependency and improving quality.

CN120318381APending Publication Date: 2025-07-15HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410543832.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-13
Filing Date
2024-04-30
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The lack of intelligence in digital live broadcasts leads to poor quality and relying on manual control, which is limited by the experience of business personnel.

Method used

By obtaining user input information, a digital human model is configured to broadcast live videos based on the script, combining voice, keyframe and barrage response feature learning to reduce manual control dependence.

Benefits of technology

Improve the degree of anthropomorphism and flexibility of digital human video live broadcasts, reduce dependence on manual control, and enhance intelligence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318381A_ABST
    Figure CN120318381A_ABST
Patent Text Reader

Abstract

The invention discloses a video live broadcast method and device based on a digital human technology, and the method comprises the steps: obtaining first information inputted by a user, configuring a first digital human model, generating a script according to the first information, and configuring the first digital human model to carry out the video live broadcast according to the script. Wherein the first information is information used for being displayed during live broadcast, the script comprises at least one piece of text information, and each piece of text information in the at least one piece of text information is associated with performance style characteristics, so that the first digital person has a performance style learned based on the first digital person model during video live broadcast, and the performance of the first digital person is improved. Therefore, the anthropomorphic degree of the digital human is improved, and the dependence of the digital human on manual control is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the priority of a Chinese patent application with the application number 202410050578.1 and the application title "An Interactive Method and Device Based on Digital Human Technology", which was filed with the State Intellectual Property Office of the People's Republic of China on January 13, 2024. The entire content of this Chinese patent application is incorporated herein by reference. Technical Field

[0003] This application relates to the field of digital human technology, and in particular, to a video live - streaming method and device based on digital human technology. Background Art

[0004] With the continuous development of computer technology, the Internet live - streaming industry has become a new type of industry. Among them, Internet live - streaming refers to: through various types of video - recording devices, collecting the images and audio of the live - streamer, and pushing them to the terminal devices of each user in the form of a video stream via a server.

[0005] Based on this, digital humans (Digital Human / Meta Human) designed based on real people are gradually applied to the Internet live - streaming industry due to their advantages such as realistic images, strong reusability, and wide application scenarios. Among them, digital humans are virtual images created through scientific and technological means such as modeling, motion capture, and artificial intelligence (AI), with human or human - like appearance features and behavior patterns, and presented through display devices.

[0006] In the business scenario of digital human live - streaming, the lack of intelligence of digital humans leads to poor quality of digital human live - streaming. In order to make digital humans more anthropomorphic and closer to real - person live - streaming, it is often necessary to manually control digital humans, resulting in a dependence of digital human live - streaming on manual control and being limited by the live - streaming experience of the control personnel themselves. Summary of the Invention

[0007] Embodiments of this application provide a video live - streaming method and device based on digital human technology, which are used to increase the intelligence and flexibility of digital humans and reduce the dependence on manual control during digital human video live - streaming.

[0008] In a first aspect, the present application provides a video live streaming method based on digital human technology. This method can be applied to a cloud management platform, which is used to manage infrastructure. The infrastructure can be a device, or a module of a device (such as a chip), or a system corresponding to the device. The device can be a terminal device or a network device (such as a server). A first digital human model of a user is stored in the infrastructure, and the first digital human model is associated with a first digital human. Based on this, the method may include the following steps: obtaining first information input by the user, where the first information is information to be displayed during live streaming; then configuring the first digital human model to generate a script according to the first information, and further configuring the first digital human to perform video live streaming according to the script. Wherein, the script includes at least one piece of text information, and each piece of text information in the at least one piece of text information is associated with a performance style feature.

[0009] In this method, the text content of the script can be segmented and displayed according to semantics, words, symbols (commas, periods, etc.). That is, each piece of text can be understood as a piece of text information. Therefore, it can be understood that the script is composed of at least one piece of text information. Any piece of text information can be a word, a character, a sentence, a symbol, etc., and no specific limitation is made here. Each piece of text information being associated with a performance style feature can be understood as this piece of text information being associated with a performance style label. During video live streaming, the first digital human can be controlled to display the corresponding text information according to the performance style label, thereby increasing the intelligence and flexibility of the digital human, making the digital human in the video live streaming vivid, and improving the anthropomorphic degree of the digital human video live streaming. Moreover, the digital human is completely controlled by the script, thereby reducing the dependence on manual control during digital human video live streaming.

[0010] In a possible implementation manner, before configuring the first digital human model to generate a script according to the first information, second information input by the user is obtained, and the second information is associated with the live video; then the voice and key frames in the live video are identified according to the second information, and the voice in the live video is converted into a first text, and the first text is converted into a structured text according to the structure of the first text; then the text features of the structured text, the picture style features of the key frames, and the sound style features of the voice are extracted, and finally the first digital human model is learned according to the text features, the picture style features, and the sound style features. Wherein, the key frames include one or more of the following: key frames showing the actions of the anchor, key frames showing the expressions of the anchor, key frames where the picture changes due to camera movement. The text features of the structured text include one or more of the following: structural features, semantic preference features, language style features. The picture style features of the key frames include one or more of the following: action features, expression features, picture features.

[0011] In this implementation manner, the structured text represents text information with structural features and semantic preference features. Among them, the structural features can be understood as the structural descriptions of multiple objects in the text. For example, the structural description includes the description order corresponding to multiple objects in the text, the composition structure of each object, and the description order of each component in the composition structure of each object, etc. The semantic preference features can be understood as the description preferences for each object. The language style features represent the language usage preferences in the structured text. Or rather, the language style features represent the language preferences of the host during the live broadcast in the live video. For example, the host is accustomed to using dialects, poems, Chinese two-part allegorical sayings, a combination of Chinese and English, famous quotes, etc. during the live broadcast. Therefore, the structured text can be understood as the literal content that can reflect the text structure.

[0012] The voice style features can include loudness, tone, etc., and can be understood as the voice characteristics of the host in the live video. For example, the host's accent, the tone when describing products, etc.

[0013] The live video can be understood as a training sample. Therefore, the first digital human model is used to learn the relevant features of the live video (including text features, picture style features, and voice style features). The first digital human model is also used to output a script, and further, the relevant features of the output script are associated with (or rather, feature-similar to) the relevant features of the live video, so as to make the features of the video live broadcast controlled according to the script similar to the features of the live video. Therefore, the features of the digital human in the video live broadcast are similar to the features of the host in the live video, and further improve the anthropomorphic degree of the digital human video live broadcast.

[0014] In a possible implementation manner, the text features of the script are learned based on the structural features of the structured text; or, the text features of the script are learned based on the structural features and semantic preference features of the structured text; or, the text features of the script are learned based on the structural features, semantic preference features, and language style features of the structured text.

[0015] In this implementation, the text features of the script include one or more of the following: structural features, semantic preference features. It can be understood that the structural features of the script are learned based on the structural features of the structured text, the semantic preference features of the script are learned based on the semantic preference features of the structured text, and the language style features of the script are learned based on the language style features of the structured text. That is to say, the structural features of the script are similar to the structural features of the structured text, the semantic preference features of the script are similar to the semantic preference features of the structured text, and the language style features of the script are similar to the language style features of the structured text. That is, the script is similar to the structured text of the live video. Therefore, according to this script, the digital human in the video live broadcast is similar to the host in the live video, thereby improving the anthropomorphic degree of the digital human video live broadcast.

[0016] In a possible implementation, the performance style features of the script include screen style features, and the screen style features of the script are learned based on the action features of the key frames; or, the screen style features of the script are learned based on the action features and expression features of the key frames; or, the screen style features of the script are learned based on the action features, expression features and screen features of the key frames; Configuring the first digital human to perform video live broadcast according to the script includes one or more of the following: the actions of the first digital human during video live broadcast, the expressions of the first digital human during video live broadcast, and the display screen of the first digital human during video live broadcast.

[0017] In this implementation, it can be understood that the screen style features of the script are similar to the screen style features of the live video. That is to say, during the video live broadcast, the actions of the first digital human are similar to those of the host, the expressions of the first digital human are similar to those of the host, and the displayed screen is similar to the display screen of the live video, thereby further improving the anthropomorphic degree of the digital human video live broadcast.

[0018] In a possible implementation, the performance style features of the script include voice style features, and the voice style features of the script are learned based on the voice style features of the speech; Configuring the first digital human to perform video live broadcast according to the script includes the voice of the first digital human during video live broadcast.

[0019] In this implementation, it can be understood that the voice style features of the script are similar to the voice style features of the live video. Therefore, when controlling the digital human to perform video live broadcast according to the script, the voice characteristics of the digital human are similar to the voice characteristics of the host in the live video, thereby further improving the anthropomorphic degree of the digital human video live broadcast.

[0020] In a possible implementation, a second digital human model of the user is also stored in the infrastructure, and the second digital human model is associated with a first digital human. Therefore, the first barrage text during the video live broadcast of the first digital human can be obtained, and the first barrage text triggers an audience interaction event; then, configure the second digital human model to generate a first response text according to the first barrage text, and configure the first digital human to perform a video live broadcast according to the first response text. Among them, the first response text is associated with the barrage response preference feature in the live video.

[0021] In this implementation, the barrage response preference feature can be understood as the response preference of the anchor to the barrage during the live broadcast in the live video. For example, the anchor is used to answering barrages about product selling points, and the anchor is used to answering barrages about chatting. Based on this, when the digital human conducts a video live broadcast, the barrages it responds to are close to the barrages responded to by the anchor in the live video, so as to improve the anthropomorphic degree of the digital human during the video live broadcast process. Optionally, the first response text is also associated with a performance style feature.

[0022] In a second aspect, the present application provides a video live broadcast device based on digital human technology. The device is used to manage the infrastructure, and a first digital human model of the user is stored in the infrastructure. The first digital human model is associated with a first digital human. The device may include an acquisition module and a processing module. Among them, the acquisition module is used to acquire the first information input by the user, and the first information is the information to be displayed during the live broadcast; the processing module is used to configure the first digital human model to generate a script according to the first information, and configure the first digital human to perform a video live broadcast according to the script. The script includes at least one piece of text information, and each piece of text information in the at least one piece of text information is associated with a performance style feature.

[0023] In a possible implementation, the acquisition module is further used to: before configuring the first digital human model to generate a script according to the first information, acquire the second information input by the user, and the second information is associated with the live video;

[0024] The processing module is further configured to: identify the speech and key frames in the live video according to the second information, convert the speech in the live video into a first text, and convert the first text into a structured text according to the structure of the first text; extract the text features of the structured text, the picture style features of the key frames, and the voice style features of the speech, and learn the first digital human model according to the text features, the picture style features, and the voice style features. Wherein, the key frames include one or more of the following: key frames showing the actions of the anchor, key frames showing the expressions of the anchor, and key frames where the picture changes due to camera changes; the text features of the structured text include one or more of the following: structural features, semantic preference features, and language style features; the picture style features of the key frames include one or more of the following: action features, expression features, and picture features;

[0025] In a possible implementation manner, the text features of the script are learned based on the structural features of the structured text; or, the text features of the script are learned based on the structural features and semantic preference features of the structured text; or, the text features of the script are learned based on the structural features, semantic preference features, and language style features of the structured text.

[0026] In a possible implementation manner, the performance style features of the script include picture style features, and the picture style features of the script are learned based on the action features of the key frames; or, the picture style features of the script are learned based on the action features and expression features of the key frames; or, the picture style features of the script are learned based on the action features, expression features, and picture features of the key frames; The processing module is specifically configured to perform one or more of the following: control the actions of the first digital human during video live broadcast, control the expressions of the first digital human during video live broadcast, and control the display picture of the first digital human during video live broadcast.

[0027] In a possible implementation manner, the performance style features of the script include voice style features, and the voice style features of the script are learned based on the voice style features of the speech; The processing module is specifically configured to: control the voice of the first digital human during video live broadcast.

[0028] In a possible implementation, a second digital human model of the user is further stored in the infrastructure. The second digital human model is associated with a first digital human. The obtaining module is further configured to: obtain first bullet screen text during a video live broadcast of the first digital human, where the first bullet screen text triggers an audience interaction event; the processing module is further configured to: configure the second digital human model to generate a first response text according to the first bullet screen text, and configure the first digital human to perform a video live broadcast according to the first response text, where the first response text is associated with bullet screen response preference features in the live video.

[0029] In a third aspect, the present application provides a computer device, including a processor, a memory, a communication interface, and a bus. The above-mentioned processor, memory, and communication interface are connected through the bus and complete communication with each other. The memory is used to store computer execution instructions. When the computer device runs, the processor executes the computer execution instructions in the memory to execute the operation steps of the method described in the first aspect or any possible implementation manner of the first aspect by using the hardware resources in the computer device.

[0030] In a fourth aspect, the present application provides a computer device cluster, including a plurality of computer devices as described in the third aspect. Each computer device is used to independently or jointly execute the operation steps of the method described in the first aspect or any possible implementation manner of the first aspect.

[0031] In a fifth aspect, the present application provides a chip system, where the chip system includes at least one chip and a memory. The at least one chip is used to read and execute a program stored in the memory to implement the operation steps of the method described in the above-mentioned first aspect or any possible implementation manner of the first aspect.

[0032] In a sixth aspect, a non-volatile computer-readable storage medium is provided. The non-volatile computer-readable storage medium includes a program. When the program runs on a device, the device is caused to execute the operation steps of the method described in the above-mentioned first aspect or any possible implementation manner of the first aspect.

[0033] In a seventh aspect, a computer program product is provided. When the computer program product runs on a device, the device is caused to execute the operation steps of the method described in the above-mentioned first aspect or any possible implementation manner of the first aspect.

[0034] Based on the implementations provided in the above aspects of the present application, further combinations can be made to provide more implementations. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a schematic diagram of an application scenario provided by the present application;

[0036] Figure 2 Schematic flowchart of a video live streaming method provided by this application;

[0037] Figure 3 Schematic flowchart of another video live streaming method provided by this application;

[0038] Figure 4 Schematic flowchart of a process for learning to obtain a first digital human model provided by this application;

[0039] Figure 5 Schematic diagram of the system architecture of a cloud management platform provided by this application;

[0040] Figure 6 Schematic diagram of the structure of a video live streaming device based on digital human technology provided by this application;

[0041] Figure 7 Schematic diagram of the structure of a computing device provided by this application;

[0042] Figure 8 Schematic diagram of the structure of another computing device provided by this application;

[0043] Figure 9 Schematic diagram of the structure of yet another computing device provided by this application. Detailed implementation manners

[0044] With the continuous development of computer technology, the Internet live streaming industry has become a major new industry. Among them, Internet live streaming refers to: through various types of video recording devices, collecting the images and audio of the live streamer, and pushing them to the terminal devices of each user in the form of a video stream via the server. Based on this, digital humans (Digital Human / MetaHuman), which are designed based on real people, are gradually being applied to the Internet live streaming industry due to their advantages such as lifelike images, strong reusability, and wide application scenarios. Among them, digital humans are virtual images created through technological means such as modeling, motion capture, and AI, with human or human-like appearance features and behavior patterns, and presented through display devices.

[0045] In the related digital human live streaming technology, the intelligence of digital human live streaming is insufficient, resulting in poor quality of digital human live streaming. In order to make digital humans more anthropomorphic and closer to real person live streaming, it is often necessary to manually control the digital humans, including the following content:

[0046] Set parameters based on the experience of business personnel, and use large models to generate scripts and lines during the live streaming process.

[0047] During the live streaming process, the business personnel control the digital humans to make real-time responses. For example, audience bullet screens, the effects of live commerce, and temporary field control.

[0048] During the live broadcast, business personnel control the actions and expressions of the digital human according to preset requirements.

[0049] In summary, in the real business scenario of digital human live broadcast, the digital human shows a strong dependence on business personnel. The generation source of the script lines, the decision-making of the script arrangement, and the real-time control during the script performance are all realized manually. Therefore, digital human live broadcast is largely limited by the experience accumulation of business personnel themselves.

[0050] Therefore, the present application provides a video live broadcast method based on digital human technology, which is used for digital human live broadcast, increases the intelligence and flexibility of digital human live broadcast, makes the digital human in the video live broadcast vivid, improves the anthropomorphic degree of digital human live broadcast, and reduces the dependence on manual control during digital human video live broadcast.

[0051] The following describes the present application in detail with reference to the accompanying drawings.

[0052] First, the application scenarios applicable to the present application are introduced. As shown Figure 1 in the figure, the user can send a first piece of information to the server 120 through the client running on the terminal device 110, and the first piece of information is the information to be displayed during the live broadcast. After obtaining the first piece of information, the cloud management platform (or the server side) running on the server 120 inputs the first piece of information into the user's first digital human model, and then obtains a script output by the first digital human model. The script includes at least one piece of text information, and each piece of text information in at least one piece of text information is associated with a performance style feature. Then, the first digital human is controlled to perform a video live broadcast according to the script, and the video live broadcast is displayed on the terminal device 110. That is to say, the terminal device 110 can display the live broadcast room, and the live broadcast room includes a video live broadcast according to the script. Optionally, the video live broadcast can be displayed by the client running on the terminal device 110.

[0053] In order to make the purpose, technical solution and advantages of the present application clearer, some terms involved in the embodiments of the present application are first described below.

[0054] (1) Digital human technology

[0055] Digital human technology is a virtual image with human or human-like appearance features and behavior patterns produced through scientific and technological means such as modeling, motion capture, and artificial intelligence (AI), and the virtual image is presented through a display device. The control of the digital human is generally realized by a script, and the script can be understood as specific program logic. The script can include the actions, languages, texts, etc. shown by the digital human.

[0056] (2) Structured text

[0057] The structured text is obtained by transforming the first text in the live video based on the structure of the first text. The structure of the structured text can be understood as a tree structure, where the child nodes in the tree structure represent live structure words, and any child node is a component of its corresponding parent node. In other words, the structured text includes the structural content of each object in the first text and the literal content of each object in the first text.

[0058] Exemplarily, the live video is a video corresponding to a live sales promotion. The first text corresponding to the live video includes: the first product and the second product, the origin, price, and color of the first product, and the volume, performance, and appearance of the second product. Based on this, the structural content of the structured text includes: the root node, child node 1, child node 2,..., child node 8. The literal content corresponding to the structural content includes the following:

[0059] The root node represents the live video.

[0060] Child node 1 represents the first product, child node 2 represents the second product, the live order of child node 1 is before the live order of child node 2, and the root node is the parent node of child node 1 and child node 2.

[0061] Child node 3 represents the origin of the first product, child node 4 represents the price of the first product, child node 5 represents the color of the first product. That is, child node 1 is the parent node of child node 3, child node 4, and child node 5, and the live order is child node 3, child node 4, child node 5; child node 6 represents the volume of the second product, child node 7 represents the performance of the second product, child node 8 represents the appearance of the second product. That is, child node 2 is the parent node of child node 6, child node 7, and child node 8, and the live order is child node 6, child node 7, child node 8.

[0062] Based on the above description, the structural features of the structured text and the semantic preference features of the structured text can be extracted from the structured text. Among them, the semantic preference features can be understood as the description preferences (or description habits) for each object. For example, for the first product, the description preference means describing three selling points of the first product, namely the origin, price, and color.

[0063] (III) Language style features

[0064] The language style features represent the language usage preferences of the structured text. The language style features can include the poetry style, allusion style, music style, etc. Taking the poetry style and allusion style as examples, for the description of the origin of the first product, the structured text prefers to use the poetry corresponding to the origin to describe the origin, and prefers to use the allusion corresponding to the origin to describe the origin.

[0065] (4) Screen style features

[0066] The screen style features represent the screen preferences used for the screen style of the live video. The screen style features may include action features, expression features, or screen features. Among them, the action features represent the body movements made by the host in the live video. For example, the body movements may include tilting the head, raising the hand, etc. The expression features represent the facial expressions made by the host in the live video. For example, the facial expressions include smiling, laughing, raising eyebrows, etc. The screen features represent the live display screen that the host needs to show in the live video.

[0067] Exemplarily, the live display screen may include the screen after the camera moves or switches. For example, when describing the color of a product, the live display screen is a close-up of the product. When describing the volume of the product, the live display screen is a comparison close-up of the product with a reference object. The live display screen may also include the screen with the change of the camera focal length. For example, when describing the appearance of a product, the live display screen is the screen with the change of the camera focal length for the product. It should be understood that the live display screen may also include the screen with the change of distance, etc. Here, the live display screen is not specifically limited.

[0068] Optionally, the screen style features are extracted from the key frames in the live video. The key frames include one or more of the following: the key frames showing the host's actions, the key frames showing the host's expressions, and the key frames with the change of the screen due to the camera change. Exemplarily, the scenarios with the change of the screen due to the camera change include: the camera moves / switches resulting in the screen switching, the change of the camera focal length resulting in the change of the screen depth of field, etc. It should be understood that the key frames may be multiple consecutive image frames. Therefore, the key frames can also be understood as the video segments that need to be learned.

[0069] (5) Voice style features

[0070] The voice style features represent the voice preferences used by the host in the live video. The voice style features may include tone features, accent features, volume features, etc. For example, the host's accent, the tone when describing the product, etc.

[0071] (6) Bullet screen response preference features

[0072] The bullet screen response preference features represent the response preferences of the host to the bullet screens during the live broadcast in the live video. For example, the bullet screen response preference features indicate that the host prefers to respond to the bullet screens about the product selling points or prefers to respond to the chatting bullet screens during the live broadcast.

[0073] In a possible implementation, the bullet chat response preference feature is extracted based on the second bullet chat text and the second response text in the live video, and the second response text is obtained from the first text according to the second bullet chat text. Among them, the bullet chat response preference feature, the second response text, and the second bullet chat text are used to train the second digital human model.

[0074] (VII) The first digital human model

[0075] The first digital human model is learned based on the live video as a training sample. The first digital human model is used to output a script, so that the characteristics of the output script are similar to those of the training sample.

[0076] Optionally, the training process of the first digital human model may include the following:

[0077] S1: Obtain the second information and determine the live video according to the second information. Among them, the second information is associated with the live video. For example, the second information is information such as the link or identifier of the live video, or the second information is the source file of the live video. Therefore, the live video can be determined according to the second information.

[0078] S2: Recognize the voice and key frames in the live video. The key frames include one or more of the following: key frames showing the actions of the host, key frames showing the expressions of the host, and key frames showing the changes in the picture due to camera changes.

[0079] The live video can be a video corresponding to a live broadcast of a real person (hereinafter referred to as the host). The voice can be understood as the voice of the host in the live video. There are multiple key frames, so the key frames can be understood as video segments that need to be learned. Exemplarily, the key frames showing the actions of the host may include image frames of the physical actions made by the host in the live video, such as the image frame of the host tilting the head, the image frame of the host raising the hand, etc.; the key frames showing the expressions of the host may include image frames of the facial expressions made by the host in the live video, such as the image frame of the host laughing, the image frame of the host raising the eyebrows, etc.; the key frames showing the changes in the picture due to camera changes may include the picture after the camera moves or switches. For example, when describing the color of a product, the live display picture is a close-up of the product, and when describing the volume of the product, the live display picture is a comparison close-up of the product with a reference object.

[0080] Optionally, the key frames showing the changes in the picture due to camera changes may further include the picture of the camera focal length change. For example, when describing the appearance of a product, the live display picture is an image frame of the camera focal length change for the product. It should be understood that the key frames showing the changes in the picture due to camera changes may further include the image frames of the distance change, etc., and no specific limitation is made here.

[0081] S3: Convert the speech in the live video into a first text, and convert the first text into a structured text according to the structure of the first text.

[0082] The first text can be understood as the speech text content of the anchor. The structure of the first text can be understood as the structural description of multiple objects in the first text. This structural description may include the description order corresponding to multiple objects in the first text, the composition structure of each object, and the description order of the composition components of each object, etc.

[0083] In a possible implementation, the speech in the live video can be converted into a first text through an Automatic Speech Recognition (ASR) model. Among them, the ASR model includes but is not limited to DeepSpeech, ESPnet, etc., and this application does not make any limitations in this regard.

[0084] After obtaining the first text, convert the first text into a structured text. Refer to the above (i), and details will not be elaborated here.

[0085] S4: Extract the text features of the structured text (which can also be called the text features of the live video), the picture style features of the key frames (which can also be called the picture style features of the live video), and the sound style features of the speech in the live video (which can also be called the sound style features of the live video), and learn a first digital human model according to the text features, picture style features, and sound style features of the live video.

[0086] Referring to the introductions in the above (i) and (ii), the text features of the live video include one or more of the following: structural features, semantic preference features, and language style features. In a possible implementation, a Large Language Model (LLM) can extract structural features, semantic preference features, language style features, picture style features, and sound style features, and the LLM can learn a first digital human model based on these features. Among them, the LLM refers to a model with a large number of parameters and a complex structure. These models can process a large amount of data and generate high-quality prediction results. During the process of digital human live broadcast, these models can be used to generate scripts and lines (i.e., scripts) for live performances, making the content spoken by the digital human more rich. The specific learning process is not limited in this application.

[0087] It can be understood that the first digital human model can also be obtained after learning based on one or more of these features (including structural features, semantic preference features, language style features, picture style features, and sound style features). This application does not make any limitations on the feature combinations used for learning the first digital human model.

[0088] In one possible implementation, the first digital human model is used to obtain first output information based on first input information, such that the features of the first output information are similar to the text features, picture style features, and sound style features of the live video. The first output information is also the script for controlling the digital human to conduct a live broadcast. Thus, when the digital human controlled by the script conducts a video live broadcast, the digital human is similar to the host in the live video, making the digital human vivid, increasing the intelligence and flexibility of the digital human, and improving the anthropomorphic degree of the digital human. Moreover, since the digital human is controlled by the script, the dependence of the digital human on manual control is reduced.

[0089] (VIII) The second digital human model

[0090] The second digital human model is learned based on the second response text and the second bullet screen text with an associated relationship in the live video as training samples. The second digital human model is used to obtain second output information based on second input information, such that the features of the second output information are similar to the bullet screen response preference features of the second response text. The second input information is the first bullet screen text when the digital human conducts a video live broadcast, and the second output information is the first response text corresponding to the first bullet screen text. That is to say, the bullet screen response preference features of the output first response text are similar to the bullet screen response preference features of the second response text, so that the bullet screen responded by the digital human is similar to the bullet screen responded by the host in the live video, improving the anthropomorphic degree of the digital human.

[0091] In one possible implementation, the training process of the second digital human model includes the following content:

[0092] An associated relationship is established between the second response text with bullet screen response preference features and the second bullet screen text. For example, the second bullet screen text is "Is it free of shipping?", and the second response text is "Yes". Then, based on the second response text and the second bullet screen text, the second digital human model is learned according to the LLM. Optionally, the first response text output by the second digital human model is also associated with the picture style features and sound style features of the live video, so that the performance style of the digital human when presenting the first response text is similar to the performance style of the host in the live video when responding to the bullet screen, improving the anthropomorphic degree of the digital human.

[0093] Based on the above introduction, Figure 2 FIG. is a schematic flowchart of a video live broadcast method based on digital human technology provided by this application. In this process, after configuring the first digital human model to generate a script according to the information to be presented during the live broadcast, configure the first digital human to conduct a video live broadcast according to this script.

[0094] As Figure 2 shown, the method may include the following steps:

[0095] Step 210: Obtain the first piece of information input by the user, where the first piece of information is the information to be displayed during the live broadcast.

[0096] In a possible implementation, the type of information displayed in the first piece of information is the same as that in the live video used to train the first digital human model (including fruits, electronic products, furniture, etc.). For example, both the first piece of information and the information displayed in the live video are about fruits, or both are about electronic products.

[0097] Step 220: Configure the first digital human model to generate a script based on the first piece of information. The script includes at least one piece of text information, and each piece of text information in the at least one piece of text information is associated with a performance style feature.

[0098] In a possible implementation, the script includes at least one piece of text information. The text features of the at least one piece of text information (i.e., the text features of the script) are learned based on the structured text of the live video. Each piece of text information in the at least one piece of text information is associated with a performance style feature, and this style feature is learned based on the performance style of the live video.

[0099] Among them, each piece of text information can be obtained by segmenting the text of the script according to semantics, words, symbols (commas, periods, etc.). For example, the text of the script includes "This apple is very big and sweet". Based on semantics, the following 3 pieces of text information are obtained: "This apple", "very big", "and sweet". The association of each piece of text information with a style feature can be understood as the performance style label when displaying this piece of text information. For example, the label corresponding to "This apple" indicates that the display screen shows the characteristics of an apple, the label corresponding to "very big" indicates that the digital human makes an action reflecting the size of the apple, and the label corresponding to "and sweet" indicates that the digital human makes an expression reflecting that the apple is very sweet.

[0100] In a possible implementation, the text features of the script include structural features and semantic preference features. Among them, the structural features of the script are learned based on the structural features of the structured text, and the semantic preference features of the script are learned based on the semantic preference features of the structured text. Based on the description of the above first digital human model, it can be understood that the structural features of the script are similar to the structural features of the structured text, and the semantic preference features of the script are similar to the semantic preference features of the structured text. Therefore, the structural features and semantic preference features of the script are respectively similar to the structural features and semantic preference features of the structured text. For example, the product description structure and product description semantic preference performed during the video live broadcast according to this script are similar to the product description structure and product description semantic preference of the live video, so as to improve the anthropomorphic degree of the digital human video live broadcast.

[0101] In a possible implementation, the text feature further includes the language style feature of the script, and the language style feature of the script is learned based on the language style feature of the structured text. Based on the description of the above first digital human model, it can be understood that the language style feature of the script is similar to the language style feature of the structured text, that is to say, the language usage preference of the script is similar to the language usage preference of the structured text. Therefore, the language preference used by the digital human controlled according to the script is similar to the language preference used by the anchor in the live video, further improving the anthropomorphic degree of the digital human video live broadcast.

[0102] In a possible implementation, the performance style feature associated with each piece of text information includes the screen style feature of the script, and the screen style feature of the script is learned based on the screen style feature of the live video. Among them, the screen style feature includes one or more of the following: action feature, expression feature or screen feature. For specific descriptions, refer to the introduction of the screen style feature in the above related technologies and will not be elaborated here. Based on the description of the above first digital human model, it can be understood that the screen style feature of the script is similar to the screen style feature of the live video, that is to say, controlling the digital human to perform a video live broadcast according to the script includes one or more of the following: the action of the digital human, the expression of the digital human or the live display screen. In this way, during the video live broadcast of the digital human, the performance style of the digital human is similar to the performance style of the anchor in the live video, and the screen changes during the video live broadcast are close to the screen changes in the live video, further improving the anthropomorphic degree of the digital human video live broadcast.

[0103] In a possible implementation, the performance style feature associated with each piece of text information further includes the language style feature of the script, and the language style feature of the script is learned based on the sound style feature of the live video. For specific descriptions, refer to the introduction of the sound style feature in the above related technologies and will not be elaborated here. Based on the description of the above first digital human model, it can be understood that the language style feature of the script is similar to the sound style feature of the live video, that is to say, the sound of the digital human can be controlled according to the script so that the sound style of the digital human is similar to the sound style of the anchor in the live video, thereby further improving the anthropomorphic degree of the digital human video live broadcast.

[0104] It can be understood that the text features of the script can include one or more of the structural features of the script, the semantic preference features of the script, and the language style features of the script; similarly, the performance style features of the script can include one or more of the screen style features of the script and the language style features of the script.

[0105] Step 230: Configure the first digital human to perform a video live broadcast according to the script.

[0106] In the above technical solution, during video live streaming, the corresponding text information can be controlled to be displayed by the first digital human according to the performance style characteristics associated with the script, thereby increasing the intelligence and flexibility of the digital human, making the digital human in the video live streaming vivid, and improving the anthropomorphic degree of the digital human video live streaming. Moreover, the digital human is completely controlled by the script, thereby reducing the dependence on manual control during digital human video live streaming.

[0107] In a possible implementation, during the process of video live streaming by the first digital human, the first bullet screen text in the live stream is also obtained. Then, when it is determined that the first bullet screen text triggers an audience interaction event, the second digital human model is configured to generate a first response text according to the first bullet screen text (that is, the first bullet screen text is input into the second digital human model, and the second response text output by the second digital human model is obtained), and the first digital human is controlled to display the second response text. For example, during the process of video live streaming by the first digital human, the first bullet screen text obtained is "Do you support free shipping?", and the first response text obtained through the second digital human model is "Yes", and then the digital human is controlled to reply "Yes" to respond to the bullet screen.

[0108] Optionally, the second response text can be associated with the language style characteristics and / or voice style characteristics of the script in step 220 above, so as to display the second response text according to the language style characteristics and / or voice style characteristics of the script.

[0109] Optionally, before controlling the digital human to display the first response text, it can be determined whether the same first bullet screen text has been repeatedly responded to multiple times. If so, the first response text is discarded; otherwise, the digital human is controlled to display the first response text. It should be understood that for different products, any two first bullet screen texts are considered different.

[0110] In the above technical solution, during the process of video live streaming by the digital human, the bullet screen is responded to in real time through the second digital human model, thereby increasing the intelligence and flexibility of the digital human, making the digital human in the video live streaming vivid, and improving the anthropomorphic degree of the digital human during the video live streaming process. And no human intervention is required when responding to the bullet screen, thereby reducing the dependence of the digital human on manual control during digital human video live streaming.

[0111] In a possible implementation, the client running on the terminal device and the server running on the server (such as a cloud management platform) are used as the execution entities to implement the above video live streaming method based on digital human technology. Refer to Figure 3 , the method may include the following steps:

[0112] Step 310: The cloud management platform obtains the first information input by the user on the client, and the first information is the information to be displayed during live streaming.

[0113] Refer to step 210 above, which will not be elaborated here.

[0114] Step 320: The cloud management platform configures the first digital human model to generate a script according to the first information. The script includes at least one piece of text information, and each piece of text information in the at least one piece of text information is associated with a performance style feature.

[0115] Refer to step 220 above, which will not be elaborated here.

[0116] Step 330: The cloud management platform configures the first digital human to conduct a video live broadcast according to the script.

[0117] Refer to step 230 above, which will not be elaborated here.

[0118] In a possible implementation, a client running on a terminal device and a server (such as a cloud management platform) running on a server are used as the execution entities to implement the method of learning the first digital human model. Refer to Figure 4 , this method may include the following steps:

[0119] Step 410: The cloud management platform obtains second information input by the user on the client, and the second information is associated with the live video.

[0120] In this process, the second information is information such as the link and identifier of the live video. This second information is used to determine the source file of the live video, and the live video is used as a training sample for machine learning.

[0121] Step 420: The cloud management platform determines the relevant features of the live video. The relevant features include one or more of the following: the text feature of the live video, the picture style feature, and the sound style feature.

[0122] In this process, the cloud management platform converts the host voice in the training sample (i.e., the live video) into the first text and converts the first text into structured text. Then, the cloud management platform obtains video frames such as the host actions, host expressions, and live pictures in the live video. Obtain the sound information (such as volume, pitch, etc.) of the host describing the product in the live video.

[0123] Then, the cloud management platform extracts the structural features, semantic preference features, and language style features of the structured text, extracts the action features, expression features, and picture features corresponding to these video frames, and extracts the sound style features of the sound information.

[0124] Step 430: The cloud management platform learns the first digital human model according to the relevant features.

[0125] In this process, the cloud management platform learns the first digital human model based on these features (including the structural features, semantic preference features, and language style features of structured text, the action features, expression features, and picture features of video frames, and the sound style features). For the specific learning process, please refer to the description in (7) above and will not be elaborated here. Based on this, after the cloud management platform receives the first information input by the user on the client side, it executes the above Figure 2 or Figure 3 method process.

[0126] To better elaborate on the above technical solution, Figure 5 a schematic diagram of the system architecture of the cloud management platform is provided. This system can implement the above Figures 2 to 4 method. Referring to Figure 5 , this system includes an identification module 510, a learning module 520, and an imitation module 530. Optionally, the identification module 510 and the learning module 520 can be applied to the server side, and the imitation module 530 is applied to the client side; or, the identification module 510, the learning module 520, and the imitation module 530 are all applied to the server side, and the client displays the video live broadcast.

[0127] The identification module 510 is used to convert the host voice in the live video of the real person live broadcast room (hereinafter simply referred to as the first live video) into the first text and convert the first text into structured text. For example, the host voice includes: the introduction of the first product and the second product, the origin, taste, price, and color of the first product are introduced respectively for the first product, and the volume, performance, and appearance of the second product are introduced respectively for the second product. That is, the first text includes this information.

[0128] The identification module 510 is also used to obtain video frames such as the host's actions, host's expressions, and live broadcast pictures in the first live video. For example, when describing the taste of the first product, the host's action is "thumbs up", the host's expression is "smiling", and the live broadcast picture is "a close-up of the host's upper body". Optionally, the identification module 510 is also used to obtain the host's voice information in the first live video. For example, when describing the taste of the first product, the host has a loud volume and a high pitch.

[0129] The learning module 520 is used to extract the structural features, semantic preference features, and language style features of the structured text, and extract the action features, expression features, and picture features corresponding to these video frames. Optionally, the learning module 520 also extracts the sound style features of the voice information. Then, based on these features (including the structural features, semantic preference features, and language style features of the structured text, the action features, expression features, and picture features of the video frames, and the sound style features), the first digital human model is learned.

[0130] Optionally, the learning module 520 may send the first digital human model to the imitation module 530, or the imitation module 530 may actively request the first model from the learning module 520, or the imitation module 530 may obtain the first digital human model according to a preset storage address, where the preset storage address is used by the learning module 520 to store the first digital human model.

[0131] The imitation module 520 is configured to input the first information into the first digital human model to obtain a script output by the first digital human model, and control the first digital human to perform a video live broadcast according to the script (for ease of description, the video live broadcast by the first digital human is hereinafter referred to as the second live video). For example, the first information includes product information corresponding to the first object, the second object, and the third object. That is, when the first digital human performs a video live broadcast, the second live video introduces the first object, the second object, and the third object respectively. For the first object, the origin, taste, price, and color of the first object are introduced respectively. For the second object and the third object, the volume, performance, and appearance of the second object and the third object are introduced respectively. Among them, the first product in the second live video is of the same type as the first object, such as both being fruits; the second product is of the same type as the second object and the third object, such as both being electronic devices. Moreover, both the first live video and the second live video use a poetic style for description during the live broadcast. Moreover, when the first digital human describes the taste of the first object, it makes a "thumbs up" gesture and a "smiling" expression, and the picture is a "close-up of the upper body of the first digital human". Moreover, when the first digital human describes the taste of the first object, the volume is high and the voice is loud.

[0132] In other words, the product description structure and the product description semantic preference performed by the first digital human in the second live video are similar to those performed by the anchor in the first live video. The language style performed by the first digital human is similar to the language style of the anchor. Moreover, the actions and expressions made by the first digital human are similar to the actions and expressions made by the anchor, and the picture displayed in the second live video is similar to the picture displayed in the first live video. Moreover, the voice style characteristics of the first digital human when describing the product are similar to the voice style characteristics of the anchor when describing the product. Therefore, the anthropomorphic degree of the digital human video live broadcast is improved, and the intelligence and flexibility of the digital human are increased. Moreover, the digital human is completely controlled by the script, thereby reducing the dependence of the digital human on manual control during the digital human video live broadcast.

[0133] In a possible implementation manner, during the process of the first digital human performing a video live broadcast, the recognition module 510 further obtains the first bullet screen text in the live broadcast, and then determines whether the first bullet screen text triggers an audience interaction event. For example, if the first bullet screen text is "Is free shipping supported?", it is determined that the first bullet screen text is an audience interaction event for the product selling point, that is, it is determined that the first bullet screen text triggers an audience interaction event.

[0134] Then, the imitation module 520 inputs the first barrage text into the second digital human model, obtains the first response text output by the second digital human model, and controls the digital human to display the first response text. For example, during the video live broadcast by the first digital human, the cloud management platform obtains the first barrage text as "Is free shipping supported?", obtains the first response text as "Supported" through the second digital human model, and then controls the first digital human to reply "Supported" to respond to the barrage.

[0135] In the above text, in combination with Figures 2 to 5 , the video live broadcast method based on digital human technology provided by the present application is described in detail. Next, in combination with Figures 4 to 7 , the video live broadcast device and computing device for executing the above method provided by the present application will be described.

[0136] Figure 6 FIG. is a schematic structural diagram of a video live broadcast device provided by the present application. This video live broadcast device can be used to implement the above video live broadcast method based on digital human technology, and thus can also achieve the beneficial effects of the above method.

[0137] As Figure 6 shown, the video live broadcast device 600 includes an acquisition module 610 and a processing module 620; the acquisition module 610 is used to acquire the first information input by the user, and the first information is the information to be displayed during the live broadcast; the processing module 620 is used to configure the first digital human model to generate a script according to the first information, and configure the first digital human to perform a video live broadcast according to the script. The script includes at least one piece of text information, and each piece of text information in the at least one piece of text information is associated with a performance style feature.

[0138] Among them, both the acquisition module 610 and the processing module 620 can be implemented by software or can be implemented by hardware. Exemplarily, next, taking the acquisition module 610 as an example, the implementation manner of the acquisition module 610 will be introduced. Similarly, the implementation manner of the processing module 620 can refer to the implementation manner of the acquisition module 610.

[0139] As an example of a software functional unit, the acquisition module 610 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the acquisition module 610 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers for running the code may be distributed in the same availability zone (AZ) or in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Usually, one region may include multiple AZs.

[0140] Similarly, the multiple hosts / virtual machines / containers for running the code may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is achieved through the communication gateway.

[0141] As an example of a hardware functional unit, the acquisition module 610 may include at least one computing device, such as a server. Alternatively, the acquisition module 610 may also be a device implemented using a central processing unit (CPU), or may be implemented using an application-specific integrated circuit (ASIC), a programmable logic device (PLD), etc. The above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a data processing unit (DPU), a neural network processing unit (NPU), a system on chip (SoC), an offload card, an acceleration card, or any combination thereof.

[0142] The multiple computing devices included in the obtaining module 610 can be distributed in the same region or in different regions. The multiple computing devices included in the obtaining module 610 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the obtaining module 610 can be distributed in the same VPC or in multiple VPCs. Among them, the multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, GALs, DPUs, NPUs, SoCs, offloading cards, and acceleration cards.

[0143] It should be noted that in other embodiments, the obtaining module 610 and the processing module 620 can be respectively used to execute any steps in the video live broadcast method based on digital human technology.

[0144] In a possible implementation, the obtaining module 610 is further configured to: before configuring the first digital human model to generate a script according to the first information, obtain second information input by the user, where the second information is associated with the live video; the processing module 620 is further configured to: identify the voice and key frames in the live video according to the second information, convert the voice in the live video into a first text, and convert the first text into a structured text according to the structure of the first text; extract the text features of the structured text, the picture style features of the key frames, and the sound style features of the voice, and learn the first digital human model according to the text features, the picture style features, and the sound style features. Among them, the key frames include one or more of the following: key frames showing the actions of the anchor, key frames showing the expressions of the anchor, and key frames where the picture changes due to camera changes; the text features of the structured text include one or more of the following: structural features, semantic preference features, and language style features; the picture style features of the key frames include one or more of the following: action features, expression features, and picture features;

[0145] In a possible implementation, the text features of the script are learned based on the structural features of the structured text; or, the text features of the script are learned based on the structural features and semantic preference features of the structured text; or, the text features of the script are learned based on the structural features, semantic preference features, and language style features of the structured text.

[0146] In a possible implementation, the performance style features of the script include screen style features, and the screen style features of the script are learned based on the action features of the key frames; or, the screen style features of the script are learned based on the action features and expression features of the key frames; or, the screen style features of the script are learned based on the action features, expression features, and screen features of the key frames; the processing module 620 is specifically configured to perform one or more of the following: control the actions of the first digital human during video live streaming, control the expressions of the first digital human during video live streaming, and control the display screen of the first digital human during video live streaming.

[0147] In a possible implementation, the performance style features of the script include sound style features, and the sound style features of the script are learned based on the sound style features of the speech; the processing module 620 is specifically configured to: control the sound of the first digital human during video live streaming.

[0148] In a possible implementation, a second digital human model of the user is further stored in the infrastructure, the second digital human model is associated with the first digital human, and the obtaining module 610 is further configured to: obtain a first bullet screen text during video live streaming of the first digital human, and the first bullet screen text triggers an audience interaction event; the processing module 620 is further configured to: configure the second digital human model to generate a first response text according to the first bullet screen text, and configure the first digital human to perform video live streaming according to the first response text, and the first response text is associated with the bullet screen response preference features in the live video.

[0149] Figure 7 FIG. is a schematic structural diagram of a computing device provided by the present application. The computing device can be used to implement the above video live streaming method based on digital human technology, and thus can also achieve the beneficial effects of the above method.

[0150] As Figure 7 shown, the computing device 100 includes a processor 104 and a communication interface 108. The processor 104 and the communication interface 108 are coupled to each other. It can be understood that the communication interface 108 can be a transceiver or an input / output interface. Optionally, the computing device 100 may further include a memory 106 for storing instructions executed by the processor 104 or storing input data required for the processor 104 to run the instructions or storing data generated after the processor 104 runs the instructions.

[0151] It can be understood that the processor 104 in the embodiments of the present application may include any one or more of computing devices such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), an ASIC, an FPGA, a CPLD, an NPU, a SoC, an offloading card, an acceleration card, etc.

[0152] The memory 106 may include a volatile memory, such as a random access memory (RAM). The processor 104 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD). In addition, the memory 106 may also be implemented by a storage class memory (SCM), a phase change memory (PCM), or other types of storage media.

[0153] It is worth noting that the same type of storage medium can be configured in the same computing device to implement the function of the memory 106, or two or more types of storage media can be configured to implement the function of the memory 106. The present application does not make any limitations in this regard.

[0154] The memory 106 stores executable program codes, and the processor 104 executes the executable program codes to respectively implement the functions of the foregoing acquisition module 610 and processing module 620, so as to implement the video live broadcast method based on the digital human technology. That is to say, the memory 106 stores instructions for executing the video live broadcast method based on the digital human technology.

[0155] Alternatively, the memory 106 stores executable codes, and the processor 104 executes the executable codes to respectively implement the functions of the foregoing video live broadcast device 600, so as to implement the video live broadcast method based on the digital human technology. That is to say, the memory 106 stores instructions for executing the video live broadcast method based on the digital human technology.

[0156] The communication interface 103 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement the communication between the computing device 100 and other devices or communication networks.

[0157] In a possible implementation, the computing device 100 may also include a chip system. The chip system includes a processor and a power supply circuit. The power supply circuit is used to supply power to the processor, and the processor is used to execute the operation steps corresponding to the video live broadcast method based on digital human technology. For the sake of brevity, it will not be elaborated here. Among them, the processor can be implemented by a GPU, or can be implemented by computing devices or AI chips such as DPU, NPU, XPU, SoC, offloading card, and acceleration card.

[0158] In a possible implementation, the computing device 100 may include multiple types of processors 104, that is, the computing device 100 is a heterogeneous device. For example, the computing device 100 includes a CPU and a GPU, and at least one of the processors 104 can execute the operation steps corresponding to the video live broadcast method based on digital human technology. For the sake of brevity, it will not be elaborated here.

[0159] As a possible implementation, the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0160] As a possible implementation, as Figure 8 shown, the present application also provides a computing device cluster including at least one computing device 100. Instructions for executing the video live broadcast method based on digital human technology may be stored in the same manner in the memories 106 of one or more of the computing devices 100 in the computing device cluster.

[0161] In some possible implementations, partial instructions for executing the video live broadcast method based on digital human technology may also be stored separately in the memories 106 of one or more of the computing devices 100 in the computing device cluster. In other words, a combination of one or more computing devices 100 can jointly execute the instructions of the video live broadcast method based on digital human technology.

[0162] It should be noted that the memories 106 in different computing devices 100 in the computing device cluster may store different instructions, which are respectively used to execute partial functions of the video live broadcast device 600. That is, the instructions stored in the memories 106 of different computing devices 100 can implement the functions of one or more of the acquisition module 610 and the processing module 620.

[0163] In some possible implementations, one or more of the computing devices in the computing device cluster can be connected through a network. Among them, the network can be a wide area network or a local area network, etc. Figure 9A possible implementation is shown. As Figure 9 shown, two computing devices 100A and 100B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation, the instructions for implementing the function of the acquisition module 610 are stored in the memory 106 of the computing device 100A. At the same time, the instructions for implementing the function of the processing module 620 are stored in the memory 106 of the computing device 100B.

[0164] Figure 9 The connection method between the computing device clusters shown can be considered that since the video live streaming method based on digital human technology provided in this application requires a large amount of storage for the questions in the questionnaire and the options associated with each question, it is considered to hand over the function implemented by the processing module 620 to the computing device 100B for execution.

[0165] It should be understood that Figure 9 the functions of the computing device 100A shown in can also be completed by multiple computing devices 100. Similarly, the functions of the computing device 100B can also be completed by multiple computing devices 100.

[0166] This application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similarly referred to the Figure 7 and Figure 8 connection method of the described computing device cluster. The difference is that the memory 106 in one or more computing devices 100 in this computing device cluster can store the same instructions for implementing the video live streaming method based on digital human technology.

[0167] In some possible implementations, the memory 106 of one or more computing devices 100 in this computing device cluster can also store partial instructions for implementing the video live streaming method based on digital human technology respectively. In other words, a combination of one or more computing devices 100 can jointly execute the instructions for implementing the video live streaming method based on digital human technology.

[0168] It should be noted that the memory 106 in different computing devices 100 in the computing device cluster can store different instructions for implementing the functions of the video live streaming device 600.

[0169] This application also provides a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it causes at least one computing device to execute the video live streaming method based on digital human technology.

[0170] The present application also provides a computer-readable storage medium. The computer-readable storage medium may be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute a video live streaming method based on digital human technology.

[0171] It can be understood that the various digital numbers involved in the embodiments of the present application are only for convenience of description and are not used to limit the scope of the embodiments of the present application. The magnitudes of the serial numbers of the above processes do not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic.

Claims

1. A video live streaming method based on digital human technology, characterized in that, The method is applied to a cloud management platform for managing infrastructure in which a first digital human model of a user is stored. The first digital human model is associated with a first digital human. The method includes: Obtain first information input by the user, where the first information is information for display during a live broadcast; Configure the first digital human model to generate a script according to the first information. The script includes at least one piece of text information, and each piece of text information in the at least one piece of text information is associated with a performance style feature; Configure the first digital human to conduct a video live broadcast according to the script.

2. The method according to claim 1, wherein, Before configuring the first digital human model to generate a script according to the first information, the method further includes: Obtain second information input by the user, where the second information is associated with the live broadcast video; Identify the speech and key frames in the live broadcast video according to the second information. The key frames include one or more of the following: key frames showing the actions of the anchor, key frames showing the expressions of the anchor, and key frames where the picture changes due to a camera change; Convert the speech in the live broadcast video into a first text, and convert the first text into a structured text according to the structure of the first text; Extract the text features of the structured text, the picture style features of the key frames, and the sound style features of the speech. Among them, the text features of the structured text include one or more of the following: structural features, semantic preference features, language style features. The picture style features of the key frames include one or more of the following: action features, expression features, picture features; Learn the first digital human model according to the text features, the picture style features, and the sound style features.

3. The method according to claim 1 or 2, characterized in that, The text features of the script are learned based on the structural features of the structured text; or the text features of the script are learned based on the structural features and semantic preference features of the structured text; or the text features of the script are learned based on the structural features, semantic preference features, and language style features of the structured text.

4. The method according to claim 1 or 2, characterized in that, The performance style features of the script include picture style features. The picture style features of the script are learned based on the action features of the key frames; or the picture style features of the script are learned based on the action features and expression features of the key frames; or the picture style features of the script are learned based on the action features, expression features, and picture features of the key frames; The configuring the first digital human to conduct a video live broadcast according to the script includes one or more of the following: the actions of the first digital human during the video live broadcast, the expressions of the first digital human during the video live broadcast, and the display picture of the first digital human during the video live broadcast.

5. The method according to claim 1 or 2, characterized in that, The performance style features of the script include sound style features. The sound style features of the script are learned based on the sound style features of the speech; The configuring the first digital human to conduct a video live broadcast according to the script includes the sound of the first digital human during the video live broadcast.

6. The method according to any one of claims 1-5, characterized in that, The infrastructure also stores a second digital human model of the user, and the second digital human model is associated with a first digital human. The method further includes: Obtaining first barrage text when the first digital human conducts a video live broadcast, where the first barrage text triggers an audience interaction event; Configuring the second digital human model to generate a first response text according to the first barrage text, and the first response text is associated with barrage response preference features in the live video; Configuring the first digital human to conduct a video live broadcast according to the first response text.

7. A video live broadcast device based on digital human technology, characterized in that, The device is used to manage the infrastructure, and the infrastructure stores a first digital human model of the user, and the first digital human model is associated with a first digital human. The device includes: An acquisition module, configured to acquire first information input by the user, where the first information is information to be displayed during the live broadcast; A processing module, configured to configure the first digital human model to generate a script according to the first information, where the script includes at least one piece of text information, and each piece of text information in the at least one piece of text information is associated with a performance style feature; Configuring the first digital human to conduct a video live broadcast according to the script.

8. The device according to claim 7, characterized in that, The acquisition module is further configured to: Before configuring the first digital human model to generate a script according to the first information, acquire second information input by the user, where the second information is associated with the live video; The processing module is further configured to: Identify the voice and key frames in the live video according to the second information, where the key frames include one or more of the following: key frames showing the actions of the anchor, key frames showing the expressions of the anchor, key frames where the picture changes due to camera changes; Convert the voice in the live video into a first text, and convert the first text into a structured text according to the structure of the first text; Extract the text features of the structured text, the picture style features of the key frames, and the sound style features of the voice, where the text features of the structured text include one or more of the following: structural features, semantic preference features, language style features, and the picture style features of the key frames include one or more of the following: action features, expression features, picture features; Learning the first digital human model according to the text features, the picture style features, and the sound style features.

9. The device according to claim 7 or 8, characterized in that, The text features of the script are learned based on the structural features of the structured text; or, the text features of the script are learned based on the structural features and semantic preference features of the structured text; or, the text features of the script are learned based on the structural features, semantic preference features, and language style features of the structured text.

10. The device according to claim 7 or 8, characterized in that, The performance style features of the script include picture style features, and the picture style features of the script are learned based on the action features of the key frames; or, the picture style features of the script are learned based on the action features and expression features of the key frames; or, the picture style features of the script are learned based on the action features, expression features, and picture features of the key frames; The processing module is specifically configured to perform one or more of the following: Control the actions of the first digital human during video live streaming, control the expressions of the first digital human during video live streaming, and control the display screen of the first digital human during video live streaming.

11. The device according to claim 7 or 8, characterized in that, The performance style characteristics of the script include voice style characteristics, and the voice style characteristics of the script are learned based on the voice style characteristics of the speech. The processing module is specifically configured to: Control the voice of the first digital human during video live streaming.

12. The device according to any one of claims 7-11, characterized in that The infrastructure also stores a second digital human model of the user, and the second digital human model is associated with the first digital human. The acquisition module is further configured to: Acquire the first barrage text during the video live streaming of the first digital human, and the first barrage text triggers an audience interaction event. The processing module is further configured to: Configure the second digital human model to generate a first response text according to the first barrage text, and the first response text is associated with the barrage response preference characteristics in the live video. Configure the first digital human to perform video live streaming according to the first response text.

13. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium includes a program, and when the program runs on the device, the device is caused to execute the method according to any one of claims 1-6.

14. A cluster of computing devices, characterized in that, The computing device cluster includes at least one computing device, and the at least one computing device includes at least one chip and a memory. The at least one chip is configured to read and execute program instructions stored in the memory to implement the method according to any one of claims 1-6.

15. A program product, characterized in that, When the program product runs on the device, the device is caused to execute the method according to any one of claims 1-6.