Video live streaming method and apparatus based on digital human technology
By generating scripts to control digital human live streams and utilizing structured text and style feature learning, the problem of insufficient intelligence in digital human live streams is solved, achieving a more human-like and flexible live stream effect.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
- Filing Date
- 2024-09-27
- Publication Date
- 2026-05-07
AI Technical Summary
The lack of intelligence in digital human live streaming results in poor quality and heavy reliance on human control, which affects the live streaming effect.
Scripts are generated by acquiring user input information, and digital human models are configured to conduct live video broadcasts based on the scripts. Scripts are generated by learning from structured text, visual style features, and audio style features, reducing reliance on human control.
It improves the anthropomorphism and flexibility of digital human live streaming, reduces reliance on human control, and enhances the intelligence and performance of digital humans.
Smart Images

Figure CN2024121996_07052026_PF_FP_ABST
Abstract
Description
A video live streaming method and device based on digital human technology
[0001] Cross-references to related applications
[0002] This application claims priority to Chinese Patent Application No. 202410050578.1, filed on January 13, 2024, entitled "An Interactive Method and Apparatus Based on Digital Human Technology", the entire contents of which are incorporated herein by reference; and Chinese Patent Application No. 202410543832.1, filed on April 30, 2024, entitled "A Video Live Streaming Method and Apparatus Based on Digital Human Technology", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of digital human technology, and in particular to a video live streaming method and apparatus based on digital human technology. Background Technology
[0004] With the continuous development of computer technology, the internet live streaming industry has become a major emerging industry. Internet live streaming refers to the process of capturing the streamer's image and audio using various types of video recording equipment and then streaming the video stream to users' terminal devices via a server.
[0005] Based on this, digital humans (or meta humans) designed based on real people are gradually being applied to the internet live streaming industry due to their advantages such as lifelike appearance, high reusability, and wide range of applications. Digital humans are virtual images created using technologies such as modeling, motion capture, and artificial intelligence (AI) to possess human-like or humanoid physical characteristics and behavioral patterns, and presented through display devices.
[0006] In the business scenario of digital human live streaming, the insufficient intelligence of digital humans leads to poor quality of live streams. In order to make digital humans more human-like and closer to real-life live streams, human control is often required, which makes digital human live streaming dependent on human control and limited by the live streaming experience of the control personnel.
[0007] Summary of the Invention
[0008] This application provides a video live streaming method and apparatus based on digital human technology, which increases the intelligence and flexibility of digital humans and reduces the reliance on human control during digital human video live streaming.
[0009] Firstly, this application provides a video live streaming method based on digital human technology. This method can be applied to a cloud management platform for managing infrastructure. The infrastructure can be a device, a module of the device (such as a chip), or a system corresponding to the device. The device can be a terminal device or a network device (such as a server). The infrastructure stores a first digital human model of the user, which is associated with a first digital human. Based on this, the method can include the following steps: obtaining first information input by the user, which is information to be displayed during the live stream; then configuring the first digital human model to generate a script based on the first information, and further configuring the first digital human to perform a live video stream according to the script. The script includes at least one segment of text information, and each segment of text information is associated with performance style features.
[0010] In this method, the text content of the script can be segmented and displayed according to semantics, words, and symbols (commas, periods, etc.). That is, each segment can be understood as a piece of text information, thus the script can be understood as consisting of at least one piece of text information. Any piece of text information can be a word, a character, a sentence, a symbol, etc., without specific limitations. The performance style features associated with each piece of text information can be understood as performance style tags associated with that piece of text information. During live video streaming, the first digital human can be controlled to display corresponding text information based on these performance style tags, thereby increasing the intelligence and flexibility of the digital human, making the digital human in the live video stream more lifelike, and improving the anthropomorphism of the digital human in the live video stream. Furthermore, the digital human is entirely controlled by the script, thus reducing the reliance on human control during the digital human live video stream.
[0011] In one possible implementation, before configuring the first digital human model to generate a script based on the first information, second information input by the user is obtained, which is associated with the live video. Then, based on the second information, speech and keyframes in the live video are identified, and the speech is converted into first text. The first text is then converted into structured text based on its structure. Next, text features of the structured text, visual style features of the keyframes, and vocal style features of the speech are extracted. Finally, the first digital human model is learned based on the text features, visual style features, and vocal style features. The keyframes include one or more of the following: keyframes showing the anchor's actions, keyframes showing the anchor's expressions, and keyframes showing changes in the visuals due to camera changes. The text features of the structured text include one or more of the following: structural features, semantic preference features, and language style features. The visual style features of the keyframes include one or more of the following: action features, facial expression features, and visual features.
[0012] In this implementation, structured text represents text information with structural features and semantic preference features. Structural features can be understood as the structural description of multiple objects in the text, such as the order of description of multiple objects, the compositional structure of each object, and the order of description of each component within that structure. Semantic preference features can be understood as the descriptive preferences for each object. Language style features represent language usage preferences in the structured text; or, in other words, language style features represent the language preferences used by the broadcaster during a live stream, such as the broadcaster's preference for using dialects, poetry, proverbs, a combination of Chinese and English, or famous quotes. Therefore, structured text can be understood as textual content that reflects the textual structure.
[0013] Voice style characteristics can include loudness, tone of voice, etc., and can be understood as the characteristics of the anchor's voice in a live video, such as the anchor's accent and tone of voice when describing products.
[0014] Live video can be understood as training samples. Therefore, the first digital human model is used to learn the relevant features of the live video (including text features, visual style features, and audio style features). The first digital human model is also used to output a script, so that the relevant features of the output script are associated with (or similar to) the relevant features of the live video. This enables the features of the live video controlled by the script to be similar to the features of the live video. Therefore, the features of the digital human in the live video are similar to the features of the anchor in the live video, thereby improving the anthropomorphism of the digital human in the live video.
[0015] In one possible implementation, the text features of the script are learned based on the structural features of the structured text; or, the text features of the script are learned based on the structural features and semantic preference features of the structured text; or, the text features of the script are learned based on the structural features, semantic preference features, and language style features of the structured text.
[0016] In this implementation, the textual features of the script include one or more of the following: structural features and semantic preference features. This can be understood as follows: the script's structural features are learned based on the structural features of structured text; the script's semantic preference features are learned based on the semantic preference features of structured text; and the script's language style features are learned based on the language style features of structured text. In other words, the script's structural features are similar to the structural features of structured text; the script's semantic preference features are similar to the semantic preference features of structured text; and the script's language style features are similar to the language style features of structured text. That is, the script is similar to the structured text of the live video. Therefore, based on this script, the digital human in the live video broadcast is similar to the host in the live video, thereby improving the anthropomorphism of the digital human live video broadcast.
[0017] In one possible implementation, the performance style features of the script include visual style features, which are learned based on the motion features of the keyframes; or, the visual style features of the script are learned based on the motion and facial expression features of the keyframes; or, the visual style features of the script are learned based on the motion, facial expression, and visual features of the keyframes; configuring the first digital human to perform live video streaming according to the script includes one or more of the following: the actions of the first digital human during live video streaming, the facial expressions of the first digital human during live video streaming, and the display screen of the first digital human during live video streaming.
[0018] In this implementation, it can be understood that the visual style of the script is similar to that of the live video. In other words, during the live video broadcast, the actions of the first digital human are similar to those of the host, the expressions of the first digital human are similar to those of the host, and the displayed images are similar to those of the live video, thereby further improving the anthropomorphism of the digital human live video broadcast.
[0019] In one possible implementation, the performance style features of the script include voice style features, which are learned based on the voice style features of the speech; configuring the first digital human to conduct live video streaming according to the script includes the voice of the first digital human during live video streaming.
[0020] In this implementation, the voice style characteristics of the script are similar to those of the live video. Therefore, when the digital human is controlled to conduct live video broadcasts according to the script, the voice characteristics of the digital human are similar to those of the anchor in the live video, thereby further improving the anthropomorphism of the digital human's live video broadcasts.
[0021] In one possible implementation, the infrastructure also stores a second digital human model of the user, which is associated with a first digital human. To this end, the infrastructure can acquire the first bullet screen text during a live video broadcast by the first digital human, which triggers a viewer interaction event. Then, the second digital human model is configured to generate a first response text based on the first bullet screen text, and the first digital human is configured to broadcast the video based on the first response text. The first response text is associated with bullet screen response preference features in the live video.
[0022] In this implementation, the bullet screen response preference features can be understood as the streamer's response preferences to bullet screen comments during a live video broadcast. For example, a streamer might habitually respond to bullet screen comments about product selling points or casual conversation. Based on this, the bullet screen comments responded to by the digital human during a live video broadcast closely resemble those responded to by the streamer, thereby enhancing the anthropomorphism of the digital human during the live video broadcast. Optionally, the first response text may also be associated with performance style features.
[0023] Secondly, this application provides a video live streaming device based on digital human technology. This device manages infrastructure in which a first digital human model of a user is stored. The first digital human model is associated with a first digital human. The device may include an acquisition module and a processing module. The acquisition module is used to acquire first information input by the user, which is information to be displayed during live streaming. The processing module is used to configure the first digital human model to generate a script based on the first information and to configure the first digital human to perform live video streaming according to the script. The script includes at least one segment of text information, and each segment of text information is associated with performance style characteristics.
[0024] In one possible implementation, the acquisition module is further configured to: acquire second information input by the user before configuring the first digital human model to generate a script based on the first information, wherein the second information is associated with the live video;
[0025] The processing module is further configured to: identify speech and keyframes in the live video based on the second information; convert the speech in the live video into first text; convert the first text into structured text based on the structure of the first text; extract text features of the structured text, visual style features of the keyframes, and vocal style features of the speech; and learn the first digital human model based on the text features, visual style features, and vocal style features. The keyframes include one or more of the following: keyframes showing the anchor's actions, keyframes showing the anchor's expressions, and keyframes showing changes in the visuals due to camera changes; the text features of the structured text include one or more of the following: structural features, semantic preference features, and language style features; the visual style features of the keyframes include one or more of the following: action features, expression features, and visual features.
[0026] In one possible implementation, the text features of the script are learned based on the structural features of the structured text; or, the text features of the script are learned based on the structural features and semantic preference features of the structured text; or, the text features of the script are learned based on the structural features, semantic preference features, and language style features of the structured text.
[0027] In one possible implementation, the performance style features of the script include visual style features, which are learned based on the action features of the keyframes; or, the visual style features of the script are learned based on the action and facial expression features of the keyframes; or, the visual style features of the script are learned based on the action, facial expression, and visual features of the keyframes; the processing module is specifically used for one or more of the following: controlling the actions of the first digital human during live video streaming, controlling the facial expressions of the first digital human during live video streaming, and controlling the display screen of the first digital human during live video streaming.
[0028] In one possible implementation, the performance style features of the script include voice style features, which are learned based on the voice style features of the speech; the processing module is specifically used to control the voice of the first digital human during live video streaming.
[0029] In one possible implementation, the infrastructure further stores a second digital human model of the user, which is associated with a first digital human. The acquisition module is further configured to: acquire a first bullet screen text during a live video broadcast by the first digital human, wherein the first bullet screen text triggers a viewer interaction event; the processing module is further configured to: configure the second digital human model to generate a first response text based on the first bullet screen text, and configure the first digital human to conduct a live video broadcast based on the first response text, wherein the first response text is associated with bullet screen response preference features in the live video.
[0030] Thirdly, this application provides a computer device, including a processor, a memory, a communication interface, and a bus. The processor, memory, and communication interface are connected via the bus and communicate with each other. The memory stores computer execution instructions. When the computer device is running, the processor executes the computer execution instructions in the memory to perform the operation steps of the method described in the first aspect or any possible implementation of the first aspect using the hardware resources in the computer device.
[0031] Fourthly, this application provides a computer device cluster, including multiple computer devices as described in the third aspect, each computer device being used independently or jointly to perform the operational steps of the method described in the first aspect or any possible implementation of the first aspect.
[0032] Fifthly, this application provides a chip system comprising at least one chip and a memory, wherein the at least one chip is used to read and execute a program stored in the memory to implement the operational steps of the method described in the first aspect or any possible implementation thereof.
[0033] A sixth aspect provides a non-volatile computer-readable storage medium comprising a program that, when executed on a device, causes the device to perform the operational steps of the method described in the first aspect or any possible implementation thereof.
[0034] In a seventh aspect, a computer program product is provided, which, when run on a device, causes the device to perform the operational steps of the method described in the first aspect or any possible implementation thereof.
[0035] Based on the implementations provided in the above aspects, this application can be further combined to provide more implementations. Attached Figure Description
[0036] Figure 1 is a schematic diagram of an application scenario provided in this application;
[0037] Figure 2 is a flowchart illustrating a video live streaming method provided in this application;
[0038] Figure 3 is a flowchart illustrating another video live streaming method provided in this application;
[0039] Figure 4 is a schematic diagram of a process for learning a first digital human model provided in this application;
[0040] Figure 5 is a schematic diagram of the system architecture of a cloud management platform provided in this application;
[0041] Figure 6 is a structural schematic diagram of a video live streaming device based on digital human technology provided in this application;
[0042] Figure 7 is a schematic diagram of the structure of a computing device provided in this application;
[0043] Figure 8 is a schematic diagram of another computing device provided in this application;
[0044] Figure 9 is a structural schematic diagram of another computing device provided in this application. Detailed Implementation
[0045] With the continuous development of computer technology, the internet live streaming industry has become a major emerging industry. Internet live streaming refers to the process of capturing images and audio of a streamer using various types of video recording equipment and then streaming the video stream to users' terminal devices via a server. Based on this, digital humans (or meta-humans) designed based on real people are increasingly being used in the internet live streaming industry due to their lifelike appearance, high reusability, and wide range of applications. Digital humans are virtual avatars created using technologies such as modeling, motion capture, and AI, possessing human-like or humanoid physical characteristics and behavioral patterns, and presented through display devices.
[0046] In current digital human live streaming technologies, the lack of intelligence in digital human live streams leads to poor quality. To make digital humans more human-like and closer to real-life live streams, human control of the digital humans is often required, including the following:
[0047] Parameters are set based on the experience of business personnel, and scripts and talking points are generated during the live broadcast using a large model.
[0048] During the live stream, business personnel control a digital human to respond in real time. This includes responding to viewer comments, monitoring the effectiveness of live-stream sales, and handling temporary event management.
[0049] During the live stream, business personnel control the digital human's movements and expressions according to preset requirements.
[0050] In summary, in real-world business scenarios involving digital human live streaming, digital humans exhibit a strong reliance on business personnel. The generation of scripted dialogue, the decision-making process for script arrangement, and the real-time control during the performance are all achieved manually. Therefore, digital human live streaming is largely limited by the experience accumulated by the business personnel themselves.
[0051] Therefore, this application provides a video live streaming method based on digital human technology. This method is used to conduct digital human live streaming, increase the intelligence and flexibility of digital human live streaming, make the digital human in the video live streaming lifelike, improve the anthropomorphism of digital human live streaming, and reduce the dependence on human control during digital human video live streaming.
[0052] The present application will now be described in detail with reference to the accompanying drawings.
[0053] First, the application scenario to which this application applies will be introduced. Referring to Figure 1, a user can send first information to a server 120 via a client running on terminal device 110. This first information is for display during live streaming. After obtaining the first information, the cloud management platform (or server) running on server 120 inputs the first information into the user's first digital human model, thereby obtaining a script output by the first digital human model. This script includes at least one piece of text information, and each piece of text information is associated with performance style features. Then, the first digital human is controlled to conduct a live video stream according to the script, and the live video stream is displayed on terminal device 110. In other words, terminal device 110 can display a live streaming room, which includes a live video stream based on the script. Optionally, the live video stream can be displayed by a client running on terminal device 110.
[0054] To make the purpose, technical solution and advantages of this application clearer, some terms involved in the embodiments of this application will be explained below.
[0055] (I) Digital Human Technology
[0056] Digital human technology is a technique that uses modeling, motion capture, and artificial intelligence (AI) to create virtual avatars with human-like or humanoid physical features and behavioral patterns, which are then displayed on a screen. The control of a digital human is typically achieved through a script, which can be understood as specific program logic. The script may include the actions, language, and text displayed by the digital human.
[0057] (II) Structured Text
[0058] Structured text is obtained by transforming the first text in the live video according to its structure. The structure of structured text can be understood as a tree structure, where child nodes represent structured words in the live video, and each child node is a component of its parent node. In other words, structured text includes both the structural content of each object in the first text and the textual content of each object.
[0059] For example, the live video is a video corresponding to a product-selling live stream. The first text corresponding to this live video includes: the first product and the second product; the origin, price, and color of the first product; and the size, performance, and appearance of the second product. Based on this, the structured text includes: root node, child node 1, child node 2, ..., child node 8. The correspondence between the structured content and the text content is as follows:
[0060] The root node represents the live video.
[0061] Child node 1 represents the first product, and child node 2 represents the second product. The live streaming order of child node 1 is before that of child node 2. The root node is the parent node of child node 1 and child node 2.
[0062] Child node 3 represents the origin of the first product, child node 4 represents the price of the first product, and child node 5 represents the color of the first product. That is, child node 1 is the parent node of child nodes 3, 4, and 5, and the live streaming order is child node 3, child node 4, and child node 5. Child node 6 represents the volume of the second product, child node 7 represents the performance of the second product, and child node 8 represents the appearance of the second product. That is, child node 2 is the parent node of child nodes 6, 7, and 8, and the live streaming order is child node 6, child node 7, and child node 8.
[0063] Based on the above description, structural features and semantic preference features of structured text can be extracted from it. Semantic preference features can be understood as descriptive preferences (or descriptive habits) for each object. For example, for product number one, the descriptive preference represents describing three selling points of product number one: place of origin, price, and color.
[0064] (III) Linguistic Style Features
[0065] Language style features represent the language usage preferences of structured text. Language style features can include poetry style, allusion style, music style, etc. Taking poetry style and allusion style as examples, for the description of the origin of the first product, the structured text prefers to use the poetry corresponding to the origin to describe the origin, and prefers to use the allusion corresponding to the origin to describe the origin.
[0066] (iv) Visual style characteristics
[0067] Visual style features indicate the visual preferences used in the live stream video. These features can include action features, facial expression features, or overall visual characteristics. Action features refer to the body movements made by the broadcaster in the live stream video, such as tilting the head or raising a hand. Facial expression features refer to the facial expressions made by the broadcaster in the live stream video, such as smiling, laughing, or raising an eyebrow. Visual characteristics refer to the live stream display that the broadcaster needs to show.
[0068] For example, the live stream display may include footage after camera movement or switching. For instance, when describing the color of a product, the live stream display might be a close-up of the product; when describing the size of the product, it might be a close-up comparing the product to a reference object. The live stream display may also include footage showing changes in camera focus, such as when describing the appearance of a product, the live stream display might show changes in camera focus related to the product. It should be understood that the live stream display may also include footage showing changes in distance, etc., but no specific limitations are imposed on the live stream display here.
[0069] Optionally, the visual style features are extracted from keyframes in the live video. These keyframes include one or more of the following: keyframes showing the broadcaster's actions, keyframes showing the broadcaster's facial expressions, and keyframes where the visual changes due to camera changes. Examples of scenarios where the visual changes due to camera changes include: camera movement / switching causing visual transitions, and changes in camera focal length causing changes in depth of field. It should be understood that keyframes can be multiple consecutive image frames; therefore, keyframes can also be understood as video segments that need to be learned.
[0070] (V) Characteristics of vocal style
[0071] Voice style features refer to the voice preferences used by the host in a live video. Voice style features can include tone of voice, accent, volume, etc. For example, the host's accent and the tone of voice when describing products.
[0072] (vi) Characteristics of bullet screen response preferences
[0073] The bullet comment response preference feature indicates the streamer's preference in responding to bullet comments during a live stream. For example, the bullet comment response preference feature indicates that the streamer prefers to respond to bullet comments that highlight product selling points, or prefers to respond to bullet comments that contain casual chat.
[0074] In one possible implementation, the bullet screen response preference features are extracted from the second bullet screen text and the second response text in the live video. The second response text is obtained from the first text based on the second bullet screen text. The bullet screen response preference features, the second response text, and the second bullet screen text are then used to train the second digital human model.
[0075] (VII) The First Digital Human Model
[0076] The first digital human model was learned based on live video as training samples. This model was used to output scripts, ensuring that the features of the output scripts were similar to those of the training samples.
[0077] Optionally, the training process for the first digital human model may include the following:
[0078] S1: Obtain the second information and determine the live video based on the second information. The second information is associated with the live video; for example, the second information may be the link or identifier of the live video, or it may be the source file of the live video. Therefore, the live video can be determined based on the second information.
[0079] S2: Recognize the audio and keyframes in the live video. Keyframes include one or more of the following: keyframes showing the broadcaster's actions, keyframes showing the broadcaster's facial expressions, and keyframes showing changes in the image due to camera changes.
[0080] Live video can be a video streamed by a real person (hereinafter referred to as the streamer), and the audio can be understood as the streamer's voice in the live video. There are multiple keyframes, so keyframes can be understood as video segments that need to be learned. For example, keyframes showing the streamer's actions can include image frames of the streamer's body movements in the live video, such as the streamer tilting their head or raising their hand; keyframes showing the streamer's facial expressions can include image frames of the streamer's facial expressions in the live video, such as the streamer laughing or raising their eyebrows; keyframes that change the scene due to camera changes can include the scene after the camera moves or switches, for example, when describing the color of a product, the live video shows a close-up of the product, and when describing the size of the product, the live video shows a close-up of the product compared to a reference object.
[0081] Optionally, keyframes that cause image changes due to lens changes may also include images showing changes in lens focal length. For example, when describing the appearance of a product, the live stream displays image frames showing changes in lens focal length related to the product. It should be understood that keyframes that cause image changes due to lens changes may also include image frames showing changes in distance, etc., and are not specifically limited here.
[0082] S3: Convert the audio in the live video into first text, and convert the first text into structured text based on its structure.
[0083] The first text can be understood as the audio-text content of the broadcaster. The structure of the first text can be understood as a structural description of multiple objects in the first text. This structural description may include the description order of the multiple objects in the first text, the composition structure of each object, and the description order of the components of each object, etc.
[0084] In one possible implementation, the speech in the live video can be converted into the first text using an Automatic Speech Recognition (ASR) model. The ASR model includes, but is not limited to, DeepSpeech, ESPnet, etc., and this application does not limit it.
[0085] After obtaining the first text, convert the first text into structured text, as described in (II) above, which will not be repeated here.
[0086] S4: Extract the text features of the structured text (also known as the text features of the live video), the visual style features of the keyframes (also known as the visual style features of the live video), and the audio style features of the speech in the live video (also known as the audio style features of the live video), and learn the first digital human model based on the text features, visual style features, and audio style features of the live video.
[0087] Referring to the descriptions in (ii) and (iii) above, the text features of the live video include one or more of the following: structural features, semantic preference features, and language style features. In one possible implementation, a Large Language Model (LLM) can extract structural features, semantic preference features, language style features, visual style features, and vocal style features, and the LLM can learn a first digital human model based on these features. Here, LLM refers to a model with a large number of parameters and a complex structure; these models can process large amounts of data and generate high-quality prediction results. During the digital human live streaming process, these models can be used to generate scripts and dialogue (i.e., scripts) for the live performance, enriching the content of the digital human's speech. The specific learning process is not limited in this application.
[0088] It is understood that the first digital human model can also be obtained by learning one or more of these features (including structural features, semantic preference features, language style features, visual style features and voice style features). This application does not limit the combination of features used in the learning of the first digital human model.
[0089] In one possible implementation, the first digital human model is used to obtain first output information based on the first input information, such that the features of the first output information are similar to the text features, visual style features, and audio style features of the live video. The first output information is essentially the script used to control the digital human during the live stream. This allows the digital human, controlled according to the script, to resemble the host in the live video, making the digital human more lifelike, increasing its intelligence and flexibility, and enhancing its anthropomorphic nature. Furthermore, since the digital human is controlled by the script, its dependence on human control is reduced.
[0090] (VIII) The Second Digital Human Model
[0091] The second digital human model is learned based on the related second response text and second bullet screen text in the live video as training samples. The second digital human model is used to obtain second output information based on the second input information, such that the features of the second output information are similar to the bullet screen response preference features of the second response text. The second input information is the first bullet screen text during the digital human's live video broadcast, and the second output information is the first response text corresponding to the first bullet screen text. In other words, the bullet screen response preference features of the output first response text are similar to those of the second response text, so that the bullet screen responses of the digital human are similar to those of the broadcaster in the live video, improving the anthropomorphism of the digital human.
[0092] In one possible implementation, the training process of the second digital human model includes the following:
[0093] A correlation is established between the second response text and the second bullet screen text, which possess characteristics of bullet screen response preferences. For example, if the second bullet screen text is "Free shipping?", the second response text is "Yes". Then, based on the second response text and the second bullet screen text, a second digital human model is obtained through LLM learning. Optionally, the first response text output by the second digital human model is also associated with the visual style features and audio style features of the live video, so that the performance style of the digital human when displaying the first response text is similar to the performance style of the anchor responding to bullet screen comments in the live video, thereby improving the anthropomorphism of the digital human.
[0094] Based on the above description, Figure 2 is a flowchart of a video live streaming method based on digital human technology provided in this application. In this process, after configuring the first digital human model to generate a script based on the information to be displayed during the live stream, the first digital human is configured to perform a video live stream according to the script.
[0095] As shown in Figure 2, the method may include the following steps:
[0096] Step 210: Obtain the first information input by the user, which is the information to be displayed during the live broadcast.
[0097] In one possible implementation, the first information is of the same type as the information displayed in the live video used to train the first digital human model (including fruits, electronic products, furniture, etc.). For example, the first information and the information displayed in the live video are both about fruits, or both are about electronic products.
[0098] Step 220: Configure the first digital human model to generate a script based on the first information. The script includes at least one piece of text information, and each piece of text information is associated with performance style features.
[0099] In one possible implementation, the script includes at least one piece of text information, the text features of the at least one piece of text information (i.e., the text features of the script) are obtained based on structured text learning from the live video, and each piece of text information in the at least one piece of text information is associated with performance style features, which are obtained based on performance style learning from the live video.
[0100] Each segment of text information can be obtained by segmenting the script text according to semantics, words, and symbols (commas, periods, etc.). For example, if the script text includes "This apple is big and sweet," based on semantics, the following three segments of text information are obtained: "This apple," "Big," and "Sweet." The style features associated with each segment of text information can be understood as performance style labels when displaying that segment of text information. For example, the label corresponding to "This apple" indicates that the displayed image has apple features, the label corresponding to "Big" indicates that the digital human makes an action that reflects the size of the apple, and the label corresponding to "Sweet" indicates that the digital human makes an expression that reflects the sweetness of the apple.
[0101] In one possible implementation, the script's textual features include structural features and semantic preference features. The script's structural features are learned based on the structural features of structured text, and the script's semantic preference features are learned based on the semantic preference features of structured text. Based on the description of the first digital human model above, it can be understood that the script's structural features are similar to the structural features of structured text, and the script's semantic preference features are similar to the semantic preference features of structured text. Therefore, the script's structural features and semantic preference features are similar to those of structured text. For example, the product description structure and semantic preference of a product description performed in a live video broadcast based on this script are similar to the product description structure and semantic preference of a live video performance, thereby improving the anthropomorphism of the digital human's live video broadcast.
[0102] In one possible implementation, the text features also include the language style features of the script, which are learned based on the language style features of the structured text. Based on the description of the first digital human model above, it can be understood that the language style features of the script are similar to those of the structured text; that is, the language usage preferences of the script are similar to those of the structured text. Therefore, the language preferences used by the digital human controlled by the script are similar to those used by the broadcaster in the live video, further enhancing the anthropomorphism of the digital human live video broadcast.
[0103] In one possible implementation, the performance style features associated with each piece of text information include the screen style features of the script. These screen style features are learned based on the screen style features of the live video. The screen style features include one or more of the following: action features, facial expression features, or visual features. For a detailed description, refer to the introduction of screen style features in the aforementioned related technologies; it will not be repeated here. Based on the description of the first digital human model above, it can be understood that the screen style features of the script are similar to the screen style features of the live video. That is, controlling the digital human to conduct live video streaming according to the script includes one or more of the following: the digital human's actions, facial expressions, or the live video display. This ensures that during the live video streaming process, the digital human's performance style is similar to that of the host in the live video, and the changes in the visuals during the live stream closely resemble those in the live video, further enhancing the anthropomorphism of the digital human's live video streaming.
[0104] In one possible implementation, the performance style features associated with each segment of text information also include the language style features of the script. These language style features are learned based on the audio style features of the live video. For a detailed description, please refer to the introduction of audio style features in the aforementioned related technologies; details will not be repeated here. Based on the description of the first digital human model above, it can be understood that the language style features of the script are similar to the audio style features of the live video. In other words, the digital human's voice can be controlled according to the script to make its voice style similar to that of the anchor in the live video, thereby further enhancing the anthropomorphism of the digital human's live video broadcast.
[0105] It is understood that the textual features of a script can include one or more of the script's structural features, semantic preference features, and language style features; similarly, the performance style features of a script can include one or more of the script's visual style features and language style features.
[0106] Step 230: Configure the first digital human to conduct a live video broadcast according to the script.
[0107] In the above technical solution, during live video streaming, the first digital human can be controlled to display corresponding text information based on the performance style characteristics associated with the script. This increases the intelligence and flexibility of the digital human, making it more lifelike and enhancing the anthropomorphism of the live video stream. Furthermore, the digital human is entirely controlled by the script, thereby reducing reliance on manual control during live video streaming.
[0108] In one possible implementation, during a live video broadcast, the first digital human also acquires the first bullet comment text. Then, when it's determined that this first bullet comment text triggers a viewer interaction event, the second digital human model is configured to generate a first response text based on the first bullet comment text (i.e., inputting the first bullet comment text into the second digital human model and obtaining the second response text output by the second digital human model), and the first digital human is then controlled to display the second response text. For example, during a live video broadcast, if the first digital human acquires the first bullet comment text "Support or not, free shipping?", the second digital human model obtains the first response text as "Support", and then controls the digital human to reply with "Support" to respond to the bullet comment.
[0109] Optionally, the second response text can be associated with the language style features and / or sound style features of the script in step 220 above, so as to display the second response text according to the language style features and / or sound style features of the script.
[0110] Optionally, before controlling the digital human to display the first response text, it can be determined whether the same first comment text has been displayed multiple times. If so, the first response text is discarded; otherwise, the digital human is controlled to display the first response text. It should be understood that, for different products, any two first comment texts are considered to be different.
[0111] In the aforementioned technical solution, during live video streaming, the digital human responds to bullet comments in real time through a second digital human model, thereby increasing the intelligence and flexibility of the digital human, making it more lifelike and enhancing its anthropomorphic nature. Furthermore, responding to bullet comments requires no human intervention, thus reducing the digital human's reliance on manual control during live video streaming.
[0112] In one possible implementation, the aforementioned video live streaming method based on digital human technology is implemented using a client running on a terminal device and a server running on a server (such as a cloud management platform) as the execution entities. Referring to Figure 3, the method may include the following steps:
[0113] Step 310: The cloud management platform obtains the first information entered by the user in the client. The first information is the information to be displayed during the live broadcast.
[0114] Referring to step 210 above, it will not be repeated here.
[0115] Step 320: The cloud management platform configures the first digital human model to generate a script based on the first information. The script includes at least one piece of text information, and each piece of text information is associated with performance style characteristics.
[0116] Referring to step 220 above, it will not be repeated here.
[0117] Step 330: Configure the cloud management platform to enable the first digital human to conduct live video streaming according to the script.
[0118] Referring to step 230 above, it will not be repeated here.
[0119] In one possible implementation, a method for learning and obtaining a first digital human model is implemented using a client running on a terminal device and a server running on a server (such as a cloud management platform) as the execution entities. Referring to Figure 4, the method may include the following steps:
[0120] Step 410: The cloud management platform obtains the second information entered by the user in the client, and the second information is associated with the live video.
[0121] In this process, the second piece of information is the link and identifier of the live video. This second piece of information is used to determine the source file of the live video, which is used as a training sample for machine learning.
[0122] Step 420: The cloud management platform determines the relevant characteristics of the live video, which include one or more of the following: text features of the live video, visual style features, and audio style features.
[0123] In this process, the cloud management platform converts the anchor's speech in the training samples (i.e., live video) into first text, and then converts the first text into structured text. Next, the cloud management platform acquires video frames from the live video, including the anchor's actions, facial expressions, and the overall live feed. It also acquires the audio information (such as volume and tone) of the anchor describing the product in the live video.
[0124] Then, the cloud management platform extracts the structural features, semantic preference features, and language style features of the structured text, extracts the action features, facial expression features, and image features corresponding to these video frames, and extracts the sound style features of the sound information.
[0125] Step 430: The cloud management platform learns the first digital human model based on relevant features.
[0126] In this process, the cloud management platform learns the first digital human model based on these features (including structural features, semantic preference features, and language style features of structured text; action features, facial expression features, and image features of video frames; and voice style features). For the specific learning process, please refer to the description in (VII) above, which will not be repeated here. Based on this, after receiving the first information input by the user on the client, the cloud management platform executes the method flow shown in Figure 2 or Figure 3 above.
[0127] To better illustrate the above technical solution, Figure 5 provides a schematic diagram of the system architecture of a cloud management platform. This system can implement the methods shown in Figures 2 to 4. Referring to Figure 5, the system includes a recognition module 510, a learning module 520, and an imitation module 530. Optionally, the recognition module 510 and the learning module 520 can be applied to the server side, and the imitation module 530 can be applied to the client side; or, the recognition module 510, the learning module 520, and the imitation module 530 can all be applied to the server side, with the client displaying the live video stream.
[0128] The recognition module 510 is used to convert the anchor's voice in the live video of the live broadcast room (hereinafter referred to as the first live video for ease of description) into first text, and then convert the first text into structured text. For example, the anchor's voice includes: introductions to the first product and the second product, introducing the origin, taste, price and color of the first product, and introducing the size, performance and appearance of the second product. That is, the first text includes this information.
[0129] The recognition module 510 is also used to acquire video frames such as the anchor's actions, expressions, and live feed from the first live video. For example, when describing the taste of the first product, the anchor's action is a "thumbs up," the anchor's expression is a "smile," and the live feed is a "close-up of the anchor's upper body." Optionally, the recognition module 510 is also used to acquire the anchor's audio information from the first live video. For example, when describing the taste of the first product, the anchor's volume is loud and the pitch is high.
[0130] Learning module 520 is used to extract structural features, semantic preference features, and language style features from the structured text, and to extract action features, facial expression features, and image features corresponding to these video frames. Optionally, learning module 520 also extracts the voice style features of the audio information. Then, based on these features (including structural features, semantic preference features, and language style features of the structured text, action features, facial expression features, and image features of the video frames, and voice style features), a first digital human model is learned.
[0131] Optionally, the learning module 520 may send the first digital human model to the imitation module 530, or the imitation module 530 may actively request the first model from the learning module 520, or the imitation module 530 may obtain the first digital human model according to a preset storage address, wherein the preset storage address is used by the learning module 520 to store the first digital human model.
[0132] [Corrected according to Rule 91, 14.11.2024] The imitation module 530 is used to input first information into the first digital human model, obtain a script output by the first digital human model, and control the first digital human to conduct a live video broadcast according to the script (for ease of description, the video broadcast by the first digital human is referred to as the second live video below). For example, the first information includes product information corresponding to the first object, the second object, and the third object. That is, when the first digital human conducts a live video broadcast, the second live video introduces the first object, the second object, and the third object respectively. For the first object, the second live video introduces the place of origin, taste, price, and color respectively. For the second object and the third object, the second live video introduces the size, performance, and appearance respectively. Among them, the first product in the second live video is the same type as the first object, such as both being fruits; the second product is the same type as the second object and the third object, such as both being electronic devices. Moreover, both the first live video and the second live video use a poetic style for description during the live broadcast. Furthermore, when the first digital human describes the taste of the first object, it makes a "thumbs up" gesture and a "smiling" expression, and the screen shows a "close-up of the upper body of the first digital human". Furthermore, when the first digital person describes the taste of the first object, the volume is high and the voice is loud.
[0133] In other words, the product description structure and semantic preferences performed by the first digital human in the second live stream are similar to those of the product description performed by the anchor in the first live stream. The language style of the first digital human is also similar to that of the anchor. Furthermore, the actions and expressions of the first digital human are similar to those of the anchor, and the visuals displayed in the second live stream are similar to those in the first live stream. Additionally, the voice style of the first digital human when describing the product is similar to that of the anchor. Therefore, increasing the anthropomorphism of the digital human in live video streaming enhances its intelligence and flexibility. Moreover, the digital human is entirely controlled by the script, thus reducing its reliance on human control during live video streaming.
[0134] In one possible implementation, during the live video broadcast, the recognition module 510 also acquires the first bullet screen text in the broadcast and then determines whether the first bullet screen text triggers a viewer interaction event. For example, if the first bullet screen text is "Does it support free shipping?", it is determined that the first bullet screen text is a viewer interaction event targeting the product's selling points, that is, it is determined that the first bullet screen text has triggered a viewer interaction event.
[0135] [Corrected according to Rule 91, 14.11.2024] Then, the imitation module 530 inputs the first bullet screen text into the second digital human model, obtains the first response text output by the second digital human model, and controls the digital human to display the first response text. For example, during a live video broadcast, the cloud management platform obtains the first bullet screen text as "Support or not, free shipping?", obtains the first response text as "Support" through the second digital human model, and then controls the first digital human to reply with "Support" to respond to the bullet screen.
[0136] The video live streaming method based on digital human technology provided in this application has been described in detail above with reference to Figures 2 to 5. The video live streaming apparatus and computing device for performing the above method provided in this application will be described below with reference to Figures 4 to 7.
[0137] Figure 6 is a schematic diagram of a video live streaming device based on digital human technology provided in this application. This video live streaming device can be used to implement the aforementioned video live streaming method based on digital human technology, and therefore can also achieve the beneficial effects of the aforementioned method.
[0138] As shown in Figure 6, the video live streaming device 600 includes an acquisition module 610 and a processing module 620. The acquisition module 610 is used to acquire the first information input by the user, which is information to be displayed during the live stream. The processing module 620 is used to configure the first digital human model to generate a script based on the first information and to configure the first digital human to perform a live video stream based on the script. The script includes at least one piece of text information, and each piece of text information is associated with performance style features.
[0139] Both the acquisition module 610 and the processing module 620 can be implemented in software or in hardware. For example, the implementation of the acquisition module 610 will be described below. Similarly, the implementation of the processing module 620 can refer to the implementation of the acquisition module 610.
[0140] As an example of a software functional unit, module 610 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, module 610 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed within the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed within the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.
[0141] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0142] As an example of a hardware functional unit, the acquisition module 610 may include at least one computing device, such as a server. Alternatively, the acquisition module 610 may be implemented using a central processing unit (CPU), an application-specific integrated circuit (ASIC), or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a data processing unit (DPU), a neural network processing unit (NPU), a system-on-chip (SoC), an offload card, an accelerator card, or any combination thereof.
[0143] The multiple computing devices included in the acquisition module 610 can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the acquisition module 610 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the acquisition module 610 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, GALs, DPUs, NPUs, SoCs, offloading cards, and accelerator cards.
[0144] It should be noted that, in other embodiments, the acquisition module 610 and the processing module 620 can be used to execute any step in the video live streaming method based on digital human technology.
[0145] In one possible implementation, the acquisition module 610 is further configured to: acquire second information input by the user before configuring the first digital human model to generate a script based on the first information, wherein the second information is associated with the live video; the processing module 620 is further configured to: identify speech and keyframes in the live video based on the second information, convert the speech in the live video into first text, convert the first text into structured text based on the structure of the first text; extract text features of the structured text, visual style features of the keyframes, and vocal style features of the speech, and learn the first digital human model based on the text features, visual style features, and vocal style features. Wherein, the keyframes include one or more of the following: keyframes showing the anchor's actions, keyframes showing the anchor's expressions, and keyframes showing changes in the visuals due to camera changes; the text features of the structured text include one or more of the following: structural features, semantic preference features, and language style features; the visual style features of the keyframes include one or more of the following: action features, expression features, and visual features;
[0146] In one possible implementation, the text features of the script are learned based on the structural features of the structured text; or, the text features of the script are learned based on the structural features and semantic preference features of the structured text; or, the text features of the script are learned based on the structural features, semantic preference features, and language style features of the structured text.
[0147] In one possible implementation, the performance style features of the script include visual style features, which are learned based on the action features of the keyframes; or, the visual style features of the script are learned based on the action and facial expression features of the keyframes; or, the visual style features of the script are learned based on the action, facial expression, and visual features of the keyframes; the processing module 620 is specifically used for one or more of the following: controlling the actions of the first digital human during live video streaming, controlling the facial expressions of the first digital human during live video streaming, and controlling the display screen of the first digital human during live video streaming.
[0148] In one possible implementation, the performance style features of the script include voice style features, which are learned based on the voice style features of the speech; the processing module 620 is specifically used to control the voice of the first digital human during live video streaming.
[0149] In one possible implementation, the infrastructure further stores a second digital human model of the user, which is associated with a first digital human. The acquisition module 610 is further configured to: acquire a first bullet screen text during a live video broadcast by the first digital human, wherein the first bullet screen text triggers a viewer interaction event; the processing module 620 is further configured to: configure the second digital human model to generate a first response text based on the first bullet screen text, and configure the first digital human to conduct a live video broadcast based on the first response text, wherein the first response text is associated with bullet screen response preference features in the live video.
[0150] Figure 7 is a schematic diagram of a computing device provided in this application. This computing device can be used to implement the above-mentioned video live streaming method based on digital human technology, and therefore can also achieve the beneficial effects of the above method.
[0151] As shown in Figure 7, the computing device 100 includes a processor 104 and a communication interface 108. The processor 104 and the communication interface 108 are coupled to each other. It is understood that the communication interface 108 can be a transceiver or an input / output interface. Optionally, the computing device 100 may also include a memory 106 for storing instructions executed by the processor 104, or storing input data required by the processor 104 to execute instructions, or storing data generated after the processor 104 executes instructions.
[0152] Optionally, the processor 104, communication interface 108, and memory 106 are interconnected via bus 102. Bus 102 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into address bus, data bus, control bus, etc.
[0153] It is understood that the processor 104 in the embodiments of this application may include any one or more of the following computing devices: central processing unit (CPU), graphics processing unit (GPU), microprocessor (MP) or digital signal processor (DSP), ASIC, FPGA, CPLD, NPU, SoC, offload card, accelerator card, etc.
[0154] Memory 106 may include volatile memory, such as random access memory (RAM). Processor 104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD). Furthermore, memory 106 may also be implemented using storage class memory (SCM), phase change memory (PCM), or other types of storage media.
[0155] It is worth noting that the same type of storage medium can be configured in the same computing device to realize the function of memory 106, or two or more types of storage media can be configured to realize the function of memory 106. This application does not limit this.
[0156] The memory 106 stores executable program code, which the processor 104 executes to implement the functions of the aforementioned acquisition module 610 and processing module 620, thereby realizing the video live streaming method based on digital human technology. That is, the memory 106 stores instructions for executing the video live streaming method based on digital human technology.
[0157] Alternatively, the memory 106 stores executable code, which the processor 104 executes to implement the functions of the aforementioned video live streaming device 600, thereby realizing the video live streaming method based on digital human technology. That is, the memory 106 stores instructions for executing the video live streaming method based on digital human technology.
[0158] [Revised according to Rule 91, 14.11.2024] Communication interface 108 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between computing device 100 and other devices or communication networks.
[0159] In one possible implementation, the computing device 100 may also include a chip system, which includes a processor and a power supply circuit. The power supply circuit supplies power to the processor, which executes the operation steps corresponding to the video live streaming method based on digital human technology. For simplicity, further details are omitted here. The processor can be implemented using a GPU, or it can be implemented using computing devices or AI chips such as a DPU, NPU, XPU, SoC, offload card, or accelerator card.
[0160] In one possible implementation, the computing device 100 may include multiple types of processors 104, i.e., the computing device 100 is a heterogeneous device. For example, the computing device 100 includes a CPU and a GPU, and at least one of the processors 104 can execute the operation steps corresponding to the video live streaming method based on digital human technology. For the sake of brevity, further details will not be elaborated here.
[0161] As one possible implementation, this application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0162] As one possible implementation, as shown in Figure 8, this application also provides a computing device cluster including at least one computing device 100. The memory 106 of one or more computing devices 100 in the computing device cluster may store the same instructions for executing a video live-streaming method based on digital human technology.
[0163] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the video live streaming method based on digital human technology. In other words, a combination of one or more computing devices 100 can jointly execute the instructions for the video live streaming method based on digital human technology.
[0164] It should be noted that the memory 106 in different computing devices 100 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the video live streaming device 600. That is, the instructions stored in the memory 106 of different computing devices 100 can implement the functions of one or more modules in the acquisition module 610 and processing module 620.
[0165] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 9 illustrates one possible implementation. As shown in Figure 9, two computing devices 100A and 100B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 106 in computing device 100A stores instructions for executing the functions of the acquisition module 610. Simultaneously, the memory 106 in computing device 100B stores instructions for executing the functions of the processing module 620.
[0166] The connection method between the computing device clusters shown in Figure 9 can be considered in light of the fact that the video live streaming method based on digital human technology provided in this application requires a large amount of storage of the questions in the questionnaire and the options associated with each question. Therefore, it is considered that the functions implemented by the processing module 620 are handed over to the computing device 100B for execution.
[0167] It should be understood that the functions of computing device 100A shown in Figure 9 can also be performed by multiple computing devices 100. Similarly, the functions of computing device 100B can also be performed by multiple computing devices 100.
[0168] This application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similar to the connection method of the computing device clusters described in Figures 7 and 8. The difference is that the memory 106 in one or more computing devices 100 in this computing device cluster can store the same instructions for executing the video live streaming method based on digital human technology.
[0169] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing the video live streaming method based on digital human technology. In other words, a combination of one or more computing devices 100 can jointly execute instructions for executing the video live streaming method based on digital human technology.
[0170] It should be noted that the memory 106 in different computing devices 100 in the computing device cluster can store different instructions to implement the functions of the video live streaming device 600.
[0171] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform a video live-streaming method based on digital human technology.
[0172] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium capable of being stored by a computing device, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct a computing device to perform a video live-streaming method based on digital human technology.
[0173] It is understood that the various numerical designations used in the embodiments of this application are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application. The order of the process numbers described above does not imply the order of execution; the execution order of each process should be determined by its function and internal logic.
Claims
1. A video live streaming method based on digital human technology, characterized in that, The method is applied to a cloud management platform, which manages infrastructure. The infrastructure stores a first digital human model of a user, and the first digital human model is associated with a first digital human. The method includes: Obtain the first information input by the user, which is information to be displayed during the live broadcast; The first digital human model is configured to generate a script based on the first information. The script includes at least one piece of text information, and each piece of text information is associated with performance style features. Configure the first digital human to conduct a live video broadcast according to the script.
2. The method according to claim 1, characterized in that, Before configuring the first digital human model to generate a script based on the first information, the method further includes: Obtain the second information input by the user, which is associated with the live video. Based on the second information, the audio and keyframes in the live video are identified. The keyframes include one or more of the following: keyframes showing the anchor's actions, keyframes showing the anchor's expressions, and keyframes that cause changes in the image due to changes in the camera. The audio in the live video is converted into first text, and the first text is converted into structured text according to its structure. Extract the text features of the structured text, the visual style features of the keyframes, and the vocal style features of the speech. The text features of the structured text include one or more of the following: structural features, semantic preference features, and language style features. The visual style features of the keyframes include one or more of the following: action features, facial expression features, and visual features. The first digital human model is learned based on the text features, the visual style features, and the sound style features.
3. The method according to claim 1 or 2, characterized in that, The text features of the script are learned based on the structural features of the structured text; or, the text features of the script are learned based on the structural features and semantic preference features of the structured text; or, the text features of the script are learned based on the structural features, semantic preference features, and language style features of the structured text.
4. The method according to claim 1 or 2, characterized in that, The performance style features of the script include visual style features, which are learned based on the motion features of the keyframes; or, the visual style features of the script are learned based on the motion features and facial expression features of the keyframes; or, the visual style features of the script are learned based on the motion features, facial expression features, and visual features of the keyframes. The configuration of the first digital human to conduct live video streaming according to the script includes one or more of the following: the actions of the first digital human during live video streaming, the facial expressions of the first digital human during live video streaming, and the display screen of the first digital human during live video streaming.
5. The method according to claim 1 or 2, characterized in that, The performance style features of the script include voice style features, which are learned based on the voice style features of the speech. The configuration of the first digital human to conduct live video streaming according to the script includes the sound of the first digital human during the live video streaming.
6. The method according to any one of claims 1-5, characterized in that, The infrastructure also stores a second digital human model of the user, which is associated with a first digital human. The method further includes: Get the first bullet screen text when the first digital human is live streaming video, and the first bullet screen text triggers a viewer interaction event; The second digital human model is configured to generate a first response text based on the first bullet screen text, and the first response text is associated with the bullet screen response preference features in the live video. Configure the first digital human to perform a live video broadcast based on the first response text.
7. A video live streaming device based on digital human technology, characterized in that, The device is used to manage infrastructure in which a first digital human model of a user is stored, the first digital human model being associated with a first digital human, the device comprising: The acquisition module is used to acquire the first information input by the user, the first information being information to be displayed during the live broadcast; The processing module is configured to configure the first digital human model to generate a script based on the first information, wherein the script includes at least one piece of text information, and each piece of text information is associated with performance style features; Configure the first digital human to conduct a live video broadcast according to the script.
8. The apparatus according to claim 7, characterized in that, The acquisition module is also used for: Before configuring the first digital human model to generate a script based on the first information, the second information input by the user is obtained, and the second information is associated with the live video. The processing module is also used for: Based on the second information, the audio and keyframes in the live video are identified. The keyframes include one or more of the following: keyframes showing the anchor's actions, keyframes showing the anchor's expressions, and keyframes that cause changes in the image due to changes in the camera. The audio in the live video is converted into first text, and the first text is converted into structured text according to its structure. Extract the text features of the structured text, the visual style features of the keyframes, and the vocal style features of the speech. The text features of the structured text include one or more of the following: structural features, semantic preference features, and language style features. The visual style features of the keyframes include one or more of the following: action features, facial expression features, and visual features. The first digital human model is learned based on the text features, the visual style features, and the sound style features.
9. The apparatus according to claim 7 or 8, characterized in that, The text features of the script are learned based on the structural features of the structured text; or, the text features of the script are learned based on the structural features and semantic preference features of the structured text; or, the text features of the script are learned based on the structural features, semantic preference features, and language style features of the structured text.
10. The apparatus according to claim 7 or 8, characterized in that, The performance style features of the script include visual style features, which are learned based on the motion features of the keyframes; or, the visual style features of the script are learned based on the motion features and facial expression features of the keyframes; or, the visual style features of the script are learned based on the motion features, facial expression features, and visual features of the keyframes. The processing module is specifically used for one or more of the following: Control the actions of the first digital human during live video streaming, control the facial expressions of the first digital human during live video streaming, and control the display screen of the first digital human during live video streaming.
11. The apparatus according to claim 7 or 8, characterized in that, The performance style features of the script include voice style features, which are learned based on the voice style features of the speech. The processing module is specifically used for: Control the sound of the first digital human during live video streaming.
12. The apparatus according to any one of claims 7-11, characterized in that, The infrastructure also stores a second digital human model of the user, which is associated with a first digital human. The acquisition module is further used for: Get the first bullet screen text when the first digital human is live streaming video, and the first bullet screen text triggers a viewer interaction event; The processing module is also used for: The second digital human model is configured to generate a first response text based on the first bullet screen text, and the first response text is associated with the bullet screen response preference features in the live video. Configure the first digital human to perform a live video broadcast based on the first response text.
13. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium includes a program that, when run on the device, causes the device to perform the method as described in any one of claims 1-6.
14. A computing device cluster, characterized in that, The computing device cluster includes at least one computing device, the at least one computing device including at least one chip and a memory, the at least one chip being used to read and execute program instructions stored in the memory to implement the method as described in any one of claims 1-6.
15. A program product, characterized in that, When the program product is run on the device, the device causes the device to perform the method as described in any one of claims 1-6.