Live streaming method and apparatus, electronic device, and storage medium

WO2026200836A1PCT designated stage Publication Date: 2026-10-01BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/085366
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2026-03-23
Publication Date
2026-10-01

Smart Images

  • Figure CN2026085366_01102026_PF_FP_ABST
    Figure CN2026085366_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of data processing, particularly to the fields of large models, artificial intelligence, and human-machine interaction, and provides a live streaming method and apparatus, an electronic device, and a storage medium. The specific implementation comprises: sending comment information of a user to a server, the comment information being used for acquiring a corresponding target video material and corresponding digital human narration information; receiving the target video material and the digital human narration information sent by the server; on the basis of a shot type of a current live streaming image of a digital human, determining a target element from the target video material and the digital human narration information; and updating the live streaming image of the digital human on the basis of the target element.
Need to check novelty before this filing date? Find Prior Art

Description

Live streaming methods, devices, electronic equipment and storage media

[0001] Cross-reference to related applications

[0002] This disclosure is based on and claims priority to Chinese Patent Application No. 202510371165.8, filed on March 26, 2025, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure relates to the field of data processing technology, specifically to the fields of large models, artificial intelligence and human-computer interaction, and particularly to a live streaming method, device, electronic device and storage medium. Background Technology

[0004] Currently, when live-streaming e-commerce features live-streaming with real people, the product's appearance, interior, and specifications are comprehensively presented based on different audience questions in the comments section. However, digital human live-streaming rooms respond to audience questions by intelligently generating spoken responses or using instant messaging (IM) bullet comments. This approach sometimes fails to address the audience's actual concerns, resulting in poor interactive effects. Summary of the Invention

[0005] This disclosure provides a live streaming method, apparatus, electronic device, and storage medium.

[0006] According to one aspect of this disclosure, a live streaming method is provided, the method comprising:

[0007] Send user comment information to the server; the comment information is used to obtain the corresponding target video material and digital voiceover information.

[0008] Receive the target video material and the digital voice broadcast information sent by the server;

[0009] Based on the camera type of the current live stream of the digital human, the target element is determined from the target video material and the digital human's voice broadcast information;

[0010] Update the live stream of the digital human based on the target element.

[0011] According to another aspect of this disclosure, an alternative live streaming method is provided, the method comprising:

[0012] Retrieve user comment information sent by the client;

[0013] Obtain the target video material and digital voiceover information corresponding to the comment information;

[0014] Based on the camera type of the current live stream of the digital human, at least one target element from the target video material and the digital human's voice broadcast information is sent to the client.

[0015] According to a third aspect of this disclosure, a live streaming device is provided, comprising:

[0016] The sending module is used to send user comment information to the server, and the comment information is used to obtain the corresponding target video material and digital voice broadcast information;

[0017] The receiving module is used to receive the target video material and the digital voice broadcast information sent by the server;

[0018] The acquisition module is used to determine the target element from the target video material and the digital human's voice broadcast information based on the camera type of the current live broadcast screen of the digital human;

[0019] The update module is used to update the live stream of the digital human based on the target element.

[0020] According to the fourth aspect of this disclosure, another live streaming device is provided, comprising:

[0021] The first acquisition module is used to acquire user comment information sent by the client.

[0022] The second acquisition module is used to acquire the target video material and digital voice broadcast information corresponding to the comment information;

[0023] The sending module is used to send at least one target element from the target video material and the digital human's voice broadcast information to the client based on the camera type of the current live broadcast screen of the digital human.

[0024] According to a fifth aspect of this disclosure, an electronic device is provided, comprising:

[0025] At least one processor; and

[0026] A memory communicatively connected to the at least one processor; wherein,

[0027] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described in the first or second aspect embodiment.

[0028] According to a sixth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method described in the first or second aspect embodiments.

[0029] According to a seventh aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps of the method described in the embodiments of the first or second aspect.

[0030] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0031] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0032] Figure 1 is a schematic diagram of a live streaming method provided in an embodiment of this disclosure;

[0033] Figure 2 is a schematic diagram of another live streaming method provided in an embodiment of this disclosure;

[0034] Figure 3 is a schematic diagram of another live streaming method provided in an embodiment of this disclosure;

[0035] Figure 4 is a schematic diagram of another live streaming method provided by an embodiment of this disclosure;

[0036] Figure 5 is a schematic diagram of another live streaming method provided by an embodiment of this disclosure;

[0037] Figure 6 is an interactive schematic diagram of a live streaming method provided in an embodiment of this disclosure;

[0038] Figure 7 is an interactive schematic diagram of another live streaming method provided in an embodiment of this disclosure;

[0039] Figure 8 is a logical flowchart of a live streaming method provided in an embodiment of this disclosure;

[0040] Figure 9 is a schematic diagram of the structure of a live streaming device provided in an embodiment of this disclosure;

[0041] Figure 10 is a schematic diagram of another live streaming device provided in an embodiment of this disclosure;

[0042] Figure 11 shows a schematic block diagram of an electronic device used to implement embodiments of the present disclosure. Detailed Implementation

[0043] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0044] Data processing is the collection, storage, retrieval, processing, transformation, and transmission of data. Its basic purpose is to extract and deduce data that is valuable and meaningful to certain specific people from a large amount of data that may be disorganized and difficult to understand.

[0045] A large model is a machine learning model with a large number of parameters and a complex computational structure. It is usually built from deep neural networks and has billions or even hundreds of billions of parameters. Its purpose is to improve the model's expressive power and predictive performance, and to be able to handle more complex tasks and data.

[0046] Artificial intelligence (AI) is a new technological science that studies and develops theories, methods, technologies, and application systems to simulate, extend, and expand human intelligence. It is an important component of the discipline of intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. Artificial intelligence can simulate the information processes of human consciousness and thought.

[0047] Human-computer interaction (HCI) is a discipline that studies the interaction between a system and a user. The system can be various kinds of machines, as well as computerized systems and software. The HCI usually refers to the part that is visible to the user. The user communicates with the system and performs operations through the HCI.

[0048] The following describes the live streaming method applied to the client side. This live streaming method can be executed by a live streaming device, which can be a client or configured within the client; this disclosure does not limit this.

[0049] Figure 1 is a schematic diagram of a live streaming method provided in an embodiment of this disclosure. As shown in Figure 1, the method includes the following steps:

[0050] S101, send the user's comment information to the server.

[0051] The comment information consists of text messages posted by users in the comment area while watching the live stream. The comment area can be a bullet screen or a comment box, etc. The comment information includes the user's questions and the content the user wants to know. It can also be used to obtain corresponding target video materials and digital voice broadcast information. The video materials are videos that provide detailed introductions to product prices, appearance, or functions. The target video materials are videos that explain the user's questions. The digital voice broadcast information is the content that the virtual digital human will verbally express, such as answers to user questions, which are introduced through verbal broadcasting.

[0052] In some implementations, the server can store video footage and voiceover information. After sending user comments to the server, the server can match the corresponding target video footage and digital voiceover information based on the comments.

[0053] S102 receives the target video material and digital voice broadcast information sent by the server.

[0054] It is understandable that the target video material and digital voiceover information are the video material and digital voiceover information corresponding to the user's comment information, and the video material and voiceover information include answers to the questions in the user's comment information.

[0055] S103, Based on the camera type of the current live stream of the digital human, determine the target element from the target video material and the digital human's voiceover information.

[0056] Optionally, the shot type can be either a regular shot or a split-screen shot. A regular shot can be understood as a page shot on the live streaming page, which can display more comprehensive information, such as complete video playback and video display. Therefore, when the shot type is a regular shot, the target elements are the target video material and the digital voiceover information, and user questions can be answered based on the linkage of video and voiceover playback. A split-screen shot can be understood as a split-screen shot or a partial slice shot, which does not include the digital human and cannot display the digital human-related video material. Therefore, when the shot type is a split-screen shot, the target element is the digital voiceover information, that is, only the text reply information to user questions is displayed.

[0057] Understandably, the server can send both the target video footage and digital speech information corresponding to the comment information to the client. The client then determines whether the target element is the target video footage and digital speech information or just digital speech information based on the actual shot type of the current live stream, and displays the target element accordingly.

[0058] S104, Update the live stream of the digital human based on the target element.

[0059] In some implementations, the target element can be displayed in the live stream of the digital human. This means displaying the target video material and the digital human's voiceover information, or just the digital human's voiceover information, in the live stream. This allows for a more comprehensive answer to the user through the target video material and the digital human's voiceover information, enabling more intelligent human-computer interaction and improving the user experience.

[0060] In some embodiments, the server can also determine the target element from the target video material and the digital human's voiceover information based on the camera type of the current live stream of the digital human, and send the target element to the client. That is, the client can directly receive the target element sent by the server. Correspondingly, S102 and S103 can be replaced by receiving the target element sent by the server.

[0061] In this embodiment, by sending user comment information to the server, the server matches the target video material and digital voiceover information that match the user's question in the comment information. After receiving the target video material and digital voiceover information sent by the server, the target element is determined according to the shot type of the current live broadcast of the digital human. The live broadcast of the digital human is then updated based on the target element. The target video material and digital voiceover information are displayed in conjunction with each other for different shot types, or only the digital voiceover information is displayed, to ensure the display effect of the target element in the live broadcast. This provides a richer and more comprehensive answer to the user's question, achieves more intelligent human-computer interaction, and improves the user experience.

[0062] Figure 2 is a schematic diagram of another live streaming method provided by an embodiment of this disclosure. As shown in Figure 2, the method includes the following steps:

[0063] S201, Send the user's comment information to the server.

[0064] In this embodiment of the disclosure, the method for implementing step S201 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0065] S202, receive the target video material and digital voice broadcast information sent by the server.

[0066] In this embodiment of the disclosure, the method for implementing step S202 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0067] S203, Based on the camera type of the current live stream of the digital human, determine the target element from the target video footage and the digital human's voiceover information.

[0068] In some implementations, in response to the shot type being the first type, the target video footage and digital voiceover information are identified as target elements; where the first type refers to ordinary shots, which are capable of comprehensive video playback and display, thus the target video footage and digital voiceover information are identified as target elements.

[0069] In some implementations, in response to the shot type being the second type, the digital human voiceover information is determined as the target element; where the second type refers to a storyboard shot, which does not include a digital human and cannot play video related to the digital human, so only the digital human voiceover information is used as the target element to adapt to the shot type for smooth target element display.

[0070] In this embodiment of the disclosure, the method for implementing step S203 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0071] S204, in response to the target elements including target video footage and digital human voice broadcast information, interrupt the current live broadcast of the digital human and record the interruption position.

[0072] It's understandable that when the target element includes target video footage and digital voiceover information, displaying the target footage is equivalent to playing the target video footage and digital voiceover information. Since video playback occupies the client's live stream screen, the current live stream of the digital human can be interrupted during video display, and the interruption point can be recorded. Therefore, after the target video footage and digital voiceover information are displayed, the user can resume the live stream from the interrupted point, allowing for responses and answers to user comments without interrupting the live stream viewing, thus improving the user experience.

[0073] S205: Based on the target video material and digital human voice broadcast information, generate and display the digital human's interstitial screen.

[0074] In some implementations, the spoken video of the digital human can be rendered based on the spoken information of the digital human; alternatively, the spoken video of the digital human speaking or expressing spoken information can be generated by graphics rendering techniques such as animation rendering or real-time rendering.

[0075] The target video material is rendered based on its rendering parameters. Optionally, the rendering parameters may include parameters such as screen size, contrast, and brightness. The target video material is then adjusted for brightness and contrast, and its size is cropped.

[0076] The rendered target video footage can be associated with the spoken footage to generate digital human interstitial footage. This means matching the spoken content of the digital human with the display of the video footage. For example, when the digital human mentions a specific feature, that feature can appear simultaneously or shortly after in the video footage, thus achieving the effect of demonstrating and explaining the product features when displaying interstitial footage.

[0077] In some implementations, the digital human's voiceover can be displayed, and during this display, target video footage can be played at a designated location on the live stream interface to update the digital human's live stream feed. This designated location on the live stream interface can be a corner, the center, or any other preset position. By playing target video footage at a designated location on the live stream interface during the voiceover display, users can obtain answers to relevant questions based on the digital human's voiceover, and simultaneously see video content corresponding to the voiceover content. This allows for a more intuitive understanding of the voiceover information, and the video footage display is more flexible, providing users with a richer and more personalized viewing experience.

[0078] Optionally, the display progress of the voice-over video can be monitored, and the playback of the target video material can be stopped when the display of the voice-over video reaches its end time, so as to maintain the consistency between the display of the voice-over video and the target video material. Furthermore, the video can return to the interrupted position and continue to display the subsequent live video of the digital human from the interrupted position, so as to ensure the integrity and continuity of the live broadcast for users and improve the user's viewing experience.

[0079] In some implementations, when the target element only includes digital human voice broadcast information, a text-based instant message is generated based on the digital human voice broadcast information, and the instant message is replied to in the interactive area on the live broadcast interface. The instant message is used to update the live broadcast screen of the digital human. It can be understood that the digital human voice broadcast information is used to answer user questions. When the shot type does not meet the display requirements of digital human voice broadcast and video material, the digital human voice broadcast information is generated into a text-based instant message and replied to in the interactive area, such as the comment section of the live broadcast screen. This allows the text-based instant message to be displayed to the user, and the live broadcast screen of the digital human is updated according to the content of the instant message. Combining the live broadcast screen and the text-based instant message helps users understand product-related information more promptly and provides users with a richer interactive experience.

[0080] In this embodiment, by sending user comment information to the server, the server matches the user's question in the comment information to obtain target video material and digital voiceover information. After receiving the target video material and digital voiceover information from the server, the target element is determined from the target video material and digital voiceover information according to the current camera type of the digital human's live stream. When the target element is the target video material and digital voiceover information, the live stream is interrupted and the interruption position is recorded. After the target element is displayed, the subsequent live stream of the digital human continues from the interruption position, ensuring the integrity and continuity of the user's live stream viewing. When the target element is displayed, the voiceover information of the digital human is rendered to obtain the voiceover screen. The target video material is rendered and associated with the voiceover screen to generate a digital human interstitial screen for display. This ensures that the product features are well displayed and explained when the interstitial screen is displayed. The target video material can also be played at a designated position on the live broadcast interface, improving the flexibility of the target video material display. When the target material only includes digital voiceover information, a text-based instant message is generated based on the digital voiceover information and displayed in the interactive area to help users understand product-related information in a timely manner and provide users with a richer interactive experience.

[0081] The following describes a live streaming method applied to the server side. This live streaming method can be executed by a live streaming device, which can be a server or configured within a server; this disclosure does not limit this. Figure 3 is a schematic diagram of another live streaming method provided by an embodiment of this disclosure. As shown in Figure 3, the method includes the following steps:

[0082] S301, retrieve user comment information sent by the client.

[0083] Understandably, the client can be the application corresponding to the live streaming page. Users can post comments in the comment area or bullet screen area of ​​the live streaming page. The comments can include the user's questions and the content the user wants to know. The client can obtain the user's comment information and send it to the server. The server obtains the user's comment information sent by the client and analyzes it.

[0084] S302, Obtain the target video material and digital voiceover information corresponding to the comment information.

[0085] Optionally, keyword extraction can be performed based on comment information to determine the user's main questions and the main content they want to know in the comment information. Matching can then be performed from a pre-built media library based on the keywords. The media library can include video materials and response materials. This allows the video materials in the media library that match the keywords to be identified as target video materials, and the response materials in the media library that match the keywords to be identified as digital voice broadcast information.

[0086] S303, based on the camera type of the current live stream of the digital human, send at least one target element from the target video footage and the digital human's voiceover information to the client.

[0087] In some implementations, the camera type of the current live stream of the digital human can include a first type and a second type. The first type can be a normal camera shot, and the second type is a split-screen camera shot. When the camera type is the first type, which is a normal camera shot, the target elements are the target video footage and the digital human's voiceover information. When the camera type is the second type, which is a split-screen camera shot, the target elements are the digital human's voiceover information.

[0088] Optionally, the client can determine the camera type of the current live stream of the digital human and send the camera type of the current live stream to the server. When the server receives the camera type, it determines the target element based on the camera type and sends the target element to the client for display, so as to ensure that the client can display the target element smoothly and reduce redundant resource transmission consumption.

[0089] In some implementations, the server can send all the target video footage and digital voiceover information to the client, which then determines the target elements based on the shot type and displays them. This makes the display and presentation of video footage and digital voiceover information more flexible and improves the user experience.

[0090] In this embodiment, the system receives user comment information sent by the client, matches it with the material library based on the comment information, and obtains the target video material and digital voiceover information corresponding to the comment information. The target video material and digital voiceover information are used to provide a more comprehensive answer to the user's questions, thereby improving the user's interactive experience. Based on the camera type of the current live broadcast of the digital human, the system determines the target element from the target video material and digital voiceover information, and sends the target element to the client to ensure that the client can display the target element smoothly and reduce redundant resource transmission consumption.

[0091] Figure 4 is a schematic diagram of another live streaming method provided by an embodiment of this disclosure. As shown in Figure 4, the method includes the following steps:

[0092] S401, retrieve user comment information sent by the client.

[0093] In this embodiment of the disclosure, the method for implementing step S401 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0094] S402, retrieve the first question-and-answer pair associated with the comment information from the pre-built question-and-answer database.

[0095] Optionally, the question-and-answer database can be pre-configured. In some embodiments, the process of building the question-and-answer database includes: obtaining configuration operations and determining question-and-answer pairs based on the configuration operations; obtaining video footage and associating the video footage with the question-and-answer pairs to generate the question-and-answer database.

[0096] In some implementations, configuration operations are used to configure question-answer pairs in the question-answer database, such as the generation method and content of the question-answer pairs. Based on the configuration operations, questions and corresponding answers are determined, thus forming question-answer pairs. Optionally, video materials can be pre-generated videos in various forms, such as explanations and demonstrations. The video materials are associated with the question-answer pairs, that is, question-answer pairs with the same content are associated with the video materials, thereby generating a question-answer database. This allows users to determine the associated video materials when retrieving text information from the question-answer database, enhancing the richness and intuitiveness of the question-answer pairs.

[0097] In some implementations, during the construction of the question-and-answer database, material rendering configuration operations can also be obtained, and based on the material rendering configuration operations, the image rendering parameters of the video material can be obtained. The material rendering configuration operations are used to perform operations such as editing or rendering of the video material. The image rendering parameters of the video material are determined according to the material rendering configuration operations. The image rendering parameters can be parameters such as size, contrast and brightness, to ensure that the video material maintains high image quality after rendering and improves the display effect of the video material.

[0098] Furthermore, the first question-and-answer pair associated with the comment information is retrieved from the question-and-answer database. For example, keywords in the comment information are extracted, and keyword matching is performed between the keywords and the question-and-answer pairs in the question-and-answer database to obtain the first question-and-answer pair that meets the similarity requirements.

[0099] In some implementations, the target user identifier corresponding to the comment information can be obtained, and the target user type of the user can be determined based on the target user identifier; based on the target user type, the target acquisition process of the target video material and digital voiceover information can be determined; and the target video material and digital voiceover information can be obtained according to the target acquisition process.

[0100] Optionally, in this embodiment of the disclosure, the target user identifier can be a unique identifier such as a user account. The target user type can include authorized users and unauthorized users. The target user identifier can be determined by a user whitelist, where authorized users are those on the whitelist, thus determining the target user type. For ordinary unauthorized users, the target acquisition process can be to match comments from a question-and-answer database. When a question-and-answer pair is matched, the target video material and digital voiceover information are determined based on the pair. When no question-and-answer pair is matched, a response is generated using a large model, and question-and-answer pair matching is performed again based on the response information to determine the target video material and digital voiceover information. For special authorized users, such as third-party users, the target video material and digital voiceover information can be obtained through their unique platform material library, which provides greater flexibility in acquiring the target video material and digital voiceover information.

[0101] S403, determine the first video material associated with the first question-and-answer pair as the target video material.

[0102] S404, determine the first response information in the first question-and-answer pair as digital population broadcast information.

[0103] In some implementations, in response to the failure to find the first question-answer pair in the question-answer database, a second response corresponding to the comment information is generated based on a large model. Keyword matching is then performed between the second response and the question-answer pairs in the question-answer database. In other words, if no first question-answer pair meeting the similarity requirement is found in the question-answer data, the user's comment information is input into a pre-trained large model, which generates the second response corresponding to the comment information. Furthermore, keyword matching is performed between the second response and the question-answer pairs in the question-answer database to determine whether a second question-answer pair corresponding to the second response exists. This improves the accuracy of acquiring target video material and digital voiceover information and avoids failure to acquire target video material and digital voiceover information due to keyword matching errors.

[0104] In response to matching the second question-and-answer pair with the second response information, determine the second video material associated with the second question-and-answer pair as the target video material; obtain the second response information in the second question-and-answer pair as digital voice broadcast information.

[0105] In response to the failure to find a second question-and-answer pair corresponding to the second response information, the second response information is determined and used as digital voice broadcast information.

[0106] In some embodiments, if no second question-and-answer pair corresponding to the second response information is matched, a default video material may be determined as the target video material, or the target video material may be obtained by other means, or the target video material may be determined to be empty. This disclosure does not limit this.

[0107] S405, based on the camera type of the current live stream of the digital human, sends at least one target element from the target video footage and the digital human's voiceover information to the client.

[0108] In this embodiment of the disclosure, the method for implementing step S405 can be implemented in any of the various embodiments of the disclosure, without limitation or further description.

[0109] In this embodiment, user comment information sent by the client is received, and a first question-and-answer pair associated with the comment information is queried from a pre-built question-and-answer database based on the comment information. During the construction of the question-and-answer database, question-and-answer pairs are determined through configuration operations and associated with video materials. Furthermore, image rendering parameters can be determined through material rendering configuration operations to ensure that the video materials maintain high image quality after rendering, thereby improving the display effect of the video materials. When the first question-and-answer pair is found, the first video material corresponding to the first question-and-answer pair is used as the target video material, and the first reply information in the first question-and-answer pair is used as a digital population broadcast message. When no first question-and-answer pair is found, a second response is generated using a large model. Based on this second response, keyword matching is performed in the question-and-answer database to determine if a matching second question-and-answer pair exists. If a matching second question-and-answer pair is found, the associated second video material is designated as the target video material. The second response in the second question-and-answer pair is used as digital voiceover information. If no matching second question-and-answer pair is found, the second response is used as digital voiceover information. This process ensures accurate target video material and digital voiceover information corresponding to the comment information, improving the accuracy of user responses and enhancing the user interaction experience.

[0110] Figure 5 is a schematic diagram of another live streaming method provided by an embodiment of this disclosure. As shown in Figure 5, the method includes the following steps:

[0111] S501, retrieve user comment information sent by the client.

[0112] In this embodiment of the disclosure, the method for implementing step S501 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0113] S502, retrieve the first question-and-answer pair associated with the comment information from the pre-built question-and-answer database.

[0114] In this embodiment of the disclosure, the method for implementing step S502 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0115] S503, determine the first video material associated with the first question-and-answer pair as the target video material.

[0116] In this embodiment of the disclosure, the method for implementing step S503 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0117] S504, determine the first response information in the first question-and-answer pair as digital population broadcast information.

[0118] In this embodiment of the disclosure, the method for implementing step S504 can be implemented in any of the various embodiments of this disclosure, and no limitation is made here, nor will it be described in detail.

[0119] S505, in response to the lens type being Type 1, identifies the target video footage and digital voiceover information as target elements.

[0120] Optionally, the first type can be a regular shot, which includes a digital human capable of digital voiceovers and video playback and display, thus identifying the target video footage and digital voiceover information as target elements.

[0121] S506, in response to the lens type being type two, determines the digital voice broadcast information as the target element.

[0122] Optionally, the second type can be a storyboard shot, which does not have a digital human and cannot perform digital voiceover and video playback, so digital voiceover information is determined as the target element.

[0123] S507 sends the target element to the client.

[0124] It is understood that in the embodiments of this disclosure, when the server receives the client's comment information, it can determine the client's camera type and determine the target element according to the camera type. When providing feedback to the client based on the comment information, the server only sends the target element to the client for the client to display. In other embodiments, the server can send both digital voice information and target video material to the client, and the client can determine the target element according to the camera type and display the target element on the client.

[0125] In this embodiment, user comment information sent by the client is received, and a first question-and-answer pair associated with the comment information is queried from a pre-built question-and-answer database based on the comment information. During the construction of the question-and-answer database, question-and-answer pairs are determined through configuration operations and associated with video materials. Furthermore, rendering parameters can be determined through material rendering configuration operations to ensure high image quality of the video materials after rendering, thereby improving the display effect of the video materials. When a first question-and-answer pair is found, the first video material corresponding to the first question-and-answer pair is used as the target video material, and the first reply information in the first question-and-answer pair is used as digital voiceover information. When a first question-and-answer pair is not found, a second reply information is generated through a large model, and the second reply information is used as the basis for the query. Keyword matching is performed in the question-and-answer database to determine if a matching second question-and-answer pair exists. If a second question-and-answer pair is found, the associated second video material is designated as the target video material, and the second response information in the second question-and-answer pair is used as digital voiceover information. If no second question-and-answer pair is found, the second response information is used as digital voiceover information. This process obtains accurate target video material and digital voiceover information corresponding to the comment information, improving the accuracy of user responses and enhancing the user interaction experience. Furthermore, the target element is determined based on whether the shot type is Type 1 or Type 2, and the target element is sent to the client to ensure that the client can successfully display the target element and reduce redundant resource transmission consumption.

[0126] The following describes the interaction process between the client and the server.

[0127] Figure 6 is an interactive schematic diagram of a live streaming method provided in an embodiment of this disclosure. As shown in Figure 6, the live streaming method includes:

[0128] S601, the client sends the user's comment information to the server.

[0129] In this embodiment of the disclosure, the method for implementing step S601 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0130] S602, the server retrieves the first question-and-answer pair associated with the comment information from the pre-built question-and-answer database.

[0131] In this embodiment of the disclosure, the method for implementing step S602 can be implemented in any of the various embodiments of this disclosure, and no limitation is made here, nor will it be described in detail.

[0132] S603, the server determines the first video material associated with the first question-and-answer pair as the target video material.

[0133] In this embodiment of the disclosure, the method for implementing step S603 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0134] S604, the server determines the first response information in the first question-and-answer pair as digital population broadcast information.

[0135] In this embodiment of the disclosure, the method for implementing step S604 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0136] S605, in response to the lens type being Type 1, the server determines the target video material and digital voiceover information as target elements.

[0137] In this embodiment of the disclosure, the method for implementing step S605 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0138] S606, in response to the lens type being type two, the server determines the digital population broadcast information as the target element.

[0139] In this embodiment of the disclosure, the method for implementing step S606 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0140] S607, the server sends the target element to the client.

[0141] In this embodiment of the disclosure, the method for implementing step S607 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0142] S608, in response to the target elements including target video footage and digital human broadcast information, the client interrupts the current live broadcast of the digital human and records the interruption position.

[0143] In this embodiment of the disclosure, the method for implementing step S608 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0144] S609: The client generates and displays the digital human's interstitial footage based on the target video material and digital human voiceover information.

[0145] In this embodiment of the disclosure, the method for implementing step S609 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0146] In this embodiment, the client obtains the user's comment information and sends it to the server. The server queries a pre-built question-and-answer database based on the comment information to find the first question-and-answer pair associated with the comment information. When the first question-and-answer pair is found, the first video material corresponding to the first question-and-answer pair is used as the target video material, and the first reply information in the first question-and-answer pair is used as the digital human's voiceover information. This obtains the accurate target video material and digital human's voiceover information corresponding to the comment information, improving the accuracy of the response to the user and enhancing the user's interactive experience. Furthermore, the target element is determined based on whether the shot type is a first type or a second type, and the target element is sent to the client. In response to the target element including the target video material and the digital human's voiceover information, the client interrupts the current live broadcast of the digital human and records the interruption position. Based on the target video material and the digital human's voiceover information, an interstitial shot of the digital human is generated and displayed, helping users to understand product-related information in a timely manner and providing users with a richer interactive experience.

[0147] Figure 7 is an interactive schematic diagram of another live streaming method provided in this embodiment of the present disclosure. As shown in Figure 7, the live streaming method includes:

[0148] S701, the client sends the user's comment information to the server.

[0149] In this embodiment of the disclosure, the method for implementing step S701 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0150] S702, the server retrieves the first question-and-answer pair associated with the comment information from a pre-built question-and-answer database.

[0151] In this embodiment of the disclosure, the method for implementing step S702 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0152] S703, the server determines the first video material associated with the first question-and-answer pair as the target video material.

[0153] In this embodiment of the disclosure, the method for implementing step S703 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0154] S704, the server determines the first response information in the first question-and-answer pair as digital population broadcast information.

[0155] In this embodiment of the disclosure, the method for implementing step S704 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0156] S705: The client receives the target video material and digital voice broadcast information sent by the server.

[0157] In this embodiment of the disclosure, the method for implementing step S705 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0158] S706: The client determines the target element from the target video footage and the digital human's voiceover information based on the camera type of the current live stream of the digital human.

[0159] In this embodiment of the disclosure, the method for implementing step S706 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0160] S707, in response to the target elements including target video footage and digital human voice broadcast information, the client interrupts the current live broadcast of the digital human and records the interruption position.

[0161] In this embodiment of the disclosure, the method for implementing step S707 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0162] S708: The client generates and displays the digital human's interstitial footage based on the target video material and digital human voiceover information.

[0163] In this embodiment of the disclosure, the method for implementing step S708 can be implemented in any of the various embodiments of the disclosure, and no limitation is made here, nor will it be described in detail.

[0164] In this embodiment, the client obtains user comment information and sends it to the server. The server queries a pre-built question-and-answer database based on the comment information to find the first question-and-answer pair associated with the comment information. When the first question-and-answer pair is found, the first video material corresponding to the first question-and-answer pair is used as the target video material, and the first reply information in the first question-and-answer pair is used as the digital voiceover information. This obtains accurate target video material and digital voiceover information corresponding to the comment information, improving the accuracy of the response to the user and enhancing the user's interactive experience. Furthermore, the target video material and digital voiceover information are sent to the client. The client determines the target element based on whether the shot type is a first type or a second type. If the target element includes the target video material and digital voiceover information, the current live broadcast of the digital human is interrupted, and the interruption position is recorded. Based on the target video material and digital voiceover information, an interstitial shot of the digital human is generated and displayed to help users understand product-related information in a timely manner and provide users with a richer interactive experience.

[0165] Figure 8 is a logical flowchart of a live streaming method provided in an embodiment of this disclosure. As shown in Figure 8, after a user posts a comment on the live stream interface, the system matches question-and-answer pairs from the question-and-answer database based on the comment information. If a matching pair is found, the system renders the associated video footage and voiceover, determines the current shot type, and identifies the target elements in the video footage and voiceover based on the shot type. If the shot type is a standard shot (i.e., a normal shot), the live stream is interrupted, and a response is given via voiceover linked to the video footage. If the shot type is a cut shot (i.e., a split-shot shot), a text response to the voiceover content is given. If no matching pair is found, a response (i.e., the second response information in the aforementioned embodiment) is generated using a large model. Keyword matching is performed in the question-and-answer database based on the response information. If a matching pair is found, the system renders the associated video footage and voiceover, determines the current shot type, and identifies the target elements in the video footage and voiceover based on the shot type for corresponding display. If no matching pair is found, the response information is displayed on the live stream screen, enhancing the interaction between the user and the digital human anchor and improving the user experience.

[0166] Figure 9 is a schematic diagram of a live streaming device provided in an embodiment of this disclosure. In some embodiments, the live streaming device can be a client or can be configured in a client; this disclosure does not limit this. As shown in Figure 9, the live streaming device 900 includes:

[0167] The sending module 901 is used to send the user's comment information to the server. The comment information is used to obtain the corresponding target video material and digital voice broadcast information.

[0168] The receiving module 902 is used to receive the target video material and digital voice broadcast information sent by the server;

[0169] The acquisition module 903 is used to determine the target element from the target video material and the digital human's voice broadcast information based on the camera type of the current live broadcast screen of the digital human;

[0170] Update module 904 is used to update the live stream of the digital human based on the target element.

[0171] In some implementations, module 903 is used for:

[0172] In response to the shot type being Type 1, the target video footage and digital voiceover information are identified as target elements;

[0173] In response to the lens type being type two, the digital voice broadcast information is identified as the target element.

[0174] In some implementations, module 904 is updated for:

[0175] In response to target elements including target video footage and digital human voice broadcast information, the current live stream of the digital human is interrupted, and the interruption position is recorded;

[0176] Based on the target video footage and digital human voiceover information, generate and display the digital human's interstitial footage.

[0177] In some implementations, module 904 is updated for:

[0178] Render the spoken audio of the digital human based on the digital human's spoken audio information;

[0179] Based on the rendering parameters of the target video material, the target video material is rendered, and then associated with the voiceover screen to generate the digital human's interstitial screen.

[0180] In some implementations, module 904 is updated for:

[0181] Display the digital human's voiceover, and during the display of the voiceover, play target video material at a designated location on the live stream interface to update the digital human's live stream.

[0182] In some implementations, update module 904 is also used for:

[0183] Monitor the display progress of the voice-over video and stop playing the target video material when the voice-over video reaches its end time;

[0184] Return to the point where the video was interrupted, and resume the live stream from that point.

[0185] In some implementations, module 904 is updated for:

[0186] In response to target elements including digital human voice broadcast information, instant messages in text form are generated based on the digital human voice broadcast information, and instant messages are replied to in the interactive area on the live broadcast interface. The instant messages are used to update the live broadcast screen of the digital human.

[0187] In this embodiment, by sending user comment information to the server, the server matches the user's question in the comment information to obtain target video material and digital voiceover information. After receiving the target video material and digital voiceover information from the server, the target element is determined from the target video material and digital voiceover information according to the current camera type of the digital human's live stream. When the target element is the target video material and digital voiceover information, the live stream is interrupted and the interruption position is recorded. After the target element is displayed, the subsequent live stream of the digital human continues from the interruption position, ensuring the integrity and continuity of the user's live stream viewing. When the target element is displayed, the voiceover information of the digital human is rendered to obtain the voiceover screen. The target video material is rendered and associated with the voiceover screen to generate a digital human interstitial screen for display. This ensures that the product features are well displayed and explained when the interstitial screen is displayed. The target video material can also be played at a designated position on the live broadcast interface, improving the flexibility of the target video material display. When the target material only includes digital voiceover information, a text-based instant message is generated based on the digital voiceover information and displayed in the interactive area to help users understand product-related information in a timely manner and provide users with a richer interactive experience.

[0188] Figure 10 is a schematic diagram of another live streaming device provided in an embodiment of this disclosure. In some embodiments, the live streaming device can be a server or can be configured in a server; this disclosure does not limit this. As shown in Figure 10, the live streaming device 1000 includes:

[0189] The first acquisition module 1001 is used to acquire user comment information sent by the client.

[0190] The second acquisition module 1002 is used to acquire the target video material and digital voice broadcast information corresponding to the comment information;

[0191] The sending module 1003 is used to send at least one target element from the target video material and the digital human's voice broadcast information to the client based on the camera type of the current live broadcast screen of the digital human.

[0192] In some implementations, the second acquisition module 1002 is used for:

[0193] From a pre-built question-and-answer database, query the first question-and-answer pair associated with the comment information;

[0194] Identify the first video clip associated with the first question-and-answer pair as the target video clip;

[0195] The first response in the first question-and-answer pair is identified as the digital population broadcast information.

[0196] In some implementations, the second acquisition module 1002 is also used for:

[0197] In response to the failure to find the first question-answer pair in the question-answer database, a second reply is generated based on the large model to correspond to the comment information.

[0198] Keyword matching is performed between the second response information and the question-and-answer pairs in the question-and-answer database;

[0199] In response to matching the second question-and-answer pair with the second response information, determine the second video material associated with the second question-and-answer pair as the target video material;

[0200] Obtain the second response information from the second question-and-answer pair as digital population broadcast information.

[0201] In some implementations, the second acquisition module 1002 is also used for:

[0202] In response to the failure to find a second question-and-answer pair corresponding to the second response information, the second response information is determined and used as digital voice broadcast information.

[0203] In some implementations, the sending module 1003 is used for:

[0204] In response to the shot type being Type 1, the target video footage and digital voiceover information are identified as target elements;

[0205] In response to the lens type being type two, the digital voice broadcast information is identified as the target element.

[0206] In some implementations, the second acquisition module 1002 is used for:

[0207] Obtain configuration operations and determine question-answer pairs based on the configuration operations;

[0208] Acquire video footage and associate it with question-and-answer pairs to generate a question-and-answer database.

[0209] In some implementations, the second acquisition module 1002 is also used for:

[0210] Get the material rendering configuration operation, and based on the material rendering configuration operation, get the video material's image rendering parameters.

[0211] In some implementations, the second acquisition module 1002 is used for:

[0212] Obtain the target user identifier corresponding to the comment information, and determine the target user type of the user based on the target user identifier;

[0213] Based on the target user type, determine the target acquisition process for target video materials and digital voice broadcast information;

[0214] Following the target acquisition process, acquire the target video footage and digital voiceover information.

[0215] In this embodiment, user comment information sent by the client is received, and a first question-and-answer pair associated with the comment information is queried from a pre-built question-and-answer database based on the comment information. During the construction of the question-and-answer database, question-and-answer pairs are determined through configuration operations and associated with video materials. Furthermore, rendering parameters can be determined through material rendering configuration operations to ensure high image quality of the video materials after rendering, thereby improving the display effect of the video materials. When a first question-and-answer pair is found, the first video material corresponding to the first question-and-answer pair is used as the target video material, and the first reply information in the first question-and-answer pair is used as digital voiceover information. When a first question-and-answer pair is not found, a second reply information is generated through a large model, and the second reply information is used as the basis for the query. Keyword matching is performed in the question-and-answer database to determine if a matching second question-and-answer pair exists. If a second question-and-answer pair is found, the associated second video material is designated as the target video material, and the second response information in the second question-and-answer pair is used as digital voiceover information. If no second question-and-answer pair is found, the second response information is used as digital voiceover information. This process obtains accurate target video material and digital voiceover information corresponding to the comment information, improving the accuracy of user responses and enhancing the user interaction experience. Furthermore, the target element is determined based on whether the shot type is Type 1 or Type 2, and the target element is sent to the client to ensure that the client can successfully display the target element and reduce redundant resource transmission consumption.

[0216] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0217] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0218] Figure 11 shows a schematic block diagram of an electronic device used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0219] As shown in Figure 11, device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1102 or a computer program loaded from storage unit 1108 into random access memory (RAM) 1103. The RAM 1103 may also store various programs and data required for the operation of device 1100. The computing unit 1101, ROM 1102, and RAM 1103 are interconnected via bus 1104. Input / output (I / O) interface 1105 is also connected to bus 1104.

[0220] Multiple components in device 1100 are connected to I / O interface 1105, including: input unit 1106, such as keyboard, mouse, etc.; output unit 1107, such as various types of monitors, speakers, etc.; storage unit 1108, such as disk, optical disk, etc.; and communication unit 1109, such as network card, modem, wireless transceiver, etc. Communication unit 1109 allows device 1100 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0221] The computing unit 1101 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above, such as the live streaming method. For example, in some embodiments, the live streaming method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1108. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1100 via ROM 1102 and / or communication unit 1109. When the computer program is loaded into RAM 1103 and executed by the computing unit 1101, one or more steps of the live streaming method described above may be performed. Alternatively, in other embodiments, the computing unit 1101 may be configured to perform the live streaming method by any other suitable means (e.g., by means of firmware).

[0222] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0223] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0224] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0225] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0226] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0227] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0228] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0229] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

[0230] All embodiments disclosed herein can be executed individually or in combination with other embodiments, and are all considered to be within the scope of protection claimed by this disclosure.

Claims

1. A live streaming method, wherein, The method includes: Send user comment information to the server; the comment information is used to obtain the corresponding target video material and digital voiceover information. Receive the target video material and the digital voice broadcast information sent by the server; Based on the camera type of the current live stream of the digital human, the target element is determined from the target video material and the digital human's voice broadcast information; Update the live stream of the digital human based on the target element.

2. The method according to claim 1, wherein, The step of determining the target element from the target video material and the digital human's voiceover information based on the camera type of the current live stream of the digital human includes: In response to the lens type being a first type, the target video material and the digital voiceover information are determined to be the target elements; In response to the lens type being the second type, the digital voice broadcast information is determined to be the target element.

3. The method according to claim 1 or 2, wherein, The process of updating the live stream of the digital human based on the target element includes: In response to the target element including the target video material and the digital human's broadcast information, the current live broadcast of the digital human is interrupted, and the interruption position is recorded; Based on the target video material and the digital human's voiceover information, the digital human's interstitial screen is generated and displayed.

4. The method according to claim 3, wherein, The step of generating the interstitial footage of the digital human based on the target video material and the digital human's voiceover information includes: Based on the digital human's speech information, render the speech animation of the digital human; Based on the rendering parameters of the target video material, the target video material is rendered, and after rendering, it is associated with the voice-over screen to generate the interstitial screen of the digital human.

5. The method according to claim 4, wherein, The display of the inserted video includes: The system displays the narration of the digital human and, during the display of the narration, plays the target video material at a designated location on the live streaming interface to update the live streaming screen of the digital human.

6. The method according to claim 5, wherein, The method further includes: The display progress of the spoken video is monitored, and playback of the target video material is stopped when the spoken video reaches its end time. Return to the point where the video was interrupted, and resume displaying the subsequent live video feed of the digital human from that point.

7. The method according to any one of claims 1-6, wherein, The process of updating the live stream of the digital human based on the target element includes: In response to the target element including the digital human's broadcast information, an instant message in text form is generated based on the digital human's broadcast information, and the instant message is replied to in the interactive area on the live broadcast interface. The instant message is used to update the live broadcast screen of the digital human.

8. A live streaming method, wherein, The method includes: Retrieve user comment information sent by the client; Obtain the target video material and digital voiceover information corresponding to the comment information; Based on the camera type of the current live stream of the digital human, at least one target element from the target video material and the digital human's voice broadcast information is sent to the client.

9. The method according to claim 8, wherein, The step of obtaining the target video material and digital voiceover information corresponding to the comment information includes: From a pre-built question-and-answer database, query the first question-and-answer pair associated with the comment information; The first video material associated with the first question-and-answer pair is identified as the target video material; The first response information in the first question-and-answer pair is determined as the digital population broadcast information.

10. The method according to claim 9, wherein, The method further includes: In response to the failure to find the first question-answer pair in the question-answer database, a second reply information corresponding to the comment information is generated based on the large model; Keyword matching is performed between the second response information and the question-and-answer pairs in the question-and-answer database; In response to matching a second question-and-answer pair corresponding to the second response information, a second video material associated with the second question-and-answer pair is determined as the target video material; Obtain the second response information from the second question-and-answer pair, and use it as the digital population broadcast information.

11. The method according to claim 10, wherein, The method further includes: In response to the failure to find a second question-and-answer pair corresponding to the second response information, the second response information is determined as the digital speech information.

12. The method according to any one of claims 8-11, wherein, Based on the camera type of the current live stream footage of the digital human, at least one target element is determined from the target video material and the digital human's voiceover information, including: In response to the lens type being a first type, the target video material and the digital voiceover information are determined to be the target elements; In response to the lens type being the second type, the digital voice broadcast information is determined to be the target element.

13. The method according to any one of claims 9-11, wherein, The process of constructing the question-and-answer database includes: Obtain configuration operations and determine question-answer pairs based on the configuration operations; Acquire video materials and associate the video materials with the question-and-answer pairs to generate the question-and-answer database.

14. The method according to claim 13, wherein, The method further includes: Obtain the material rendering configuration operation, and based on the material rendering configuration operation, obtain the image rendering parameters of the video material.

15. The method according to any one of claims 8-14, wherein, The step of obtaining the target video material and digital voiceover information corresponding to the comment information includes: Obtain the target user identifier corresponding to the comment information, and determine the target user type of the user based on the target user identifier; Based on the target user type, determine the target acquisition process for the target video material and digital voice broadcast information; According to the target acquisition process, the target video material and digital voice broadcast information are acquired.

16. A live streaming device, wherein, include: The sending module is used to send user comment information to the server, and the comment information is used to obtain the corresponding target video material and digital voice broadcast information; The receiving module is used to receive the target video material and the digital voice broadcast information sent by the server; The acquisition module is used to determine the target element from the target video material and the digital human's voice broadcast information based on the camera type of the current live broadcast screen of the digital human; The update module is used to update the live stream of the digital human based on the target element.

17. A live streaming device, wherein, include: The first acquisition module is used to acquire user comment information sent by the client. The second acquisition module is used to acquire the target video material and digital voice broadcast information corresponding to the comment information; The sending module is used to send at least one target element from the target video material and the digital human's voice broadcast information to the client based on the camera type of the current live broadcast screen of the digital human.

18. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7 or 8-15.

19. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-7 or 8-15.

20. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-7 or 8-15.