Live broadcast method and device, electronic equipment and storage medium

By sending user comment information to the server in the live broadcast room, obtaining target video materials and digital population broadcast information, and updating the live broadcast screen according to the camera type, the problem that the live broadcast room of real people and digital people cannot solve the audience problems in a positive manner, and the interactive effect is improved.

CN120201256APending Publication Date: 2025-06-24BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510371165.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When dealing with audience questions in the live broadcast and digital live broadcast rooms in the e-commerce industry, there are problems that cannot be solved positively in the audience's actual concern, and the interaction effect is poor.

Method used

By sending user comment information to the server, the target video material and digital population broadcast information are obtained, and the target element is determined from this information based on the lens type of the current live broadcast screen of the digital person, and the live broadcast screen of the digital person is updated.

Benefits of technology

It has achieved a more comprehensive answer to the questions asked by the audience, improved the user's interactive experience, and enhanced the interactive effect of the live broadcast room.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201256A_ABST
    Figure CN120201256A_ABST
Patent Text Reader

Abstract

The invention provides a live broadcast method and device, electronic equipment and a storage medium, and relates to the technical field of data processing, in particular to the fields of large models, artificial intelligence and human-computer interaction. According to the specific implementation scheme, comment information of a user is sent to a server side, wherein the comment information is used for obtaining corresponding target video materials and digital population broadcast information; receiving a target video material and digital population broadcast information sent by the server; determining a target element from the target video material and the digital human broadcast information according to the lens type of the current live broadcast picture of the digital human; and updating the live picture of the digital person based on the target element.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technologies, specifically to the fields of large models, artificial intelligence, and human-computer interaction, and particularly to a live broadcast method, device, electronic device, and storage medium. Background Art

[0002] Currently, when introducing products in real-person live broadcasts in the e-commerce industry, information such as the appearance, interior, and data of the product is comprehensively introduced according to different questions from the audience in the comment area. In a digital human live broadcast room, the intelligent acquisition of the audience's comments and questions is through voiceovers or bullet screen replies in the instant messaging (IM) area, and there is a situation where the actual concerns of the audience cannot be positively addressed, resulting in poor interaction effects. Summary of the Invention

[0003] The present disclosure provides a live broadcast method, device, electronic device, and storage medium.

[0004] According to one aspect of the present disclosure, a live broadcast method is provided, the method including:

[0005] Sending the comment information of the user to the server, where the comment information is used to obtain the corresponding target video material and digital human voiceover information;

[0006] Receiving the target video material and the digital human voiceover information sent by the server;

[0007] Determining the target element from the target video material and the digital human voiceover information according to the lens type of the current live broadcast screen of the digital human;

[0008] Updating the live broadcast screen of the digital human based on the target element.

[0009] According to another aspect of the present disclosure, another live broadcast method is provided, the method including:

[0010] Obtaining the comment information of the user sent by the client;

[0011] Obtaining the target video material and the digital human voiceover information corresponding to the comment information;

[0012] Sending at least one target element of the target video material and the digital human voiceover information to the client according to the lens type of the current live broadcast screen of the digital human.

[0013] According to a third aspect of the present disclosure, a live broadcast device is provided, including:

[0014] A sending module, configured to send the comment information of the user to the server, where the comment information is used to obtain the corresponding target video material and digital human voiceover information;

[0015] A receiving module, configured to receive the target video material and the digital human voiceover information sent by the server;

[0016] An obtaining module, configured to determine the target element from the target video material and the digital human voiceover information according to the lens type of the current live broadcast picture of the digital human;

[0017] An updating module, configured to update the live broadcast picture of the digital human based on the target element.

[0018] According to a fourth aspect of the present disclosure, another live broadcast device is provided, including:

[0019] A first obtaining module, configured to obtain the comment information of the user sent by the client;

[0020] A second obtaining module, configured to obtain the target video material and the digital human voiceover information corresponding to the comment information;

[0021] A sending module, configured to send at least one target element of the target video material and the digital human voiceover information to the client according to the lens type of the current live broadcast picture of the digital human.

[0022] According to a fifth aspect of the present disclosure, an electronic device is provided, including:

[0023] At least one processor; and

[0024] A memory communicatively connected to the at least one processor; wherein,

[0025] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the embodiments of the first aspect or the second aspect.

[0026] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method described in the embodiments of the first aspect or the second aspect.

[0027] According to a seventh aspect of the present disclosure, a computer program product is provided, including a computer program, and the computer program implements the steps of the method described in the embodiments of the first aspect or the second aspect when executed by a processor.

[0028] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0029] The accompanying drawings are used to better understand the present solution and do not constitute a limitation to the present disclosure. Among them:

[0030] Figure 1 is a schematic diagram of a live broadcast method provided by an embodiment of the present disclosure;

[0031] Figure 2 is a schematic diagram of another live broadcast method provided by an embodiment of the present disclosure;

[0032] Figure 3 is a schematic diagram of another live broadcast method provided by an embodiment of the present disclosure;

[0033] Figure 4 is a schematic diagram of another live broadcast method provided by an embodiment of the present disclosure;

[0034] Figure 5 is a schematic diagram of another live broadcast method provided by an embodiment of the present disclosure;

[0035] Figure 6 is an interaction schematic diagram of a live broadcast method provided by an embodiment of the present disclosure;

[0036] Figure 7 is an interaction schematic diagram of another live broadcast method provided by an embodiment of the present disclosure;

[0037] Figure 8 is a logic flowchart of a live broadcast method provided by an embodiment of the present application;

[0038] Figure 9 is a schematic structural diagram of a live broadcast device provided by an embodiment of the present disclosure;

[0039] Figure 10 is a schematic structural diagram of another live broadcast device provided by an embodiment of the present disclosure;

[0040] Figure 11 shows a schematic block diagram of an electronic device used to implement an embodiment of the present disclosure. Detailed Embodiments

[0041] The following makes an explanation of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.

[0042] Data Processing refers to the collection, storage, retrieval, processing, transformation, and transmission of data. The basic purpose is to extract and derive valuable and meaningful data for certain specific people from a large amount of data that may be chaotic and difficult to understand.

[0043] A Large Model is a machine learning model with a large number of parameters and a complex computational structure. It is usually constructed by a deep neural network and has billions or even hundreds of billions of parameters. The purpose is to improve the model's expressive ability and prediction performance and be able to handle more complex tasks and data.

[0044] Artificial Intelligence (AI) is a new technical science that studies, develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. It is an important part of the intelligent discipline. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial Intelligence can simulate the information process of human consciousness and thinking.

[0045] Human–Computer Interaction is a discipline that studies the interaction relationship between a system and a user. The system can be various machines or computerized systems and software. The human–computer interaction interface usually refers to the part visible to the user. The user communicates with the system and operates through the human–computer interaction interface.

[0046] Figure 1 It is a schematic diagram of a live broadcast method provided by an embodiment of the present disclosure. As Figure 1 shown, the method includes the following steps:

[0047] S101, sending the user's comment information to the server.

[0048] Among them, the comment information is the text information sent by the user in the comment area when watching the live broadcast. The comment area can be a bullet screen or a comment box, etc. The comment information includes the user's questions and the content the user wants to know. It can also be used to obtain the corresponding target video material and digital human voice-over information. The video material is a video with a detailed introduction of the product price, appearance, or function, etc. The target video material is the video explaining the user's questions. The digital human voice-over information is the content that the virtual digital human will orally express, such as the answer information to the user's questions, and it is introduced in a voice-over manner.

[0049] In some implementations, content such as video materials and digital human voiceover information can be stored in the server. After sending the user's comment information to the server, the corresponding target video material and digital human voiceover information can be matched in the server according to the comment information.

[0050] S102. Receive the target video material and digital human voiceover information sent by the server.

[0051] It can be understood that the target video material and digital human voiceover information are the video material and digital human voiceover information corresponding to the user's comment information, and the video material and voiceover information include the answers to the questions in the user's comment information.

[0052] S103. Determine the target element from the target video material and digital human voiceover information according to the lens type of the current live broadcast screen of the digital human.

[0053] Optionally, the lens type can be a normal lens or a split-screen lens. The normal lens can be understood as the page lens of the live broadcast page, which can display more comprehensive information, such as complete video playback and video display. Therefore, when the lens type is a normal lens, the target element is the target video material and digital human voiceover information, and the user's question can be answered based on the linked playback of the video and the voiceover; the split-screen lens can be understood as a split-screen lens or a partial slice lens, which does not include the digital human and cannot display the video material related to the digital human. Therefore, when the lens type is a split-screen lens, the target element is the digital human voiceover information, that is, only the text reply information to the user's question is displayed.

[0054] It can be understood that the server can send both the target video material and digital human voiceover information corresponding to the comment information to the client, and the client determines whether the target element is the target video material and digital human voiceover information or only the digital human voiceover information according to the actual current live broadcast lens type, so as to display the target element.

[0055] S104. Update the live broadcast screen of the digital human based on the target element.

[0056] In some implementations, the target element can be displayed in the live broadcast screen of the digital human, that is, the video material and digital human voiceover information or only the digital human voiceover information are displayed in the live broadcast screen, so as to answer the user more comprehensively through the video material and digital human voiceover information, realize more intelligent human-computer interaction, and improve the user experience.

[0057] In this embodiment, by sending the user's comment information to the server, the server matches to obtain the target video material and the digital human voiceover information that match the user's question in the user's comment information. After receiving the target video material and the digital human voiceover information sent by the server, the target element is determined according to the lens type of the current live broadcast picture of the digital human, so as to update the live broadcast picture of the digital human based on the target element, and perform the linked display of the target video material and the digital human voiceover information for different lens types, or only display the digital human voiceover information, ensuring the display effect of the target element on the live broadcast picture, answering the user's question more richly and comprehensively, realizing more intelligent human-computer interaction, and improving the user experience.

[0058] Figure 2 It is a schematic diagram of another live broadcast method provided by an embodiment of the present disclosure. As Figure 2 shown, the method includes the following steps:

[0059] S201, send the user's comment information to the server.

[0060] In the embodiment of the present disclosure, the implementation method of step S201 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and no further description is given.

[0061] S202, receive the target video material and the digital human voiceover information sent by the server.

[0062] In the embodiment of the present disclosure, the implementation method of step S202 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and no further description is given.

[0063] S203, determine the target element from the target video material and the digital human voiceover information according to the lens type of the current live broadcast picture of the digital human.

[0064] In some implementations, in response to the lens type being the first type, determine the target video material and the digital human voiceover information as the target element; where the first type refers to a normal lens that can perform comprehensive video playback and display, so the target video material and the digital human voiceover information are determined as the target element.

[0065] In some implementations, in response to the lens type being the second type, determine the digital human voiceover information as the target element; where the second type refers to a split-screen lens that does not include the digital human and cannot perform digital human-related video playback, so only the digital human voiceover information is used as the target element to adapt to the lens type for smooth display of the target element.

[0066] In the embodiments of the present disclosure, the implementation method of step S203 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made herein and will not be elaborated further.

[0067] S204. In response to the target element including the target video material and the digital human voiceover information, interrupt the current live broadcast screen of the digital human and record the position where the screen is interrupted.

[0068] It can be understood that when the target element includes the target video material and the digital human voiceover information, the display of the target material is to play the target video material and the digital human voiceover information. The playback of the video occupies the live broadcast screen of the client. Therefore, when the video is displayed, it is necessary to interrupt the current live broadcast screen of the digital human and record the position where the screen is interrupted. After the display of the target video material and the digital human voiceover information is completed, the user can continue to play the live broadcast screen from the position where the screen is interrupted, reply to and answer the user's comments without disturbing the user's viewing of the live broadcast, and improve the user experience.

[0069] S205. Generate an inserted broadcast screen of the digital human based on the target video material and the digital human voiceover information, and display the inserted broadcast screen.

[0070] In some implementations, the voiceover screen of the digital human can be rendered based on the digital human voiceover information; optionally, graphic rendering technologies such as animation rendering or real-time rendering can be used to generate the voiceover screen of the digital human speaking or expressing the voiceover information.

[0071] Render the target video material based on the screen rendering parameters of the target video material; optionally, the screen rendering parameters can include parameter information such as screen size, contrast, and brightness. Adjust the brightness and contrast of the target video material and perform operations such as size cropping through the screen rendering parameters, and associate with the voiceover screen after rendering to generate the inserted broadcast screen of the digital human, that is, match the voiceover content of the digital human with the display of the video material. For example, when the digital human mentions a certain specific feature, the feature can appear in the video material simultaneously or later, so as to achieve the effect of displaying and explaining the product features when the inserted broadcast screen is displayed.

[0072] In some implementations, the mouthpiece broadcast screen of the digital human can be displayed, and during the display of the mouthpiece broadcast screen, the target video material can be played at a specified position on the live broadcast interface to update the live broadcast screen of the digital human. The specified position on the live broadcast interface can be a corner, the middle, or any other preset position. That is, during the display of the mouthpiece broadcast screen, the target video material is played at the specified position on the live broadcast interface. Users can obtain answers to relevant questions based on the mouthpiece broadcast screen of the digital human, and at the same time, they can also see the video content corresponding to the mouthpiece broadcast content, understand the mouthpiece broadcast information more intuitively, the video material display is more flexible, and a richer and more personalized viewing experience is provided for users.

[0073] Optionally, the display progress of the mouthpiece broadcast screen can also be monitored. In response to the mouthpiece broadcast screen reaching the end time, the playback of the target video material is stopped to maintain the display consistency of the mouthpiece broadcast screen and the target video material; the position where the screen is interrupted is returned, and the subsequent live broadcast screen of the digital human is continued to be displayed from the interrupted position, ensuring the integrity and coherence of the user's live broadcast viewing and enhancing the user's viewing experience.

[0074] In some implementations, when the target element only includes the digital human's mouthpiece broadcast information, an instant message in text form is generated based on the digital human's mouthpiece broadcast information and replied in the interaction area on the live broadcast interface. The instant message is used to update the live broadcast screen of the digital human; it can be understood that the digital human's mouthpiece broadcast information is the information used to answer users' questions. When the camera type does not meet the requirements for digital human mouthpiece broadcast and video material display, the digital human's mouthpiece broadcast information is generated into an instant message in text form and replied to the interaction area, such as the comment area of the live broadcast screen, to display the instant message in text form to the user, and at the same time, the live broadcast screen of the digital human is updated according to the content of the instant message. Combining the live broadcast screen and the instant message in text form helps users understand product-related information more timely and provides a richer interactive experience for users.

[0075] In this embodiment, by sending the user's comment information to the server, the server matches to obtain the target video material and the digital human voiceover information that match the user's question in the user's comment information. After receiving the target video material and the digital human voiceover information sent by the server, according to the lens type of the current live broadcast screen of the digital human, the target element is determined from the target video material and the digital human voiceover information. When the target element is the target video material and the digital human voiceover information, the live broadcast screen is terminated and the screen interruption position is recorded. After the target element is displayed, the subsequent live broadcast of the digital human continues from the screen interruption position, ensuring the integrity and coherence of the user's live broadcast viewing; when displaying the target element, the voiceover screen is obtained by rendering the digital human voiceover information, the target video material is rendered and the rendered video material is associated with the voiceover screen to generate the inserted screen of the digital human for display, ensuring that a better display and explanation effect of the product features can be achieved when displaying the inserted screen. The target video material can also be played at a specified position on the live broadcast interface, improving the flexibility of the target video material display; when the target material only includes the digital human voiceover information, an instant message in text form is generated based on the digital human voiceover information and displayed in the interaction area to help users understand the product-related information in a timely manner and provide a richer interactive experience for users.

[0076] Figure 3 It is a schematic diagram of another live broadcast method provided by an embodiment of the present application. As Figure 3 shown, the method includes the following steps:

[0077] S301, Obtain the user's comment information sent by the client.

[0078] It can be understood that the client can be the application program corresponding to the live broadcast page. The user can post comment information in the comment area or the bullet screen area of the live broadcast page. The comment information can include the user's questions and the content the user wants to know. The client can obtain the user's comment information and send the comment information to the server, and the server analyzes the user's comment information sent by the client.

[0079] S302, Obtain the target video material and the digital human voiceover information corresponding to the comment information.

[0080] Optionally, keyword extraction can be performed based on the comment information to determine the user's main questions and the main content the user wants to know in the comment information, and match according to the keyword from the pre-constructed material library. The material library can include video materials and reply materials, so as to determine the video material content in the material library that matches the keyword as the target video material, and determine the reply material in the material library that matches the keyword as the digital human voiceover information.

[0081] S303. Send at least one target element among the target video material and the digital human's voice-over information to the client according to the shot type of the current live broadcast image of the digital human.

[0082] In some implementations, the shot types of the current live broadcast image of the digital human may include a first type and a second type. The first type may be a normal shot, and the second type is a segmented shot. When the shot type is the first type, that is, a normal shot, the target elements are the target video material and the digital human's voice-over information; when the shot type is the second type, that is, a segmented shot, the target element is the digital human's voice-over information.

[0083] Optionally, the client can determine the shot type of the current live broadcast image of the digital human and send the shot type of the current live broadcast image to the server. When the server receives the shot type, it determines the target elements according to the shot type and sends the target elements to the client for display, so as to ensure that the client can smoothly display the target elements and reduce the consumption of redundant resource transmission.

[0084] In some implementations, the server can also send all the target video material and the digital human's voice-over information to the client. The client determines the target elements according to the shot type and displays the target elements. The display and presentation of the video material and the digital human's voice-over information are more flexible, improving the user experience.

[0085] In this embodiment, receive the comment information of the user sent by the client, match it from the material library according to the comment information, and obtain the target video material and the digital human's voice-over information corresponding to the comment information. Answer the user's questions more comprehensively through the target video material and the digital human's voice-over information, improving the user interaction experience; determine the target elements from the target video material and the digital human's voice-over information according to the shot type of the current live broadcast image of the digital human, and send the target elements to the client to ensure that the client can smoothly display the target elements and reduce the consumption of redundant resource transmission.

[0086] Figure 4 It is a schematic diagram of another live broadcast method provided by the embodiments of the present application. As Figure 4 shown, the method includes the following steps:

[0087] S401. Obtain the comment information of the user sent by the client.

[0088] In the embodiments of the present disclosure, the implementation method of step S401 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and no further description is given.

[0089] S402. Query the first Q&A pair associated with the comment information from the pre-constructed Q&A database.

[0090] Optionally, the Q&A database can be pre-configured. Obtain the configuration operation and determine Q&A pairs based on the configuration operation; obtain video materials and associate the video materials with the Q&A pairs to generate a Q&A database.

[0091] In some implementations, the configuration operation is used to configure the Q&A pairs in the Q&A database, such as the generation method and content of the Q&A pairs, etc.; determine questions and corresponding answers based on the configuration operation to form Q&A pairs; optionally, the video materials can be videos in various forms such as pre-generated explanations and demonstrations. Associating the video materials with the Q&A pairs means associating the Q&A pairs with the same content and the video materials, thereby generating a Q&A database, so that when the user obtains text information in the Q&A database, they can also determine the associated video materials, enhancing the richness and intuitiveness of the Q&A pairs.

[0092] In some implementations, during the construction process of the Q&A database, it is also possible to obtain a material rendering configuration operation and, based on the material rendering configuration operation, obtain the screen rendering parameters of the video materials; the material rendering configuration operation is used to perform operations such as editing or rendering on the video materials, and determine the screen rendering parameters of the video materials according to the material rendering configuration operation. The screen rendering parameters can be parameters such as size, contrast, and brightness, ensuring that the video materials maintain a high image quality after rendering and improving the display effect of the video materials.

[0093] Furthermore, query the first Q&A pair associated with the comment information from the Q&A database, for example, extract the keywords in the comment information, and perform keyword matching between the keywords and the Q&A pairs in the Q&A database to obtain the first Q&A pair that meets the similarity requirements.

[0094] In some implementations, it is also possible to determine the target user type of the user according to the target user identifier corresponding to the comment information; according to the target user type, determine the target acquisition process of the target video materials and the digital human voiceover information; according to the target acquisition process, obtain the target video materials and the digital human voiceover information.

[0095] Optionally, in this embodiment, the target user identifier may be a unique identifier such as a user account. The target user types may include authorized users and unauthorized users. The target user identifier can be judged through a user whitelist. Those within the user whitelist are authorized users, thereby determining the target user type of the user. For ordinary unauthorized users, the target acquisition process may be to match according to the comment information from the Q&A database. When a Q&A pair is matched, the target video material and the digital human voice-over information are determined according to the Q&A pair. When no Q&A pair is matched, a reply message is generated through a large model, and the Q&A pair matching is performed again according to the reply message to determine the target video material and the digital human voice-over information. For special authorized users, such as third-party users, the target video material and the digital human voice-over information can be obtained through their unique platform material library, and the acquisition flexibility of the target video material and the digital human voice-over information is stronger.

[0096] S403. Determine the first video material associated with the first Q&A pair as the target video material.

[0097] S404. Determine the first reply message in the first Q&A pair as the digital human voice-over information.

[0098] In some implementations, in response to not querying the first Q&A pair from the Q&A database, a second reply message corresponding to the comment information is generated based on the large model; keyword matching is performed on the second reply message and the Q&A pairs in the Q&A database. That is, if no first Q&A pair meeting the similarity requirement is queried from the Q&A data, the user's comment information is input into the pre-trained large model, and the large model generates a second reply message corresponding to the comment information. Further, keyword matching is performed on the second reply message and the Q&A pairs in the Q&A database to determine whether there is a second Q&A pair corresponding to the second reply message, so as to improve the accuracy of obtaining the target video material and the digital human voice-over information and avoid the failure of obtaining the target video material and the digital human voice-over information caused by keyword matching errors.

[0099] In response to matching a second Q&A pair corresponding to the second reply message, determine the second video material associated with the second Q&A pair as the target video material; obtain the second reply message in the second Q&A pair as the digital human voice-over information.

[0100] In response to not matching a second Q&A pair corresponding to the second reply message, determine the second reply message as the digital human voice-over information.

[0101] S405. According to the lens type of the current live broadcast screen of the digital human, send at least one target element in the target video material and the digital human voice-over information to the client.

[0102] In the embodiments of the present disclosure, the implementation method of step S405 can be implemented in any one of the embodiments of the present disclosure, without limitation thereto and will not be elaborated herein.

[0103] In this embodiment, the comment information of the user sent by the client is received, and the first Q&A pair associated with the comment information is queried from the pre-constructed Q&A database. When constructing the Q&A database, the Q&A pair is determined through configuration operations, and the Q&A pair is associated with the video material. The picture rendering parameters can also be determined through material rendering configuration operations to ensure that the video material maintains a high picture quality after rendering and improve the display effect of the video material. When the first Q&A pair is queried, the first video material corresponding to the first Q&A pair is used as the target video material, and the first reply information in the first Q&A pair is used as the digital voiceover information. When the first Q&A pair is not queried, the second reply information is generated through a large model, and keyword matching is performed in the Q&A database according to the second reply information to determine whether there is a matching second Q&A pair. When the second Q&A pair is matched, the second video material associated with the second Q&A pair is determined as the target video material, and the second reply information in the second Q&A pair is the digital voiceover information. When the second Q&A pair is not matched, the second reply information is used as the digital voiceover information, so as to obtain the target video material and digital voiceover information corresponding to the accurate comment information, improve the accuracy of the reply to the user, and enhance the user interaction experience.

[0104] Figure 5 It is a schematic diagram of another live broadcast method provided by the embodiments of the present application. As Figure 5 shown, the method includes the following steps:

[0105] S501, Obtain the comment information of the user sent by the client.

[0106] In the embodiments of the present disclosure, the implementation method of step S501 can be implemented in any one of the embodiments of the present disclosure, without limitation thereto and will not be elaborated herein.

[0107] S502, Query the first Q&A pair associated with the comment information from the pre-constructed Q&A database.

[0108] In the embodiments of the present disclosure, the implementation method of step S502 can be implemented in any one of the embodiments of the present disclosure, without limitation thereto and will not be elaborated herein.

[0109] S503, Determine the first video material associated with the first Q&A pair as the target video material.

[0110] In the embodiments of the present disclosure, the implementation method of step S503 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made thereto herein, nor will it be elaborated further.

[0111] S504. Determine the first reply information in the first Q&A pair as the digital human voiceover information.

[0112] In the embodiments of the present disclosure, the implementation method of step S504 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made thereto herein, nor will it be elaborated further.

[0113] S505. In response to the lens type being the first type, determine the target video material and the digital human voiceover information as the target elements.

[0114] Optionally, the first type can be a normal lens, which includes that the digital human can perform digital human voiceover as well as video playback and display. Therefore, the target video material and the digital human voiceover information are determined as the target elements.

[0115] S506. In response to the lens type being the second type, determine the digital human voiceover information as the target element.

[0116] Optionally, the second type can be a storyboard lens, in which there is no digital human and thus no digital human voiceover and video playback can be performed. Therefore, the digital human voiceover information is determined as the target element.

[0117] S507. Send the target elements to the client.

[0118] It can be understood that in this embodiment, when the server receives the client comment information, it can determine the lens type of the client and determine the target elements according to the lens type. When feeding back to the client based on the comment information, only the target elements are sent to the client for the client to display; in other embodiments, the server can send both the digital human voiceover information and the target video material to the client, and the client determines the target elements according to the lens type and displays the target elements on the client side.

[0119] In this embodiment, the comment information of the user sent by the client is received, and the first Q&A pair associated with the comment information is queried from the pre-constructed Q&A database. When constructing the Q&A database, the Q&A pair is determined through configuration operations, and the Q&A pair is associated with the video material. The screen rendering parameters can also be determined through material rendering configuration operations to ensure that the video material maintains a high image quality after rendering and improve the display effect of the video material. When the first Q&A pair is queried, the first video material corresponding to the first Q&A pair is used as the target video material, and the first reply information in the first Q&A pair is used as the digital voiceover information. When the first Q&A pair is not queried, the second reply information is generated through a large model, and keyword matching is performed in the Q&A database according to the second reply information to determine whether there is a matching second Q&A pair. When the second Q&A pair is matched, the second video material associated with the second Q&A pair is determined as the target video material, and the second reply information in the second Q&A pair is the digital voiceover information. When the second Q&A pair is not matched, the second reply information is used as the digital voiceover information, so as to obtain the target video material and digital voiceover information corresponding to the accurate comment information, improve the accuracy of the reply to the user, and enhance the user interaction experience. Further, the target element is determined according to whether the lens type is the first type or the second type, and the target element is sent to the client to ensure that the client can smoothly display the target element and reduce the consumption of redundant resource transmission.

[0120] Figure 6 It is an interaction schematic diagram of a live broadcast method provided by an embodiment of the present application. As Figure 6 shown, the live broadcast method includes:

[0121] S601. The client sends the comment information of the user to the server.

[0122] In the embodiment of the present disclosure, the implementation method of step S601 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and no further description is given.

[0123] S602. The server queries the first Q&A pair associated with the comment information from the pre-constructed Q&A database.

[0124] In the embodiment of the present disclosure, the implementation method of step S602 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and no further description is given.

[0125] S603. Determine the first video material associated with the first Q&A pair as the target video material.

[0126] In the embodiment of the present disclosure, the implementation method of step S603 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and no further description is given.

[0127] S604. Determine the first response information in the first Q&A pair as the digital human voiceover information.

[0128] In the embodiments of the present disclosure, the implementation method of step S604 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and will not be elaborated further.

[0129] S605. In response to the lens type being the first type, determine the target video material and the digital human voiceover information as the target elements.

[0130] In the embodiments of the present disclosure, the implementation method of step S605 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and will not be elaborated further.

[0131] S606. In response to the lens type being the second type, determine the digital human voiceover information as the target element.

[0132] In the embodiments of the present disclosure, the implementation method of step S606 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and will not be elaborated further.

[0133] S607. The server sends the target elements to the client.

[0134] In the embodiments of the present disclosure, the implementation method of step S607 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and will not be elaborated further.

[0135] S608. In response to the target elements including the target video material and the digital human voiceover information, the client interrupts the current live broadcast screen of the digital human and records the position where the screen is interrupted.

[0136] In the embodiments of the present disclosure, the implementation method of step S608 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and will not be elaborated further.

[0137] S609. Generate an inserted broadcast screen of the digital human based on the target video material and the digital human voiceover information, and display the inserted broadcast screen.

[0138] In the embodiments of the present disclosure, the implementation method of step S609 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and will not be elaborated further.

[0139] In this embodiment, the client obtains the user's comment information and sends it to the server. The server queries the first Q&A pair associated with the comment information from the pre-constructed Q&A database. When the first Q&A pair is queried, the first video material corresponding to the first Q&A pair is used as the target video material, and the first reply information in the first Q&A pair is used as the digital human voiceover information, obtaining the target video material and the digital human voiceover information corresponding to the accurate comment information, improving the accuracy of the user's reply and enhancing the user interaction experience; further determining the target element according to whether the lens type is the first type or the second type, and sending the target element to the client. In response to the target element including the target video material and the digital human voiceover information, the client interrupts the current live broadcast screen of the digital human and records the position where the screen is interrupted. Based on the target video material and the digital human voiceover information, an inserted broadcast screen of the digital human is generated and the inserted broadcast screen is displayed, helping the user to timely understand the product-related information and providing a richer interaction experience for the user.

[0140] Figure 7 It is an interaction schematic diagram of another live broadcast method provided by an embodiment of the present application. As Figure 7 shown, this live broadcast method includes:

[0141] S701, the client sends the user's comment information to the server.

[0142] In the embodiments of the present disclosure, the implementation method of step S701 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and no further description is given.

[0143] S702, the server queries the first Q&A pair associated with the comment information from the pre-constructed Q&A database.

[0144] In the embodiments of the present disclosure, the implementation method of step S702 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and no further description is given.

[0145] S703, determine the first video material associated with the first Q&A pair as the target video material.

[0146] In the embodiments of the present disclosure, the implementation method of step S703 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and no further description is given.

[0147] S704, determine the first reply information in the first Q&A pair as the digital human voiceover information.

[0148] In the embodiments of the present disclosure, the implementation method of step S704 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and no further description is given.

[0149] S705, the client receives the target video material and the digital voiceover information sent by the server.

[0150] In the embodiments of the present disclosure, the implementation method of step S705 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and will not be elaborated further.

[0151] S706, determine the target elements from the target video material and the digital voiceover information according to the shot type of the current live broadcast screen of the digital human.

[0152] In the embodiments of the present disclosure, the implementation method of step S706 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and will not be elaborated further.

[0153] S707, in response to the target elements including the target video material and the digital voiceover information, interrupt the current live broadcast screen of the digital human and record the position where the screen is interrupted.

[0154] In the embodiments of the present disclosure, the implementation method of step S707 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and will not be elaborated further.

[0155] S708, generate an inserted broadcast screen of the digital human based on the target video material and the digital voiceover information, and display the inserted broadcast screen.

[0156] In the embodiments of the present disclosure, the implementation method of step S708 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and will not be elaborated further.

[0157] In this embodiment, the client obtains the user's comment information and sends it to the server. The server queries the first Q&A pair associated with the comment information from the pre-constructed Q&A database. When the first Q&A pair is queried, the first video material corresponding to the first Q&A pair is used as the target video material, and the first reply information in the first Q&A pair is used as the digital voiceover information, so as to obtain the target video material and the digital voiceover information corresponding to the accurate comment information, improve the accuracy of the user's reply, and enhance the user interaction experience; further send the target video material and the digital voiceover information to the client. The client determines the target elements according to whether the shot type is the first type or the second type, and when the target elements include the target video material and the digital voiceover information, interrupt the current live broadcast screen of the digital human and record the position where the screen is interrupted. Generate an inserted broadcast screen of the digital human based on the target video material and the digital voiceover information, and display the inserted broadcast screen, helping the user to timely understand the product-related information and providing a richer interaction experience for the user.

[0158] Figure 8 is a logical flowchart of a live broadcast method provided by an embodiment of the present application. As Figure 8 shown, after the user publishes comment information in the live broadcast interface, question-answer pairs are matched from the question-answer database based on the comment information, and it is determined whether corresponding question-answer pairs are matched. When corresponding question-answer pairs are matched, the video material and the voice-over screen associated with the question-answer pairs are rendered, and the target elements in the video material and the voice-over screen are determined according to the current camera type. During normal shot segmentation, the live broadcast screen is interrupted and the video screen is linked with the voice-over for reply. During cut-shot segmentation, a text reply of the voice-over content is made; when no question-answer pairs are matched, a reply message is generated through a large model, and keyword matching is performed in the question-answer database according to the reply message. If corresponding question-answer pairs are matched, the video material and the voice-over screen associated with the question-answer pairs are rendered, and the target elements in the video material and the voice-over screen are determined for corresponding display; if no corresponding question-answer pairs are matched, the reply message is displayed in the live broadcast screen, enhancing the interaction between the user and the digital human anchor and improving the user experience.

[0159] Figure 9 is a schematic structural diagram of a live broadcast device provided by an embodiment of the present application. As Figure 9 shown, the live broadcast device 900 includes:

[0160] A sending module 901, configured to send the user's comment information to the server, where the comment information is used to obtain corresponding target video material and digital human voice-over information;

[0161] A receiving module 902, configured to receive the target video material and digital human voice-over information sent by the server;

[0162] An obtaining module 903, configured to determine target elements from the target video material and digital human voice-over information according to the camera type of the current live broadcast screen of the digital human;

[0163] An updating module 904, configured to update the live broadcast screen of the digital human based on the target elements.

[0164] In some implementations, the obtaining module 903 includes:

[0165] In response to the camera type being the first type, determining the target video material and digital human voice-over information as target elements;

[0166] In response to the camera type being the second type, determining the digital human voice-over information as the target element.

[0167] In some implementations, the updating module 904 includes:

[0168] In response to the target element including the target video material and digital voiceover information, interrupt the current live broadcast screen of the digital human and record the position where the screen is interrupted;

[0169] Generate an inserted broadcast screen of the digital human based on the target video material and digital voiceover information, and display the inserted broadcast screen.

[0170] In some implementations, the update module 904 includes:

[0171] Render the voiceover screen of the digital human based on the digital voiceover information;

[0172] Render the target video material based on the screen rendering parameters of the target video material, and associate it with the voiceover screen after rendering to generate an inserted broadcast screen of the digital human.

[0173] In some implementations, the update module 904 includes:

[0174] Display the voiceover screen of the digital human, and during the display of the voiceover screen, play the target video material at a specified position on the live broadcast interface to update the live broadcast screen of the digital human.

[0175] In some implementations, the device 900 further includes:

[0176] Monitor the display progress of the voiceover screen, and in response to the voiceover screen reaching the end time, stop playing the target video material;

[0177] Return the screen interruption position, and continue to display the subsequent live broadcast screen of the digital human from the interruption position.

[0178] In some implementations, the update module 904 includes:

[0179] In response to the target element including digital voiceover information, generate an instant message in text form based on the digital voiceover information, and reply to the instant message in the interaction area on the live broadcast interface. The instant message is used to update the live broadcast screen of the digital human.

[0180] In this embodiment, by sending the user's comment information to the server, the server performs matching to obtain the target video material and the digital human voiceover information that match the user's question in the user's comment information. After receiving the target video material and the digital human voiceover information sent by the server, according to the lens type of the current live broadcast screen of the digital human, the target element is determined from the target video material and the digital human voiceover information. When the target element is the target video material and the digital human voiceover information, the live broadcast screen is terminated and the screen interruption position is recorded. After the target element is displayed, the subsequent live broadcast of the digital human continues from the screen interruption position, ensuring the integrity and coherence of the user's live broadcast viewing; when displaying the target element, the voiceover screen is obtained by rendering the digital human voiceover information, the target video material is rendered, and the rendered video material is associated with the voiceover screen to generate the inserted screen of the digital human for display, ensuring that the effect of better displaying and explaining the product features can be achieved when displaying the inserted screen. The target video material can also be played at a specified position on the live broadcast interface, improving the flexibility of the target video material display; when the target material only includes the digital human voiceover information, an instant message in text form is generated based on the digital human voiceover information and displayed in the interaction area to help the user timely understand the product-related information and provide a richer interaction experience for the user.

[0181] Figure 10 It is a schematic structural diagram of another live broadcast device provided by an embodiment of the present application. As Figure 10 shown, the live broadcast device 1000 includes:

[0182] A first acquisition module 1001, configured to acquire the user's comment information sent by the client;

[0183] A second acquisition module 1002, configured to acquire the target video material and the digital human voiceover information corresponding to the comment information;

[0184] A sending module 1003, configured to send at least one target element in the target video material and the digital human voiceover information to the client according to the lens type of the current live broadcast screen of the digital human.

[0185] In some implementations, the second acquisition module 1002 includes:

[0186] Query the first question-and-answer pair associated with the comment information from the pre-constructed question-and-answer database;

[0187] Determine the first video material associated with the first question-and-answer pair as the target video material;

[0188] Determine the first reply information in the first question-and-answer pair as the digital human voiceover information.

[0189] In some implementations, the device 1000 further includes:

[0190] In response to the first Q&A pair not being found in the Q&A database, generate second response information corresponding to the comment information based on the large model;

[0191] Perform keyword matching on the second response information and the Q&A pairs in the Q&A database;

[0192] In response to matching the second Q&A pair corresponding to the second response information, determine the second video material associated with the second Q&A pair as the target video material;

[0193] Obtain the second response information in the second Q&A pair as the digital human voiceover information.

[0194] In some implementations, the apparatus 1000 further includes:

[0195] In response to the second Q&A pair corresponding to the second response information not being matched, determine the second response information as the digital human voiceover information.

[0196] In some implementations, the sending module 1003 includes:

[0197] In response to the lens type being the first type, determine the target video material and the digital human voiceover information as target elements;

[0198] In response to the lens type being the second type, determine the digital human voiceover information as the target element.

[0199] In some implementations, the second acquisition module 1002 includes:

[0200] Obtain a configuration operation and determine a Q&A pair based on the configuration operation;

[0201] Obtain video materials and associate the video materials with the Q&A pairs to generate a Q&A database.

[0202] In some implementations, the apparatus 1000 further includes:

[0203] Obtain a material rendering configuration operation and obtain the screen rendering parameters of the video material based on the material rendering configuration operation.

[0204] In some implementations, the second acquisition module 1002 includes:

[0205] Obtain the target user identifier corresponding to the comment information and determine the target user type of the user according to the target user identifier;

[0206] According to the target user type, determine the target acquisition process of the target video material and the digital human voiceover information;

[0207] According to the target acquisition process, obtain the target video material and the digital human voiceover information.

[0208] In this embodiment, the comment information of the user sent by the client is received, and the first Q&A pair associated with the comment information is queried from the pre-constructed Q&A database. When constructing the Q&A database, the Q&A pair is determined through configuration operations, and the Q&A pair is associated with the video material. The screen rendering parameters can also be determined through the material rendering configuration operation to ensure that the video material maintains a high picture quality after rendering and improve the display effect of the video material. When the first Q&A pair is queried, the first video material corresponding to the first Q&A pair is used as the target video material, and the first reply information in the first Q&A pair is used as the digital voiceover information. When the first Q&A pair is not queried, the second reply information is generated through a large model, and keyword matching is performed in the Q&A database according to the second reply information to determine whether there is a matching second Q&A pair. When the second Q&A pair is matched, the second video material associated with the second Q&A pair is determined as the target video material, and the second reply information in the second Q&A pair is the digital voiceover information. When the second Q&A pair is not matched, the second reply information is used as the digital voiceover information, so as to obtain the target video material and digital voiceover information corresponding to the accurate comment information, improve the accuracy of the reply to the user, and enhance the user interaction experience. Further, the target element is determined according to whether the lens type is the first type or the second type, and the target element is sent to the client to ensure that the client can smoothly display the target element and reduce the consumption of redundant resource transmission.

[0209] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0210] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0211] Figure 11 A schematic block diagram of an electronic device for implementing the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described herein and / or claimed.

[0212] As Figure 11As shown, device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the device 1100 can also be stored. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.

[0213] Multiple components in the device 1100 are connected to the I / O interface 1105, including: an input unit 1106, such as a keyboard, a mouse, etc.; an output unit 1107, such as various types of displays, speakers, etc.; a storage unit 1108, such as a magnetic disk, an optical disc, etc.; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1109 allows the device 1100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0214] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1101 executes the various methods and processes described above, such as the live broadcast method. For example, in some embodiments, the live broadcast method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the computing unit 1101, one or more steps of the live broadcast method described above can be executed. Alternatively, in other embodiments, the computing unit 1101 can be configured to execute the live broadcast method in any other appropriate manner (e.g., by means of firmware).

[0215] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0216] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on the remote machine or server.

[0217] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0218] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0219] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0220] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.

[0221] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0222] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A live broadcast method, wherein: The method comprises: Sending the user's comment information to the server, wherein the comment information is used to obtain the corresponding target video material and digital voice broadcast information; Receiving the target video material and the digital voice broadcast information sent by the server; According to the shot type of the current live broadcast screen of the digital human, determining the target element from the target video material and the live broadcast information of the digital human; The live broadcast screen of the digital human is updated based on the target element.

2. The method according to claim 1, wherein: The step of determining the target element from the target video material and the digital human live broadcast information according to the shot type of the current live broadcast screen of the digital human includes: In response to the shot type being the first type, determining the target video material and the digital voice broadcast information as the target elements; In response to the shot type being the second type, determining the digital human voice broadcast information as the target element.

3. The method according to claim 1, wherein: The updating of the live broadcast screen of the digital human based on the target element comprises: In response to the target element including the target video material and the digital human live broadcast information, interrupting the current live broadcast screen of the digital human and recording the screen interruption position; Based on the target video material and the digital human broadcast information, an insert screen of the digital human is generated and displayed.

4. The method according to claim 3, wherein: The step of generating the digital human insert screen based on the target video material and the digital human broadcast information includes: Rendering the digital person's spoken broadcast picture based on the digital person's spoken broadcast information; Based on the picture rendering parameters of the target video material, the target video material is rendered, and after rendering, it is associated with the oral picture to generate the insert picture of the digital human.

5. The method according to claim 4, wherein: The displaying of the interstitial screen includes: The spoken broadcast screen of the digital human is displayed, and during the display of the spoken broadcast screen, the target video material is played at a designated position of the live broadcast interface to update the live broadcast screen of the digital human.

6. The method according to claim 5, wherein: The method further comprises: Monitoring the display progress of the oral broadcast picture, and stopping the playback of the target video material in response to the oral broadcast picture reaching the end time; Return to the interrupted position of the screen, and continue to display the subsequent live screen of the digital human from the interrupted position.

7. The method according to claim 1, wherein: The updating of the live broadcast screen of the digital human based on the target element comprises: In response to the target element including the digital human broadcast information, an instant message in text form is generated based on the digital human broadcast information, and the instant message is replied in the interactive area on the live broadcast interface, and the instant message is used to update the live broadcast screen of the digital human.

8. A live broadcast method, wherein: The method comprises: Get the user's comment information sent by the client; Obtaining target video material and digital voice broadcast information corresponding to the comment information; According to the lens type of the current live broadcast screen of the digital human, the target video material and at least one target element in the digital human live broadcast information are sent to the client.

9. The method according to claim 8, wherein: The step of obtaining the target video material and digital voice broadcast information corresponding to the comment information includes: Querying a first question-answer pair associated with the review information from a pre-built question-answer database; Determining a first video material associated with the first question-answer pair as the target video material; Determine the first answer information in the first question-answer pair as the digital population broadcast information.

10. The method according to claim 9, wherein: The method further comprises: In response to not finding the first question-answer pair in the question-answer database, generating second reply information corresponding to the comment information based on the big model; Performing keyword matching on the second reply information and the question-answer pairs in the question-answer database; In response to matching a second question-answer pair corresponding to the second reply information, determining a second video material associated with the second question-answer pair as the target video material; The second answer information in the second question and answer pair is obtained as the digital population broadcast information.

11. The method according to claim 10, wherein: The method further comprises: In response to not matching a second question-answer pair corresponding to the second reply information, determining the second reply information as the digital population broadcast information.

12. The method according to any one of claims 8 to 11, wherein: According to the shot type of the current live broadcast screen of the digital human, at least one target element is determined from the target video material and the live broadcast information of the digital human, including: In response to the shot type being the first type, determining the target video material and the digital voice broadcast information as the target elements; In response to the shot type being the second type, determining the digital human voice broadcast information as the target element.

13. The method according to any one of claims 9 to 11, wherein: The process of constructing the question-answer database includes: Obtaining a configuration operation, and determining a question-answer pair based on the configuration operation; Video material is acquired, and the video material is associated with the question-answer pair to generate the question-answer database.

14. The method according to claim 13, wherein: The method further comprises: Obtain a material rendering configuration operation, and based on the material rendering configuration operation, obtain a picture rendering parameter of the video material.

15. The method according to any one of claims 8 to 11, wherein: The step of obtaining the target video material and digital voice broadcast information corresponding to the comment information includes: Obtaining a target user identifier corresponding to the comment information, and determining a target user type of the user according to the target user identifier; Determining a target acquisition process of the target video material and digital voice broadcast information according to the target user type; According to the target acquisition process, the target video material and digital voice broadcast information are acquired.

16. A live broadcast device, wherein: include: A sending module, used to send the user's comment information to the server, wherein the comment information is used to obtain the corresponding target video material and digital voice broadcast information; A receiving module, used for receiving the target video material and the digital voice broadcast information sent by the server; An acquisition module, used to determine the target element from the target video material and the digital human live broadcast information according to the shot type of the current live broadcast screen of the digital human; An updating module is used to update the live broadcast screen of the digital human based on the target element.

17. A live broadcast device, wherein: include: The first acquisition module is used to acquire the user's comment information sent by the client; The second acquisition module is used to acquire the target video material and digital voice broadcast information corresponding to the comment information; The sending module is used to send the target video material and at least one target element in the digital human broadcast information to the client according to the lens type of the current live broadcast screen of the digital human.

18. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7 or 8-15.

19. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7 or 8-15.

20. A computer program product comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1-7 or 8-15.