Digital human live broadcast method, apparatus and device, and storage medium
By extracting key information in the live broadcast room and using big models to generate digital broadcast text and voices, the problem of time-consuming manpower in the existing technology is solved, and the broadcast content that is automated and quickly generated matching live broadcast content is improved, which improves the real-time and accuracy of live broadcasts.
Patent Information
- Application Number
- CN202510352512.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-25
AI Technical Summary
The existing technology requires a lot of manpower and professional ability to keep up with hot content in live broadcasts, and it is time-consuming and difficult to efficiently generate live broadcast manuscripts of hot content.
By extracting key information in the live broadcast room, using a big model to generate digital broadcast text and sounds, and adjusting live broadcast actions, combining text-to-speech technology and image recognition technology, broadcast content that matches live broadcast content.
It realizes automation and rapid generation of broadcast texts and actions related to live broadcast content, reduces manpower demand and improves the real-time and accuracy of live broadcast content.
Smart Images

Figure CN120374806A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and particularly to the fields of large models, information flow, content generation, etc. Background Art
[0002] When conducting a live broadcast on the host side of hot content, it usually requires a large amount of manpower to collect materials and process them into manuscripts of hot content. Then, the host conducts the live broadcast according to the manuscript. This method not only requires a lot of manpower but also strong professional capabilities, which is time-consuming and difficult to keep up with the hot topics. Summary of the Invention
[0003] This disclosure provides a digital human live broadcast method, device, equipment, and storage medium.
[0004] According to one aspect of this disclosure, there is provided a digital human live broadcast method, including:
[0005] Extracting key information according to the background data stream of the live broadcast room;
[0006] Generating a first prompt word according to the key information;
[0007] Using a large model to generate a broadcast text of the digital human based on the first prompt word;
[0008] Generating the voice of the digital human according to the broadcast text;
[0009] Adjusting the live broadcast actions of the digital human according to the voice of the digital human.
[0010] According to another aspect of this disclosure, there is provided a digital human live broadcast device, including:
[0011] An extraction module for extracting key information according to the background data stream of the live broadcast room;
[0012] A first prompt word generation module for generating a first prompt word according to the key information;
[0013] A broadcast text generation module for using a large model to generate a broadcast text of the digital human based on the first prompt word;
[0014] A voice generation module for generating the voice of the digital human according to the broadcast text;
[0015] An adjustment module for adjusting the live broadcast actions of the digital human according to the voice of the digital human.
[0016] According to another aspect of this disclosure, there is provided an electronic device, including:
[0017] At least one processor; and
[0018] A memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute any method in the embodiments of the present disclosure.
[0020] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute any method in the embodiments of the present disclosure.
[0021] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements any method in the embodiments of the present disclosure.
[0022] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings
[0023] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0024] Figure 1 is a schematic flowchart of a digital human live broadcast method according to an embodiment of the present disclosure;
[0025] Figure 2 is a schematic flowchart of a digital human live broadcast method according to another embodiment of the present disclosure;
[0026] Figure 3 is a schematic flowchart of a digital human live broadcast method according to another embodiment of the present disclosure;
[0027] Figure 4 is a schematic flowchart of a digital human live broadcast method according to another embodiment of the present disclosure;
[0028] Figure 5 is a schematic diagram of live broadcast material superimposition according to an embodiment of the present disclosure;
[0029] Figure 6 is a schematic flowchart of a digital human live broadcast method according to another embodiment of the present disclosure;
[0030] Figure 7 is a schematic flowchart of a digital human live broadcast method according to another embodiment of the present disclosure;
[0031] Figure 8 is an application scenario diagram of a live broadcast screen;
[0032] Figure 9 It is a schematic flowchart of a digital life generation solution;
[0033] Figure 10 It is a schematic structural diagram of a digital human live broadcast device according to an embodiment of the present disclosure;
[0034] Figure 11 It is a schematic structural diagram of a digital human live broadcast device according to another embodiment of the present disclosure;
[0035] Figure 12 It is a schematic structural diagram of a digital human live broadcast device according to another embodiment of the present disclosure;
[0036] Figure 13 It is a block diagram of an electronic device for implementing the digital human live broadcast method of the embodiments of the present disclosure. Detailed implementation manners
[0037] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0038] Figure 1 It is a schematic flowchart of a digital human live broadcast method 100 according to an embodiment of the present disclosure, and the method includes:
[0039] S110. Extract key information according to the background data stream of the live broadcast room;
[0040] S120. Generate a first prompt word according to the key information;
[0041] S130. Use a large model to generate a broadcast text of the digital human based on the first prompt word;
[0042] S140. Generate the voice of the digital human according to the broadcast text;
[0043] S150. Adjust the live broadcast actions of the digital human according to the voice of the digital human.
[0044] In the embodiments of the present disclosure, a live streaming room includes a virtual space for real-time video and / or audio live streaming via the Internet, which is a real-time interactive platform built on the Internet through live streaming technology. The host can display various contents to the audience in the live streaming room. While displaying the contents in the live streaming room, it is usually necessary for the host to introduce the displayed contents. Some live streaming rooms are introduced by human hosts, and some live streaming rooms are introduced by digital human hosts. Digital humans can also be called virtual humans, virtual characters, digital hosts, etc. Digital humans can include virtual characters created by one or more technologies such as artificial intelligence, computer graphics, speech synthesis, virtual reality, etc. Digital humans can be virtual characters with human appearance, behavior, and interaction capabilities.
[0045] The background data stream of the live streaming room can include video data streams and / or audio data streams collected by devices such as cameras at the live streaming end. Some video data streams include audio data streams, and some video data streams do not include audio data streams. Various key information can be extracted from the live audio stream and / or live video stream of the live streaming room. In some cases, the key information can also be called keywords. There can be multiple extraction methods, such as extracting key information from the background data stream through one or more of large models, dedicated models, or dedicated tools. In some application scenarios, the digital human in the live streaming room can broadcast contents with high real-time requirements, such as news hotspots and other contents. In this case, the key information of news hotspots and other contents can be extracted from the background data stream.
[0046] In the embodiments of the present disclosure, prompts of large models, such as large language models, can be used to guide the large models to generate the required contents. For example, the prompts can include contents such as the text input when interacting with the large models. The first prompt of the large model can be generated using the key information extracted from the background data stream. For example, one or more key information such as the theme of the news event, the involved people, the location where the event occurred, the time of the event, the nature of the event, and the popular topic tags extracted from the news background data stream can be used to generate the first prompt.
[0047] In the embodiments of the present disclosure, key information can be input into a large model, and the large model can generate a first prompt word in combination with a prompt word template; alternatively, a first prompt word can be generated using a prompt word model first, and then the first prompt word can be input into the large model. Based on the key information of the background data stream in the first prompt word, the large model can generate the broadcast text of the digital human. For example, the large model uses the materials learned by itself that are related to the key information of the background data stream to generate a broadcast text with strong relevance to the content of the background data stream currently being played in the live room. For example, the broadcast text applicable to the hot background data stream can include "Hello everyone, today is xx month xx day, welcome to listen to today's hot live broadcast. The current live content is mainly about the progress of Hot Topic A. Specifically, it includes...". Another example is that the broadcast text applicable to the science popularization live room can include "Hello everyone, welcome to listen to today's science popularization live broadcast. The current live content is mainly about the detailed introduction of astronomical knowledge B. Specifically, it includes...".
[0048] In the embodiments of the present disclosure, the broadcast text can be converted into the voice of the digital human by means of text-to-speech technology (TTS), etc., and the live actions of the digital human can be adjusted according to the voice of the digital human. The live actions of the digital human can include the limb actions, lip actions, other expressions, etc. of the digital human. For example, the voice features in the voice of the digital human can be converted into the lip parameters of the digital human, such as the opening and closing degree and shape change of the digital human's lips. Through these lip parameters, the lip actions of the digital human can be controlled. The broadcast content of the digital human can include the voice of the digital human and the live actions of the digital human. The live actions of the digital human can also be referred to as the image of the digital human, the picture of the digital human, etc.
[0049] According to the embodiments of the present disclosure, a large model can be used to quickly and accurately generate the broadcast text of the digital human based on the key information of the background data stream in the live room, and then generate the voice of the digital human and adjust the live actions of the digital human in the live room, so that the broadcast content of the digital human is more adapted to the background data stream in the live room.
[0050] Figure 2 FIG. 200 is a schematic flowchart of a digital human live broadcast method 200 according to another embodiment of the present disclosure. This method 200 can be used to implement the step S110 in the digital human live broadcast method 100. In one implementation, the method 200 includes: extracting key information from the background data stream of the live room, which further includes:
[0051] S210. Recognize the voice and / or frame images of the background data stream of the live room to obtain the live content text;
[0052] S220. Extract the key information from the live content text.
[0053] For example, keyword extraction can be performed on audio through speech-to-text methods; keyword extraction can be performed on video through image recognition and key frame extraction methods; keyword extraction can be performed on the text in the live broadcast room through text recognition.
[0054] In the embodiments of the present disclosure, voice recognition can be performed on the sound in the audio data stream through a voice recognition tool. Then, after processing the recognition results of the audio data stream, such as proofreading, segmentation, and typesetting, the live broadcast content text is obtained. The frame images in the video data stream can be transmitted to an image recognition tool for image recognition through the image recognition tool. Image recognition can be performed on consecutive frame images, or on the frame images obtained by sampling the video data stream. After processing the image recognition results, such as proofreading, segmentation, and typesetting, the live broadcast content text is obtained. Then, key information can be extracted from the live broadcast content text.
[0055] According to the embodiments of the present disclosure, the live broadcast content text is obtained from the sound and / or frame images of the background data stream, and then key information is extracted. The extracted key information has a strong correlation with the background data stream, and the broadcast text of the digital human generated using this key information also has a strong correlation with the background data stream, which is beneficial to improving the accuracy of the broadcast content of the digital human in the live broadcast room.
[0056] In one implementation, the method further includes one or more of the following steps:
[0057] In the case where the key information cannot be extracted from the background data stream, the large model is used to generate the broadcast text based on a preset second prompt word;
[0058] In the case where the key information in the background data stream cannot generate the broadcast text that meets the broadcast requirements, the large model is used to generate the broadcast text based on a preset second prompt word;
[0059] In the case where a preset trigger condition is met, the large model is used to generate the broadcast text based on a preset second prompt word.
[0060] In the embodiments of the present disclosure, the second prompt word can be preset manually or generated automatically. For example, manually preset prompt words related to the live broadcast content are collected through a prompt word preset interface. Another example is to analyze the title, brief introduction, etc. of the live broadcast room through a large model to generate a preset second prompt word. The preset second prompt word can include one or more of, for example, the title of the live broadcast room, the focus content of attention, the expected live broadcast effect, the purpose of this live broadcast, the content that needs to be summarized and reported, etc.
[0061] In some scenarios, when obtaining key information through background data streams, there may be a variety of situations where key information cannot be extracted. For example, the traffic data of the background data stream is too large or the live broadcast delay is too high, resulting in the failure to support key information extraction from the background data stream, the inability to extract valid voice from the audio data stream, and the inability to extract valid text from the video data stream. In these cases, the preset second prompt word can be input into the large model to generate the broadcast text of the digital human.
[0062] In some scenarios, although key information can be extracted from the background data stream, there may be too few key information, insufficient information provided by the key information, or inaccurate key information. In this case, it may be impossible to generate a broadcast text that meets the broadcast requirements. In these cases, the preset second prompt word can be input into the large model to generate the broadcast text of the digital human.
[0063] In some scenarios, some trigger conditions can be preset. When the trigger conditions are met, it is not necessary to extract the key information of the background data stream or ignore the extracted key information. The large model directly generates the broadcast text based on the preset second prompt word corresponding to the trigger condition. Different trigger conditions can have different second prompt words. For example, time conditions (predetermined time trigger, periodic trigger, etc.) correspond to the second prompt word Prompt1, interaction conditions (voice trigger, click trigger, gesture trigger, etc.) correspond to the second prompt word Prompt2, data conditions (the number of viewers exceeds the threshold, the amount of barrage reaches the threshold, etc.) correspond to the second prompt word Prompt3, etc.
[0064] According to the embodiments of the present disclosure, the use of preset prompt words can support the generation of broadcast texts in different situations, thereby improving the flexibility and accuracy of generating broadcast texts.
[0065] Figure 3 3 is a flow chart of a digital human live broadcast method 300 according to another embodiment of the present disclosure, and the method 300 can be used to implement step S150 in the digital human live broadcast method 100. In one implementation, the method 300 includes: adjusting the live broadcast action of the digital human according to the voice of the digital human, and further includes:
[0066] S310, extracting voice features from the voice of the digital person;
[0067] S320, obtaining the lip shape parameters of the digital human corresponding to the sound feature;
[0068] S330: Adjust the lip movement of the digital human according to the sound feature and the corresponding lip shape parameter of the digital human.
[0069] In the embodiments of the present disclosure, voice features are extracted from the voice of the digital human through a voice processing tool, a voice processing model, etc. For example, the voice features may include one or more of the loudness, pitch, timbre, frequency, and duration of the voice, etc.
[0070] In the embodiments of the present disclosure, the lip parameters of the digital human can be determined according to the voice features. The lip parameters of the digital human may include one or more of the lip opening degree, lip width, lip protrusion, lip position, etc. For example, the lip opening degree of the digital human can be determined according to the loudness and pitch of the voice, etc. According to the lip parameters of the digital human, the lip movement of the digital human can be adjusted. For example, the lip opening degree of the digital human is adjusted according to the difference between the lip opening degree of the digital human determined according to the extracted voice features and the current lip opening degree.
[0071] According to the embodiments of the present disclosure, adjusting the picture of the digital human according to the voice features can improve the accuracy of the digital human picture.
[0072] Figure 4 FIG. 400 is a schematic flowchart of a digital human live broadcast method 400 according to another embodiment of the present disclosure. The method 400 can be used to implement the steps in the digital human live broadcast method 100. In one implementation, the method 400 further includes:
[0073] S410, obtaining the background data stream of the live broadcast room;
[0074] S420, superimposing the picture and voice of the digital human on the background data stream.
[0075] In the embodiments of the present disclosure, the live broadcast end camera can capture the live broadcast picture and / or voice on the scene to form a background data stream. The live broadcast platform can superimpose the background data stream with the picture and voice of the digital human to generate a live broadcast stream and push the live broadcast stream to the client of the audience for playback. A schematic diagram of superimposing the background data stream of a live broadcast room and a digital human is shown in Figure 5 as shown.
[0076] According to the embodiments of the present disclosure, by superimposing the picture and voice of the digital human on the background data stream of the live broadcast room, the generated live broadcast stream can simulate the live broadcast effect of a real live broadcaster.
[0077] In one implementation, as shown in Figure 4 shown, the method 400 further includes:
[0078] S430, superimposing the additional materials of the live broadcast room on the background data stream.
[0079] In the embodiments of the present disclosure, the additional materials in the live broadcast room may include: the watermark of the live broadcast room, the logo of the live broadcast platform, special effects for displaying the live broadcast effect, charts for understanding the live broadcast content, the pictures captured by the additional camera, and other materials. By superimposing the background data stream, the digital human, and the additional materials, a live stream can be generated.
[0080] According to the embodiments of the present disclosure, by generating a live stream by superimposing a digital human and additional materials in the background data stream, the flexibility and interest of the specific content of the live stream can be improved.
[0081] Figure 6 FIG. 600 is a schematic flowchart of a digital human live broadcast method 600 according to another embodiment of the present disclosure, and this method 600 can be used to implement the steps in the digital human live broadcast method 100. In one implementation, the method 600 further includes:
[0082] S610. Provide a prompt word setting interface in the live broadcast client; wherein, the content allowed to be provided by the user in the prompt word setting interface includes one or more of the following: title, focus hotspots, expected effects, live broadcast purpose, and content to be summarized;
[0083] S620. Generate the second prompt word in response to an editing instruction for the prompt word setting interface.
[0084] In the embodiments of the present disclosure, the prompt word setting interface can be displayed in the live broadcast client, and the prompt word setting interface may include the content allowed to be provided by the user. The content provided by the user can reflect the user's live broadcast requirements. The user can provide the required content through various operation methods such as input and selection. For example, the user can enter the title AA of the live broadcast room in the title input box of the prompt word setting interface, enter the hot keyword K1, K2 in the focus hotspots input box, and select the effect E1 in the expected effect option. After collecting the content provided by the user through the prompt word setting interface, the preset second prompt word can be generated according to these contents. The prompt word setting interface can also be referred to as a prompt word setting section, a prompt word setting module, a prompt word preset interface, etc.
[0085] According to the embodiments of the present disclosure, the content provided by the user collected using the prompt word setting interface can be used to generate a prompt word that better meets the live broadcast requirements, and further generate a more accurate broadcast text for the digital human, improving the accuracy of the digital human's broadcast content.
[0086] Figure 7 FIG. 700 is a schematic flowchart of a digital human live broadcast method 700 according to another embodiment of the present disclosure, and this method 700 can be used to implement the steps in the digital human live broadcast method 100. In one implementation, the method 700 further includes:
[0087] S710. Record the voice of a person, which is used to support converting target text into the voice of a digital human using TTS technology;
[0088] S720. Record the video of a person, which is used to generate the appearance of a digital human;
[0089] Among them, the voice of the digital human is used to adjust the lip movement in the appearance of the digital human.
[0090] In the embodiments of the present disclosure, the voice file of a real person, that is, the voice of a person, can be recorded through audio recording software. Through technologies such as voice data collection, feature extraction, deep learning model training, text-to-speech technology (TTS), and personalized voice generation, the features of the voice of a person are digitized to generate a digital voice. The digital voice is used as the voice source of the digital human. This can make the voice emitted by the digital human close to the voice of a person, making it more realistic.
[0091] In the embodiments of the present disclosure, the picture of a real person, that is, the video of a person, can be recorded through video recording software. For the recorded video of a person, artificial intelligence technology is used, combined with technologies such as three-dimensional modeling, motion capture, expression recognition, and rendering, to digitize the appearance, movements, and expressions of the real person, etc., to generate the virtual appearance of a digital human. This can make the appearance of the digital human close to the appearance of a person, making it more vivid and vivid.
[0092] In the embodiments of the present disclosure, the target text, such as the broadcast text in a live broadcast room, can be converted into the voice of a digital human through text-to-speech technology (TTS). Then, the lip movement (lip action) of the appearance of the digital human is adjusted using the voice of the digital human.
[0093] According to the embodiments of the present disclosure, by recording the voice and video of a real person, a more realistic voice and movements of the digital human can be obtained, making the digital human more vivid.
[0094] In an application scenario, on the host side, the production cost of content such as news hotspots is relatively high, and the production of content is time-consuming and cannot keep up with the hotspots. The embodiments of the present disclosure, combined with the AI large model and digital human technology, can automatically produce content.
[0095] For example: There is a slow live camera set up at a certain location. By pulling the relevant data of the camera, using the large model to generate a live broadcast script, combining TTS broadcast to drive the digital human, and synthesizing the relevant pictures of the camera and the digital human to generate a live broadcast for users to watch.
[0096] Such as Figure 8As shown, for example, for a certain news hot topic, relevant data (background data stream) can be captured in real time through a camera, and a script (broadcast text) can be generated based on key information. The script is converted into speech, and the digital human is driven by the speech. The picture of the digital human and the real-time news hot topic are merged together to generate a live broadcast (or called live stream).
[0097] As Figure 9 shown, an AI digital human live broadcast method may include the following steps:
[0098] S901. Record the voice of the anchor: Find an anchor to record the audio, supporting TTS broadcast (for example, input text and the audio reads the relevant text).
[0099] S902. Record video resources: Find an anchor to record video to generate a digital human model (for example, input audio, drive lip movement to form a digital human).
[0100] S903. Upload other materials: Such as watermarks, station logos, etc.
[0101] S904. Sort out key information such as keywords.
[0102] S905. Input the keywords into the large model.
[0103] S906. The large model can feedback the speaking lines, or called the anchor script. For example, the interpretation text of hot topic A.
[0104] S907. Pull the camera picture, for example, the slow live stream captured by the camera at the live broadcast end, including video data stream and / or audio data stream, etc.
[0105] S908. Input the interpretation text of the large model into TTS broadcast to generate the broadcast voice (or called audio).
[0106] S909. Input the broadcast voice into the digital human model to generate the digital human picture.
[0107] S910. Overlay the digital human picture and the camera picture, and overlay other materials, and encode and push the stream (push the live stream to the client).
[0108] S911. The client (C-side) user pulls the relevant stream to watch the live broadcast.
[0109] Figure 10 FIG. is a schematic structural diagram of a digital human live broadcast device 1000 according to an embodiment of the present disclosure. The device 1000 may include:
[0110] An extraction module 1010, configured to extract key information according to the background data stream of the live broadcast room;
[0111] The first prompt word generation module 1020 is configured to generate a first prompt word according to the key information;
[0112] The first broadcast text generation module 1030 is configured to use a large model to generate a broadcast text for the digital human based on the first prompt word;
[0113] The voice generation module 1040 is configured to generate the voice of the digital human according to the broadcast text;
[0114] The adjustment module 1050 is configured to adjust the live broadcast actions of the digital human according to the voice of the digital human.
[0115] Figure 11 FIG. 13 is a schematic structural diagram of a digital human live broadcast device 1100 according to another embodiment of the present disclosure. The device 1100 includes: an extraction module 1110, a first prompt word generation module 1120, a first broadcast text generation module 1130, a voice generation module 1140, and an adjustment module 1150. The functions of the above modules can refer to the functions of the respective modules of the digital human live broadcast device 1000 in the above embodiment. In one implementation manner, the extraction module 1110 includes:
[0116] The recognition sub-module 1111 is configured to recognize the voice and / or frame images of the background data stream of the live broadcast room to obtain a live broadcast content text;
[0117] The extraction sub-module 1112 is configured to extract the key information from the live broadcast content text.
[0118] Figure 12 FIG. 23 is a schematic structural diagram of a digital human live broadcast device 1200 according to another embodiment of the present disclosure. The device 1200 includes one or more features of any of the above device embodiments. In one implementation manner, the device further includes:
[0119] The second broadcast text generation module 1210 is configured to use the large model to generate the broadcast text based on a preset second prompt word when the key information cannot be extracted from the background data stream;
[0120] The second broadcast text generation module 1210 is further configured to use the large model to generate the broadcast text based on a preset second prompt word when the key information in the background data stream cannot generate the broadcast text that meets the broadcast requirements;
[0121] The second broadcast text generation module 1210 is further configured to use the large model to generate the broadcast text based on a preset second prompt word when a preset trigger condition is met.
[0122] In one implementation manner, the adjustment module 1150 includes:
[0123] An acquisition sub-module 1151, configured to acquire the voice feature according to the voice of the digital human;
[0124] A generation sub-module 1152, configured to acquire the lip parameter of the digital human corresponding to the voice feature;
[0125] An adjustment sub-module 1153, configured to adjust the lip movement of the digital human according to the voice feature and the corresponding lip parameter of the digital human.
[0126] In one embodiment, the apparatus further includes:
[0127] An acquisition module 1220, configured to acquire the background data stream of the live broadcast room;
[0128] A superimposition module 1230, configured to superimpose the picture and voice of the digital human on the background data stream.
[0129] In one embodiment, the superimposition module 1240 is further configured to superimpose the additional materials of the live broadcast room on the background data stream.
[0130] In one embodiment, the apparatus further includes:
[0131] A setting module 1240, configured to provide a prompt word setting interface in the live broadcast client; wherein, the content allowed to be provided by the user in the prompt word setting interface includes one or more of the following: title, focus of attention, expected effect, purpose of the live broadcast, and content to be summarized;
[0132] A second prompt word generation module 1250, configured to generate the second prompt word in response to an editing instruction for the prompt word setting interface.
[0133] In one embodiment, the apparatus further includes:
[0134] A text conversion module 1260, configured to record the voice of a person and convert the voice of the person into text;
[0135] A voice conversion module 1270, configured to convert the target text into the voice of the digital human by using TTS technology;
[0136] A digital human generation module 1280, configured to record a person's video and generate a digital human according to the person's video;
[0137] A digital human adjustment module 1290, configured to adjust the lip movement of the digital human according to the voice of the digital human.
[0138] For the specific functions and example descriptions of each module and sub-module of the apparatus according to the embodiments of the present disclosure, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated herein.
[0139] In the technical solutions of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0140] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0141] Figure 13 FIG. shows a schematic block diagram of an exemplary electronic device 1300 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital assistant, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0142] As Figure 13 shown, the device 1300 includes a computing unit 1301 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1302 or a computer program loaded from a storage unit 1308 into a random access memory (RAM) 1303. In the RAM 1303, various programs and data required for the operation of the device 1300 can also be stored. The computing unit 1301, the ROM 1302, and the RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.
[0143] A plurality of components in the device 1300 are connected to the I / O interface 1305, including: an input unit 1306, such as a keyboard, a mouse, etc.; an output unit 1307, such as various types of displays, speakers, etc.; a storage unit 1308, such as a magnetic disk, an optical disk, etc.; and a communication unit 1309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1309 allows the device 1300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0144] The computing unit 1301 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1301 executes the various methods and processes described above, such as the digital human live broadcast method. For example, in some embodiments, the digital human live broadcast method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1308. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1300 via the ROM 1302 and / or the communication unit 1309. When the computer program is loaded into the RAM 1303 and executed by the computing unit 1301, one or more steps of the digital human live broadcast method described above can be executed. Alternatively, in other embodiments, the computing unit 1301 can be configured to execute the digital human live broadcast method in any other suitable manner (e.g., by means of firmware).
[0145] The various embodiments of the systems and techniques described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0146] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.
[0147] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0148] For purposes of providing an interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0149] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0150] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0151] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0152] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A digital human live broadcast method, comprising: Extracting key information according to the background data stream of the live broadcast room; Generating a first prompt word according to the key information; Using a large model to generate a broadcast text of the digital human based on the first prompt word; Generating the voice of the digital human according to the broadcast text; Adjusting the live broadcast actions of the digital human according to the voice of the digital human.
2. The method according to claim 1, wherein, Extracting key information according to the background data stream of the live broadcast room, including: Identifying the voice and / or frame images of the background data stream of the live broadcast room to obtain a live broadcast content text; Extracting the key information from the live broadcast content text.
3. The method according to claim 2, the method further comprising one or more of the following steps: In the case where the key information cannot be extracted from the background data stream, using the large model to generate the broadcast text based on a preset second prompt word; In the case where the key information in the background data stream cannot generate the broadcast text that meets the broadcast requirements, using the large model to generate the broadcast text based on a preset second prompt word; In the case where a preset trigger condition is met, using the large model to generate the broadcast text based on a preset second prompt word.
4. The method according to any one of claims 1 to 3, wherein Adjusting the live broadcast actions of the digital human according to the voice of the digital human, including: Obtaining the voice characteristics according to the voice of the digital human; Obtaining the lip shape parameters of the digital human corresponding to the voice characteristics; Adjusting the lip movement of the digital human according to the voice characteristics and the corresponding lip shape parameters of the digital human.
5. The method according to any one of claims 1 to 4, the method further comprising: Obtaining the background data stream of the live broadcast room; Overlaying the picture and voice of the digital human in the background data stream.
6. The method according to any one of claims 1 to 5, the method further comprising: Overlaying additional materials of the live broadcast room in the background data stream.
7. The method according to any one of claims 1 to 6, the method further comprising: Providing a prompt word setting interface in the live broadcast client; wherein, the content allowed to be provided by the user in the prompt word setting interface includes one or more of the following: title, focus hotspots, expected effects, live broadcast purposes, and content to be summarized; Generating a second prompt word in response to an editing instruction for the prompt word setting interface.
8. The method according to any one of claims 1 to 7, the method further comprising: Recording a human voice, which is used to support converting target text into the voice of the digital human using TTS technology; Recording a human video, which is used to generate the appearance of the digital human; Wherein, the voice of the digital human is used to adjust the lip movement in the appearance of the digital human.
9. A digital human live broadcast device, comprising: An extraction module for extracting key information according to the background data stream of the live broadcast room; A first prompt word generation module for generating a first prompt word according to the key information; A first broadcast text generation module for using a large model to generate a broadcast text of the digital human based on the first prompt word; A voice generation module for generating the voice of the digital human according to the broadcast text; An adjustment module, configured to adjust the live broadcast actions of the digital human according to the voice of the digital human.
10. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-8.
11. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.
12. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-8.