Generation method, device and system for digital human live broadcast materials and electronic equipment
Through multi-modal understanding of object information data, live copy is generated and converted into voice data and timestamp matching information, combined with lip-driven video materials, the problem of complex preparation and automatic switching of digital live broadcast materials is solved, and automated material switching and synchronous display is realized, reducing labor costs.
Patent Information
- Application Number
- CN202510700035.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-08-29
AI Technical Summary
In the prior art, the preparation process of digital live broadcast materials is complicated and the operation is cumbersome, and the automatic switching of materials cannot be achieved, resulting in high labor costs.
Through multi-modal understanding of object information data, live copy is generated and converted into voice data and timestamp matching information, combined with lip-driven video material, the automatic switching and synchronous display of the material is realized.
The preparation process of digital live broadcast materials is simplified, the automatic switching of materials is realized, labor costs are reduced, and the degree of automation of live broadcasts is improved.
Smart Images

Figure CN120568086A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and more particularly to a method, device system, electronic device, computer-readable storage medium, and computer program product for generating digital human live broadcast materials. The present application also relates to a method for rendering digital human live broadcast images. Background Art
[0002] With the rapid development of internet technology, live streaming has been applied in a variety of fields and scenarios. Digital human live streaming, a form of live streaming where digital humans replace real people, reduces labor costs while also providing greater controllability for live streaming technology.
[0003] However, in the existing technology, the preparation process of digital human live broadcast materials is complicated and the operation is cumbersome, resulting in a lot of time and energy required to prepare materials in the early stage, and the corresponding materials cannot be automatically switched during the live broadcast. Summary of the Invention
[0004] The present application provides a method, device, system, electronic device, computer-readable storage medium, and computer program product for generating digital human live broadcast materials, which can simplify the preparation process of digital human live broadcast materials and realize automatic switching of materials during live broadcast.
[0005] In a first aspect, the present application provides a method for generating digital human live broadcast materials, comprising:
[0006] According to the digital human live broadcast material generation request, object information data corresponding to the digital human live broadcast material is obtained, where the object information data includes at least one of text, image or video;
[0007] Perform multimodal understanding on the object information data and generate live broadcast copy;
[0008] Matching different parts of the live broadcast copy with corresponding images or videos as first matching information;
[0009] Converting the live broadcast text into voice data, and generating matching information between the live broadcast text and the timestamp as second matching information;
[0010] Generating lip-driven video material from the voice data;
[0011] The lip-driven video material, the first matching information and the second matching information are used as materials for digital human live broadcast.
[0012] In a second aspect, the present application provides a method for rendering a digital human live broadcast image, comprising:
[0013] Obtaining object information data, lip-driven video material, and a timestamp, and rendering the object information data and the lip-driven video material according to the timestamp to generate a first digital human live broadcast screen, wherein the first digital human live broadcast screen is used to display the object information data in real time;
[0014] The digital human image is rendered according to the size and streaming layout of the product material to obtain an adapted digital human image.
[0015] In a third aspect, the present application provides a device for generating digital human live broadcast materials, comprising:
[0016] An object information data acquisition unit is configured to acquire object information data corresponding to the digital human live broadcast material according to a digital human live broadcast material generation request, wherein the object information data includes at least one of text, image or video;
[0017] A live broadcast copy generating unit, configured to perform multimodal understanding on the object information data and generate a live broadcast copy;
[0018] A first matching information generating unit is configured to match different parts of the live broadcast text with corresponding images or videos as first matching information;
[0019] A second matching information generating unit, configured to convert the live broadcast text into voice data, and generate matching information between the live broadcast text and a timestamp as second matching information;
[0020] a lip-actuated video material generating unit, configured to generate lip-actuated video material from the voice data;
[0021] The lip-driven video material, the first matching information and the second matching information are used as digital human live broadcast materials.
[0022] Fourthly, this application provides a system for generating digital human live broadcast materials, including a server, a cloud, a large language model, and a text-to-speech module:
[0023] The server is used to send a request for generating digital human live broadcast materials; match different parts of the live broadcast text with corresponding images or videos as first matching information;
[0024] The cloud is used to generate lip-driven video material based on the voice data, obtain object information data and a timestamp, and render the object information data and the lip-driven video material according to the timestamp to generate a first digital human live broadcast screen, which is used to display the object information data in real time; the digital human screen is rendered according to the size and push layout of the product material to obtain an adapted digital human screen;
[0025] The large language model is used to obtain object information data corresponding to the digital human live broadcast material according to the digital human live broadcast material generation request, and the object information data includes at least one of text, image or video;
[0026] The text-to-speech module is used to convert the live broadcast text into voice data and generate matching information between the live broadcast text and the timestamp as second matching information.
[0027] In a fifth aspect, the present application provides an electronic device, characterized in that it includes: a processor, a memory, and computer program instructions stored in the memory and executable on the processor; when the processor executes the computer program instructions, it implements any one of the methods described in the first aspect.
[0028] In a sixth aspect, the present application provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer execution instructions, and when the computer execution instructions are executed by a processor, they are used to implement any one of the methods described in the first aspect.
[0029] In a seventh aspect, the present application provides a computer program product, comprising a computer program, which implements the method described in any one of the first aspects when executed by a processor.
[0030] Compared with the prior art, this application has the following advantages:
[0031] The method for generating digital human live broadcast material provided in the embodiment of the present application obtains object information data corresponding to the digital human live broadcast material based on a digital human live broadcast material generation request, performs multimodal understanding on different forms of content such as text, images, and videos included in the object information data, and generates corresponding live broadcast copy based on the results of the multimodal understanding; in order to match the live broadcast copy with other materials, the live broadcast copy is divided into different parts, and the live broadcast copy of each different part is matched with the corresponding image and video to form first matching information; the live broadcast copy is content that may be presented through voice during the live broadcast process, so the live broadcast copy is converted into voice data, and matching information of the live broadcast copy and the timestamp is generated as second matching information; in order to match the voice data with the digital human's lip movements, the digital human's lip video material is generated based on the voice data.
[0032] The method for generating digital human live broadcast materials provided in the embodiment of the present application comprehensively and accurately analyzes various types of information of the objects corresponding to the digital human live broadcast materials through multimodal understanding of object information data, and the live broadcast copy generated based on the analysis and understanding results is also richer and more accurate; by dividing the live broadcast copy into different parts and matching them with corresponding images or videos, each part corresponds to an image or video that is most consistent with its content as the first matching information; the live broadcast copy is also converted into voice data, and a timestamp is generated at the same time. There is second matching information between the timestamp and the live broadcast copy; combined with the first matching information and the second matching information, the voice data can be matched with the corresponding image, video, etc., and contains the timestamp information; it can be seen that the first matching information and the second matching information are conducive to the automatic switching of subsequent digital human live broadcast materials; the digital human's lip-driven video material is generated based on the voice data, and the lip-driven video material is combined with the first matching information and the second matching information to jointly generate the digital human live broadcast material.
[0033] The rendering method for a digital human live broadcast screen provided in an embodiment of the present application obtains live broadcast text, object information data, lip-actuated video material, and a timestamp, renders the object information data and lip-actuated data based on the obtained content, and obtains a first digital human screen. The first digital human screen is then rendered based on the size and resolution of the object information data to obtain an adapted digital human screen. The digital human live broadcast screen provided in an embodiment of the present application synchronously renders the digital human's lip-actuated data and the object information data to be displayed during the live broadcast using a timestamp. This allows the digital human's voice, lip movements, and object information data to be displayed in the live broadcast screen seen by the audience to be presented synchronously, achieving the integration of text, screen, and material. Since object information data is typically displayed as a live broadcast background, while the digital human's movements are displayed through a foreground video, the rendering method of the present application can achieve automatic switching between the front video and the back background, improving the richness of the digital human live broadcast screen and information. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 This is a schematic diagram of an application scenario of the method for generating digital human live broadcast materials provided by this application;
[0035] Figure 2 This is a flow chart of a method for generating digital human live broadcast materials provided in an embodiment of the present application;
[0036] Figure 3 This is a schematic diagram of a process for generating digital human live broadcast materials provided by an embodiment of the present application;
[0037] Figure 4 This is a flow chart of a method for rendering a digital human live broadcast screen provided in an embodiment of the present application;
[0038] Figure 5 This is a block diagram of a device for generating digital human live broadcast materials provided by an embodiment of the present application;
[0039] Figure 6 This is a system structure diagram for generating digital human live broadcast materials provided by an embodiment of the present application;
[0040] Figure 7 This is a structural block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0041] In order to enable those skilled in the art to better understand the technical solutions of this application, the following clearly and completely describes this application in conjunction with the drawings in the embodiments of this application. However, this application can be implemented in many other ways different from the following description. Therefore, based on the embodiments provided in this application, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of this application.
[0042] It should be noted that the terms "first", "source domain", "third", etc. in the claims, description and drawings of the present application are used to distinguish similar objects and are not used to describe a specific order or sequence. The data used in this way are interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including", "having" and their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0043] Live streaming, as an innovative way to disseminate information, is very popular among consumers. In real life, live streaming has become the main form of product promotion and introduction for many electronic product service platforms. Electronic product service platforms use live streaming to display products to viewers. The host and users interact with each other through likes, follows, comments, and bullet screens to learn about merchants (stores) and products, which also increases the purchase conversion rate. Electronic product service platforms can provide good support for live streaming. Electronic product service platforms can support live streaming platforms. Live streaming platforms are also electronic product service platforms that can recommend products. Therefore, unless otherwise specified, live streaming platforms and electronic product service platforms (referred to as e-commerce platforms) have the same meaning below. Currently, most live broadcasts are conducted by real-person hosts throughout the entire process, which is a relatively common form of live streaming. Because traditional live streaming methods require a large amount of manpower investment, and the working time and physical strength of real-person hosts in traditional live streaming forms are limited, it is impossible to adjust the live streaming time throughout the day or at will. Therefore, the existing technology proposes digital humans to replace real-person hosts for live streaming.
[0044] Digital human live broadcast, also known as virtual anchor or digital human live broadcast, is a new type of live broadcast that combines virtual humans (i.e. digital humans, virtual characters or virtual anchors) with live video. Digital human live broadcast can, to a certain extent, overcome the problem of high labor costs in traditional live broadcast. Currently, digital human live broadcast in existing technologies cannot achieve the live broadcast effect of automatically acquiring and switching materials. In order to solve the above-mentioned problems existing in digital human live broadcast, this application proposes a method for generating digital human live broadcast materials, further reducing the labor costs of live broadcast. This application proposes a live broadcast data processing method that automatically acquires object information data, automatically generates live broadcast copy, and automatically switches live broadcast materials. The technical solution reduces labor costs in multiple aspects.
[0045] In order to facilitate understanding of the method embodiments of the present application, the application scenarios of the embodiments of the present application are introduced. Figure 1 , Figure 1 Schematic diagram of an application scenario of the method for generating live broadcast materials for a digital human provided in an embodiment of the present application. The embodiment of the present application can be applied to a live broadcast terminal 101, which can be a mobile phone, tablet computer, wearable device, laptop computer, desktop computer, or other device capable of configuring a digital human live broadcast.
[0046] In this embodiment, merchants can input product information, such as product identification or product link, on the live broadcast terminal 101. The live broadcast terminal 101 receives the data input by the merchant, processes the data, obtains the generated digital human live broadcast material, and displays the digital human live broadcast material on the live broadcast terminal 101 for the merchant to preview. The merchant can adjust and modify the input data or preview content, and can initiate secondary data processing to regenerate the digital human live broadcast material.
[0047] Example 1
[0048] The first embodiment of this application provides a method for generating digital human live broadcast materials. Figure 2 , Figure 2 This is a flow chart of a method for generating digital human live broadcast materials provided in an embodiment of the present application.
[0049] S201: According to a request for generating a digital human live broadcast material, object information data corresponding to the digital human live broadcast material is obtained, where the object information data includes at least one of text, image or video.
[0050] The digital human live broadcast material generation request in this step carries object information. For example, if a certain product is to be explained, the object information is product information, including product identification (product ID) or product link. The basic information of the product can be uniquely determined through the product information.
[0051] The server generates a request for generating live streaming materials based on product information. Specifically, merchants can enter the product information they want to broadcast on the live streaming console. The server receives the product information and generates a request for generating live streaming materials based on it. The merchant can be the user who uploaded the product information on the live streaming platform. The server then sends the request to the LLM module.
[0052] It is understandable that the LLM (Large Language Model) module is a large language model module, on which a large language model is deployed. A large language model refers to a pre-trained model based on deep learning technology, especially in the field of natural language processing, with a large number of parameters and the ability to handle complex language tasks. By training a large amount of text data, the large language model can generate, understand, translate, answer questions, summarize, write and other language-related tasks. In the first embodiment of the present application, the large language model is mainly used to generate live copy.
[0053] The LLM module can obtain corresponding object information data according to the generation request of the digital human live broadcast material. The object information data can usually be obtained from the platform's internal database, which accumulates information data related to the object in the historical process.
[0054] S202: Perform multimodal understanding on the object information data and generate live broadcast copy.
[0055] This step is used to perform multimodal understanding on the object information data after the LLM module receives the object information data, input the results of the multimodal understanding into the LLM module, and further generate the live broadcast copy.
[0056] The object information data can reflect the detailed information of the product from different angles, collect object information data in different forms, and screen and analyze the object information data to achieve a comprehensive and in-depth understanding of the product, so as to accurately grasp the characteristics of the product. It can be seen that the multimodal understanding of the object information data helps to generate more accurate live broadcast copy that conforms to the characteristics of the product itself. In the fields of e-commerce, advertising, social media, etc., the multimodal understanding of object information data can improve the accuracy and efficiency of product information analysis.
[0057] Understanding images can be accomplished by screening for eligible images and extracting key image information and image categories. Specifically, images are first converted into text using optical character recognition (OCR). This textual information is then used to analyze and interpret the images. This textual information determines whether the images are eligible for selection. Eligible images are then input into the Live Streaming Manager (LLM) as eligible images for livestream copy generation. However, since the OCR textual information for most images is complex, abstract extraction is required before entering the LLM. Furthermore, since a product often has a large number of associated images, it is essential to select images based on the desired product information. Therefore, before entering the LLM, images are labeled with their categories so that a screening strategy can be designed based on the image classification results. Based on the image classification, the most relevant images are selected, and the corresponding textual information is input into the LLM. Based on the extracted images and corresponding textual information, the LLM module determines the selling point information that best matches the product, thereby generating the explanatory speech as the first livestream copy.
[0058] Understanding the video can be accomplished in a step-by-step process. First, the video is split into different frames, each frame being treated as an image. Similar processing is performed based on the aforementioned understanding of images, ultimately outputting the video's category information and the features corresponding to the textual information within it. Based on the textual features within the category information, the selling point information that best matches the product is determined, generating explanation scripts as the second live broadcast copy.
[0059] LLM not only processes language but also combines visual information for deep reasoning, generating understanding, summaries, and analysis of videos. However, to understand content from videos using LLM, it is usually necessary to first extract text and visual information from the video and then combine these multimodal data for analysis. Therefore, the acquisition of the second live broadcast copy can also be achieved through the following methods:
[0060] For videos, key frames, audio, subtitles, and other potentially informative elements can be extracted. Computer vision techniques (such as convolutional neural networks) can then be used to identify objects, scenes, and actions within the video. If audio is involved, speech recognition can be used to convert this information into text. Automatic speech recognition can also be used to convert subtitles or audio within the video into text, providing key information. Image processing and video analysis techniques can then be used to understand objects, people, and actions within the video frames. Image features can also be combined with language models to understand the context and situation within the video. Furthermore, because videos are dynamic, temporal information is crucial. Recurrent neural networks and long short-term memory networks can be used to process this time series data, capturing the sequence and causal relationships of events within the video. Once the audio is converted to text and combined with visual information, LLMs can be used to perform deep text understanding. For example, LLMs can analyze dialogue content, character interactions, and context inferred from visual content. LLMs can extract key information from multimodal information to summarize and interpret the video content. LLMs can also provide more context-aware summaries of the content through natural language processing. LLM can also help analyze emotions in videos, identify characters' emotional expressions, and understand emotional changes and their role in the story. In the above processing, the model for multimodal understanding can adopt the Transformer architecture or the CLIP model. The Transformer is a deep learning model architecture based on the self-attention mechanism. It can process sequential data and is widely used in natural language processing tasks. Compared with traditional recurrent neural networks and long short-term memory networks, the Transformer has significant advantages in processing long-range dependencies. Commonly used Transformer architectures include models such as GPT (Generative Pre-trained Transformer) and BERT (Bidirectional Encoder Representation Transformer), which can be extended to process multimodal data. The CLIP (Contrastive Language-Image Pre-training) model can be trained by combining text and image content, thereby achieving the ability to extract information from images and videos and integrate it with language models. In fact, the Transformer architecture combined with vision (such as CLIP) and language models can achieve good results in vision-language tasks. The specific implementation of multimodal data processing of the material can be adjusted according to actual conditions and is not limited by this application.
[0061] It is worth noting that when the object information data corresponding to a certain product includes content in multiple different forms, the first live broadcast copy and the second live broadcast copy can be analyzed and integrated again through LLM to generate a whole as a live broadcast copy. Alternatively, all object information data can be screened separately, and the text information of the different object information data after screening can be obtained separately. The text information can then be processed uniformly to generate a live broadcast copy through LLM. In the process of fusing the first live broadcast copy and the second live broadcast copy, a Transformer architecture or CLIP similar to the previous paragraph can be used. Through multimodal models and spatiotemporal information processing technology, richer and more accurate information can be comprehensively analyzed from static images and dynamic videos to generate an overall understanding and summary. The specific operation process can refer to the above content. In addition, in the process of generating live broadcast copy based on object information data, the judgment and screening conditions of the object information data can be added or adjusted according to actual needs, so that the screened content continues to be close to actual needs. As a supplement, the information obtained from the object information data can actually also directly include text. The text is usually a detailed introduction related to the product or information about similar categories. Therefore, the text can also be screened and analyzed, and used as a live broadcast copy to be integrated with the first live broadcast copy and the second live broadcast copy to generate an overall understanding and summary. Finally, LLM outputs a complete live broadcast copy.
[0062] S203: Match different parts of the live broadcast text with corresponding images or videos as first matching information.
[0063] The live broadcast copy obtained according to the above steps may be quite long. In addition, during the playback of a live broadcast copy, multiple text, videos, images, and other materials may be matched and displayed simultaneously. Therefore, in this step, the live broadcast copy is divided into different parts to facilitate matching different parts of the live broadcast copy with corresponding images and videos. As can be seen, this step can reduce the size and complexity of data processing during operation, thereby achieving faster processing speed.
[0064] S204: Convert the live broadcast text into voice data, and generate matching information between the live broadcast text and the timestamp as second matching information.
[0065] During the live broadcast, the live copy is presented in audio form. In this step, the live copy is converted into voice data, which can be played to the audience as audio during the live broadcast. The live copy is usually in text form, and the process of generating speech from text includes the timestamp information of the speech output. The matching information between the voice data, the live copy, and the timestamp is used as the second matching information. After generating the second matching information, the TTS reasoning module stores the second matching information for subsequent timeline matching and material switching. It can be understood that TTS (Text-to-Speech) is a technology that converts text into speech, which allows a computer or device to read the input text aloud through synthesized speech. TTS supports processes such as text analysis, language generation, and speech synthesis. Specifically, it can first analyze the input text to understand the sentence structure, grammar, punctuation, etc., then generate an appropriate language model based on the text to determine how to express the speech. Finally, the synthesizer converts this language data into natural speech and plays it out. In the embodiment of the present application, TTS supports the production of audio data based on the segmented copy content and outputs the start and end timestamps of each segment for material switching matching.
[0066] S205: Generate lip-driven video material from the voice data.
[0067] Binding the voice data to the digital human can generate lip-driven video material of the digital human. The lip-driven video material includes the digital human's lip shape video and the voice data emitted synchronously by the digital human. Generating the lip-driven video material from the voice data includes: using the digital human's lip-driven inference capability to generate the digital human's lip-driven video material based on the voice data and the timestamp.
[0068] The lip-driven video material, the first matching information, and the second matching information obtained in the above steps can be presented as digital human live broadcast material.
[0069] In a real-time manner, the matching of different parts of the live broadcast text with corresponding images or videos as first matching information includes: segmenting the live broadcast text to obtain segmented live broadcast text; matching the segmented live broadcast text with the image or video and generating a correspondence between the live broadcast text and the object information data; and using the correspondence obtained from the matching as the first matching information. The segmentation process can assist in judgment based on punctuation marks in the live broadcast text, while extracting feature information such as live broadcast text keywords to divide the live broadcast text into different parts. For example, the content of the live broadcast copy is "A solid wooden table, made mainly of oak. Oak is known for its durability and beautiful grain. The natural grain of oak is uniform and unique, with a warm and natural color, showing a pattern of light and dark wood grain, making it an ideal choice for making high-end furniture. The tabletop is treated with a high-quality environmentally friendly paint finish, which is wear-resistant, stain-resistant, waterproof and moisture-proof, effectively extending the service life of the tabletop. The legs are designed with a combination of oak and beech, which has greater stability and load-bearing capacity. The mortise and tenon structure technology is used to firmly connect the legs to the tabletop, reducing the use of screws, ensuring a more stable table and avoiding the loose joints common in modern furniture. This solid wooden table provides users with a great experience in terms of texture, load-bearing capacity, and design details, making it a core piece of furniture in a home or office space." The live broadcast copy can be roughly divided into the following different parts based on semantics: A. A solid wooden table, made mainly of oak. B. Oak is known for its durability and beautiful grain. The natural grain of oak is uniform and unique, with a warm and natural color, showing a pattern of light and dark wood grain, making it an ideal choice for making high-end furniture. C. The tabletop is treated with a high-quality, environmentally friendly paint finish that is wear-resistant, stain-resistant, waterproof, and moisture-resistant, effectively extending its lifespan. D. The legs are a combination of oak and beech, providing greater stability and load-bearing capacity. E. The mortise and tenon joints securely connect the legs to the tabletop, reducing the need for screws, ensuring a more stable table and avoiding the loose joints common in modern furniture. F. This solid wood table offers a superior user experience in terms of texture, load-bearing capacity, and design details, making it a perfect centerpiece for any home or office space. The images or videos included in the acquired object information data may include the following: 1. A three-dimensional display video of a table, 2. A picture of unfelled oak, 3. A picture of an oak tabletop, 4. A video of the high-quality environmentally friendly paint treatment process, 5. A picture of the tabletop treated with environmentally friendly paint, 6. A picture of a tabletop waterproof test, 7. An overall picture of the table legs, 8. A picture of the joint between the oak and beech wood of the table legs, 9. A picture showing the mortise and tenon structure, 10. Four pictures of the table placed in a bedroom, a study, an office, and a library respectively.The first matching information contains the matching relationship between the different parts of the above-mentioned live copy and the corresponding image or video: copy A corresponds to 1, copy B corresponds to 2, 3, copy C corresponds to 4, 5, 6, copy D corresponds to 7, 8, copy E corresponds to 9, and copy F corresponds to 10. The above explains the content contained in the first matching information in the form of examples. Of course, different parts of the live copy can also be split into different parts according to different conditions according to actual needs. Since this step splits the live copy into multiple smaller parts, the understanding of semantics is more detailed, and therefore, more accurate first matching information can be obtained.
[0070] The segmentation processing of the live broadcast copy to obtain the segmented live broadcast copy includes: segmenting the live broadcast copy according to semantics, and formulating a copy theme for each segmented live broadcast copy; obtaining the matching relationship between the segmented live broadcast copy and the copy theme, and storing the matching relationship. In order to match the content explained by the live broadcast copy with the content in the object information data, the live broadcast copy content is segmented according to semantics, and a corresponding theme is formulated for each segment of the live broadcast copy. During preview, the theme can provide a concise and concise copy theme for merchants on the live broadcast end, which helps merchants to quickly understand and review the live broadcast materials. Continuing with the live broadcast copy listed above, the theme of segment A is the table as a whole, and the theme of copy B is that oak is the main material of the table... The theme can actually also be used to obtain more accurate first matching information.
[0071] In the aforementioned steps, the LLM module generates live broadcast copy corresponding to the product information based on its understanding. The LLM module also segments the live broadcast copy, dividing it into multiple different parts based on semantics, and assigning a corresponding theme to each part. The theme is convenient for merchants to quickly understand when previewing. The LLM module also matches different parts of the live broadcast copy with pictures and videos, so that different parts of the live broadcast copy correspond to corresponding pictures and videos, achieving the effect of synchronously displaying pictures and videos when the live broadcast copy is played.
[0072] At this point, the live broadcast content, including live broadcast images, and videos, has been obtained, as well as the matching relationship between the live broadcast text and the corresponding live broadcast images and videos. The matching dimension at this time is the segmentation of the live broadcast text, and which live broadcast images and videos correspond to the first segment of the live broadcast text, and which image and video correspond to the second segment. The matching relationship between the segmented live broadcast text and the live broadcast images and videos is the first matching information.
[0073] Converting the live copywriting into voice data and generating the matching information between the live copywriting and the timestamps as the second matching information includes: generating timestamps at the character dimension based on the voice data and obtaining the corresponding relationship between the live copywriting and the timestamps; using the corresponding relationship between the timestamps and the live copywriting as the second matching information. For example, when the text of the live copywriting is "Hello", the corresponding voice data and timestamps may be that "Hello" is played at 0.1s, "Hello" is played at 0.2s, and "Hello" is played from 0.3s to 0.5s. It can be seen that during the generation of the voice data, the corresponding timestamps are generated synchronously, and the voice data and the live copywriting are relatively supplied. Therefore, there is also a corresponding relationship between the live copywriting and the timestamps. The copywriting "Hello" corresponds to 0.1s, the copywriting "Hello" corresponds to 0.2s, and the copywriting "Hello" corresponds to 0.3s, 0.4s, and 0.5s. The matching information among the voice data, the live copywriting, and the timestamps is used as the second matching information. To improve the matching accuracy, the precision of the timestamps can be increased during the generation of the voice data. For example, timestamps at the syllable dimension or the phoneme dimension can be output. Among them, outputting timestamps at the syllable dimension means that if there are multi-syllable words in the text, the system will generate timestamps for each syllable, which helps for more refined control of speech synthesis (such as adjusting the speech rate or optimizing the accent). A phoneme is the smallest pronunciation unit in a language. Some advanced TTS systems (especially neural network-based TTS models such as Tacotron or WaveNet) can generate phoneme-level timestamps. This method is very precise, but it may increase the complexity of the generation process. Therefore, to improve the matching accuracy without being overly complex, timestamps at the character dimension or the syllable dimension are usually adopted to provide support for subsequent material matching. When the TTS module generates audio, the time alignment technology can automatically mark the start and end times of each syllable, word, or phoneme.
[0074] In one implementation, the first matching information and the second matching information are matched to obtain the corresponding relationship between the live copywriting, the voice data, the corresponding image or video, and the timestamps, and digital human live materials are generated according to the corresponding relationship, including steps S301 - S304. Please refer to Figure 3 , Figure 3 is a schematic flowchart of an example of generating digital human live materials provided by an embodiment of the present application:
[0075] S301: Obtain the corresponding relationship between the live copywriting and the corresponding image or video in the first matching information;
[0076] S302: Obtain the corresponding relationship between the live copywriting and the timestamps in the second matching information;
[0077] S303: Acquire the correspondence between the lip-actuated video material and the voice data;
[0078] S304: Match the object information data, lip driving video material and the timestamp to generate digital human live broadcast material.
[0079] After obtaining the first matching information in S201, the first matching information is formatted and uniformly converted into JSON format for storage. Before S301, the information formatted into JSON format is parsed, and then the first matching information stored therein is obtained, thereby extracting the corresponding relationship between the live broadcast copy and the image or video.
[0080] JSON is a lightweight, structured data format with a clear key-value structure. This makes stored data not only easy to read and understand, but also convenient for exchanging data across different application systems. When LLM generates content, storing it in JSON format allows the content to be organized according to a specific structure, making it easier to extract, query, and process the data. Using JSON can help unify the storage format and ensure structural consistency across all generated content. JSON also offers advantages such as strong cross-platform compatibility and scalability. Therefore, using JSON format for data transmission improves data transmission efficiency while ensuring data accuracy.
[0081] After obtaining the second matching information in S202, the corresponding relationship in the second matching information is stored for subsequent timeline matching and material switching. In S302, the corresponding relationship between the live broadcast copy and the timestamp is obtained from the second matching information.
[0082] Through S301-S304, the object information data, lip-driven video material and timestamp matching can be achieved to generate digital human live broadcast material. For example, continuing the above example about the solid wood table, the live broadcast copy C, the desktop is treated with high-quality environmentally friendly paint, which is wear-resistant and stain-resistant, and waterproof and moisture-proof, effectively extending the service life of the desktop. The images or videos included in the corresponding object information data include: 4, a video of the high-quality environmentally friendly paint treatment process, 5, a picture of the desktop treated with environmentally friendly paint, 6, a picture of the desktop waterproof test. The first matching information contains the matching relationship between the different parts of the above live broadcast copy and the corresponding images or videos. Copy C corresponds to 4, 5, and 6. The second matching information includes the process of playing the live broadcast copy C: 0.1s-0.5s corresponds to "the desktop is treated with high-quality environmentally friendly paint", 0.6s-1s corresponds to "wear-resistant and stain-resistant", and 1s-2s corresponds to "and waterproof and moisture-proof, effectively extending the service life of the desktop". Therefore, through rendering, the final digital human live broadcast material obtained is that the digital human's lips will change synchronously with the voice data while outputting voice data, and as the semantics of the text in the voice data change, the displayed pictures, videos, etc. will also automatically switch, thereby realizing the automatic switching of images, videos and other materials according to the live broadcast text.
[0083] In one embodiment, the digital human live broadcast material generation request carries product information, which includes at least one of a product identifier and a product link. The product identifier or ID, along with the product link, can be used to obtain basic product information. This basic product information can then be used to obtain more detailed information related to the product stored in the content database. This detailed product information can then be used to generate richer and more accurate live broadcast content for the product.
[0084] The object information data corresponding to the digital human live broadcast material generation request is obtained, including: obtaining product information according to the digital human live broadcast material generation request, obtaining product material information corresponding to the product information in the content database according to the product information; using the product material information as the object information data; wherein the content database stores the product material information corresponding to the product information and has the ability to generate product material information. The product information included in the digital human live broadcast material generation request can obtain the corresponding product details and product materials related to the product information in the database, and a material generation module can also be deployed in the content database to supplement the product materials. For example, when the number of original product images in the content database corresponding to the product information is insufficient, or the content fit of the product images is low, the material generation module can generate new images and other materials that are more closely aligned with the product information based on the product information.
[0085] The method for generating live broadcast copy also includes: parsing the object information data from at least one of the domain, rights and interests information, product characteristics, and type dimensions, and generating live broadcast copy based on the parsed content. Since the object information data may contain materials in different forms such as text, images, and videos, different forms require different parsing methods, so these materials are multimodally understood. On this basis, in the process of analyzing the object information data, it is also necessary to parse from different angles such as the domain, rights and interests information, product characteristics, and type dimensions corresponding to the product information. For example, if the product belongs to the furniture category, the relevant preferences and hot products in the furniture field can be analyzed, and the common points between the product and the hot products can be extracted and analyzed. It can also be combined with information such as the current promotional activities of the product to explain the product itself to the audience while providing corresponding promotional activities.
[0086] The method for generating digital human live broadcast materials provided in the embodiment of the present application can automatically obtain corresponding materials through the product information uploaded by the user, significantly reducing the cost of material preparation and simplifying the merchant's live broadcast process. At the same time, because the technical solution of the embodiment of the present application provides matching information between timestamps and live broadcast copy, as well as matching information between live broadcast copy and live broadcast materials, it is possible to automatically switch the materials in the live broadcast at the appropriate time. Furthermore, it is possible to display the material information of the product in real time during the product explanation process. It can be seen that the technical solution provided by the embodiment of the present application provides a more convenient live broadcast operation method, so that the obtained digital human live broadcast materials have richer product details and more automated material switching effects, thereby reducing the labor cost and time cost required for live broadcast.
[0087] Example 2
[0088] The second embodiment of this application provides a method for rendering a digital human live broadcast screen. For details on the part involving the digital human live broadcast material, please refer to the first embodiment. Figure 4 , Figure 4 A flowchart of a method for rendering a digital human live broadcast screen provided in an embodiment of the present application.
[0089] The method for rendering a digital human live broadcast screen includes:
[0090] S401: Obtain live broadcast text, object information data, lip-actuated video material, and timestamps, and render the object information data and lip-actuated video material to generate a first digital human live broadcast image. Timestamps not only provide precise control but also enable dynamic adjustment, integration, and presentation of these materials based on actual needs.
[0091] S402: Rendering the first digital human live broadcast picture according to the size and resolution layout of the object information data to obtain an adapted digital human live broadcast picture.
[0092] In one embodiment, the method for rendering a digital human live broadcast screen further includes: adding switching and special effects rendering to at least one first digital human live broadcast screen to generate a second digital human live broadcast screen; and using the second digital human live broadcast screen as the digital human live broadcast screen. The digital human live broadcast material generated in Example 1 can be understood as segmented material, and during the live broadcast, these segmented materials can be connected and presented to the audience as a whole. Therefore, switching and special effects can be added between each segment of material for rendering, and the rendered live broadcast material is played as a whole live broadcast screen. The rendering of the object information data and the lip-driven video material includes: performing nonlinear rendering on the object information data and the lip-driven video material. Specifically, nonlinear rendering is performed on the first matching information, the second matching information, and the lip-driven video material obtained in Example 1 to achieve synchronous display of the material and the live broadcast copy.
[0093] Non-linear rendering refers to the dynamic arrangement and adjustment of materials on the timeline based on timestamps, rather than playing or displaying them in a fixed order. By synchronizing the timestamps of audio, video, and image materials, materials can be flexibly combined and switched as needed, creating a variety of presentation effects. To perform non-linear rendering of audio, video, and image materials at the same timestamp, the timelines of the different materials must be controlled to ensure that they start, end, or switch synchronously at the same timestamp.
[0094] For example, in video and audio editing, content can be played back in a non-sequential manner. Instead, it can be flexibly controlled based on timestamps, and even the order, speed, acceleration, or reverse playback can be changed. In non-linear video rendering, by adjusting timestamps, video materials can dynamically switch lenses as needed. For example, multiple lenses can be set to play the same scene from different angles simultaneously, or a different scene can be quickly switched to a different scene at a certain point in time to achieve non-linear jumps in time and space. Videos and pictures can be mixed or transitioned at the same timestamp. For example, during video playback, certain picture materials can gradually change from transparent to fully displayed, creating a visual gradient effect. Special effects based on timestamps can also be added to present gradually changing scenes in the video. Specifically, through timestamps, the start time, duration, and end time of special effects can be precisely controlled to achieve complex effects such as explosions, smoke, and light and shadow changes.
[0095] For example, the display order of image assets can be determined by timestamp. For example, a series of image assets can be displayed one by one according to their timestamps, and the timestamp can be used to precisely control the display duration of each image. Images can also transition with timestamp changes through special effects such as gradients, scaling, and rotation, or switch to a new image at a certain point in time. In other words, by adjusting the precise control of timestamps, these changes can be ensured to match other assets (such as audio or video).
[0096] This rendering method allows different types of materials (voice, video, pictures, etc.) to be processed collaboratively according to timestamps within the same time frame, achieving dynamic changes, interactivity, and rich performance effects.
[0097] Example 3
[0098] The third embodiment of the present application provides a device for generating digital human live broadcast materials. Figure 5 , Figure 5 This is a block diagram of a device for generating digital human live broadcast materials provided by an embodiment of the present application, the device comprising:
[0099] The object information data acquisition unit 501 is used to acquire object information data corresponding to the digital human live broadcast material according to the digital human live broadcast material generation request, wherein the object information data includes at least one of text, image or video;
[0100] A live broadcast copy generating unit 502 is configured to perform multimodal understanding on the object information data and generate a live broadcast copy;
[0101] A first matching information generating unit 503 is configured to match different parts of the live broadcast text with corresponding images or videos as first matching information;
[0102] A second matching information generating unit 504 is configured to convert the live broadcast text into voice data and generate matching information between the live broadcast text and the timestamp as second matching information;
[0103] A lip-actuated video material generating unit 505 is configured to generate lip-actuated video material from the speech data;
[0104] The lip-driven video material, the first matching information and the second matching information are used as digital human live broadcast materials.
[0105] Example 4
[0106] Embodiment 4 of the present application provides a system for generating digital human live broadcast materials, the system including the method for generating digital human live broadcast materials described in Embodiment 1. The following description of the embodiment of the system for generating digital human live broadcast materials is merely illustrative, and the relevant implementation details can be understood by referring to Embodiments 1 and 2.
[0107] Please see Figure 6 , Figure 6 : This is a system structure diagram for generating digital human live broadcast materials provided by an embodiment of the present application. The system for generating digital human live broadcast materials includes a live broadcast terminal 101, a server 602, a cloud 603, an LLM module 604, and a text-to-speech module 605:
[0108] The live broadcast terminal 101 is used by merchants to upload product information, preview digital human live broadcast materials, and perform secondary editing. Through the live broadcast terminal 101, merchants can generate copy based on the uploaded product information with one click. The preview content will also be displayed in the form of segmented explanation copy and corresponding product materials. If the merchant is not satisfied with the results, they can modify, add or delete the explanation copy and materials, or further modify and add product information to obtain more suitable explanation copy and materials. In this case, the product materials displayed in the preview can be the segmented live broadcast copy corresponding to the product information and matching images, videos, and other materials.
[0109] The server 602 is used to send a request for generating digital human live broadcast materials, match different parts of the live broadcast text with corresponding images or videos, and use the matching information as the first matching information. The server is also used to send the request for generating digital human live broadcast materials to the LLM module.
[0110] The cloud 603 is used to generate lip-driven video materials based on voice data, and obtain object information data and timestamps, and render the object information data and the lip-driven video materials according to the timestamp to generate a first digital human live broadcast screen, which is used to display the object information data in real time; according to the size of the product material and the push flow layout, the digital human screen is rendered to obtain an adapted digital human screen; the push flow layout can make the screen seen by the user have a better presentation effect. For example, if the viewing device has a resolution of 1080P or 720P, dynamic adaptation is achieved according to the push flow resolution size and the size of the picture or other materials.
[0111] The LLM module 604 is used to obtain object information data corresponding to the digital human live broadcast material according to the digital human live broadcast material generation request, and the object information data includes at least one of text, image or video.
[0112] The text-to-speech module 605 is used to convert the live broadcast text into voice data and generate matching information between the live broadcast text and the timestamp as second matching information.
[0113] The server is also responsible for issuing product information production instructions and storing the materials and results generated by the cloud, TTS, LLM and lip-driven algorithms. It also uploads the product information uploaded by users to the content database, transcodes videos, and reduces the cost of video storage and playback.
[0114] An AI audio and video cross-platform software development toolkit has been deployed on the cloud. This toolkit can be deployed simultaneously on the cloud and on clients, including live broadcast clients and viewer user clients. The AI audio and video cross-platform software development toolkit primarily includes two modules: a non-linear editing module and a rendering module. The purpose of the non-linear editing module is to support the parallel rendering of multiple materials, including functions such as timelines, timeline matching, and sentence matching. The rendering module primarily implements the rendering and switching of materials, including the ability to preload materials, render and switch images / materials, and add special effects. Basic computing power facilities are also deployed on the cloud, including relevant cloud machines and Linux images. Basic dependencies are deployed within the images for the rapid deployment of algorithms such as lip drive and TTS. This shows that the cloud computing power facilities provide computing power services for upper-level modules, ensuring that the upper-level algorithms have sufficient computing resources.
[0115] In this system, merchants upload product information to the live broadcast end, where the product information may include product ID, product link, and may also include product pictures, videos, etc. After the live broadcast end receives the product information, it is synchronized to the server end. When necessary, the server end converts the format of the product information and triggers the LLM module material generation link to start running. The server end also sends a digital human live broadcast material generation request to the LLM module, and the digital human live broadcast material generation request contains the product ID. LLM returns the matching information of the pictures, videos and texts it generates to the server end, and the server end converts it into JSON format. The server end sends the data sent by LLM to the cloud in JSON format, that is, in Figure 6 The protocol information is sent to the cloud in the form of a live broadcast protocol. Therefore, the protocol information actually includes the text, text segmentation information, and the product image corresponding to the text. The cloud parses the formatted data into the original image, video, and text content and matching information, and sends the segmented live broadcast text to the TTS inference module. The TTS inference module outputs the corresponding segmented audio files and timestamps based on the segmented live broadcast text and returns them to the cloud.
[0116] The cloud renders the digital human live video material based on the digital human's lip-driven video material. As the live foreground video, the live text and live pictures, videos and other materials are non-linearly rendered through the non-linear rendering module and the picture rendering module. Based on the first matching information and the second matching information, it is determined when to display a certain live picture or video. At the same time, it also realizes the corresponding display of a certain picture or video when the voice data of the live text is played to a certain position. Moreover, when it is necessary to highlight a certain picture or video, the key content such as the rear background picture can be enlarged, and the secondary content such as the front video can be reduced, or it can be directly switched to only display the key content such as the rear background picture. In this way, the switching between the front video camera and the rear background is realized, and the display effect of the live material and the intelligent synchronization of the text and the material are optimized.
[0117] The cloud sends the content generated in the above steps to the server, which formats it and sends it to the live broadcast client. The live broadcast client generates product preview information and product explanation preview effects for users to preview and modify. If the user makes secondary edits on the live broadcast client, the edited content is retrieved and processed again through the above steps.
[0118] When the product preview information is returned to the live broadcast end, the rendering of the material and copy can be realized on the live broadcast end. That is, the live broadcast end uses the non-linear rendering module and the screen rendering module to non-linearly render the live broadcast copy and live broadcast pictures, videos and other materials. Based on the first matching information and the second matching information, it determines when to display the corresponding live broadcast picture or video. At the same time, it also realizes the corresponding display of the corresponding picture or video when the voice data of the live broadcast copy is played to the corresponding position. Moreover, when it is necessary to focus on a certain picture or video, it can be achieved by enlarging the key content such as the rear background picture, reducing the secondary content such as the front video, or directly switching to only displaying the key content such as the rear background picture, so as to realize the switching between the front video and the rear background, optimizing the display effect of the live broadcast material and the intelligent synchronization of the copy and material.
[0119] In addition, when the above content is to be presented to the audience client, the rendering of materials and copy can be realized on the audience client, that is, the client uses the non-linear rendering module and the picture rendering module to perform non-linear rendering on the live copy and live pictures, videos and other materials. According to the first matching information and the second matching information, it is determined to display a certain live picture or video at a certain moment. At the same time, it also realizes the corresponding display of a certain picture or video when the voice data of the live copy is played to a certain position.
[0120] The system for generating live broadcast materials for digital humans provided in this application provides a solid technical foundation for copy generation and material matching through the LLM module's capabilities for content understanding, copy generation, and segment summarization. Furthermore, combined with TTS's timestamp matching capabilities, the system integrates materials and copy based on a cloud cluster. The resulting matched materials are rendered and streamed to the audience, allowing them to see the digital human live broadcast automatically acquiring materials, generating copy, and switching between materials.
[0121] Example 5
[0122] The fifth embodiment of the present application also provides an electronic device embodiment corresponding to the method for generating digital human live broadcast materials provided in the first embodiment. The following description of the electronic device embodiment is merely illustrative. The electronic device embodiment is as follows:
[0123] Please see Figure 7 Understand the above electronic devices, Figure 7 Schematic diagram of an electronic device provided in Example 4. The electronic device provided in Example 4 includes: a processor 701, a memory 702, a communication bus 703, and a communication interface 704;
[0124] The memory 702 is used to store computer instructions for data processing. When the computer instructions are read and executed by the processor 701, the following steps are performed:
[0125] According to the digital human live broadcast material generation request, object information data corresponding to the digital human live broadcast material is obtained, where the object information data includes at least one of text, image or video;
[0126] Perform multimodal understanding on the object information data and generate live broadcast copy;
[0127] Matching different parts of the live broadcast copy with corresponding images or videos as first matching information;
[0128] Converting the live broadcast text into voice data, and generating matching information between the live broadcast text and the timestamp as second matching information;
[0129] Generating lip-driven video material from the voice data;
[0130] The lip-driven video material, the first matching information and the second matching information are used as digital human live broadcast materials.
[0131] Example 6
[0132] Embodiment 6 of the present application provides a computer-readable storage medium for implementing the methods described in Embodiments 1 and 2. The computer-readable storage medium embodiment is described briefly; for relevant details, please refer to the corresponding description of the above method embodiments. The embodiments described below are merely illustrative. The computer-readable storage medium stores computer-executable instructions that, when executed by a processor, implement the methods described in Embodiments 1 and 2.
[0133] Example 7
[0134] The seventh embodiment of the present application also provides a computer program product for implementing the methods described in Embodiments 1 and 2. The computer program product embodiments provided in this application are described in a relatively simple manner. For relevant parts, please refer to the corresponding descriptions of the above method embodiments. The embodiments described below are merely illustrative.
[0135] The computer program product provided in this embodiment includes a computer program, which, when executed by a processor, implements the following steps:
[0136] According to the digital human live broadcast material generation request, object information data corresponding to the digital human live broadcast material is obtained, where the object information data includes at least one of text, image or video;
[0137] Perform multimodal understanding on the object information data and generate live broadcast copy;
[0138] Matching different parts of the live broadcast copy with corresponding images or videos as first matching information;
[0139] Converting the live broadcast text into voice data, and generating matching information between the live broadcast text and the timestamp as second matching information;
[0140] Generating lip-driven video material from the voice data;
[0141] The lip-driven video material, the first matching information and the second matching information are used as digital human live broadcast materials.
[0142] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0143] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0144] 1. Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include non-transitory media such as modulated data signals and carrier waves.
[0145] 2. Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0146] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0147] Although the present application is disclosed as above with the preferred embodiments, it is not intended to limit the present application. Any person skilled in the art may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.
Claims
1. A method for generating digital human live broadcast materials, characterized in that: include: According to the digital human live broadcast material generation request, object information data corresponding to the digital human live broadcast material is obtained, where the object information data includes at least one of text, image or video; Perform multimodal understanding on the object information data and generate live broadcast copy; Matching different parts of the live broadcast copy with corresponding images or videos as first matching information; Converting the live broadcast text into voice data, and generating matching information between the live broadcast text and the timestamp as second matching information; Generating lip-driven video material from the voice data; The lip-driven video material, the first matching information and the second matching information are used as digital human live broadcast materials.
2. The method for generating digital human live broadcast materials according to claim 1, characterized in that: The matching of different parts of the live broadcast copy with corresponding images or videos as first matching information includes: Segmenting the live broadcast copy to obtain segmented live broadcast copy; Matching the segmented live broadcast copy with the image or video, and generating a corresponding relationship between the live broadcast copy and the object information data; The corresponding relationship obtained by the matching is used as the first matching information.
3. The method for generating digital human live broadcast materials according to claim 1, characterized in that: The converting the live broadcast text into voice data and generating matching information between the live broadcast text and the timestamp as second matching information includes: Generate a timestamp of a word dimension according to the voice data, and obtain a corresponding relationship between the live broadcast copy and the timestamp; The correspondence between the timestamp and the live broadcast copy is used as the second matching information.
4. The method for generating digital human live broadcast materials according to claim 2, characterized in that: The segmenting of the live broadcast copy to obtain the segmented live broadcast copy includes: Segmenting the live broadcast copy according to semantics, and formulating a copy theme for each segmented live broadcast copy; Obtain the matching relationship between the segmented live broadcast copy and the copy theme, and store the matching relationship.
5. The method for generating digital human live broadcast materials according to claim 3, characterized in that: The step of generating lip-driven video material from the voice data includes: By using the digital human's lip-driven reasoning capability, the digital human's lip-driven video material is generated according to the voice data and the timestamp.
6. The method for generating digital human live broadcast materials according to claim 1, characterized in that: include: Matching the first matching information with the second matching information, obtaining a correspondence between the live broadcast text, voice data, corresponding image or video, and timestamp, and generating digital human live broadcast materials according to the correspondence, including: Obtaining a correspondence between the live broadcast copy and the corresponding image or video in the first matching information; Obtaining a correspondence between the live broadcast copy and the timestamp in the second matching information; Obtaining a correspondence between the lip-actuated video material and the voice data; The object information data, lip-driven video material and the timestamp are matched to generate digital human live broadcast material.
7. The method for generating digital human live broadcast materials according to claim 1, characterized in that: The digital human live broadcast material generation request carries product information, and the product information includes at least one of a product identifier and a product link.
8. The method for generating digital human live broadcast materials according to claim 7, characterized in that: The obtaining of object information data corresponding to the generation of digital human live broadcast materials includes: Generating a request for obtaining product information according to the digital human finger material, and obtaining product material information corresponding to the product information in a content database according to the product information; Using the product material information as the object information data; The content database stores commodity material information corresponding to commodity information and has the ability to generate commodity material information.
9. The method for generating digital human live broadcast materials according to claim 1, characterized in that: Also includes: The object information data is parsed from at least one of the dimensions of domain, rights and interests information, product characteristics, and type, and a live broadcast copy is generated based on the parsed content.
10. A rendering method for a digital human live broadcast image, characterized in that: include: Acquire live broadcast text, object information data, lip-driven video material, and a timestamp, render the object information data and the lip-driven video material, and generate a first digital human live broadcast screen; The first digital human live broadcast picture is rendered according to the size and resolution layout of the object information data to obtain an adapted digital human live broadcast picture.
11. The method for rendering a digital human live broadcast image according to claim 10, characterized in that: Also includes: Add switching and special effects rendering to at least one first digital human live broadcast screen to generate a second digital human live broadcast screen; The second digital human live broadcast screen is used as the digital human live broadcast screen.
12. The method for rendering a digital human live broadcast image according to claim 10, characterized in that: The rendering of the object information data and the lip-driven video material includes: performing nonlinear rendering on the object information data and the lip-driven video material.
13. A device for generating digital human live broadcast materials, characterized in that: include: An object information data acquisition unit is configured to acquire object information data corresponding to the digital human live broadcast material according to a digital human live broadcast material generation request, wherein the object information data includes at least one of text, image or video; A live broadcast copy generating unit, configured to perform multimodal understanding on the object information data and generate a live broadcast copy; A first matching information generating unit is configured to match different parts of the live broadcast text with corresponding images or videos as first matching information; A second matching information generating unit, configured to convert the live broadcast text into voice data, and generate matching information between the live broadcast text and a timestamp as second matching information; a lip-actuated video material generating unit, configured to generate lip-actuated video material from the voice data; The lip-driven video material, the first matching information and the second matching information are used as digital human live broadcast materials.
14. A system for generating digital human live broadcast materials, characterized in that: Including server, cloud, large language model and text-to-speech module: The server is used to send a request for generating digital human live broadcast materials; match different parts of the live broadcast text with corresponding images or videos as first matching information; The cloud is used to generate lip-driven video material based on the voice data, obtain object information data and a timestamp, and render the object information data and the lip-driven video material according to the timestamp to generate a first digital human live broadcast screen, which is used to display the object information data in real time; the digital human screen is rendered according to the size and push layout of the product material to obtain an adapted digital human screen; The large language model is used to obtain object information data corresponding to the digital human live broadcast material according to the digital human live broadcast material generation request, and the object information data includes at least one of text, image or video; The text-to-speech module is used to convert the live broadcast text into voice data and generate matching information between the live broadcast text and the timestamp as second matching information.
15. An electronic device, characterized in that: include: a processor, a memory, and computer program instructions stored on the memory and executable on the processor; When the processor executes the computer program instructions, the method according to any one of claims 1 to 12 is implemented.
16. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method according to any one of claims 1 to 12.
17. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Multimedia information playing method and device, equipment and storage medium
CN113891133A
Mouth shape animation synthesis method and device and electronic equipment
CN117115318A
Live broadcast method, apparatus and device, and storage medium
CN119484879A
Digital character audio and video generation method and digital character live broadcast interaction method
CN119562140A
Intelligent explanation method and device for short video copywriting
CN119579250A