Live streaming control method and apparatus, device and storage medium
By automatically creating a digital human live stream room when the live stream room with a real person closes, and using artificial intelligence to generate image elements and interactive information, the problem of long manual configuration time when live streams with real people end is solved, and a quick switch to a digital human live stream room is achieved, improving live streaming efficiency and experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2026-03-19
AI Technical Summary
In existing technologies, there is still room for improvement in how to effectively utilize digital human live streaming to improve live streaming efficiency, especially in quickly switching to a digital human live streaming room when a live stream ends, reducing manual configuration time, and providing a rich and personalized live streaming experience.
When the live streaming room is closed, a digital human live streaming room is automatically created based on the information from the live streaming room. Artificial intelligence technology is used to generate image elements, digital human images, voices, scripts and interactive information to achieve a seamless switch to digital human live streaming.
It enables the rapid setup of digital human live streaming rooms, saving the costs associated with live streamers, improving the startup speed and overall efficiency of live streaming, providing a rich and personalized live streaming experience, reducing costs, and increasing the flexibility and reach of live streaming.
Smart Images

Figure CN2025120532_19032026_PF_FP_ABST
Abstract
Description
Live broadcast control method and device, equipment and storage medium
[0001] The present disclosure claims priority to a Chinese patent application No. 202411276282.8, filed on September 12, 2024, and entitled "Live broadcast control method, device, equipment and storage medium", the entire contents of which are incorporated herein by reference. TECHNICAL FIELD
[0002] The present disclosure relates to the field of computer technology, and in particular to the fields of artificial intelligence, information flow, digital human, and smart e-commerce. BACKGROUND
[0003] With the development of AI (Artificial Intelligence), digital human live broadcast has become a potential new live broadcast method and has been gradually applied and developed in multiple industries.
[0004] As an innovative interactive method, a perfect digital human live broadcast room can not only attract the attention of users, but also save labor costs and improve work efficiency. SUMMARY
[0005] The present disclosure provides a live broadcast control method, device, equipment and storage medium.
[0006] According to an aspect of the present disclosure, a live broadcast control method is provided, comprising:
[0007] creating a digital human live broadcast room based on information of the live human broadcast room in the case that the live human broadcast room is closed;
[0008] live broadcasting based on the digital human live broadcast room.
[0009] According to another aspect of the present disclosure, a live broadcast control device is provided, comprising:
[0010] a creating module configured to create a digital human live broadcast room based on information of the live human broadcast room in the case that the live human broadcast room is closed;
[0011] a live broadcasting module configured to live broadcast based on the digital human live broadcast room.
[0012] According to another aspect of the present disclosure, an electronic device is provided, comprising:
[0013] at least one processor; and
[0014] a memory in communication with the at least one processor; wherein
[0015] The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any of the embodiments of the present disclosure.
[0016] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method according to any of the embodiments of the present disclosure.
[0017] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method according to any of the embodiments of the present disclosure.
[0018] In the embodiments of the present disclosure, by intelligently creating a digital person live room and rendering it, a digital person live room can be quickly started to shorten the manual configuration time required when manually switching to a digital person live room when a real person live room is off-air, and a more rich and personalized live experience is provided.
[0019] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0020] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them:
[0021] FIG. 1 is a flowchart of a live control method according to an embodiment of the present disclosure.
[0022] FIG. 2 is a flowchart of creating a digital person live room according to an embodiment of the present disclosure.
[0023] FIG. 3 is a flowchart of generating an image type element according to an embodiment of the present disclosure.
[0024] FIG. 4 is a flowchart of obtaining a target template based on a diffusion model according to an embodiment of the present disclosure.
[0025] FIG. 5 is a flowchart of diffusion model iteration according to an embodiment of the present disclosure.
[0026] FIG. 6 is a timing diagram of switching from a real person live room to a digital person live room according to an embodiment of the present disclosure.
[0027] FIG. 7 is a timing diagram of switching from a digital person live room to a real person live room according to an embodiment of the present disclosure.
[0028] FIG. 8 is a structural diagram of a live control apparatus according to an embodiment of the present disclosure.
[0029] FIG. 9 is a block diagram of an electronic device for implementing a live broadcast control method according to an embodiment of the disclosure. DETAILED DESCRIPTION
[0030] Exemplary embodiments of the disclosure are described herein with reference to the accompanying drawings, which are presented for the purpose of illustration and description. It is to be understood that the disclosure is not limited to the disclosed embodiments, and various changes and modifications can be made to the embodiments described herein without departing from the scope of the disclosure. Also, for the sake of brevity and clarity, descriptions of well-known functions and constructions are omitted herein.
[0031] The terms "first", "second", and the like in the disclosure are used to distinguish similar objects, and are not necessarily used to describe a particular order or sequence. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, inclusion of a series of steps or units. The method, system, product or device is not necessarily limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product or device.
[0032] With the wide application of digital humans in fields including but not limited to entertainment, education, commercial marketing, etc., most real anchors will use digital humans for live broadcast to break the limitations of region and time. Digital human live broadcast provides a more rich and novel live broadcast experience for the audience. However, how to effectively use digital human live broadcast and improve live broadcast efficiency still needs to be improved.
[0033] Therefore, the disclosure provides a live broadcast control method. The method can intelligently and automatically create a digital human live broadcast room and render the digital human live broadcast room. The method can quickly start a digital human live broadcast room to shorten the manual configuration time required for manually switching to a digital human live broadcast room when a real human live broadcast is ended.
[0034] As shown in FIG. 1, a flowchart of a live broadcast control method according to an embodiment of the disclosure is shown, which includes:
[0035] S101, in the case that a real human live broadcast room is closed, a digital human live broadcast room is created based on information of the real human live broadcast room.
[0036] That is, the closing of the real human live broadcast room or the stopping of the real human live broadcast is a trigger signal for triggering the creation of the digital human live broadcast room. The closing of the real human live broadcast room triggers the creation of the digital human live broadcast room to realize seamless switching between the real human live broadcast and the digital human live broadcast.
[0037] In implementation, a virtual live room can be created based on digital human technology to continue the demand and / or style of real-time live streaming, so that live streaming can continue without manual intervention. Digital human live streaming is a new form of live streaming that uses virtual characters (digital humans) generated by artificial intelligence to conduct live streaming activities. Digital humans can be completely fictional characters or images created based on real models.
[0038] S102, live streaming based on a digital human live room.
[0039] As the name implies, live streaming based on a digital human live room refers to using highly humanized and interactive virtual images created by artificial intelligence to conduct live streaming activities. When using a digital human live room for live streaming, the live streaming content and form of the real-time live streaming room can be referred to for reference.
[0040] In the embodiments of the present disclosure, when the real-time live streaming room is closed, a digital human live room is automatically created and live streaming is conducted. The live streaming demand can be responded to immediately without waiting for the preparation time of the real-time anchor, and the real-time anchor does not need to perform tedious configuration on the digital human live room. This way can save the relevant costs of the real-time anchor while greatly improving the startup speed and overall efficiency of live streaming. In addition, digital humans are not limited by time, place, and physiological conditions, and can meet the viewing needs of different audience groups, increasing the flexibility and coverage of live streaming. Based on the digital human live room, live streaming can be quickly adjusted and configured according to different live streaming content and scenes of the real-time live streaming room, meet the diversified needs of the audience, and provide a more rich and personalized live streaming experience.
[0041] Establishing a digital human live room is a complex and meticulous process that involves multiple aspects of configuration and customization. For example, designing a unique image for the digital human, writing a professional script for the digital human, selecting appropriate voices for the digital human, and configuring the decoration of the digital human live room. In order to improve the efficiency of creating a digital human live room as much as possible, in the embodiments of the present disclosure, the second live streaming elements of the digital human live room can be generated based on the first live streaming elements of the real-time live streaming room that are important or time-consuming during the live streaming process of the real-time anchor. In the case that the real-time live streaming room is closed, as shown in FIG. 2, the creation of the digital human live room can be accelerated based on the following steps:
[0042] S201, obtaining the second live streaming elements generated based on the first live streaming elements of the real-time live streaming room.
[0043] During the live streaming of the real-time anchor, there are various first live streaming elements, which are the basis for building live streaming content. The first live streaming elements mainly include but are not limited to at least one of the following:
[0044] Live content, such as text explained by the host, displayed goods, performed performances, etc.
[0045] Activity information, such as preferential activities, etc.
[0046] Host image: such as the appearance characteristics of the host, tone, live habits, etc.
[0047] Interactive elements: such as the way and content of interaction with the audience.
[0048] Live streaming can be captured and transmitted by the live streaming platform, recorded and saved in real time to facilitate the parsing of the first live streaming elements from the screen recording. This provides basic data for subsequent digital person live room creation.
[0049] Product information may be dynamically updated, such as optimizing activities, listing and delisting products. In the embodiments of the present disclosure, an access interface for product information can be provided to facilitate real-time access to product information.
[0050] In implementation, to improve the efficiency of generating a digital person live room, the first live streaming elements obtained can be analyzed and processed to generate second live streaming elements. Among them, the second live streaming elements include at least one of the following:
[0051] (1) Image elements of the digital person live room;
[0052] Image elements are visual elements generated for the digital person live room based on the first live streaming elements of the real person live room. For example, obtaining the product information of the historical real person live room, generating digital person live room pendants according to the preferential, promotional, and activity information. Generate digital person live room stickers according to the product name and price. Generate digital person live room background pictures according to the product detail page.
[0053] Among them, the digital person live room pendant refers to a functional icon on the live room interface. In the digital person live room, the pendant may include high-definition pictures of goods, preferential information of goods, promotional information of goods, etc., to build a more rich live streaming scene.
[0054] Digital person live room stickers refer to images or text information superimposed on the video screen of the live room, used to display brand information such as product name, product price, etc. In the digital person live room, the sticker can be dynamic and updated with the change of live streaming content, such as real-time adjustment of the price of the goods being promoted, activity countdown, etc., to improve marketing effect, attract audience attention, and increase interaction.
[0055] Digital human live streaming background picture refers to the background picture of digital human live streaming, which sets the scene and atmosphere for live streaming content. In digital human live streaming, background pictures can be very rich and varied, from static images to dynamic videos, and even interactive virtual environments. A good background picture not only attracts the audience's visual attention, but also enhances the emotional expression and information transmission of live streaming content.
[0056] In the embodiments of the present disclosure, through image elements such as digital human live streaming pendant, digital human live streaming patch, and digital human live streaming background picture, the visual effect of digital human live streaming can be greatly enriched, the attention of the audience can be attracted, and the live streaming content can be made more lively and interesting.
[0057] (2) Digital human;
[0058] Digital human, i.e. virtual anchor of digital human live streaming, is a virtual person with human form and behavior simulated by artificial intelligence. In implementation, video clips of real person live streaming can be intercepted, clear faces in the video clips are detected, and digital human image of digital human live streaming is generated according to the face of the anchor. In the case that clear faces cannot be detected in the video clips, high-quality faces are selected from a public image library as digital human images according to the gender of the anchor.
[0059] In addition, popular real person anchors can be selected to construct digital human anchors according to the traffic, sales, and other conditions of historical real person live streaming.
[0060] In the case of closing of real person live streaming and seamless switching to digital human live streaming, in order to continue the live streaming experience, the image, action, and expression of digital human live streaming can be created as much as possible according to the last real person anchor before the closing of real person live streaming, so as to imitate the real person anchor.
[0061] (3) Digital human voice;
[0062] The voice of digital human is generated by voice synthesis technology, which can simulate voices with different timbres and emotions, making the expression of digital human more rich and real. In implementation, voice clips of anchors in real person live streaming can be intercepted, and corresponding digital human voices can be generated according to the voice information of the anchors.
[0063] (4) Digital human live streaming script;
[0064] The digital human live streaming room script is a script or dialogue content followed by the digital human during live streaming, determines the theme, content and process of the digital human live streaming, and is the basis for ensuring the smooth progress of live streaming. When implemented, the product information and anchor's oral broadcast information of the real person live streaming room can be obtained, the anchor's oral broadcast information is converted into text based on the ASR (Automatic Speech Recognition) technology, the detailed information of the product is combined, and multiple scripts conforming to the anchor's style are generated through a large model.
[0065] (5) Digital human live streaming room warm-up question and answer information.
[0066] That is, questions and answers used for interaction and warm-up before or during live streaming. It can help the digital human to establish interaction with the audience and increase the activity of the live streaming room and the audience's sense of participation. For example, according to the product information of the real person live streaming room and the historical live streaming room user question and answer information, the question and answer pairs of the digital human live streaming room are generated through a large model, which are used for the warm-up question and answer of the digital human live streaming room.
[0067] In the embodiments of the present disclosure, the image type element of the digital human live streaming room can create different live streaming atmospheres, such as holiday atmosphere, activity atmosphere, etc. This makes it easier for the audience to immerse themselves in the live streaming scene and improves the interactivity and participation of the live streaming. The digital human, as the core of the digital human live streaming room, can interact with the audience, demonstrate products, explain product features and answer audience questions. Generating digital human voice based on real anchor voice through artificial intelligence helps to improve the audience's viewing experience and makes information transmission more effective. The digital human live streaming room script can ensure that the live streaming content is rich and smooth, while maximizing the audience's interest in the live streaming room. The digital human live streaming room warm-up question and answer information can increase the activity of the live streaming room. In summary, these elements work together to make the digital human live streaming room provide a novel, interactive and efficient experience, while reducing costs and improving work efficiency.
[0068] S202, creating a digital human live streaming room based on the second live streaming element.
[0069] Based on the generated second live streaming element, the generated second live streaming element is imported into the digital human live streaming room. And configure various parameters of the digital human live streaming room on the live streaming platform, such as live streaming room name, cover, tag, etc.
[0070] In implementation, a plurality of live room widgets can be generated in advance. During the live broadcast, the corresponding live room widget is dynamically called for display. Since the digital human live room widget usually contains product discount information, when the discount information in the live room widget is invalid, effective countermeasures need to be taken in time to avoid misleading the audience. In implementation, when the second live element includes a digital human live room widget, the validity of the discount information in the digital human live room widget can be detected; when the discount information is invalid, the output of the digital human live room widget to the digital human live room is stopped.
[0071] That is, during the process of digital human live room live broadcast, the validity of the discount activities displayed in the widget needs to be checked regularly or in real time according to the product information recorded in the database. For example, during the digital human live process, the operator can modify the product information in the database at any time according to the needs. During the process of digital human live, the product with discount activities to be displayed can be verified based on the database to check whether the discount activities are valid. Once it is found that a certain discount has expired or is invalid for other reasons, it can be immediately shielded to optimize the corresponding live room widget and not displayed to the audience. For example, directly delete the links and data related to the widget, or set the widget to transparent when rendering to avoid the audience seeing the outdated discount information causing live broadcast failure. In this way, the content of the digital human live room is always consistent with the expectations, and the audience's viewing experience is also improved. In addition, the correct information processing mechanism not only helps to establish the audience's trust, but also may have a positive impact on key performance indicators such as the conversion rate of the live room.
[0072] In the embodiments of the present disclosure, the second live element of the digital human live room is generated through the first live element of the real person live room, and the simulation is performed through artificial intelligence technology, which can ensure the consistency of the element information of the digital human live room and the real person live room. Based on the second live element, the digital human live room is created, which can improve the efficiency of live broadcast and reduce the cost of live broadcast.
[0073] In the embodiments of the present disclosure, based on the first live element of the real person live room, the specific implementation steps of generating the image type element are as shown in FIG. 3, including the following contents:
[0074] S301, obtaining a target image and a to-be-processed template. The to-be-processed template is a template corresponding to a target image type required by the target image in the template set.
[0075] The target image is an existing product image, and the template set is a set containing multiple different templates. These templates correspond to different image categories or scene requirements. The target image is the basis for subsequent processing, and its content and features will determine the selection of templates in subsequent steps. For example, if the target image is a scene of a tourist attraction, a landscape template may be selected as the template to be processed. In addition, it can be determined whether it is a holiday according to the current time, and a template that can render the theme or activities of the corresponding holiday is selected during the holiday.
[0076] S302, under the guidance of the target image, adjusting the template to be processed based on the diffusion model to obtain a target template that is visually adapted to the target image.
[0077] The diffusion model is a generative model that is trained by gradually adding noise to the image and gradually removing noise, so that the resulting diffusion model can be used to adjust the image to generate high-quality and diverse images. Under the guidance of the target image, the diffusion model iteratively adjusts the template to be processed to continuously adjust the template so that the template can better highlight the target image and visually assist the target image in conveying information, and can also help the audience better understand and focus on the generated image category elements.
[0078] S303, synthesizing the target image into the target template to obtain an image category element.
[0079] After obtaining the target template that is adapted to the target image, the target image is synthesized into the template. The final image category element is a perfect combination of the target image and the target template in vision. It not only retains the content and features of the target image, but also integrates the style and layout of the target template.
[0080] In the embodiments of the present disclosure, the target template is adjusted by the diffusion model, which can make the target template visually closer to the target image, thereby improving the quality and watchability of the generated image category element. Synthesizing the target image into the visually adapted template can make the image and the background more naturally integrated, reduce the sense of discomfort, and improve the quality of the generated digital person live room.
[0081] In the embodiments of the present disclosure, under the guidance of the target image, the diffusion model can be used to iteratively adjust the template to be processed to obtain the target template in the manner shown in FIG. 4, including:
[0082] S401, superimposing a first Gaussian noise on the target image to obtain an initial guide image; and superimposing a second Gaussian noise on the template to be processed to obtain an intermediate template.
[0083] As shown in FIG. 5, the initial guide image is obtained by adding the first Gaussian noise to the target image so that the target image has a certain randomness. Similarly, the second Gaussian noise is added to the template to be processed so that the template and the target image are at the same noise level, so as to facilitate subsequent denoising processing to optimize the template.
[0084] S402, merge the initial guide image into the target position of the intermediate template to obtain a first synthesized image.
[0085] By combining the content of the target image and the layout of the template to be processed, the target image with Gaussian noise (initial guide image) is merged into the specified position of the intermediate template. In implementation, the merging can be based on a mask. The mask image corresponds to the template, and the mask image is used to define the position of the target image in the template.
[0086] S403, gradually denoise the first synthesized image using a diffusion model to obtain a target template.
[0087] That is, the first synthesized image is iterated by the diffusion model to gradually repair the entire template while maintaining the original style and important information of the template.
[0088] In implementation, the diffusion model is used to denoise the first synthesized image step by step to obtain the target template, and the process is shown in FIG. 5, including the following steps:
[0089] Step A1, predict the predicted noise corresponding to the current time step based on the diffusion model.
[0090] The time step represents the number of stages passed between the completely noisy state and the completely data state in the diffusion process. In the iteration process of the diffusion model, noise prediction is performed on the current image state at each time step. Noise prediction is performed based on the understanding of the image noise pattern by the model and the rules learned from the training data. Noise prediction can estimate the amount of noise present in the current first synthesized image, providing a basis for repairing the template in the next step.
[0091] In implementation, to ensure that the target image does not change, and that the target template obtained by denoising is as consistent as possible with the main content conveyed by the original template image when the target image is fused. Therefore, in implementation, the time step used in the reverse process of the diffusion model is limited, and can be a relatively small value. In addition, the addition of the first Gaussian noise to the target image for initialization in the embodiments of the present disclosure can avoid the problem that the addition of pure Gaussian noise at the initial time step causes undesirable results. For example, it can avoid the problem that the pure Gaussian noise causes the target image and the template to be processed to not retain any original information.
[0092] Step A2, denoise the first synthesized image based on the predicted noise to obtain a second synthesized image.
[0093] Based on the estimated prediction noise in step A1, adjust each pixel in the first composite image to achieve the effect of fine-tuning the first composite image, so that the template part of the first composite image after noise reduction is more harmonious with the target image.
[0094] Step A3, the image part of the target image is segmented from the second composite image to obtain the first intermediate guide image, and the template part to be processed is segmented to update the intermediate template.
[0095] As shown in FIG. 5, the second composite image is segmented, and the part of the template to be processed is also extracted to obtain a new intermediate template. At the same time, the part of the target image is extracted to form the first intermediate guide image, so as to add the third Gaussian noise to the first intermediate guide image in step A4 to obtain the second intermediate guide image.
[0096] In implementation, a larger value is taken for the initial time step, and the value of the time step symbolizes the noise level corresponding to the generated Gaussian noise. As the time step decreases, the noise level decreases, thereby gradually achieving fine-tuning of the template. Therefore, in implementation, the third Gaussian noise is related to the next time step of the current time step, and the noise level of the third Gaussian noise is smaller than that of the third Gaussian noise of the previous time step.
[0097] Step A5, the second intermediate guide image is merged into the target position of the updated intermediate template to obtain a new first composite image, and the step of predicting the prediction noise corresponding to the current time step based on the diffusion model is returned to be executed until the current time step reaches the target time step, and the updated intermediate template is obtained as the target template.
[0098] That is, the second intermediate guide image is merged into the corresponding position of the updated intermediate template to form a new first composite image. And this image will be used as the basis for the next iteration. Repeat the above steps until the set number of time steps is reached. For example, the initial time step T is set to 60, and the adjustment process of the template is completed when the time step T=0 is executed. Among them, in each time step, the model will perform noise prediction, noise reduction, segmentation, noise addition and merging operations on the image. Finally, when the target time step is reached, the updated intermediate template obtained is the target template, which can ensure that the target image and the template to be processed have low noise levels while retaining key features.
[0099] In summary, through the entire iteration process, the diffusion model can gradually reduce the noise in the image while maintaining the key features in the template to be processed, and finally obtain a high-quality target template.
[0100] In the embodiments of the present disclosure, by adding Gaussian noise to the target image and the template to be processed, a certain degree of randomness can be introduced, so that a more natural and richer target template can be generated based on the diffusion model. The initial guide image is merged into the target position of the intermediate template to obtain a first synthesized image, and the addition of the initial guide image can make the generated target template more consistent with the target image. Through the process of gradually reducing noise, the details in the target template can be gradually optimized to obtain a high-quality target template. For example, when the color of the template to be processed in the initial state is extremely close to the color of the product in the target image, in order to highlight the product, the method can be used to increase the difference between the color of the template and the color of the product, so that the template and the product image are more harmonious, and the product theme is highlighted.
[0101] In the embodiments of the present disclosure, in order to more effectively learn and generate high-quality samples, a deterministic denoising method proposed by DDIM (Denoising Diffusion Implicit Models, a kind of diffusion model) is adopted, and by modifying the sampling strategy, the same high-quality data can be generated in fewer steps, thereby improving the generation speed.
[0102] In implementation, the DDIM diffusion model is used to gradually denoise the first synthesized image to obtain the target template, which can also be implemented based on the following steps:
[0103] Step B1, determine the step length of the time step, and determine the current time step based on the step length.
[0104] The step length of the time step determines the speed and quality of the diffusion process. Generally, the smaller the step length, the more refined the diffusion process, but the larger the amount of calculation; the larger the step length, the more rough the diffusion process, but the amount of calculation is relatively small. Therefore, before starting the diffusion process, a time step variable needs to be initialized, for example, in each iteration period, the value of the current time step is updated according to the preset step length.
[0105] For example, in the case of a step length of dt and a previous time step of t1, the current time step is (t1-dt).
[0106] Step B2, predict noise information in the first synthesized image based on the current time step in the diffusion model.
[0107] In each current time step that needs to be processed, the current first synthesized image is input into the diffusion model, and the diffusion model will predict the noise information in the first synthesized image based on the current time step.
[0108] Step B3, denoise the first synthesized image based on the noise information to obtain a third synthesized image.
[0109] The first synthesized image is denoised using the predicted noise information, and the denoised image is a third synthesized image, which has a more suitable visual performance for the target image than the first synthesized image.
[0110] Step B4: The image part of the target image is segmented from the third synthesized image to obtain a third intermediate guide image, and the template part to be processed is segmented to update the intermediate template.
[0111] The target image part and the template part to be processed are segmented from the third synthesized image. The intermediate template is updated using the segmented template part to be processed to prepare for the next merging.
[0112] Step B5: A fourth Gaussian noise is added to the third intermediate guide image to obtain a fourth intermediate guide image; the fourth Gaussian noise is related to the next time step of the current time step, and the noise level of the third Gaussian noise is smaller than the noise level of the third Gaussian noise of the previous time step.
[0113] Step B6: The third intermediate guide image is merged into the target position of the updated intermediate template to obtain a new first synthesized image, and the step of determining the step length of the time step is returned to be executed until the current time step reaches the target time step, and the updated intermediate template is obtained as the target template.
[0114] The third intermediate guide image with the added fourth Gaussian noise is merged into the target position of the updated intermediate template. The above steps are repeatedly executed until the current time step reaches the target time step, and the updated intermediate template is finally obtained as the target template.
[0115] In the embodiments of the present disclosure, the DDIM diffusion model is used to generate the target template, which can generate a high-quality target template in fewer iteration steps, thereby speeding up the efficiency of generating the second live element of the digital human live room and saving processing resources.
[0116] In the embodiments of the present disclosure, in order to ensure that the target template obtained based on the diffusion model has a better adaptation effect with the target image, the step length of the time step in the diffusion model can be dynamically adjusted based on the time step.
[0117] The dynamic adjustment of the step length of the time step refers to dynamically adjusting the step length of the time step according to the current iteration state or a specific rule in the iteration process of the diffusion model, so as to optimize the diffusion process. For example, a larger step length (such as greater than a step length threshold) can be used in the early stage of the diffusion process to quickly remove noise, and the step length can be reduced in the later stage to finely adjust image details. For example, all time steps can be divided into multiple time step intervals, different time step intervals correspond to different step lengths, and the larger the time step interval value is, the larger the corresponding step length can be. In this way, the template can be gradually and finely repaired to adapt to the target image.
[0118] In some embodiments, the new first synthesized image obtained in step B6 can also be evaluated during each iteration. For example, the neural network model can be used to evaluate the aesthetic degree, the clarity, the content relevance, the fitting degree, and the like between the target image and the template, to obtain a final evaluation result based on different dimensions. Then, the step size can be dynamically adjusted according to the final evaluation result. For example, in the case where the final evaluation result fed back by the model is not ideal (for example, the evaluation result is represented by a score, and a score lower than a preset score indicates that it is not ideal), the step size can be increased to speed up the process. In the case where the final evaluation result is good (for example, the score of the evaluation result is higher than the preset score), the step size can be reduced to achieve fine adjustment.
[0119] The specific adjustment manner can be determined based on actual conditions, and the present disclosure does not limit the same.
[0120] In the embodiments of the present disclosure, by dynamically adjusting the step size of the time step, the diffusion model used to process the synthesized image can adopt different step sizes at different iteration stages, which not only ensures the processing efficiency, but also improves the processing precision at the key stage to optimize the quality of the synthesized image.
[0121] In the embodiments of the present disclosure, in the case where the live streaming room of the real person is started again, the live streaming room identifier of the target object is detected. In the case where the live streaming room identifier indicates that the live streaming room of the digital person is live streaming, the live streaming room of the digital person is stopped. The live streaming room of the real person is switched to for live streaming.
[0122] In the embodiments of the present disclosure, by detecting the live streaming room identifier of the target object, when the system detects that the live streaming room of the real person is started again, the broadcast of the live streaming room of the digital person is paused and automatically switched to the live streaming room of the real person. This provides better flexibility for switching between the live streaming of the real person and the live streaming of the digital person in the live streaming platform, and can adjust the live streaming plan according to real-time conditions to ensure that better and more diverse live streaming content is provided.
[0123] In summary, in the embodiments of the present disclosure, the live streaming of the real person and the live streaming of the digital person can be seamlessly switched between each other. The operation of manually configuring the generation of the live streaming room of the digital person can be omitted, and the live streaming efficiency can be improved.
[0124] The process of switching from the live streaming of the real person to the live streaming of the digital person is shown in FIG. 6.
[0125] In S601, the real person anchor off-broadcasting case triggers the real person off-broadcasting signal. In S602, the live broadcast center receives the real person off-broadcasting signal, obtains the stream address of the real person live broadcast room, and sends the stream address to the digital person live broadcast (i.e. digital person live broadcast management platform). In S603, based on the obtained stream address of the real person live broadcast room, the live broadcast content, anchor image, interactive content and other key information are extracted from the audio and video data contained in the stream address, and the digital person live broadcast room is intelligently created by analyzing and processing the key information. In S604, the real person live broadcast room disconnects the stream address. In S605, the digital person live broadcast room is pushed.
[0126] Correspondingly, in the case that the real person anchor starts the real person live broadcast room again, the digital person live broadcast switching to the real person live broadcast flow is shown in FIG. 7:
[0127] In S701, the real person anchor starts to enter the live broadcast room for real person live broadcast, and triggers the real person on-broadcasting signal. In S702, the live broadcast center receives the real person on-broadcasting signal, and determines the current live broadcast state by detecting the live broadcast room identifier of the target object. In S703, in the case that the current live broadcast room is a digital person live broadcast room, and the real person live broadcast room is ready. In S704, the digital person live broadcast room stream is cut off through the stream address. In S705, the real person live broadcast room stream is pushed to the address stream of the digital person live broadcast room.
[0128] Based on the same technical concept, the disclosure embodiments also provide a live broadcast control device 800, as shown in FIG. 8, which comprises:
[0129] The creation module 801 is configured to create a digital person live broadcast room based on the information of the real person live broadcast room in the case that the real person live broadcast room is closed.
[0130] The live broadcast module 802 is configured to live broadcast based on the digital person live broadcast room.
[0131] In some embodiments, the creation module comprises:
[0132] The acquisition unit is configured to acquire a first live broadcast element generated based on the real person live broadcast room, and generate a second live broadcast element.
[0133] The live broadcast unit is configured to create a digital person live broadcast room based on the second live broadcast element.
[0134] In some embodiments, the second live broadcast element comprises at least one of the following:
[0135] The image element of the digital person live broadcast room, the digital person, the digital person sound, the digital person live broadcast room script, and the digital person live broadcast room warm-up question and answer information.
[0136] In some embodiments, the image element of the digital person live broadcast room comprises at least one of the following:
[0137] A digital human live streaming room pendant, a digital human live streaming room sticker, and a digital human live streaming room background picture.
[0138] In some embodiments, the method further comprises:
[0139] The acquisition module is configured to acquire a target image and a to-be-processed template, the to-be-processed template being a template corresponding to a target image class required by the target image in a template set;
[0140] The processing module is configured to adjust the to-be-processed template based on a diffusion model under the guidance of the target image, to obtain a target template that is visually adapted to the target image.
[0141] The synthesis module is configured to synthesize the target image into the target template to obtain an image class element.
[0142] In some embodiments, the processing module comprises:
[0143] The superposition unit is configured to superimpose a first Gaussian noise into the target image to obtain an initial guidance image, and superimpose a second Gaussian noise into the to-be-processed template to obtain an intermediate template.
[0144] The merging unit is configured to merge the initial guidance image into a target position of the intermediate template to obtain a first synthesis image.
[0145] The noise reduction unit is configured to gradually reduce noise of the first synthesis image using the diffusion model to obtain the target template.
[0146] In some embodiments, the noise reduction unit is specifically configured to:
[0147] predict a predicted noise corresponding to a current time step based on the diffusion model;
[0148] reduce noise of the first synthesis image based on the predicted noise to obtain a second synthesis image.
[0149] segment an image part of the target image from the second synthesis image to obtain a first intermediate guidance image, and segment a to-be-processed template part to update the intermediate template.
[0150] add a third Gaussian noise to the first intermediate guidance image to obtain a second intermediate guidance image.
[0151] merge the second intermediate guidance image into a target position of the updated intermediate template to obtain a new first synthesis image, and return to the step of predicting the predicted noise corresponding to the current time step based on the diffusion model until the current time step reaches a target time step, to obtain the updated intermediate template as the target template.
[0152] In some embodiments, the noise reduction unit is specifically configured to:
[0153] determine a step length of a time step;
[0154] determine a current time step based on the step length;
[0155] predict noise information in the first synthesized image based on the current time step in the diffusion model;
[0156] de-noise the first synthesized image based on the noise information to obtain a third synthesized image;
[0157] segment an image part of the target image from the third synthesized image to obtain a third intermediate guide image, and segment the to-be-processed template part to update the intermediate template;
[0158] add fourth Gaussian noise to the third intermediate guide image to obtain a fourth intermediate guide image;
[0159] merge the third intermediate guide image into the target position of the updated intermediate template to obtain a new first synthesized image, and return to execute the step of determining the step length of the time step until the current time step reaches a target time step, and obtain the updated intermediate template as the target template.
[0160] In some embodiments, the step length of the time step is dynamically adjusted based on the time step.
[0161] In some embodiments, the live broadcast unit is specifically configured to:
[0162] In the case where the second live broadcast element includes a digital human live broadcast room pendant, detecting the validity of the preferential information in the digital human live broadcast room pendant;
[0163] In the case where the preferential information is invalid, stop outputting the digital human live broadcast room pendant to the digital human live broadcast room.
[0164] In some embodiments, further comprising a switching module configured to:
[0165] In the case where the real person live broadcast room is started again, detecting the live broadcast room identifier of the target object;
[0166] In the case where the live broadcast room identifier indicates that the digital human live broadcast room is live broadcasting, stopping the digital human live broadcast room;
[0167] Switching to the real person live broadcast room for live broadcasting.
[0168] The specific functions and examples of each module and sub-module of the device of the embodiments of the present disclosure are described in the above method embodiments, and will not be described here.
[0169] In the technical solutions of the present disclosure, the acquisition, storage and application of user personal information comply with relevant laws and regulations and do not violate public order and good customs.
[0170] According to embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0171] FIG. 9 shows a schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.
[0172] As shown in FIG. 9, the device 900 includes a computing unit 901 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded into a random access memory (RAM) 903 from a storage unit 908. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0173] Various components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunications networks.
[0174] The computing unit 901 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 901 performs various methods and processes described above, such as the live streaming control method. For example, in some embodiments, the live streaming control method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded onto the RAM 903 and executed by the computing unit 901, one or more steps of the live streaming control method described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the live streaming control method by any other appropriate means, such as by means of firmware.
[0175] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0176] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0177] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0178] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0179] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0180] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0181] It should be understood that the various forms of flow shown above can be re-ordered, steps added or removed, etc. For example, the steps recited in the present disclosure can be performed in parallel, in series, in a different order, etc., as long as the desired results of the technology disclosed in the present disclosure are achieved, which is not limited herein.
[0182] The above detailed description does not constitute a limitation of the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A live broadcast control method, comprising: in the case of closing a live broadcast in a real person live room, creating a digital person live room based on information of the real person live room; live broadcast based on the digital person live room.
2. The method of claim 1, wherein, The digital person live room based on the information of the real person live room comprises: obtaining a first live element based on the real person live room in advance, generating a second live element; based on the second live element, create the digital person live room.
3. The method of claim 2, wherein, The second live element comprises at least one of the following: image elements of digital person live room, digital person, digital person sound, digital person live room script, digital person live room warm-up question and answer information.
4. The method of claim 3, wherein, The image elements of the digital person live room include at least one of the following: digital person live room pendant, digital person live room patch, digital person live room background picture.
5. The method of claim 3, wherein, Based on the first live element of the real person live room, the image element is generated, comprising: obtaining a target image and a to-be-processed template; the to-be-processed template is a template corresponding to a target image class required by the target image in a template set; under the guidance of the target image, adjust the to-be-processed template based on the diffusion model to obtain a target template that is visually adapted to the target image; synthesizing the target image into the target template to obtain the image element.
6. The method of claim 5, wherein, The target template that is visually adapted to the target image is obtained by adjusting the to-be-processed template based on the diffusion model under the guidance of the target image, comprising: adding first Gaussian noise to the target image to obtain an initial guide image; and, adding second Gaussian noise to the to-be-processed template to obtain an intermediate template; merge the initial guide image into the target position of the intermediate template to obtain a first synthesized image; gradually denoise the first synthesized image using a diffusion model to obtain the target template.
7. The method of claim 6, wherein, The target template is obtained by gradually denoising the first synthesized image using a diffusion model, comprising: predicting the prediction noise corresponding to the current time step based on the diffusion model; based on the prediction noise, denoise the first synthesized image to obtain a second synthesized image; segment the image part of the target image from the second synthesized image to obtain a first intermediate guide image, and segment the to-be-processed template part to update the intermediate template; add third Gaussian noise to the first intermediate guide image to obtain a second intermediate guide image; merge the second intermediate guide image into the target position of the updated intermediate template to obtain a new first synthesized image, and return to execute the step of predicting the prediction noise corresponding to the current time step based on the diffusion model until the current time step reaches the target time step, and obtain the updated intermediate template as the target template.
8. The method of claim 6, wherein, The target template is obtained by gradually denoising the first synthesized image using a diffusion model, comprising: determine the step length of the time step; determine the current time step based on the step length; in the diffusion model, predict the noise information in the first synthesized image based on the current time step; based on the noise information, denoise the first synthesized image to obtain a third synthesized image; segmenting an image part of the target image from the third synthetic image to obtain a third intermediate guide image, and segmenting the template part to be processed to update the intermediate template; adding fourth Gaussian noise to the third intermediate guide image to obtain a fourth intermediate guide image; merging the third intermediate guide image into a target position of the updated intermediate template to obtain a new first synthetic image, and returning to perform the step of determining the step length of the time step until the current time step reaches a target time step, to obtain an updated intermediate template as the target template.
9. The method of claim 8, wherein, The step length of the time step is dynamically adjusted based on the time step.
10. The method of claim 2, wherein, The creating of the digital human live room based on the second live element includes: In the case that the second live element includes a digital human live room widget, detecting the validity of the preferential information in the digital human live room widget; In the case that the preferential information is invalid, stopping outputting the digital human live room widget to the digital human live room.
11. The method of any one of claims 1-10, further comprising: In the case that the real person live room is started again, detecting a live room identifier of a target object; In the case that the live room identifier indicates that the digital human live room is live streaming, stopping the digital human live room from live streaming; Switching to the real person live room for live streaming.
12. A live streaming control apparatus, comprising: a creating module configured to create a digital human live room based on information of a real person live room in the case that the real person live room stops live streaming; a live streaming module configured to live stream based on the digital human live room.
13. The apparatus of claim 12, wherein, The creating module includes: an obtaining unit configured to obtain a second live element generated based on a first live element of the real person live room in advance; a live streaming unit configured to create the digital human live room based on the second live element.
14. The apparatus of claim 13, wherein, The second live element includes at least one of the following: an image type element of the digital human live room, a digital human, a digital human sound, a digital human live room script, and digital human live room warm-up question and answer information.
15. The apparatus of claim 14, wherein, The image type element of the digital human live room includes at least one of the following: a digital human live room widget, a digital human live room sticker, and a digital human live room background image.
16. The apparatus of claim 14, wherein, Further comprising: an obtaining module configured to obtain a target image and a template to be processed; The template to be processed is a template corresponding to a target image type required by the target image in a template set; a processing module configured to adjust the template to be processed based on a diffusion model under guidance of the target image to obtain a target template visually adapted to the target image; a synthesizing module configured to synthesize the target image into the target template to obtain the image type element.
17. The apparatus of claim 16, wherein, The processing module includes: a superimposing unit configured to superimpose first Gaussian noise into the target image to obtain an initial guide image, and superimpose second Gaussian noise into the template to be processed to obtain an intermediate template; a merging unit configured to merge the initial guide image into a target position of the intermediate template to obtain a first synthetic image; a noise reducing unit configured to gradually reduce noise of the first synthetic image using a diffusion model to obtain the target template.
18. The apparatus of claim 17, wherein, The noise reduction unit is specifically configured to: predict a predicted noise corresponding to a current time step based on the diffusion model; perform noise reduction on the first synthesized image based on the predicted noise to obtain a second synthesized image; segment an image part of the target image from the second synthesized image to obtain a first intermediate guide image, and segment the template part to be processed to update the intermediate template; add third Gaussian noise to the first intermediate guide image to obtain a second intermediate guide image; merge the second intermediate guide image into a target position of the updated intermediate template to obtain a new first synthesized image, and return to perform the step of predicting a predicted noise corresponding to a current time step based on the diffusion model until the current time step reaches a target time step, to obtain an updated intermediate template as the target template.
19. The apparatus of claim 17, wherein, The noise reduction unit is specifically configured to: determine a step length of a time step; determine a current time step based on the step length; predict noise information in the first synthesized image based on the current time step in the diffusion model; perform noise reduction on the first synthesized image based on the noise information to obtain a third synthesized image; segment an image part of the target image from the third synthesized image to obtain a third intermediate guide image, and segment the template part to be processed to update the intermediate template; add fourth Gaussian noise to the third intermediate guide image to obtain a fourth intermediate guide image; merge the third intermediate guide image into a target position of the updated intermediate template to obtain a new first synthesized image, and return to perform the step of determining a step length of a time step until the current time step reaches a target time step, to obtain an updated intermediate template as the target template.
20. The apparatus of claim 19, wherein, The step length of the time step is dynamically adjusted based on the time step.
21. The apparatus of claim 13, wherein, The live streaming unit is specifically configured to: in a case where the second live streaming element includes a digital human live streaming room widget, detect validity of preferential information in the digital human live streaming room widget; in a case where the preferential information is invalid, stop outputting the digital human live streaming room widget to the digital human live streaming room.
22. The apparatus of any one of claims 12-21, further comprising a switching module configured to: in a case where the real person live streaming room is started again, detect a live streaming room identifier of a target object; in a case where the live streaming room identifier indicates that the digital human live streaming room is live streaming, stop the digital human live streaming room; switch to the real person live streaming room for live streaming.
23. An electronic device, comprising: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-11.
24. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-11.
25. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-11.
Citation Information
Patent Citations
Method and system for mixed use of virtual real persons during live video broadcasting
CN111866529A
Live broadcast method and device based on digital human, electronic equipment and readable storage medium
CN117579851A
Live broadcast room voice driving method and device, electronic equipment and storage medium
CN118075559A
Live broadcast control method, apparatus and device, and storage medium
CN119583831A
Multimodal diffusion models
US20240265505A1