A live video generation method and device, electronic equipment and storage medium

By combining audio data and 3D scene data, and using a digital human image generation model to render live stream images, the problem of insufficient 2D live stream backgrounds is solved, achieving a realistic and vivid live stream effect.

CN118870042BActive Publication Date: 2026-01-27NANJING SILICON INTELLIGENCE TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310486468.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-28
Publication Date
2026-01-27
Estimated Expiration
2043-04-28

AI Technical Summary

Technical Problem

Most existing live streaming backgrounds are composed of two-dimensional images or textures, resulting in insufficient sense of space and realism, which affects the live streaming effect and the streamer's performance.

Method used

By acquiring audio data and 3D scene data, a pre-trained digital human image generation model is used to process the audio data, generate multiple frames of digital human images, and fuse them with the 3D scene data for rendering, producing realistic and vivid live broadcast images.

Benefits of technology

It improves the aesthetics and realism of the live stream background, enhances the viewing experience of the live stream video, provides more effective information, and improves the overall live stream effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118870042B_ABST
    Figure CN118870042B_ABST
Patent Text Reader

Abstract

The application provides a live video generation method and device, an electronic device and a storage medium, relates to the technical field of network live broadcast, and can make the virtual background in the live video more real and improve the live broadcast effect. The method is applied to a content expression device in a live broadcast system and includes the following steps: acquiring audio data; acquiring three-dimensional scene data; processing the audio data based on a digital person image generation model to obtain multiple frames of digital person images corresponding to the audio data; the digital person image generation model has the capability of obtaining multiple frames of digital person images corresponding to the audio data based on the audio data; rendering multiple frames of live broadcast images based on the multiple frames of digital person images and the three-dimensional scene data; the live broadcast images include digital person images and a virtual three-dimensional scene indicated by the three-dimensional scene data; and sending the audio data and the multiple frames of live broadcast images to a live broadcast data generation device, so that the live broadcast data generation device generates live broadcast video data based on the audio data and the multiple frames of live broadcast images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of live streaming technology, and in particular to a method and apparatus for generating live video, an electronic device, and a storage medium. Background Technology

[0002] With the development of live streaming technology, more and more users are choosing live streaming for various purposes, such as live gaming, video streaming, or product sales. During live streaming, to save on the cost of changing backgrounds, easily changeable virtual backgrounds are often generated. However, most current live streaming backgrounds are composed of images or textures, which only provide basic aesthetic enhancement. Because they contain only immutable two-dimensional image information, they lack spatial depth and realism. This limits the overall aesthetic appeal of the live stream and makes the broadcast less realistic and engaging, resulting in a less effective and less engaging experience. Summary of the Invention

[0003] To address the aforementioned technical problems, this application provides a method for generating live video, a method for processing live video parameter data, and an electronic device, which can improve data organization efficiency.

[0004] The technical solution of this application is as follows:

[0005] Firstly, this application provides a method for generating live video, applied to a content expression device in a live streaming system. The live streaming system also includes a live data generation device connected to the content expression device. The method includes: acquiring audio data; acquiring three-dimensional scene data; processing the audio data based on a digital human image generation model to obtain multiple frames of digital human images corresponding to the audio data; wherein the digital human image generation model has the ability to obtain multiple frames of digital human images corresponding to the audio data based on the audio data; rendering multiple frames of live images based on the multiple frames of digital human images and the three-dimensional scene data; the live images include a virtual three-dimensional scene indicated by the digital human images and the three-dimensional scene data; and sending the audio data and the multiple frames of live images to the live data generation device, so that the live data generation device generates live video data based on the audio data and the multiple frames of live images.

[0006] In one possible implementation of the first aspect, acquiring three-dimensional scene data includes: acquiring scene images of the target three-dimensional scene from multiple perspectives; and generating three-dimensional scene data based on the scene images from multiple perspectives.

[0007] In one possible implementation of the first aspect, generating three-dimensional scene data based on scene images from multiple perspectives includes: generating a three-dimensional scene model to be determined based on scene images from multiple perspectives; receiving an adjustment operation, adjusting the three-dimensional scene model to be determined in response to the adjustment operation to obtain a three-dimensional scene model, and determining the three-dimensional scene data by using the three-dimensional scene model and its feature parameters.

[0008] In one possible implementation of the first aspect, the three-dimensional scene data includes: a three-dimensional scene model and feature parameters of the three-dimensional scene model; the feature parameters include: depth information of each object and coordinates of each object; based on multiple frames of digital human images and three-dimensional scene data, rendering multiple frames of live images includes: determining the first depth of the digital human in the virtual three-dimensional scene in the multiple frames of live images according to the depth information of each object in the three-dimensional scene model; determining the first coordinate data of the first depth of the digital human in the virtual three-dimensional scene in the multiple frames of live images according to the coordinates of each object in the three-dimensional scene model; and rendering the three-dimensional scene model and the multiple frames of digital human images based on the first depth and first coordinate data of the digital human in the three-dimensional scene in the multiple frames of live images to obtain multiple frames of live images.

[0009] In one possible implementation of the first aspect, the three-dimensional scene data includes: a three-dimensional scene model and feature parameters of the three-dimensional scene model; the feature parameters also include lighting information; rendering the three-dimensional scene model and the multi-frame digital human images based on the first depth and first coordinate data of the digital human in the virtual three-dimensional scene from the multi-frame live images includes: adjusting the color values ​​in the multi-frame digital human images based on the lighting information to obtain multi-frame target digital human images; and rendering the three-dimensional scene model and the multi-frame target digital human images based on the first depth and first coordinate data of the digital human in the three-dimensional scene from the multi-frame live images to obtain multi-frame live images.

[0010] In one possible implementation of the first aspect, generating live video data based on audio data and multiple live images includes: determining the correspondence between each live image in the multiple live images and multiple audio segments in the audio data; and generating live video data based on the correspondence.

[0011] Secondly, this application provides a live video generation apparatus, which includes an acquisition module, a processing module, a rendering module, and a generation module. The acquisition module is used to acquire audio data; it is also used to acquire 3D scene data. The processing module is used to process the audio data acquired by the acquisition module based on a digital human image generation model to obtain multiple frames of digital human images corresponding to the audio data; wherein the digital human image generation model has the ability to obtain multiple frames of digital human images corresponding to the audio data. The rendering module is used to render multiple frames of live images based on the multiple frames of digital human images obtained by the processing module and the 3D scene data acquired by the acquisition module; the live images include digital human images and a virtual 3D scene indicated by the 3D scene model. The sending module is used to send the audio data acquired by the acquisition module and the multiple frames of live images obtained by the rendering module to a live data generation device, so that the live data generation device generates live video data based on the audio data and the multiple frames of live images.

[0012] In one possible implementation of the second aspect, the acquisition module is specifically used to: acquire scene images of the target 3D scene from multiple perspectives; and generate 3D scene data based on the scene images from multiple perspectives.

[0013] Thirdly, this application provides an electronic device including a processor and a memory; the memory is used to store computer instructions, and when the electronic device is running, the processor executes the computer instructions stored in the memory to enable the electronic device to implement the live video generation method provided in the first aspect above.

[0014] Fourthly, this application also provides a computer-readable storage medium storing instructions that, when executed on an electronic device, cause the electronic device to implement the live video generation method provided in the first aspect above.

[0015] Fifthly, this application also provides a computer program product containing computer instructions, which, when executed on an electronic device, cause the electronic device to implement the live video generation method provided in the first aspect above.

[0016] In this application, the aforementioned names do not limit the devices or functional modules themselves. In actual implementation, these devices or functional modules may appear under other names. As long as the functions of each device or functional module are similar to those in this application, they fall within the scope of the claims of this application and their equivalents.

[0017] These or other aspects of this application will become more readily apparent in the following description.

[0018] The technical solution provided in this application firstly processes audio data based on a pre-trained digital human image generation model to obtain multiple frames of digital human images corresponding to the audio data. Then, based on these multiple frames of digital human images and the 3D scene data, multiple frames of live-stream images can be rendered. Since the live-stream images are obtained by fusing and rendering digital human images and 3D scene data, there is no significant disconnect between the digital human images and the 3D scene indicated by the 3D scene data, and the 3D scene also provides a good aesthetic enhancement for the live stream. In other words, the live-stream images are relatively realistic and vivid for viewers. Furthermore, after sending the audio data and live-stream images to a live-stream video generation device, the device can generate realistically generated live-stream video data based on the audio data and the multiple frames of live-stream images. During the use of this live-stream video data, the live-stream background can better enhance the user's viewing experience, providing more effective information (such as realistically and vividly demonstrating the usage scenario of the goods being sold through a 3D scene), thus improving the live-stream effect. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of the structure of an implementation environment provided in an embodiment of this application;

[0022] Figure 2 This application provides a schematic diagram of the structure of a live streaming system according to an embodiment of the present application.

[0023] Figure 3 A flowchart illustrating a method for generating live video provided in this application embodiment. Figure 1 ;

[0024] Figure 4 A flowchart illustrating a method for generating live video provided in this application embodiment. Figure 2 ;

[0025] Figure 5 A flowchart illustrating a method for generating live video provided in this application embodiment. Figure 3 ;

[0026] Figure 6A flowchart illustrating a method for generating live video provided in this application embodiment. Figure 4 ;

[0027] Figure 7 A schematic diagram of a digital phreography image provided in an embodiment of this application;

[0028] Figure 8 A schematic diagram of a digital human template image provided in an embodiment of this application;

[0029] Figure 9 A flowchart illustrating a method for generating live video provided in this application embodiment. Figure 5 ;

[0030] Figure 10 A schematic diagram of a three-dimensional scene model provided in an embodiment of this application;

[0031] Figure 11 A schematic diagram of a live image provided in an embodiment of this application;

[0032] Figure 12 A flowchart illustrating a method for generating live video provided in this application embodiment. Figure 6 ;

[0033] Figure 13 A flowchart illustrating a method for generating live video provided in this application embodiment. Figure 7 ;

[0034] Figure 14 A flowchart illustrating a method for generating live video provided in this application embodiment. Figure 8 ;

[0035] Figure 15 A scene illustration of a movable ornament provided in an embodiment of this application;

[0036] Figure 16 This is a schematic diagram of a mobile digital human scenario provided in an embodiment of this application;

[0037] Figure 17 A schematic diagram of the structure of a live video generation device provided in an embodiment of this application;

[0038] Figure 18 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0039] To better understand the above-mentioned objectives, features, and advantages of this application, the solution of this application will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0040] Many specific details are set forth in the following description in order to provide a full understanding of this application, but this application may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some embodiments of this application, and not all embodiments.

[0041] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0042] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0043] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0044] In existing online live streaming scenarios, the virtual backgrounds of live streaming rooms are mostly composed of two-dimensional images or textures. The resulting virtual backgrounds are fixed and cannot be changed. This limits the degree of beautification of the live streaming room. At the same time, the sense of separation between the two-dimensional background and the three-dimensional anchor makes the live streaming process less realistic and vivid, resulting in poor live streaming effects.

[0045] To address the aforementioned issues, this application provides a method for generating live video. In this method, after acquiring audio data and 3D scene data, the audio data is first processed using a pre-trained digital human image generation model to obtain multiple frames of digital human images corresponding to the audio data. Then, multiple frames of live video images are rendered based on these multiple frames of digital human images and the 3D scene data. Since the live video images are obtained through fusion rendering of digital human images and 3D scene data, there is no significant disconnect between the digital human images and the 3D scene indicated by the 3D scene data, and the 3D scene provides a good aesthetic enhancement for the live stream. In other words, the live video images are relatively realistic and vivid for viewers. Furthermore, based on the audio data and the multiple frames of live video images, truly generated live video data can be generated. Using this live video data, the live background can better enhance the user's viewing experience, providing more effective information (such as realistically and vividly depicting the usage scenario of the goods being sold through the 3D scene), thus improving the live stream effect.

[0046] The technical solutions provided in the embodiments of this application can be applied to, for example... Figure 1 The implementation environment shown may include a live streaming system 01, a live streaming platform 02, and a user terminal 03. The live streaming system 01 and the live streaming platform 02 can communicate via wired or wireless communication methods, and the live streaming platform 02 and the user terminal 03 can also communicate via wired or wireless communication methods.

[0047] The live streaming system 01 is mainly used to generate live streaming videos, which are then pushed to the live streaming platform 02 for live streaming. In this embodiment, there may be one or more live streaming platforms 02. Figure 1 This application uses only one live streaming platform, 02, as an example, and does not impose any specific restrictions on it.

[0048] The live streaming platform 02 is mainly used to push the live streaming data to the user terminal 03 according to specific rules after receiving the live streaming video from the live streaming system 01, so that the user terminal 03 can play the live streaming data corresponding to the live streaming video for the user to watch. In this embodiment, there may be one or more user terminals 03. The user terminal 03 can be an electronic device with a live streaming application corresponding to the live streaming platform 02 installed.

[0049] Reference Figure 2 As shown in the figure, in this application, the live streaming system 01 may include a decision-making device 011, a process-driven device 012, a content expression device 013, and a live streaming data generation device 014.

[0050] Among them, the decision-making device 011 establishes a communication connection with the process-driven device 012 and the content expression device 013, while the process-driven device 012 and the content expression device 013 establish a communication connection with the live data generation device 014.

[0051] It should be noted that in this application, the decision-making device 011, process-driven device 012, content expression device 013, and live data generation device 014 can be separate devices, different parts of the same device, or modules in a data center implementing different functions, depending on actual needs. This application does not impose specific limitations in this regard. Specifically, the process-driven device 012 can acquire user process requirement data and generate live flow data and live rule data based on it. The content expression device 013 can acquire user content requirement data and generate live content data based on it. The live content data may include audio data and multiple frames of live images. The live data generation device 014 can generate live video based on the live flow data, live rule data, and live content data.

[0052] Decision-making device 011 can generate live streaming decision data based on the playback data of the live streaming video, and adjust the live streaming process data and live streaming rule data generated by process driver device 012 based on the live streaming decision data, and / or adjust the live streaming content data generated by content expression device 013.

[0053] The method for generating live video provided in the embodiments of this application is described below. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0054] based on Figure 1 The implementation environment shown in this application embodiment illustrates a method for generating live video. This method can be implemented by a live video generation device, which may be the content expression device or a part thereof described in the foregoing embodiments. (See reference...) Figure 3 As shown, the method may include S301-S305:

[0055] S301, Obtain audio data.

[0056] In this embodiment, the audio data can specifically be the script that the user needs the host to say in the live stream. Taking live-streaming e-commerce as an example, the audio data can include audio of the host introducing and promoting the products during the live-streaming e-commerce process, audio of the host interacting with viewers, etc.

[0057] In this application embodiment, audio data acquisition can be achieved in the following ways:

[0058] In the first feasible approach, the audio data can be directly input by the user into the live video generation device, and the input method can be any feasible method.

[0059] In some embodiments, the content that the user directly inputs into the live video generation device can be the spoken content (specifically, text content) that the user needs, and the live video generation device can use a specific text conversion model to convert the spoken content into audio data.

[0060] In the second feasible approach, the user first provides audio request data to the live video generation device. This audio request data can reflect the user's needs regarding the content spoken by the host in the live stream. For example, the audio request data can be a short descriptive phrase such as "I want a live stream selling maternity and baby products, and the host should be female," or it can include keywords such as "live stream for product sales," "maternity and baby products," and "mature female host." The specific content of this request can be in text or audio format. If the content request is in audio format, the live video generation device needs to first convert the audio request data into text data before further processing.

[0061] After receiving the audio request data, the live video generation device can determine the audio data corresponding to the audio request data based on the audio request data.

[0062] Specifically, after receiving audio request data, the live video generation device can use this data to obtain audio request features. Based on these features, the device can then determine audio data from an audio library that matches the user's needs. When the audio request data is a short descriptive sentence, the audio request features can be keywords within the data. For example, if the audio request data is "I want a live stream selling baby products, and the host must be female," then the audio request features would be "selling products," "baby products," and "female host." If the audio request data itself is a keyword, then the audio request features can be the data itself. This application does not impose specific limitations in this regard. The audio library can include various types of audio data and the relationships between these audio data and the audio request features.

[0063] Of course, in practice, audio data can be acquired in any other feasible way, and this application does not impose any specific restrictions on this.

[0064] S302. Obtain 3D scene data.

[0065] The 3D scene data can be used to indicate a virtual 3D scene, and may include a 3D scene model and its feature parameters. For example, the feature parameters may include one or more of the following: depth information of each object, coordinates of each object, and lighting information.

[0066] In this embodiment of the application, the acquisition of three-dimensional scene data can be achieved in the following ways:

[0067] In the first feasible approach, the 3D scene data can be directly input by the user into the live video generation device, and the input method can be any feasible method. For example, in some embodiments, the user can directly draw to obtain the 3D scene data using 3D painting software.

[0068] In the second feasible approach, the user first provides scene requirement data to the live video generation device. This scene requirement data can reflect the user's needs for the background of the live stream. For example, the scene requirement data can be a descriptive sentence like "I want a live stream room selling maternity and baby products, with warm lighting and decorations all made of maternity and baby products," or it can include keywords such as "live stream room for product sales," "maternity and baby products," "warm lighting," and "maternity and baby decorations." The specific content requirement can be in text or audio format. If the content requirement is in audio format, the live video generation device needs to first convert the requirement data into text data before further processing.

[0069] After receiving the scene requirement data, the live video generation device can determine the corresponding three-dimensional scene data based on the scene requirement data.

[0070] Specifically, after receiving scene requirement data, the live video generation device can use this data to obtain scene requirement features. Based on these features, the device can then determine 3D scene data from a scene library that meets the user's needs. When the scene requirement data is a short descriptive sentence, the scene requirement features can be keywords within the data. For example, if the scene requirement data is "I want a live stream room selling baby products, with warm lighting and baby products as decorations," then the scene requirement features are "selling goods," "baby products," "warm lighting," and "baby decorations." If the scene requirement data itself is a keyword, then the requirement features can be the data itself. This application does not impose specific limitations on this. The scene library may include various types of 3D scene data and the relationships between these data and scene requirement features.

[0071] In the third feasible approach, the user can provide images of the target 3D scene from multiple perspectives to the live video generation device, thereby enabling the device to generate corresponding 3D scene data for the target 3D scene. Based on this, combined with... Figure 3 , refer to Figure 4 As shown, S302 may specifically include S401 and S402:

[0072] S401. Obtain scene images of the target 3D scene from multiple perspectives.

[0073] In this embodiment, the target 3D scene is an actual 3D scene, such as a baby store scene, a kitchenware store scene, an electronics store scene, an office supplies store scene, etc. The target 3D scene can also be a virtual 3D scene, such as a store scene in a game world. Scene images from multiple perspectives of the target scene refer to the 2D images of the target scene from different perspectives.

[0074] Taking a real-world 3D scene as an example, in this embodiment of the application, the live video generation device can obtain scene images of the target 3D scene from multiple perspectives from a user terminal (e.g., a mobile phone) (specifically, the user inputs the images into the live video generation device using the user terminal), and can also search for captured images of the target 3D scene from multiple perspectives on the Internet as scene images of the target 3D scene from multiple perspectives.

[0075] Taking the acquisition of scene images of a target 3D scene from multiple perspectives by a live video generation device from a user terminal as an example, in some embodiments, the user terminal may store scene images of the target 3D scene from multiple perspectives and send them to the live video generation device when needed by the user. The scene images of the target 3D scene stored in the user terminal may be those captured by the user through the user terminal, or captured using any feasible shooting device (e.g., a camera) and then stored on the user terminal; or they may be those obtained by the user from a network search. This application does not impose specific limitations in this regard.

[0076] In one possible implementation, the user terminal can also extract scene images from multiple perspectives from a 360-degree video of the target scene stored locally. Specifically, the user terminal can extract keyframes from the 360-degree video of the target scene as scene images from multiple perspectives according to a predetermined rule. This predetermined rule can be to extract frames at specific time intervals, or any other feasible method, and this application does not impose any specific limitations on it.

[0077] When the data video generating device searches for scene images of a target 3D scene from multiple perspectives on the network, the user needs to input the name or description of the target scene into the data video generating device so that the data video generating device can obtain scene images of the target 3D scene from multiple perspectives on the network.

[0078] Of course, if the target 3D scene is a virtual 3D scene, the data video generation device can obtain scene images of the target 3D scene from multiple possible websites based on the name or description information of the target 3D scene input by the user.

[0079] S402. Generate the 3D scene data based on scene images from multiple perspectives.

[0080] After acquiring scene images from multiple perspectives, the live video generation device can use any feasible 3D reconstruction method to generate a 3D scene model and its feature parameters corresponding to the target 3D scene, thus obtaining the 3D scene data. In this embodiment, 3D reconstruction refers to the mathematical process and computer technology of recovering the 3D information (shape, etc.) of an object using 2D projection. This application does not impose specific limitations on the 3D reconstruction method.

[0081] Based on the technical solutions corresponding to S401 and S402, the live video generation device can generate three-dimensional scene data corresponding to the target three-dimensional scene based on scene images from multiple perspectives of the target three-dimensional scene, providing data support for the subsequent generation of a live room with a virtual three-dimensional scene.

[0082] In some embodiments, the live video generation device may be capable of generating 3D scene data using scene images from multiple perspectives. However, due to the limitations of the multiple perspective scene images input by the user (e.g., insufficient perspective), the 3D scene model automatically generated by the live video generation device cannot well meet the user's needs. Therefore, after the initial 3D scene model is generated, relevant operators (e.g., art team members) are needed to adjust the initial 3D scene model to obtain a 3D scene model that meets the user's requirements. Based on this, combined with Figure 4 , refer to Figure 5 As shown, S402 may specifically include S501 and S502:

[0083] S501. Generate a three-dimensional scene model to be determined based on scene images from multiple perspectives.

[0084] For example, taking the ability of a live video generation device to generate three-dimensional scene data from scene images from multiple perspectives as an example, which is achieved through a 3D reconstruction application (such as Maya, 3D Max, etc.), S501 can specifically be: the live video generation device performs 3D reconstruction on scene images from multiple perspectives of the target scene based on the 3D reconstruction application installed on itself, so as to obtain a three-dimensional scene model to be determined.

[0085] Taking Maya as an example of a 3D reconstruction application, due to Maya's specific requirements, the scene images from multiple perspectives can include three views of the target 3D scene, namely, the front scene image, the side scene image, and the rear scene image. Of course, in practice, depending on the different 3D reconstruction applications, the scene images from multiple perspectives can include scene images corresponding to any other possible perspectives, and this application does not impose any specific restrictions on this.

[0086] S502, Receive adjustment operation, respond to the adjustment operation to adjust the three-dimensional scene model to obtain a three-dimensional scene model, and determine the three-dimensional scene data by combining the three-dimensional scene model and the feature parameters of the three-dimensional scene model.

[0087] Specifically, after the live video generation device generates a pending 3D scene model, it can display the model. Once the operator sees the pending 3D scene model, they can adjust it based on the scene image input by the user. The adjustment process can be as follows:

[0088] 1. The operator performs a first adjustment operation based on the scene image input by the user to adjust the geometric parameters such as the size, height, and width of the three-dimensional model to be determined. After receiving the first adjustment operation, the live video generation device will adjust the geometric parameters such as the size, height, and width of the three-dimensional model to be determined accordingly to obtain the first three-dimensional scene model to be adjusted.

[0089] This first step can be implemented using a 3D reconstruction application (such as Maya).

[0090] 2. The operator imports the first 3D scene model to be adjusted into a 3D design application (e.g., ZBrush) in the live video generation device and inputs design operations. The live video generation device receives the design operations and, in response, calls the functions in the 3D design application to adjust the number of faces and details of the first 3D scene model to be adjusted, resulting in a more realistic second 3D scene to be adjusted.

[0091] 3. The operator imports the second 3D scene to be adjusted into the 3D reconstruction application and inputs a topology operation. The live video generation device receives the topology operation and, in response, calls the function in the 3D reconstruction application to adjust the number of faces in the second 3D model to be adjusted, thereby obtaining a third 3D scene to be adjusted that conforms to the configuration parameters of the live video generation device.

[0092] 4. The operator performs UV splitting on the third 3D scene to be adjusted. Here, U represents the distribution on the horizontal coordinate, V represents the distribution on the vertical coordinate, and UV can be understood as the skin of the 3D model. UV splitting specifically refers to unfolding the 3D model into a 2D image. The live video generation device receives this UV splitting operation and, in response, invokes functions in the 3D reconstruction application to perform UV splitting on the third 3D model to be adjusted, obtaining a 2D unfolded image of the third 3D model to be adjusted.

[0093] 5. The operator performs image retouching on the 2D unfolded image of the third 3D model to be adjusted. The live video generation device receives the image retouching operation and, in response, adjusts the texture material of the 2D unfolded image of the third 3D model to be adjusted to obtain a more realistic 3D scene model and determines the feature parameters of the 3D scene model.

[0094] The implementation of these 5 steps can be based on image editing applications (such as Photoshop, Substance Painter, etc.).

[0095] Furthermore, in this embodiment, to make the ornaments (or non-fixed objects, such as potted plants, desktop decorations, etc.) in the virtual 3D scene indicated by the 3D scene data replaceable and addable, the user can also input multiple perspective images of the target ornament into the live video generation device, so that the live video generation device can generate a 3D model of the target ornament. The specific generation method can be the same as the generation method of the 3D scene data corresponding to the target 3D scene in the foregoing embodiments, and will not be repeated here.

[0096] Based on the technical solutions corresponding to S501 and S502 above, the live video generation device can successfully obtain three-dimensional scene data, providing data support for the subsequent generation of live images with virtual three-dimensional scenes as the live background.

[0097] S303. Based on the digital human image generation model, process the audio data to obtain multiple frames of digital human images corresponding to the audio data.

[0098] Among them, the digital human image generation model has the ability to obtain multiple frames of digital human images corresponding to the audio data based on the audio data.

[0099] After the live video generation device acquires the audio data and 3D scene data, it first needs to determine the multi-frame digital human images corresponding to the audio data. At this point, the live video generation device can process the audio data based on a pre-trained digital human image generation model to obtain multi-frame digital human images. Each frame of these multi-frame digital human images corresponds to a specific audio segment in the audio data, and the lip movements of the digital human in the image corresponding to that audio segment match the pronunciation of that audio segment.

[0100] In one possible implementation, combining Figure 3 , refer to Figure 6 As shown, S303 may specifically include S3031 and S3032:

[0101] S3031. Based on the processing layer in the digital human image generation model, determine the multi-frame digital human mouth shape images corresponding to the audio data.

[0102] In this system, the digital lip-sync image may consist only of the digital human's head, while the lip movements on the digital face correspond to the pronunciation of a specific audio segment in the audio data. For example, a frame of the digital lip-sync image may look like this: Figure 7 As shown.

[0103] In this embodiment, since the audio samples used during the training of the digital human image generation model may differ from the actual audio data format, directly inputting the original audio data into the digital human image generation model may result in significant errors. Therefore, during S3031, the audio data is first converted to a standard format before being processed by the processing layer in the digital human image generation model. For example, the standard format could be mono audio with a 16kHz sampling rate.

[0104] In some embodiments, the processing layer in the digital human image generation model can be a lip-sync sub-model, which specifically describes the ability to generate multiple frames of digital human lip-sync images based on audio data. For example, the training process of this lip-sync sub-model is as follows:

[0105] S1, obtain the sample image and the corresponding sample audio.

[0106] Specifically, the sample images and audio can be obtained from sample video data. The sample video can be obtained by having a video model read text as required. The sample image is specifically a frontal image of the video model reading the text in the sample video, and the sample audio can be the audio data corresponding to that frontal image in the sample video.

[0107] S3, extract the sample audio features of the sample audio, and extract the sample face image from the sample image corresponding to the sample audio.

[0108] The extraction of sample audio features can be performed in any feasible manner, such as using a specific audio feature extraction algorithm or a pre-trained audio feature extraction model. This application does not impose specific limitations in this regard. In this embodiment, the sample audio features can specifically be a 20*256 two-dimensional array.

[0109] To extract sample faces from sample images, a face recognition library (dlib) can be used. Alternatively, any other feasible method can be used to extract sample faces from sample images.

[0110] S4 uses the sample audio features of the sample audio as training data and the sample face image corresponding to the sample audio as supervision information to train the initial sub-model to obtain the lip shape sub-model.

[0111] Specifically, S4 can include S41-S43:

[0112] S41. Input the sample audio features of the sample audio into the initial sub-model to obtain the predicted digital morphology image.

[0113] In the embodiments of this application, the specific framework of the initial sub-model can be any feasible neural network architecture, such as generative adversarial networks (GANs). Generally, the bias parameters in the initial sub-model are 0, while the weight parameters can be random values.

[0114] S42. Determine the loss value based on the predicted digital phagocytic image and sample face images.

[0115] S43. Iteratively update the initial sub-model based on the loss value until the loss value is less than the preset threshold, and then determine the latest updated initial sub-model as the lip shape sub-model.

[0116] The process of iteratively updating the initial sub-model based on the loss value involves adjusting the intrinsic parameters in the initial sub-model according to the loss value, and then re-inputting the sample audio features, sample facial key point data, and sample facial images with the mouth area covered into the updated initial sub-model. This process is repeated until the loss value is less than a preset threshold.

[0117] Of course, in practice, the lip shape sub-model can be any other feasible training method, and this application does not impose any specific restrictions on it.

[0118] S3032. Based on the fusion layer in the digital human image generation model, multiple frames of digital human mouth shape images are fused with digital human template images to obtain multiple frames of digital human images.

[0119] Specifically, the digital human template image can refer to a fixed image of a digital human, including the entire body of the digital human, for example... Figure 8 ( Figure 8 (Using only a half-body portrait as an example). The fusion layer of the digital human image generation model can first determine the specific region coordinates of the corresponding lip-sync image in the digital human template image. Based on these coordinates, the multiple frames of digital human lip-sync images can be fused into the corresponding region of the digital human template, thereby obtaining multiple frames of digital human images.

[0120] In addition, in practice, there can be multiple digital human template images, and the actions of the digital human in different digital human template images can be different. During the implementation of S3032, multiple digital human template images can be sorted according to certain rules, and at least one frame of digital human mouth shape image can be fused into each digital human template image to obtain multi-frame digital human images with richer digital human actions.

[0121] Of course, in practice, the digital human image generation model can also be a single model, which can be trained using any feasible training method. This application does not impose any specific restrictions on this.

[0122] Based on the technical solutions of S3031 and S3032 above, the live video generation device can successfully obtain multiple frames of digital human images corresponding to the audio data based on the digital human image generation model, so that subsequent live videos can be successfully generated based on the multiple frames of digital human images.

[0123] S304. Based on multiple frames of digital human images and 3D scene data, multiple frames of live images are rendered.

[0124] The live stream images include images of digital humans and target 3D scenes indicated by 3D scene models.

[0125] In this embodiment of the application, S304 may specifically be the live video generation device sending multiple frames of digital human images and 3D scene data to, for example, a device for generating live video. Figure 1 The server 02 shown in the figure renders the video and then returns multiple frames of live video images to the live video generation device.

[0126] In some embodiments, the 3D scene data includes: a 3D scene model and feature parameters of the 3D scene model; the feature parameters include: depth information of each object and coordinates of each object. At this time, combined with... Figure 3 , refer to Figure 9 As shown, S304 may specifically include S901-S903:

[0127] S901. Based on the depth information of each object in the 3D scene model, determine the first depth of the digital human in the virtual 3D scene in multiple live images.

[0128] In practice, refer to Figure 10 As shown, the virtual 3D scene indicated by the 3D scene model is a three-dimensional scene containing multiple objects. Therefore, digital human images cannot be arbitrarily placed in this virtual 3D scene for rendering. Based on this, the live video generation device needs to determine which positions are suitable for placing digital humans (specifically, digital human images) based on the depth information of each object in the 3D scene model, that is, to determine the first depth of the digital human in the virtual 3D scene in each frame of the live video.

[0129] S902. Based on the coordinates of each object in the 3D scene model, determine the first coordinate data of the first depth of the digital human in the virtual 3D scene in the multi-frame live broadcast images.

[0130] After determining the first depth of the digital human, refer to Figure 11 As shown, digital human 101 can still move under this first depth condition. To make the digital human's position in the virtual 3D scene more reasonable and convenient for viewers during live streaming, it is generally placed in the center of the virtual 3D scene. However, the specific determination of this center position needs to be based on the coordinates of objects in the virtual 3D scene to prevent the digital human and objects in the virtual 3D scene from overlapping, which would result in an unrealistic live stream image.

[0131] S903. Based on the first depth and first coordinate data of the digital human in the three-dimensional scene from the multi-frame live images, render the three-dimensional scene model and the multi-frame digital human images to obtain the multi-frame live images.

[0132] After determining the first depth and first coordinate data of the digital human in the virtual 3D scene, the live video generation device can successfully obtain real multi-frame live images.

[0133] In some embodiments, lighting varies in different virtual 3D scenes, and real lighting alters the RGB information (color values) in a person's image. The RGB information of the digital human in multi-frame images generated by the digital human image generation model is related to the RGB information of the digital human in the training samples of the model. Directly fusing the 3D scene model and the digital human image results in a disconnect between the digital human and the virtual 3D scene. Therefore, when fusing the 3D scene model and the digital human image, it is necessary to adjust the RGB information of the digital human image based on the lighting information of the 3D scene model to ensure a seamless and realistic rendering of the virtual 3D scene and the digital human image in the resulting live stream. Based on this, combined with... Figure 9 , refer to Figure 12 As shown, S903 may include S1201 and S1202:

[0134] S1201. Based on illumination information, adjust the color values ​​in multiple frames of digital human images to obtain multiple frames of target digital human images.

[0135] The feature parameters of the 3D scene model in the 3D scene data may include lighting information. Specifically, this lighting information may include the position, direction, and intensity of the light source in the virtual 3D scene.

[0136] Specifically, adjusting the color values ​​of multiple frames of digital human images based on illumination information can be done in any feasible way. For example, a pre-trained style transfer model can be used to adjust the color values ​​of the digital human image based on the illumination information. This style transfer model has the capability to adjust the color values ​​of human images using illumination information.

[0137] S1202. Based on the first depth and first coordinate data of the digital human in the three-dimensional scene in the multi-frame live images, render the three-dimensional scene model and the multi-frame target digital human images to obtain multi-frame live images.

[0138] In the embodiments of this application, the method by which the live video generation device renders the three-dimensional scene model and multi-frame target digital human images may include UI rendering and patch rendering.

[0139] In the live stream image rendered by the UI, the digital human image will be placed on the top layer, meaning that all actions of the digital human image will not affect other objects in the virtual 3D scene.

[0140] In the rendered live stream image, other objects (such as a table) can exist in front of the digital human image, and the digital human's movements will affect the display of the table. For example, if the digital human's arm stretches forward, it will block part of the table's display.

[0141] Based on the technical solutions in S1201 and S1202 described above, the live video generation device can generate more realistic live images. This results in more realistic and vivid live videos, improving the viewing experience for viewers, enhancing the live streaming effect, and increasing the retention rate of the live stream.

[0142] S305. Send audio data and the multi-frame live images to the live data generation device so that the live data generation device generates live video data based on the audio data and the multi-frame live images.

[0143] When live video data is generated by a live data generation device, to ensure smoother playback, audio data needs to be synchronized with multiple frames of live video images. This means that when each live image in the live video data is played, the audio segment played is the audio segment corresponding to that image. Based on this, combined with... Figure 3 , refer to Figure 13 As shown, after S305, the live streaming data generation device generates live streaming video data based on audio data and multiple frames of live streaming images. Specifically, this can include S1301 and S1302:

[0144] S1301, The live data generation device determines the correspondence between each frame of the live image in the multi-frame live image and multiple audio segments in the audio data.

[0145] In this embodiment, the multi-frame live stream images are obtained from multiple frames of digital human images and 3D scene rendering. Each frame of the live stream image corresponds to one digital human image. The multiple frames of digital human images are generated based on audio data, and their generation order is fixed. Therefore, based on the generation order of the multiple frames of digital human images, it is possible to know which segment of audio data each frame of digital human images corresponds to. For example, if the total duration of the audio data is 10 seconds and there are 250 frames of digital human images, then according to the generation time order, every ten frames of digital human images correspond to a 1-second audio segment in the audio data, and one frame of digital human images corresponds to a 40ms audio segment.

[0146] Based on the correspondence between each audio segment in the multi-frame digital human image and audio data, the correspondence between each frame of the live image and multiple audio segments in the audio data can be obtained.

[0147] In one possible implementation, after determining the correspondence between each frame of the live stream image and multiple audio segments in the audio data, a timestamp can be set for each frame of the live stream image. Similarly, for each audio segment as an audio frame, a timestamp corresponding to the live stream image can also be set for each audio frame. For example, if there are 250 frames of live stream image and the total duration of the audio data is 10 seconds, with each audio segment lasting 40ms, then the timestamps for the multiple live stream images can be 40ms, 80ms…10000ms. Likewise, the timestamps for the multiple audio frames can be 40ms, 80ms…10000ms.

[0148] S1302. The live streaming data generation device generates live streaming video data based on the corresponding relationship.

[0149] For example, taking a live video stream with 250 frames, a total audio data duration of 10 seconds, and each audio segment (or audio frame) lasting 40ms as an example, the live video generation device based on this correspondence can be configured with a clock drive. Every 40ms, the clock drive synchronizes one live video frame and one audio frame, creating a 40ms segment of audio and video data within the live video stream. Subsequently, during the live video generation device's live streaming of this data, the clock drive can synchronize and play one live video frame and one audio frame every 40ms.

[0150] In this embodiment of the application, in order to improve the audio quality in the live video data, the audio data can be converted into high-fidelity audio data when generating the live video data. For example, the high-fidelity format can be 48000 sampling rate single-channel 15-bit PCM (pulse code modulation) data.

[0151] The specific implementation of S1302 above can be carried out by the global mixer Mixer in the live data generation device.

[0152] In some embodiments, when a live streaming data generation device uses live video data for live streaming, it can convert the format of the live image in the live video data from RGBA to VP8 (a video compression format or video compression specification), and the audio data from PCM to Opus. Then, the live video data is transmitted to the live streaming terminal for display via a Real-Time Communication (RTC) protocol stack. In this embodiment, the RTC can specifically be WebRTC.

[0153] In other embodiments, the audio data may include audio corresponding to the digital human anchor and audio corresponding to the assistant anchor (or narrator). These two types of audio may be partially played simultaneously during the live stream. Therefore, when generating the live stream video, the live stream data generation device mixes the two types of audio that need to be played simultaneously according to a certain volume ratio, and then generates the final live stream video based on the correspondence between the audio data and the live stream image.

[0154] Based on the technical solutions corresponding to S1301 and S1302 above, the live video generation device can successfully generate live video data with synchronized audio and images, ensuring the smooth progress of subsequent live broadcasts.

[0155] In some embodiments, the positions of the digital human and the objects in the virtual 3D scene generally remain unchanged in each frame of the live video. However, due to limitations in the rendering method, the live video generation device may cause partial occlusion of the digital human and objects in the virtual 3D scene after rendering the live image. To eliminate these occlusions and provide a better user experience, the live video generation device can also display a preview interface of the live image after it is generated, and provide users with the function of adjusting the content in the live image. Based on this, referring to... Figure 14 As shown, the live video generation method provided in this application embodiment may further include S1401 and S1402:

[0156] S1401, Receive user's live room adjustment operation.

[0157] The adjustments made to the live stream room may include one or more of the following: adjusting decorations, adjusting lighting, adjusting the position of digital characters, etc.

[0158] For example, the operation of adjusting an ornament can be adding an ornament, deleting an ornament, or moving an ornament. The operation of adjusting lighting can be reducing the number of light sources, adding a light source, adjusting the direction of the light source, or adjusting the position of the light source. The operation of adjusting the position of a digital human can be adjusting the position of the digital human (including depth and position at the same depth).

[0159] S1402. In response to the adjustment operation of the live broadcast room, adjust the live broadcast image.

[0160] For example, taking the adjustment operation in the live broadcast room as a moving decoration as an example, refer to... Figure 15 As shown, users can drag the ornament 1501 to any possible location.

[0161] Taking adjusting the position of the digital human in the live stream as an example, refer to... Figure 16 As shown, users can drag Digital Man 1601 to move it left and right or forward and backward. Specifically, if Digital Man 1601 moves forward and backward, it will increase or decrease in size accordingly; it will increase in size when moving forward and decrease in size when moving backward.

[0162] Based on the technical solutions corresponding to S1401 and S1402 above, users can more easily adjust the content in the live image, avoid the possible occlusion of digital people and virtual 3D scenes in the live image generated by the live video generation device, so that the final live video better meets the user's needs and improves the user experience.

[0163] The foregoing mainly describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, it includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0164] This application embodiment can divide the live video generation device into functional modules based on the above method example. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.

[0165] Corresponding to the methods in the foregoing embodiments, this application also provides a live video generation apparatus. This live video generation apparatus is used to implement the aforementioned live video generation method. The functions of this live video generation apparatus can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions.

[0166] For example, Figure 17 A schematic diagram of a live video generation device is shown. This device can be the content expression device or a part thereof in the foregoing embodiments. (See Figure 17) Figure 17 As shown, the device includes an acquisition module 171, a processing module 172, a rendering module 173, and a sending module 174.

[0167] The system includes an acquisition module 171 for acquiring audio data and 3D scene data; a processing module 172 for processing the audio data acquired by the acquisition module 171 based on a digital human image generation model to obtain multiple frames of digital human images corresponding to the audio data; wherein the digital human image generation model has the ability to obtain multiple frames of digital human images corresponding to the audio data; a rendering module 173 for rendering multiple frames of live images based on the multiple frames of digital human images obtained by the processing module 172 and the 3D scene data acquired by the acquisition module 171; the live images include digital human images and a virtual 3D scene indicated by the 3D scene model; and a sending module 174 for sending the audio data acquired by the acquisition module 171 and the multiple frames of live images obtained by the rendering module 173 to a live data generation device, so that the live data generation device can generate live video data based on the audio data and the multiple frames of live images.

[0168] Optionally, the acquisition module 171 is specifically used to: acquire scene images of the target 3D scene from multiple perspectives; and generate 3D scene data based on the scene images from multiple perspectives.

[0169] Further optionally, the acquisition module 171 is specifically used to: generate a three-dimensional scene model to be determined based on scene images from multiple perspectives; receive adjustment operations, adjust the three-dimensional scene model to be determined in response to the adjustment operations to obtain a three-dimensional scene model, and determine the three-dimensional scene data by combining the three-dimensional scene model and the feature parameters of the three-dimensional scene model.

[0170] Optionally, the 3D scene data includes: a 3D scene model and feature parameters of the 3D scene model; the feature parameters include: depth information of each object and coordinates of each object; the rendering module 173 is specifically used to: determine the first depth of the digital human in the virtual 3D scene in the multi-frame live images based on the depth information of each object in the 3D scene model obtained by the acquisition module 171; determine the first coordinate data of the first depth of the digital human in the virtual 3D scene in the multi-frame live images based on the coordinates of each object in the 3D scene model; and render the 3D scene model and the multi-frame digital human images based on the first depth and first coordinate data of the digital human in the 3D scene in the multi-frame live images to obtain the multi-frame live images.

[0171] Further optionally, the three-dimensional scene data includes: a three-dimensional scene model and feature parameters of the three-dimensional scene model; the feature parameters also include lighting information; the rendering module 173 is specifically used to: adjust the color values ​​in the multi-frame digital human images based on the lighting information to obtain multi-frame target digital human images; and render the three-dimensional scene model and the multi-frame target digital human images based on the first depth and first coordinate data of the digital human in the three-dimensional scene in the multi-frame live images to obtain multi-frame live images.

[0172] Optionally, the processing module 172 is specifically used to: determine the multi-frame digital lip-shape images corresponding to the audio data acquired by the acquisition module 171 based on the processing layer in the digital human image generation model; and fuse the multi-frame digital lip-shape images with the digital human template image based on the fusion layer in the digital human image generation model to obtain multi-frame digital human images.

[0173] The effects achievable by the aforementioned live video generation device can be referenced from the effects of the live video generation method described in the foregoing embodiments, and will not be repeated here.

[0174] Figure 18 This is a structural diagram of an electronic device provided in an embodiment of this disclosure. Specifically, this electronic device can be the content expression device described in the foregoing embodiments. For example... Figure 18 As shown, the electronic device includes one or more processors 1801 and memory 1802.

[0175] The processor 1801 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.

[0176] The memory 1802 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 1701 may execute the program instructions to implement the live video generation methods of the various embodiments of this disclosure described above, or other desired functions.

[0177] In one example, the electronic device may also include an input device 1703 and an output device 1704, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0178] Of course, for the sake of simplicity, Figure 18 Only some of the components of the electronic device relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device may include any other suitable components depending on the specific application.

[0179] In addition to the methods and devices described above, embodiments of this disclosure may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the live video stream generation method described in the foregoing embodiments.

[0180] Computer program products can be written in any combination of one or more programming languages ​​to perform the operations of embodiments of this disclosure. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0181] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the live video stream generation method described in the foregoing embodiments.

[0182] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0183] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0184] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0185] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.

[0186] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.

[0187] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A method for generating live video, applied to a content expression device in a live streaming system, wherein the live streaming system further includes a live data generation device connected to the content expression device, characterized in that, The live streaming system includes a decision-making device, a process-driven device, a content expression device, and a live streaming data generation device; wherein, the decision-making device establishes a communication connection with the process-driven device and the content expression device, and the process-driven device and the content expression device establish a communication connection with the live streaming data generation device; The process-driven device is used to acquire user process requirement data and generate live streaming process data and live streaming rule data based on the user process requirement data. The content expression device is used to acquire user content demand data and generate live content data based on the content demand data; The live streaming data generation device is used to generate live streaming videos and perform live streaming based on the live streaming process data, the live streaming rule data, and the live streaming content data. The decision-making device is used to generate live decision data based on the playback data of the live video, and adjust the live process data and the live rule data, and / or adjust the live content data based on the live decision data; The method for generating the live video includes: Acquire audio data; Acquire scene images of the target 3D scene from multiple perspectives; the scene images are captured by a user terminal or a shooting device; Generate a three-dimensional scene model based on scene images from multiple perspectives; The system receives an adjustment operation, responds to the adjustment operation by adjusting a given 3D scene model to obtain a 3D scene model, and determines the 3D scene data by defining the 3D scene model and its feature parameters. The steps of adjusting the given 3D scene model to obtain the 3D scene model include: adjusting the geometric parameters of the given 3D scene model based on a first adjustment operation to obtain a first 3D scene model to be adjusted; adjusting the number of faces and details of the model in the first 3D scene model to be adjusted based on a design operation input in a 3D design application to obtain a second 3D scene model to be adjusted; adjusting the number of faces in the second 3D scene model to be adjusted based on a topology operation input in a 3D reconstruction application to obtain a third 3D scene model to be adjusted; performing a UV unwrapping operation on the third 3D scene model to obtain a 2D unfolded image of the third 3D scene model; and adjusting the texture material of the 2D unfolded image based on a retouching operation performed on the 2D unfolded image to obtain the 3D scene model and its feature parameters. The audio data is processed based on a digital human image generation model to obtain multiple frames of digital human images corresponding to the audio data; wherein, the digital human image generation model has the ability to obtain multiple frames of digital human images corresponding to the audio data based on the audio data. Based on the multi-frame digital human images and the 3D scene data, multi-frame live images are rendered; the live images include the virtual 3D scene indicated by the digital human images and the 3D scene data; The audio data and the multi-frame live images are sent to the live data generation device so that the live data generation device can generate live video data based on the audio data and the multi-frame live images; The 3D scene data includes: a 3D scene model and feature parameters of the 3D scene model; the feature parameters include: depth information of each object and coordinates of each object; the rendering of multiple frames of live-stream images based on the multiple frames of digital human images and the 3D scene data includes: Based on the depth information of each object in the three-dimensional scene model, determine the first depth of the digital human in the virtual three-dimensional scene in the multi-frame live broadcast images; Based on the coordinates of each object in the three-dimensional scene model, determine the first coordinate data of the first depth of the digital human in the virtual three-dimensional scene in the multi-frame live image; Based on the first depth and first coordinate data of the digital human in the three-dimensional scene in the multi-frame live images, the three-dimensional scene model and the multi-frame digital human images are rendered to obtain the multi-frame live images. The method further includes: Receive user's live stream adjustment request; the adjustment request includes adjusting the position of the digital human. In response to adjustments made in the live stream, the position and size of the digital figure in the live stream image are adjusted; the digital figure becomes larger when it moves forward and smaller when it moves backward.

2. The method according to claim 1, characterized in that, The 3D scene data includes: a 3D scene model and feature parameters of the 3D scene model; the feature parameters also include lighting information. The rendering of the 3D scene model and the multi-frame digital human images based on the first depth and first coordinate data of the digital human in the virtual 3D scene from the multi-frame live images includes: Based on the illumination information, the color values ​​in the multi-frame digital human images are adjusted to obtain multi-frame target digital human images; Based on the first depth and first coordinate data of the digital human in the three-dimensional scene in the multi-frame live images, the three-dimensional scene model and the multi-frame target digital human images are rendered to obtain multi-frame live images.

3. The method according to claim 1 or 2, characterized in that, The process of processing the audio data based on the digital human image generation model to obtain multiple frames of digital human images corresponding to the audio data includes: Based on the processing layer in the digital human image generation model, determine the multi-frame digital lip-shape images corresponding to the audio data; Based on the fusion layer in the digital human image generation model, the multi-frame digital human mouth shape images are fused with the digital human template images to obtain the multi-frame digital human images.

4. A live video generation apparatus, applied to a content expression device in a live streaming system, wherein the live streaming system further includes a live data generation device connected to the content expression device, characterized in that, include: The acquisition module is used to acquire audio data; The acquisition module is also used to acquire scene images of the target 3D scene from multiple perspectives; The scene images are captured by the user terminal or the shooting device; Generate a three-dimensional scene model based on scene images from multiple perspectives; The system receives an adjustment operation, responds to the adjustment operation by adjusting a given 3D scene model to obtain a 3D scene model, and determines the 3D scene data by defining the 3D scene model and its feature parameters. The steps of adjusting the given 3D scene model to obtain the 3D scene model include: adjusting the geometric parameters of the given 3D scene model based on a first adjustment operation to obtain a first 3D scene model to be adjusted; adjusting the number of faces and details of the model in the first 3D scene model to be adjusted based on a design operation input in a 3D design application to obtain a second 3D scene model to be adjusted; adjusting the number of faces in the second 3D scene model to be adjusted based on a topology operation input in a 3D reconstruction application to obtain a third 3D scene model to be adjusted; performing a UV unwrapping operation on the third 3D scene model to obtain a 2D unfolded image of the third 3D scene model; and adjusting the texture material of the 2D unfolded image based on a retouching operation performed on the 2D unfolded image to obtain the 3D scene model and its feature parameters. The processing module is used to process the audio data acquired by the acquisition module based on the digital human image generation model to obtain multiple frames of digital human images corresponding to the audio data; wherein, the digital human image generation model has the ability to obtain multiple frames of digital human images corresponding to the audio data based on the audio data; The rendering module is used to render multiple frames of live images based on the multiple frames of digital human images obtained by the processing module and the three-dimensional scene data obtained by the acquisition module; the live images include the digital human images and the virtual three-dimensional scene indicated by the three-dimensional scene model; The sending module is used to send the audio data acquired by the acquisition module and the multi-frame live images obtained by the rendering module to the live data generation device, so that the live data generation device can generate live video data based on the audio data and the multi-frame live images; The 3D scene data includes: a 3D scene model and its feature parameters; the feature parameters include: depth information of each object and coordinates of each object; the rendering module is specifically used to: determine the first depth of the digital human in the virtual 3D scene in the multi-frame live images based on the depth information of each object in the 3D scene model obtained by the acquisition module; determine the first coordinate data of the first depth of the digital human in the virtual 3D scene in the multi-frame live images based on the coordinates of each object in the 3D scene model; and render the 3D scene model and the multi-frame digital human images based on the first depth and first coordinate data of the digital human in the 3D scene in the multi-frame live images to obtain the multi-frame live images. The content presentation device is used to receive adjustments made by users in their live stream; these adjustments include adjusting the position of the digital human. The live streaming data generation device is used to adjust the position and size of the digital man in the live streaming image in response to the adjustment operation in the live streaming room; the digital man becomes larger when it moves forward and smaller when it moves backward.

5. An electronic device, characterized in that, include: A memory and a processor, the memory being used to store a computer program; the processor being used to cause the electronic device to perform the live video generation method as described in any one of claims 1-3 when executing the computer program.

6. A computer-readable storage medium, characterized in that, It includes at least one instruction that, when executed on an electronic device, causes the electronic device to perform the live video generation method as described in any one of claims 1-3.

Citation Information

Patent Citations

  • Data annotation method and device based on three-dimensional simulation scene and storage medium

    CN113223146A

  • Digital human video generation method and device, electronic equipment and storage medium

    CN113987269A

  • Cross-live-broadcasting-platform interactive virtual anchor implementation method and cross-live-broadcasting-platform interactive virtual anchor implementation system

    CN114302245A

  • Live broadcast method and device, storage medium, electronic equipment and product

    CN115442658A