A live broadcast background fusion method and system, a storage medium and a live broadcast device

By segmenting and converting recorded footage into 3D point clouds, and combining the 3D point clouds of real-time portraits with background images, the problem of unrealistic background replacement in live broadcasts was solved, achieving higher quality synthesis effects and a more immersive experience.

CN117750043BActive Publication Date: 2026-04-28ZHUHAI SHIXI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHUHAI SHIXI TECH CO LTD
Filing Date
2023-11-20
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Current live streaming technology does not produce realistic background replacement effects, requiring technicians to manually adjust parameters. Furthermore, the spatial relationships are lost after background replacement, resulting in unnatural effects.

Method used

By acquiring recorded footage and performing image segmentation, the data is converted into 3D point clouds. These points are then combined with the 3D point clouds of real-time human figures and background images for synthesis. The depth information of the 3D point clouds is used to ensure the realism of the relative position and distance between foreground objects and the environment.

Benefits of technology

It achieves higher quality composite effects, making live images more realistic, increasing the flexibility and interactivity of post-editing, and enhancing the realism and immersion of live content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117750043B_ABST
    Figure CN117750043B_ABST
Patent Text Reader

Abstract

The application discloses a live broadcast background fusion method and system, a storage medium and a live broadcast device, which are used for realizing higher quality of synthesis effect and enhancing the reality and immersion of live broadcast content. The method comprises the following steps: acquiring a recording and broadcasting material, and performing image segmentation on the recording and broadcasting material to obtain a foreground object and a background image in the recording and broadcasting material; performing image understanding on the foreground object, and converting the foreground object into a first 3D point cloud according to the result of image understanding; acquiring a real-time portrait, and acquiring a second 3D point cloud of the real-time portrait; and synthesizing a target live broadcast image according to the first 3D point cloud, the second 3D point cloud and the background image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and in particular to a method, system, storage medium, and live streaming device for background blending in live streaming. Background Technology

[0002] With the continuous advancement of digital technology and the growth of user demands, live streaming has emerged as a new form of communication and entertainment in people's daily lives. Virtual backgrounds are a common feature in live streaming, allowing broadcasters to create various scenes and atmospheres without altering the real-world environment, thus enhancing the visual effects and viewing experience.

[0003] To make viewers feel that the streamer is broadcasting in a real setting, the background is replaced using green screen keying, blending the streamer into a pre-recorded background video. However, if the background needs to be changed during the live stream, technicians need to readjust the camera angles for the new background video and adjust the streamer's image parameters, including adjusting skin tone and clothing color. Furthermore, current background replacement technology simply overlays the streamer's image onto the background layer, lacking spatial relationship, resulting in an unrealistic effect. Summary of the Invention

[0004] This application provides a background blending method, system, storage medium, and live streaming equipment for achieving higher quality composite effects and enhancing the realism and immersion of live streaming content.

[0005] The first aspect of this application provides a method for background blending in live streaming, including:

[0006] Acquire recorded video footage and perform image segmentation on the recorded video footage to obtain foreground objects and background images in the recorded video footage;

[0007] The foreground object is subjected to image understanding, and the foreground object is converted into a first 3D point cloud based on the image understanding result;

[0008] Acquire a real-time portrait and acquire a second 3D point cloud of the real-time portrait;

[0009] The target live image is synthesized based on the first 3D point cloud, the second 3D point cloud, and the background image.

[0010] Optionally, the step of performing image understanding on the foreground object and converting the foreground object into a first 3D point cloud based on the image understanding result includes:

[0011] Image understanding is performed on the foreground object to obtain its segmentation mask and depth information;

[0012] The foreground object is converted into 3D point cloud information based on the segmentation mask and depth information;

[0013] The first conversion coefficient is calculated based on the depth information, and the 3D point cloud information of the foreground object is scaled using the first conversion coefficient to obtain the first 3D point cloud.

[0014] Optionally, calculating the target conversion coefficient based on the depth information includes:

[0015] A first depth value is determined based on the depth information to indicate the center position of the foreground object.

[0016] Determine a second depth value for the center position of the foreground object based on the 3D point cloud information;

[0017] The first conversion coefficient is determined based on the ratio of the first depth value to the second depth value.

[0018] Optionally, acquiring real-time portraits includes:

[0019] The camera continuously captures N images of the anchor within a preset time period using a shooting device;

[0020] The target image at the intermediate moment is determined from N images of the anchor, and the real-time portrait is obtained by cutting out the target image.

[0021] Optionally, acquiring the second 3D point cloud of the real-time portrait includes:

[0022] The real-time portrait is converted into 3D point cloud information;

[0023] Determine the target distance between the broadcaster and the filming device;

[0024] The second conversion coefficient is calculated based on the target distance, and the 3D point cloud information of the real-time portrait is scaled using the second conversion coefficient to obtain the second 3D point cloud.

[0025] Optionally, when the foreground object contains a human figure, the step of synthesizing the target live-stream image based on the first 3D point cloud, the second 3D point cloud, and the background image includes:

[0026] In the first 3D point cloud, a target point cloud belonging to the human figure is determined, and the first center position of the target point cloud is determined;

[0027] Determine the second center position of the second 3D point cloud;

[0028] The target point cloud is replaced with the second 3D point cloud, and after the replacement, the second center position is consistent with the first center position.

[0029] The target live image is synthesized based on the replaced first 3D point cloud and the background image.

[0030] Optionally, before synthesizing the target live image based on the first 3D point cloud, the second 3D point cloud, and the background image, the method further includes:

[0031] Denoising and smoothing operations are performed on the first 3D point cloud and the second 3D point cloud.

[0032] A second aspect of this application provides a background blending system for live streaming, comprising:

[0033] An acquisition unit is used to acquire recorded video material and perform image segmentation on the recorded video material to obtain foreground objects and background images in the recorded video material;

[0034] The first conversion unit is used to perform image understanding on the foreground object and convert the foreground object into a first 3D point cloud based on the image understanding result.

[0035] The second conversion unit is used to acquire a real-time portrait and acquire a second 3D point cloud of the real-time portrait;

[0036] The synthesis unit is used to synthesize a target live image based on the first 3D point cloud, the second 3D point cloud, and the background image.

[0037] Optionally, the first conversion unit includes:

[0038] The processing module is used to perform image understanding on the foreground object to obtain the segmentation mask and depth information of the foreground object;

[0039] The conversion module is used to convert the foreground object into 3D point cloud information based on the segmentation mask and depth information;

[0040] The scaling module is used to calculate a first conversion coefficient based on the depth information, and to scale the 3D point cloud information of the foreground object using the first conversion coefficient to obtain a first 3D point cloud.

[0041] Optionally, the scaling module is specifically used for:

[0042] A first depth value is determined based on the depth information to indicate the center position of the foreground object.

[0043] Determine a second depth value for the center position of the foreground object based on the 3D point cloud information;

[0044] The first conversion coefficient is determined based on the ratio of the first depth value to the second depth value.

[0045] Optionally, the second conversion unit is specifically used for:

[0046] The camera continuously captures N images of the anchor within a preset time period using a shooting device;

[0047] The target image at the intermediate moment is determined from N images of the anchor, and the real-time portrait is obtained by cutting out the target image.

[0048] Optionally, the second conversion unit is further configured to: convert the real-time portrait into 3D point cloud information;

[0049] Determine the target distance between the broadcaster and the filming device;

[0050] The second conversion coefficient is calculated based on the target distance, and the 3D point cloud information of the real-time portrait is scaled using the second conversion coefficient to obtain the second 3D point cloud.

[0051] Optionally, when the foreground object contains a human figure, the compositing unit is specifically used for:

[0052] In the first 3D point cloud, a target point cloud belonging to the human figure is determined, and the first center position of the target point cloud is determined;

[0053] Determine the second center position of the second 3D point cloud;

[0054] The target point cloud is replaced with the second 3D point cloud, and after the replacement, the second center position is consistent with the first center position.

[0055] The target live image is synthesized based on the replaced first 3D point cloud and the background image.

[0056] Optionally, the system further includes:

[0057] The post-processing unit is used to perform noise reduction and smoothing operations on the first 3D point cloud and the second 3D point cloud.

[0058] A third aspect of this application provides a background blending device for live streaming, the device comprising:

[0059] Processor, memory, input / output units, and bus;

[0060] The processor is connected to the memory, the input / output unit, and the bus;

[0061] The memory stores a program, which the processor invokes to execute the first aspect and any optional live background blending method of the first aspect.

[0062] The fourth aspect of this application provides a computer-readable storage medium storing a program that, when executed on a computer, performs the first aspect and any optional method of background blending in live streaming.

[0063] The fifth aspect of this application provides a live streaming device, which includes an integrated or separate camera and a host, wherein the host performs the method as described in any of the first aspects during operation.

[0064] As can be seen from the above technical solutions, this application has the following advantages:

[0065] The recorded live stream footage is segmented and understood using a large visual model. Based on the image understanding results, foreground objects in the recorded footage are converted into a first 3D point cloud, and real-time portraits are converted into a second 3D point cloud. The first 3D point cloud of the foreground objects, the second 3D point cloud of the real-time portrait, and the background image from the recorded footage are then synthesized to obtain the target live stream image. Because the 3D point cloud contains depth information of the objects, the relative positions and distances of the foreground objects and real-time portraits with their surroundings are more realistic, resulting in a higher quality synthesis effect and a more realistic synthesized live stream image. Furthermore, the introduction of 3D point clouds increases the flexibility of post-editing, provides perspective changes and interactivity, enhances the realism and immersion of the live stream content, and provides users with a richer viewing experience. Attached Figure Description

[0066] To more clearly illustrate the technical solutions in this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0067] Figure 1 A schematic flowchart of an embodiment of the background blending method in live streaming provided in this application;

[0068] Figure 2 A schematic diagram of an embodiment of image segmentation in the live streaming background fusion method provided in this application;

[0069] Figure 3 A schematic flowchart of another embodiment of the background blending method in live streaming provided in this application;

[0070] Figure 4 A schematic diagram of the structure of an embodiment of the live streaming background blending system provided in this application;

[0071] Figure 5A schematic diagram of another embodiment of the live streaming background blending system provided in this application;

[0072] Figure 6 A schematic diagram of an embodiment of the background blending device for live streaming provided in this application. Detailed Implementation

[0073] This application provides a background blending method, system, storage medium, and live streaming equipment for achieving higher quality composite effects and enhancing the realism and immersion of live streaming content.

[0074] It should be noted that the background blending method for live streaming provided in this application can be applied to terminals as well as servers. For example, the terminal can be a live streaming device, smartphone, computer, tablet, smart TV, smartwatch, portable computer terminal, or a fixed terminal such as a desktop computer. For ease of explanation, this application uses the terminal as the implementing entity for illustrative purposes.

[0075] Please see Figure 1 , Figure 1 An embodiment of the background blending method in live streaming provided in this application includes:

[0076] 101. Obtain the recorded video footage and perform image segmentation on the recorded video footage to obtain the foreground objects and background images in the recorded video footage;

[0077] The terminal first needs to acquire pre-recorded footage, which can be images or videos. This footage can include various props, background sets, decorations, etc., to enrich the scene. It can also contain human images for later replacement with real-time human images during the fusion process. The terminal uses a large visual model to perform image segmentation on the pre-recorded footage. Image segmentation separates foreground objects and background images for subsequent image processing and compositing. Foreground objects in the footage mainly refer to objects or elements that occupy a prominent position and are clearly distinguishable from the background. These can be physical objects, such as furniture, decorations, tools, vehicles, etc., or human objects, such as anchors, actors, lecturers, etc. Please refer to [link to relevant documentation]. Figure 2 , Figure 2 This is a schematic diagram of image segmentation, which includes foreground objects such as clothes, shoes, and bags, as well as a background image of a wall.

[0078] In some specific embodiments, the large visual model can employ the Segment Anything Model (SAM). The Segment Anything Model consists of image segmentation tools, datasets, and models, and can be used for applications that require finding and segmenting any object in any image.

[0079] 102. Perform image understanding on the foreground object and convert the foreground object into a first 3D point cloud based on the image understanding results;

[0080] The terminal uses a deep learning model to perform image understanding of foreground objects and converts the foreground objects into a first 3D point cloud using the results of image understanding. Image understanding refers to using artificial intelligence to perform semantic understanding of images, analyzing what objects are present in the image, the relationships between objects, etc. A 3D point cloud is a collection of three-dimensional points used to represent the geometry and topology of objects, scenes, or structures in three-dimensional space. Each point in the point cloud has its own three-dimensional coordinates (X, Y, Z) to describe its position in space. Converting foreground objects into a first 3D point cloud can very accurately represent the external shape of the foreground objects, including details such as curves, edges, and surface bumps.

[0081] In some specific embodiments, the One-2-3-45 scheme can be used to convert the foreground object into the first 3D point cloud. That is, a multi-view image of the foreground object is generated using a 2D diffusion model, then 2D image features are extracted from the multi-view image, and then the 3D model of the foreground object is reconstructed to obtain the first 3D point cloud of the foreground object.

[0082] 103. Acquire the real-time portrait and obtain the second 3D point cloud of the real-time portrait;

[0083] After processing the foreground objects in the recorded footage, the terminal also needs to acquire real-time human images, which can be images of anchors, actors, lecturers, etc. The terminal also needs to acquire a second 3D point cloud of this real-time human image to capture the real-time position and movement of the human body.

[0084] It should be noted that when acquiring the second 3D point cloud of a real-time portrait, one can directly use a device with 3D imaging capabilities to capture the second 3D point cloud, or one can first capture a 2D real-time portrait with a regular camera, and then use the One-2-3-45 solution to convert the 2D real-time portrait into a second 3D point cloud. The specific acquisition method is not limited here.

[0085] 104. Synthesize the target live image based on the first 3D point cloud, the second 3D point cloud, and the background image.

[0086] The terminal synthesizes the target live-stream image based on the first 3D point cloud, the second 3D point cloud, and the background image. Specifically, the terminal uses the first 3D point cloud, the second 3D point cloud, and camera parameters (viewpoint, focal length, etc.) to project 3D points onto a 2D image space, generating the position of each point in the 3D point cloud on the background image. Then, the color information of each point is synthesized with the corresponding pixels on the background image to obtain the target live-stream image. The target live-stream image seamlessly integrates foreground objects, real-time human figures, and the background, transmitting the target live-stream image to the audience in real time and providing a more realistic viewing experience.

[0087] In some specific embodiments, after obtaining the first 3D point cloud of the foreground object and the second 3D point cloud of the real-time portrait, the terminal can first fuse the first and second 3D point clouds. Specifically, the first and second 3D point clouds are placed in the same coordinate system, ensuring that the second 3D point cloud covers the first 3D point cloud. Then, redundant points are removed (e.g., points in overlapping areas are removed). Based on the processing result, a unified point cloud data is created to obtain the target 3D point cloud. The target 3D point cloud fuses the 3D features of the foreground object and the real-time portrait, creating a unified 3D scene. The target 3D point cloud and the background image are then synthesized to obtain the target live image, thereby further improving the consistency and realism of subsequent synthesis.

[0088] In this embodiment, a large visual model is used to perform image segmentation and image understanding on the recorded broadcast footage. Based on the image understanding results, foreground objects in the recorded broadcast footage are converted into a first 3D point cloud, and real-time portraits captured in real-time are converted into a second 3D point cloud. The first 3D point cloud of the foreground objects, the second 3D point cloud of the real-time portrait, and the background image from the recorded broadcast footage are then synthesized to obtain the target live broadcast image. Because the 3D point cloud contains depth information of the objects, the relative positions and distances of the foreground objects and real-time portraits with their surrounding environment are more realistic, thus achieving a higher quality synthesis effect and making the synthesized live broadcast image more realistic. Furthermore, the introduction of 3D point clouds increases the flexibility of post-editing, provides perspective changes and interactivity, enhances the realism and immersion of the live broadcast content, and provides users with a richer viewing experience.

[0089] The background blending method for live streaming provided in this application is described in detail below. Please refer to [link / reference]. Figure 3 , Figure 3 An embodiment of the background blending method in live streaming provided in this application includes:

[0090] 301. Obtain the recorded video footage and perform image segmentation on the recorded video footage to obtain the foreground objects and background images in the recorded video footage;

[0091] In this embodiment, step 301 is similar to step 101 in the previous embodiment, and will not be described again here.

[0092] 302. Perform image understanding on the foreground object to obtain the segmentation mask and depth information of the foreground object;

[0093] The terminal uses image understanding to assign each pixel in an image to a different category or instance, thereby generating a segmentation mask and corresponding object category information. It then estimates the depth value of each pixel to the camera to obtain the depth information of the foreground object. After obtaining the segmentation mask and depth information of the foreground object, the terminal can also perform post-processing on the segmentation mask and depth information, such as noise removal, hole filling, or other image processing operations, to improve the quality of the results.

[0094] It should be noted that the terminal can use a single model to simultaneously perform image segmentation and image understanding of recorded video footage. Such a model can share the underlying feature extractor, allowing for a tighter integration of information between the two tasks, thus improving the model's efficiency and generalization ability. Specifically, after the terminal inputs the recorded video footage into a pre-trained image processing model, the model can directly output the background image, the segmentation mask of the foreground object, the depth information of the foreground object, and other information that may be used in subsequent synthesis processes, thereby improving processing efficiency.

[0095] 303. Convert the foreground object into 3D point cloud information based on the segmentation mask and depth information;

[0096] The terminal uses the results of image understanding—namely, the segmentation mask and depth information of the foreground object—to convert the foreground object into a 3D point cloud. Specifically, for each foreground pixel in the segmentation mask, the terminal calculates its 3D coordinates (X, Y, Z) using the depth information and pixel coordinates. If color information is available, it is associated with the corresponding 3D coordinates to add a color attribute to each point cloud point. This process is repeated until the entire foreground object area is covered. When converting the foreground object into 3D point cloud information, the object category information corresponding to that foreground object can also be considered, thereby enhancing the semantics and information richness of the 3D point cloud information.

[0097] 304. Calculate the first conversion coefficient based on the depth information, and scale the 3D point cloud information of the foreground object using the first conversion coefficient to obtain the first 3D point cloud;

[0098] The 3D point cloud information obtained in step 303 does not have an absolute size, but only a relative size consistent with the foreground object. At this time, the terminal can calculate a conversion coefficient e1 based on the depth information of the foreground object to convert the relative size into a more realistic absolute size. That is, the terminal calculates the first conversion coefficient based on the depth information of the foreground object, and then scales the 3D point cloud information of the foreground object using the first conversion coefficient to obtain the real first 3D point cloud.

[0099] Specifically, the first conversion factor is calculated as follows:

[0100] A. Determine the first depth value of the center position of the foreground object based on the depth information;

[0101] The terminal obtains the depth information of the foreground object through image understanding, and acquires the first depth value Z0 of the center position [X_center, Y_center] of the foreground object. The specific center position of the foreground object can be determined by the average value of all foreground object pixels, or by manual selection by the user. The specific method is not limited here.

[0102] It should be noted that if the actual depth value of the center position of the foreground object relative to the shooting device is known (marked in advance when shooting the recording footage), the first depth value can be determined directly based on the actual depth value.

[0103] B. Determine the second depth value of the center position of the foreground object based on the 3D point cloud information;

[0104] The terminal obtains the point cloud depth at the same center position [X_center, Y_center] through the 3D point cloud information of the foreground object, which is the second depth value Z1.

[0105] C. Determine the first conversion coefficient based on the ratio of the first depth value to the second depth value.

[0106] The terminal obtains the first conversion coefficient e1 through Z0 / Z1.

[0107] After calculating the first transformation coefficient e1, the 3D point cloud information of the foreground object is scaled using the first transformation coefficient e1. Assuming the coordinates of any point in the foreground object point cloud are [Xi,Yi,Zi], then the 3D coordinates after the transformation coefficient are e1*[Xi,Yi,Zi].

[0108] 305. Take N images of the anchor continuously within a preset time period using a shooting device, determine the target image at the middle moment among the N anchor images, and obtain a real-time portrait based on the target image by cutting out the image.

[0109] In this embodiment, a 2D real-time portrait is first captured using a regular camera, and then the One-2-3-45 scheme is used to convert the 2D real-time portrait into a second 3D point cloud. When acquiring the real-time portrait, images are continuously stored using the capturing device. N images of the broadcaster are continuously captured within a preset time period, and then one image from the middle of the sequence is used for image cutout to obtain the real-time portrait. This is done to avoid slight shaking of the person at the beginning of image storage, ensuring the accuracy of the cutout.

[0110] 306. Convert the real-time portrait into 3D point cloud information, determine the target distance from the anchor to the shooting device, calculate the second conversion coefficient based on the target distance, and scale the 3D point cloud information of the real-time portrait using the second conversion coefficient to obtain the second 3D point cloud;

[0111] The terminal can obtain real-time 3D point cloud information of the human face using the One-2-3-45 scheme. Each point has three-dimensional coordinate information, and these points can represent the surface of the real-time human face. At this point, the converted 3D point cloud information of the real-time human face does not have absolute dimensions, only relative dimensions. Therefore, the terminal needs to calculate another conversion coefficient e2 based on the anchor's depth information to convert this relative dimension into a more realistic absolute dimension. Specifically, first, determine the target distance Z2 from a certain position on the anchor's body to the shooting device. To ensure quality, the absolute deviation of this target distance Z2 should be controlled within 1%. Then, based on this target distance and the point cloud depth value Z3 at the corresponding position in the 3D point cloud information, calculate the second conversion coefficient e2, e2 = Z2 / Z3.

[0112] After calculating the second transformation coefficient e2, the 3D point cloud information of the real-time portrait is scaled using the second transformation coefficient e1. Assuming the coordinates of any point in the real-time portrait point cloud are [Xq,Yq,Zq], the 3D coordinates after transformation coefficients are e2*[Xq,Yq,Zq].

[0113] 307. Perform denoising and smoothing operations on the first and second 3D point clouds;

[0114] In this embodiment, before synthesizing the first 3D point cloud, the second 3D point cloud, and the background image, denoising and smoothing operations are required for both the first and second 3D point clouds. Denoising removes outliers and noise from both clouds to improve their quality. Methods include statistical filtering, Gaussian filtering, and distance-based filtering, which can eliminate isolated or outlier points that do not conform to the point cloud structure. Smoothing reduces high-frequency noise and irregularities in the point cloud. Smoothing can be achieved through sliding windowing, average value filtering, or curve-based methods, resulting in a more continuous and natural point cloud surface. These processing steps all contribute to improving the quality of the 3D point cloud, thereby enhancing the quality of the subsequent synthesized target live image.

[0115] 308. Synthesize the target live image based on the first 3D point cloud, the second 3D point cloud, and the background image.

[0116] The terminal synthesizes the target live-stream image based on the first 3D point cloud, the second 3D point cloud, and the background image. Specifically, the terminal uses the first 3D point cloud, the second 3D point cloud, and camera parameters (viewpoint, focal length, etc.) to project 3D points onto a 2D image space, generating the position of each point in the 3D point cloud on the background image. Then, the color information of each point is synthesized with the corresponding pixels on the background image to obtain the target live-stream image. The target live-stream image seamlessly integrates foreground objects, real-time human figures, and the background, transmitting the target live-stream image to the audience in real time and providing a more realistic viewing experience.

[0117] In some specific embodiments, when the broadcaster is already present in the pre-recorded footage, the broadcaster can refer to the human posture in the pre-recorded footage (ideally a simple standing or sitting posture) to achieve a more realistic background blending. In this case, the point cloud in the first 3D point cloud that is determined to belong to the human image can be directly replaced with the second 3D point cloud of the real-time human image before the live image is synthesized. Specifically, the terminal first determines the target point cloud in the first 3D point cloud that belongs to the human image and determines the first center position of the target point cloud (the center of the person in the footage); then it determines the second center position of the second 3D point cloud (the center of the real-time human image); finally, the target point cloud is replaced with the second 3D point cloud, and the second center position after replacement is kept consistent with the first center position to ensure the accuracy of the fusion of the first and second 3D point clouds. The replaced first 3D point cloud is then synthesized with the background image to obtain the target live image.

[0118] In this embodiment, a large visual model is used to perform image segmentation and image understanding on the recorded footage. Based on the image understanding results, foreground objects in the recorded footage are converted into a first 3D point cloud, and real-time portraits captured in real-time are converted into a second 3D point cloud. The first 3D point cloud of the foreground objects, the second 3D point cloud of the real-time portrait, and the background image from the recorded footage are then synthesized to obtain the target live image. A conversion coefficient is considered during the 3D point cloud conversion process to ensure that the 3D point cloud accurately reflects the size of the actual objects and portraits, making the relative positions and distances of the foreground objects and real-time portraits to their surroundings more realistic. This results in a higher quality synthesis effect and a more realistic synthesized live image. Furthermore, the introduction of 3D point clouds increases the flexibility of post-editing, provides perspective changes and interactivity, enhances the realism and immersion of the live content, and provides users with a richer viewing experience.

[0119] Please see Figure 4 , Figure 4 An embodiment of the live streaming background blending system provided in this application includes:

[0120] The acquisition unit 401 is used to acquire recorded broadcast materials and perform image segmentation on the recorded broadcast materials to obtain foreground objects and background images in the recorded broadcast materials;

[0121] The first conversion unit 402 is used to perform image understanding on the foreground object and convert the foreground object into a first 3D point cloud based on the result of the image understanding.

[0122] The second conversion unit 403 is used to acquire the real-time portrait and acquire the second 3D point cloud of the real-time portrait;

[0123] The synthesis unit 404 is used to synthesize a target live image based on the first 3D point cloud, the second 3D point cloud, and the background image.

[0124] In this embodiment, the acquisition unit 401 performs image segmentation on the recorded footage, and the first conversion unit 402 converts the foreground objects in the recorded footage into a first 3D point cloud. The second conversion unit 403 converts the real-time captured portrait into a second 3D point cloud. Then, the synthesis unit 404 synthesizes the first 3D point cloud of the foreground objects, the second 3D point cloud of the real-time portrait, and the background image from the recorded footage to obtain the target live image. Because the 3D point cloud contains depth information of the objects, the relative positions and distances of the foreground objects and the real-time portrait with their surrounding environment are more realistic, thereby achieving a higher quality synthesis effect and making the synthesized live image more realistic. Furthermore, the introduction of 3D point clouds can increase the flexibility of post-editing, provide perspective changes and interactivity, enhance the realism and immersion of the live content, and provide users with a richer viewing experience.

[0125] The following is a detailed description of the live streaming background blending system provided in this application. Please refer to [link / reference]. Figure 5 , Figure 5 Another embodiment of the live streaming background blending system provided in this application includes:

[0126] The acquisition unit 501 is used to acquire recorded broadcast materials and perform image segmentation on the recorded broadcast materials to obtain foreground objects and background images in the recorded broadcast materials;

[0127] The first conversion unit 502 is used to perform image understanding on the foreground object and convert the foreground object into a first 3D point cloud based on the result of the image understanding.

[0128] The second conversion unit 503 is used to acquire the real-time portrait and acquire the second 3D point cloud of the real-time portrait;

[0129] The synthesis unit 504 is used to synthesize a target live image based on the first 3D point cloud, the second 3D point cloud, and the background image.

[0130] Optionally, the first conversion unit 502 includes:

[0131] The processing module 5021 is used to perform image understanding on the foreground object and obtain the segmentation mask and depth information of the foreground object;

[0132] The conversion module 5022 is used to convert the foreground object into 3D point cloud information based on the segmentation mask and depth information;

[0133] The scaling module 5023 is used to calculate the first conversion coefficient based on the depth information, and to scale the 3D point cloud information of the foreground object using the first conversion coefficient to obtain the first 3D point cloud.

[0134] Optionally, the scaling module 5023 is specifically used for:

[0135] The first depth value is used to determine the center position of the foreground object based on the depth information;

[0136] Determine the second depth value of the center position of the foreground object based on 3D point cloud information;

[0137] The first conversion coefficient is determined based on the ratio of the first depth value to the second depth value.

[0138] Optionally, the second conversion unit 503 is specifically used for:

[0139] The camera continuously captures N images of the anchor within a preset time period using a shooting device;

[0140] Determine the target image at the intermediate moment from N images of the anchor, and obtain the real-time portrait based on the target image through image matting.

[0141] Optionally, the second conversion unit 503 is further configured to: convert real-time portraits into 3D point cloud information;

[0142] Determine the target distance between the broadcaster and the filming equipment;

[0143] The second conversion coefficient is calculated based on the target distance, and the 3D point cloud information of the real-time portrait is scaled using the second conversion coefficient to obtain the second 3D point cloud.

[0144] Optionally, when the foreground object contains a human figure, the compositing unit 504 is specifically used for:

[0145] In the first 3D point cloud, identify the target point cloud belonging to the human figure and determine the first center position of the target point cloud;

[0146] Determine the second center position of the second 3D point cloud;

[0147] Replace the target point cloud with the second 3D point cloud. After the replacement, the position of the second center is consistent with the position of the first center.

[0148] The target live image is synthesized based on the replaced first 3D point cloud and the background image.

[0149] Optionally, the system may also include:

[0150] The post-processing unit 505 is used to perform noise reduction and smoothing operations on the first 3D point cloud and the second 3D point cloud.

[0151] In this embodiment, the functions of each unit are the same as described above. Figure 3 The steps in the method embodiments shown correspond to those in the examples, and will not be repeated here.

[0152] This application also provides a background blending device for live streaming; please refer to [link / reference]. Figure 6 , Figure 6 One embodiment of the background blending apparatus for live streaming provided in this application includes:

[0153] Processor 601, memory 602, input / output unit 603, bus 604;

[0154] The processor 601 is connected to the memory 602, the input / output unit 603, and the bus 604;

[0155] The memory 602 stores a program, and the processor 601 calls the program to execute any of the above-mentioned background blending methods in live streaming.

[0156] This application also relates to a computer-readable storage medium storing a program, characterized in that, when the program is run on a computer, it causes the computer to execute any of the above-mentioned live streaming background blending methods.

[0157] This application also relates to a live streaming device, which includes an integrated or separate camera and a host, the host executing any of the above-described live streaming background blending methods during operation.

[0158] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0159] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0160] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0161] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0162] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A method for background blending in live streaming, characterized in that, The method includes: Acquire recorded video footage and perform image segmentation on the recorded video footage to obtain foreground objects and background images in the recorded video footage; The foreground object is subjected to image understanding, and the foreground object is converted into a first 3D point cloud based on the image understanding result; Acquire a real-time portrait and acquire a second 3D point cloud of the real-time portrait; A target live image is synthesized based on the first 3D point cloud, the second 3D point cloud, and the background image; The step of performing image understanding on the foreground object and converting the foreground object into a first 3D point cloud based on the image understanding result includes: Image understanding is performed on the foreground object to obtain its segmentation mask and depth information; The foreground object is converted into 3D point cloud information based on the segmentation mask and depth information; The first conversion coefficient is calculated based on the depth information, and the 3D point cloud information of the foreground object is scaled using the first conversion coefficient to obtain the first 3D point cloud. The acquisition of the second 3D point cloud of the real-time portrait includes: The real-time portrait is converted into 3D point cloud information; Determine the target distance between the broadcaster and the filming equipment; The second conversion coefficient is calculated based on the target distance, and the 3D point cloud information of the real-time portrait is scaled using the second conversion coefficient to obtain the second 3D point cloud.

2. The method according to claim 1, characterized in that, The calculation of the target conversion coefficient based on the depth information includes: A first depth value is determined based on the depth information to indicate the center position of the foreground object. Determine a second depth value for the center position of the foreground object based on the 3D point cloud information; The first conversion coefficient is determined based on the ratio of the first depth value to the second depth value.

3. The method according to claim 1, characterized in that, The acquisition of real-time human images includes: The camera continuously captures N images of the anchor within a preset time period using a shooting device; The target image at the intermediate moment is determined from N images of the anchor, and the real-time portrait is obtained by cutting out the target image.

4. The method according to claim 1, characterized in that, When the foreground object contains a human figure, the process of synthesizing the target live-stream image based on the first 3D point cloud, the second 3D point cloud, and the background image includes: In the first 3D point cloud, a target point cloud belonging to the human figure is determined, and the first center position of the target point cloud is determined; Determine the second center position of the second 3D point cloud; The target point cloud is replaced with the second 3D point cloud, and after the replacement, the second center position is consistent with the first center position. The target live image is synthesized based on the replaced first 3D point cloud and the background image.

5. The method according to any one of claims 1 to 4, characterized in that, Before synthesizing the target live image based on the first 3D point cloud, the second 3D point cloud, and the background image, the method further includes: Denoising and smoothing operations are performed on the first 3D point cloud and the second 3D point cloud.

6. A background blending system for live streaming, characterized in that, The system includes: An acquisition unit is used to acquire recorded video material and perform image segmentation on the recorded video material to obtain foreground objects and background images in the recorded video material; The first conversion unit is used to perform image understanding on the foreground object and convert the foreground object into a first 3D point cloud based on the image understanding result. The second conversion unit is used to acquire a real-time portrait and acquire a second 3D point cloud of the real-time portrait; A synthesis unit is used to synthesize a target live image based on the first 3D point cloud, the second 3D point cloud, and the background image; The first conversion unit includes: The processing module is used to perform image understanding on the foreground object to obtain the segmentation mask and depth information of the foreground object; The conversion module is used to convert the foreground object into 3D point cloud information based on the segmentation mask and depth information; The scaling module is used to calculate a first conversion coefficient based on the depth information, and to scale the 3D point cloud information of the foreground object using the first conversion coefficient to obtain a first 3D point cloud. The second conversion unit is further configured to: convert the real-time portrait into 3D point cloud information; Determine the target distance between the broadcaster and the filming equipment; The second conversion coefficient is calculated based on the target distance, and the 3D point cloud information of the real-time portrait is scaled using the second conversion coefficient to obtain the second 3D point cloud.

7. A background blending device for live streaming, characterized in that, The device includes: Processor, memory, input / output units, and bus; The processor is connected to the memory, the input / output unit, and the bus; The memory stores a program, which the processor invokes to perform the method as described in any one of claims 1 to 5.

8. A computer-readable storage medium having a program stored thereon, the program performing the method as described in any one of claims 1 to 5 when executed on a computer.

9. A live streaming device, characterized in that, The live streaming device includes a camera and a host unit that are either integrated or separate, and the host unit performs the method as described in any one of claims 1 to 5 when it is running.

Citation Information

Patent Citations

  • Live streaming method for real-time three-dimensional image display

    CN113891101A

  • Virtual light rendering method and device for live broadcasting room and storage medium

    CN116668732A