Media content processing method and device and program product

By determining the depth map and color depth fill map of a 2D image and combining it with 3D rendering technology, 3D/6-DOF media content is generated, solving the problems of high cost and storage difficulties, and realizing low-cost and high-efficiency 3D media content rendering and display.

CN121664967APending Publication Date: 2026-03-13BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

In existing technologies, the acquisition equipment for 3D/6-DOF media content is costly, generates large amounts of data, and is difficult to store, making it difficult to popularize and effectively render and display.

Method used

By determining the depth map of a 2D image, a foreground mask map and a color depth fill map are generated. Combined with 3D rendering technology, display characteristics of 3D/6-DOF media content are generated, and the 3D information of traditional 2D images is used for rendering.

Benefits of technology

No specialized acquisition equipment is required, reducing costs and storage pressure, expanding the audience reach of 3D/6DOF media content, and improving the viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121664967A_ABST
    Figure CN121664967A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a media content processing method and device and a program product, and the method comprises the steps: determining a depth image of a two-dimensional image, determining a foreground mask image of the two-dimensional image according to the depth image, and enabling the two-dimensional image to be a media frame obtained from a media stream; layering the depth map to obtain at least one depth layering map, and determining a color depth filling map of each depth layering map; and performing three-dimensional rendering under the first visual angle according to the two-dimensional image, the depth map, the foreground mask map and the depth filling map of each color, and generating media content displayed at the first visual angle. By utilizing the method, the media content of which the display view angle is changed along with the adjustment of the first view angle can be rendered and generated only by determining the three-dimensional information corresponding to the two-dimensional image and combining the first view angle without using special acquisition equipment; the cost, storage and end-side rendering and display pressure caused by playing of the three-dimensional / six-degree-of-freedom media content in related technologies are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer application technology, and in particular to a media content processing method, apparatus, and program product. Background Technology

[0002] With the widespread application of head-mounted virtual reality devices and glasses-free 3D devices, media content is shifting from traditional planar playback to 3D and higher degrees of freedom (such as 6DoF). For 3D / 6DoF media content playback, the main objective is the acquisition and generation of 3D / high-degree-of-freedom media content.

[0003] However, compared to capturing planar media content, both 3D and higher-resolution 6DoF capture require specialized hardware, which is more expensive and has higher configuration requirements, making it difficult to popularize among consumers. Furthermore, the greater freedom of media content capture necessitates massive amounts of data, leading to storage difficulties and placing immense pressure on media rendering and display. Summary of the Invention

[0004] This disclosure provides a media content processing method, apparatus, and program product.

[0005] In a first aspect, embodiments of this disclosure provide a media content processing method, the method comprising:

[0006] Determine the depth map of a two-dimensional image, and determine the foreground mask map of the two-dimensional image based on the depth map, wherein the two-dimensional image is a media frame obtained from a media stream;

[0007] The depth map is layered to obtain at least one depth layer map, and the color depth fill map of each depth layer map is determined.

[0008] Based on the two-dimensional image, the depth map, the foreground mask map, and each of the color depth fill maps, a three-dimensional rendering is performed from a first perspective to generate media content displayed from the first perspective.

[0009] Secondly, embodiments of this disclosure also provide a media content processing apparatus, the apparatus comprising:

[0010] The first determining module is used to determine the depth map of the two-dimensional image and determine the foreground mask map of the two-dimensional image based on the depth map. The two-dimensional image is a media frame obtained from a media stream.

[0011] The second determining module is used to perform layering processing on the depth map to obtain at least one depth layer map, and to determine the color depth fill map of each depth layer map.

[0012] The content rendering module is used to perform three-dimensional rendering based on the two-dimensional image, the depth map, the foreground mask map, and each of the color depth fill maps from a first perspective, to generate media content displayed from the first perspective.

[0013] Thirdly, this disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the media content processing method provided in any embodiment of this disclosure. Attached Figure Description

[0014] To more clearly illustrate the technical methods of the exemplary embodiments of this disclosure, the accompanying drawings used in describing the embodiments are briefly introduced below. Obviously, the accompanying drawings described are only a portion of the embodiments to be described in this disclosure, and not all of them. Those skilled in the art can derive other drawings from these drawings without any creative effort.

[0015] Figure 1 A schematic flowchart illustrating a media content processing method provided in an embodiment of this disclosure;

[0016] Figure 2 This is a schematic diagram of the structure of a media content processing device provided in an embodiment of the present disclosure;

[0017] Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of this disclosure. Detailed Implementation

[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0019] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0020] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0021] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules, or units, and are not used to limit the order of functions performed by these devices, modules, or units or their interdependencies. It should also be noted that the modifications of "a" and "a plurality of" mentioned in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0022] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0023] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0024] It is understood that before using the technical methods disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0025] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as the electronic device, application, server, or storage medium performing the operations of this disclosed technology.

[0026] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0027] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0028] It should be noted that the application scenarios involved in this embodiment can be scenarios involving real-shot images and / or videos, or live streaming scenarios. In these application scenarios, two-dimensional media content can be well displayed on terminal devices. However, with the development and widespread application of head-mounted VR devices, AR devices, and glasses-free 3D devices, the experience demand is shifting from two-dimensional media content to three-dimensional media content or higher degrees of freedom (such as 6 degrees of freedom).

[0029] However, it's understandable that if traditional monocular cameras are used to capture images or videos, the resulting and played-out content will also be traditional two-dimensional images or videos, depending on the technology used. Therefore, at present, experiencing live-action 3D / 6DoF media content requires specialized capture equipment, which incurs high costs and configuration requirements, making it unsuitable for the widespread adoption of 3D / 6DoF media content in the aforementioned scenarios. Furthermore, because 3D / 6DoF media content requires massive amounts of data, it also presents more challenges and problems for storage, rendering, and display.

[0030] Based on this Figure 1 This is a flowchart illustrating a media content processing method provided in an embodiment of the present disclosure. This embodiment is applicable to situations involving media content processing, particularly to situations where two-dimensional media content is determined to be in a three-dimensional or 6DoF media content format. The method can be executed by a media content processing device, which can be implemented through software and / or hardware and can be configured in a terminal and / or server to implement the media content processing method in this embodiment of the present disclosure.

[0031] like Figure 1 As shown, a media content processing method provided in this embodiment may include:

[0032] S101. Determine the depth map of the two-dimensional image, and determine the foreground mask map of the two-dimensional image based on the depth map. The two-dimensional image is a media frame obtained from the media stream.

[0033] In this embodiment, the media frame to be processed at the current execution moment can first be obtained from the acquired or received media stream, such as a live stream, a camera live stream, or a video file stream, and the media frame can be determined as a two-dimensional image with processing requirements.

[0034] The media content processing method provided in this embodiment can be specifically considered as a method for determining the three-dimensional information of two-dimensional images in a media stream, and generating playable video frames with the display characteristics of 3D / 6DoF media content through three-dimensional rendering processing based on the three-dimensional information. It is understood that the display characteristics of 3D / 6DoF media content generally involve adjusting the display perspective according to the pose of the execution device during video display, thereby achieving a more spatial display effect. In related technical implementations, the spatial sense presented during 3D / 6DoF media content playback is mainly due to the acquisition device capturing more scene image information.

[0035] Therefore, if we want a 2D image to be rendered into playable media content with a spatial display effect based on the corresponding 3D information, we can first perform depth estimation on the 2D image using this step in this embodiment to determine the depth map of the 2D image, thereby adding depth information to the 2D image. In the depth map, each pixel in the 2D image records a corresponding depth value. The value of each pixel represents the spatial distance information from the scene point corresponding to that pixel to the camera. Therefore, the depth map can be used to distinguish the distance of pixels from the camera. The specific method for determining the depth map of a 2D image can be: processing the 2D image using a trained neural network or a preset depth estimation algorithm to infer the depth value of each pixel in the 2D image, and using the depth value as pixel information to determine the depth map of the 2D image.

[0036] To ensure accurate locking of the core subject during 3D rendering and avoid issues such as high computational load, background obfuscation, and distorted occlusion relationships, this embodiment uses depth information from the depth map to evaluate the depth visibility confidence of each pixel in the depth domain. A higher depth visibility confidence indicates a higher likelihood of foreground, thus determining the foreground mask of the 2D image. The foreground mask can be understood as an image with the same size as a 2D video frame, used to distinguish the foreground from the background in a 2D image.

[0037] In an optional embodiment, if the two-dimensional image contains a pre-defined subject object, when determining the foreground mask map of the two-dimensional image, the matting result obtained by matting the subject object region can also be introduced. The foreground mask map is determined together based on the depth map and the matting result, further improving the accuracy of the determined foreground mask map.

[0038] S102. Perform layering processing on the depth map to obtain at least one depth layer map, and determine the color depth fill map of each depth layer map.

[0039] Understandably, the display characteristics of 3D / 6DoF media content generally involve adjusting the viewing angle according to the device's pose during video presentation. However, since 2D images only contain a single viewing angle, when the angle deviates from the original angle, background areas occluded by the foreground in the 2D image will be exposed. If these areas are not pre-filled with pixels, texture holes, such as black or blurry blocks, will appear during rendering, severely damaging the user experience and immersion. Furthermore, directly performing complete 3D modeling and rendering on a 2D image requires processing every single pixel, resulting in a massive number of pixels and extremely high computational costs. Due to the significant differences in computational power among different terminal devices, 3D / 6DoF media content will be unplayable on low-computing-power devices (such as mobile phones).

[0040] Therefore, in this embodiment, the depth map can be divided into multiple layers according to depth, such as foreground layer, midground layer and background layer, and the occlusion area caused by the spatial position relationship between layers can be identified. This determines the area in each layer that needs to be filled with color and generates the corresponding color and depth information through the filling algorithm, thereby determining the color depth filling map of each depth layer.

[0041] It should be noted that this embodiment mainly proposes the implementation logic of a media content processing method. The execution subject is not specifically limited. Different execution devices can be used as execution subjects to implement the logic execution of the media content processing method, depending on the application scenario in which the method is used. Alternatively, one or more execution devices can be used together to implement the logic execution of the media content processing method. In one optional embodiment, steps S101-S103 can be completed by a single execution device. In another optional embodiment, considering the data processing and computational workload of steps S101 and S102, they can be completed by an execution device with strong storage and computational capabilities. However, considering the desired spatial visual effect when the two-dimensional image ultimately presents three-dimensional / 6DoF media content, step S103 can be completed by an execution device connected to the previous execution device for playing the generated three-dimensional / 6DoF media content.

[0042] S103. Based on the two-dimensional image, depth map, foreground mask map, and various color depth fill maps, perform three-dimensional rendering from the first perspective to generate media content displayed from the first perspective.

[0043] Understandably, two-dimensional images are typically flat scenes with a fixed perspective, such as when viewing a photograph, where the perspective is fixed at the photographer's position. A first-person perspective, on the other hand, can be considered the perspective of a pre-set virtual camera, which is dynamically changing. The position of this virtual camera is controlled by the device's sensors, simulating human eye movement. For example, for devices like mobile phones, a gyroscope is used to control the movement of the virtual camera; for VR headsets, the user's head position (which the headset can obtain) is used for control.

[0044] 3D rendering in a first-person perspective can be understood as the rendering process following the viewer's first-person viewpoint, achieving 3D deformation and mapping in any viewing direction to create a 3D stereoscopic effect from different perspectives. This transforms the viewer's experience from viewing a flat plane to feeling as if they are in a 3D scene, with the scene's presentation angle completely following the viewer's first-person perspective. For example, when viewing an object from the side, you would see the object's side profile. By performing 3D rendering in a first-person perspective, the image from the front view can be transformed into an image from the corresponding side view.

[0045] Based on this, to ensure the effectiveness and user experience of 3D rendering, in this embodiment, the first viewpoint can be determined based on device sensors, and 3D rendering can be performed based on the 2D image, depth map, foreground mask map, and various color depth fill maps under the first viewpoint. The 2D image can provide the basic color texture for rendering; the depth map can provide the "spatial distance information" of each pixel; the foreground mask map can lock the foreground to ensure that the foreground can block the background during rendering, rather than the background penetrating the foreground; the various depth layer maps can provide the spatial hierarchy under the first viewpoint; and the various color depth fill maps can fill the occluded areas exposed when the first viewpoint is switched, thereby generating media content displayed from the first viewpoint. Here, media content can be considered as content with 3D / 6DoF media content display characteristics.

[0046] For example, when the device for performing 3D rendering is a mobile terminal such as a mobile phone, the first viewpoint can be determined by the gyroscope deployed in the mobile terminal; when the device for performing 3D rendering is a VR terminal, the first viewpoint can be determined by the head-mounted device sensor deployed in the VR terminal.

[0047] This embodiment provides a media content processing method. It determines a depth map of a two-dimensional image, and then determines a foreground mask map of the two-dimensional image based on the depth map. The two-dimensional image is a media frame acquired from a media stream. The depth map is layered to obtain at least one depth layer map, and a color depth fill map is determined for each depth layer map. Based on the two-dimensional image, depth map, foreground mask map, and color depth fill maps, three-dimensional rendering is performed from a first-viewpoint to generate media content displayed from that perspective. Using this method, two-dimensional images acquired by a monocular image acquisition device can be reused. By determining the corresponding three-dimensional information of the two-dimensional image and combining it with a first-viewpoint perspective and three-dimensional rendering technology, media content corresponding to the two-dimensional image and displayed from a first-viewpoint can be generated. This ensures that the media content can mimic the display characteristics of three-dimensional / six-degree-of-freedom media content, expanding gameplay, enriching the three-dimensional / six-degree-of-freedom media content, and improving the viewing experience. Compared to other technologies for playing 3D / 6DOF media content, this solution eliminates the need for specialized acquisition equipment. By simply determining the 3D information of a traditional 2D image and performing 3D rendering based on that information and a first-person perspective before playback, the solution ensures that the acquired media content visually changes its playback perspective according to the device's pose. This avoids the cost and storage issues associated with other technologies for playing 3D / 6DOF media content, while also reducing the rendering and display burden and thus expanding the audience reach of 3D / 6DOF media content.

[0048] As a first optional embodiment of this example, based on the above embodiment, the step of determining the foreground mask map of the two-dimensional image according to the depth map can be optimized as follows:

[0049] a1) Normalize the depth map within a set floating-point range to generate a disparity map of the depth map.

[0050] In this embodiment, each pixel in the depth map can be normalized, and the normalized value obtained after normalization can be used as the disparity value of the pixel. A disparity map can be constructed using the disparity values ​​corresponding to each pixel. For example, the floating-point range can be set to [0,1].

[0051] In an optional embodiment, after generating the disparity map, the disparity map can also be filtered, such as by median filtering, to suppress discrete noise in the disparity map and improve its accuracy and effectiveness.

[0052] b1) Without performing image matting on the 2D image, determine the depth visibility confidence map of the 2D image based on the disparity map, and use the depth visibility confidence map as the foreground mask map.

[0053] It is understandable that if no object requiring image matting is preset, or if the object is preset but does not exist in the 2D image, then image matting will not be performed on the 2D image. In this embodiment, when no image matting is performed on the 2D image, the depth visibility confidence map of the 2D image is determined solely based on the disparity map. The depth visibility confidence map includes the depth visibility confidence of each pixel in the 2D image, which can be considered as the probability value of whether a pixel is reliably visible in the depth domain. Since the depth visibility confidence can characterize the reliable visibility of a pixel, the depth visibility confidence map constructed from the depth visibility confidence can be determined as the foreground mask map, used to distinguish between foreground and background regions in the image.

[0054] As one implementation method, the determination of the depth visibility confidence map of the 2D image based on the disparity map can be further optimized into the following steps:

[0055] b11) Perform gradient calculation on the pixels in the disparity map to obtain the first gradient magnitude map of the disparity map, and obtain the second gradient magnitude map after smoothing the first gradient magnitude map.

[0056] In this embodiment, the gradient values ​​of each pixel in the disparity map in both the horizontal and vertical directions can be determined using operators for edge detection, such as the 3×3 Scharr operator. These gradient values ​​characterize the rate of change of disparity values; therefore, a larger gradient value indicates a steeper depth change at that pixel. For example, at the boundary between a foreground object and the background, the disparity value changes abruptly, resulting in a high gradient value. The first gradient magnitude map can be considered as an image obtained by integrating the intensity of gradients in the horizontal and vertical directions, directly reflecting the degree of depth change in each region of the disparity map.

[0057] However, the first gradient magnitude map may contain spurious gradients caused by image texture noise, such as fine lines on the object's surface or depth estimation errors. These spurious gradients can interfere with subsequent depth visibility judgments, leading to these noise points being misclassified as steep depth regions. Therefore, this step can also smooth the first gradient magnitude map to obtain a processed second gradient magnitude map. Smoothing makes the gradient values ​​of adjacent pixels more balanced, preserving true depth boundaries (such as obvious gradient differences between foreground and background) while eliminating minor gradient fluctuations caused by noise (such as meaningless gradients generated by object surface textures). For example, global normalization and Gaussian smoothing can be performed on each pixel in the first gradient magnitude map.

[0058] b12) Determine the depth visibility confidence level corresponding to each gradient magnitude in the second gradient magnitude map, and generate a depth visibility confidence map of the two-dimensional image based on each depth visibility confidence level.

[0059] In this embodiment, the gradient magnitudes in the second gradient magnitude map, which reflect the steepness of depth changes, are converted into probability values ​​to measure whether a pixel is reliably visible in the depth domain. These probability values ​​can be considered as depth visibility confidence scores. By summing the depth visibility confidence scores corresponding to each pixel, a depth visibility confidence map of the two-dimensional image can be generated. This depth visibility confidence map can be used to indicate which pixels have reliable depth.

[0060] It's important to note that gradient magnitude and the corresponding depth visibility confidence are negatively correlated. A smaller gradient magnitude corresponds to a larger depth visibility confidence, meaning the pixel corresponding to that gradient magnitude has higher visibility in the depth domain, and thus a lower probability of being occluded by other pixels during 3D rendering. Conversely, a larger gradient magnitude corresponds to a smaller depth visibility confidence, meaning the pixel corresponding to that gradient magnitude has lower visibility in the depth domain, and thus a higher probability of being occluded by other pixels during 3D rendering.

[0061] In an optional embodiment, the relationship between gradient magnitude and corresponding depth visibility confidence can be represented by an exponential decay function. Specifically, for each pixel, the depth visibility confidence of that pixel is equal to the target value raised to the power of the natural constant e. The target value is determined by multiplying the decay coefficient by the square of the gradient magnitude and taking the negative of the product as the target value. The decay coefficient is an empirical parameter used to control the strength of the negative correlation between gradient and visibility. A larger decay coefficient results in a faster decrease in depth visibility confidence for the same gradient magnitude; a smaller decay coefficient results in a slower decrease in depth visibility confidence for the same gradient magnitude. This can be flexibly adjusted according to different scenarios. The depth visibility confidence ranges from (0,1], with a larger value indicating higher visibility of the corresponding pixel in the depth domain.

[0062] The above-described technical solution in this embodiment determines the gradient magnitude of pixels in the disparity map and converts each gradient magnitude into depth visibility confidence, thereby identifying pixels with high visibility in the depth domain and providing support for subsequent accurate 3D rendering.

[0063] c1) When performing image matting on a two-dimensional image, obtain the first matting result for the first object in the two-dimensional image.

[0064] Optionally, the first object can be considered as a character object.

[0065] In this embodiment, when a two-dimensional image contains a human figure, the two-dimensional image can be processed by a trained matting model or a preset matting algorithm to separate the human figure region from the two-dimensional image and obtain a mask image for the first object in the two-dimensional image. The mask image of the first object is determined as the first matting result, that is, the first matting result is essentially a human figure mask image.

[0066] d1) Determine the disparity occlusion map of the two-dimensional image based on the disparity map, and determine the foreground mask map based on the first matting result, the disparity occlusion map and the depth visibility confidence map.

[0067] Among them, the parallax occlusion map can be considered as a map used to characterize the probability that a pixel is occluded.

[0068] In this embodiment, the disparity difference between neighboring pixels and the current pixel can be determined based on each pixel in the disparity map under different set directions and different time lengths. The disparity occlusion intensity of each pixel is then determined based on these disparity differences. The disparity occlusion intensity can be understood as a metric representing the probability of similar points being occluded. A disparity occlusion map is then determined based on the disparity occlusion intensity of each pixel in the two-dimensional image. Next, by combining the first matting result, the disparity occlusion map, and the depth visibility confidence map, the human object is initially located based on the first matting result. The occlusion status of pixels is determined based on the disparity occlusion map, and the visibility of pixels in the depth domain is determined based on the depth visibility confidence map, thereby determining the foreground mask map. This solves problems such as blurred human boundaries and misjudgment of the background in the foreground mask map.

[0069] In this embodiment, as one implementation, the step of determining the disparity occlusion map of the two-dimensional image based on the disparity map can be further optimized into the following steps:

[0070] a2) In at least one set direction, perform neighborhood difference processing on the pixels in the disparity map with at least one set step size to obtain the joint difference value of the pixels under each set direction and set step size combination.

[0071] For example, for each pixel in the disparity map, neighborhood differencing can be performed sequentially in the four main horizontal and vertical directions with multiple set step sizes, calculating the disparity difference between the neighboring pixels and the pixel at each step size. Furthermore, a distance penalty coefficient can be introduced to minimize the occlusion effect of more distant neighboring pixels on the pixel. Based on the disparity difference determined for the pixel at the corresponding step size in each set direction, combined with the distance penalty coefficient, the joint difference value corresponding to the pixel in each set direction and set step size combination is determined. The specific method for determining the joint difference value corresponding to the pixel in a set direction and set step size combination can be expressed as follows:

[0072] ;

[0073] in, For pixels In a set direction and a set step size The joint difference value corresponding to the combination; For disparity map pixels in ; For disparity map In a set direction and pixel Distance set step size The neighboring pixels; This is an empirical coefficient. This constitutes the distance penalty coefficient.

[0074] b2) The maximum value among the joint difference values ​​is determined as the maximum difference value of the pixel, and the disparity occlusion intensity of the pixel is determined based on the maximum difference value.

[0075] In this embodiment, the joint difference values ​​corresponding to each pixel under all combinations of direction and step size are obtained. The maximum value is determined from all joint difference values, and the maximum value is determined as the maximum difference value of the pixel. The disparity occlusion intensity of the pixel can be determined based on the maximum difference value by: performing a nonlinear mapping function, such as the tanh function, on the maximum difference value and truncating negative values, that is, converting the maximum difference value into a value between [0,1], thereby obtaining the disparity occlusion intensity of the pixel. The closer the disparity occlusion intensity is to 0, the higher the probability that the pixel is not occluded (mostly the first object region or the unoccluded background region). The closer the disparity occlusion intensity is to 1, the higher the probability that the pixel is occluded by other pixels (mostly the region in the background blocked by the first object).

[0076] c2) Generate a disparity occlusion map of the two-dimensional image based on the disparity occlusion intensity of each pixel.

[0077] In this embodiment, a disparity occlusion map of a two-dimensional image can be generated based on information such as the position of each pixel in the two-dimensional image and the disparity occlusion intensity of each pixel.

[0078] In this embodiment, a disparity occlusion map is generated by determining the joint difference value of each pixel under each direction and step size combination, and determining the disparity occlusion intensity of the pixel based on the maximum difference value. This enables accurate identification of the occluded pixel region based on the disparity occlusion map, avoiding subsequent misjudgment of the blocked background as the foreground.

[0079] In this embodiment, as another implementation, the step of determining the foreground mask map based on the first matting result, the disparity occlusion map, and the depth visibility confidence map can be further optimized into the following steps:

[0080] a3) Normalize the first cutout result within a set floating-point range to generate a second cutout result based on the first cutout result.

[0081] In this embodiment, after obtaining the first matting result, the first matting result can be normalized to the [0,1] interval to generate a second matting result that clearly defines the first object area and the background area.

[0082] b3) Based on the set structuring elements, the pixels in the second cutout result are dilated to obtain the dilated map of the second cutout result.

[0083] In this embodiment, the pixels in the second cutout result are dilated to slightly expand the range of the first object area, covering the narrow edge area of ​​the first object, such as the blurred boundary area between hair strands, clothing edges and the background.

[0084] c3) Determine the difference map between the expansion map and the second matting result, and determine the edge narrow band region of the first object from the difference map.

[0085] In this embodiment, the difference between the expansion map and the second matting result is determined as a difference map. Based on the difference map, the narrow edge region of the first object is determined. This narrow edge region is the area where the first object is most likely to be confused with the background. For example, in the difference map, only the pixels in the narrow edge region outside the edge of the first object have a pixel value of 1, while the pixels in other regions, such as the pure first object region and the pure background, have a pixel value of 0.

[0086] d3) Based on the parallax occlusion map, determine the suppression weights of pixels in the narrow edge region, update the depth visibility confidence map based on the suppression weights, and use the updated depth visibility confidence map as the foreground mask map.

[0087] In this embodiment, the occlusion relationship of the edge region of the first object is corrected and refined based on the disparity occlusion map and the edge narrowband region. The method for determining the suppression weight of pixels within the edge narrowband region based on the disparity occlusion map can be as follows: determine the disparity occlusion intensity corresponding to the pixel in the disparity occlusion map; determine the suppression weight of the pixel based on the product of the complement of the disparity occlusion intensity and the difference map. The method for updating the depth visibility confidence map based on the suppression weight can be as follows: for each pixel in the depth visibility confidence map, determine the product between the depth visibility confidence of the pixel and the complement of the suppression weight, determine the product value as the new depth visibility confidence corresponding to the pixel, and update the depth visibility confidence map based on the new depth visibility confidence.

[0088] For example, updating the depth-visible confidence graph based on suppression weights can be represented as follows:

[0089] ;

[0090] in, This is the updated depth-visibility confidence map; This is the depth-visibility confidence plot before the update; This represents the parallax occlusion intensity in the parallax occlusion map. It is a difference plot; To suppress weights.

[0091] The above formula shows that the value in the difference plot is only 1 in the narrow edge region, meaning it is only preserved in the narrow edge region. The value is 0 in all other regions. Within the narrow edge band region, The closer the value is to 1 (indicating a higher probability that a pixel is not occluded), the closer the suppression weight is to 1. The closer the confidence level is to 0 (indicating a higher probability of pixel occlusion), the closer the suppression weight is to 0. Therefore, the purpose of this step is to reduce the depth visibility confidence of unoccluded pixels within the narrow edge region to avoid misclassifying edge background as foreground; while retaining the depth visibility confidence of pixels outside the narrow edge region; and for occluded pixels within the narrow edge region, their depth visibility confidence can be retained because these pixels, due to their steep depth, have sufficiently high disparity occlusion intensity in the disparity occlusion map and will ultimately be accurately classified as background. By updating the depth visibility confidence map, the updated map can accurately distinguish between the foreground and background of the first object, thus the updated depth visibility confidence map can be used as the foreground mask map.

[0092] As a second optional embodiment of this example, based on the above embodiment, the depth map can be layered to obtain at least one depth layer map, specifically optimized as follows:

[0093] a4) Use a deep clustering strategy to determine at least one first depth layer map of the depth map, and obtain the depth interval and rectangular bounding box corresponding to each first depth layer map.

[0094] The first depth layer map can be understood as an initial depth layer map determined by simply splitting the depth map according to depth values ​​based on a preset number of layers. For example, the method of determining at least one first depth layer map of the depth map using a depth clustering strategy can be as follows: perform depth clustering on the depth image using a depth clustering strategy to obtain a preset number of first depth layer maps, wherein the first depth layer map of each layer is jointly determined by the corresponding lower depth bound and upper depth bound.

[0095] In this embodiment, the depth range corresponding to each first depth layer can be obtained based on the lower and upper depth bounds of each first depth layer. Pixels within each first depth layer need to satisfy a basic spatial relationship, that is, the depth of a pixel in the depth map that is within the first depth layer is greater than the lower depth bound and less than the upper depth bound.

[0096] As described above, after depth layering, a minimum rectangular region can be defined as a bounding box for each first depth layer image. This bounding box can completely enclose all pixels in the first depth layer image without any extra blank areas. At the same time, it can clearly define the spatial location range of the first depth layer image in the two-dimensional image. That is, the core function of the bounding box is to limit the processing boundary of the first depth layer image.

[0097] b4) For each first depth layer map except the first first depth layer map, determine the occlusion area from the depth map based on the lower limit value of the corresponding depth interval and the rectangular bounding box, and obtain the second depth layer map after filtering out the occlusion area from the first depth layer map.

[0098] It should be noted that since the first depth layer image is the foremost image, it is not occluded. However, starting from the second depth layer image, when performing layering processing on the depth map, in addition to satisfying the basic spatial relationships, it is also necessary to include the occlusion areas of the current depth layer image by the previous depth layer image.

[0099] Therefore, in this embodiment, in order to accurately determine pixels belonging only to each first depth layer map, it is necessary to identify the occlusion region and filter it out in the first depth layer map to obtain a second depth layer map, which retains the unoccluded region. Specifically, for each first depth layer map other than the first one, the occlusion region within the first depth layer map can be determined from the depth map based on the lower depth limit and the bounding box in the corresponding depth interval of the first depth layer map, and then filtered out. ( Occlusion region of the first depth layer map The method can be expressed as:

[0100] ;

[0101] in, This is a depth map; For the first The lower bound of the depth of the first depth layer map; For the first The rectangular bounding box of the first depth layer map; To negate (or counteract). Through... It can be determined that the first The region corresponding to the first depth layer map in the previous first depth layer map will be defined as the region corresponding to the rectangular bounding box region in the previous first depth layer map. The occlusion area of ​​the first depth layer map.

[0102] c4) The first first depth layer map is determined as a depth layer map, and each second depth layer map is determined as a depth layer map.

[0103] In this embodiment, the first first depth layer map can be directly determined as a depth layer map, while for each of the other first depth layer maps, the corresponding second depth layer maps are determined as depth layer maps. This avoids interference between the depth layer maps, improves the accuracy of depth layer map division, and ensures the independence of each depth layer map.

[0104] As a third optional embodiment of this embodiment, based on the above embodiments, the determination of the color depth fill map of each depth layer map can be specified as follows:

[0105] a5) Obtain the disparity occlusion map determined by the relative depth map.

[0106] b5) Perform a logical AND operation between the parallax occlusion map and each depth layer map to obtain the processed layered occlusion map.

[0107] In this embodiment, for each depth layer map, a logical AND operation is performed between the parallax occlusion map and the depth layer map to determine the occluded regions that belong to the current depth layer map. The main characteristics of the pixels in these regions are that they are exposed when the viewpoint changes, and they originally did not have depth and color data in the current depth layer map. Based on these regions, a layered mask map is determined, which is used to represent regions that require filling.

[0108] c5) Based on the two-dimensional image and the depth map, fill each layer of the mask map with color data and depth data respectively to obtain the color depth fill map of each depth layer map.

[0109] Optionally, this step can be performed by a trained RGBD filling network. Specifically, the layered masking image, the corresponding 2D image, and the depth map can be input into the trained RGBD filling network. The RGBD filling network generates and fills the occluded areas in the layered masking image that require filling with color and depth data, and outputs a color depth filling map of the depth layered image.

[0110] In one alternative embodiment, a padding training dataset can be constructed and used to train an initial RGBD padding network. This allows the padding training data to learn to generate natural and depth-logical missing color and depth data based on existing color data, depth information, and occlusion relationships. The padding training dataset can be constructed by first determining the corresponding training depth map and training disparity occlusion map from a large number of training 2D images, and then organizing them into training data with occluded regions and corresponding complete color and depth data.

[0111] The above-described technical solution in this embodiment obtains each layered masking map by performing a logical AND operation between the parallax occlusion map and each depth layered map. Then, it fills each layered masking map with color data and depth data according to the two-dimensional image and the depth map. This ensures that when the viewpoint is switched during subsequent 3D rendering, the exposed area can directly call the filled color data and depth data, and there will be no black screen or blur.

[0112] As a fourth optional embodiment of this example, based on the above embodiments, the step of performing 3D rendering based on the 2D image, depth map, foreground mask map, and various color depth fill maps in a first-view perspective to generate media content displayed in a first-view perspective can be optimized as follows:

[0113] a6) Divide the 2D image into color layers of the same number as the depth layers, and construct the same number of rendering channels.

[0114] Understandably, to accelerate rendering, a separate rendering pass can be built for each depth layer map, allowing each pass to handle information specific to its corresponding layer. However, to complete the 3D rendering of the corresponding layer by combining the corresponding depth layer map within each rendering pass, it is necessary to convert planar colors into layered colors.

[0115] Specifically, based on the number of depth layers, the two-dimensional image is divided into color layers corresponding to the depth layers, so that each depth layer has a corresponding color layer.

[0116] b6) Input the color layer map, depth layer map, and color depth fill map as input information and input them to the corresponding rendering channels according to the hierarchical relationship, and input the foreground mask map as input information to the associated rendering channel.

[0117] In this embodiment, the hierarchical relationship can be considered as the hierarchical relationship based on which the depth layer map is divided, and can be determined according to the lower and upper bounds of the depth. Each rendering channel will receive the color layer map, depth layer map, and color depth fill map of the corresponding layer. At the same time, the foreground mask map can be input as input information to the associated rendering channel, which is usually the first rendering channel, that is, the rendering channel corresponding to the first layer or the rendering channel corresponding to the first depth layer map. In an optional embodiment, it can also be the rendering channel corresponding to the layer where the first object is located.

[0118] c6) Determine the first view vector representing the first view based on the current terminal pose information.

[0119] In this embodiment, the terminal pose information can be understood as the pose information of the media content playback device that performs 3D rendering. The terminal pose information can come from the sensors of the terminal device, such as the mobile phone gyroscope or VR headset sensor, and includes position information (such as forward and backward movement, left and right movement, or up and down movement) and posture information (such as angle information such as rotation, head turning, head raising, or head lowering).

[0120] Following the above description, the terminal pose information is converted into a first-view vector representing the first-person perspective, indicating the viewer's starting point and the direction of their gaze. By determining the first-view vector representing the first-person perspective, the direction of image rendering is clarified.

[0121] d6) Through the rendering channel, perform 3D rendering on the input information based on the first-view vector to obtain the rendered color depth layer.

[0122] In this embodiment, each rendering channel determines the visible area of ​​the layer from the first viewpoint, as well as the color data of the visible area, based on the determined first viewpoint vector and the input information, thereby performing 3D rendering to obtain the rendered color depth layer. The input information can be considered as 3D information, that is, information used to render a 2D image into 3D or higher-degree-of-freedom media content, which may include a color layer map, a depth layer map, and a color depth fill map.

[0123] e6) Integrate the color depth layers according to their hierarchical relationship to generate media content presented from a first-person perspective.

[0124] In this embodiment, according to the hierarchical relationship, the color depth layers output by each rendering channel can be superimposed in the order of foreground in front and background behind to generate media content displayed from a first-person perspective. This presents the effect of scene synchronous change with the change of the first-person perspective or the adjustment of the pose of the media content playback device, ensuring that the generated media content can imitate the display characteristics of three-dimensional / six-degree-of-freedom media content.

[0125] Based on the fourth optional embodiment described above, as one implementation method, the step of performing three-dimensional rendering on the input information according to the first view vector through the rendering channel to obtain the rendered color depth layer can be further optimized to the following steps:

[0126] a7) Determine the texture positioning coordinates that match the first viewpoint based on the first viewpoint vector and the depth layer map in the input information.

[0127] Understandably, to ensure that the rendering result of each rendering pass accurately matches the viewing angle, it is necessary to first determine the texture positioning coordinates that match the first-viewpoint based on the first-viewpoint vector representing the direction of the first-viewpoint's line of sight. Texture positioning coordinates can be considered as pixel position indices in the color layer map, used to establish the pixel correspondence between the depth layer map and the color layer map. For example, texture positioning coordinates can be represented by floating-point numbers in the range [0,1]. The U-axis corresponds to the left-right direction of the texture, with 0 being the leftmost point and 1 the rightmost point; the V-axis corresponds to the up-down direction of the texture, with 0 being the bottommost point and 1 the topmost point. The texture positioning coordinates formed by the U-axis and V-axis can accurately locate any pixel in the color layer map.

[0128] Optionally, determining the texture positioning coordinates matching the first viewpoint based on the first viewpoint vector and the depth layer map in the input information includes:

[0129] a71) Determine the pixel corresponding to the first-view vector in the color layer map of the input information.

[0130] In this embodiment, the pixel to be rendered can be determined based on the first view vector, and the pixel corresponding to the pixel to be rendered and its initial texture coordinates can be determined in the color layer map.

[0131] a72) For each pixel, determine the depth value of the pixel on the depth layer map, and determine the depth sampling interval with the same direction as the first view vector starting from the depth value.

[0132] In this embodiment, the corresponding depth value can be determined based on the projection of each pixel on the depth layer map, and its depth value is set as the near clipping plane. The preset far clipping plane is used as the sampling boundary, and the depth sampling interval along the first view vector is constructed with the depth value as the starting point, thereby limiting the sampling range, avoiding invalid search in areas outside the line, and improving computational efficiency.

[0133] a73) Determine at least one intermediate texture coordinate contained in the depth sampling interval according to the set sampling step size, and construct at least one texture coordinate interval based on adjacent intermediate texture coordinates.

[0134] In this embodiment, equal-step sampling is performed within the depth sampling interval, generating a series of candidate sampling points based on a set sampling step size. Each candidate sampling point corresponds to an intermediate texture coordinate. A texture coordinate interval can be determined based on the two intermediate texture coordinates corresponding to every two adjacent candidate sampling points. For example, if there are 5 intermediate texture coordinates within the depth sampling interval, 4 consecutive texture coordinate intervals will be formed, covering the entire sampling range.

[0135] Alternatively, the depth value of each candidate sampling point can be determined based on its projection onto the depth layer map.

[0136] a74) Determine the target texture coordinate range, which is the texture coordinate range where the depth layer map and the first view vector intersect; determine the texture coordinates of the intersection points, and use the texture coordinates as the texture positioning coordinates of the pixel points in the first view.

[0137] In this embodiment, each intermediate texture coordinate can be back-projected to world space or tangent space to obtain its corresponding three-dimensional spatial position. The geometric relationship between this three-dimensional spatial position and the current pixel's position in the depth layer map is calculated to construct a determinant value, which is used to determine whether the first-view vector traverses the depth layer map. Next, the sign change of the determinant value is monitored. For example, when the determinant value changes from negative to positive, it is determined that the first-view vector has traversed the depth layer map, and the two adjacent candidate sampling points before and after the sign change are recorded. The texture coordinate interval corresponding to the two candidate sampling points is determined as the target texture coordinate interval. Within the target texture coordinate interval, there is an intersection between the first-view vector and the depth layer map.

[0138] Following the above description, the texture coordinates of the intersection point can be further determined from the target texture coordinate range. The method for determining the texture coordinates of the intersection point is as follows: perform a binary search within the target texture coordinate range, iteratively reduce the sampling step size, and gradually approach the intersection point of the first-view vector and the depth layer map until a preset convergence condition is met. After meeting the preset convergence condition, perform linear interpolation based on the two nearest neighbor candidate sampling points obtained from the binary search, output the texture coordinates of the intersection point, and use these texture coordinates as the texture positioning coordinates of the pixel point under the first view.

[0139] The above-described technical solution in this embodiment determines a depth sampling interval with the same direction as the first view vector, and divides the depth sampling interval into multiple texture coordinate intervals according to the sampling step size. By determining the target texture coordinate interval where the intersection point is located from multiple texture coordinate intervals, point-by-point traversal is avoided, thus improving search efficiency. Then, based on the bisection method and linear interpolation operation, the accurate texture coordinates of the intersection point are determined, ensuring that the final texture coordinates are accurate enough and match the first view, thus avoiding blurry or misaligned rendering results.

[0140] b7) Determine the edge marker information of texture positioning coordinates according to the set edge control strategy.

[0141] It is understandable that texture positioning coordinates may fall on edge areas that are prone to artifacts such as stretching, clipping, and / or color bleeding during 3D rendering. In order to ensure that the edge transitions of the media content rendered in the first-person perspective are natural and distortion-free, it is necessary to first determine and mark the texture positioning coordinates that may cause the above problems.

[0142] Optionally, determining the edge marker information for texture positioning coordinates according to the set edge control strategy includes:

[0143] b71) Determine the texture coordinate range in which the texture positioning coordinates are located, and determine the coordinate offset of the texture positioning coordinates.

[0144] In this embodiment, the coordinate range where the texture positioning coordinates are located during the last binary search can be determined as the texture coordinate range where the texture positioning coordinates are located, and the difference between the texture positioning coordinates of the pixel and its initial texture coordinates can be determined as the coordinate offset.

[0145] b72) Determine the depth gradient slope of the texture positioning coordinates based on the texture coordinate range and coordinate offset.

[0146] In this embodiment, two texture coordinates constituting the texture coordinate interval are determined according to the texture coordinate interval. The height difference can be determined according to the difference between the two texture coordinates, and the quotient of the height difference and the coordinate offset is determined as the depth gradient slope of the texture positioning coordinate. The depth gradient slope is used to characterize the steepness of the surface undulation in the depth layer map.

[0147] (b73) If the depth gradient slope exceeds the set slope threshold, and / or if the texture positioning coordinates fall within the set boundary coordinate range, the set marker value is used as the edge marker information of the texture positioning coordinates.

[0148] In this embodiment, if the depth gradient slope exceeds a set slope threshold, the region can be identified as a steep edge region, posing a risk of parallax stretching artifacts. Conversely, if the texture positioning coordinates fall within a set boundary coordinate range, these coordinates may experience color bleeding or seam issues due to texture wrapping or background color sampling, thus affecting the spatial continuity of the rendering result. By using a set marker value as edge marker information for the texture positioning coordinates, texture positioning coordinates that are likely to affect the rendering result are identified, aiding in subsequent processing and reducing their impact on the rendering outcome. Optionally, different marker values ​​can be set as edge marker information for the cases where the depth gradient slope exceeds the set slope threshold and the cases where the texture positioning coordinates fall within the set boundary coordinate range.

[0149] c7) Based on the edge marker information and the color layer map in the input information, determine the pixel information of the pixel point under the texture positioning coordinate.

[0150] In this embodiment, determining the pixel information of a pixel at the texture positioning coordinates based on the edge marker information and the color layer map in the input information includes: if the edge marker information is a set marker value, determining the set transparency as the pixel information of the pixel at the texture positioning coordinates; otherwise, determining the texture color value corresponding to the texture positioning coordinates in the color layer map, and determining the texture color value as the pixel information of the pixel at the texture positioning coordinates.

[0151] In this embodiment, when the edge marker information corresponds to a set marker value where the depth gradient slope exceeds a set slope threshold, a smooth transition function can be introduced to non-linearly attenuate the original disparity scaling coefficient, simultaneously adjust the disparity offset intensity, and adjust the existing blend transparency through a set transparency. This gradually weakens the disparity effect in the edge region, thereby suppressing texture stretching or sampling distortion caused by abrupt depth changes. Conversely, when the edge marker information corresponds to a set marker value where the texture positioning coordinates fall within the set boundary coordinate range, the sampling result can be directly discarded. This avoids color penetration or seam problems caused by texture wrapping or background color sampling, ensuring the spatial continuity of the rendering result, and the set transparency is determined as the pixel information of the pixel at the texture positioning coordinates.

[0152] In this embodiment, when the edge marker information is not a set marker value or there is no edge marker information, the texture color value corresponding to the texture positioning coordinate is determined from the color layer map, and the texture color value is determined as the pixel information of the pixel point under the texture positioning coordinate.

[0153] d7) Generate a color depth layer based on the information of each pixel.

[0154] In this embodiment, the pixels are rendered according to the information of each pixel to generate a color depth layer, which ensures that the edges of the rendering results in the first view are clear and natural, without deformation, breakage and color abnormality, thus improving the rendering effect and enabling better simulation of the display effect of three-dimensional or higher degree of freedom media content.

[0155] Based on any of the above embodiments, as one implementation, after determining the color depth fill map of each depth layer map, the optimization may further include:

[0156] a8) Perform texture compression on each color depth fill map to generate a texture compression map; stitch the texture compression map, two-dimensional image, depth map and foreground mask map according to the first stitching format to generate a three-dimensional information map.

[0157] It should be noted that when two different execution devices are used as the execution entities to implement the logic of this media content processing method, the first execution device, typically a cloud server with powerful storage and computing capabilities, is usually chosen to complete the initial steps such as determining the color depth fill maps for each depth layer. After determining the color depth fill maps for each depth layer, the 2D image, depth map, foreground mask map, each depth layer map, and each color depth fill map need to be sent to the second execution terminal for media content rendering, generation, and playback. The second execution terminal is usually a widely available mobile terminal device, such as a mobile phone or tablet, with limited computing and storage capabilities.

[0158] Considering the performance limitations and communication transmission pressure of the second execution terminal, in this embodiment, after determining the color depth fill map of each depth layer map, the first execution device also needs to perform texture compression on each color depth fill map to generate a texture compression map, and then stitch the texture compression map, the two-dimensional image, the depth map and the foreground mask map according to the preset first stitching format to generate a three-dimensional information map. This three-dimensional information map is suitable for two-dimensional distribution link transmission and can reduce transmission bandwidth and the rendering and display pressure on the second execution terminal side.

[0159] In this embodiment, the method for generating a texture compressed map by compressing the color depth fill maps can be as follows: First, the data in each color depth fill map and the data generated in the previous step can be obtained, and preliminary stitching processing can be performed. For example, the number of layers is... A multi-layered color depth fill map, composed of various color depth fill maps, can include .

[0160] Two-dimensional images can be obtained. And obtain the corresponding color layer map based on the multi-layer color depth fill map. Multi-layered color images are created by splicing them together. Next, obtain the disparity map converted from the two-dimensional image. And obtain the layer map of each depth based on the multi-layer color depth fill map. Multi-layer depth image is formed by stitching together Furthermore, obtain the original layer mask image corresponding to the two-dimensional image. and each layer of masking image The images are stitched together to form a multi-layered filled mask image. A value of 1 indicates that the original 2D image has no areas that need to be filled.

[0161] It should be noted that all data corresponds one-to-one with the hierarchical sequence number (e.g., ... correspond and All of these are information from the same depth layer.

[0162] Next, based on the multi-layered filled mask images in hierarchical order, the transparency weights of each layer are generated through adjacent mask difference processing, clarifying the contribution of each layer to the final image, such as higher weights for closer layers and lower weights for farther layers. For example, the transparency weight of the first layer can be directly... The transparency weights for the second and higher layers can be determined by subtracting the previous layer's masking image from the current layer's masking image, and then using the clip(0,1) function to limit the range, ensuring that the transparency weights are between 0 and 1. This method retains only the region weights unique to the current layer, avoiding duplication with previous layers.

[0163] The transparency weights of the original layers are determined by subtracting the maximum transparency weight from all layers from 1, and then using the clip(0,1) function to limit the range, thus preserving the weights of areas not covered by any layers (i.e., areas in the original image that are unoccluded and do not require filling). The total transparency weight is determined by summing the transparency weights of all layers.

[0164] In this embodiment, after data collection is completed and transparency weights are determined for each layer, texture compression is performed on the multi-layer color image based on the transparency weights of each layer, the total transparency weight, and the multi-layer color image to obtain a first compression result. Specifically, each layer color image is multiplied by the transparency weight of its corresponding layer, and then the product results are added together. The quotient of the sum and the total transparency weight is determined as the first compression result.

[0165] For example, determine the first compression result. The specific method can be expressed as:

[0166] ;

[0167] in, This represents the total transparency weight.

[0168] Next, texture compression can be performed on the multi-layer depth image based on the transparency weights of each layer, the total transparency weight, and the multi-layer depth image to obtain a second compression result. The specific method can refer to the method used to determine the first compression result. For example, determining the second compression result... The specific method can be expressed as:

[0169] ;

[0170] Next, the first compression result and the second compression result are determined as a texture compression map, and a three-dimensional information map is stitched together according to the texture compression map, the two-dimensional image and the depth map in accordance with the first stitching format.

[0171] In an alternative embodiment, a depth-visible confidence map can be incorporated when stitching according to the first stitching format to further improve the accuracy of the 3D information map.

[0172] c8) Encode the 3D information map to generate a 3D information bitstream for transmission of the 3D information map.

[0173] It should be noted that 3D infographics can be transmitted through 2D transmission links, enabling them to be transmitted to any general-purpose device. This lowers the equipment threshold and transmission requirements, and better expands the audience and application scenarios of 3D / six degrees of freedom media content.

[0174] Based on the above implementation, optionally, before performing 3D rendering based on the 2D image, depth map, foreground mask map, and each color depth fill map in the first viewpoint, the following method is also included:

[0175] a9) Receive and decode the three-dimensional information stream to obtain a three-dimensional information map.

[0176] b9) Split the 3D information map to obtain the split 2D image, depth map, foreground mask map and texture compression map, and obtain the color fill map from the texture compression map.

[0177] Understandably, since the first execution device that determines the color depth fill map of each depth layer map and the second execution device that performs 3D rendering and media content display are different devices, after the first execution device encodes the generated 3D information map into a 3D information bitstream and transmits it to the second execution device, the second execution device needs to receive the 3D information bitstream and decode it to obtain the 3D information map and extract the 2D image, depth map, foreground mask map and each color fill map that can be used for 3D rendering, so as to provide support for 3D rendering and ensure the rendering effect.

[0178] Figure 2This is a schematic diagram of a media content processing apparatus provided in an embodiment of this disclosure. This embodiment is applicable to media content processing, particularly to determining the three-dimensional or 6DoF media content format of two-dimensional media content. The apparatus can be implemented through software and / or hardware and can be configured in a terminal and / or server to implement the media content processing method in this embodiment. Figure 2 As shown, the device may specifically include: a first determining module 21, a second determining module 22, and a content rendering module 23.

[0179] The first determining module 21 is used to determine the depth map of the two-dimensional image and determine the foreground mask map of the two-dimensional image based on the depth map. The two-dimensional image is a media frame obtained from a media stream.

[0180] The second determining module 22 is used to perform layering processing on the depth map to obtain at least one depth layer map, and to determine the color depth fill map of each depth layer map.

[0181] The content rendering module 23 is used to perform three-dimensional rendering based on the two-dimensional image, the depth map, the foreground mask map, and each of the color depth fill maps in a first viewpoint, and generate media content displayed in the first viewpoint.

[0182] This disclosure provides a media content processing apparatus that determines a depth map of a two-dimensional image, and then determines a foreground mask map of the two-dimensional image based on the depth map. The two-dimensional image is a media frame acquired from a media stream. The depth map is layered to obtain at least one depth layer map, and a color depth fill map is determined for each depth layer map. Based on the two-dimensional image, depth map, foreground mask map, and each color depth fill map, three-dimensional rendering is performed from a first-viewpoint to generate media content displayed from that first-viewpoint. Using this apparatus, two-dimensional images acquired by a monocular image acquisition device can be reused. By determining the corresponding three-dimensional information of the two-dimensional image and combining it with the first-viewpoint and three-dimensional rendering technology, media content corresponding to the two-dimensional image and displayed from the first-viewpoint can be generated. This ensures that the media content can mimic the display characteristics of three-dimensional / six-degree-of-freedom media content, expanding gameplay, enriching the three-dimensional / six-degree-of-freedom media content, and improving the viewing experience. Compared to other technologies for playing 3D / 6DOF media content, this solution eliminates the need for specialized acquisition equipment. By simply determining the 3D information of a traditional 2D image and performing 3D rendering based on that information and a first-person perspective before playback, the solution ensures that the acquired media content visually changes its playback perspective according to the device's pose. This avoids the cost and storage issues associated with other technologies for playing 3D / 6DOF media content, while also reducing the rendering and display burden and thus expanding the audience reach of 3D / 6DOF media content.

[0183] Furthermore, the first determining module 21 may specifically include:

[0184] The disparity map generation unit is used to normalize the depth map within a set floating-point range to generate a disparity map of the depth map.

[0185] The first foreground mask determination unit is used to determine the depth visibility confidence map of the two-dimensional image based on the disparity map without performing image matting on the two-dimensional image, and to determine the depth visibility confidence map as the foreground mask map.

[0186] The image matting unit is used to obtain a first matting result for a first object in the two-dimensional image when performing image matting processing on the two-dimensional image;

[0187] The second foreground mask determination unit is used to determine the disparity occlusion map of the two-dimensional image based on the disparity map, and to determine the foreground mask map based on the first matting result, the disparity occlusion map, and the depth visibility confidence map.

[0188] Furthermore, the first foreground mask determination unit can specifically be used for:

[0189] Gradient calculation is performed on the pixels in the disparity map to obtain a first gradient magnitude map of the disparity map, and a second gradient magnitude map is obtained after smoothing the first gradient magnitude map.

[0190] Determine the depth visibility confidence level corresponding to each gradient magnitude in the second gradient magnitude map, and generate a depth visibility confidence map of the two-dimensional image based on each depth visibility confidence level.

[0191] Furthermore, the second foreground mask determination unit can specifically be used for:

[0192] In at least one set direction, the pixels in the disparity map are subjected to neighborhood difference processing with at least one set step size to obtain the joint difference value of the pixels under each combination of the set direction and the set step size.

[0193] The maximum value among the joint difference values ​​is determined as the maximum difference value of the pixel, and the disparity occlusion intensity of the pixel is determined based on the maximum difference value.

[0194] A disparity occlusion map of the two-dimensional image is generated based on the disparity occlusion intensity of each pixel.

[0195] Furthermore, the second foreground mask determination unit can also be used for:

[0196] The first cutout result is normalized within a set floating-point range to generate a second cutout result of the first cutout result;

[0197] Based on the set structural elements, the pixels in the second cutout result are dilated to obtain the dilated image of the second cutout result;

[0198] Determine the difference map between the expansion map and the second matting result, and determine the edge narrow band region of the first object from the difference map;

[0199] Based on the parallax occlusion map, the suppression weights of pixels in the edge narrow band region are determined, and the depth visibility confidence map is updated based on the suppression weights. The updated depth visibility confidence map is then used as the foreground mask map.

[0200] Furthermore, the second determining module 22 can specifically be used for:

[0201] A deep clustering strategy is used to determine at least one first depth layer map of the depth map, and the depth interval and rectangular bounding box corresponding to each first depth layer map are obtained.

[0202] For each first depth layer map except the first first depth layer map, the occlusion area is determined from the depth map according to the lower limit value of the corresponding depth interval and the rectangular bounding box. After filtering out the occlusion area from the first depth layer map, a second depth layer map is obtained.

[0203] The first first depth layer map is determined as a depth layer map, and each of the second depth layer maps is determined as a depth layer map.

[0204] Furthermore, the second determining module 22 can also be used specifically for:

[0205] Obtain a disparity occlusion map determined relative to the depth map;

[0206] The parallax occlusion map is logically ANDed with each of the depth layer maps to obtain the processed layered occlusion map.

[0207] Based on the two-dimensional image and the depth map, color data and depth data are filled into each of the layered mask images to obtain the color depth filled image of each depth layered image.

[0208] Furthermore, the content rendering module 23 may specifically include:

[0209] The image segmentation and channel construction unit is used to segment the two-dimensional image into color layer maps with the same number as the depth layer maps based on the depth map, and construct rendering channels with the same number as the depth layer maps.

[0210] The input unit is used to input the color layer map, the depth layer map, and the color depth fill map as input information and input them to the corresponding rendering channels according to the hierarchical relationship, and to input the foreground mask map as input information to the associated rendering channel.

[0211] The view vector determination unit is used to determine a first view vector representing the first view based on the current terminal pose information.

[0212] The rendering unit is used to perform three-dimensional rendering on the input information based on the first view vector through the rendering channel to obtain the rendered color depth layer.

[0213] The media content generation unit is used to integrate the various color depth layers according to their hierarchical relationship to generate media content displayed from the first perspective.

[0214] Furthermore, the rendering unit may specifically include:

[0215] The coordinate determination subunit is used to determine the texture positioning coordinates that match the first viewpoint based on the first viewpoint vector and the depth layer map in the input information.

[0216] The marking information determination subunit is used to determine the edge marking information of the texture positioning coordinates according to the set edge control strategy;

[0217] The pixel information determination subunit is used to determine the pixel information of the pixel point under the texture positioning coordinates based on the edge marking information and the color layer map in the input information.

[0218] The color depth layer generation subunit is used to generate the color depth layer based on the pixel information.

[0219] Furthermore, the coordinate determination subunit can specifically be used for:

[0220] Determine the pixel point corresponding to the first-view vector in the color layer map of the input information;

[0221] For each pixel, determine the depth value of the pixel on the depth layer map, and determine a depth sampling interval with the same direction as the first view vector starting from the depth value;

[0222] The depth sampling interval is determined according to the set sampling step size, and at least one intermediate texture coordinate is formed based on adjacent intermediate texture coordinates.

[0223] Determine the target texture coordinate range, which is the texture coordinate range where the depth layer map intersects with the first view vector;

[0224] Determine the texture coordinates of the intersection point, and use the texture coordinates as the texture positioning coordinates of the pixel point matched in the first viewpoint.

[0225] Furthermore, the marking information determining the subunit can specifically be used for:

[0226] Determine the texture coordinate range in which the texture positioning coordinates are located, and determine the coordinate offset of the texture positioning coordinates;

[0227] The depth gradient slope of the texture positioning coordinates is determined based on the texture coordinate range and the coordinate offset.

[0228] If the depth gradient slope exceeds a set slope threshold, and / or if the texture positioning coordinates fall within a set boundary coordinate range, a set marker value is used as the edge marker information for the texture positioning coordinates.

[0229] Furthermore, the pixel information determining subunit can specifically be used for:

[0230] If the edge marker information is a set marker value, the set transparency is determined as the pixel information of the pixel point under the texture positioning coordinates; otherwise,

[0231] Determine the texture color value corresponding to the texture positioning coordinates in the color layer map, and determine the texture color value as the pixel information of the pixel point under the texture positioning coordinates.

[0232] Furthermore, the device also includes a compression module, which can specifically be used for:

[0233] After determining the color depth fill map of each of the depth layer maps, texture compression is performed on each of the color depth fill maps to generate a texture compression map;

[0234] The texture compression map, the two-dimensional image, the depth map, and the foreground mask map are stitched together according to the first stitching format to generate a three-dimensional information map.

[0235] The three-dimensional information map is encoded to generate a three-dimensional information code stream for transmission of the three-dimensional information map.

[0236] Furthermore, the device also includes a decoding and splitting module, which can be specifically used for:

[0237] Before performing 3D rendering based on the 2D image, the depth map, the foreground mask map, and each of the color depth fill maps in the first viewpoint, the 3D information stream is received and decoded to obtain the 3D information map.

[0238] The three-dimensional information map is split to obtain the split two-dimensional image, depth map, foreground mask map and texture compression map, and each color fill map is obtained from the texture compression map.

[0239] The above-described apparatus can execute the media content processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.

[0240] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.

[0241] Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of this disclosure. Reference is made below. Figure 3It illustrates a computer device suitable for implementing embodiments of the present disclosure (e.g., Figure 3 The diagram below shows the structure of the terminal device or server 30. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 3 The computer device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0242] like Figure 3 As shown, the computer device 30 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 31, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 32 or a program loaded from a storage device 38 into a random access memory (RAM) 33. The RAM 33 also stores various programs and data required for the operation of the computer device 30. The processing unit 31, ROM 32, and RAM 33 are interconnected via a bus 35. An edit / output (I / O) interface 34 is also connected to the bus 35.

[0243] Typically, the following devices can be connected to I / O interface 34: input devices 36 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 37 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 38 including, for example, magnetic tapes, hard disks, etc.; and communication devices 39. Communication device 39 allows computer device 30 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 A computer device 30 with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have instead.

[0244] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 39, or installed from a storage device 38, or installed from a ROM 32. When the computer program is executed by the processing device 31, it performs the functions defined in the methods of embodiments of this disclosure.

[0245] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0246] The computer device provided in this embodiment and the media content processing method provided in the above embodiments belong to the same concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0247] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the media content processing method provided in the above embodiments.

[0248] It should be noted that the computer-readable medium described above in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0249] In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0250] According to one or more embodiments of this disclosure, [Example 1] provides a media content processing method, including: determining a depth map of a two-dimensional image; determining a foreground mask map of the two-dimensional image based on the depth map, wherein the two-dimensional image is a media frame obtained from a media stream; performing layered processing on the depth map to obtain at least one depth layer map, and determining a color depth fill map for each of the depth layer maps; performing three-dimensional rendering based on the two-dimensional image, the depth map, the foreground mask map, and each of the color depth fill maps in a first viewpoint to generate media content displayed in the first viewpoint.

[0251] According to one or more embodiments of this disclosure, [Example 2] provides the method of Example 1. Optionally, determining the foreground mask map of the two-dimensional image based on the depth map includes: normalizing the depth map in a set floating-point range to generate a disparity map of the depth map; determining a depth visibility confidence map of the two-dimensional image based on the disparity map when no image matting is performed on the two-dimensional image, and determining the depth visibility confidence map as the foreground mask map; obtaining a first matting result for a first object in the two-dimensional image when image matting is performed on the two-dimensional image; determining a disparity occlusion map of the two-dimensional image based on the disparity map; and determining the foreground mask map based on the first matting result, the disparity occlusion map, and the depth visibility confidence map.

[0252] According to one or more embodiments of this disclosure, [Example 3] provides the method of Example 2. Optionally, determining the depth visibility confidence map of the two-dimensional image based on the disparity map includes: performing gradient calculation on the pixels in the disparity map to obtain a first gradient magnitude map of the disparity map, and obtaining a second gradient magnitude map after smoothing the first gradient magnitude map; determining the depth visibility confidence level corresponding to each gradient magnitude in the second gradient magnitude map, and generating the depth visibility confidence map of the two-dimensional image based on each depth visibility confidence level.

[0253] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 2. Optionally, determining the disparity occlusion map of the two-dimensional image based on the disparity map includes: performing neighborhood difference processing on pixels in the disparity map at least one set step size in at least one set direction to obtain joint difference values ​​corresponding to the pixels under each combination of the set direction and the set step size; determining the maximum value among the joint difference values ​​as the maximum difference value of the pixel, and determining the disparity occlusion intensity of the pixel based on the maximum difference value; and generating the disparity occlusion map of the two-dimensional image based on the disparity occlusion intensity of each pixel.

[0254] According to one or more embodiments of this disclosure, Example 5 provides a method of Example 2. Optionally, determining the foreground mask map based on the first matting result, the disparity occlusion map, and the depth visibility confidence map includes: normalizing the first matting result in a set floating-point range to generate a second matting result of the first matting result; dilating the pixels in the second matting result based on a set structuring element to obtain an expanded map of the second matting result; determining a difference map between the expanded map and the second matting result, and determining the edge narrowband region of the first object from the difference map; determining the suppression weight of the pixels in the edge narrowband region based on the disparity occlusion map, updating the depth visibility confidence map based on the suppression weight, and determining the updated depth visibility confidence map as the foreground mask map.

[0255] According to one or more embodiments of this disclosure, Example Six provides a method of Example One. Optionally, the step of performing layering processing on the depth map to obtain at least one depth layer map includes: using a depth clustering strategy to determine at least one first depth layer map of the depth map, obtaining a depth interval and a rectangular bounding box corresponding to each first depth layer map; for each first depth layer map other than the first first depth layer map, determining an occlusion region from the depth map based on the lower limit value of the corresponding depth interval and the rectangular bounding box, and filtering out the occlusion region from the first depth layer map to obtain a second depth layer map; determining the first first depth layer map as a depth layer map, and determining each second depth layer map as a depth layer map.

[0256] According to one or more embodiments of this disclosure, Example 7 provides the method of Example 1. Optionally, determining the color depth fill map of each of the depth layer maps includes: obtaining a disparity occlusion map determined relative to the depth map; performing a logical AND operation between the disparity occlusion map and each of the depth layer maps to obtain a processed layered mask map; and filling each of the layered mask maps with color data and depth data according to the two-dimensional image and the depth map to obtain the color depth fill map of each of the depth layer maps.

[0257] According to one or more embodiments of this disclosure, Example 8 provides a method of Example 1. Optionally, the step of performing 3D rendering based on the 2D image, the depth map, the foreground mask map, and each of the color depth fill maps in a first viewpoint to generate media content displayed in the first viewpoint includes: dividing the 2D image into color layer maps of the same number as the depth layer maps based on the depth map, and constructing rendering channels of the same number; inputting the color layer maps, the depth layer maps, and the color depth fill maps as input information and inputting them into the corresponding rendering channels according to the hierarchical relationship, and inputting the foreground mask map as input information into the associated rendering channel; determining a first viewpoint vector representing the first viewpoint based on the current terminal pose information; performing 3D rendering on the input information based on the first viewpoint vector through the rendering channels to obtain the rendered color depth layer; and integrating each of the color depth layers according to the hierarchical relationship to generate media content displayed in the first viewpoint.

[0258] According to one or more embodiments of this disclosure, [Example Nine] provides the method of Example Eight, wherein, optionally, the step of performing three-dimensional rendering on the input information based on the first viewpoint vector through the rendering channel to obtain a rendered color depth layer includes: determining texture positioning coordinates matching the first viewpoint based on the first viewpoint vector and the depth layer map in the input information; determining edge marker information of the texture positioning coordinates according to a set edge control strategy; determining pixel information of the pixels at the texture positioning coordinates based on the edge marker information combined with the color layer map in the input information; and generating the color depth layer based on each pixel information.

[0259] According to one or more embodiments of this disclosure, Example 10 provides the method of Example 9. Optionally, determining the texture positioning coordinates matching the first viewpoint based on the first viewpoint vector and the depth layer map in the input information includes: determining the pixel point corresponding to the first viewpoint vector in the color layer map of the input information; for each pixel point, determining the depth value of the pixel point on the depth layer map, and determining a depth sampling interval with the same direction as the first viewpoint vector starting from the depth value; determining at least one intermediate texture coordinate contained in the depth sampling interval according to a set sampling step size, and constructing at least one texture coordinate interval based on adjacent intermediate texture coordinates; determining a target texture coordinate interval, the target texture coordinate interval being a texture coordinate interval where the depth layer map and the first viewpoint vector intersect; determining the texture coordinates of the intersection point, and using the texture coordinates as the texture positioning coordinates matching the pixel point under the first viewpoint.

[0260] According to one or more embodiments of this disclosure, Example 11 provides the method of Example 9. Optionally, determining the edge marker information of the texture positioning coordinates according to a set edge control strategy includes: determining the texture coordinate interval in which the texture positioning coordinates are located, and determining the coordinate offset of the texture positioning coordinates; determining the depth gradient slope of the texture positioning coordinates according to the texture coordinate interval and the coordinate offset; and using a set marker value as the edge marker information of the texture positioning coordinates when the depth gradient slope exceeds a set slope threshold, and / or when the texture positioning coordinates fall within a set boundary coordinate range.

[0261] According to one or more embodiments of this disclosure, Example Twelve provides the method of Example Nine, wherein, optionally, determining the pixel information of a pixel at the texture positioning coordinates based on the edge marker information and the color layer map in the input information includes: if the edge marker information is a set marker value, determining the set transparency as the pixel information of the pixel at the texture positioning coordinates; otherwise, determining the texture color value corresponding to the texture positioning coordinates in the color layer map, and determining the texture color value as the pixel information of the pixel at the texture positioning coordinates.

[0262] According to one or more embodiments of this disclosure, Example Thirteen provides a method of any one of Examples One to Twelve. Optionally, after determining the color depth fill map of each of the depth layer maps, the method further includes: performing texture compression on each of the color depth fill maps to generate a texture compression map; stitching the texture compression map, the two-dimensional image, the depth map, and the foreground mask map according to a first stitching format to generate a three-dimensional information map; and encoding the three-dimensional information map to generate a three-dimensional information bitstream for transmission of the three-dimensional information map.

[0263] According to one or more embodiments of this disclosure, Example Fourteen provides the method of Example Thirteen, which, optionally, before performing three-dimensional rendering based on the two-dimensional image, the depth map, the foreground mask map, and each of the color depth fill maps in a first viewpoint, further includes: receiving and decoding the three-dimensional information stream to obtain a three-dimensional information map; splitting the three-dimensional information map to obtain the split two-dimensional image, depth map, foreground mask map, and texture compression map, and obtaining each color fill map from the texture compression map.

[0264] In some implementations, the data requesting end or server can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0265] The aforementioned computer-readable medium may be included in the aforementioned computer device; or it may exist independently and not assembled into the computer device.

[0266] The aforementioned computer-readable medium carries one or more programs that, when executed by the computer device, cause the computer device to:

[0267] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0268] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0269] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".

[0270] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0271] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0272] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to the specific combination of the above-described technical features, but should also cover other technical methods formed by any combination of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical methods formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0273] Furthermore, although the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while some specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0274] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A media content processing method, comprising: Determine the depth map of a two-dimensional image, and determine the foreground mask map of the two-dimensional image based on the depth map, wherein the two-dimensional image is a media frame obtained from a media stream; The depth map is layered to obtain at least one depth layer map, and the color depth fill map of each depth layer map is determined. Based on the two-dimensional image, the depth map, the foreground mask map, and each of the color depth fill maps, a three-dimensional rendering is performed from a first perspective to generate media content displayed from the first perspective.

2. The method according to claim 1, wherein determining the foreground mask map of the two-dimensional image based on the depth map comprises: The depth map is normalized within a set floating-point range to generate a disparity map of the depth map; Without performing image matting on the two-dimensional image, the depth visibility confidence map of the two-dimensional image is determined based on the disparity map, and the depth visibility confidence map is determined as the foreground mask map; When performing image matting on the two-dimensional image, a first matting result is obtained for the first object in the two-dimensional image; The disparity occlusion map of the two-dimensional image is determined based on the disparity map, and the foreground mask map is determined based on the first matting result, the disparity occlusion map, and the depth visibility confidence map.

3. The method according to claim 2, wherein determining the depth visibility confidence map of the two-dimensional image based on the disparity map comprises: Gradient calculation is performed on the pixels in the disparity map to obtain a first gradient magnitude map of the disparity map, and a second gradient magnitude map is obtained after smoothing the first gradient magnitude map. Determine the depth visibility confidence level corresponding to each gradient magnitude in the second gradient magnitude map, and generate a depth visibility confidence map of the two-dimensional image based on each depth visibility confidence level.

4. The method according to claim 2, wherein determining the disparity occlusion map of the two-dimensional image based on the disparity map comprises: In at least one set direction, the pixels in the disparity map are subjected to neighborhood difference processing with at least one set step size to obtain the joint difference value of the pixels under each combination of the set direction and the set step size. The maximum value among the joint difference values ​​is determined as the maximum difference value of the pixel, and the disparity occlusion intensity of the pixel is determined based on the maximum difference value. A disparity occlusion map of the two-dimensional image is generated based on the disparity occlusion intensity of each pixel.

5. The method according to claim 2, wherein determining the foreground mask map based on the first matting result, the disparity occlusion map, and the depth visibility confidence map comprises: The first cutout result is normalized within a set floating-point range to generate a second cutout result of the first cutout result; Based on the set structural elements, the pixels in the second cutout result are dilated to obtain the dilated image of the second cutout result; Determine the difference map between the expansion map and the second matting result, and determine the edge narrow band region of the first object from the difference map; Based on the parallax occlusion map, the suppression weights of pixels in the edge narrow band region are determined, and the depth visibility confidence map is updated based on the suppression weights. The updated depth visibility confidence map is then used as the foreground mask map.

6. The method according to claim 1, wherein performing layering processing on the depth map to obtain at least one depth layer map comprises: A deep clustering strategy is used to determine at least one first depth layer map of the depth map, and the depth interval and rectangular bounding box corresponding to each first depth layer map are obtained. For each first depth layer map except the first first depth layer map, the occlusion area is determined from the depth map according to the lower limit value of the corresponding depth interval and the rectangular bounding box. After filtering out the occlusion area from the first depth layer map, a second depth layer map is obtained. The first first depth layer map is determined as a depth layer map, and each of the second depth layer maps is determined as a depth layer map.

7. The method according to claim 1, wherein the step of performing three-dimensional rendering based on the two-dimensional image, the depth map, the foreground mask map, and each of the color depth fill maps in a first viewpoint to generate media content displayed in the first viewpoint comprises: The two-dimensional image is divided into color layers of the same number as the depth layers, and the same number of rendering channels are constructed based on the depth map. The color layer map, the depth layer map, and the color depth fill map are used as input information and input to the corresponding rendering channels according to the hierarchical relationship, and the foreground mask map is used as input information and input to the associated rendering channel. Based on the current terminal pose information, determine the first view vector representing the first view; Through the rendering channel, the input information is rendered in three dimensions based on the first view vector to obtain the rendered color depth layer. The color depth layers are integrated according to their hierarchical relationship to generate media content displayed from the first perspective.

8. The method according to claim 7, wherein the step of performing three-dimensional rendering on the input information based on the first view vector through the rendering channel to obtain a rendered color depth layer includes: Based on the first viewpoint vector and the depth layer map in the input information, determine the texture positioning coordinates that match the first viewpoint; Based on the set edge control strategy, the edge marker information of the texture positioning coordinates is determined; Based on the edge marker information and the color layer map in the input information, determine the pixel information of the pixel point under the texture positioning coordinates; The color depth layer is generated based on the pixel information.

9. The method according to claim 8, wherein determining the texture positioning coordinates matching the first viewpoint based on the first viewpoint vector and the depth layer map in the input information comprises: Determine the pixel point corresponding to the first-view vector in the color layer map of the input information; For each pixel, determine the depth value of the pixel on the depth layer map, and determine a depth sampling interval with the same direction as the first view vector starting from the depth value; The depth sampling interval is determined according to the set sampling step size, and at least one intermediate texture coordinate is formed based on adjacent intermediate texture coordinates. Determine the target texture coordinate range, which is the texture coordinate range where the depth layer map intersects with the first view vector; Determine the texture coordinates of the intersection point, and use the texture coordinates as the texture positioning coordinates of the pixel point matched in the first viewpoint.

10. The method according to claim 8, wherein determining the edge marker information of the texture positioning coordinates according to a set edge control strategy includes: Determine the texture coordinate range in which the texture positioning coordinates are located, and determine the coordinate offset of the texture positioning coordinates; The depth gradient slope of the texture positioning coordinates is determined based on the texture coordinate range and the coordinate offset. If the depth gradient slope exceeds a set slope threshold, and / or if the texture positioning coordinates fall within a set boundary coordinate range, a set marker value is used as the edge marker information for the texture positioning coordinates.

11. The method according to claim 8, wherein determining the pixel information of the pixel point at the texture positioning coordinates based on the edge marker information and the color layer map in the input information comprises: When the edge marker information is a set marker value, the set transparency is determined as the pixel information of the pixel point under the texture positioning coordinates; otherwise, Determine the texture color value corresponding to the texture positioning coordinates in the color layer map, and determine the texture color value as the pixel information of the pixel point under the texture positioning coordinates.

12. The method according to any one of claims 1-11, further comprising, after determining the color depth fill map of each of the depth layer maps: Each of the aforementioned color depth fill maps is texture compressed to generate a texture compressed map; The texture compression map, the two-dimensional image, the depth map, and the foreground mask map are stitched together according to the first stitching format to generate a three-dimensional information map. The three-dimensional information map is encoded to generate a three-dimensional information code stream for transmission of the three-dimensional information map.

13. The method of claim 12, further comprising, before performing 3D rendering based on the 2D image, the depth map, the foreground mask map, and each of the color depth fill maps in the first viewpoint: Receive and decode the three-dimensional information stream to obtain a three-dimensional information map; The three-dimensional information map is split to obtain the split two-dimensional image, depth map, foreground mask map and texture compression map, and each color fill map is obtained from the texture compression map.

14. A media content processing apparatus, comprising: The first determining module is used to determine the depth map of the two-dimensional image and determine the foreground mask map of the two-dimensional image based on the depth map. The two-dimensional image is a media frame obtained from a media stream. The second determining module is used to perform layering processing on the depth map to obtain at least one depth layer map, and to determine the color depth fill map of each depth layer map. The content rendering module is used to perform three-dimensional rendering based on the two-dimensional image, the depth map, the foreground mask map, and each of the color depth fill maps from a first perspective, to generate media content displayed from the first perspective.

15. A computer program product comprising a computer program that, when executed by a processor, implements the media content processing method according to any one of claims 1-13.