Depth map and video processing, reconstruction method, device, equipment and storage medium
By using quantization parameters matching the viewing angle to perform depth map quantization processing and performing appropriate sampling processing, the problem of depth map quantization in the prior art limiting image quality is solved, and the image quality and data efficiency of free viewpoint video are improved.
Patent Information
- Application Number
- CN202010630749.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-03
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-07-03
AI Technical Summary
The depth map quantization processing method in the prior art limits the image quality of the reconstructed free viewpoint video.
Using quantization parameters that match the actual situation of the corresponding viewing angle, the pixel depth values in the estimated depth map are quantized to obtain the quantized depth map, and downsampling or upsampling are performed when necessary to optimize image quality.
By fully utilizing the expression space of deep quantization bits, the image quality of the reconstructed free viewpoint video is improved, and data storage and transmission resources are saved.
Smart Images

Figure CN113963094B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of video processing technology, and in particular to depth map and video processing, reconstruction methods, devices, equipment and storage media. Background Art
[0002] Free viewpoint video is a technology that can provide a high-freedom viewing experience. Users can adjust the viewing angle through interactive operations during the viewing process and watch from the free viewpoint they want, which can greatly improve the viewing experience.
[0003] In large-scale scenes, such as sports games, using depth-based image rendering (DIBR) technology to achieve high-degree-of-freedom viewing is a solution with great potential and feasibility. The expression of free viewpoint video is generally to stitch the texture map collected by multiple cameras with the corresponding depth map to form a stitched image.
[0004] At present, after the server estimates the depth map of the scene and object based on the texture map, the depth value is quantized into 8-bit binary data to express it as a depth map. The texture maps of multiple synchronized perspectives and the depth maps of the corresponding perspectives are spliced to obtain a spliced image, and then the spliced image and the corresponding parameter data are compressed according to the frame sequence to obtain a free viewpoint video and transmit it, so that the terminal device can reconstruct the free viewpoint image based on the obtained free viewpoint video stream.
[0005] The inventors have found through research that the quality of the free viewpoint image reconstructed by the current depth map quantization processing method is limited by the current depth map quantization processing method. Summary of the invention
[0006] In view of this, the embodiments of the present specification provide a depth map and video processing, reconstruction method, device, equipment and storage medium, which can improve the image quality of the reconstructed free viewpoint video.
[0007] The embodiment of this specification provides a depth map processing method, including:
[0008] Obtaining an estimated depth map generated based on a plurality of frame-synchronized texture maps, wherein the plurality of texture maps have different viewing angles;
[0009] Obtaining a depth value of a pixel in the estimated depth map;
[0010] The quantization parameter data corresponding to the estimated depth map viewing angle is obtained and quantized on the depth values of the pixels in the estimated depth map to obtain quantized depth values of the corresponding pixels in the quantized depth map.
[0011] Optionally, the acquiring and, based on the quantization parameter data corresponding to the estimated depth map viewing angle, quantizing the depth values of corresponding pixels in the estimated depth map to obtain quantized depth values of corresponding pixels in the quantized depth map includes:
[0012] Obtaining a minimum value of the depth distance from the optical center and a maximum value of the depth distance from the optical center corresponding to the viewing angle of the estimated depth map;
[0013] Based on the minimum depth distance from the optical center and the maximum depth distance from the optical center of the corresponding viewing angle of the estimated depth map, the corresponding quantization formula is used to quantize the depth value of the corresponding pixel in the estimated depth map to obtain the quantized depth value of the corresponding pixel in the quantized depth map.
[0014] Optionally, the step of quantizing the depth values of corresponding pixels in the estimated depth map using a corresponding quantization formula based on the minimum depth distance from the optical center and the maximum depth distance from the optical center of the viewing angle corresponding to the estimated depth map to obtain a quantized depth value of the corresponding pixel in the quantized depth map includes:
[0015] The depth value of the corresponding pixel in the estimated depth map is quantized using the following quantization formula:
[0016]
[0017] Among them, M is the quantization bit of the pixel corresponding to the estimated depth map, range is the depth value of the corresponding pixel in the estimated depth map, Depth is the quantized depth value of the corresponding pixel in the estimated depth map, N is the viewing angle corresponding to the estimated depth map, depth_range_near_N is the minimum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N, and depth_range_far_N is the maximum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N.
[0018] Optionally, the method further comprises:
[0019] Downsampling the quantized depth map to obtain a first depth map;
[0020] The texture maps of the multiple perspectives of the frame synchronization and the first depth maps of the corresponding perspectives are spliced according to a preset splicing method to obtain a spliced image.
[0021] The embodiment of this specification also provides a free viewpoint video reconstruction method, the method comprising:
[0022] Acquire a free viewpoint video, the free viewpoint video comprising a plurality of spliced images at frame moments and parameter data corresponding to the spliced images, the spliced images comprising texture maps of a plurality of synchronized viewing angles and a quantized depth map of a corresponding viewing angle, the parameter data corresponding to the spliced images comprising: quantized parameter data of an estimated depth map of a corresponding viewing angle and camera parameter data;
[0023] Obtaining a quantized depth value of a pixel in the quantized depth map;
[0024] Obtaining and, based on quantization parameter data of an estimated depth map corresponding to a viewing angle of the quantized depth map, performing dequantization processing on quantized depth values of pixels in the quantized depth map to obtain a corresponding estimated depth map;
[0025] Based on the synchronized texture maps of multiple viewing angles and the estimated depth maps of corresponding viewing angles, the image of the virtual viewpoint is reconstructed according to the position information of the virtual viewpoint and the camera parameter data corresponding to the stitched image.
[0026] Optionally, the acquiring and, based on the quantization parameter data of the estimated depth map corresponding to the viewing angle of the quantized depth map, performing dequantization processing on the quantized depth value in the quantized depth map to obtain the estimated depth map corresponding to the viewing angle includes:
[0027] Obtaining a minimum value of the depth distance from the optical center and a maximum value of the depth distance from the optical center corresponding to the viewing angle of the estimated depth map;
[0028] Based on the minimum depth distance from the optical center and the maximum depth distance from the optical center of the estimated depth map corresponding to the viewing angle, the quantized depth value in the quantized depth map is dequantized using a corresponding dequantization formula to obtain the depth value of the pixel corresponding to the estimated depth map of the corresponding viewing angle.
[0029] Optionally, the step of performing inverse quantization processing on the quantized depth value in the quantized depth map based on the minimum value of the depth distance from the optical center and the maximum value of the depth distance from the optical center of the viewing angle corresponding to the estimated depth map, using a corresponding inverse quantization formula to obtain a depth value of a pixel corresponding to the estimated depth map of the corresponding viewing angle includes:
[0030] The quantized depth values in the quantized depth map are dequantized using the following dequantization formula to obtain corresponding pixel values in the estimated depth map:
[0031]
[0032]
[0033]
[0034] Among them, M is the quantization bit of the pixel corresponding to the quantized depth map, range is the depth value of the corresponding pixel in the estimated depth map, Depth is the quantized depth value of the corresponding pixel in the quantized depth map, N is the viewing angle corresponding to the estimated depth map, depth_range_near_N is the minimum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N, depth_range_far_N is the maximum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N, maxdisp is the maximum quantized depth distance corresponding to the viewing angle N, and mindisp is the minimum quantized depth distance corresponding to the viewing angle N.
[0035] Optionally, the resolution of the quantized depth map is smaller than the resolution of the texture map corresponding to the viewing angle; before reconstructing the image of the virtual viewpoint, the method further includes:
[0036] The estimated depth map corresponding to the viewing angle is upsampled to obtain a second depth map for reconstructing the virtual viewpoint image.
[0037] Optionally, upsampling the estimated depth map of the corresponding viewing angle to obtain a second depth map for reconstructing the virtual viewpoint image includes:
[0038] Obtaining depth values of pixels in the estimated depth map as pixel values of corresponding even-numbered rows and even-numbered columns in the second depth map;
[0039] For the depth values of the pixels in the even rows and odd columns of the second depth map, determining the corresponding pixel in the corresponding texture map as the middle pixel, based on the relationship between the brightness channel value of the middle pixel in the corresponding texture map and the brightness channel values of the left pixel and the right pixel corresponding to the middle pixel;
[0040] For the depth values of odd-numbered rows of pixels in the second depth map, the corresponding pixels in the corresponding texture map are determined as intermediate pixels based on the relationship between the brightness channel value of the intermediate pixels in the corresponding texture map and the brightness channel values of the upper pixels and the brightness channel values of the lower pixels corresponding to the intermediate pixels.
[0041] Optionally, the texture maps based on multiple perspectives and the estimated depth maps of corresponding perspectives, reconstructing the image of the virtual viewpoint according to the acquired position information of the virtual viewpoint and the camera parameter data corresponding to the stitched image, comprises:
[0042] According to the position information of the virtual viewpoint and the camera parameter data corresponding to the stitched image, a plurality of target texture maps and target depth maps are selected from the synchronized texture maps of the plurality of perspectives and the estimated depth maps of the corresponding perspectives;
[0043] The target texture map and the target depth map are combined and rendered to obtain an image of the virtual viewpoint.
[0044] The embodiment of this specification also provides a free viewpoint video processing method, the method comprising:
[0045] Acquire a free viewpoint video, the free viewpoint video comprising a plurality of spliced images at frame moments and parameter data corresponding to the spliced images, the spliced images comprising texture maps of a plurality of synchronized viewing angles and a quantized depth map of a corresponding viewing angle, the parameter data corresponding to the spliced images comprising: quantized parameter data of an estimated depth map of a corresponding viewing angle and camera parameter data;
[0046] Obtaining a quantized depth value of a pixel in the quantized depth map;
[0047] Obtaining and, based on quantization parameter data of an estimated depth map corresponding to a viewing angle of the quantized depth map, performing dequantization processing on quantized depth values of pixels in the quantized depth map to obtain a corresponding estimated depth map;
[0048] In response to the user interaction behavior, determining the position information of the virtual viewpoint;
[0049] Based on the synchronized texture maps of multiple viewing angles and the estimated depth maps of corresponding viewing angles, the image of the virtual viewpoint is reconstructed according to the position information of the virtual viewpoint and the camera parameter data corresponding to the stitched image.
[0050] Optionally, determining the position information of the virtual viewpoint in response to the user interaction behavior includes: determining corresponding virtual viewpoint path information in response to the user's gesture interaction operation;
[0051] The method reconstructs the image of the virtual viewpoint based on the texture maps of the synchronized multiple viewpoints and the estimated depth maps of the corresponding viewpoints according to the acquired position information of the virtual viewpoint and the camera parameter data corresponding to the stitched image, including:
[0052] According to the virtual viewpoint path information, a texture map in the spliced image at a corresponding frame moment and an estimated depth map of a corresponding viewing angle are selected as a target texture map and a target depth map;
[0053] The target texture map and the target depth map are combined and rendered to obtain an image of the virtual viewpoint.
[0054] Optionally, the method further comprises:
[0055] Acquire a virtual rendering target object in the image of the virtual viewpoint;
[0056] Acquire a virtual information image generated based on augmented reality special effect input data of the virtual rendering target object;
[0057] The virtual information image and the image of the virtual viewpoint are synthesized and displayed.
[0058] Optionally, the acquiring of a virtual information image generated based on augmented reality special effect input data of the virtual rendering target object comprises:
[0059] According to the position of the virtual rendering target object in the image of the virtual viewpoint obtained by three-dimensional calibration, a virtual information image matching the position of the virtual rendering target object is obtained.
[0060] Optionally, acquiring a virtual rendering target object in the image of the virtual viewpoint includes:
[0061] In response to the special effect generation interaction control instruction, a virtual rendering target object in the image of the virtual viewpoint is obtained.
[0062] The embodiment of this specification provides a depth map processing device, the device comprising:
[0063] An estimated depth map acquisition unit, adapted to acquire an estimated depth map generated based on a plurality of frame-synchronized texture maps, wherein the plurality of texture maps have different viewing angles;
[0064] A depth value acquisition unit, adapted to acquire a depth value of a pixel in the depth map;
[0065] a quantization parameter data acquisition unit, adapted to acquire quantization parameter data corresponding to the estimated depth map viewing angle;
[0066] The quantization processing unit is adapted to perform quantization processing on the depth values of the pixels in the estimated depth map based on the quantization parameter data corresponding to the viewing angle of the estimated depth map, so as to obtain the quantized depth values of the corresponding pixels in the quantized depth map.
[0067] The embodiment of this specification also provides a free viewpoint video reconstruction device, the device comprising:
[0068] A first video acquisition unit is adapted to acquire a free viewpoint video, wherein the free viewpoint video includes a stitched image at a plurality of frame moments and parameter data corresponding to the stitched image, wherein the stitched image includes a texture map of a plurality of synchronized viewing angles and a quantized depth map of a corresponding viewing angle, and the parameter data corresponding to the stitched image includes: quantized parameter data of an estimated depth map of a corresponding viewing angle and camera parameter data;
[0069] A first quantized depth value acquisition unit, adapted to obtain a quantized depth value of a pixel in the quantized depth map;
[0070] A first quantization parameter data acquisition unit, adapted to acquire quantization parameter data corresponding to the quantized depth map viewing angle;
[0071] A first depth map dequantization processing unit, adapted to perform dequantization processing on the quantized depth map of the corresponding viewing angle based on the quantization parameter data corresponding to the viewing angle of the quantized depth map, to obtain a corresponding estimated depth map;
[0072] The first image reconstruction unit is adapted to reconstruct an image of the virtual viewpoint based on texture maps of multiple viewpoints and estimated depth maps of corresponding viewpoints, according to the acquired position information of the virtual viewpoint and camera parameter data corresponding to the stitched image.
[0073] The embodiment of this specification also provides a free viewpoint video processing device, the device comprising:
[0074] A second video acquisition unit is adapted to acquire a free viewpoint video, wherein the free viewpoint video includes a stitched image at a plurality of frame moments and parameter data corresponding to the stitched image, wherein the stitched image includes a texture map of a plurality of synchronized viewing angles and a quantized depth map of a corresponding viewing angle, and the parameter data corresponding to the stitched image includes: quantized parameter data of an estimated depth map of a corresponding viewing angle and camera parameter data;
[0075] A second quantized depth value acquisition unit, adapted to acquire a quantized depth value of a pixel in the quantized depth map;
[0076] A second depth map dequantization processing unit, adapted to obtain and, based on quantization parameter data of an estimated depth map corresponding to a viewing angle of the quantized depth map, dequantize quantized depth values of pixels in the quantized depth map to obtain a corresponding estimated depth map;
[0077] A virtual viewpoint position determination unit, adapted to determine position information of a virtual viewpoint in response to user interaction behavior;
[0078] The second image reconstruction unit is adapted to reconstruct the image of the virtual viewpoint based on the synchronized texture maps of multiple viewpoints and the estimated depth maps of the corresponding viewpoints, according to the position information of the virtual viewpoint and the camera parameter data corresponding to the stitched image.
[0079] An embodiment of the present specification also provides an electronic device, including a memory and a processor, wherein the memory stores computer instructions that can be executed on the processor, wherein the processor executes the steps of the method described in any of the above embodiments when executing the computer instructions.
[0080] The embodiment of this specification also provides a server device, including a processor and a communication component, wherein:
[0081] The processor is adapted to execute the steps of the depth map processing method described in any of the foregoing embodiments to obtain a quantized depth map, splice the texture maps of multiple perspectives of frame synchronization and the first depth map of the corresponding perspective in a preset splicing manner to obtain a spliced image, and encapsulate the spliced images of multiple frames and the corresponding parameter data to obtain a free viewpoint video;
[0082] The communication component is suitable for transmitting the free viewpoint video.
[0083] The embodiment of this specification also provides a terminal device, including a communication component, a processor and a display component, wherein:
[0084] The communication component is adapted to acquire free viewpoint video;
[0085] The processor is adapted to execute the steps of the free viewpoint video reconstruction method or the free viewpoint video processing method described in any of the above embodiments;
[0086] The display component is suitable for displaying the reconstructed image obtained by the processor.
[0087] The embodiments of the present specification also provide a computer-readable storage medium on which computer instructions are stored, wherein the computer instructions, when executed, execute the steps of the method described in any of the aforementioned embodiments.
[0088] Compared with the prior art, the technical solution of the embodiment of this specification has the following beneficial effects:
[0089] By using the depth map processing method in the embodiment of the present specification, during the depth map quantization process, a quantization parameter that matches the actual situation of the corresponding viewing angle is used to quantize the depth values of the pixels in the estimated depth map, so that for the depth map of each viewing angle, the expression space of the depth quantization bit can be fully utilized, thereby improving the image quality of the reconstructed free viewpoint video.
[0090] Furthermore, by downsampling the quantized depth map to obtain a first depth map, and splicing the first depth map and the texture map of the corresponding perspective according to a preset splicing method, the overall data volume of the spliced image can be reduced, thereby saving storage resources and transmission resources of the spliced image.
[0091] Furthermore, on the one hand, when the decoding resolution of the overall stitched image is limited, by setting the resolution of the quantized depth map smaller than the resolution of the texture map of the corresponding perspective, a texture map with a higher resolution can be transmitted, and then by upsampling the estimated depth map of the corresponding perspective to obtain a second depth map, and based on the texture maps of multiple perspectives synchronized in the stitched image and the second depth map of the corresponding perspective, free viewpoint video reconstruction is performed, so that a free viewpoint image with higher clarity can be obtained, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] Figure 1 This is a schematic diagram of a specific application system for free viewpoint video display in an embodiment of this specification;
[0093] Figure 2 This is a schematic diagram of an interactive interface of a terminal device in an embodiment of this specification;
[0094] Figure 3 It is a schematic diagram of a collection device setting method in an embodiment of this specification;
[0095] Figure 4 This is another schematic diagram of the interactive interface of a terminal device in an embodiment of this specification;
[0096] Figure 5 This is a schematic diagram of a field of view application scenario in an embodiment of this specification;
[0097] Figure 6 is a flowchart of a depth map processing method in an embodiment of this specification;
[0098] Figure 7 is a schematic diagram of a free viewpoint video data generation process in an embodiment of this specification;
[0099] Figure 8 is a schematic diagram of generation and processing of 6DoF video data in an embodiment of this specification;
[0100] Fig. 9 It is a schematic diagram of the structure of a data header file in an embodiment of this specification;
[0101] Fig.10 is a schematic diagram of 6DoF video data processing on the user side in an embodiment of this specification;
[0102] Fig.11 is a structural schematic diagram of a spliced image in an embodiment of this specification;
[0103] Fig.12 is a flow chart of a free viewpoint video reconstruction method in an embodiment of this specification;
[0104] Fig.13is a flow chart of a combined rendering method in an embodiment of this specification;
[0105] Fig.14 is a flow chart of a free viewpoint video processing method in an embodiment of this specification;
[0106] Fig.15 is a flow chart of another free viewpoint video processing method in an embodiment of this specification;
[0107] Figures 16 to 20 It is a schematic diagram of a display interface of an interactive terminal in an embodiment of this specification;
[0108] Fig.21 is a schematic diagram of the structure of a depth map processing device in an embodiment of this specification;
[0109] Fig. 22 is a schematic structural diagram of a free viewpoint video reconstruction device in an embodiment of this specification;
[0110] Fig.23 is a structural schematic diagram of a free viewpoint video processing device in an embodiment of this specification;
[0111] Fig.24 It is a structural schematic diagram of an electronic device in an embodiment of this specification;
[0112] Fig.25 It is a structural diagram of a server device in an embodiment of this specification;
[0113] Fig.26 It is a structural diagram of a terminal device in an embodiment of this specification. DETAILED DESCRIPTION
[0114] In order to enable those skilled in the art to better understand and implement the embodiments in this specification, the following first provides an exemplary introduction to the implementation of free viewpoint video in conjunction with the accompanying drawings and specific application scenarios.
[0115] refer to Figure 1 In the embodiment of the present invention, a specific application system for free viewpoint video display may include a collection system 11, a server 12 and a display device 13 of multiple collection devices, wherein the collection system 11 may collect images of the viewing area; the collection system 11 or the server 12 may process the acquired synchronized multiple texture maps to generate multi-angle free viewpoint data that can support the display device 13 to switch virtual viewpoints. The display device 13 may display a reconstructed image generated based on the multi-angle free viewpoint data, the reconstructed image corresponds to a virtual viewpoint, and may display reconstructed images corresponding to different virtual viewpoints according to user instructions, switching the viewing position and viewing angle.
[0116] In a specific implementation, the process of reconstructing an image and obtaining a reconstructed image can be implemented by the display device 13, or by a device located in a content delivery network (Content Delivery Network, CDN) in an edge computing manner. It is understandable that Figure 1 This is only an example and is not intended to limit the collection system, server, terminal device, or specific implementation method.
[0117] Continue to refer Figure 1 , the user can view the area to be viewed through the display device 13. In this embodiment, the area to be viewed is a basketball court. As mentioned above, the viewing position and viewing angle can be switched.
[0118] For example, the user can slide on the screen to switch the virtual viewpoint. Figure 2 , the user's finger moves along D 22 When you slide the screen in the right direction, you can switch the virtual viewpoint for viewing. Figure 3 , the position of the virtual viewpoint before sliding may be VP1, and after sliding the screen to switch the virtual viewpoint, the position of the virtual viewpoint may be VP2. Figure 4 After sliding the screen, the reconstructed image displayed on the screen can be Figure 4 The reconstructed image can be obtained by reconstructing the image based on multi-angle free view data generated by images collected by multiple collection devices in an actual collection scenario.
[0119] It is understandable that the image viewed before switching may also be a reconstructed image. The reconstructed image may be a frame image in a video stream. In addition, the method of switching the virtual viewpoint according to the user's instruction may be various, which is not limited here.
[0120] In a specific implementation, the virtual viewpoint can be represented by 6 degrees of freedom (DoF) coordinates, where the spatial position of the virtual viewpoint can be represented as (x, y, z), and the viewing angle can be represented as three rotation directions (θ, ).
[0121] Virtual viewpoint is a three-dimensional concept, and three-dimensional information is required to generate a reconstructed image. In a specific implementation, the multi-angle free viewpoint data may include depth map data to provide the third dimension information outside the plane image. Compared with other implementations, such as providing three-dimensional information through point cloud data, the amount of depth map data is relatively small.
[0122] In the embodiment of the present invention, the switching of the virtual viewpoint can be performed within a certain range, which is the multi-angle free viewing angle range. That is, within the multi-angle free viewing angle range, the virtual viewpoint position and viewing angle can be switched arbitrarily.
[0123] The multi-angle free viewing angle range is related to the arrangement of the acquisition equipment. The wider the shooting coverage of the acquisition equipment, the larger the multi-angle free viewing angle range. The image quality displayed by the terminal device is related to the number of acquisition devices. Generally, the more acquisition devices are set, the fewer the hollow areas in the displayed image.
[0124] In addition, the range of the multi-angle free viewing angle is related to the spatial distribution of the acquisition device. The range of the multi-angle free viewing angle and the interaction mode with the display device on the terminal side can be set based on the spatial distribution relationship of the acquisition device.
[0125] like Figure 1 and Figure 3 As shown, at a height H above the basket LK , several collection devices are arranged along a certain path, for example, 6 collection devices, namely, collection devices CJ1 to CJ6, can be arranged along an arc. It can be understood that the arrangement position, quantity and support method of the collection devices can be various, and are not limited here.
[0126] It is understandable that the above specific application scenario examples are used to better understand the embodiments of this specification, however, the embodiments of this specification are not limited to the above specific application scenarios. The inventors have found through research that the current depth map processing method still has some limitations, resulting in the quality of the image in the reconstructed free viewpoint video being affected.
[0127] In response to the above problems, the embodiments of this specification provide corresponding depth map processing methods and free viewpoint video reconstruction methods. In order to make the purpose, scheme, principle and effect of the embodiments of this specification clearer, the following refers to the accompanying drawings and provides a detailed description through specific embodiments.
[0128] In image processing, the sampled image value is represented by a number. The process of converting the continuous value of the image function into its digital equivalent is quantization. Image quantization assigns an integer number to each continuous sample value.
[0129] At present, after the server device (such as server 12) estimates the depth map of the scene and the object based on the texture map, it quantizes the depth value into 8-bit binary data to express it as a depth map, which is called the estimated depth map for the convenience of description. The texture maps of multiple synchronized perspectives and the estimated depth maps of the corresponding perspectives are spliced to obtain a spliced image. The spliced image and the corresponding parameter data are compressed according to the frame timing to obtain a free viewpoint video. The terminal device can reconstruct the free viewpoint image based on the freely acquired free viewpoint video.
[0130] However, the inventors have discovered through research that the current depth map quantization processing method quantizes each depth map in the stitched image based on a same set of quantization parameter data, and the quality of the reconstructed free viewpoint image is limited by the current depth map quantization processing method.
[0131] Specifically, based on the maximum and minimum depth values in the predefined field of view and the depth value of each pixel in the estimated depth map estimated by the server, quantization is performed through the following formula to obtain an 8-bit binary quantized depth value with a value range of 0-255.
[0132]
[0133] Among them, range is the depth value of the corresponding pixel in the estimated depth map, Depth is the quantized depth value of the corresponding pixel in the estimated depth map, depth_range_near is the minimum depth distance from the optical center in the preset field of view, and depth_range_far is the maximum depth distance from the optical center in the preset field of view.
[0134] However, in specific application scenarios, using the same set of fixed quantization parameter data to quantize the depth values of pixels in all estimated depth maps in the stitched image may result in the inability to fully utilize the depth map's expression space. For example, the depth_range_near, that is, the nearest object, of the estimated depth map corresponding to some perspectives is larger than that of the estimated depth map corresponding to other perspectives. Therefore, after quantization using the same quantization parameter data, the entire 8-bit binary expression space is not fully utilized, and the maximum depth value of the pixel in some perspectives will be far away from 255, while the minimum depth value of the pixel in some perspectives will be far away from 0.
[0135] like Figure 5 The schematic diagram of the field of view application scenario shown in FIG. 5 shows a scene area 50, which includes an object R and is provided with a plurality of acquisition devices P1, P2, ... n …P N , collection equipment P1~P N The optical centers are arranged in an arc shape, and are C1, C2, ... n …C N , each collection device P1~P N Corresponding optical axes L1, L2...L n …L N from Figure 5 Object R and each collection device P1~P N The optical center C1~C NThe spatial relationship between them can be seen intuitively. The minimum and maximum distances between the object R and the optical center of each acquisition device are different. Therefore, based on the acquisition devices P1~P N The estimated depth map obtained by estimating the collected texture map has different minimum and maximum depth distances from the optical center.
[0136] Based on this, in the depth map quantization process, the embodiments of this specification use quantization parameters that match the actual situation of the corresponding viewing angle to quantize the depth values of the pixels in the estimated depth map, so that for the depth map of each viewing angle, the expression space of the depth quantization bits can be fully utilized, thereby improving the image quality of the reconstructed free viewpoint video.
[0137] Reference Figure 6 The flowchart of the depth map processing method shown in the figure, the embodiment of this specification may specifically include the following quantization processing steps:
[0138] S61, obtaining an estimated depth map generated based on a plurality of frame-synchronized texture maps, wherein the plurality of texture maps have different viewing angles.
[0139] In the specific implementation, Figure 1 As shown, a collection system composed of multiple collection devices can synchronously collect images to obtain the multiple frame-synchronized texture maps.
[0140] The origin of the acquisition device (such as a camera) coordinate system can be used as the optical center, and the depth value can be the distance from each point in the field of view to the optical center along the optical axis. In a specific implementation, an estimated depth map corresponding to each texture map can be obtained based on the multiple texture maps of the frame synchronization.
[0141] S62: Obtain depth values of pixels in the estimated depth map.
[0142] S63, acquiring and quantizing the depth values of the pixels in the estimated depth map based on the quantization parameter data corresponding to the viewing angle of the estimated depth map, to obtain quantized depth values of the corresponding pixels in the quantized depth map.
[0143] In a specific implementation, the quantization parameter estimation value may include: the minimum depth distance from the optical center and the maximum depth distance from the optical center of the estimated depth map corresponding to the viewing angle. In order to quantize the depth value of the pixel in the estimated depth map, the minimum depth distance from the optical center and the maximum depth distance from the optical center of the estimated depth map corresponding to the viewing angle can be first obtained. Then, based on the minimum depth distance from the optical center and the maximum depth distance from the optical center of the estimated depth map corresponding to the viewing angle, the corresponding quantization formula can be used to quantize the depth value of the corresponding pixel in the estimated depth map to obtain the quantized depth value of the corresponding pixel in the quantized depth map.
[0144] In some embodiments of the present specification, the depth value of the corresponding pixel in the estimated depth map is quantized using the following quantization formula:
[0145]
[0146] Among them, M is the quantization bit of the pixel corresponding to the estimated depth map, range is the depth value of the corresponding pixel in the estimated depth map, Depth is the quantized depth value of the corresponding pixel in the estimated depth map, N is the viewing angle corresponding to the estimated depth map, depth_range_near_N is the minimum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N, and depth_range_far_N is the maximum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N.
[0147] After the pixel depth map in the estimated depth map is quantized by the above embodiment, the object closest to the camera (optical center) in the quantized depth map corresponding to each viewing angle can be quantized to be closer to 2 M In a specific implementation, M can be 8 bits, 16 bits, etc. If M is 8 bits, the depth value of the object closest to the optical center in the quantized depth map corresponding to each viewing angle is quantized to be closer to 255.
[0148] In a specific implementation, the texture maps of multiple synchronized viewpoints can be spliced with the quantized depth maps of the corresponding viewpoints to obtain a spliced image, and then based on the spliced images at multiple frame times and the parameter data corresponding to the spliced images, a free viewpoint video can be obtained. Considering the limitation of the transmission bandwidth, the free viewpoint video can be compressed and transmitted to the terminal device for image reconstruction of the free viewpoint video.
[0149] Combined with reference Figure 7In order to reconstruct free viewpoint video, texture image acquisition and depth map calculation are required, which includes three main steps, namely multi-camera video acquisition (Multi-camera Video Capturing), camera internal and external parameter calculation (Camera Parameter Estimation), and depth map calculation (Depth Map Calculation). For multi-camera acquisition, the videos acquired by each camera are required to be aligned at the frame level. Among them, the texture image (Texture Image) can be obtained through multi-camera video acquisition; the camera parameters (Camera Parameter) can be obtained through camera internal and external parameter calculation, which can include camera internal parameter data and external parameter data; the depth map (Depth Map) can be obtained through depth map calculation. Multiple synchronized texture maps and depth maps of corresponding viewing angles and camera parameters form 6DoF video data.
[0150] In the embodiments of this specification, no special camera, such as a light field camera, is required to capture the video. Similarly, no complicated camera calibration is required before the capture. The positions of multiple cameras can be laid out and arranged to better capture the objects or scenes that need to be captured.
[0151] After the above three steps are completed, the texture map collected from multiple cameras, the camera parameters of all cameras, and the depth map of each camera are obtained. These three parts of data can be called data files in multi-angle free-view video data, or 6-degree-of-freedom video data (6DoF video data). With this data, the user end can generate virtual viewpoints based on the virtual 6-degree-of-freedom (DoF) position, thereby providing a 6DoF video experience.
[0152] Combined with reference Figure 8 , 6DoF video data and indicative data can be compressed and transmitted to the user side, and the user side can obtain the user side 6DoF expression, that is, the aforementioned 6DoF video data and metadata, based on the received data. Among them, the indicative data can also be called metadata.
[0153] Metadata can be used to describe the data mode of 6DoF video data, which may include: Stitching Pattern metadata, which is used to indicate the storage rules of pixel data of multiple texture maps and quantized depth map data in the stitched image; Padding pattern metadata, which can be used to indicate the way to perform edge protection in the stitched image, quantization parameter metadata corresponding to the viewing angle, and other metadata. Metadata can be stored in the data header file, and the specific storage order can be as follows: Fig. 9 as shown, or stored in another order.
[0154] Combined with reference Fig.10 , the user side obtains 6DoF video data, which includes camera parameters, texture maps, quantized depth maps, and metadata, in addition to the user's interactive behavior data. With this data, the user side can use 6DoF rendering based on depth image rendering (DIBR, Depth Image-Based Rendering) to generate a virtual viewpoint image at a specific 6DoF position generated by the user's interactive behavior, that is, according to the user's instruction, determine the virtual viewpoint at the 6DoF position corresponding to the instruction.
[0155] At present, any video frame in free viewpoint video data is generally expressed as a spliced image formed by texture maps collected by multiple cameras and corresponding depth maps. Fig.11 The structural schematic diagram of the stitched image shown in the figure, wherein the upper half of the stitched image is the texture map area, which is divided into 8 texture map sub-areas, respectively storing the pixel data of 8 synchronized texture maps, and each texture map has a different shooting angle, that is, a different viewing angle. The lower half of the stitched image is the depth map area, which is divided into 8 depth map sub-areas, respectively storing the corresponding quantized depth maps of the above 8 texture maps. Among them, the texture map of viewing angle N and the quantized depth map of viewing angle N are pixel-to-pixel corresponding, and the stitched image is compressed and transmitted to the terminal for decoding and DIBR, so that the image can be interpolated at the user's interactive viewpoint.
[0156] The inventor has found through research that Fig.11 As shown, for each texture map, there is a corresponding quantized depth map of the same resolution, so that the resolution of the overall stitched image is twice that of the texture map set. Since the video decoding resolution of the terminal (such as a mobile terminal) is generally limited, the above-mentioned free viewpoint video data expression method can only be achieved by reducing the resolution of the texture map, which results in a decrease in the clarity of the reconstructed image perceived by the user on the terminal side.
[0157] To address the above problem, in some embodiments of the present specification, the quantized depth map can be first downsampled to obtain a first depth map, and the synchronized texture maps of multiple perspectives and the first depth map of the corresponding perspective can be spliced in a preset splicing method to obtain a spliced image.
[0158] In order to enable those skilled in the art to better understand and implement the embodiments of this specification, two specific downsampling processing examples are given below:
[0159] One is to perform sampling processing on the pixels in the quantized depth map to obtain the first depth map. For example, one pixel may be sampled from every other pixel in the quantized depth map to obtain the first depth map, and the resolution of the obtained first depth map is 50% of the original depth map.
[0160] Another method is to filter the pixels in the quantized depth map based on the corresponding texture map to obtain the first depth map.
[0161] In order to save data storage resources and data transmission resources, the stitched image can be rectangular.
[0162] In order to enable those skilled in the art to better understand and implement the embodiments of this specification, the free viewpoint video reconstruction method on the terminal side after the above-mentioned depth map processing is introduced below through a specific embodiment.
[0163] Reference Fig.12 The flowchart of the free viewpoint video reconstruction method is shown in FIG. 1 . In the embodiment of this specification, the free viewpoint video reconstruction can be performed by specifically adopting the following steps:
[0164] S121, obtaining a free viewpoint video, wherein the free viewpoint video includes stitched images at multiple frame moments and parameter data corresponding to the stitched images, wherein the stitched images include texture maps of multiple synchronized perspectives and quantized depth maps of corresponding perspectives, and the parameter data corresponding to the stitched images include: quantized parameter data of the estimated depth map of the corresponding perspective and camera parameter data.
[0165] In a specific implementation, the free viewpoint video may be in the form of a video compression file or may be transmitted in the form of a video stream. The parameter data of the spliced image may be stored in a header file of the free viewpoint video data, and the specific form may refer to the introduction of the above embodiment.
[0166] In some embodiments of the present specification, the quantization parameter data of the estimated depth map corresponding to the viewing angle may be stored in the form of an array. For example, for a free viewpoint video with 16 sets of texture maps and quantized depth maps in the spliced image, the quantization parameter data may be represented in sequence as follows:
[0167] Array Z = [view 0 parameter value, view 2 parameter value…view 15 quantization parameter value].
[0168] S122: Obtain quantized depth values of pixels in the quantized depth map.
[0169] S123, obtaining and dequantizing the quantized depth values of the pixels in the quantized depth map based on the quantized depth map's quantized parameter data of the estimated depth map corresponding to the viewing angle, to obtain a corresponding estimated depth map.
[0170] S124, based on the synchronized texture maps of multiple viewing angles and the estimated depth maps of corresponding viewing angles, reconstructing the image of the virtual viewpoint according to the position information of the virtual viewpoint and the camera parameter data corresponding to the stitched image.
[0171] For step S123, in some embodiments of this specification, the following method is adopted:
[0172] Obtaining a minimum value of the depth distance from the optical center and a maximum value of the depth distance from the optical center corresponding to the viewing angle of the estimated depth map;
[0173] Based on the minimum depth distance from the optical center and the maximum depth distance from the optical center of the estimated depth map corresponding to the viewing angle, the quantized depth value in the quantized depth map is dequantized using a corresponding dequantization formula to obtain the depth value of the pixel corresponding to the estimated depth map of the corresponding viewing angle.
[0174] In a specific embodiment of the present specification, the quantized depth value in the quantized depth map is dequantized using the following dequantization formula to obtain the corresponding pixel value in the estimated depth map:
[0175]
[0176]
[0177]
[0178] Among them, M is the quantization bit of the pixel corresponding to the quantized depth map, range is the depth value of the corresponding pixel in the estimated depth map, Depth is the quantized depth value of the corresponding pixel in the quantized depth map, N is the viewing angle corresponding to the estimated depth map, depth_range_near_N is the minimum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N, depth_range_far_N is the maximum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N, maxdisp is the maximum quantized depth distance corresponding to the viewing angle N, and mindisp is the minimum quantized depth distance corresponding to the viewing angle N.
[0179] In some embodiments of the present specification, corresponding to the aforementioned embodiments, if the resolution of the quantized depth map is smaller than the resolution of the texture map of the corresponding perspective, for example, on the server side, the quantized depth map has been downsampled, then on the terminal device side, the estimated depth map of the corresponding perspective can be upsampled to obtain a second depth map, and then the second depth map is used to reconstruct the virtual viewpoint image.
[0180] In a specific implementation, there may be a variety of upsampling methods, some examples of which are given below:
[0181] In the example of the first method, the estimated depth map that has been downsampled by 1 / 4 is upsampled to obtain a second depth map with the same resolution as the texture map. Based on different rows and columns, the following different processing methods are specifically included:
[0182] (1) obtaining a depth value of a pixel in the estimated depth map as a pixel value of a corresponding even-numbered row and even-numbered column in the second depth map;
[0183] (2) for the depth values of pixels in even rows and odd columns of the second depth map, determining the corresponding pixel in the corresponding texture map as the middle pixel, based on the relationship between the brightness channel value of the middle pixel in the corresponding texture map and the brightness channel values of the left pixel and the right pixel corresponding to the middle pixel;
[0184] Specifically, for the depth values of odd-numbered rows of pixels in the second depth map, the corresponding pixels in the corresponding texture map are determined as intermediate pixels based on the relationship between the brightness channel value of the intermediate pixels in the corresponding texture map and the brightness channel values of the upper pixels and the brightness channel values of the lower pixels corresponding to the intermediate pixels.
[0185] Specifically, based on the relationship between the brightness channel value of the middle pixel in the corresponding texture map and the brightness channel values of the left pixel and the right pixel corresponding to the middle pixel, there are three cases:
[0186] a1. If the absolute value of the difference between the brightness channel value of the middle pixel in the corresponding texture map and the brightness channel value of the right pixel corresponding to the middle pixel is less than the quotient of the absolute value of the difference between the brightness channel value of the middle pixel and the brightness channel value of the left pixel and a preset threshold, then the depth value corresponding to the right pixel is selected as the depth value of the corresponding pixel in the even row and odd column of the second depth map, that is:
[0187] a2. If the absolute value of the difference between the luminance channel value of the middle pixel in the corresponding texture map and the luminance channel value of the left pixel corresponding to the middle pixel is less than the quotient of the absolute value of the difference between the luminance channel value of the middle pixel and the luminance channel value of the right pixel and the preset threshold, then select the depth value corresponding to the left pixel as the depth value of the corresponding pixel in the odd columns of the even rows in the second depth map;
[0188] a3. Otherwise, select the maximum value of the depth values corresponding to the left pixel and the right pixel as the depth value of the corresponding pixel in the odd columns of the even rows in the second depth map.
[0189] (3) For the depth values of the pixels in the odd rows of the second depth map, determine the corresponding pixel in the corresponding texture map as the middle pixel, and determine based on the relationship between the luminance channel value of the middle pixel in the corresponding texture map and the luminance channel values of the upper pixel and the lower pixel corresponding to the middle pixel.
[0190] b1. If the absolute value of the difference between the luminance channel value of the middle pixel in the corresponding texture map and the luminance channel value of the lower pixel corresponding to the middle pixel is less than the quotient of the absolute value of the difference between the luminance channel value of the middle pixel and the luminance channel value of the upper pixel and the preset threshold, then select the depth value corresponding to the lower pixel as the depth value of the corresponding pixel in the odd rows in the second depth map;
[0191] b2. If the absolute value of the difference between the luminance channel value of the middle pixel in the corresponding texture map and the luminance channel value of the upper pixel corresponding to the middle pixel is less than the quotient of the absolute value of the difference between the luminance channel value of the middle pixel and the luminance channel value of the lower pixel and the preset threshold, then select the depth value corresponding to the upper pixel as the depth value of the corresponding pixel in the odd rows in the second depth map;
[0192] b3. Otherwise, select the maximum value of the depth values corresponding to the upper pixel and the lower pixel as the depth value of the corresponding pixel in the odd columns of the even rows in the second depth map.
[0193] The three cases a1 to a3 in the above step (2) can be expressed by the formula:
[0194] If abs(pix_C - pix_R) < abs(pix_C - pix_L) / THR, then select Dep_R;
[0195] If abs(pix_C - pix_L) < abs(pix_C - pix_R) / THR, then select Dep_L;
[0196] Otherwise, for other cases, select Max(Dep_R, Dep_L).
[0197] In the above step (3), the three cases of b1 to b3 can be expressed by the formula:
[0198] If abs(pix_C - pix_D) < abs(pix_C - pix_U) / THR, then Dep_D is adopted;
[0199] If abs(pix_C - pix_U) < abs(pix_C - pix_D) / THR, then Dep_U is selected;
[0200] Otherwise, for other cases, Max(Dep_D, Dep_U) is selected.
[0201] In the above formula, pix_C is the luminance channel value (Y value) of the middle pixel in the texture map at the position corresponding to the depth value in the second depth map, pix_L is the luminance channel value of the pixel to the left of pix_C, pix_R is the luminance channel value of the pixel to the right of pix_C, pix_U is the luminance channel value of the pixel above pix_C, pix_D is the luminance channel value of the pixel below pix_C, Dep_R is the depth value corresponding to the pixel to the right of the middle pixel in the texture map at the position corresponding to the depth value in the second depth map, Dep_L is the depth value corresponding to the pixel to the left of the middle pixel in the texture map at the position corresponding to the depth value in the second depth map, Dep_D is the depth value corresponding to the pixel below the middle pixel in the texture map at the position corresponding to the depth value in the second depth map, and Dep_U is the depth value corresponding to the pixel above the middle pixel in the texture map at the position corresponding to the depth value in the second depth map. abs represents the absolute value, and THR is a settable threshold. In an embodiment of this specification, THR is set to 2.
[0202] Example of Method 2:
[0203] Obtain the depth value of the pixel in the estimated depth map as the pixel value of the corresponding row and column in the second depth map; for the pixels in the second depth map that have no corresponding relationship with the pixels in the estimated depth map, filter them based on the difference between the corresponding pixel in the corresponding texture map and the pixel values of the surrounding pixels of the corresponding pixel.
[0204] Among them, for the pixels in the second depth map that have no corresponding relationship with the pixels in the estimated depth map, filter them based on the difference between the corresponding pixel in the corresponding texture map and the pixel values of the surrounding pixels of the corresponding pixel.
[0205] There can be various specific filtering methods, and the following gives two specific embodiments.
[0206] Specific Embodiment 1, Nearest Neighbor Filtering Method
[0207] Specifically, for pixels in the second depth map that do not correspond to pixels in the estimated depth map, the pixel values of the corresponding pixels in the texture map can be compared with the pixel values of the four diagonal pixels around the corresponding pixels to obtain the pixel point closest to the pixel value of the corresponding pixel, and the depth value in the estimated depth map corresponding to the pixel point closest to the pixel value is used as the depth value of the pixel corresponding to the corresponding pixel in the texture map in the second depth map.
[0208] Specific embodiment 2: weighted filtering method
[0209] Specifically, the corresponding pixels in the texture map can be compared with the pixels around the corresponding pixels, and the depth values in the estimated depth map corresponding to the surrounding pixels can be weighted according to the similarity of the pixel values to obtain the depth values of the corresponding pixels in the texture map in the second depth map.
[0210] The above shows some methods for upsampling the estimated depth map to obtain the second depth map. It can be understood that the above is only an example, and the specific upsampling method is not limited in the embodiments of this specification. In addition, the method for upsampling the estimated depth map in any video frame may correspond to the method for downsampling the quantized depth map to obtain the first depth map, or there may be no corresponding relationship. In addition, the upsampling ratio and the downsampling ratio may be the same or different.
[0211] Next, some specific examples are given for step S124.
[0212] In a specific implementation, in order to save data processing resources and improve image reconstruction efficiency while ensuring image reconstruction quality, only part of the texture map and the estimated depth map of the corresponding viewing angle in the stitched image can be selected as the target texture map and the target depth map for reconstruction of the virtual viewpoint image. Specifically:
[0213] According to the position information of the virtual viewpoint and the parameter data corresponding to the stitched image, multiple target texture maps and target depth maps can be selected from the synchronized texture maps of multiple viewpoints and the estimated depth maps of corresponding viewpoints. Afterwards, the target texture maps and target depth maps can be combined and rendered to obtain the image of the virtual viewpoint.
[0214] In a specific implementation, the position information of the virtual viewpoint can be determined based on user interaction behavior or according to a pre-setting. If it is determined based on user interaction behavior, the virtual viewpoint position at the corresponding interaction moment can be determined by obtaining the trajectory data corresponding to the user interaction operation. In some embodiments of this specification, the position information of the virtual viewpoint corresponding to the corresponding video frame can also be pre-set on the server side (such as a server or cloud), and the set virtual viewpoint position information can be transmitted in the header file of the free viewpoint video.
[0215] In a specific implementation, the spatial position relationship between each texture map and the estimated depth map of the corresponding perspective and the virtual viewpoint position can be determined based on the virtual viewpoint position and the parameter data corresponding to the stitched image. In order to save data processing resources, the texture map and the estimated depth map of the corresponding perspective that satisfy a preset position relationship and / or quantity relationship with the virtual viewpoint position can be selected from the synchronized texture maps of multiple perspectives and the estimated depth maps of the corresponding perspective as the target texture map and target depth map according to the position information of the virtual viewpoint and the parameter data corresponding to the stitched image.
[0216] For example, the texture maps and estimated depth maps corresponding to the 2 to N viewpoints closest to the virtual viewpoint position may be selected. Wherein, N is the number of texture maps in the stitched image, that is, the number of acquisition devices corresponding to the texture maps. In a specific implementation, the quantitative relationship value may be fixed or variable.
[0217] Reference Fig.13 The flowchart of the combined rendering method shown in the figure may specifically include the following steps in some embodiments of the present specification:
[0218] S131 , forward mapping the target depth maps in the selected stitched image to the virtual positions.
[0219] S132, post-processing the target depth map after forward mapping.
[0220] In a specific implementation, there may be multiple post-processing methods. In some embodiments of this specification, at least one of the following methods may be used to post-process the target depth map:
[0221] 1) Perform foreground edge protection processing on the target depth map after forward mapping;
[0222] 2) Perform pixel-level filtering on the target depth map after forward mapping.
[0223] S133, reverse mapping the selected target texture images in the spliced image respectively.
[0224] S134, fusing the virtual texture images generated after reverse mapping to obtain a fused texture image.
[0225] Through the above steps S131 to S134, a reconstructed image can be obtained.
[0226] In a specific implementation, the fused texture image may be hole-filled to obtain a reconstructed image corresponding to the virtual viewpoint position at the user interaction moment. The quality of the reconstructed image may be improved by hole-filling.
[0227] In a specific implementation, before image reconstruction, the target depth map can be preprocessed (for example, upsampling), or the estimated depth map obtained after all inverse quantization processes in the stitched image can be preprocessed (for example, upsampling), and then image reconstruction based on the virtual viewpoint can be performed.
[0228] In order to enable those skilled in the art to better understand and implement, the embodiments of this specification also provide specific embodiments such as devices and equipment corresponding to the aforementioned method embodiments, which are described below with reference to the accompanying drawings.
[0229] The present specification also provides a corresponding free viewpoint video processing method. Fig.14 , which may specifically include the following steps:
[0230] S141, obtaining a free viewpoint video, wherein the free viewpoint video includes stitched images at multiple frame moments and parameter data corresponding to the stitched images, wherein the stitched images include texture maps of multiple synchronized perspectives and quantized depth maps of corresponding perspectives, and the parameter data corresponding to the stitched images include: quantized parameter data of the estimated depth map of the corresponding perspective and camera parameter data.
[0231] In a specific implementation, by acquiring a free viewpoint video and decoding the free viewpoint video, the spliced images at the multiple frame moments and the parameter data corresponding to the spliced images can be obtained.
[0232] The free viewpoint video may be in the form of a multi-angle free viewpoint video as exemplified in the aforementioned embodiment, such as a 6DoF video.
[0233] By downloading the free viewpoint video stream or obtaining the stored free viewpoint video data file, a video frame sequence can be obtained. Each video frame may include a spliced image formed by synchronized texture maps of multiple viewpoints and a first depth map of the corresponding viewpoint. A structure of a spliced image is as follows: Fig.11As shown. It is understandable that other structures of stitched images can be used. For example, different stitching methods can be used according to the different ratios of the resolutions of the texture map and the first depth map of the corresponding viewing angle. For example, one texture map can correspond to multiple first depth maps (for example, the first depth map is a depth map after 25% downsampling).
[0234] In addition to the stitched image, the free viewpoint video data file may also include metadata describing the stitched image. In a specific implementation, parameter data of the stitched image may be obtained from the metadata, for example, one or more of the information such as camera parameters of the stitched image, stitching rules of the stitched image, and resolution information of the stitched image may be obtained.
[0235] In a specific implementation, the parameter information of the stitched image can be transmitted in combination with the stitched image, for example, can be stored in a video file header. The embodiments of this specification do not limit the specific format of the stitched image, nor the specific type and storage location of the parameter information of the stitched image, and it is sufficient to be able to obtain a reconstructed image of the corresponding virtual viewpoint position based on the virtual viewpoint video.
[0236] In a specific implementation, the free viewpoint video may be in the form of a video compression file or may be transmitted in the form of a video stream. The parameter data of the spliced image may be stored in a header file of the free viewpoint video data, and the specific form may refer to the introduction of the above embodiment.
[0237] In some embodiments of the present specification, the quantization parameter data of the estimated depth map corresponding to the viewing angle may be stored in the form of an array. For example, for a free viewpoint video with 16 sets of texture maps and quantized depth maps in the spliced image, the quantization parameter data may be represented in sequence as follows:
[0238] Array Z = [view 0 parameter value, view 2 parameter value...view 15 quantization parameter value].
[0239] S142: Obtain quantized depth values of pixels in the quantized depth map.
[0240] S143, obtaining and, based on quantization parameter data of an estimated depth map corresponding to a viewing angle of the quantized depth map, performing dequantization processing on quantized depth values of pixels in the quantized depth map to obtain a corresponding estimated depth map.
[0241] The specific quantization parameter data used in the inverse quantization processing in step S143 and the specific inverse quantization processing method can be found in the introduction of the above-mentioned embodiment, and will not be repeated here.
[0242] S144, determining position information of the virtual viewpoint in response to the user interaction behavior.
[0243] In a specific implementation, if the free viewpoint video adopts the 6DoF expression method, the virtual viewpoint position information based on user interaction can be expressed as coordinates (x, y, z, θ, ), the virtual viewpoint position information can be generated under one or more preset user interaction modes. For example, it can be the coordinates input by the user operation, such as manual click or gesture path, or the virtual position determined by voice input, or a customized virtual viewpoint can be provided for the user (for example: the user can input the position or perspective in the scene, such as the basket, the sidelines, the referee's perspective, the coach's perspective, etc.). Or based on a specific object (for example, a player on the court, an actor or guest in the image, a host, etc., the perspective of the object can be switched after the user clicks on the corresponding object). It can be understood that the specific user interaction behavior mode is not limited in the embodiments of the present invention, as long as the virtual viewpoint position information based on user interaction can be obtained.
[0244] As an optional example, in response to the user's gesture interaction operation, the corresponding virtual viewpoint path information can be determined. In terms of gesture interaction, the corresponding virtual viewpoint path can be planned based on different forms of gesture interaction, so that the corresponding virtual viewpoint path information can be determined based on the user's specific gesture operation. For example, it can be pre-planned that the user's finger slides left and right relative to the touch screen, corresponding to the left and right movement of the viewing angle; the user's finger slides up and down relative to the touch screen, corresponding to the up and down movement of the viewpoint position; the finger zoom operation corresponds to the zooming in and out of the viewpoint position.
[0245] It is understandable that the virtual viewpoint path planned based on the gesture form above is only for exemplary purposes, and virtual viewpoint paths based on other gesture forms may be predefined, or the user may customize the settings, thereby enhancing the user experience.
[0246] S145, based on the synchronized texture maps of multiple viewing angles and the estimated depth maps of corresponding viewing angles, according to the position information of the virtual viewpoint and the camera parameter data corresponding to the stitched image, reconstruct the image of the virtual viewpoint.
[0247] In a specific implementation, the texture map at the corresponding frame moment and the estimated depth map of the corresponding viewing angle can be selected as the target texture map and the target depth map according to the virtual viewpoint path information, and the target texture map and the target depth map are combined and rendered to obtain the image of the virtual viewpoint.
[0248] The specific selection method can be referred to the above-mentioned embodiment, which will not be described in detail here.
[0249] It should be noted that based on the virtual viewpoint path information, part of the texture map and the second depth map of the corresponding perspective in one frame or multiple consecutive frames of stitched images can be selected in time sequence as the target texture map and target depth map to reconstruct the image of the corresponding virtual viewpoint.
[0250] In a specific implementation, the reconstructed free viewpoint image may be further processed. An exemplary expansion method is given below.
[0251] In order to enrich the user's visual experience, augmented reality (AR) special effects can be embedded in the reconstructed free viewpoint image. Fig.15 The flowchart of the free viewpoint video processing method shown in the figure implements the implantation of AR special effects in the following manner:
[0252] S151, obtaining a virtual rendering target object in the image of the virtual viewpoint.
[0253] In a specific implementation, certain objects in the image of the free viewpoint video can be determined as virtual rendering target objects based on certain indication information, and the indication information can be generated based on user interaction, or can be obtained based on certain preset trigger conditions or third-party instructions. In an optional embodiment of the present specification, in response to a special effect generation interactive control instruction, a virtual rendering target object in the image of the virtual viewpoint can be obtained.
[0254] S152, obtaining a virtual information image generated based on the augmented reality special effect input data of the virtual rendering target object.
[0255] In the embodiment of the present specification, the implanted AR special effects are presented in the form of a virtual information image. The virtual information image can be generated based on the augmented reality special effects input data of the target object. After determining the virtual rendering target object, the virtual information image generated based on the augmented reality special effects input data of the virtual rendering target object can be obtained.
[0256] In the embodiments of the present specification, the virtual information image corresponding to the virtual rendering target object may be generated in advance, or may be generated immediately in response to a special effect generation instruction.
[0257] In a specific implementation, a virtual information image that matches the position of the virtual rendering target object can be obtained based on the position of the virtual rendering target object in the reconstructed image obtained by three-dimensional calibration, so that the obtained virtual information image can be more closely matched with the position of the virtual rendering target object in the three-dimensional space, and the displayed virtual information image is more consistent with the real state in the three-dimensional space, so that the displayed composite image is more realistic and vivid, thereby enhancing the user's visual experience.
[0258] In a specific implementation, based on the augmented reality special effect input data of the virtual rendering target object, a virtual information image corresponding to the target object can be generated according to a preset special effect generation method.
[0259] In specific implementations, a variety of special effect generation methods can be used.
[0260] For example, the augmented reality special effect input data of the target object may be input into a preset three-dimensional model, and based on the position of the virtual rendering target object in the image obtained by three-dimensional calibration, a virtual information image matching the virtual rendering target object may be output;
[0261] For example, the augmented reality special effects input data of the virtual rendering target object can be input into a preset machine learning model, and based on the position of the virtual rendering target object in the image obtained by three-dimensional calibration, a virtual information image matching the virtual rendering target object is output.
[0262] S153, synthesizing and displaying the virtual information image and the image of the virtual viewpoint.
[0263] In a specific implementation, there are multiple ways to synthesize and display the virtual information image and the image of the virtual viewpoint. Two specific achievable examples are given below:
[0264] Example 1: Fusing the virtual information image with the corresponding image to obtain a fused image, and displaying the fused image;
[0265] Example 2: superimposing the virtual information image on the corresponding image to obtain a superimposed composite image, and displaying the superimposed composite image.
[0266] In a specific implementation, the obtained composite image can be directly displayed, or the obtained composite image can be inserted into the video stream to be played for display. For example, the fused image can be inserted into the video stream to be played for display.
[0267] The free viewpoint video may include a special effects display identifier. In a specific implementation, the superimposition position of the virtual information image in the image of the virtual viewpoint may be determined based on the special effects display identifier, and then the virtual information image may be superimposed and displayed at the determined superimposition position.
[0268] In order to enable those skilled in the art to better understand and implement, the following is a detailed description of the image display process of an interactive terminal. Figures 16 to 20 The interactive terminal T1 plays the video in real time. Fig.16, displaying the video frame P1. Next, the video frame P2 displayed by the interactive terminal includes a plurality of special effect display identifiers such as the special effect display identifier I1. The video frame P2 is indicated by an inverted triangle symbol pointing to the target object, such as Fig.17 As shown. It is understandable that other methods may also be used to display the special effect display mark. When the terminal user touches and clicks the special effect display mark I1, the system automatically obtains the virtual information image corresponding to the special effect display mark I1, and superimposes the virtual information image on the video frame P3, as shown in FIG. Fig.18 As shown in FIG. 1 , a three-dimensional ring R1 is rendered with the position of the field where the athlete Q1 stands as the center. Next, Fig.19 and Fig. 20 As shown, the terminal user touches and clicks the special effect display mark I2 in the video frame P3, and the system automatically obtains the virtual information image corresponding to the special effect display mark I2, and superimposes and displays the virtual information image on the video frame P3 to obtain a superimposed image, i.e., the video frame P4, in which the hit rate information display board M0 is displayed. The hit rate information display board M0 displays the position, name, and hit rate information of the target object, i.e., the athlete Q2.
[0269] like Figures 16 to 20 As shown, the terminal user can continue to click on other special effect display logos displayed in the video frame to watch videos showing the AR special effects corresponding to each special effect display logo.
[0270] It is understandable that different types of embedded special effects can be distinguished by different types of special effect display logos.
[0271] Reference Fig.21 The structural diagram of the depth map processing device shown in FIG. 2 , wherein the depth map processing device 210 may include: an estimated depth map acquisition unit 211, a depth value acquisition unit 212, a quantization parameter data acquisition unit 213 and a quantization processing unit 214, specifically:
[0272] The estimated depth map acquisition unit 211 is adapted to acquire an estimated depth map generated based on a plurality of frame-synchronized texture maps, wherein the plurality of texture maps have different viewing angles;
[0273] A depth value acquisition unit 212, adapted to acquire a depth value of a pixel in the depth map;
[0274] A quantization parameter data acquisition unit 213, adapted to acquire quantization parameter data corresponding to the estimated depth map viewing angle;
[0275] The quantization processing unit 214 is adapted to perform quantization processing on the depth values of the pixels in the estimated depth map based on the quantization parameter data corresponding to the viewing angle of the estimated depth map, so as to obtain quantized depth values of the corresponding pixels in the quantized depth map.
[0276] In a specific implementation, the quantization parameter data acquisition unit 213 is suitable for obtaining the minimum depth distance from the optical center and the maximum depth distance from the optical center of the estimated depth map corresponding to the viewing angle; accordingly, the quantization processing unit 214 can use the corresponding quantization formula to quantize the depth value of the corresponding pixel in the estimated depth map based on the minimum depth distance from the optical center and the maximum depth distance from the optical center of the viewing angle corresponding to the estimated depth map, so as to obtain the quantized depth value of the corresponding pixel in the quantized depth map.
[0277] The specific quantization principle of the quantization processing unit 214 and the specific quantization formula that can be used can all refer to the description in the above embodiments.
[0278] As an alternative example, continue to refer to Fig.21 The depth map processing device 210 may further include: a downsampling processing unit 215 and a splicing unit 216, wherein:
[0279] The downsampling processing unit 215 is adapted to perform downsampling processing on the quantized depth map to obtain a first depth map;
[0280] The stitching unit 216 is adapted to stitch the synchronized texture maps of the multiple viewing angles and the first depth maps of the corresponding viewing angles in a preset stitching manner to obtain a stitched image.
[0281] Reference Fig. 22 The free viewpoint video reconstruction device 220 is a schematic diagram of the structure of the free viewpoint video reconstruction device shown in the figure, wherein the free viewpoint video reconstruction device 220 may include: a first video acquisition unit 221, a first quantized depth value acquisition unit 222, a first quantization parameter data acquisition unit 223, a first depth map inverse quantization processing unit 224 and a first image reconstruction unit 225, specifically:
[0282] The first video acquisition unit 221 is adapted to acquire a free viewpoint video, wherein the free viewpoint video includes a spliced image at a plurality of frame moments and parameter data corresponding to the spliced image, wherein the spliced image includes a texture map of a plurality of synchronized viewing angles and a quantized depth map of a corresponding viewing angle, and the parameter data corresponding to the spliced image includes: quantized parameter data of an estimated depth map of a corresponding viewing angle and camera parameter data;
[0283] The first quantized depth value acquiring unit 222 is adapted to obtain the quantized depth value of the pixel in the quantized depth map;
[0284] The first quantization parameter data acquisition unit 223 is adapted to acquire quantization parameter data corresponding to the quantized depth map viewing angle;
[0285] The first depth map dequantization processing unit 224 is adapted to perform dequantization processing on the quantized depth map of the corresponding viewing angle based on the quantization parameter data corresponding to the viewing angle of the quantized depth map to obtain a corresponding estimated depth map;
[0286] The first image reconstruction unit 225 is adapted to reconstruct the image of the virtual viewpoint based on the texture maps of multiple viewpoints and the estimated depth maps of the corresponding viewpoints, according to the acquired position information of the virtual viewpoint and the camera parameter data corresponding to the stitched image.
[0287] In a specific implementation, in a specific implementation, the first quantization parameter data acquisition unit 223 is suitable for obtaining the minimum depth distance from the optical center and the maximum depth distance from the optical center of the estimated depth map corresponding to the viewing angle; accordingly, the first depth map dequantization processing unit 224 is suitable for dequantizing the quantized depth value in the quantized depth map based on the minimum depth distance from the optical center and the maximum depth distance from the optical center of the estimated depth map corresponding to the viewing angle, using the corresponding dequantization formula to obtain the depth value of the pixel corresponding to the estimated depth map of the corresponding viewing angle.
[0288] In some embodiments of the present specification, the specific dequantization formula used by the first depth map dequantization processing unit 224 can refer to the aforementioned embodiment and will not be described again here.
[0289] Reference Fig.23 The embodiment of this specification also provides a free viewpoint video processing device, such as Fig.23 As shown, the free viewpoint video processing device 230 may include: a second video acquisition unit 231, a second quantized depth value acquisition unit 232, a second depth map inverse quantization processing unit 233, a virtual viewpoint position determination unit 234, and a second image reconstruction unit 235, wherein:
[0290] The second video acquisition unit 231 is adapted to acquire a free viewpoint video, wherein the free viewpoint video includes a spliced image at a plurality of frame moments and parameter data corresponding to the spliced image, wherein the spliced image includes a texture map of a plurality of synchronized viewing angles and a quantized depth map of a corresponding viewing angle, and the parameter data corresponding to the spliced image includes: quantized parameter data of an estimated depth map of a corresponding viewing angle and camera parameter data;
[0291] The second quantized depth value obtaining unit 232 is adapted to obtain the quantized depth value of the pixel in the quantized depth map;
[0292] The second depth map dequantization processing unit 233 is adapted to obtain and perform dequantization processing on the quantized depth values of the pixels in the quantized depth map based on the quantization parameter data of the estimated depth map corresponding to the viewing angle of the quantized depth map, so as to obtain a corresponding estimated depth map;
[0293] The virtual viewpoint position determination unit 234 is adapted to determine the position information of the virtual viewpoint in response to the user interaction behavior;
[0294] The second image reconstruction unit 235 is adapted to reconstruct the image of the virtual viewpoint based on the synchronized texture maps of multiple viewpoints and the estimated depth maps of the corresponding viewpoints, according to the position information of the virtual viewpoint and the camera parameter data corresponding to the stitched image.
[0295] The specific implementation of the free viewpoint video processing device in the embodiments of this specification can refer to the above-mentioned free viewpoint video processing method, which will not be repeated here.
[0296] This specification also provides an electronic device, referring to Fig.24 The structural schematic diagram of the electronic device shown, wherein the electronic device 240 may include a memory 241 and a processor 242, the memory 241 stores computer instructions that can be executed on the processor 242, wherein the processor 242 can execute the steps of the method described in any of the aforementioned embodiments when running the computer instructions, and the specific steps, principles, etc. can be found in the corresponding method embodiments mentioned above, which will not be repeated here.
[0297] In a specific implementation, the electronic device can be set up on the service side as a server or cloud device based on a specific solution, or can be set up on the user side as a terminal device.
[0298] The embodiment of this specification also provides a corresponding server device, referring to Fig.25 The structural diagram of the server device shown in FIG. 1 is a schematic diagram of the structure of the server device shown in FIG. 1 . In a specific implementation, as shown in FIG. 1 , Fig.25 As shown, the server device 250 may include a processor 251 and a communication component 252, wherein:
[0299] The processor 251 is adapted to execute the steps of the depth map processing method described in any of the foregoing embodiments to obtain a quantized depth map, stitch the texture maps of multiple synchronized perspectives and the first depth map of the corresponding perspective in a preset stitching manner to obtain a stitched image, and encapsulate the stitched images of multiple frames and the corresponding parameter data to obtain a free viewpoint video;
[0300] The communication component 252 is adapted to transmit the free viewpoint video.
[0301] The present specification also provides a terminal device, referring to Fig.26 The schematic diagram of the structure of the terminal device shown in FIG. Fig.26 As shown, the terminal device 260 may include a communication component 261, a processor 262 and a display component 263, wherein:
[0302] The communication component 261 is adapted to obtain free viewpoint video;
[0303] The processor 262 is suitable for executing the steps of the free viewpoint video reconstruction method or the free viewpoint video processing method described in any of the aforementioned embodiments. The specific steps can be found in the description of the aforementioned free viewpoint video reconstruction method and free viewpoint video processing method embodiments, which will not be repeated here.
[0304] The display component 263 is suitable for displaying the reconstructed image obtained by the processor 262 .
[0305] In the embodiments of this specification, the terminal device may be a mobile terminal such as a mobile phone, a tablet computer, a personal computer, a television, or a combination of any terminal device and an external display device.
[0306] The embodiments of the present specification also provide a computer-readable storage medium having computer instructions stored thereon, wherein the computer instructions, when executed, execute the steps of the free viewpoint video reconstruction method or the free viewpoint video processing method described in any of the aforementioned embodiments. For details, please refer to the aforementioned specific embodiments, which will not be repeated here.
[0307] In a specific implementation, the computer-readable storage medium may be any appropriate readable storage medium such as an optical disk, a mechanical hard disk, a solid-state hard disk, etc.
[0308] Although the embodiments of this specification are disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of this specification, so the protection scope of the present invention shall be subject to the scope defined by the claims.
Claims
1. A depth map processing method, wherein: include: Obtaining an estimated depth map generated based on a plurality of frame-synchronized texture maps, wherein the plurality of texture maps have different viewing angles; Obtaining a depth value of a pixel in the estimated depth map; Acquiring and quantizing the depth values of pixels in the estimated depth map based on the quantization parameter data corresponding to the estimated depth map viewing angle to obtain quantized depth values of corresponding pixels in the quantized depth map, including: Obtaining a minimum value of the depth distance from the optical center and a maximum value of the depth distance from the optical center corresponding to the viewing angle of the estimated depth map; Based on the minimum depth distance from the optical center and the maximum depth distance from the optical center of the viewing angle corresponding to the estimated depth map, a corresponding quantization formula is used to quantize the depth value of the corresponding pixel in the estimated depth map to obtain a quantized depth value of the corresponding pixel in the quantized depth map, including: The depth value of the corresponding pixel in the estimated depth map is quantized using the following quantization formula: Among them, M is the quantization bit of the pixel corresponding to the estimated depth map, range is the depth value of the corresponding pixel in the estimated depth map, Depth is the quantized depth value of the corresponding pixel in the estimated depth map, N is the viewing angle corresponding to the estimated depth map, depth_range_near_N is the minimum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N, and depth_range_far_N is the maximum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N.
2. The method according to claim 1, wherein: Also includes: Downsampling the quantized depth map to obtain a first depth map; The texture maps of the multiple perspectives of the frame synchronization and the first depth maps of the corresponding perspectives are spliced according to a preset splicing method to obtain a spliced image.
3. A free viewpoint video reconstruction method, wherein: include: Acquire a free viewpoint video, the free viewpoint video comprising a plurality of spliced images at frame moments and parameter data corresponding to the spliced images, the spliced images comprising texture maps of a plurality of synchronized viewing angles and a quantized depth map of a corresponding viewing angle, the parameter data corresponding to the spliced images comprising: quantized parameter data of an estimated depth map of a corresponding viewing angle and camera parameter data; Obtaining a quantized depth value of a pixel in the quantized depth map; Obtaining and, based on quantization parameter data of an estimated depth map corresponding to a viewing angle of the quantized depth map, performing dequantization processing on quantized depth values of pixels in the quantized depth map to obtain a corresponding estimated depth map; Based on the synchronized texture maps of multiple viewing angles and the estimated depth maps of the corresponding viewing angles, the image of the virtual viewpoint is reconstructed according to the position information of the virtual viewpoint and the camera parameter data corresponding to the stitched image; The acquiring and performing inverse quantization processing on the quantized depth values in the quantized depth map based on the quantized depth map's quantized parameter data of the estimated depth map corresponding to the viewing angle to obtain the estimated depth map corresponding to the viewing angle includes: Obtaining a minimum value of the depth distance from the optical center and a maximum value of the depth distance from the optical center corresponding to the viewing angle of the estimated depth map; Based on the minimum value of the depth distance from the optical center and the maximum value of the depth distance from the optical center of the estimated depth map corresponding to the viewing angle, the quantized depth value in the quantized depth map is dequantized using a corresponding dequantization formula to obtain a depth value of a pixel corresponding to the estimated depth map of the corresponding viewing angle, including: The quantized depth values in the quantized depth map are dequantized using the following dequantization formula to obtain corresponding pixel values in the estimated depth map: Among them, M is the quantization bit of the pixel corresponding to the quantized depth map, range is the depth value of the corresponding pixel in the estimated depth map, Depth is the quantized depth value of the corresponding pixel in the quantized depth map, N is the viewing angle corresponding to the estimated depth map, depth_range_near_N is the minimum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N, depth_range_far_N is the maximum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N, maxdisp is the maximum quantized depth distance corresponding to the viewing angle N, and mindisp is the minimum quantized depth distance corresponding to the viewing angle N.
4. The method according to claim 3, wherein: The resolution of the quantized depth map is smaller than the resolution of the texture map corresponding to the viewing angle; Before reconstructing the image of the virtual viewpoint, the method further includes: The estimated depth map corresponding to the viewing angle is upsampled to obtain a second depth map for reconstructing the image of the virtual viewpoint.
5. The method according to claim 4, wherein: The upsampling of the estimated depth map corresponding to the viewing angle to obtain a second depth map for reconstructing the virtual viewpoint image includes: Obtaining depth values of pixels in the estimated depth map as pixel values of corresponding even-numbered rows and even-numbered columns in the second depth map; For the depth values of the pixels in the even rows and odd columns of the second depth map, determining the corresponding pixel in the corresponding texture map as the middle pixel, based on the relationship between the brightness channel value of the middle pixel in the corresponding texture map and the brightness channel values of the left pixel and the right pixel corresponding to the middle pixel; For the depth values of odd-numbered rows of pixels in the second depth map, the corresponding pixels in the corresponding texture map are determined as intermediate pixels based on the relationship between the brightness channel value of the intermediate pixels in the corresponding texture map and the brightness channel values of the upper pixels and the brightness channel values of the lower pixels corresponding to the intermediate pixels.
6. The method according to claim 3, wherein: The method reconstructs the image of the virtual viewpoint based on the texture maps of the synchronized multiple viewpoints and the estimated depth maps of the corresponding viewpoints according to the acquired position information of the virtual viewpoint and the camera parameter data corresponding to the stitched image, including: According to the position information of the virtual viewpoint and the camera parameter data corresponding to the stitched image, a plurality of target texture maps and target depth maps are selected from the synchronized texture maps of the plurality of perspectives and the estimated depth maps of the corresponding perspectives; The target texture map and the target depth map are combined and rendered to obtain an image of the virtual viewpoint.
7. A free viewpoint video processing method, wherein: include: Acquire a free viewpoint video, the free viewpoint video comprising a plurality of spliced images at frame moments and parameter data corresponding to the spliced images, the spliced images comprising texture maps of a plurality of synchronized viewing angles and a quantized depth map of a corresponding viewing angle, the parameter data corresponding to the spliced images comprising: quantized parameter data of an estimated depth map of a corresponding viewing angle and camera parameter data; Obtaining a quantized depth value of a pixel in the quantized depth map; Acquiring and dequantizing the quantized depth values of pixels in the quantized depth map based on the quantized depth map's quantized parameter data of the estimated depth map corresponding to the viewing angle, to obtain a corresponding estimated depth map, including: Obtaining a minimum value of the depth distance from the optical center and a maximum value of the depth distance from the optical center corresponding to the viewing angle of the estimated depth map; Based on the minimum value of the depth distance from the optical center and the maximum value of the depth distance from the optical center of the estimated depth map corresponding to the viewing angle, the quantized depth value in the quantized depth map is dequantized using a corresponding dequantization formula to obtain a depth value of a pixel corresponding to the estimated depth map of the corresponding viewing angle, including: The quantized depth values in the quantized depth map are dequantized using the following dequantization formula to obtain corresponding pixel values in the estimated depth map: Wherein, M is the quantization bit of the pixel corresponding to the quantized depth map, range is the depth value of the corresponding pixel in the estimated depth map, Depth is the quantized depth value of the corresponding pixel in the quantized depth map, N is the viewing angle corresponding to the estimated depth map, depth_range_near_N is the minimum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N, depth_range_far_N is the maximum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N, maxdisp is the maximum quantized depth distance corresponding to the viewing angle N, and mindisp is the minimum quantized depth distance corresponding to the viewing angle N; In response to the user interaction behavior, determining the position information of the virtual viewpoint; Based on the synchronized texture maps of multiple viewing angles and the estimated depth maps of corresponding viewing angles, the image of the virtual viewpoint is reconstructed according to the position information of the virtual viewpoint and the camera parameter data corresponding to the stitched image.
8. The method according to claim 7, wherein: The determining the position information of the virtual viewpoint in response to the user interaction behavior includes: determining the corresponding virtual viewpoint path information in response to the user's gesture interaction operation; The method reconstructs the image of the virtual viewpoint based on the texture maps of the synchronized multiple viewpoints and the estimated depth maps of the corresponding viewpoints according to the acquired position information of the virtual viewpoint and the camera parameter data corresponding to the stitched image, including: According to the virtual viewpoint path information, a texture map in the spliced image at a corresponding frame moment and an estimated depth map of a corresponding viewing angle are selected as a target texture map and a target depth map; The target texture map and the target depth map are combined and rendered to obtain an image of the virtual viewpoint.
9. The method according to claim 7 or 8, wherein: Also includes: Acquire a virtual rendering target object in the image of the virtual viewpoint; Acquire a virtual information image generated based on augmented reality special effect input data of the virtual rendering target object; The virtual information image and the image of the virtual viewpoint are synthesized and displayed.
10. The method according to claim 9, wherein: The step of acquiring a virtual information image generated based on augmented reality special effect input data of the virtual rendering target object comprises: According to the position of the virtual rendering target object in the image of the virtual viewpoint obtained by three-dimensional calibration, a virtual information image matching the position of the virtual rendering target object is obtained.
11. The method according to claim 9, wherein: The step of acquiring a virtual rendering target object in the image of the virtual viewpoint includes: In response to the special effect generation interaction control instruction, a virtual rendering target object in the image of the virtual viewpoint is acquired.
12. A depth map processing device, wherein: include: An estimated depth map acquisition unit, adapted to acquire an estimated depth map generated based on a plurality of frame-synchronized texture maps, wherein the plurality of texture maps have different viewing angles; A depth value acquisition unit, adapted to acquire a depth value of a pixel in the depth map; a quantization parameter data acquisition unit, adapted to acquire quantization parameter data corresponding to the estimated depth map viewing angle; A quantization processing unit, adapted to quantize the depth values of pixels in the estimated depth map based on the quantization parameter data corresponding to the estimated depth map viewing angle to obtain quantized depth values of corresponding pixels in the quantized depth map, comprising: Obtaining a minimum value of the depth distance from the optical center and a maximum value of the depth distance from the optical center corresponding to the viewing angle of the estimated depth map; Based on the minimum depth distance from the optical center and the maximum depth distance from the optical center of the viewing angle corresponding to the estimated depth map, a corresponding quantization formula is used to quantize the depth value of the corresponding pixel in the estimated depth map to obtain a quantized depth value of the corresponding pixel in the quantized depth map, including: The depth value of the corresponding pixel in the estimated depth map is quantized using the following quantization formula: Among them, M is the quantization bit of the pixel corresponding to the estimated depth map, range is the depth value of the corresponding pixel in the estimated depth map, Depth is the quantized depth value of the corresponding pixel in the estimated depth map, N is the viewing angle corresponding to the estimated depth map, depth_range_near_N is the minimum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N, and depth_range_far_N is the maximum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N.
13. A free viewpoint video reconstruction device, wherein: include: A first video acquisition unit is adapted to acquire a free viewpoint video, wherein the free viewpoint video includes a stitched image at a plurality of frame moments and parameter data corresponding to the stitched image, wherein the stitched image includes a texture map of a plurality of synchronized viewing angles and a quantized depth map of a corresponding viewing angle, and the parameter data corresponding to the stitched image includes: quantized parameter data of an estimated depth map of a corresponding viewing angle and camera parameter data; A first quantized depth value acquisition unit, adapted to obtain a quantized depth value of a pixel in the quantized depth map; A first quantization parameter data acquisition unit, adapted to acquire quantization parameter data corresponding to the quantized depth map viewing angle; The first depth map dequantization processing unit is adapted to perform dequantization processing on the quantized depth map of the corresponding viewing angle based on the quantization parameter data corresponding to the viewing angle of the quantized depth map to obtain the corresponding estimated depth map, including: obtaining the minimum depth distance from the optical center and the maximum depth distance from the optical center of the viewing angle corresponding to the estimated depth map; based on the minimum depth distance from the optical center and the maximum depth distance from the optical center of the viewing angle corresponding to the estimated depth map, using the corresponding dequantization formula to perform dequantization processing on the quantized depth value in the quantized depth map to obtain the depth value of the corresponding pixel of the estimated depth map of the corresponding viewing angle, including: using the following dequantization formula to perform dequantization processing on the quantized depth value in the quantized depth map to obtain the corresponding pixel value in the estimated depth map: Wherein, M is the quantization bit of the pixel corresponding to the quantized depth map, range is the depth value of the corresponding pixel in the estimated depth map, Depth is the quantized depth value of the corresponding pixel in the quantized depth map, N is the viewing angle corresponding to the estimated depth map, depth_range_near_N is the minimum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N, depth_range_far_N is the maximum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N, maxdisp is the maximum quantized depth distance corresponding to the viewing angle N, and mindisp is the minimum quantized depth distance corresponding to the viewing angle N; The first image reconstruction unit is adapted to reconstruct an image of the virtual viewpoint based on texture maps of multiple viewpoints and estimated depth maps of corresponding viewpoints, according to the acquired position information of the virtual viewpoint and camera parameter data corresponding to the stitched image.
14. A free viewpoint video processing device, wherein: include: A second video acquisition unit is adapted to acquire a free viewpoint video, wherein the free viewpoint video includes a stitched image at a plurality of frame moments and parameter data corresponding to the stitched image, wherein the stitched image includes a texture map of a plurality of synchronized viewing angles and a quantized depth map of a corresponding viewing angle, and the parameter data corresponding to the stitched image includes: quantized parameter data of an estimated depth map of a corresponding viewing angle and camera parameter data; A second quantized depth value acquisition unit, adapted to acquire a quantized depth value of a pixel in the quantized depth map; The second depth map dequantization processing unit is adapted to obtain and perform dequantization processing on the quantized depth values of the pixels in the quantized depth map based on the quantization parameter data of the estimated depth map corresponding to the viewing angle of the quantized depth map to obtain the corresponding estimated depth map, including: Obtaining a minimum value of the depth distance from the optical center and a maximum value of the depth distance from the optical center corresponding to the viewing angle of the estimated depth map; Based on the minimum value of the depth distance from the optical center and the maximum value of the depth distance from the optical center of the estimated depth map corresponding to the viewing angle, the quantized depth value in the quantized depth map is dequantized using a corresponding dequantization formula to obtain a depth value of a pixel corresponding to the estimated depth map of the corresponding viewing angle, including: The quantized depth values in the quantized depth map are dequantized using the following dequantization formula to obtain corresponding pixel values in the estimated depth map: Wherein, M is the quantization bit of the pixel corresponding to the quantized depth map, range is the depth value of the corresponding pixel in the estimated depth map, Depth is the quantized depth value of the corresponding pixel in the quantized depth map, N is the viewing angle corresponding to the estimated depth map, depth_range_near_N is the minimum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N, depth_range_far_N is the maximum depth distance from the optical center in the estimated depth map corresponding to the viewing angle N, maxdisp is the maximum quantized depth distance corresponding to the viewing angle N, and mindisp is the minimum quantized depth distance corresponding to the viewing angle N; A virtual viewpoint position determination unit, adapted to determine position information of a virtual viewpoint in response to user interaction behavior; The second image reconstruction unit is adapted to reconstruct the image of the virtual viewpoint based on the synchronized texture maps of multiple viewpoints and the estimated depth maps of the corresponding viewpoints, according to the position information of the virtual viewpoint and the camera parameter data corresponding to the stitched image.
15. An electronic device comprising a memory and a processor, wherein the memory stores computer instructions executable on the processor, wherein: When the processor executes the computer instructions, the steps of the method of any one of claims 1 to 2, claims 3 to 6, or claims 7 to 11 are performed.
16. A server device, comprising a processor and a communication component, wherein: The processor is adapted to execute the steps of the method according to any one of claims 1 to 2, obtain a quantized depth map, splice the texture maps of multiple synchronized perspectives and the first depth map of the corresponding perspective in a preset splicing manner to obtain a spliced image, and encapsulate the spliced images of multiple frames and the corresponding parameter data to obtain a free viewpoint video; The communication component is suitable for transmitting the free viewpoint video.
17. A terminal device, comprising a communication component, a processor and a display component, wherein: The communication component is adapted to acquire free viewpoint video; The processor is adapted to perform the steps of the method according to any one of claims 3 to 6 or any one of claims 7 to 11; The display component is suitable for displaying the reconstructed image obtained by the processor.
18. A computer-readable storage medium having computer instructions stored thereon, wherein: When the computer instructions are executed, the steps of the method described in any one of claims 1 to 2, claims 3 to 6, or claims 7 to 11 are executed.
Citation Information
Patent Citations
Fast image drafting method based on depth drawing
CN101271583A
Method for generating depth maps from monocular images and systems using the same
CN102741879A