A depth map synthesis method for realizing high-resolution wide field angle virtual viewpoint generation
By converting the depth map from the side view to the main view reference frame and making it consistent with the color texture map, the problems of holes and reconstruction quality in the generation of virtual viewpoints with wide field of view are solved, and efficient and high-quality virtual viewpoint generation is achieved.
Patent Information
- Application Number
- CN202310272311.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-20
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2043-03-20
AI Technical Summary
Existing virtual viewpoint generation methods suffer from reconstruction quality issues, such as holes, artifacts, cracks, and overlaps, under wide field of view, and are computationally complex, making it difficult to efficiently generate high-quality, high-resolution images.
By collecting image information from several angles of the target scene, the depth map from the side view is transformed into the reference frame of the main view, so that the depth sub-maps at different positions have the same structural depth information and are consistent with the color texture map. The virtual viewpoint generation algorithm is then used to fill in the holes in the image.
It achieves efficient generation of high-quality, high-resolution images, reduces padding errors, improves computational efficiency, expands the observation range, and maintains image quality in real-time computation.
Smart Images

Figure CN116704114B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to a depth map synthesis method for realizing high-resolution wide-view-angle virtual viewpoint generation, and belongs to the technical field of three-dimensional reconstruction. BACKGROUND
[0002] Visual information accounts for 80% of daily information inflow and is the main means for humans to obtain external information. In daily life, the collection, transmission and imaging unit of visual information collects a three-dimensional scene into two-dimensional information, which is quite different from human familiar stereoscopic vision in information capacity and quality. The binocular structure of the human eye provides a stereoscopic vision system for humans to measure depth, estimate position and reconstruct stereoscopic vision, which is the result of adaptive inheritance in the evolution process of primates and carnivorous mammals to adapt to the needs of collection and hunting. There are many ways for humans to perceive stereoscopic vision, which can be roughly classified into psychological factors and physiological factors. The stereoscopic vision formed by seven psychological factors such as spatial perspective, light and shade relationship and occlusion relationship is closely related to the learning experience of the observer and is difficult to realize stereoscopic imaging through a unified method. In the stereoscopic vision formed by physiological factors such as binocular disparity, binocular convergence, eye accommodation, monocular motion parallax and color difference, binocular disparity is considered to be the most critical factor in producing stereoscopic vision. There are mainly two ways to realize stereoscopic vision through binocular disparity, namely double-viewpoint image generation and multi-viewpoint image generation. Double-viewpoint image generation produces a stereoscopic effect by presenting two fixed disparity images to the left and right eyes, which is a kind of “pseudo three-dimensional” effect; multi-viewpoint image is obtained by synthesizing a sequence of disparity images, which can provide continuous disparity images to the human eye to form stereoscopic vision. In the process of forming multi-view images, the light field encoding of multiple images collected under the target scene is needed to complete the three-dimensional reconstruction. Direct use of multiple images for light field encoding requires a large amount of storage and computing resources, which greatly limits real-time transmission and reconstruction. Using virtual viewpoint generation algorithm to reconstruct the light field can significantly save computing cost and improve computing efficiency. The virtual reconstruction algorithm needs to use the depth information of the target scene to guide the reconstruction in the process of reconstructing the scene, but as the field of view angle increases, the depth information of different structures cannot be easily recovered. Although the use of iterative updating can reduce part of the calculation error, there are still obvious problems such as holes, artifacts, cracks and overlaps in the rendering process, which have a negative impact on visual experience. In order to effectively solve the problem of reconstruction quality caused by the difference in scene structure depth under a wide field of view angle, a new depth map format is proposed. This depth map format twists the depth information obtained at different angles to a certain reference depth according to the scene structure, which can realize virtual viewpoint generation and three-dimensional hole filling under a wide field of view angle. SUMMARY
[0003] The application aims to provide a depth map synthesis method for realizing high-resolution wide field angle virtual viewpoint generation, which collects image information of a target scene at several angles, selects a certain position as a main viewing angle and converts depth maps at the remaining positions to a reference system of the main viewing angle, so that the depth sub-maps at different positions have the same structural depth information and correspond to the color sub-maps in structural outline. Figure One The changed depth sub-maps can correspond to the corresponding color texture Figure One maps. The changed depth maps can then fill the holes of the high-resolution scene at any observation direction within a wide field angle range, which can expand the observation range and ensure the quality of the filled scene.
[0004] In order to achieve the above-mentioned purpose, the technical scheme of the application is as follows: a depth map synthesis method for realizing high-resolution wide field angle virtual viewpoint generation, the method comprising the following steps:
[0005] Step 1) obtaining image information of a target scene at several angles;
[0006] Step 2) taking a certain angle of capturing image information as a main viewing angle and other angles as side viewing angles, converting the original depth maps at the side viewing angles to a depth reference system of the main viewing angle as side viewing depth sub-maps, so that the side viewing depth sub-maps are the same as the main viewing depth sub-maps in structural depth and the same as the side viewing color sub-maps in structural outline;
[0007] Step 3) using the several pairs of completed images to reconstruct the scene by a virtual viewpoint generation algorithm and realizing the completion of image holes in virtual viewpoint generation relying on the new depth maps.
[0008] As an improvement of the application, the step 1) obtaining image information of a target scene at several angles, the camera array form for capturing scene image information can be parallel cameras, the main optical axes of each camera being perpendicular to the plane of the scene; or can be converging cameras, the main optical axes of each camera converging at a common point in the scene; or can be skew cameras, the main optical axes of each camera converging at a common point in the scene and the display plane showing the same range by calculation. Or other camera arrays, the selection of the camera array can be arbitrary.
[0009] The step 1) obtaining image information of a target scene at several angles, the camera model for capturing scene image information can be a perspective camera model; can be an orthogonal camera model; can be a weak perspective camera model; or can be other camera models. The selection of the camera model can be arbitrary.
[0010] As an improvement of the present application, the step 1) acquires image information of the target scene at several angles, and the number of images acquired can be arbitrary under the condition that the selected angles do not coincide.
[0011] As an improvement of the present application, in the step 2), the angle at which the image information is captured is the main view angle, and the other angles are side view angles. The image information at the side view angles can be directly captured by the camera, or can be generated by an algorithm from the main view texture map, or can be generated by a deep learning model according to the captured image information. The acquisition method of the side view image information can be arbitrary.
[0012] As an improvement of the present application, in the step 2), the original depth map at the side view angle is converted to the reference system at the main view angle as a side view depth sub-map. The original depth map used here can be directly captured by an RGBD camera, or can be generated by an algorithm from the texture map, or can be generated by a deep learning model according to the captured image information. The acquisition method of the original depth map can be arbitrary.
[0013] As an improvement of the present application, in the step 2), the original depth map at the side view angle is converted to the depth reference system at the main view angle as a side view depth sub-map. The conversion to the main view reference system can be calculated using an analytical method, i.e. by changing the rotation angle and pixel position, or can be calculated using a deep learning method, i.e. by learning through a deep learning network to obtain the target depth map. The calculation method is arbitrary.
[0014] As an improvement of the present application, in the step 2), the original depth map at the side view angle is converted to the depth reference system at the main view angle as a side view depth sub-map. The depth information in the depth sub-map is used to index the pixels at the same spatial position in the two side view images, which are used to fill the missing pixels. The calculation formula can be written as:
[0015] D sMax =D m +(D m -D zero )tan(α max ) (1)
[0016] wherein, represents the depth information at the side view angle generated by conversion, hereinafter referred to as side view depth; represents the main view depth sub-map image information, hereinafter referred to as main view depth; D zero represents the depth information of the zero parallax plane; α max represents the deflection of the side view angle relative to the main view angle.
[0017] As an improvement of the present application, the side-view depth subgraph in step 2) is the same as the front-view depth subgraph in structure depth and the same as the side-view color subgraph in structure contour, the generated side-view depth subgraph and the side-view color subgraph should be one-to-one correspondence, and the partial pixel content in the generated side-view depth subgraph needs to be completed according to the relative structure information of the original side-view depth subgraph or the texture relationship of the side-view color subgraph. For the generated depth information, the completion formula can be written as:
[0018]
[0019] Wherein represents the completed generated side-view depth information, hereinafter referred to as the completed side-view depth; d m represents the gradient information of the same structure at the generated side-view depth hole position corresponding to the front-view depth subgraph; the completed side-view depth can also be obtained by the gradient of the texture correlation; it can be obtained by a depth learning model; the completion calculation method is arbitrary.
[0020] As an improvement of the present application, in step 3), the transmission format of the image can be arbitrary, and multiple sets of depth images can be placed in different electrical frequencies of the same channel to save transmission resources, or multiple sets of depth images and their corresponding RGB images can be placed adjacent to speed up the calculation. The transmission format of the image information can be arbitrary.
[0021] As an improvement of the present application, in step 3), the virtual viewpoint generation algorithm is used to reconstruct the scene by using the completed several pairs of images, and the image hole completion in the virtual viewpoint generation is realized relying on the new depth map, and the calculation formula can be written as:
[0022] P V1 =P m +(D m -D zero )tan(α) (3)
[0023] P V2 =P s -(D sf -D zero )[tan(α max )-tan(α)] (4)
[0024] P V =(1-β)P V1 +βP V2 (5)
[0025] Wherein represents the color image information generated at any generated viewing angle by the front-view color subgraph, hereinafter referred to as the front-view generated information; represents color image information at any generated view angle generated by the side-view color subgraph, hereinafter referred to as side-view generated information; represents main-view color subgraph information, hereinafter referred to as main-view color information; represents side-view color subgraph information, hereinafter referred to as side-view color information; α represents a distortion parameter obtained from a relative main-view subgraph angle position; represents color image information at any generated view angle generated by fusion, hereinafter referred to as generated information; β represents a fusion coefficient.
[0026] Beneficial effects: 1. The depth image synthesis method first proposed in the application can efficiently generate high-quality high-resolution images when applied to virtual viewpoint generation. The structural depth of images in the traditional virtual viewpoint generation method is only related to the image capture position, which causes the same color texture information in images captured at different positions to correspond to different structural depths, resulting in many defects in the generated images and a complex image texture information matching process. By unifying the depth image information at different collection positions to the same reference system, the application can efficiently complete high-quality virtual viewpoint generation.
[0027] 2. The depth image synthesis method first proposed in the application can efficiently complete the hole filling problem caused by view angle deflection and generate high-quality and credible image pixels when applied to virtual viewpoint generation, thereby reducing the filling error caused by the mismatch of structural depth information in traditional images and avoiding the complex and tedious steps caused by the use of an iterative method.
[0028] 3. The depth image synthesis method first proposed in the application can efficiently complete image feature registration, image information fusion, and new viewpoint synthesis when applied to virtual viewpoint generation, which is different from the complex calculation of the traditional method using feature value calculation, polar line correction, and triangulation. By unifying the input images to the same reference system, the features between multiple input images only differ by an offset related to the camera deflection angle. The calculation complexity of the offset is a constant order of magnitude, which is significantly improved compared to the calculation of the traditional method with an image size of an order of magnitude.
[0029] 4. The depth image synthesis method first proposed in the application can make color texture information available for image hole filling at a larger field of view angle by aligning multiple color texture images under the same depth reference, so that the observation range of the scene is no longer limited to a very small observation angle. At the same time, because the input image information is sparse, this generation method can perform real-time calculation and observation more quickly and efficiently on a rendering device. BRIEF DESCRIPTION OF DRAWINGS
[0030] Figure 1 is a schematic diagram of the convergent camera array described in step 1 of the application;
[0031] Figure 2 is the other camera array schematic diagram described in step 1 of the present application;
[0032] Figure 3 is the schematic diagram of the original depth map unified to the main view angle depth reference system described in step 2 of the present application;
[0033] Figure 4 is the side view depth subgraph structure depth and structure contour schematic diagram described in step 2 of the present application;
[0034] Figure 5 is the schematic diagram of the side view depth subgraph completion calculation process described in step 2 of the present application;
[0035] Figure 6 is a schematic diagram of one arrangement of input image information in step 3 of the present application;
[0036] Figure 7 is the schematic diagram of completing virtual viewpoint generation and image hole completion described in step 3 of the present application. DETAILED DESCRIPTION
[0037] The application will be further illustrated below in conjunction with specific embodiments, which should be understood as merely illustrating the present application and not limiting the scope of the present application. After reading the present application, those skilled in the art will make various equivalent transformations of the present application, which all fall within the scope defined by the claims attached hereto.
[0038] Embodiment 1: see Figures 1-7 A depth map synthesis method for realizing high-resolution wide field angle virtual viewpoint generation, the method comprising the following steps:
[0039] Step 1) obtaining image information at several angles of a target scene;
[0040] Step 2) taking an angle at which image information is captured as a main view angle, and other angles as side view angles, converting the original depth map at the side view angles to the depth reference system at the main view angle, as a side view depth subgraph, so that the side view depth subgraph is the same as the main view depth subgraph in structural depth and the same as the side view color subgraph in structural contour;
[0041] Step 3) reconstructing the scene by a virtual viewpoint generation algorithm using the completed several pairs of images, and relying on the new depth map to complete the image hole in virtual viewpoint generation.
[0042] The step 1) obtains image information of the target scene at several angles, and the camera array arranged to capture the image information of the scene can be parallel cameras, the main optical axis of each camera being perpendicular to the plane of the scene; or can be converging cameras, the main optical axis of each camera converging at a common point in the scene; or can be skew cameras, the main optical axis of each camera converging at a common point in the scene and the imaging plane being calculated to display the same range. Or other camera arrays, the selection of the camera array can be arbitrary.
[0043] The step 1) obtains image information of the target scene at several angles, and the camera model for capturing the image information of the scene can be a perspective camera model; can be an orthogonal camera model; can be a weak perspective camera model; or can be other camera models. The selection of the camera model can be arbitrary.
[0044] The step 1) obtains image information of the target scene at several angles, and the number of images obtained can be arbitrary under the condition that the selected angles do not coincide.
[0045] In the step 2), the angle at which the image information is captured is the main viewing angle, and the image information at the side viewing angle can be directly captured by the camera; or can be generated by the texture map at the main viewing angle through an algorithm; or can be generated by a deep learning model according to the captured image information. The acquisition method of the image information at the side viewing angle can be arbitrary.
[0046] In the step 2), the original depth map at the side viewing angle is converted to the reference system at the main viewing angle, serving as a side viewing depth sub-map. The original depth map used here can be directly captured by an RGBD camera; or can be generated by a texture map through an algorithm; or can be generated by a deep learning model according to the captured image information. The acquisition method of the original depth map can be arbitrary.
[0047] In the step 2), the original depth map at the side viewing angle is converted to the depth reference system at the main viewing angle, serving as a side viewing depth sub-map. The conversion to the main viewing angle reference system can be calculated using an analytical method, i.e., by changing the rotation angle and pixel position; or can be calculated using a deep learning method, i.e., by learning through a deep learning network to obtain the target depth map. The calculation method can be arbitrary.
[0048] In the step 2), the original depth map at the side viewing angle is converted to the depth reference system at the main viewing angle, serving as a side viewing depth sub-map. The depth information in the depth sub-map is used to index the pixels at the same spatial position in the two side views, serving as the filling of the missing pixels, and the calculation formula can be written as:
[0049] D sMax =Dm +(D m -D zero )tan(α max ) (1)
[0050] wherein, denotes the converted depth information at the side view angle, hereinafter referred to as side view depth, wherein d sx , d sy are the depth values of the original side view depth converted to the x-axis and y-axis of the main view coordinate system, respectively; denotes the main view depth sub-image information, hereinafter referred to as main view depth, wherein d mx , d my are the depth values of the x-axis and y-axis of the main view depth map, respectively; D zero denotes the depth information where the zero parallax plane is located; a max denotes the deflection of the side view angle relative to the main view angle.
[0051] The side view depth sub-image in the step 2) is the same as the main view depth sub-image in the structural depth and the same as the side view color sub-image in the structural contour. The generated side view depth sub-image and the side view color sub-image should be one-to-one correspondence. If the partial pixel content in the generated side view depth sub-image needs to be completed according to the relative structural information of the original side view depth sub-image or the texture relationship of the side view color sub-image, for the distorted generated depth information, the completion formula can be written as:
[0052]
[0053] wherein denotes the completed generated side view depth information, hereinafter referred to as completed side view depth, wherein d sfx , d sfy are the depth values of the x-axis and y-axis of the completed side view depth D sMax in the main view coordinate system, respectively; d m denotes the gradient information of the same structure of the main view depth sub-image at the generated side view depth hole position; the completed side view depth can also be obtained by the texture related gradient; can be obtained by a deep learning model; the completion calculation mode is arbitrary.
[0054] In the step 3), a plurality of images after completion are used, and the transmission format of the image can be arbitrary. A plurality of depth images can be placed in different electrical frequencies of the same channel to save transmission resources, or a plurality of depth images and their corresponding RGB images can be placed adjacent to each other to accelerate calculation. The transmission format of the image information can be arbitrary.
[0055] As an improvement of the present application, the step 3) reconstructs the scene by using the several pairs of images after the completion through the virtual viewpoint generation algorithm, and relies on the new depth map to realize the completion of the image hole in the virtual viewpoint generation, and the calculation formula can be written as:
[0056] P V1 m +(D m -D zero )tan(α) (3)
[0057] P V2 s -(D sf -D zero )[tan(α max )-tan(α)] (4)
[0058] P V V1 +βP V2 (5)
[0059] Wherein represents the color image information at any generated viewing angle generated by the main view color subgraph, which is called the main view generated information below, wherein x v1 , y v1 respectively represent the coordinate information of the main view generated information on the x-axis and y-axis in the main view coordinate system; represents the color image information at any generated viewing angle generated by the side view color subgraph, which is called the side view generated information below, wherein x v2 , y v2 respectively represent the coordinate information of the side view generated information on the x-axis and y-axis in the main view coordinate system; represents the main view color subgraph information, which is called the main view color information below, wherein x m , y m respectively represent the coordinate information of the main view color information on the x-axis and y-axis in the main view coordinate system; represents the side view color subgraph information, which is called the side view color information below, wherein x s , y s respectively represent the coordinate information of the side view color information on the x-axis and y-axis in the side view coordinate system; α represents the distortion parameter obtained from the relative main view subgraph angle position; represents the color image information at any generated viewing angle generated by the fusion, which is called the generated information below, wherein x v , y v respectively represent the coordinate information of the generated information on the x-axis and y-axis in the main view coordinate system; β represents the fusion coefficient.
[0060] Embodiment 2: asFigure 1 As shown, the depth map synthesis method for realizing high-resolution wide field angle virtual viewpoint generation disclosed by the embodiment of the present application mainly comprises the following steps:
[0061] Step 1) setting a camera model and a camera array
[0062] Setting different camera models and different camera arrays for a target scene will bring different influences on the image information collected by the camera. Here, we choose a relatively convenient combination, that is, using three orthogonal cameras and setting them as a parallel converging camera array.
[0063] The specific method is to use three RGBD orthogonal cameras with consistent internal parameters and capable of acquiring depth information, to set a main view RGBD camera with the scene center as the origin O and the scene normal as the 0° basis, and to set two left and right side view RGBD cameras with consistent internal parameters at ±β angles, and the three cameras intersect at the origin O along the main optical axis.
[0064] This way can simplify the calculation and imaging process. If not set in this way, other camera arrays can also be set. Only the captured subgraphs need to be transformed in the preprocessing stage, and the images can be scaled according to the distance ratio to ensure that the projection of each pixel of the image to the scene is the same or similar scale.
[0065] Step 2) acquiring image information of the target scene.
[0066] Obtaining image information of the target scene from the three RGBD cameras, that is, acquiring three pairs of color subgraphs and depth subgraphs of the target scene, which are left and right side view subgraph pairs composed of color subgraphs and depth subgraphs captured by RGBD cameras with the maximum imaging angle as the main optical axis, denoted as left view color subgraph L, left view depth subgraph D L , right view color subgraph R, right view depth subgraph D R ; and main view subgraph pairs composed of color subgraphs and depth subgraphs captured at the angle bisector of the maximum angle as the scene normal, denoted as main view color subgraph M and main view depth subgraph D M .
[0067] Step 3) twisting the main view depth subgraph according to the twist parameter to change it into the twisted maximum left and right depth subgraph.
[0068] In order to simplify the calculation, orthogonal cameras are used in step 1) to capture images in a converging array arrangement, which avoids the perspective distortion caused by using perspective cameras and reduces the calculation complexity. If other camera models and camera array combinations are used, attention should be paid to the correction of the images to make the left and right side view subgraph pairs similar to the main view subgraph pairs in texture contour.
[0069] Let The maximum observation view depth information generated by the warping can be represented as:
[0070] D sMax = D m + (D m - D zero ) tan (a max ) (1)
[0071] The main view depth information is represented. Specifically, d mx and d mx are the depth information on the x-axis and y-axis at the main view reference position, respectively.
[0072] D zero represents the depth at the depth baseline position, which is set at the depth at which the main optical axes of the cameras converge, and can be adjusted according to the differences in the scene. This variable super parameter affects the in-and-out screen positions of the images but does not affect the image reconstruction.
[0073] a max represents the maximum deflection angle. If the position of the selected main view camera is not on the angle bisector of the observation view, the maximum deflection angle needs to be adjusted, which can be represented as:
[0074]
[0075] where Angle represents the maximum imaging angle, a re represents the angle at which the main view camera is located.
[0076] Considering that the main view image using an orthogonal camera cannot be completely aligned with the left and right side view subgraphs when the warping is changed, an edge matching needs to be completed. Here, a simple position difference can be used as the matching standard.
[0077] Step 4) Complete the depth subgraph in step 3).
[0078] The maximum observation view depth information generated by the warping will have a large number of holes and structural losses due to the change in angle. The background holes are easy to repair because the background depth can be considered consistent or smooth. Common image repair algorithms or simple interpolation can be used to achieve this. However, the repair of structural scene loss is difficult. Here, a simple calculation method can be used to complete the repair using the structural gradient of the main view depth subgraph.
[0079] Let represent the completed maximum observation view depth information generated by the warping, which can be represented as:
[0080]
[0081] where d m represents the gradient information at the same structure of the main depth submap corresponding to the maximum observed depth hole position of the twist generated. The depth gradient change can be obtained, where d x , d y are the depth of the main depth submap at the x-axis direction corresponding to the depth hole and the depth of the main depth submap at the x-axis direction corresponding to the depth hole, respectively. Here, the calculation method of the gradient is arbitrary, as long as it can describe the change trend of the generated depth map at the same structure of the corresponding main depth submap.
[0082] Step 5) Complete the virtual generation and image hole completion using the depth submap completed in step 4).
[0083] Let represent the color image information at any generated viewing angle generated by the main color submap, which can be represented as:
[0084] P V1 = P m + (D m - D zero ) tan (a) (4)
[0085] Let represent the color image information at any generated viewing angle generated by the calibrated side color submap, which can be represented as:
[0086] P V2 = P s - (D sf - D zero ) [tan (a max ) - tan (a)] (5)
[0087] P V1 and P V2 Two sets of images are generated from two different image sources, which can greatly expand the image information, so that the holes caused by the change of viewing angle can be filled in the reconstruction process, because the final generated image is consistent, here we choose to use the weighted fusion method to simplify the calculation process.
[0088] Let represent the color image information at any generated viewing angle, which can be represented as:
[0089] P V = (1-β) P V1 + β P V2 (6)
[0090] represents the main reference color submap image information, a represents the relative angle of the target generated position, and β represents the fusion coefficient. By presenting the viewing angle range, β can be initialized as The fusion manner herein is arbitrary, and other manners such as iteration, voting, optimization, and the like can be selected for fusion calculation.
[0091] The above merely describes the preferred embodiments of the present application, and it should be noted that, for those skilled in the art, some improvements and refinements can be made without departing from the principles of the present application, and these improvements and refinements should also be considered as falling within the protection scope of the present application.
Claims
1. A depth map compositing method for implementing high-resolution wide field-of-view virtual viewpoint generation, characterized by, The method comprises the following steps: Step 1) obtaining image information of a target scene from several angles; Step 2) taking an angle at which image information is captured as a main view angle, and other angles as side view angles, converting original depth maps at the side view angles to a depth reference system at the main view angle to obtain side view depth sub-maps, so that the side view depth sub-maps are the same as the main view depth sub-maps in structural depth and the same as side view color sub-maps in structural outlines; Step 3) reconstructing the scene by using the several pairs of images after completion through a virtual view point generation algorithm, and completing image holes in virtual view point generation by relying on new depth maps; In step 2), the side view depth sub-maps are the same as the main view depth sub-maps in structural depth and the same as the side view color sub-maps in structural outlines, the generated side view depth sub-maps correspond to the side view color sub-maps one by one, and part of pixel content in the generated side view depth sub-maps is completed according to relative structural information of the original side view depth sub-maps or texture relationship of the side view color sub-maps; for distorted generated depth information, a completion formula is written as: wherein represents the completed generated side-view depth information, hereinafter referred to as the completed side-view depth, wherein d sfx , d sfy respectively represent the completed side-view depth D sMax depth values on the x-axis and y-axis in the main-view coordinate system; d m represents the gradient information at the same structure of the generated measured-view depth hole position corresponding to the main-view depth subgraph; the completed side-view depth is obtained through the main-view depth gradient; or obtained through the corresponding texture graph gradient; or obtained through a deep learning model; under the condition of ensuring that the completed side-view depth is the same as the main-view depth subgraph in the structural depth and the same as the side-view color subgraph in the structural contour, this completion mode is arbitrary; In step 3), the transmission format of the images is arbitrary, and multiple sets of depth images are stored in different channels of the same space to save transmission resources; or the multiple sets of depth images and their corresponding RGB images are placed adjacently to accelerate calculation; In step 3), the scene is reconstructed by using the several pairs of images after completion through the virtual view point generation algorithm, and image holes in virtual view point generation are completed by relying on the synthesized depth maps, and a calculation formula is written as: P V1 = P m + (D m - D zero ) tan(a) (3) P V2 = P s - D sf - D zero ) [tan(a max ) - tan(a)] (4) P V = (1 - β)P V1 + βP V2 (5) wherein represents color image information at an arbitrary generated view angle generated from the main view color subgraph, hereinafter referred to as main view generated information, wherein x v1 , y v1 represent coordinate information of the main view generated information on the x-axis and y-axis in the main view coordinate system, respectively; represents color image information at an arbitrary generated view angle generated from the side view color subgraph, hereinafter referred to as side view generated information, wherein x v2 , y v2 represent coordinate information of the side view generated information on the x-axis and y-axis in the main view coordinate system, respectively; represents main view color subgraph information, hereinafter referred to as main view color information, wherein x m , y m represent coordinate information of the main view color information on the x-axis and y-axis in the main view coordinate system, respectively; represents side view color subgraph information, hereinafter referred to as side view color information, wherein x s , y s represent coordinate information of the side view color information on the x-axis and y-axis in the side view coordinate system, respectively; a represents a warping parameter obtained from the relative main view subgraph angle position; represents color image information at an arbitrary generated view angle generated from the fusion, hereinafter referred to as generated information, wherein x v , y v represent coordinate information of the generated information on the x-axis and y-axis in the main view coordinate system, respectively; b represents a fusion coefficient.
2. The depth map synthesis method for generating a high-resolution wide field of view virtual viewpoint according to claim 1, wherein: In step 1), the camera array form for capturing scene image information is parallel cameras, the main optical axes of the cameras are perpendicular to the plane formed by the scene; or converging cameras, the main optical axes of the cameras converge at a common point in the scene; or skew cameras, the main optical axes of the cameras converge at a common point in the scene and the imaging planes display the same range through calculation; In step 1), the camera model for capturing scene image information is a view frustum camera model; or an orthogonal camera model; or a weak perspective camera model.
3. The depth map synthesis method for generating a high-resolution wide field of view virtual viewpoint according to claim 1, wherein: In step 1), the number of images is arbitrary under the condition that the selected angles do not coincide.
4. The depth map synthesis method for generating a high-resolution wide field of view virtual viewpoint according to claim 1, wherein: In step 2), the image information at the side view angles is directly captured by the cameras; or the image information at the side view angles is generated by the texture map at the main view angle through an algorithm; or the image information at the side view angles is generated by a deep learning model according to the captured image information; under the condition that the image information is derived from the same scene, the way of obtaining the image information at the side view angles is arbitrary.
5. The depth map synthesis method for generating a high-resolution wide field of view virtual viewpoint according to claim 1, wherein: The step 2) converts the original depth map under the side view angle to the reference system where the main view angle is located as a side view depth sub-map, wherein the original depth map used herein is directly captured by an RGBD camera; or the original depth map is generated by an algorithm from a texture map; or the original depth map is generated by a deep learning model according to captured image information; and the manner of obtaining the original side view depth map is arbitrary under the condition that the side view texture map is matched with the original side view depth map.
6. The depth map synthesis method for generating a high-resolution wide field of view virtual viewpoint according to claim 1, wherein: The step 2) converts the original depth map under the side view angle to the depth reference system where the main view angle is located as a side view depth sub-map, and the conversion to the main view angle reference system is calculated by using an analytical method, that is, the rotation angle and the pixel position are changed; or the conversion to the main view angle reference system is calculated by using a deep learning method, that is, the target depth map is obtained by learning through a deep learning network.
7. The depth map synthesis method for generating a high-resolution wide field of view virtual viewpoint according to claim 6, wherein, In the step 2), the original depth map under the side view angle is converted to the depth reference system where the main view angle is located, and the formula of the analytical calculation can be written as: D sMax = D m + (D m - D zero ) tan(a max ) (1) wherein, denotes the converted depth information at the side view angle of generation, hereinafter referred to as side view depth, wherein d sx , d sy are the depth values on the x-axis and y-axis of the original side view depth converted to the main view coordinate system, respectively; denotes the main view depth sub-image information, hereinafter referred to as main view depth, wherein d mx , d my are the depth values on the x-axis and y-axis of the main view depth, respectively; D zero denotes the depth information where the zero parallax plane is located; a max denotes the deflection of the side view angle relative to the main view angle.
Citation Information
Patent Citations
Virtual viewpoint depth map processing method, device and apparatus, and storage medium
CN112581389A
Method for realizing high-resolution ultra-wide field angle three-dimensional display image format
CN114067050A