Viewpoint rendering method and viewpoint rendering device
Patent Information
- Application Number
- CN202380010124.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-08-14
AI Technical Summary
The prior art is difficult to effectively convert two-dimensional images into three-dimensional images, especially when it does not rely on high-precision depth images, and the manual post-production cost is high and the efficiency is low.
By performing initial feature extraction, channel dimension stitching and dimensionality reduction on two-dimensional images and depth images, multiple image distortions and repairs, multiple viewpoint images are finally generated.
It realizes the generation of high-quality viewpoint images without relying on high-precision depth images, which reduces the artifact problems caused by image repair algorithms and improves the speed and efficiency of viewpoint images generation.
Smart Images

Figure CN119998845A_ABST
Abstract
Description
Viewpoint rendering method and viewpoint rendering device Technical Field
[0001] The present disclosure belongs to the field of display technology, and particularly relates to a viewpoint rendering method and a viewpoint rendering device. Background Art
[0002] As user needs become more diverse and standardized, ultra-high-definition two-dimensional images are increasingly unable to meet user viewing needs. The sense of depth and visual impact provided by three-dimensional images have become a new pursuit for users. Consequently, glasses-free 3D display technology has become a hot topic in the display industry. With the introduction of glasses-free 3D display products, there is widespread market demand for consumer-grade 3D content.
[0003] Currently, there are two sources of 3D content. One is to directly capture native 3D images and videos using binocular cameras. However, binocular cameras are not widely available and are expensive. They also render existing 2D images useless, requiring new 3D content to be captured, making them unsuitable for widespread user demand. The other method involves converting existing 2D images and videos into 3D. However, this is currently mostly done manually using specialized software for post-production, primarily for cinematic productions. Manual post-production is costly, time-consuming, and labor-intensive, making it unsuitable for widespread user demand.
[0004] Summary of the Invention
[0005] The present disclosure aims to solve at least one of the technical problems existing in the prior art and provides a viewpoint rendering method and a viewpoint rendering device.
[0006] In a first aspect, an embodiment of the present disclosure provides a viewpoint rendering method, wherein the viewpoint rendering method includes:
[0007] Performing initial feature extraction on a two-dimensional image and a depth image corresponding to the two-dimensional image to obtain a first initial feature of the two-dimensional image and a second initial feature of the depth image;
[0008] The first initial feature and the second initial feature are concatenated in the channel dimension, and the channel dimension is reduced to obtain a first initial reduced-dimensional feature;
[0009] Performing multiple image distortion and repair on the first initial dimensionality reduction features to obtain fusion features;
[0010] The fusion features are fused and channel dimension reduction is performed to generate multiple viewpoint images.
[0011] Optionally, the step of fusing the fused features and reducing the channel dimension to generate multiple viewpoint images further includes:
[0012] Performing initial feature extraction on the viewpoint image and the corresponding label image to obtain a third initial feature and a fourth initial feature;
[0013] Downsampling the width and height dimensions of the third initial feature and the fourth initial feature, reducing the width and height dimensions to 1, and reducing the channel dimension to 1, to obtain a second initial reduced dimensionality feature and a third initial reduced dimensionality feature;
[0014] The second initial dimensionality reduction feature and the third initial dimensionality reduction feature are discriminated according to the label image.
[0015] Optionally, the step of fusing the fused features and reducing the channel dimension to generate multiple viewpoint images further includes:
[0016] Performing initial feature extraction on the viewpoint image and the corresponding label image to obtain a fifth initial feature and a sixth initial feature;
[0017] Downsampling the fifth initial feature and the sixth initial feature in terms of width and height, and reducing the channel dimension to 1, to obtain a fourth initial reduced dimensionality feature and a fifth initial reduced dimensionality feature;
[0018] The fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature are discriminated according to the label image.
[0019] Optionally, the fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature are multiple groups;
[0020] Each group of the fourth initial dimensionality reduction features and the fifth initial dimensionality reduction features corresponds to the discriminated pixel blocks in descending order.
[0021] Optionally, the performing initial feature extraction on the two-dimensional image and the depth image corresponding to the two-dimensional image to obtain a first initial feature of the two-dimensional image and a second initial feature of the depth image further includes:
[0022] The two-dimensional image is input into a monocular depth estimation model to generate the depth image corresponding to the two-dimensional image.
[0023] Optionally, the fusion features are fused and channel dimension reduction is performed to generate multiple viewpoint images, and then the method further includes:
[0024] splicing the plurality of viewpoint images in a wide dimension to obtain a composite viewpoint image;
[0025] The synthesized viewpoint images are interwoven to generate a three-dimensional image.
[0026] In a second aspect, an embodiment of the present disclosure provides a viewpoint rendering device, wherein the viewpoint rendering device includes:
[0027] a first initial feature extraction module, configured to perform initial feature extraction on a two-dimensional image and a depth image corresponding to the two-dimensional image, to obtain a first initial feature of the two-dimensional image and a second initial feature of the depth image;
[0028] a concatenation module configured to concatenate the first initial feature and the second initial feature in a channel dimension, and perform dimensionality reduction on the channel dimension to obtain a first initial reduced dimensionality feature;
[0029] a fusion module configured to perform multiple image distortion and repair on the first initial dimensionality reduction features to obtain fused features;
[0030] The output module is configured to fuse the fused features and reduce the channel dimension to generate multiple viewpoint images.
[0031] Optionally, the viewpoint rendering device further includes:
[0032] A second initial feature extraction module is configured to perform initial feature extraction on the viewpoint image and the corresponding label image to obtain a third initial feature and a fourth initial feature;
[0033] a first downsampling and dimensionality reduction module, configured to downsample the third initial feature and the fourth initial feature in terms of width and height, and reduce the channel dimension to 1, to obtain a second initial dimensionality reduction feature and a third initial dimensionality reduction feature;
[0034] The first discrimination module is configured to discriminate the second initial dimensionality reduction feature and the third initial dimensionality reduction feature according to the label image.
[0035] Optionally, the viewpoint rendering device further includes:
[0036] a third initial feature extraction module, configured to perform initial feature extraction on the viewpoint image and the corresponding label image to obtain a fifth initial feature and a sixth initial feature;
[0037] a second downsampling and dimensionality reduction module, configured to downsample the fifth initial feature and the sixth initial feature in terms of width and height, and reduce the channel dimension to 1, to obtain a fourth initial dimensionality reduction feature and a fifth initial dimensionality reduction feature;
[0038] The second discrimination module is configured to discriminate the fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature according to the label image.
[0039] Optionally, the viewpoint rendering device further includes:
[0040] The depth image generation module is configured to input the two-dimensional image into a monocular depth estimation model to generate the depth image corresponding to the two-dimensional image.
[0041] Optionally, the viewpoint rendering device further includes:
[0042] a synthesis module configured to stitch the plurality of viewpoint images in a wide dimension to obtain a synthesized viewpoint image;
[0043] The interleaving module is configured to interleave the synthetic viewpoint images to generate a three-dimensional image.
[0044] Optionally, the first initial feature extraction module includes: a three-layer two-dimensional convolutional neural network;
[0045] The splicing module includes: a layer of two-dimensional convolutional neural network;
[0046] The fusion module includes: a three-layer two-dimensional convolutional neural network; a residual connection is used between the first layer of the two-dimensional convolutional neural network and the third layer of the two-dimensional convolutional neural network;
[0047] The output module includes: a three-layer two-dimensional convolutional neural network.
[0048] Optionally, the second initial feature extraction module includes: a one-layer two-dimensional convolutional neural network;
[0049] The first downsampling and dimensionality reduction module includes: a two-dimensional convolutional neural network with a multi-layer length of 2 and a stride of 1 and a binary adaptive mean pooling layer;
[0050] The first discrimination module includes: a layer of two-dimensional convolutional neural network.
[0051] Optionally, the third initial feature extraction module includes: a layer of two-dimensional convolutional neural network;
[0052] The second downsampling and dimensionality reduction module includes: a two-dimensional convolutional neural network with a multi-layer length of 2 and a stride of 1;
[0053] The second discrimination module includes: a layer of two-dimensional convolutional neural network.
[0054] In a third aspect, an embodiment of the present disclosure provides an electronic device, characterized by including:
[0055] at least one processor; and
[0056] a memory communicatively connected to the at least one processor; wherein,
[0057] The memory stores one or more computer programs that can be executed by the at least one processor. The one or more computer programs are executed by the at least one processor to enable the at least one processor to perform the viewpoint rendering method provided above.
[0058] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the viewpoint rendering method provided above when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] FIG1 is a schematic diagram of an exemplary viewpoint rendering method.
[0060] FIG2 is a flow chart of a viewpoint rendering method provided in an embodiment of the present disclosure.
[0061] FIG3 is a flow chart of another viewpoint rendering method provided by an embodiment of the present disclosure.
[0062] FIG4 is a flow chart of another viewpoint rendering method provided by an embodiment of the present disclosure.
[0063] FIG5 is a flow chart of another viewpoint rendering method provided by an embodiment of the present disclosure.
[0064] FIG6 is a schematic structural diagram of a viewpoint rendering device provided by an embodiment of the present disclosure.
[0065] FIG7 is a schematic structural diagram of a first feature extraction module in a viewpoint rendering device provided by an embodiment of the present disclosure.
[0066] FIG8 is a schematic structural diagram of a splicing module in a viewpoint rendering device provided by an embodiment of the present disclosure.
[0067] FIG9 is a schematic structural diagram of a fusion module in a viewpoint rendering device provided by an embodiment of the present disclosure.
[0068] FIG10 is a schematic structural diagram of an output module in a viewpoint rendering device provided by an embodiment of the present disclosure.
[0069] FIG11 is a schematic structural diagram of another viewpoint rendering device provided by an embodiment of the present disclosure.
[0070] FIG12 is a schematic structural diagram of another viewpoint rendering device provided by an embodiment of the present disclosure.
[0071] FIG13 is a schematic structural diagram of an electronic device provided in some embodiments of the present disclosure. DETAILED DESCRIPTION
[0072] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, the present disclosure is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0073] Unless otherwise defined, the technical or scientific terms used in this disclosure should have the usual meanings understood by people with ordinary skills in the field to which this disclosure belongs. The words "first", "second" and similar words used in this disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "one", "an" or "the" do not indicate a quantity limitation, but rather indicate the existence of at least one. Words such as "include" or "comprise" mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connect" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0074] With the rise of deep learning, applying it to 3D image generation has become one of the most cutting-edge research directions in the display field. Compared to manual post-production, deep learning-based 3D image generation technology offers advantages such as high efficiency, low cost, and high quality, making it ideal for a wide range of consumer needs.
[0075] Figure 1 is a schematic diagram of an exemplary viewpoint rendering method. As shown in Figure 1, the traditional viewpoint rendering process for 3D images primarily involves two aspects: 1. an image warping algorithm; and 2. an image inpainting algorithm. The image warping algorithm involves performing pixel offsets on a 2D input image based on its corresponding depth image to generate a new viewpoint image. The image inpainting algorithm involves performing pixel inpainting on the viewpoint image, which can result in holes with no pixel values. These holes need to be repaired to ensure a smooth and natural viewpoint image.
[0076] However, current image distortion algorithms rely too heavily on the high precision of depth images. If the input depth image is not accurate enough, it will greatly affect the quality of viewpoint image generation. In addition, existing image restoration algorithms will cause serious artifact problems in viewpoint image generation.
[0077] In order to solve at least one of the above-mentioned technical problems, the embodiments of the present disclosure provide a viewpoint rendering method and a viewpoint rendering device. The viewpoint rendering method and the viewpoint rendering device provided by the embodiments of the present disclosure will be further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0078] In a first aspect, an embodiment of the present disclosure provides a viewpoint rendering method. FIG2 is a flow chart of the viewpoint rendering method provided by an embodiment of the present disclosure. As shown in FIG2 , the viewpoint rendering method includes the following steps S201 to S204 .
[0079] S201 : Perform initial feature extraction on a two-dimensional image and a depth image corresponding to the two-dimensional image to obtain a first initial feature of the two-dimensional image and a second initial feature of the depth image.
[0080] In the above step S201, the original two-dimensional image F 2D and its corresponding depth image F Dep As the input of the generative network, the two are input into the initial feature extraction module for initial feature extraction to obtain the first initial feature F of the two-dimensional image. 2D-init And the second initial feature F of the depth image Dep-init , the size is B×C×H×W (where C is the number of designed channels and B is the batch size set during training).
[0081] S202: Concatenate the first initial feature and the second initial feature in the channel dimension, and perform dimensionality reduction on the channel dimension to obtain a first initial reduced dimensionality feature.
[0082] In the above step S202, the first initial feature F 2D-init and the second initial feature F Dep-init The channel dimension is spliced to obtain a feature of size B×(2C)×H×W, and the channel dimension is reduced through a two-dimensional convolutional neural network layer to obtain the first initial dimension reduction feature F init-down , dimensions are B × C × H × W.
[0083] S203: Perform multiple image distortion and restoration on the first initial dimensionality reduction feature to obtain a fusion feature.
[0084] In the above step S203, the first initial dimension reduction feature F init-down Input into n fusion modules to distort and repair the image and obtain the fusion feature F fusion , dimensions are B×C×H×W.
[0085] S204: Fusing the fused features and reducing the channel dimension to generate multiple viewpoint images.
[0086] In the above step S204, the fusion feature F fusion The input and output modules are fused and the channel dimension is reduced to finally obtain the generated viewpoint image F out , dimensions are B×3×H×W.
[0087] The viewpoint rendering method provided by the disclosed embodiments automatically renders each viewpoint image after the user uploads a standard 2D image. This method eliminates the need for high-precision depth images and effectively improves the robustness of image distortion in situations where depth image accuracy is poor. This significantly reduces artifacts caused by image restoration algorithms and meets the high-quality requirements of viewpoint image generation. Furthermore, the combined image distortion and restoration process simplifies the process and significantly speeds up viewpoint image generation.
[0088] In some embodiments, Figure 3 is a flow chart of another viewpoint rendering method provided by an embodiment of the present disclosure. As shown in Figure 3, in this viewpoint rendering method, the fusion features are fused and the channel dimension is reduced to generate multiple viewpoint images, and then the following steps S301 to S303 are also included.
[0089] S301 , performing initial feature extraction on the viewpoint image and the corresponding label image to obtain a third initial feature and a fourth initial feature.
[0090] In the above step S301, a one-layer two-dimensional convolutional neural network is used to perform initial feature extraction on the viewpoint image output by the generation network and the corresponding label image (both with sizes of B×3×H×W), and obtain three initial features and a fourth initial feature, both with sizes of B×C×H×W.
[0091] S302 , downsampling the third initial feature and the fourth initial feature in terms of width and height dimensions, reducing the width and height dimensions to 1, and reducing the channel dimension to 1, to obtain a second initial reduced dimensionality feature and a third initial reduced dimensionality feature.
[0092] In the above step S302, the width and height dimensions are downsampled through n sampling and dimensionality reduction modules consisting of a step size of 2 and a step size of 1. Then, the width and height dimensions are reduced to 1, with a size of B×C×1×1. Then, a layer of two-dimensional convolutional neural network is used to reduce the channel dimension to 1 to generate the second initial dimensionality reduction feature and the third initial dimensionality reduction feature, both of which have a size of B×1×1×1.
[0093] S303: Discriminate the second initial dimensionality reduction feature and the third initial dimensionality reduction feature according to the label image.
[0094] In step S303, the second and third initial dimensionality reduction features are evaluated based on the label image until the second and third initial dimensionality reduction features are close in similarity to the label image. If the similarity between the two is significantly different, steps S301 to S302 are repeated.
[0095] In some embodiments, Figure 4 is a flow chart of another viewpoint rendering method provided by an embodiment of the present disclosure. As shown in Figure 4, in this viewpoint rendering method, the fusion features are fused and the channel dimension is reduced to generate multiple viewpoint images, and then the following steps S401 to S403 are also included.
[0096] S401 , performing initial feature extraction on the viewpoint image and the corresponding label image to obtain a fifth initial feature and a sixth initial feature.
[0097] In the above step S401, a one-layer two-dimensional convolutional neural network is used to perform initial feature extraction on the viewpoint image output by the generation network and the corresponding label image (both with dimensions of B×3×H×W), and obtain five initial features and a sixth initial feature, both with dimensions of B×C×H×W.
[0098] S402 , downsampling the fifth initial feature and the sixth initial feature in terms of width and height, and reducing the channel dimension to 1, to obtain a fourth initial reduced-dimensionality feature and a fifth initial reduced-dimensionality feature.
[0099] In the above step S402, after n sampling and dimensionality reduction modules consisting of a step size of 2 and a step size of 1, downsampling of the width and height dimensions is performed, and the size is B×C×(H / / n)×(W / / n). Then, a layer of two-dimensional convolutional neural network is used to reduce the channel dimension to 1 to generate the fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature, both of which have a size of B×1×(H / / n)×(W / / n).
[0100] S403: Discriminate the fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature according to the label image.
[0101] In step S403, the fourth and fifth initial dimensionality reduction features are evaluated based on the label image until the fourth and fifth initial dimensionality reduction features are close in similarity to the label image. If the similarity is significantly different, steps S401 to S402 are repeated.
[0102] In some embodiments, there are multiple groups of the fourth initial dimensionality reduction features and the fifth initial dimensionality reduction features; each group of the fourth initial dimensionality reduction features and the fifth initial dimensionality reduction features corresponds to the order of the discriminated pixel blocks from large to small.
[0103] In step S402 above, after downsampling the width and height dimensions, the width and height dimensions are not 1. This means that the image can be divided into multiple pixel blocks for separate evaluation, rather than evaluating the entire image. Furthermore, when evaluating pixel blocks, the number of sampling and dimensionality reduction modules can be varied, thereby varying the size of the pixel blocks being evaluated, thereby achieving a progressive evaluation process.
[0104] In some embodiments, Figure 5 is a flow chart of another viewpoint rendering method provided by an embodiment of the present disclosure. As shown in Figure 5, in this viewpoint rendering method, initial feature extraction is performed on a two-dimensional image and a depth image corresponding to the two-dimensional image to obtain a first initial feature of the two-dimensional image and a second initial feature of the depth image. It also includes step S501, inputting the two-dimensional image into a monocular depth estimation model to generate a depth image corresponding to the two-dimensional image.
[0105] In the above step S501, the original two-dimensional image F 2D (The size is 3×H×W, where 3 represents RGB3 channels, H represents the height of the video frame, and W represents the width of the video frame) and is input into an open source monocular depth estimation algorithm (such as Monodepth, Monodepth2, etc.) to obtain the corresponding depth image F Dep (Size is 3×H×W, which is the same as the input two-dimensional image F 2D keep the size consistent).
[0106] In some embodiments, as shown in FIG5 , in the viewpoint rendering method, the fusion features are fused and the channel dimension is reduced to generate multiple viewpoint images, and then the method further includes step S503 and step S504 .
[0107] S503 , stitching the multiple viewpoint images in a wide dimension to obtain a synthesized viewpoint image.
[0108] S504: Interweave the synthesized viewpoint images to generate a three-dimensional image.
[0109] In the above steps S601 to S602, taking two viewpoints as an example, a viewpoint image F 2D,1 and the second viewpoint image F 2D,2 , stitching in the wide dimension to obtain a synthetic 3D two-view image F 3D . Use the corresponding interleaving algorithm to interleave the 3D two-view image F 3D After interleaving, it can be displayed on a 3D display device.
[0110] In the second aspect, an embodiment of the present disclosure provides a viewpoint rendering device. Figure 6 is a structural schematic diagram of a viewpoint rendering device provided by an embodiment of the present disclosure. As shown in Figure 6, the viewpoint rendering device includes: a first initial feature extraction module 601, a splicing module 602, a fusion module 603, and an output module 604.
[0111] The first initial feature extraction module 601 is configured to perform initial feature extraction on the two-dimensional image and the depth image corresponding to the two-dimensional image, and obtain a first initial feature of the two-dimensional image and a second initial feature of the depth image.
[0112] FIG7 is a schematic diagram of the structure of the first feature extraction module in a viewpoint rendering device provided by an embodiment of the present disclosure. As shown in FIG7 , the first initial feature extraction module 601 includes: a three-layer two-dimensional convolutional neural network. 2D and its corresponding depth image F Dep As the input of the generative network, a layer of two-dimensional convolutional neural network is first used to reduce the channel dimension of the input features (size B×C×H×W) by 2 times, and obtain features of size B×(C / / 2)×H×W. Then, a layer of two-dimensional convolutional neural network is used to efficiently extract the features, and obtain features of size B×(C / / 2)×H×W. Finally, a layer of two-dimensional convolutional neural network is used to increase the channel dimension of the features, and the first initial feature F is obtained. 2D-init And the second initial feature F of the depth image Dep-init , with dimensions of B×C×H×W. The first initial feature extraction module 601 is funnel-shaped in the channel dimension. This design can achieve a good balance between computational complexity and performance.
[0113] The splicing module 602 is configured to splice the first initial feature and the second initial feature in the channel dimension, and perform dimensionality reduction on the channel dimension to obtain a first initial reduced-dimensionality feature.
[0114] FIG8 is a schematic diagram of the structure of a splicing module in a viewpoint rendering device provided by an embodiment of the present disclosure. As shown in FIG8 , the splicing module 602 includes: a layer of two-dimensional convolutional neural network. 2D-init and the second initial feature F Dep-init The channel dimension is spliced to obtain a feature of size B×(2C)×H×W, and the channel dimension is reduced through a two-dimensional convolutional neural network layer to obtain the first initial dimension reduction feature F init-down , dimensions are B×C×H×W.
[0115] The fusion module 603 is configured to perform multiple image distortions and repairs on the first initial dimensionality reduction features to obtain fused features.
[0116] FIG9 is a schematic diagram of the structure of a fusion module in a viewpoint rendering device provided by an embodiment of the present disclosure. As shown in FIG9 , the fusion module 603 includes: a three-layer two-dimensional convolutional neural network. The difference between the fusion module 603 and the first initial feature extraction module 601 is that the first layer of the two-dimensional convolutional neural network and the third layer of the two-dimensional convolutional neural network use residual connections, which is conducive to network optimization and convergence. The first initial feature F 2D-init and the second initial feature F Dep-init The channel dimension is spliced to obtain a feature of size B×(2C)×H×W, and the channel dimension is reduced through a two-dimensional convolutional neural network layer to obtain the first initial dimension reduction feature Finit-down , dimensions are B×C×H×W.
[0117] The output module 604 is configured to fuse the fused features and reduce the channel dimension to generate multiple viewpoint images.
[0118] FIG10 is a schematic diagram of the structure of an output module in a viewpoint rendering device provided by an embodiment of the present disclosure. As shown in FIG10 , the output module 604 includes: a three-layer two-dimensional convolutional neural network. The three-layer two-dimensional convolutional neural network gradually reduces the channel dimension to output the final color three-channel viewpoint image F out , dimensions are B×3×H×W.
[0119] In some embodiments, Figure 11 is a structural schematic diagram of another viewpoint rendering device provided in an embodiment of the present disclosure. As shown in Figure 11, the viewpoint rendering device also includes: a second initial feature extraction module 111, a first downsampling and dimensionality reduction module 112, and a first discrimination module 113.
[0120] The second initial feature extraction module 111 is configured to perform initial feature extraction on the viewpoint image and the corresponding label image to obtain a third initial feature and a fourth initial feature.
[0121] As shown in Figure 11, the second initial feature extraction module 111 includes: a layer of two-dimensional convolutional neural network; a layer of two-dimensional convolutional neural network is used to perform initial feature extraction on the viewpoint image output by the generation network and the corresponding label image (both with sizes of B×3×H×W), and obtain three initial features and a fourth initial feature, both with sizes of B×C×H×W.
[0122] The first downsampling and dimensionality reduction module 112 is configured to downsample the third initial feature and the fourth initial feature in terms of width and height, and reduce the channel dimension to 1, to obtain the second initial dimensionality reduction feature and the third initial dimensionality reduction feature.
[0123] As shown in Figure 11, the first downsampling and dimensionality reduction module 112 includes: multiple layers of two-dimensional convolutional neural networks with a stride of 2 and a step size of 1, and a layer of binary adaptive mean pooling layer (AdaptiveAvgPool2d); after n two-dimensional convolutional neural networks with a stride of 2 and a stride of 1, downsampling of the width and height dimensions is performed, and then a layer of binary adaptive mean pooling layer is used to reduce the width and height dimensions to 1, with a size of B×C×1×1, and then a layer of two-dimensional convolutional neural network is used to reduce the channel dimension to 1, generating a second initial dimensionality reduction feature and a third initial dimensionality reduction feature, both of which have a size of B×1×1×1.
[0124] The first discrimination module 113 is configured to discriminate the second initial dimensionality reduction feature and the third initial dimensionality reduction feature according to the label image.
[0125] As shown in FIG11 , the first discrimination module 113 includes a one-layer two-dimensional convolutional neural network. The one-layer two-dimensional convolutional neural network discriminates the second and third initial dimensionality reduction features based on the label image until the second and third initial dimensionality reduction features have similarities to the label image. If the similarities are significantly different, the above process is repeated.
[0126] In some embodiments, Figure 12 is a structural schematic diagram of another viewpoint rendering device provided in an embodiment of the present disclosure. As shown in Figure 12, the viewpoint rendering device also includes: a third initial feature extraction module 121, a second downsampling and dimensionality reduction module 122, and a second discrimination module 123.
[0127] The third initial feature extraction module 121 is configured to perform initial feature extraction on the viewpoint image and the corresponding label image to obtain a fifth initial feature and a sixth initial feature.
[0128] As shown in Figure 12, the third initial feature extraction module 121 includes a one-layer two-dimensional convolutional neural network. The one-layer two-dimensional convolutional neural network is used to perform initial feature extraction on the viewpoint image output by the generation network and the corresponding label image (both with dimensions of B×3×H×W), obtaining the fifth initial feature and the sixth initial feature, both with dimensions of B×C×H×W.
[0129] The second downsampling and dimensionality reduction module 122 is configured to downsample the fifth initial feature and the sixth initial feature in terms of width and height, and reduce the channel dimension to 1, to obtain the fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature.
[0130] As shown in FIG12 , the second downsampling and dimensionality reduction module 122 includes: multiple layers of a 2D convolutional neural network with a stride of 1 and a layer of a 2D convolutional neural network. After passing through n sampling and dimensionality reduction modules with a stride of 2 and a stride of 1, downsampling is performed in both the width and height dimensions to a size of B×C×(H / / n)×(W / / n). Then, a layer of a 2D convolutional neural network is used to reduce the channel dimension to 1, generating the fourth and fifth initial dimensionality reduction features, both with a size of B×1×(H / / n)×(W / / n).
[0131] The second discrimination module 123 is configured to discriminate the fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature according to the label image.
[0132] The second discrimination module 123 includes a two-dimensional convolutional neural network. Based on the label image, the fourth and fifth initial dimensionality reduction features are discriminated until the fourth and fifth initial dimensionality reduction features have similarities to the label image. If the similarities are significantly different, the above process is repeated.
[0133] In some embodiments, the viewpoint rendering device further includes: a depth image generation module (not shown in the figure), which is configured to input the two-dimensional image into a monocular depth estimation model to generate a depth image corresponding to the two-dimensional image.
[0134] The depth image generation module can convert the original two-dimensional image F 2D (The size is 3×H×W, where 3 represents RGB3 channels, H represents the height of the video frame, and W represents the width of the video frame) and is input into an open source monocular depth estimation algorithm (such as Monodepth, Monodepth2, etc.) to obtain the corresponding depth image F Dep (Size is 3×H×W, which is the same as the input two-dimensional image F 2D keep the size consistent).
[0135] In some embodiments, the viewpoint rendering device also includes: a synthesis module (not shown in the figure), configured to splice multiple viewpoint images in a wide dimension to obtain a synthetic viewpoint image; an interweaving module (not shown in the figure), configured to interweave the synthetic viewpoint images to generate a three-dimensional image.
[0136] Taking two viewpoints as an example, a viewpoint image F 2D,1 and the second viewpoint image F 2D,2 The synthesis module can stitch the two viewpoint images in a wide dimension to obtain a synthesized 3D two-viewpoint image F 3D The interleaving module can use the corresponding interleaving algorithm to interleave the 3D two-view image F 3D After interleaving, it can be displayed on a 3D display device.
[0137] On the third aspect, an embodiment of the present disclosure provides an electronic device. Figure 13 is a structural diagram of the electronic device provided in some embodiments of the present disclosure. As shown in Figure 13, the electronic device includes: one or more processors 131; a memory 132, on which one or more programs are stored. When the one or more programs are executed by one or more processors, the one or more processors implement the viewpoint rendering method provided in any of the above embodiments; one or more I / O interfaces 133, connected between the processor and the memory, and configured to implement information interaction between the processor and the memory.
[0138] Among them, the processor 131 is a device with data processing capabilities, including but not limited to a central processing unit (CPU); the memory 132 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory (FLASH); the I / O interface (read-write interface) 133 is connected between the processor 131 and the memory 132, and can realize information interaction between the processor 131 and the memory 132, including but not limited to a data bus (Bus), etc.
[0139] In some embodiments, the processor 131 , the memory 132 , and the I / O interface 133 are connected to each other via a bus, and further connected to other components of the computing device.
[0140] In a fourth aspect, this embodiment provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the viewpoint rendering method provided by any of the above embodiments.
[0141] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable, and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0142] It is understood that the above embodiments are merely exemplary embodiments for illustrating the principles of the present disclosure, and the present disclosure is not limited thereto. Those skilled in the art may make various modifications and improvements without departing from the spirit and substance of the present disclosure, and such modifications and improvements are also considered to be within the scope of protection of the present disclosure.
Claims
1. A viewpoint rendering method, wherein: The viewpoint rendering method comprises: Performing initial feature extraction on a two-dimensional image and a depth image corresponding to the two-dimensional image to obtain a first initial feature of the two-dimensional image and a second initial feature of the depth image; The first initial feature and the second initial feature are concatenated in the channel dimension, and the channel dimension is reduced to obtain a first initial reduced dimension feature; Performing multiple image distortion and repair on the first initial dimension reduction feature to obtain a fusion feature; The fusion features are fused and channel dimension reduction is performed to generate multiple viewpoint images.
2. The viewpoint rendering method according to claim 1, wherein: The fusion features are fused and the channel dimension is reduced to generate multiple viewpoint images, and then the following steps are further included: Performing initial feature extraction on the viewpoint image and the corresponding label image to obtain a third initial feature and a fourth initial feature; Downsampling the third initial feature and the fourth initial feature in terms of width and height dimensions, reducing the width and height dimensions to 1, and reducing the channel dimension to 1, to obtain a second initial reduced dimensionality feature and a third initial reduced dimensionality feature; The second initial dimensionality reduction feature and the third initial dimensionality reduction feature are discriminated according to the label image.
3. The viewpoint rendering method according to claim 1, wherein: The fusion features are fused and the channel dimension is reduced to generate multiple viewpoint images, and then the following steps are further included: Performing initial feature extraction on the viewpoint image and the corresponding label image to obtain a fifth initial feature and a sixth initial feature; Downsampling the fifth initial feature and the sixth initial feature in terms of width and height, and reducing the channel dimension to 1, to obtain a fourth initial reduced dimensionality feature and a fifth initial reduced dimensionality feature; The fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature are discriminated according to the label image.
4. The viewpoint rendering method according to claim 3, wherein: The number of groups of the fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature is multiple; Each group of the fourth initial dimensionality reduction features and the fifth initial dimensionality reduction features corresponds to the order of the discriminated pixel blocks from large to small.
5. The viewpoint rendering method according to claim 1, wherein: The extracting of initial features from the two-dimensional image and the depth image corresponding to the two-dimensional image to obtain a first initial feature of the two-dimensional image and a second initial feature of the depth image also includes: The two-dimensional image is input into a monocular depth estimation model to generate the depth image corresponding to the two-dimensional image.
6. The viewpoint rendering method according to claim 1, wherein: The fusion features are fused and the channel dimension is reduced to generate multiple viewpoint images, and then the following steps are further included: splicing the plurality of viewpoint images in a wide dimension to obtain a synthetic viewpoint image; The synthesized viewpoint images are interlaced to generate a three-dimensional image.
7. A viewpoint rendering device, wherein: The viewpoint rendering device comprises: A first initial feature extraction module is configured to perform initial feature extraction on a two-dimensional image and a depth image corresponding to the two-dimensional image, and obtain a first initial feature of the two-dimensional image and a second initial feature of the depth image; a concatenation module configured to concatenate the first initial feature and the second initial feature in a channel dimension, and perform dimension reduction on the channel dimension to obtain a first initial reduced dimension feature; A fusion module is configured to perform multiple image distortion and repair on the first initial dimension reduction feature to obtain a fusion feature; The output module is configured to fuse the fused features and reduce the channel dimension to generate multiple viewpoint images.
8. The viewpoint rendering device according to claim 7, wherein: The viewpoint rendering device also includes: A second initial feature extraction module is configured to perform initial feature extraction on the viewpoint image and the corresponding label image to obtain a third initial feature and a fourth initial feature; A first downsampling and dimension reduction module is configured to downsample the third initial feature and the fourth initial feature in terms of width and height, and reduce the channel dimension to 1, to obtain a second initial dimension reduction feature and a third initial dimension reduction feature; The first discrimination module is configured to discriminate the second initial dimensionality reduction feature and the third initial dimensionality reduction feature according to the label image.
9. The viewpoint rendering device according to claim 7, wherein: The viewpoint rendering device also includes: A third initial feature extraction module is configured to perform initial feature extraction on the viewpoint image and the corresponding label image to obtain a fifth initial feature and a sixth initial feature; A second downsampling and dimension reduction module is configured to downsample the fifth initial feature and the sixth initial feature in terms of width and height, and reduce the channel dimension to 1, to obtain a fourth initial dimension reduction feature and a fifth initial dimension reduction feature; The second discrimination module is configured to discriminate the fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature according to the label image.
10. The viewpoint rendering device according to claim 7, wherein: The viewpoint rendering device also includes: The depth image generation module is configured to input the two-dimensional image into a monocular depth estimation model to generate the depth image corresponding to the two-dimensional image.
11. The viewpoint rendering device according to claim 7, wherein: The viewpoint rendering device also includes: A synthesis module is configured to stitch the plurality of viewpoint images in a wide dimension to obtain a synthesized viewpoint image; The interlacing module is configured to interlace the synthetic viewpoint images to generate a three-dimensional image.
12. The viewpoint rendering device according to claim 7, wherein: The first initial feature extraction module includes: a three-layer two-dimensional convolutional neural network; The splicing module includes: a layer of two-dimensional convolutional neural network; The fusion module includes: a three-layer two-dimensional convolutional neural network; a residual connection is used between the first layer of the two-dimensional convolutional neural network and the third layer of the two-dimensional convolutional neural network; The output module includes: a three-layer two-dimensional convolutional neural network.
13. The viewpoint rendering device according to claim 8, wherein: The second initial feature extraction module includes: a layer of two-dimensional convolutional neural network; The first downsampling and dimension reduction module includes: a two-dimensional convolutional neural network with a multi-layer length of 2 and a step size of 1 and a binary adaptive mean pooling layer; The first discrimination module includes: a layer of two-dimensional convolutional neural network.
14. The viewpoint rendering device according to claim 9, wherein: The third initial feature extraction module includes: a layer of two-dimensional convolutional neural network; The second downsampling and dimension reduction module includes: a two-dimensional convolutional neural network with a multi-layer length of 2 and a step length of 1; The second discrimination module includes: a layer of two-dimensional convolutional neural network.
15. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor to enable the at least one processor to perform the viewpoint rendering method according to any one of claims 1 to 6.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the computer program implements the viewpoint rendering method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method and appratus for estimating depth, and method and apparatus for converting 2d video to 3d video
CN101754040A
Projection full-convolution network three-dimensional model segmentation method based on fusion of multi-view-angle features
CN108389251A
Image rendering method, device and system, computer readable storage medium and equipment
CN108665521A
Continuous time warp and binocular time warp for virtual and augmented reality display systems and methods
CN109863538A
Virtual reality display method, display device and computer readable medium
CN111290581A