Viewpoint rendering method and viewpoint rendering apparatus

By performing feature extraction and dimensionality reduction on 2D and depth images, and combining deep learning algorithms to generate viewpoint images, the problems of high cost and low efficiency in existing technologies are solved, and high-quality 3D image generation is achieved.

CN119998845BActive Publication Date: 2026-05-29BOE TECHNOLOGY GROUP CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BOE TECHNOLOGY GROUP CO LTD
Filing Date
2023-08-14
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In existing technologies, binocular cameras are expensive, and 2D images cannot be directly converted into 3D content. Post-production is costly and cannot meet the general needs of users.

Method used

By performing initial feature extraction, stitching, and dimensionality reduction on 2D and depth images, and combining deep learning algorithms for image distortion and restoration, multiple viewpoint images are generated, reducing the reliance on high-precision depth images.

Benefits of technology

It enables the generation of high-quality 3D images at low cost and high efficiency, reduces artifact problems, and is suitable for the needs of consumer-grade 3D content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998845B_ABST
    Figure CN119998845B_ABST
Patent Text Reader

Abstract

The application discloses a view point rendering method and device, belonging to the field of display technology, which can solve the problem that the existing image distortion algorithm is too dependent on high precision of a depth image. The view point rendering method comprises the following steps: performing initial feature extraction on a two-dimensional image and a depth image corresponding to the two-dimensional image to obtain first initial features of the two-dimensional image and second initial features of the depth image; splicing the first initial features and the second initial features in a channel dimension and performing dimension reduction in the channel dimension to obtain first initial reduced features; performing image distortion and repair on the first initial reduced features for multiple times to obtain fusion features; and performing fusion and channel dimension reduction on the fusion features to generate multiple view point images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure belongs to the field of display technology, specifically relating to a viewpoint rendering method and viewpoint rendering device. Background Technology

[0002] With the diversification and increasing standardization of user needs, ultra-high-definition 2D images are gradually becoming insufficient to meet users' viewing demands. The stereoscopic effect and visual impact brought by 3D images have become the latest pursuit of users. Therefore, glasses-free 3D display technology has become the latest hot topic in the display field. With the launch of glasses-free 3D display products, there is a broad market demand for consumer-grade 3D content.

[0003] Currently, 3D content originates from two sources. One is directly capturing native 3D images and videos using binocular cameras. However, binocular cameras are not widespread, are expensive, and existing 2D images cannot be used, requiring the capture of new 3D content, which is not suitable for general user needs. The other source is converting existing 2D images and videos into 3D. However, this is currently mostly done manually using professional software in post-production, primarily used in film-level productions. Manual post-production is expensive, time-consuming, and labor-intensive, also unsuitable for general user needs. Summary of the Invention

[0004] This disclosure aims to at least solve one of the technical problems existing in the prior art, and provides a viewpoint rendering method and viewpoint rendering apparatus.

[0005] In a first aspect, embodiments of this disclosure provide a viewpoint rendering method, wherein the viewpoint rendering method includes:

[0006] Initial feature extraction is performed on the two-dimensional image and the corresponding depth image to obtain the first initial feature of the two-dimensional image and the second initial feature of the depth image;

[0007] The first initial feature and the second initial feature are concatenated along the channel dimension, and then the channel dimension is reduced to obtain the first initial dimensionality-reduced feature;

[0008] The first initial dimensionality reduction feature is subjected to multiple image distortions and repairs to obtain fused features;

[0009] The fusion features are fused and the channel dimension is reduced to generate multiple viewpoint images.

[0010] Optionally, the step of fusing the fused features and reducing the channel dimension to generate multiple viewpoint images further includes:

[0011] Initial feature extraction is performed on the viewpoint image and the corresponding label image to obtain the third initial feature and the fourth initial feature;

[0012] The third initial feature and the fourth initial feature are downsampled in width and height dimensions, and the width and height dimensions are reduced to 1, and the channel dimension is reduced to 1, to obtain the second initial dimensionality reduction feature and the third initial dimensionality reduction feature;

[0013] Based on the labeled image, the second initial dimensionality reduction feature and the third initial dimensionality reduction feature are distinguished.

[0014] Optionally, the step of fusing the fused features and reducing the channel dimension to generate multiple viewpoint images further includes:

[0015] Initial feature extraction is performed on the viewpoint image and the corresponding label image to obtain the fifth initial feature and the sixth initial feature;

[0016] The fifth and sixth initial features are downsampled in width and height dimensions, and the channel dimension is reduced to 1 to obtain the fourth and fifth initial dimensionality-reduced features.

[0017] Based on the labeled image, the fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature are discriminated.

[0018] Optionally, the fourth and fifth initial dimensionality reduction features may consist of multiple groups;

[0019] Each set of the fourth and fifth initial dimensionality reduction features corresponds to the pixel blocks being identified in descending order.

[0020] Optionally, the step of performing initial feature extraction on the two-dimensional image and the corresponding depth image to obtain a first initial feature of the two-dimensional image and a second initial feature of the depth image further includes:

[0021] The two-dimensional image is input into a monocular depth estimation model to generate the depth image corresponding to the two-dimensional image.

[0022] Optionally, the fusion features are fused and the channel dimension is reduced to generate multiple viewpoint images, and then the process further includes:

[0023] Multiple viewpoint images are stitched together in a wide dimension to obtain a composite viewpoint image;

[0024] The synthesized viewpoint images are interwoven to generate a three-dimensional image.

[0025] Secondly, embodiments of this disclosure provide a viewpoint rendering apparatus, wherein the viewpoint rendering apparatus includes:

[0026] The first initial feature extraction module is configured to perform initial feature extraction on a two-dimensional image and a depth image corresponding to the two-dimensional image to obtain a first initial feature of the two-dimensional image and a second initial feature of the depth image.

[0027] The splicing module is configured to splice the first initial feature and the second initial feature in the channel dimension, and perform channel dimension reduction to obtain the first initial dimensionality-reduced feature;

[0028] The fusion module is configured to perform multiple image distortions and repairs on the first initial dimensionality reduction features to obtain fused features;

[0029] The output module is configured to fuse the fused features and reduce the channel dimension to generate multiple viewpoint images.

[0030] Optionally, the viewpoint rendering apparatus further includes:

[0031] The second initial feature extraction module is configured to perform initial feature extraction on the viewpoint image and the corresponding label image to obtain the third initial feature and the fourth initial feature.

[0032] The first downsampling and dimensionality reduction module is configured to downsample the width and height dimensions of the third initial feature and the fourth initial feature, and reduce the channel dimension to 1 to obtain the second initial dimensionality reduction feature and the third initial dimensionality reduction feature.

[0033] The first discrimination module is configured to discriminate the second initial dimensionality reduction feature and the third initial dimensionality reduction feature based on the label image.

[0034] Optionally, the viewpoint rendering apparatus further includes:

[0035] The third initial feature extraction module is configured to perform initial feature extraction on the viewpoint image and the corresponding label image to obtain the fifth initial feature and the sixth initial feature.

[0036] The second downsampling and dimensionality reduction module is configured to downsample the fifth initial feature and the sixth initial feature in both width and height dimensions, and reduce the channel dimension to 1 to obtain the fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature.

[0037] The second discrimination module is configured to discriminate the fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature based on the label image.

[0038] Optionally, the viewpoint rendering apparatus further includes:

[0039] The depth image generation module is configured to input the two-dimensional image into a monocular depth estimation model to generate the depth image corresponding to the two-dimensional image.

[0040] Optionally, the viewpoint rendering apparatus further includes:

[0041] The compositing module is configured to stitch together multiple viewpoint images in a wide dimension to obtain a composite viewpoint image;

[0042] An interlacing module is configured to interlace the synthesized viewpoint image to generate a three-dimensional image.

[0043] Optionally, the first initial feature extraction module includes: a three-layer two-dimensional convolutional neural network;

[0044] The splicing module includes: a one-layer two-dimensional convolutional neural network;

[0045] The fusion module includes: a three-layer two-dimensional convolutional neural network; the first layer and the third layer are connected by a residual connection;

[0046] The output module includes a three-layer two-dimensional convolutional neural network.

[0047] Optionally, the second initial feature extraction module includes: a one-layer two-dimensional convolutional neural network;

[0048] The first downsampling and dimensionality reduction module includes: a multi-layer two-dimensional convolutional neural network with a length of 2 and a stride of 1, and a single-layer binary adaptive mean convergence layer;

[0049] The first discrimination module includes: a one-layer two-dimensional convolutional neural network.

[0050] Optionally, the third initial feature extraction module includes: a one-layer two-dimensional convolutional neural network;

[0051] The second downsampling and dimensionality reduction module includes: a multi-layer two-dimensional convolutional neural network with a length of 2 and a stride of 1;

[0052] The second discrimination module includes: a one-layer two-dimensional convolutional neural network.

[0053] Thirdly, embodiments of this disclosure provide an electronic device, characterized in that it includes:

[0054] At least one processor; and

[0055] A memory communicatively connected to the at least one processor; wherein,

[0056] The memory stores one or more computer programs that can be executed by the at least one processor, the one or more of the computer programs being executed by the at least one processor to enable the at least one processor to perform the viewpoint rendering method as provided above.

[0057] Fourthly, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the viewpoint rendering method provided above. Attached Figure Description

[0058] Figure 1 This is a schematic diagram of an exemplary viewpoint rendering method.

[0059] Figure 2 This is a flowchart illustrating the viewpoint rendering method provided in an embodiment of the present disclosure.

[0060] Figure 3 This is a flowchart illustrating another viewpoint rendering method provided in an embodiment of the present disclosure.

[0061] Figure 4 This is a flowchart illustrating another viewpoint rendering method provided in an embodiment of the present disclosure.

[0062] Figure 5 This is a flowchart illustrating another viewpoint rendering method provided in an embodiment of the present disclosure.

[0063] Figure 6 This is a schematic diagram of the structure of a viewpoint rendering device provided in an embodiment of the present disclosure.

[0064] Figure 7 This is a schematic diagram of the structure of a first feature extraction module in a viewpoint rendering apparatus provided in an embodiment of the present disclosure.

[0065] Figure 8 This is a schematic diagram of the structure of a splicing module in a viewpoint rendering device provided in an embodiment of this disclosure.

[0066] Figure 9 This is a schematic diagram of the structure of a fusion module in a viewpoint rendering apparatus provided in an embodiment of the present disclosure.

[0067] Figure 10 This is a schematic diagram of the structure of an output module in a viewpoint rendering apparatus provided in an embodiment of the present disclosure.

[0068] Figure 11 This is a schematic diagram of another viewpoint rendering apparatus provided in an embodiment of the present disclosure.

[0069] Figure 12 This is a schematic diagram of another viewpoint rendering apparatus provided in an embodiment of the present disclosure.

[0070] Figure 13 This is a schematic diagram of the structure of an electronic device provided in some embodiments of this disclosure. Detailed Implementation

[0071] To enable those skilled in the art to better understand the technical solutions of this disclosure, the disclosure will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0072] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms “first,” “second,” and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an,” “a,” or “the,” and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms “including,” “comprising,” or “containing,” and similar terms mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. The terms “connected,” “linked,” or similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms “upper,” “lower,” “left,” and “right,” etc., are used only to indicate relative positional relationships, and these relative positional relationships may change accordingly when the absolute position of the described objects changes.

[0073] With the rise of deep learning, applying deep learning technology to 3D image generation has become one of the most cutting-edge research directions in the display field. Compared with manual post-production, deep learning-based 3D image generation technology has advantages such as high efficiency, low cost, and high generation quality, making it very suitable for the wide range of consumer needs.

[0074] Figure 1 This is a schematic diagram of an exemplary viewpoint rendering method, such as... Figure 1 As shown, the traditional 3D image viewpoint rendering process mainly includes two aspects: 1. Image warping algorithm; 2. Image inpainting algorithm. The image warping algorithm refers to shifting pixels in the input 2D image and its corresponding depth image to obtain a new viewpoint image. The image inpainting algorithm addresses the issue that the image warping algorithm may result in holes with no pixel values ​​in the obtained viewpoint image, requiring repair to make the entire viewpoint image smooth and natural.

[0075] However, current image warping algorithms rely too heavily on the high accuracy of depth images. If the accuracy of the input depth image is insufficient, it will greatly affect the quality of the generated viewpoint image. Furthermore, existing image inpainting algorithms cause serious artifact problems in viewpoint image generation.

[0076] To at least solve one of the aforementioned technical problems, this disclosure provides a viewpoint rendering method and a viewpoint rendering apparatus. The viewpoint rendering method and viewpoint rendering apparatus provided in this disclosure will be described in further detail below with reference to the accompanying drawings and specific embodiments.

[0077] Firstly, embodiments of this disclosure provide a viewpoint rendering method. Figure 2 This is a flowchart illustrating the viewpoint rendering method provided in the embodiments of this disclosure, as shown below. Figure 2 As shown, the viewpoint rendering method includes the following steps S201 to S204.

[0078] S201, perform initial feature extraction on the two-dimensional image and the corresponding depth image to obtain the first initial feature of the two-dimensional image and the second initial feature of the depth image.

[0079] In step S201 above, the original two-dimensional image F 2D and its corresponding depth image F Dep As inputs to the generator network, both are fed into the initial feature extraction module for initial feature extraction, resulting in the first initial feature F of the two-dimensional image. 2D-init and the second initial feature F of the depth image Dep-init The dimensions are all B×C×H×W (where C is the number of channels designed and B is the batch size set during training).

[0080] S202, the first initial feature and the second initial feature are concatenated along the channel dimension, and the channel dimension is reduced to obtain the first initial dimensionality-reduced feature.

[0081] In step S202 above, the first initial feature F is... 2D-init Second initial feature F Dep-init The features are concatenated along the channel dimension to obtain features of size B×(2C)×H×W, and then passed through a two-dimensional convolutional neural network layer to reduce the channel dimension, resulting in the first initial dimensionality-reduced feature F. init-down The dimensions are B×C×H×W.

[0082] S203, the first initial dimensionality reduction features are subjected to multiple image distortions and repairs to obtain fused features.

[0083] In step S203 above, the first initial dimensionality reduction feature F is... init-down The image is input into n fusion modules for distortion and restoration, resulting in fusion features F. fusion The dimensions are B×C×H×W.

[0084] S204 performs feature fusion and channel dimension reduction to generate multiple viewpoint images.

[0085] In step S204 above, the fused feature F fusion The input and output modules perform fusion and channel dimension reduction to finally obtain the generated viewpoint image F. out The dimensions are B×3×H×W.

[0086] In the viewpoint rendering method provided in this disclosure, after a user uploads a regular 2D image, various viewpoint images can be automatically rendered. This eliminates the need for high-precision depth images, effectively improving the robustness of image distortion in cases of poor depth image accuracy. Therefore, it significantly improves the artifact problem caused by image inpainting algorithms, achieving the high-quality requirements for viewpoint image generation. Furthermore, the combination of image distortion and inpainting simplifies the process and significantly increases the speed of viewpoint image generation.

[0087] In some embodiments, Figure 3 A flowchart illustrating another viewpoint rendering method provided in this embodiment of the disclosure is shown below. Figure 3 As shown, in this viewpoint rendering method, the fusion features are fused and the channel dimension is reduced to generate multiple viewpoint images. Then, the method includes the following steps S301 to S303.

[0088] S301, perform initial feature extraction on the viewpoint image and the corresponding label image to obtain the third initial feature and the fourth initial feature.

[0089] In step S301 above, a two-dimensional convolutional neural network is used to extract initial features from the viewpoint image and the corresponding label image (both with dimensions B×3×H×W) output by the generating network, to obtain the third initial feature and the fourth initial feature, both with dimensions B×C×H×W.

[0090] S302, perform width and height downsampling on the third and fourth initial features, reduce the width and height dimensions to 1, and reduce the channel dimension to 1 to obtain the second and third initial dimensionality-reduced features.

[0091] In step S302 above, after passing through n sampling and dimensionality reduction modules with step size of 2 and step size of 1, the width and height dimensions are downsampled. Then, the width and height dimensions are reduced to 1, with a size of B×C×1×1. Then, a two-dimensional convolutional neural network is used to reduce the channel dimension to 1, generating the second initial dimensionality reduction feature and the third initial dimensionality reduction feature, both with a size of B×1×1×1.

[0092] S303, Based on the label image, the second initial dimensionality reduction feature and the third initial dimensionality reduction feature are discriminated.

[0093] In step S303 above, the second initial dimensionality reduction feature and the third initial dimensionality reduction feature are judged based on the label image until the similarity between the second initial dimensionality reduction feature and the third initial dimensionality reduction feature and the label image is close. If the similarity between the two differs greatly, steps S301 to S302 above are repeated.

[0094] In some embodiments, Figure 4 This is a flowchart illustrating another viewpoint rendering method provided in an embodiment of the present disclosure, as shown below. Figure 4 As shown, in this viewpoint rendering method, the fusion features are fused and the channel dimension is reduced to generate multiple viewpoint images. Then, the method includes the following steps S401 to S403.

[0095] S401, perform initial feature extraction on the viewpoint image and the corresponding label image to obtain the fifth initial feature and the sixth initial feature.

[0096] In step S401 above, a two-dimensional convolutional neural network is used to extract initial features from the viewpoint image and the corresponding label image (both with dimensions B×3×H×W) output by the generating network, to obtain the fifth initial feature and the sixth initial feature, both with dimensions B×C×H×W.

[0097] S402, perform width and height downsampling on the fifth and sixth initial features, and reduce the channel dimension to 1 to obtain the fourth and fifth initial dimensionality-reduced features.

[0098] In step S402 above, after passing through n sampling and dimensionality reduction modules with step size of 2 and step size of 1, the width and height dimensions are downsampled, with a size of B×C×(H / / n)×(W / / n). Then, a two-dimensional convolutional neural network is used to reduce the channel dimension to 1, generating the fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature, both with a size of B×1×(H / / n)×(W / / n).

[0099] S403, Based on the label image, distinguish between the fourth and fifth initial dimensionality reduction features.

[0100] In step S403 above, the fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature are judged based on the label image until the similarity between the fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature and the label image is close. If the similarity between the two differs greatly, steps S401 to S402 above are repeated.

[0101] In some embodiments, the number of groups of the fourth initial dimensionality reduction features and the fifth initial dimensionality reduction features is multiple; each group of the fourth initial dimensionality reduction features and the fifth initial dimensionality reduction features corresponds to the pixel blocks being identified in descending order.

[0102] In step S402 above, after downsampling in both width and height dimensions, the width and height dimensions are not 1. This means the image can be divided into multiple pixel blocks for separate discrimination, rather than the entire image being discriminated against. Furthermore, when discriminating pixel blocks, the number of sampling and dimensionality reduction modules can be changed, thus altering the size of the discriminated pixel blocks and achieving a progressive discrimination process.

[0103] In some embodiments, Figure 5 This is a flowchart illustrating another viewpoint rendering method provided in an embodiment of the present disclosure, as shown below. Figure 5 As shown, in this viewpoint rendering method, initial feature extraction is performed on the two-dimensional image and the corresponding depth image to obtain the first initial feature of the two-dimensional image and the second initial feature of the depth image. Before this, step S501 is also included, in which the two-dimensional image is input into the monocular depth estimation model to generate the corresponding depth image of the two-dimensional image.

[0104] In step S501 above, the original two-dimensional image F can be... 2D (The dimensions are 3×H×W, where 3 represents the RGB 3 channels, H represents the height of the video frame, and W represents the width of the video frame.) This data is input into an open-source monocular depth estimation algorithm (such as Monodepth, Monodepth2, etc.) to obtain the corresponding depth image F. Dep (Dimensions are 3×H×W, compared to the input two-dimensional image F) 2D (Keep dimensions consistent).

[0105] In some embodiments, such as Figure 5 As shown, in this viewpoint rendering method, the fusion features are fused and the channel dimension is reduced to generate multiple viewpoint images, and then steps S503 and S504 are included.

[0106] S503 stitches together multiple viewpoint images in a wide dimension to obtain a composite viewpoint image.

[0107] S504 interweaves the synthesized viewpoint images to generate a 3D image.

[0108] In steps S601 to S602 above, taking two viewpoints as an example, one viewpoint image F is... 2D,1 and the second viewpoint image F 2D,2 By stitching the images together along the width dimension, a synthetic 3D two-viewpoint image F was obtained. 3D The corresponding interleaving algorithm is used to process the 3D two-viewpoint image F. 3D By interweaving, it can be displayed on a 3D display device.

[0109] Secondly, embodiments of this disclosure provide a viewpoint rendering apparatus. Figure 6 This is a schematic diagram of the structure of a viewpoint rendering device provided in an embodiment of the present disclosure, as shown below. Figure 6 As shown, the viewpoint rendering device includes: a first initial feature extraction module 601, a stitching module 602, a fusion module 603, and an output module 604.

[0110] The first initial feature extraction module 601 is configured to perform initial feature extraction on the two-dimensional image and the corresponding depth image to obtain the first initial feature of the two-dimensional image and the second initial feature of the depth image.

[0111] Figure 7 This is a schematic diagram of the structure of a first feature extraction module in a viewpoint rendering apparatus provided in an embodiment of the present disclosure, as shown below. Figure 7 As shown, the first initial feature extraction module 601 includes a three-layer two-dimensional convolutional neural network. This network extracts the original two-dimensional image F... 2D and its corresponding depth image F Dep As input to the generative network, the input features (of size B×C×H×W) are first reduced in dimensionality by a factor of 2 using a two-layer convolutional neural network, resulting in features of size B×(C / / 2)×H×W. Then, a second two-layer convolutional neural network is used to efficiently extract these features, yielding features of size B×(C / / 2)×H×W. Finally, a third two-layer convolutional neural network is used to increase the dimensionality of the features by a factor of 2, resulting in the first initial feature F. 2D-init and the second initial feature F of the depth image Dep-init The dimensions are B×C×H×W. The first initial feature extraction module 601 is funnel-shaped in the channel dimension, which achieves a good balance between computational load and performance.

[0112] The splicing module 602 is configured to splice the first initial feature and the second initial feature in the channel dimension, and perform channel dimension reduction to obtain the first initial dimension-reduced feature.

[0113] Figure 8 This is a schematic diagram of the structure of a stitching module in a viewpoint rendering device provided in an embodiment of the present disclosure, as shown below. Figure 8 As shown, the splicing module 602 includes: a one-layer two-dimensional convolutional neural network. The first initial feature F... 2D-init Second initial feature F Dep-init The features are concatenated along the channel dimension to obtain features of size B×(2C)×H×W, and then passed through a two-dimensional convolutional neural network layer to reduce the channel dimension, resulting in the first initial dimensionality-reduced feature F. init-down The dimensions are B×C×H×W.

[0114] The fusion module 603 is configured to distort and repair the first initial dimensionality reduction features multiple times to obtain fused features.

[0115] Figure 9 This is a schematic diagram of the structure of a fusion module in a viewpoint rendering apparatus provided in an embodiment of the present disclosure, as shown below. Figure 9As shown, the fusion module 603 includes a three-layer two-dimensional convolutional neural network. The difference between the fusion module 603 and the first initial feature extraction module 601 is that the first and third two-dimensional convolutional neural networks use residual connections, which is beneficial for network optimization and convergence. The first initial feature F... 2D-init Second initial feature F Dep-init The features are concatenated along the channel dimension to obtain features of size B×(2C)×H×W, and then passed through a two-dimensional convolutional neural network layer to reduce the channel dimension, resulting in the first initial dimensionality-reduced feature F. init-down The dimensions are B×C×H×W.

[0116] The output module 604 is configured to fuse the fused features and reduce the channel dimension to generate multiple viewpoint images.

[0117] Figure 10 This is a schematic diagram of the structure of an output module in a viewpoint rendering apparatus provided in an embodiment of the present disclosure, as shown below. Figure 10 As shown, the output module 604 includes a three-layer two-dimensional convolutional neural network. The three-layer two-dimensional convolutional neural network gradually reduces the channel dimension to output the final color three-channel viewpoint image F. out The dimensions are B×3×H×W.

[0118] In some embodiments, Figure 11 This is a schematic diagram of another viewpoint rendering apparatus provided in an embodiment of the present disclosure, as shown below. Figure 11 As shown, the viewpoint rendering device also includes: a second initial feature extraction module 111, a first downsampling and dimensionality reduction module 112, and a first discrimination module 113.

[0119] The second initial feature extraction module 111 is configured to perform initial feature extraction on the viewpoint image and the corresponding label image to obtain the third initial feature and the fourth initial feature.

[0120] like Figure 11 As shown, the second initial feature extraction module 111 includes: a one-layer two-dimensional convolutional neural network; the one-layer two-dimensional convolutional neural network is used to extract initial features from the viewpoint image and the corresponding label image (both with size B×3×H×W) output by the generator network, to obtain the third initial feature and the fourth initial feature, both with size B×C×H×W.

[0121] The first downsampling and dimensionality reduction module 112 is configured to downsample the third and fourth initial features in width and height dimensions, and reduce the channel dimension to 1 to obtain the second and third initial dimensionality reduction features.

[0122] like Figure 11As shown, the first downsampling and dimensionality reduction module 112 includes: a multi-layer two-dimensional convolutional neural network with a stride of 2 and a stride of 1, and a single-layer binary adaptive mean pooling layer (AdaptiveAvgPool2d); after passing through n two-dimensional convolutional neural networks with strides of 2 and 1, downsampling is performed on the width and height dimensions; then, the two dimensions of width and height are reduced to 1 using a single-layer binary adaptive mean pooling layer, with a size of B×C×1×1; then, the channel dimension is reduced to 1 using a single-layer two-dimensional convolutional neural network, generating a second initial dimensionality reduction feature and a third initial dimensionality reduction feature, both with a size of B×1×1×1.

[0123] The first discrimination module 113 is configured to discriminate the second initial dimensionality reduction feature and the third initial dimensionality reduction feature based on the label image.

[0124] like Figure 11 As shown, the first discrimination module 113 includes: a one-layer two-dimensional convolutional neural network; using the one-layer two-dimensional convolutional neural network to discriminate the second initial dimensionality reduction features and the third initial dimensionality reduction features based on the label image, until the similarity between the second initial dimensionality reduction features and the third initial dimensionality reduction features and the label image is close. If the similarity between the two differs greatly, the above process is repeated.

[0125] In some embodiments, Figure 12 This is a schematic diagram of another viewpoint rendering apparatus provided in an embodiment of the present disclosure, as shown below. Figure 12 As shown, the viewpoint rendering device also includes: a third initial feature extraction module 121, a second downsampling and dimensionality reduction module 122, and a second discrimination module 123.

[0126] The third initial feature extraction module 121 is configured to perform initial feature extraction on the viewpoint image and the corresponding label image to obtain the fifth initial feature and the sixth initial feature.

[0127] like Figure 12 As shown, the third initial feature extraction module 121 includes a one-layer two-dimensional convolutional neural network. The one-layer two-dimensional convolutional neural network is used to extract initial features from the viewpoint image and the corresponding label image (both with dimensions B×3×H×W) output by the generator network, resulting in five initial features and a sixth initial feature, both with dimensions B×C×H×W.

[0128] The second downsampling and dimensionality reduction module 122 is configured to downsample the fifth and sixth initial features in both width and height dimensions, and reduce the channel dimension to 1 to obtain the fourth and fifth initial dimensionality reduction features.

[0129] like Figure 12As shown, the second downsampling and dimensionality reduction module 122 includes: a multi-layer two-dimensional convolutional neural network with a length of 2 and a stride of 1, and a single-layer two-dimensional convolutional neural network. After passing through n sampling and dimensionality reduction modules with strides of 2 and 1, downsampling is performed on the width and height dimensions, with a size of B×C×(H / / n)×(W / / n). Then, a single-layer two-dimensional convolutional neural network is used to reduce the channel dimension to 1, generating the fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature, both with a size of B×1×(H / / n)×(W / / n).

[0130] The second discrimination module 123 is configured to discriminate the fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature based on the label image.

[0131] The second discrimination module 123 includes a one-layer two-dimensional convolutional neural network. Based on the label image, it discriminates the fourth and fifth initial dimensionality reduction features until the similarity between the fourth and fifth initial dimensionality reduction features and the label image is close. If the similarity differs significantly, the above process is repeated.

[0132] In some embodiments, the viewpoint rendering apparatus further includes a depth image generation module (not shown in the figure), configured to input a two-dimensional image into a monocular depth estimation model to generate a depth image corresponding to the two-dimensional image.

[0133] The depth image generation module can generate the original two-dimensional image F 2D (The dimensions are 3×H×W, where 3 represents the RGB 3 channels, H represents the height of the video frame, and W represents the width of the video frame.) This data is input into an open-source monocular depth estimation algorithm (such as Monodepth, Monodepth2, etc.) to obtain the corresponding depth image F. Dep (The dimensions are 3×H×W, and the input two-dimensional image F) 2D (Keep dimensions consistent).

[0134] In some embodiments, the viewpoint rendering apparatus further includes: a compositing module (not shown in the figure), configured to stitch together multiple viewpoint images in a wide dimension to obtain a composite viewpoint image; and an interlacing module (not shown in the figure), configured to interlace the composite viewpoint image to generate a three-dimensional image.

[0135] Taking a two-viewpoint example, let's consider one viewpoint image F. 2D,1 and the second viewpoint image F 2D,2 The compositing module can stitch together two viewpoint images in a wide dimension to obtain a synthesized 3D two-viewpoint image F. 3D The interleaving module can utilize the corresponding interleaving algorithm to process 3D two-viewpoint images F... 3D By interweaving, it can be displayed on a 3D display device.

[0136] Thirdly, embodiments of this disclosure provide an electronic device, Figure 13 This is a schematic diagram of the structure of an electronic device provided in some embodiments of this disclosure, such as... Figure 13 As shown, the electronic device includes: one or more processors 131; a memory 132 storing one or more programs, which, when executed by one or more processors, cause one or more processors to implement the viewpoint rendering method provided in any of the above embodiments; and one or more I / O interfaces 133 connected between the processors and the memory, configured to enable information interaction between the processors and the memory.

[0137] The processor 131 is a device with data processing capabilities, including but not limited to a central processing unit (CPU); the memory 132 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory (FLASH); the I / O interface (read / write interface) 133 is connected between the processor 131 and the memory 132, enabling information exchange between the processor 131 and the memory 132, including but not limited to a data bus (Bus).

[0138] In some embodiments, the processor 131, memory 132, and I / O interface 133 are interconnected via a bus, and thus connected to other components of the computing device.

[0139] Fourthly, this embodiment provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the viewpoint rendering method provided in any of the above embodiments.

[0140] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0141] It is understood that the above embodiments are merely exemplary embodiments used to illustrate the principles of this disclosure, and this disclosure is not limited thereto. For those skilled in the art, various modifications and improvements can be made without departing from the spirit and substance of this disclosure, and these modifications and improvements are also considered to be within the scope of protection of this disclosure.

Claims

1. A viewpoint rendering method, wherein, The viewpoint rendering method includes: Initial feature extraction is performed on the two-dimensional image and the corresponding depth image to obtain the first initial feature of the two-dimensional image and the second initial feature of the depth image; The first initial feature and the second initial feature are concatenated along the channel dimension, and then the channel dimension is reduced to obtain the first initial dimensionality-reduced feature; The first initial dimensionality reduction feature is subjected to multiple image distortions and repairs to obtain fused features; The fusion features are fused and the channel dimension is reduced to generate multiple viewpoint images; The process of fusing the fused features and reducing the channel dimension to generate multiple viewpoint images further includes: Initial feature extraction is performed on the viewpoint image and the corresponding label image to obtain the third initial feature and the fourth initial feature; The third initial feature and the fourth initial feature are downsampled in width and height dimensions, and the width and height dimensions are reduced to 1, and the channel dimension is reduced to 1, to obtain the second initial dimensionality reduction feature and the third initial dimensionality reduction feature; Based on the labeled image, the second initial dimensionality reduction feature and the third initial dimensionality reduction feature are distinguished; Alternatively, the process of fusing the fusion features and reducing the channel dimension to generate multiple viewpoint images may further include: Initial feature extraction is performed on the viewpoint image and the corresponding label image to obtain the fifth initial feature and the sixth initial feature; The fifth and sixth initial features are downsampled in width and height dimensions, and the channel dimension is reduced to 1 to obtain the fourth and fifth initial dimensionality-reduced features. Based on the labeled image, the fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature are discriminated.

2. The viewpoint rendering method according to claim 1, wherein, The fourth and fifth initial dimensionality reduction features consist of multiple groups; Each set of the fourth and fifth initial dimensionality reduction features corresponds to the pixel blocks being identified in descending order.

3. The viewpoint rendering method according to claim 1, wherein, Before performing initial feature extraction on the two-dimensional image and the corresponding depth image to obtain the first initial feature of the two-dimensional image and the second initial feature of the depth image, the process further includes: The two-dimensional image is input into a monocular depth estimation model to generate the depth image corresponding to the two-dimensional image.

4. The viewpoint rendering method according to claim 1, wherein, The fusion features are fused and the channel dimension is reduced to generate multiple viewpoint images, and then the process further includes: Multiple viewpoint images are stitched together in a wide dimension to obtain a composite viewpoint image; The synthesized viewpoint images are interwoven to generate a three-dimensional image.

5. A viewpoint rendering apparatus, wherein, The viewpoint rendering device includes: The first initial feature extraction module is configured to perform initial feature extraction on a two-dimensional image and a depth image corresponding to the two-dimensional image to obtain a first initial feature of the two-dimensional image and a second initial feature of the depth image. The splicing module is configured to splice the first initial feature and the second initial feature in the channel dimension, and perform channel dimension reduction to obtain the first initial dimension-reduced feature; The fusion module is configured to perform multiple image distortions and repairs on the first initial dimensionality reduction features to obtain fused features; The output module is configured to fuse the fused features and reduce the channel dimension to generate multiple viewpoint images; The viewpoint rendering device further includes: The second initial feature extraction module is configured to perform initial feature extraction on the viewpoint image and the corresponding label image to obtain the third initial feature and the fourth initial feature. The first downsampling and dimensionality reduction module is configured to downsample the width and height dimensions of the third initial feature and the fourth initial feature, and reduce the channel dimension to 1 to obtain the second initial dimensionality reduction feature and the third initial dimensionality reduction feature. The first discrimination module is configured to discriminate the second initial dimensionality reduction feature and the third initial dimensionality reduction feature based on the label image; Alternatively, the viewpoint rendering apparatus may further include: The third initial feature extraction module is configured to perform initial feature extraction on the viewpoint image and the corresponding label image to obtain the fifth initial feature and the sixth initial feature. The second downsampling and dimensionality reduction module is configured to downsample the fifth initial feature and the sixth initial feature in both width and height dimensions, and reduce the channel dimension to 1 to obtain the fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature. The second discrimination module is configured to discriminate the fourth initial dimensionality reduction feature and the fifth initial dimensionality reduction feature based on the label image.

6. The viewpoint rendering apparatus according to claim 5, wherein, The viewpoint rendering device further includes: The depth image generation module is configured to input the two-dimensional image into a monocular depth estimation model to generate the depth image corresponding to the two-dimensional image.

7. The viewpoint rendering apparatus according to claim 5, wherein, The viewpoint rendering device further includes: The compositing module is configured to stitch together multiple viewpoint images in a wide dimension to obtain a composite viewpoint image; An interlacing module is configured to interlace the synthesized viewpoint image to generate a three-dimensional image.

8. The viewpoint rendering apparatus according to claim 5, wherein, The first initial feature extraction module includes: a three-layer two-dimensional convolutional neural network; The splicing module includes: a one-layer two-dimensional convolutional neural network; The fusion module includes: a three-layer two-dimensional convolutional neural network; the first layer and the third layer are connected by a residual connection; The output module includes a three-layer two-dimensional convolutional neural network.

9. The viewpoint rendering apparatus according to claim 5, wherein, The second initial feature extraction module includes: a one-layer two-dimensional convolutional neural network; The first downsampling and dimensionality reduction module includes: a multi-layer two-dimensional convolutional neural network with a length of 2 and a stride of 1, and a single-layer binary adaptive mean convergence layer; The first discrimination module includes: a one-layer two-dimensional convolutional neural network.

10. The viewpoint rendering apparatus according to claim 5, wherein, The third initial feature extraction module includes: a one-layer two-dimensional convolutional neural network; The second downsampling and dimensionality reduction module includes: a multi-layer two-dimensional convolutional neural network with a length of 2 and a stride of 1; The second discrimination module includes: a one-layer two-dimensional convolutional neural network.

11. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the viewpoint rendering method as described in any one of claims 1-4.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the viewpoint rendering method as described in any one of claims 1-4.