Depth map optimization model training method and device, equipment and storage medium
Patent Information
- Application Number
- CN202380010344.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-28
- Publication Date
- 2025-06-06
AI Technical Summary
The existing depth map is determined by a single-view image, resulting in low estimation accuracy of depth information and cannot meet the high accuracy requirements of image processing services for depth information.
The depth map optimization model training method is adopted. By obtaining training samples of the same image at different viewpoints, the depth map optimization model and viewpoint rendering model are trained, and the depth map is optimized to improve the accuracy of depth information.
By learning depth information from different viewpoints, the depth map optimization model can more accurately estimate the depth information and improve the depth map quality in the image processing business.
Smart Images

Figure CN120112946A_ABST
Abstract
Description
A deep graph optimization model training method, device, equipment and storage medium Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a depth map optimization model training method, device, equipment and storage medium. Background Art
[0002] Many current image processing applications require depth maps. For example, depth maps can be used when generating stereoscopic 3D images, and when rendering images, depth maps can be used to assist with rendering, specifically to highlight foreground areas.
[0003] However, current depth maps are often determined based on single-viewpoint images using algorithms, such as monocular depth estimation algorithms. Depth maps determined in this way have low depth estimation accuracy.
[0004] Summary of the Invention
[0005] The present invention provides a depth map optimization model training method, device, equipment and storage medium to address the deficiencies in the related art.
[0006] According to a first aspect of an embodiment of the present invention, a depth map optimization model training method is provided, comprising:
[0007] Obtain depth map optimization model and viewpoint rendering model;
[0008] Acquire a training sample set; wherein the sample features of any training sample in the training sample set include: a first view image and a first depth image; and the sample label of any training sample includes: a second view image corresponding to the first view image;
[0009] The first depth map includes depth information of each pixel of the first view image; the first view image and the second view image are the same image at different viewpoints;
[0010] The depth map optimization model is used to optimize the input first depth map according to the input first viewpoint map and the first depth map to obtain an optimized depth map;
[0011] The viewpoint rendering model is used to predict the corresponding second viewpoint image according to the optimized depth map obtained by the depth map optimization model and the first viewpoint image to obtain a predicted viewpoint image;
[0012] Based on the training sample set, training the depth map optimization model and the viewpoint rendering model;
[0013] During the training process, parameters of the depth map optimization model are updated according to the loss between the second view map and the predicted view map.
[0014] Optionally, the method further comprises: determining a foreground area for the first view image;
[0015] The updating of the parameters of the depth map optimization model according to the loss between the second view map and the predicted view map includes:
[0016] updating parameters of the depth map optimization model according to the loss between the second view image and the predicted view image on the same target area;
[0017] The position of the target area in the second view image and the predicted view image is the same as the position of the determined foreground area in the first view image.
[0018] Optionally, determining the foreground area for the first view image includes any one of the following:
[0019] Performing foreground object detection on the first view image to determine a foreground area;
[0020] Performing foreground target detection on the first view image to obtain a foreground region frame, then performing target segmentation on image content in the foreground region frame, and determining the segmented image content as the foreground region;
[0021] Foreground object segmentation is performed on the first view image, and the segmented image content is determined as the foreground area.
[0022] Optionally, the depth map optimization model includes: a view map feature extraction layer and a depth map feature extraction layer;
[0023] The viewpoint map feature extraction layer is used to: extract a first viewpoint feature map for the input first viewpoint map; and then increase the dimension of the obtained first viewpoint feature map to obtain a second viewpoint feature map;
[0024] The depth map feature extraction layer is used to: extract a first depth feature map from an input first depth map; and then increase the dimension of the obtained first depth feature map to obtain a second depth feature map.
[0025] Optionally, the depth map optimization model includes: a feature fusion layer;
[0026] The depth map optimization model is used to synthesize the view feature map extracted based on the first view map and the depth feature map extracted based on the first depth map to obtain a comprehensive feature map;
[0027] The feature fusion layer is used to: reduce the dimension of the input feature map to obtain a reduced-dimensional feature map; then perform feature fusion on the reduced-dimensional feature map to obtain a fused feature map; then increase the dimension of the fused feature map to obtain an increased-dimensional feature map, and output the sum of the increased-dimensional feature map and the input feature map;
[0028] The input feature map includes any one of the following: the comprehensive feature map, a feature map extracted based on the comprehensive feature map, and a feature map output by other feature fusion layers.
[0029] Optionally, the first view image, the first depth image and the second view image are acquired based on the same three-dimensional image.
[0030] Optionally, a method for obtaining any training sample in the training sample set includes:
[0031] Acquire a three-dimensional image of the target;
[0032] Acquire images of the target three-dimensional image at two different viewpoints, and determine them as a first viewpoint image and a second viewpoint image corresponding to the first viewpoint image;
[0033] Based on the depth information of the target three-dimensional image, obtaining the depth information of each pixel in the first view image to determine a first depth map; or, based on the first view image, determining the first depth map according to a preset depth estimation algorithm;
[0034] The determined first view image and the first depth image are determined as sample features, and the determined second view image is determined as a sample label to obtain a training sample.
[0035] Optionally, the viewpoint rendering model is pre-trained; the method further comprises: freezing parameters of the viewpoint rendering model during the training of the depth map optimization model and the viewpoint rendering model.
[0036] Optionally, the method further includes:
[0037] During the training process, the parameters of the viewpoint rendering model are updated according to the loss between the second viewpoint image and the predicted viewpoint image.
[0038] Optionally, the viewpoint rendering model includes: a first dimensionality-raising fusion layer;
[0039] The viewpoint rendering model is used to: extract a third viewpoint feature map of a first size and a fourth viewpoint feature map of a second size for the first viewpoint map; wherein the first size is smaller than the second size; perform feature fusion on the third viewpoint feature map and the optimized feature map extracted based on the optimized depth map to obtain an optimized fused feature map of the first size;
[0040] The first dimensionality-raising fusion layer is used to: increase the dimension of the optimized fusion feature map to obtain a dimensionality-raising fusion feature map of the second size; perform feature fusion on the dimensionality-raising fusion feature map and the fourth viewpoint feature map, and output a fusion result of the second size.
[0041] Optionally, the viewpoint rendering model further includes: a second dimensionality-raising fusion layer;
[0042] The viewpoint rendering model is further used to: extract a fifth viewpoint feature map of a third size for the first viewpoint map; wherein the second size is smaller than the third size;
[0043] The second dimensionality-raising fusion layer is used to: increase the dimension of the fusion result output by the first dimensionality-raising fusion layer to obtain a feature map to be fused of a third size; perform feature fusion on the feature map to be fused and the fifth viewpoint feature map, and output a fusion result of a third size.
[0044] According to a second aspect of an embodiment of the present invention, a depth map optimization model training device is provided, comprising:
[0045] An acquisition unit is configured to acquire a depth map optimization model and a viewpoint rendering model; acquire a training sample set; wherein the sample features of any training sample in the training sample set include: a first viewpoint map and a first depth map; and the sample label of any training sample includes: a second viewpoint map corresponding to the first viewpoint map;
[0046] The first depth map includes depth information of each pixel of the first view image; the first view image and the second view image are the same image at different viewpoints;
[0047] The depth map optimization model is used to optimize the input first depth map according to the input first viewpoint map and the first depth map to obtain an optimized depth map;
[0048] The viewpoint rendering model is used to predict the corresponding second viewpoint image according to the optimized depth map obtained by the depth map optimization model and the first viewpoint image to obtain a predicted viewpoint image;
[0049] A training unit, configured to train the depth map optimization model and the viewpoint rendering model based on the training sample set;
[0050] During the training process, parameters of the depth map optimization model are updated according to the loss between the second view map and the predicted view map.
[0051] Optionally, the training unit is further configured to: determine a foreground area for the first view image;
[0052] The training unit is used to:
[0053] updating parameters of the depth map optimization model according to the loss between the second view image and the predicted view image on the same target area;
[0054] The position of the target area in the second view image and the predicted view image is the same as the position of the determined foreground area in the first view image.
[0055] Optionally, the training unit is configured to perform any one of the following to determine a foreground area for the first view image:
[0056] Performing foreground object detection on the first view image to determine a foreground area;
[0057] Performing foreground target detection on the first view image to obtain a foreground region frame, then performing target segmentation on image content in the foreground region frame, and determining the segmented image content as the foreground region;
[0058] Foreground object segmentation is performed on the first view image, and the segmented image content is determined as the foreground area.
[0059] Optionally, the depth map optimization model includes: a view map feature extraction layer and a depth map feature extraction layer;
[0060] The viewpoint map feature extraction layer is used to: extract a first viewpoint feature map for the input first viewpoint map; and then increase the dimension of the obtained first viewpoint feature map to obtain a second viewpoint feature map;
[0061] The depth map feature extraction layer is used to: extract a first depth feature map from an input first depth map; and then increase the dimension of the obtained first depth feature map to obtain a second depth feature map.
[0062] Optionally, the depth map optimization model includes: a feature fusion layer;
[0063] The depth map optimization model is used to synthesize the view feature map extracted based on the first view map and the depth feature map extracted based on the first depth map to obtain a comprehensive feature map;
[0064] The feature fusion layer is used to: reduce the dimension of the input feature map to obtain a reduced-dimensional feature map; then perform feature fusion on the reduced-dimensional feature map to obtain a fused feature map; then increase the dimension of the fused feature map to obtain an increased-dimensional feature map, and output the sum of the increased-dimensional feature map and the input feature map;
[0065] The input feature map includes any one of the following: the comprehensive feature map, a feature map extracted based on the comprehensive feature map, and a feature map output by other feature fusion layers.
[0066] Optionally, the first view image, the first depth image and the second view image are acquired based on the same three-dimensional image.
[0067] Optionally, a method for obtaining any training sample in the training sample set includes:
[0068] Acquire a three-dimensional image of the target;
[0069] Acquire images of the target three-dimensional image at two different viewpoints, and determine them as a first viewpoint image and a second viewpoint image corresponding to the first viewpoint image;
[0070] Based on the depth information of the target three-dimensional image, obtaining the depth information of each pixel in the first view image to determine a first depth map; or, based on the first view image, determining the first depth map according to a preset depth estimation algorithm;
[0071] The determined first view image and the first depth image are determined as sample features, and the determined second view image is determined as a sample label to obtain a training sample.
[0072] Optionally, the viewpoint rendering model is pre-trained; and the training unit is further configured to: freeze parameters of the viewpoint rendering model during the process of training the depth map optimization model and the viewpoint rendering model.
[0073] Optionally, the training unit is further configured to:
[0074] During the training process, the parameters of the viewpoint rendering model are updated according to the loss between the second viewpoint image and the predicted viewpoint image.
[0075] Optionally, the viewpoint rendering model includes: a first dimensionality-raising fusion layer;
[0076] The viewpoint rendering model is used to: extract a third viewpoint feature map of a first size and a fourth viewpoint feature map of a second size for the first viewpoint map; wherein the first size is smaller than the second size; perform feature fusion on the third viewpoint feature map and the optimized feature map extracted based on the optimized depth map to obtain an optimized fused feature map of the first size;
[0077] The first dimensionality-raising fusion layer is used to: increase the dimension of the optimized fusion feature map to obtain a dimensionality-raising fusion feature map of the second size; perform feature fusion on the dimensionality-raising fusion feature map and the fourth viewpoint feature map, and output a fusion result of the second size.
[0078] Optionally, the viewpoint rendering model further includes: a second dimensionality-raising fusion layer;
[0079] The viewpoint rendering model is further used to: extract a fifth viewpoint feature map of a third size for the first viewpoint map; wherein the second size is smaller than the third size;
[0080] The second dimensionality-raising fusion layer is used to: increase the dimension of the fusion result output by the first dimensionality-raising fusion layer to obtain a feature map to be fused of a third size; perform feature fusion on the feature map to be fused and the fifth viewpoint feature map, and output a fusion result of a third size.
[0081] According to the above embodiments, by using the same image from different viewpoints to train the depth map optimization model, the depth map optimization model can learn to mine more accurate depth information from different viewpoints during the training process, thereby facilitating the improvement of the accuracy of the depth information in the depth map.
[0082] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0084] FIG1 is a schematic diagram of a flow chart of a depth map optimization model training method according to an embodiment of the present invention;
[0085] FIG2 is a schematic structural diagram of a depth map optimization model according to an embodiment of the present invention;
[0086] FIG3 is a schematic structural diagram of a viewpoint rendering model according to an embodiment of the present invention;
[0087] FIG4 is a schematic diagram showing a flow chart of a depth map optimization method according to an embodiment of the present invention;
[0088] FIG5 is a schematic flow chart of a method for generating a three-dimensional image according to an embodiment of the present invention;
[0089] FIG6 is a schematic flow chart of a 3D video conversion method according to an embodiment of the present invention;
[0090] FIG7 is a schematic diagram showing the principle of depth map optimization model training according to an embodiment of the present invention;
[0091] FIG8 is a schematic diagram showing an image optimization result according to an embodiment of the present invention;
[0092] FIG9 is a schematic structural diagram of a depth map optimization model training device according to an embodiment of the present invention;
[0093] FIG10 is a schematic structural diagram of a depth map optimization device according to an embodiment of the present invention;
[0094] FIG11 is a schematic structural diagram of a three-dimensional image generating device according to an embodiment of the present invention;
[0095] FIG12 is a schematic structural diagram of a 3D video conversion device according to an embodiment of the present invention;
[0096] FIG13 is a schematic diagram of the hardware structure of a computer device configured with the method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0097] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.
[0098] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0099] Many current image processing applications require depth maps. For example, depth maps can be used when generating stereoscopic 3D images, and when rendering images, depth maps can be used to assist with rendering, specifically to highlight foreground areas.
[0100] However, current depth maps are often determined based on single-viewpoint images using algorithms, such as monocular depth estimation algorithms. Depth maps determined in this way have low depth estimation accuracy.
[0101] In order to solve the above problems, an embodiment of the present invention discloses a depth map optimization model training method.
[0102] In this method, a machine learning approach is introduced to train a depth map optimization model, so that the trained depth map optimization model can be used to optimize the depth map, thereby improving the accuracy of the depth information in the depth map.
[0103] In order to improve the accuracy of depth information in the depth map, this method further introduces the same image under different viewpoints. Since the same image under different viewpoints can be used to determine the three-dimensional image, it is convenient to learn and mine more accurate depth information.
[0104] The same image from different viewpoints, for example, the left and right viewpoints of the same scene, or two images captured simultaneously by a binocular camera.
[0105] The same image from different viewpoints can be used to easily mine information about three-dimensional scenes, thereby facilitating the extraction of more accurate depth information.
[0106] To train models based on images from different viewpoints, this method uses an image and depth map from one viewpoint to predict the same image from another viewpoint. This prediction between images from different viewpoints facilitates learning and mining more accurate depth information.
[0107] The specific model structure can be designed into a two-stage model structure by dividing the tasks, which can specifically include a depth map optimization model and a viewpoint rendering model.
[0108] Among them, the depth map optimization model can be used to further optimize the depth map for an image and a depth map under one viewpoint, while the viewpoint rendering model can predict the same image under another viewpoint based on the optimized depth map and the image under one viewpoint.
[0109] By training the above model structure, the depth map optimization model can be used to learn and mine the depth information of the same image from different viewpoints, thereby improving the accuracy of the depth information.
[0110] Moreover, by dividing the model structure into two stages by tasks, it is convenient to check the effect of model training.
[0111] Moreover, this model structure can realize unsupervised training of the depth map optimization model. There is no need to obtain sample labels of the optimized depth map. The model can be trained using the same image from different viewpoints, which cleverly reduces the difficulty of obtaining training samples.
[0112] The trained depth map optimization model can be used to optimize the depth map, thereby improving the accuracy of the depth information in the depth map.
[0113] The following is a detailed explanation of a depth map optimization model training method provided by an embodiment of the present invention.
[0114] As shown in FIG1 , FIG1 is a flow chart of a depth map optimization model training method according to an embodiment of the present invention.
[0115] The embodiments of the present invention do not limit the execution subject of the method process. Alternatively, the execution subject may be any computing device, for example, a server or client for optimizing depth maps, or a server or client for providing image processing functions.
[0116] The method may include the following steps.
[0117] S101: Obtain a depth map optimization model and a viewpoint rendering model.
[0118] S102: Obtain a training sample set.
[0119] The sample features of any training sample in the training sample set may include: a first view image and a first depth image; and the sample label of any training sample may include: a second view image corresponding to the first view image.
[0120] The first depth map may include depth information of each pixel of the first view image; the first view image and the corresponding second view image may be the same image at different viewpoints.
[0121] The depth map optimization model can be used to optimize the input first depth map according to the input first viewpoint map and the first depth map to obtain an optimized depth map.
[0122] The viewpoint rendering model can be used to predict the corresponding second viewpoint image based on the optimized depth map obtained by the depth map optimization model and the first viewpoint image to obtain a predicted viewpoint image.
[0123] S103: Based on the training sample set, train the depth map optimization model and the viewpoint rendering model; during the training process, update the parameters of the depth map optimization model according to the loss between the second viewpoint map and the predicted viewpoint map.
[0124] The above method process uses the same image from different viewpoints to train the depth map optimization model, which facilitates the depth map optimization model to learn to mine more accurate depth information from different viewpoints during the training process, thereby facilitating the improvement of the accuracy of the depth information in the depth map.
[0125] In addition, the above method process can also realize unsupervised training of the depth map optimization model. There is no need to obtain sample labels of the optimized depth map. The model can be trained using the same image from different viewpoints, which cleverly reduces the difficulty of obtaining training samples.
[0126] The following is a detailed explanation of each aspect.
[0127] 1. About the training sample set.
[0128] The sample features of any training sample in the training sample set may include: a first view image and a first depth image; the sample label of any training sample may include: a second view image corresponding to the first view image.
[0129] For ease of understanding, the first view image, the first depth image, and the second view image in any training sample in the training sample set are first explained.
[0130] Optionally, the first view image may be a single-view image, and the first depth map may include depth information of each pixel at the viewpoint where the first view image is located.
[0131] Optionally, the first view image and the corresponding second view image may be the same image at different viewpoints, specifically the same three-dimensional image at different viewpoints.
[0132] For example, for the same 3D image, the left view image may be determined as the first view image, and the right view image may be determined as the corresponding second view image.
[0133] For another example, for the same scene, the captured left view image can be determined as the first view image, and the right view image can be determined as the corresponding second view image. Of course, for the same scene, the captured right view image can also be determined as the first view image, and the left view image can be determined as the corresponding second view image.
[0134] Optionally, the second view image may be used to construct a three-dimensional image together with the corresponding first view image.
[0135] Optionally, the first view image and the corresponding second view image may be different images of the same scene captured simultaneously from different viewpoints.
[0136] Optionally, the first view image and the corresponding second view image may be two different images captured simultaneously by a binocular camera; of course, they may also be two different images captured simultaneously by a multi-view camera.
[0137] The above explains the features and labels of a single training sample. It is understandable that other training samples in the training sample set can refer to the above explanation.
[0138] Optionally, the sample features of each training sample in the training sample set include a first view image and a first depth image, and the sample label of the training sample includes a second view image corresponding to the first view image.
[0139] In each training sample, the first view image in the sample feature and the second view image in the sample label can be the same image under different viewpoints.
[0140] The first viewpoint images between different training samples may be different images, and the second viewpoint images between different training samples may also be different images.
[0141] This method does not limit the source of the training sample set.
[0142] Optionally, the first view image, the first depth image, and the corresponding second view image may be directly obtained and combined into a training sample.
[0143] Optionally, a first view image, a first depth image, and a corresponding second view image may be generated based on an image, and then combined into a training sample.
[0144] Optionally, the first view image, the first depth image, and the second view image may be acquired based on the same three-dimensional image.
[0145] Different single-viewpoint images under different viewpoints can be easily extracted for the same three-dimensional image, which improves the convenience of sample acquisition. The first depth map can be determined based on the depth information in the same three-dimensional image, specifically, it can be determined based on the depth information of each pixel under the viewpoint corresponding to the first viewpoint map.
[0146] This embodiment can improve the convenience of sample acquisition based on the same three-dimensional image.
[0147] The method flow is not limited to the way of obtaining samples based on three-dimensional images.
[0148] For ease of understanding, optionally, a method for obtaining any training sample in the training sample set may include: obtaining a target three-dimensional image; obtaining images of the target three-dimensional image under two different viewpoints, respectively determining them as a first viewpoint image, and a second viewpoint image corresponding to the first viewpoint image; based on the depth information of the target three-dimensional image, obtaining the depth information of each pixel in the first viewpoint image, and determining the first depth map; determining the determined first viewpoint image and first depth map as sample features, and determining the determined second viewpoint image as a sample label, to obtain a training sample.
[0149] Optionally, the method for obtaining any training sample in the training sample set may include: obtaining a target three-dimensional image; obtaining images of the target three-dimensional image under two different viewpoints, respectively determining them as a first viewpoint image and a second viewpoint image corresponding to the first viewpoint image; based on the first viewpoint image, determining a first depth map according to a preset depth estimation algorithm; determining the determined first viewpoint image and first depth map as sample features, and determining the determined second viewpoint image as a sample label to obtain a training sample.
[0150] Optionally, the preset depth estimation algorithm may be a monocular depth estimation algorithm.
[0151] This embodiment clearly discloses a specific method for acquiring samples based on three-dimensional images, which can improve the efficiency of sample acquisition.
[0152] 2. About the depth map optimization model.
[0153] The depth map optimization model can be used to optimize the input first depth map according to the input first viewpoint map and the first depth map to obtain an optimized depth map.
[0154] This method does not limit the specific structure of the depth map optimization model.
[0155] Optionally, the deep graph optimization model can adopt structures such as graph neural networks or graph convolutional networks.
[0156] 1. About feature extraction.
[0157] Optionally, for the input first view image and first depth image, a feature map may be extracted for subsequent optimization of the depth map.
[0158] Optionally, the depth map optimization model may include: a view map feature extraction layer.
[0159] Optionally, the viewpoint map feature extraction layer can be used to: extract a first viewpoint feature map from the input first viewpoint map; and then increase the dimension of the obtained first viewpoint feature map to obtain a second viewpoint feature map.
[0160] This embodiment does not limit the method of extracting the first viewpoint feature map. Optionally, the feature map can be extracted based on graph convolution or graph neural network.
[0161] This embodiment does not limit the method of increasing the dimension. Optionally, at least one of the following may be increased: channel dimension, resolution dimension, width, height, etc.
[0162] Optionally, the depth map optimization model may include: a depth map feature extraction layer.
[0163] Optionally, the depth map feature extraction layer can be used to: extract a first depth feature map from the input first depth map; and then increase the dimension of the obtained first depth feature map to obtain a second depth feature map.
[0164] This embodiment does not limit the method of extracting the first depth feature map. Optionally, the feature map can be extracted based on graph convolution or graph neural network.
[0165] This embodiment does not limit the method of increasing the dimension. Optionally, at least one of the following may be increased: channel dimension, resolution dimension, width, height, etc. Specifically, the channel dimension may be increased.
[0166] This embodiment can extract features first and then increase the dimension of the feature map, so that low-dimensional features can be extracted during feature extraction while ensuring that the dimension of the feature map is high (high precision), thereby reducing the amount of calculation and improving calculation efficiency.
[0167] For ease of understanding, an example of a view map feature extraction layer or a depth map feature extraction layer is given below.
[0168] Among them, a one-layer two-dimensional graph convolutional neural network is first used to perform preliminary feature extraction on the input first view image (size 3×H×W, where H is the height of the first view image, W is the width of the first view image, and 3 is the three channels of RGB). The obtained feature size is B×(C / 2)×H×W (where B is the batch size and C is the number of channels processed).
[0169] Then, a one-layer two-dimensional graph convolutional neural network is used to fuse the features extracted in the previous layer, and the resulting feature size is B×(C / 2)×H×W.
[0170] Finally, a one-layer two-dimensional graph convolutional neural network is used to upgrade the channel dimension of the features extracted in the previous layer, and the final feature size is B×C×H×W.
[0171] The overall feature extraction layer is funnel-shaped in the channel dimension. This design can achieve a good balance between computational complexity and performance.
[0172] 2. About feature fusion.
[0173] Optionally, the depth map optimization model needs to output an optimized first depth map for the input first view map and first depth map, so that fusion between feature maps is required.
[0174] Optionally, the depth map optimization model can be used to fuse the first viewpoint feature map and the first depth feature map, and can also be used to fuse the second viewpoint feature map and the second depth feature map.
[0175] Optionally, the depth map optimization model may include: a feature fusion layer.
[0176] Optionally, the depth map optimization model can be used to synthesize the viewpoint feature map extracted based on the first viewpoint map and the depth feature map extracted based on the first depth map to obtain a comprehensive feature map.
[0177] There is no limitation on the method of extracting the viewpoint feature map or the method of extracting the depth feature map. The first viewpoint feature map or the second viewpoint feature map may be extracted, or the first depth feature map or the second depth feature map may be extracted.
[0178] The specific integration method is not limited and can be to perform feature fusion to obtain a comprehensive feature map, or to further extract features based on the fusion result after feature fusion. For example, graph convolution can be used to extract a comprehensive feature map based on the fusion result.
[0179] Optionally, the feature fusion layer can be used to: reduce the dimension of the input feature map to obtain a reduced-dimensionality feature map; then perform feature fusion on the reduced-dimensionality feature map to obtain a fused feature map; then increase the dimension of the fused feature map to obtain an increased-dimensionality feature map, and output the increased-dimensionality feature map.
[0180] Optionally, the feature fusion layer can be used to: reduce the dimension of the input feature map to obtain a reduced-dimensionality feature map; then perform feature fusion on the reduced-dimensionality feature map to obtain a fused feature map; then increase the dimension of the fused feature map to obtain an increased-dimensionality feature map, and output the sum of the increased-dimensionality feature map and the input feature map.
[0181] Optionally, the input feature map may include any one of the following: a comprehensive feature map, a feature map extracted based on the comprehensive feature map, and a feature map output by other feature fusion layers.
[0182] Among them, the feature fusion layer can fuse low-dimensional features by reducing the dimension, reduce the amount of calculation, and then increase the dimension of the fused feature map. In this way, low-dimensional features can be extracted during feature extraction while ensuring the high dimension (high precision) of the feature map, thereby reducing the amount of calculation and improving calculation efficiency.
[0183] In addition, by introducing a residual structure and outputting the sum of the increased-dimensional feature map and the input feature map, the learning and fitting ability of the feature fusion layer and the model training effect can be improved.
[0184] In addition, the deep graph optimization model can include one or more feature fusion layers, which can directly perform further fusion on the comprehensive feature map, or first extract features from the comprehensive feature map and then perform further fusion, or further fusion can be performed on the output of other feature fusion layers.
[0185] Each feature fusion layer can be connected in series.
[0186] This embodiment does not limit the method of extracting feature maps based on comprehensive feature maps. Optionally, feature maps can be extracted based on graph convolution or graph neural networks.
[0187] This embodiment does not limit the method of reducing or increasing the dimension. Optionally, at least one of the following may be reduced or increased: channel dimension, resolution dimension, width, height, etc. Specifically, the channel dimension may be reduced or increased.
[0188] For ease of understanding, an example of a feature fusion layer is given below.
[0189] First, a one-layer two-dimensional graph convolutional neural network is used to perform preliminary channel dimensionality reduction on the input features (size B×C×H×W), and the resulting feature size is B×(C / 2)×H×W.
[0190] Then, a one-layer two-dimensional graph convolutional neural network is used to fuse the reduced-dimensional features, and the resulting feature size is B×(C / 2)×H×W.
[0191] Finally, a one-layer two-dimensional graph convolutional neural network is used to upgrade the channel dimension of the fusion feature and add it to the features of the input feature fusion layer. The final feature size is B×C×H×W.
[0192] The feature fusion layer adds a residual connection, which is beneficial for network optimization and convergence. This design can achieve a good balance between computational complexity and performance.
[0193] 3. About output.
[0194] Optionally, the depth map optimization model may further include an output layer for outputting a predicted optimized depth map based on the feature map output by the previous layer. Specifically, the output layer may be configured to output a predicted optimized depth map based on the feature map output by the previous feature fusion layer.
[0195] This embodiment does not limit the structure of the output layer, which can be a graph convolutional network or a graph neural network.
[0196] For ease of understanding, an example of an output layer is given below.
[0197] First, a one-layer two-dimensional graph convolutional neural network is used to perform preliminary channel dimensionality reduction on the input features (size B×C×H×W), and the resulting feature size is B×(C / 2)×H×W.
[0198] Then, a one-layer two-dimensional graph convolutional neural network is used to reduce the channel dimension again, and the resulting feature size is B × (C / 4) × H × W.
[0199] Finally, a one-layer two-dimensional graph convolutional neural network is used to reconstruct the features, and the final output image has a size of B×3×H×W.
[0200] The output layer gradually reduces the channel dimension through three layers of two-dimensional graph convolution to output the final color three-channel image.
[0201] 4. Example of deep graph optimization model structure.
[0202] As shown in FIG2 , FIG2 is a structural diagram of a depth map optimization model according to an embodiment of the present invention.
[0203] The depth map optimization model may include a view map feature extraction layer, a depth map feature extraction layer, a preliminary fusion layer, n feature fusion layers and an output layer.
[0204] The following explains the steps of how the model processes an image.
[0205] (1) First, the first view image (size 3×H×W, where H is the height of the image, W is the width of the image, and 3 is the three channels of RGB) and its corresponding first depth map (size 3×H×W, where H is the height of the image, W is the width of the image, and 3 is the three channels of RGB) are used as input to the depth map optimization model. The B first view images are input to the view image feature extraction layer for feature extraction to obtain the second view feature map (size B×C×H×W); the B first depth images are input to the depth map feature extraction layer for feature extraction to obtain the second depth feature map (size B×C×H×W). (Where B is the batch size and C is the number of channels processed).
[0206] (2) Then, the two features are concatenated in the channel dimension to obtain a concatenated feature with a size of B×(2C)×H×W. A preliminary fusion layer is then used to reduce the channel dimension by a factor of 2 to achieve preliminary fusion of the two features and obtain a comprehensive feature map (with a size of B×C×H×W).
[0207] (3) Secondly, the comprehensive feature map is input into n feature fusion layers for further feature fusion to obtain a fused feature map (size is B×C×H×W). The purpose is to use the edge information in the first viewpoint image to guide the optimization of the first depth map and improve the accuracy of the depth information in the depth map.
[0208] (4) Finally, the fused feature map is input into the output layer for fusion and channel dimension reduction, and finally B optimized depth maps (size is 3×H×W) are obtained.
[0209] 3. About the viewpoint rendering model.
[0210] The viewpoint rendering model can be used to predict the corresponding second viewpoint based on the optimized depth map obtained by the depth map optimization model and the first viewpoint, thereby obtaining a predicted viewpoint. Specifically, the prediction can be performed using the second viewpoint corresponding to the input first viewpoint as the prediction target to obtain the predicted viewpoint. The second viewpoint corresponding to the first viewpoint is the label of the training sample.
[0211] This method does not limit the specific structure of the viewpoint rendering model.
[0212] Optionally, the viewpoint rendering model can adopt a graph neural network or a graph convolutional network structure.
[0213] In an optional embodiment, the viewpoint rendering model may perform feature fusion on the input optimized depth map and the first viewpoint map.
[0214] Optionally, the viewpoint rendering model can be used to perform feature fusion on the input optimized depth map and the first viewpoint map, and predict the second viewpoint map corresponding to the input first viewpoint map based on the fusion result to obtain a predicted viewpoint map.
[0215] Optionally, in order to reduce the amount of calculation, the feature maps extracted from the optimized depth map and the first viewpoint map can be reduced in dimension, and then feature fusion can be performed on the reduced-dimensional feature maps, and then the fusion results can be increased in dimension.
[0216] Furthermore, optionally, multiple viewpoint feature maps of different dimensions can be extracted for the first viewpoint image. In the process of dimensionality upgrading, the extracted viewpoint feature maps of different dimensions are combined for fusion and feature extraction, so that the accuracy of predicting the corresponding second viewpoint image can be improved with the help of the viewpoint feature map of the first viewpoint image, that is, the accuracy of the predicted viewpoint image.
[0217] Of course, optionally, multiple depth feature maps of different dimensions can be extracted for the optimized depth map. In the process of dimensionality upgrading, the extracted depth feature maps of different dimensions are combined for fusion and feature extraction, so that the accuracy of predicting the corresponding second viewpoint map can be improved with the help of the depth feature map of the optimized depth map, that is, the accuracy of the predicted viewpoint map.
[0218] The above two embodiments can be combined with each other to combine viewpoint feature maps and depth feature maps of different dimensions in the dimensionality upgrading process.
[0219] For ease of understanding, in a specific example, viewpoint feature maps with resolutions of 8×8, 4×4, and 2×2 can be extracted for the first viewpoint image, and depth feature maps with resolutions of 8×8, 4×4, and 2×2 can be extracted for the optimized depth map. Feature fusion can then be performed on the viewpoint feature map with a resolution of 2×2 and the depth feature map with a resolution of 2×2. The fusion result with a resolution of 2×2 is then dimensionally upgraded to obtain a 4×4 fusion result, which is then combined with the 4×4 viewpoint feature map and / or the 4×4 depth feature map to obtain a 4×4 dimensionally upgraded fusion result. The 4×4 dimensionally upgraded fusion result can then be dimensionally upgraded to obtain an 8×8 dimensionally upgraded fusion result, which is then combined with the 8×8 viewpoint feature map and / or the 8×8 depth feature map to obtain the final result feature map. The resulting feature map can be used to reconstruct and generate a predicted viewpoint image.
[0220] Optionally, the viewpoint rendering model is used to: extract a third viewpoint feature map of a first size and a fourth viewpoint feature map of a second size for the first viewpoint map; wherein the first size is smaller than the second size; and perform feature fusion on the third viewpoint feature map and the optimized feature map extracted based on the optimized depth map to obtain an optimized fused feature map of the first size.
[0221] Optionally, the viewpoint rendering model may include: a first dimensionality-increasing fusion layer.
[0222] Optionally, the first dimensionality-upgrading fusion layer can be used to: increase the dimension of the optimized fusion feature map to obtain a dimensionality-upgrading fusion feature map of the second size; perform feature fusion on the dimensionality-upgrading fusion feature map and the fourth viewpoint feature map, and output a fusion result of the second size.
[0223] This embodiment does not limit the specific method of feature fusion, and feature fusion can be performed using graph convolution, graph neural network, graph sampling, pooling, etc.
[0224] The size of the optimized feature map extracted based on the optimized depth map may be a first size.
[0225] On the one hand, this embodiment can fuse feature maps of lower sizes and then perform dimensionality fusion. While ensuring that the dimension of the feature maps is high (high precision), low-dimensional features can be used for fusion during feature fusion, thereby reducing the amount of calculation and improving calculation efficiency.
[0226] On the other hand, by combining higher-dimensional viewpoint feature maps in the dimensionality-increasing process, the accuracy of the predicted viewpoint map can be improved.
[0227] In addition, optionally, the feature map may be directly upgraded without fusing the fourth viewpoint feature map. Optionally, the depth feature map of the second size may be fused during the dimensionality upgrade process.
[0228] Optionally, on the basis of the first dimensionality-raising fusion layer, one or more similar dimensionality-raising fusion layers may be additionally added to gradually fuse viewpoint feature maps and / or depth feature maps of higher dimensions.
[0229] Optionally, the viewpoint rendering model may further include: a second dimensionality-increasing fusion layer.
[0230] Optionally, the viewpoint rendering model may also be used to: extract a fifth viewpoint feature map of a third size for the first viewpoint map; wherein the second size is smaller than the third size.
[0231] Optionally, the second dimensionality-raising fusion layer can be used to: increase the dimension of the fusion result output by the first dimensionality-raising fusion layer to obtain a feature map to be fused of a third size; perform feature fusion on the feature map to be fused and the fifth viewpoint feature map, and output a fusion result of a third size.
[0232] This embodiment can improve the accuracy of the predicted viewpoint graph through further dimensionality fusion.
[0233] Alternatively, the feature map may be directly upgraded without fusing the fifth viewpoint feature map. Alternatively, the depth feature map of the third size may be fused during the dimensionality upgrade process.
[0234] Optionally, based on this embodiment, one or more cascaded dimensionality-increasing fusion layers can be deployed in the view rendering model through simple reasoning. Specifically, the view rendering model can include one or more cascaded dimensionality-increasing fusion layers to gradually fuse higher-dimensional view feature maps and / or depth feature maps to improve the accuracy of the predicted view map.
[0235] For the feature maps of different dimensions extracted from the first view image or the optimized depth image, the present method does not limit the specific extraction method. Optionally, the feature maps of different dimensions can be extracted by downsampling or dimensionality reduction.
[0236] For example, for the C×H×W feature map extracted from the first viewpoint image, a C×(H / 2)×(W / 2) feature map and a C×(H / 4)×(W / 4) feature map can be extracted by downsampling.
[0237] The specific way to increase the dimension of the dimensional fusion layer may be to increase at least one of the following: channel dimension, resolution dimension, width and height, etc. Specifically, the resolution dimension may be increased.
[0238] For ease of understanding, this method also provides a structural example of a viewpoint rendering model.
[0239] As shown in FIG3 , FIG3 is a structural diagram of a viewpoint rendering model according to an embodiment of the present invention.
[0240] The view rendering model may include a view map feature extraction layer, a depth map feature extraction layer, a preliminary fusion layer, a feature fusion layer, two dimension-upgrading fusion layers, and an output layer. The number of dimension-upgrading fusion layers is for illustrative purposes only; three or more layers are possible.
[0241] The following explains the steps of how the model processes an image.
[0242] (1) The first view map (with a size of 3×H×W, where H is the height of the image, W is the width of the image, and 3 is the 3 channels of RGB) is input to the view map feature extraction layer.
[0243] Specifically, the initial feature extraction can be performed by one layer of convolution, and then the H and W dimensions are downsampled by two convolutions to obtain feature map 1 with a size of B×C×H×W, feature map 2 with a size of B×C×(H / 2)×(W / 2), and feature map 3 with a size of B×C×(H / 4)×(W / 4).
[0244] (2) The optimized depth map (with a size of 3×H×W, where H is the height of the image, W is the width of the image, and 3 is the 3 channels of RGB) is input into the depth map feature extraction layer.
[0245] Specifically, one layer of convolution can be used to extract initial features, and then two convolutions can be used to downsample the H and W dimensions to obtain a feature map 4 with a size of B×C×(H / 4)×(W / 4).
[0246] (3) Feature map 3 and feature map 4 are concatenated in the channel dimension to obtain a feature map of size B×(2C)×(H / 4)×(W / 4), and then pass through a preliminary fusion layer, specifically a two-dimensional convolutional neural network layer to perform a 2-fold dimensionality reduction of the channel dimension, to achieve preliminary fusion of the two features and obtain feature map 5 (size B×C×(H / 4)×(W / 4)).
[0247] (4) Input feature map 5 into a feature fusion layer for feature fusion to obtain feature map 6 (size is B×C×(H / 4)×(W / 4)).
[0248] (5) Input feature map 6 and feature map 2 into a dimensionality-enhancing fusion layer.
[0249] First, the feature map 6 is upgraded to B×C×(H / 2)×(W / 2). Specifically, the dimension can be upgraded using a Bilinear layer, and then feature fusion is performed. After another layer of convolution, feature map 7 (with a size of B×C×(H / 2)×(W / 2)) is obtained.
[0250] (6) Input feature map 7 and feature map 1 into a dimensionality-enhancing fusion layer.
[0251] First, the feature map 7 is upgraded to B×C×H×W. Specifically, the dimension can be upgraded using a Bilinear layer, and then feature fusion is performed. After another layer of convolution, the feature map 8 (with a size of B×C×H×W) is obtained.
[0252] (7) Finally, the feature map 8 is input into the output layer for fusion and channel dimension reduction, and finally the predicted viewpoint map (size is 3×H×W) is obtained.
[0253] For the explanation of feature fusion and output layer, please refer to the explanation of the deep graph optimization model above.
[0254] 4. About the training process.
[0255] Regarding the depth map optimization model and viewpoint rendering model connected in series, the process of this method does not limit the specific training method.
[0256] The depth map optimization model and the viewpoint rendering model can be trained based on the training sample set; during the training process, the parameters of the depth map optimization model can be updated according to the loss between the second viewpoint map and the predicted viewpoint map.
[0257] Optionally, since the sample label in the training sample set is the second view image, the view rendering model needs to participate in the training, and the parameters of the view rendering model may be updated or not.
[0258] Optionally, the viewpoint rendering model may be pre-trained; the above method flow may further include: freezing the parameters of the viewpoint rendering model during the training of the depth map optimization model and the viewpoint rendering model.
[0259] This embodiment can improve the convergence speed and training efficiency by only updating the parameters of the depth map optimization model.
[0260] Optionally, the above method flow may further include: during the training process, updating parameters of the viewpoint rendering model according to the loss between the second viewpoint image and the predicted viewpoint image.
[0261] This embodiment can update the depth map optimization model and the viewpoint rendering model through comprehensive training, thereby improving the overall training effect of the two models.
[0262] Regarding the calculation method of the loss, this method process does not specifically limit it.
[0263] Optionally, the loss may specifically be the similarity between the predicted view image and the second view image.
[0264] Optionally, the model training update may be performed based on the total loss between the predicted view image corresponding to each training sample and the second view image, and specifically, the calculation may be performed using an L1 loss or an L2 loss calculation method.
[0265] Optionally, model training can be updated based on the loss of the same region between the predicted view image and the second view image. Specifically, the loss can be the similarity between the image content of the same region in the predicted view image and the second view image. This can highlight key regions and improve training efficiency.
[0266] Alternatively, for a depth map, the depth information of the background may be highly uncertain. For example, it is generally difficult to accurately determine the depth information of backgrounds such as the sky, city, and horizon. However, the foreground area is often the focus of image viewing, and its depth information has greater certainty.
[0267] Therefore, optionally, the model training can be updated based on the loss of the same foreground area between the predicted view image and the second view image, thereby highlighting the key areas and improving the training efficiency.
[0268] Optionally, the above method may further include: determining a foreground area for the first view image.
[0269] Updating the parameters of the depth map optimization model based on the loss between the second view image and the predicted view image may include: updating the parameters of the depth map optimization model based on the loss on the same target area between the second view image and the predicted view image; the position of the target area in the second view image and the predicted view image is the same as the position of the determined foreground area in the first view image.
[0270] This embodiment can determine the loss of image content at the position of the foreground area between the second view image and the predicted view image by using the foreground area of the first view image, thereby improving the efficiency of loss calculation.
[0271] Optionally, the foreground area can also be determined for the second view image, and the parameters of the depth map optimization model can be updated based on the loss between the second view image and the predicted view image. This can include: updating the parameters of the depth map optimization model based on the loss on the same target area between the second view image and the predicted view image; the position of the target area in the second view image and the predicted view image is the same as the position of the determined foreground area in the second view image.
[0272] The method flow does not limit the method of determining the foreground area.
[0273] Optionally, determining the foreground area for the first view image may include any of the following:
[0274] (1) Perform foreground target detection on the first view image and determine the foreground area.
[0275] (2) Perform foreground target detection on the first view image to obtain a foreground region frame, then perform target segmentation on the image content in the foreground region frame, and determine the segmented image content as the foreground region.
[0276] (3) Segment the foreground target on the first view image and determine the segmented image content as the foreground area.
[0277] This embodiment can determine the foreground area through target detection and target segmentation, thereby improving the efficiency and accuracy of foreground area determination.
[0278] For ease of understanding, an example of a loss function is given below.
[0279] First, according to the foreground area Fmask in the first view image, the corresponding foreground pixels in the generated predicted view image F2D-out are obtained: Fout=F2D-out⊙Fmask.
[0280] Secondly, according to the foreground area Fmask in the first view image, the corresponding foreground pixels in the second view image Fright (label) are obtained: FGT=Fright⊙Fmask.
[0281] Finally, the reconstruction loss L1 loss is used to calculate the similarity between the two: £L1=L1(Fout,FGT).
[0282] Among them, ⊙ represents the dot product operation.
[0283] 5. About application methods.
[0284] The depth map optimization model and viewpoint rendering model trained using the training sample set can be used to implement a variety of applications.
[0285] This method does not limit the application of the trained depth map optimization model and / or the trained viewpoint rendering model.
[0286] Three examples are given below for illustrative purposes.
[0287] 1. Optimize depth map.
[0288] In an optional embodiment, the depth map is often obtained by a preset algorithm, such as a monocular estimation algorithm, but the accuracy is low.
[0289] In order to facilitate the use of more accurate depth maps for image processing, for example, using the depth map to optimize the rendering of the corresponding single viewpoint image to improve the clarity of the foreground area and the smoothness of the background area; or using the depth map for more detailed foreground segmentation, etc., a depth map optimization model trained based on the above method embodiment can be used.
[0290] As shown in FIG4 , FIG4 is a schematic flow chart of a depth map optimization method according to an embodiment of the present invention.
[0291] S201: Acquire a target viewpoint map and a target depth map.
[0292] The target depth map may include the depth information of each pixel in the target view image.
[0293] S202: Inputting the target viewpoint map and the target depth map into a pre-trained depth map optimization model to obtain an optimized depth map output by the depth map optimization model.
[0294] The depth map optimization model may be obtained by training based on the above method embodiment.
[0295] The target view graph can be any view graph. For the convenience of description, any view graph is referred to as the target view graph.
[0296] This embodiment does not limit the specific method for obtaining the target depth map. Optionally, it can be determined by a monocular estimation algorithm based on the target viewpoint map.
[0297] This embodiment can optimize the depth map by using a machine learning method and a trained depth map optimization model to improve the accuracy of the depth information in the depth map.
[0298] 2. Generate a three-dimensional image.
[0299] In an optional embodiment, a three-dimensional image may be generated using the same image from different viewpoints, such as a left viewpoint image and a right viewpoint image.
[0300] Therefore, for a single viewpoint image, the depth map optimization model and viewpoint rendering model trained based on the above method embodiment can be used to generate the same image under different viewpoints for generating a three-dimensional image.
[0301] As shown in FIG5 , FIG5 is a flow chart of a method for generating a three-dimensional image according to an embodiment of the present invention.
[0302] S301: Acquire a target viewpoint map and a target depth map.
[0303] The target depth map may include depth information of each pixel in the target view image.
[0304] S302: Inputting the target viewpoint map and the target depth map into a pre-trained depth map optimization model to obtain an optimized depth map output by the depth map optimization model.
[0305] S303: Inputting the optimized depth map and the target viewpoint map into a pre-trained viewpoint rendering model to obtain a predicted viewpoint map output by the viewpoint rendering model.
[0306] Among them, the depth map optimization model and the viewpoint rendering model can be trained based on the above method embodiment.
[0307] S304: Perform interleaving processing on the target view image and the predicted view image to generate a three-dimensional image.
[0308] This embodiment can use machine learning to optimize the depth map using a trained depth map optimization model to improve the accuracy of depth information in the depth map, and further use the trained viewpoint rendering model to generate predicted viewpoint maps under different viewpoints, generate three-dimensional images through interleaving processing, realize the generation of three-dimensional images, and improve the accuracy of three-dimensional images.
[0309] The target view graph can be any view graph. For the convenience of description, any view graph is referred to as the target view graph.
[0310] This embodiment does not limit the specific method for obtaining the target depth map. Optionally, it can be determined based on the target viewpoint image using a monocular estimation algorithm. Based on this method, a depth map to be optimized can be generated for a single viewpoint image, and then a 3D image can be generated using this embodiment.
[0311] This embodiment does not limit the interleaving processing method, and may specifically employ multi-viewpoint image interleaving processing to generate a three-dimensional image.
[0312] 3. 3D video conversion.
[0313] Based on a single-viewpoint image, the above method can be used to predict the same image from another viewpoint to generate a three-dimensional image. Therefore, the above method process can be further executed for each video frame in the video (that is, a single-viewpoint image) to generate a three-dimensional image, thereby obtaining a three-dimensional video.
[0314] As shown in FIG6 , FIG6 is a schematic flow chart of a 3D video conversion method according to an embodiment of the present invention.
[0315] S401: Acquire a target video.
[0316] S402: For any video frame in the target video, perform a preset operation to obtain a corresponding 3D video frame.
[0317] The preset operations include: obtaining a target depth map corresponding to the targeted video frame; the target depth map contains depth information of each pixel of the targeted video frame; inputting the targeted video frame and the target depth map into a pre-trained depth map optimization model to obtain an optimized depth map output by the depth map optimization model; inputting the optimized depth map and the targeted video frame into a pre-trained viewpoint rendering model to obtain a predicted viewpoint map output by the viewpoint rendering model; the depth map optimization model and the viewpoint rendering model can be trained based on the above-mentioned method embodiment; and performing interleaving processing on the targeted video frame and the predicted viewpoint map to generate a corresponding three-dimensional video frame.
[0318] S403: 3D video frames corresponding to the video frames in the target video are synthesized to obtain a 3D video.
[0319] This embodiment does not limit the method for obtaining the target video, and the target video may be obtained locally or received.
[0320] The target video can be any video. For the convenience of description, any video is referred to as the target video.
[0321] This embodiment does not limit the method of synthesizing the 3D video frames. Alternatively, the corresponding 3D video frames may be combined according to the order of the video frames in the target video to obtain the 3D video.
[0322] Optionally, a preset operation may be performed on each video frame in the target video to obtain a corresponding 3D video frame. Specifically, the 3D video frames corresponding to each video frame in the target video may be integrated to obtain the 3D video.
[0323] This embodiment can use machine learning to optimize the depth map using a trained depth map optimization model to improve the accuracy of depth information in the depth map, and further use the trained viewpoint rendering model to generate predicted viewpoint maps under different viewpoints, generate three-dimensional images through interleaving processing, realize the generation of three-dimensional images, and improve the accuracy of three-dimensional images.
[0324] Based on the generation of three-dimensional images, the generation of three-dimensional videos can also be achieved, thereby improving the accuracy of three-dimensional videos.
[0325] For ease of understanding, the embodiment of the present invention further provides an application embodiment.
[0326] 1. Background.
[0327] With the increasing diversification and standardization of user needs, ultra-high-definition 2D images are increasingly unable to meet user viewing needs. The three-dimensionality and visual impact of 3D images have become a new pursuit for users. Consequently, glasses-free 3D display technology has become a hot topic in the display industry. Products such as glasses-free 3D displays, 3D photo albums, and 3D tablets already exist. It is foreseeable that glasses-free 3D will become a new trend in the display industry.
[0328] With the advent of naked-eye 3D display products, there is a widespread market demand for consumer-level 3D content.
[0329] There are two sources of 3D content:
[0330] 1. Directly capture native 3D images and videos using a binocular camera. However, binocular cameras are not widely available, are expensive, and cannot be used to reuse existing 2D content, requiring new 3D content to be captured. Therefore, this solution is not currently suitable for general user needs.
[0331] 2. Convert existing 2D images and videos to 3D. This solution currently relies heavily on manual post-production using specialized software, and is often used in film-quality productions. However, manual post-production is expensive, time-consuming, and labor-intensive, making it unsuitable for general user needs.
[0332] With the rise of deep learning, applying it to 3D image generation has become one of the most cutting-edge research directions in the display field. Compared to manual post-production, deep learning-based 3D image generation technology offers advantages such as high efficiency, low cost, and high quality, making it ideally suited to a wide range of consumer needs.
[0333] The 3D image generation process needs to use the existing 2D image to generate the corresponding depth map, and then use both as input to the 3D image generation process.
[0334] However, the application target of existing monocular depth estimation algorithms is not 3D image generation, so they are not suitable for 3D image generation.
[0335] Specifically, the depth map used for 3D image generation should be more accurate in the foreground area of the image, while the background area should be smoother.
[0336] However, existing monocular depth estimation algorithms have insufficient foreground accuracy, which causes the generated viewpoint images to have very significant artifacts subjectively.
[0337] Therefore, it is particularly necessary to study a depth map region optimization algorithm suitable for 3D image generation to adapt to depth maps of different quality as input and effectively improve the accuracy of foreground area depth estimation.
[0338] Currently, there is no open source depth map dataset for 3D image generation. This dataset is difficult to obtain and expensive.
[0339] The solution proposed in this embodiment cleverly utilizes an unsupervised learning method, which can effectively implement the training of the depth map optimization model even without accurate depth map labels.
[0340] 2. Brief introduction of the plan.
[0341] This embodiment designs a depth map region optimization algorithm based on unsupervised learning.
[0342] To address the problems that existing monocular depth estimation algorithms are not suitable for 3D image generation and the difficulty in obtaining depth maps corresponding to 3D images, this embodiment innovatively designs a depth map optimization algorithm that can be trained without a depth dataset, and proposes a progressively fused viewpoint rendering algorithm, which significantly improves the generation artifact problem caused by inaccurate foreground depth and effectively improves the robustness of 3D image generation.
[0343] 3. Explanation of the program steps.
[0344] As shown in FIG7 , FIG7 is a schematic diagram showing the principle of depth map optimization model training according to an embodiment of the present invention.
[0345] The depth map optimization algorithm proposed in this embodiment is performed under the premise of having a 3D dataset and a trained viewpoint rendering model.
[0346] A 3D dataset (such as the open source KITTY dataset) contains a left view image Fleft, a right view image Fright, and a depth map FDep corresponding to the left view image.
[0347] The functional requirement of the viewpoint rendering model used in this embodiment is to output a corresponding right viewpoint image based on the input left viewpoint image and depth map.
[0348] The left view image is the 2D image in FIG7 , and the right view image is the original view image in FIG7 .
[0349] The specific steps of the method are as follows:
[0350] 1. First, use the target detection algorithm (such as Fast R-CNN, YOLO v5, SSD, etc.) to extract the foreground area in the left view image Fleft.
[0351] 2. Then, use the instance segmentation algorithm (such as PP-Matting, Unet, etc.) to segment the foreground area and obtain Fmask (the mask image corresponding to the left view image Fleft).
[0352] 3. Then, the left view image Fleft and its corresponding depth map FDep are input into the depth map optimization model to obtain the optimized depth map FDep-out.
[0353] 4. Secondly, FDep-out and Fleft are input into the trained viewpoint rendering model (parameters are frozen), and the new viewpoint graph F2D-out is rendered according to FDep-out.
[0354] 5. Finally, according to the preset mask image, take out the corresponding area in F2D-out and the corresponding area in the right view image Fright (label), calculate the loss between the two areas and perform gradient update, thus completing a model training.
[0355] It is important to note that since the parameters of the viewpoint rendering model are frozen, training will only update the parameters of the depth map optimization model.
[0356] By cleverly utilizing the existing viewpoint rendering model and 3D dataset (used to generate left and right viewpoint map datasets), the loss is converted from the unlabeled depth space to the labeled left and right viewpoint map space, realizing a depth map optimization algorithm based on unsupervised learning.
[0357] At the same time, since the overall model is only fitted to the mask image (foreground area), there is no need to circle the foreground area in advance after training, and the model can automatically have the foreground selection capability.
[0358] The solution proposed in this embodiment has no special requirements for the viewpoint rendering algorithm and can be widely used in 3D conversion services of images / videos to further optimize the depth map obtained by the depth estimation algorithm.
[0359] Because the available object detection and instance segmentation algorithms are flexible, you can choose the appropriate one based on the specific business scenario. For example, for 3D conversion in a live broadcast scenario, you can choose an object detection and portrait cutout algorithm to achieve "hair-level" 3D conversion for the host.
[0360] For a more specific explanation of the model structure, please refer to the above.
[0361] For ease of understanding, an example of image optimization is given below.
[0362] As shown in FIG8 , FIG8 is a schematic diagram showing an image optimization result according to an embodiment of the present invention.
[0363] This includes the depth map before and after optimization, as well as the right view map before and after optimization.
[0364] Corresponding to the above method embodiments, the embodiments of the present invention also provide corresponding device embodiments.
[0365] As shown in FIG9 , FIG9 is a structural diagram of a depth map optimization model training device according to an embodiment of the present invention.
[0366] The apparatus may include the following units.
[0367] An acquisition unit 501 is configured to acquire a depth map optimization model and a viewpoint rendering model; acquire a training sample set; wherein the sample features of any training sample in the training sample set include: a first viewpoint map and a first depth map; and the sample label of any training sample includes: a second viewpoint map corresponding to the first viewpoint map;
[0368] The first depth map includes depth information of each pixel in the first view image; the first view image and the second view image are the same image at different viewpoints;
[0369] The depth map optimization model is used to optimize the input first depth map according to the input first viewpoint map and the first depth map to obtain an optimized depth map;
[0370] The viewpoint rendering model is used to predict the corresponding second viewpoint image based on the optimized depth map obtained by the depth map optimization model and the first viewpoint image to obtain a predicted viewpoint image;
[0371] A training unit 502 is configured to train a depth map optimization model and a viewpoint rendering model based on a training sample set;
[0372] During the training process, the parameters of the depth map optimization model are updated according to the loss between the second view image and the predicted view image.
[0373] Optionally, the training unit 502 is further configured to: determine a foreground area for the first view image;
[0374] The training unit 502 is used to:
[0375] Update the parameters of the depth map optimization model based on the loss between the second view image and the predicted view image on the same target area;
[0376] The position of the target area in the second view image and the predicted view image is the same as the position of the determined foreground area in the first view image.
[0377] Optionally, the training unit 502 is configured to perform any one of the following to determine a foreground area for the first view image:
[0378] Perform foreground object detection on the first view image to determine the foreground area;
[0379] Perform foreground target detection on the first view image to obtain a foreground region frame, then perform target segmentation on the image content in the foreground region frame, and determine the segmented image content as the foreground region;
[0380] The foreground object is segmented for the first view image, and the segmented image content is determined as the foreground area.
[0381] Optionally, the depth map optimization model includes: a view map feature extraction layer and a depth map feature extraction layer;
[0382] The viewpoint map feature extraction layer is used to: extract a first viewpoint feature map for the input first viewpoint map; then increase the dimension of the obtained first viewpoint feature map to obtain a second viewpoint feature map;
[0383] The depth map feature extraction layer is used to: extract a first depth feature map from the input first depth map; and then increase the dimension of the obtained first depth feature map to obtain a second depth feature map.
[0384] Optionally, the depth map optimization model includes: a feature fusion layer;
[0385] The depth map optimization model is used to synthesize the view feature map extracted based on the first view map and the depth feature map extracted based on the first depth map to obtain a comprehensive feature map;
[0386] The feature fusion layer is used to: reduce the dimension of the input feature map to obtain a reduced-dimensional feature map; then perform feature fusion on the reduced-dimensional feature map to obtain a fused feature map; then increase the dimension of the fused feature map to obtain an increased-dimensional feature map, and output the sum of the increased-dimensional feature map and the input feature map;
[0387] The input feature map includes any of the following: a comprehensive feature map, a feature map extracted based on the comprehensive feature map, and a feature map output by other feature fusion layers.
[0388] Optionally, the first view image, the first depth image, and the second view image are acquired based on the same three-dimensional image.
[0389] Optionally, a method for obtaining any training sample in the training sample set includes:
[0390] Acquire a three-dimensional image of the target;
[0391] Acquire images of the target three-dimensional image at two different viewpoints, and determine them as a first viewpoint image and a second viewpoint image corresponding to the first viewpoint image;
[0392] Based on the depth information of the target three-dimensional image, obtain the depth information of each pixel in the first view image to determine the first depth map; or, based on the first view image, determine the first depth map according to a preset depth estimation algorithm;
[0393] The determined first view image and the first depth image are determined as sample features, and the determined second view image is determined as a sample label to obtain a training sample.
[0394] Optionally, the viewpoint rendering model is pre-trained; the training unit 502 is further configured to freeze the parameters of the viewpoint rendering model during the training of the depth map optimization model and the viewpoint rendering model.
[0395] Optionally, the training unit 502 is further configured to:
[0396] During the training process, the parameters of the view rendering model are updated according to the loss between the second view image and the predicted view image.
[0397] Optionally, the viewpoint rendering model includes: a first dimensionality-raising fusion layer;
[0398] The viewpoint rendering model is used to: extract a third viewpoint feature map of a first size and a fourth viewpoint feature map of a second size for the first viewpoint map; wherein the first size is smaller than the second size; perform feature fusion on the third viewpoint feature map and the optimized feature map extracted based on the optimized depth map to obtain an optimized fused feature map of the first size;
[0399] The first dimensionality-raising fusion layer is used to: increase the dimension of the optimized fusion feature map to obtain a dimensionality-raising fusion feature map of the second size; perform feature fusion on the dimensionality-raising fusion feature map and the fourth viewpoint feature map, and output a fusion result of the second size.
[0400] Optionally, the viewpoint rendering model further includes: a second dimensionality-raising fusion layer;
[0401] The viewpoint rendering model is also used to: extract a fifth viewpoint feature map of the third size for the first viewpoint map; wherein the second size is smaller than the third size; the second dimensionality-increasing fusion layer is used to: increase the dimension of the fusion result output by the first dimensionality-increasing fusion layer to obtain a feature map to be fused of the third size; perform feature fusion on the feature map to be fused and the fifth viewpoint feature map, and output a fusion result of the third size.
[0402] For detailed explanation, please refer to the above method embodiment.
[0403] As shown in FIG10 , FIG10 is a schematic structural diagram of a depth map optimization device according to an embodiment of the present invention.
[0404] The device may include the following units:
[0405] The first image acquisition unit 601 is configured to acquire a target view image and a target depth image; the target depth image includes depth information of each pixel in the target view image.
[0406] A first optimization unit 602 is configured to input a target view map and a target depth map into a pre-trained depth map optimization model to obtain an optimized depth map output by the depth map optimization model;
[0407] The depth map optimization model is obtained through training using the above-mentioned training device or the above-mentioned method embodiment.
[0408] For detailed explanation, please refer to the above method embodiment.
[0409] As shown in FIG11 , FIG11 is a schematic structural diagram of a three-dimensional image generating device according to an embodiment of the present invention.
[0410] The device may include the following units:
[0411] The second image acquisition unit 701 is configured to acquire a target view image and a target depth image; the target depth image includes depth information of each pixel in the target view image.
[0412] The second optimization unit 702 is configured to input the target view map and the target depth map into a pre-trained depth map optimization model to obtain an optimized depth map output by the depth map optimization model.
[0413] The second prediction unit 703 is configured to input the optimized depth map and the target view map into a pre-trained view rendering model to obtain a predicted view map output by the view rendering model.
[0414] The depth map optimization model and the viewpoint rendering model are obtained through training using the above-mentioned training device or the above-mentioned method embodiment.
[0415] The second interleaving unit 704 is configured to perform interleaving processing on the target view image and the predicted view image to generate a 3D image.
[0416] For detailed explanation, please refer to the above method embodiment.
[0417] As shown in FIG12 , FIG12 is a structural diagram of a three-dimensional video conversion device according to an embodiment of the present invention.
[0418] The device may include the following units:
[0419] The video acquisition unit 801 is used to acquire a target video.
[0420] The operation unit 802 is used to perform the following operations for any video frame in the target video to obtain a corresponding three-dimensional video frame: obtain a target depth map corresponding to the targeted video frame; the target depth map contains depth information of each pixel of the targeted video frame; input the targeted video frame and the target depth map into a pre-trained depth map optimization model to obtain an optimized depth map output by the depth map optimization model; input the optimized depth map and the targeted video frame into a pre-trained viewpoint rendering model to obtain a predicted viewpoint map output by the viewpoint rendering model; the depth map optimization model and the viewpoint rendering model are obtained by training with the above-mentioned training device or the above-mentioned method embodiment; and perform interleaving processing on the targeted video frame and the predicted viewpoint map to generate a corresponding three-dimensional video frame.
[0421] The integration unit 803 is configured to integrate the 3D video frames corresponding to the video frames in the target video to obtain a 3D video.
[0422] For detailed explanation, please refer to the above method embodiment.
[0423] An embodiment of the present invention further provides a computer device, which includes at least a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any one of the above method embodiments is implemented.
[0424] An embodiment of the present invention also provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any of the above-mentioned method embodiments.
[0425] FIG13 is a schematic diagram of the hardware structure of a computer device configured to implement the method according to an embodiment of the present invention, according to an embodiment of the present invention. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other within the device via the bus 1050.
[0426] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention.
[0427] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present invention are implemented through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0428] The input / output interface 1030 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0429] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).
[0430] The bus 1050 comprises a pathway for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).
[0431] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, those skilled in the art will understand that the above device may only include the components necessary to implement the embodiments of the present invention, and does not necessarily include all the components shown in the figure.
[0432] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements any of the above method embodiments when executed by a processor.
[0433] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program implements any of the above method embodiments when executed by a processor.
[0434] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0435] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the embodiments of the present invention can be implemented by means of software plus the necessary general hardware platform. Based on this understanding, the technical solution of the embodiments of the present invention can essentially or in other words, the contributing part can be embodied in the form of a software product. The computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or certain parts of the embodiments of the present invention.
[0436] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.
[0437] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is merely illustrative, wherein the modules described as separate components may or may not be physically separated, and when implementing the embodiment of the present invention, the functions of each module can be implemented in the same one or more software and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the embodiment. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0438] The above description is only a specific implementation of the embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the embodiment of the present invention. These improvements and modifications should also be regarded as protection for the embodiment of the present invention.
[0439] In the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance. The term "plurality" refers to two or more, unless otherwise clearly defined.
[0440] Other embodiments of the present invention will readily occur to those skilled in the art after considering the specification and practicing the disclosure herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the invention being indicated by the following claims.
[0441] It should be understood that the present invention is not limited to the exact construction described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
Claims
1. A method for training a depth map optimization model, characterized in that: include: Obtain depth map optimization model and viewpoint rendering model; Acquire a training sample set; the sample features of any training sample in the training sample set include: a first view image and a first depth image; the sample label of any training sample includes: a second view image corresponding to the first view image; The first depth map contains depth information of each pixel in the first view image; the first view image and the second view image are the same image at different viewpoints; The depth map optimization model is used to optimize the input first depth map according to the input first viewpoint map and the first depth map to obtain an optimized depth map; The viewpoint rendering model is used to predict the corresponding second viewpoint image according to the optimized depth image obtained by the depth map optimization model and the first viewpoint image to obtain a predicted viewpoint image; Based on the training sample set, training the depth map optimization model and the viewpoint rendering model; During the training process, the parameters of the depth map optimization model are updated according to the loss between the second view map and the predicted view map.
2. The method according to claim 1, characterized in that Also includes: determining a foreground area for the first view image; The updating of the parameters of the depth map optimization model according to the loss between the second view map and the predicted view map comprises: updating the parameters of the depth map optimization model according to the loss between the second view map and the predicted view map on the same target area; The position of the target area in the second view image and the predicted view image is the same as the position of the determined foreground area in the first view image.
3. The method according to claim 2, characterized in that The determining of the foreground area for the first view image includes any one of the following: Performing foreground object detection on the first view image to determine a foreground area; Performing foreground target detection on the first view image to obtain a foreground area frame, and then performing target segmentation on the image content in the foreground area frame, and determining the segmented image content as the foreground area; Foreground object segmentation is performed on the first view image, and the segmented image content is determined as the foreground area.
4. The method according to claim 1, characterized in that: The depth map optimization model includes: a view map feature extraction layer and a depth map feature extraction layer; The view map feature extraction layer is used to: extract a first view feature map for the input first view map; and then increase the dimension of the obtained first view feature map to obtain a second view feature map; The depth map feature extraction layer is used to: extract a first depth feature map for an input first depth map; and then increase the dimension of the obtained first depth feature map to obtain a second depth feature map.
5. The method according to claim 1, characterized in that The depth map optimization model includes: a feature fusion layer; The depth map optimization model is used to synthesize the view feature map extracted based on the first view map and the depth feature map extracted based on the first depth map to obtain a comprehensive feature map; The feature fusion layer is used to: reduce the dimension of the input feature map to obtain a reduced-dimensional feature map; then perform feature fusion on the reduced-dimensional feature map to obtain a fused feature map; then increase the dimension of the fused feature map to obtain an increased-dimensional feature map, and output the sum of the increased-dimensional feature map and the input feature map; The input feature map includes any one of the following: the comprehensive feature map, a feature map extracted based on the comprehensive feature map, and a feature map output by other feature fusion layers.
6. The method according to claim 1, characterized in that The first view image, the first depth image and the second view image are acquired based on the same three-dimensional image.
7. The method according to claim 1, characterized in that The method of obtaining any training sample in the training sample set includes: Acquire a three-dimensional image of a target; Acquire images of the target three-dimensional image at two different viewpoints, and determine them as a first viewpoint image and a second viewpoint image corresponding to the first viewpoint image respectively; Based on the depth information of the target three-dimensional image, obtain the depth information of each pixel in the first view image to determine a first depth map; or, based on the first view image, determine the first depth map according to a preset depth estimation algorithm; The determined first view image and the first depth image are determined as sample features, and the determined second view image is determined as a sample label to obtain a training sample.
8. The method according to claim 1, characterized in that The viewpoint rendering model is pre-trained; the method further comprises: freezing parameters of the viewpoint rendering model during the training of the depth map optimization model and the viewpoint rendering model.
9. The method according to claim 1, characterized in that: Also includes: During the training process, the parameters of the viewpoint rendering model are updated according to the loss between the second viewpoint image and the predicted viewpoint image.
10. The method according to claim 1, characterized in that The viewpoint rendering model includes: a first dimensional fusion layer; The viewpoint rendering model is used to: extract a third viewpoint feature map of a first size and a fourth viewpoint feature map of a second size for the first viewpoint map; wherein the first size is smaller than the second size; perform feature fusion on the third viewpoint feature map and the optimized feature map extracted based on the optimized depth map to obtain an optimized fused feature map of the first size; The first dimensionality-upgrading fusion layer is used to: increase the dimension of the optimized fusion feature map to obtain a dimensionality-upgrading fusion feature map of a second size; perform feature fusion on the dimensionality-upgrading fusion feature map and the fourth viewpoint feature map, and output a fusion result of a second size.
11. The method according to claim 10, characterized in that The viewpoint rendering model further includes: a second dimension-upgrading fusion layer; The viewpoint rendering model is also used to: extract a fifth viewpoint feature map of a third size for the first viewpoint map; wherein the second size is smaller than the third size; The second dimensionality-raising fusion layer is used to: increase the dimension of the fusion result output by the first dimensionality-raising fusion layer to obtain a feature map to be fused of a third size; perform feature fusion on the feature map to be fused and the fifth viewpoint feature map, and output a fusion result of a third size.
12. A depth map optimization method, characterized in that: include: Acquire a target view image and a target depth image; the target depth image contains depth information of each pixel of the target view image; Inputting the target viewpoint map and the target depth map into a pre-trained depth map optimization model to obtain an optimized depth map output by the depth map optimization model; The depth map optimization model is trained based on the training method described in any one of claims 1 to 11.
13. A three-dimensional image generation method, characterized in that: include: Acquire a target view image and a target depth image; the target depth image contains depth information of each pixel of the target view image; Inputting the target viewpoint map and the target depth map into a pre-trained depth map optimization model to obtain an optimized depth map output by the depth map optimization model; Inputting the optimized depth map and the target view map into a pre-trained view rendering model to obtain a predicted view map output by the view rendering model; The depth map optimization model and the viewpoint rendering model are obtained by training based on the training method described in any one of claims 1 to 11; Interlacing processing is performed on the target view image and the predicted view image to generate a three-dimensional image.
14. A three-dimensional video conversion method, characterized in that: include: Get the target video; For any video frame in the target video, the following operations are performed to obtain the corresponding 3D video frame: Obtain a target depth map corresponding to the targeted video frame; the target depth map includes depth information of each pixel of the targeted video frame; Inputting the targeted video frame and the target depth map into a pre-trained depth map optimization model to obtain an optimized depth map output by the depth map optimization model; Inputting the optimized depth map and the targeted video frame into a pre-trained viewpoint rendering model to obtain a predicted viewpoint map output by the viewpoint rendering model; The depth map optimization model and the viewpoint rendering model are obtained by training based on the training method described in any one of claims 1 to 11; Performing interleaving processing on the targeted video frame and the predicted viewpoint map to generate a corresponding three-dimensional video frame; The three-dimensional video frames corresponding to the video frames in the target video are integrated to obtain a three-dimensional video.
15. A depth map optimization model training device, characterized in that: include: An acquisition unit is used to acquire a depth map optimization model and a viewpoint rendering model; acquire a training sample set; the sample features of any training sample in the training sample set include: a first viewpoint map and a first depth map; the sample label of any training sample includes: a second viewpoint map corresponding to the first viewpoint map; The first depth map contains depth information of each pixel in the first view image; the first view image and the second view image are the same image at different view points; The depth map optimization model is used to optimize the input first depth map according to the input first viewpoint map and the first depth map to obtain an optimized depth map; The viewpoint rendering model is used to predict the corresponding second viewpoint image according to the optimized depth image obtained by the depth map optimization model and the first viewpoint image to obtain a predicted viewpoint image; A training unit, configured to train the depth map optimization model and the viewpoint rendering model based on the training sample set; During the training process, the parameters of the depth map optimization model are updated according to the loss between the second view map and the predicted view map.
16. A depth map optimization device, characterized in that: include: A first image acquisition unit, configured to acquire a target view image and a target depth image; the target depth image includes depth information of each pixel of the target view image; A first optimization unit, configured to input the target view map and the target depth map into a pre-trained depth map optimization model to obtain an optimized depth map output by the depth map optimization model; The depth map optimization model is obtained by training based on the training device described in claim 15.
17. A three-dimensional image generating device, characterized in that: include: A second image acquisition unit, configured to acquire a target view image and a target depth image; the target depth image includes depth information of each pixel of the target view image; A second optimization unit, configured to input the target view map and the target depth map into a pre-trained depth map optimization model to obtain an optimized depth map output by the depth map optimization model; A second prediction unit, configured to input the optimized depth map and the target view map into a pre-trained view rendering model to obtain a predicted view map output by the view rendering model; The depth map optimization model and the viewpoint rendering model are obtained by training based on the training device according to claim 15; The second interleaving unit is used to perform interleaving processing on the target view image and the predicted view image to generate a three-dimensional image.
18. A three-dimensional video conversion device, characterized in that: include: A video acquisition unit, used for acquiring a target video; The operation unit is used to perform the following operations on any video frame in the target video to obtain a corresponding three-dimensional video frame: Obtain a target depth map corresponding to the targeted video frame; the target depth map includes depth information of each pixel of the targeted video frame; Inputting the targeted video frame and the target depth map into a pre-trained depth map optimization model to obtain an optimized depth map output by the depth map optimization model; Inputting the optimized depth map and the targeted video frame into a pre-trained viewpoint rendering model to obtain a predicted viewpoint map output by the viewpoint rendering model; The depth map optimization model and the viewpoint rendering model are obtained by training based on the training device according to claim 15; Performing interleaving processing on the targeted video frame and the predicted viewpoint map to generate a corresponding three-dimensional video frame; The synthesis unit is used to synthesize the 3D video frames corresponding to the video frames in the target video to obtain the 3D video.
19. An electronic device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method according to any one of claims 1 to 14.
20. A computer-readable storage medium storing a computer program, characterized in that: The computer program implements the method according to any one of claims 1 to 14 when executed by a processor.