Method for generating multi-view data of landscape painting
By pre-processing, depth estimation and self-supervising perspective generation of traditional Chinese landscape paintings, combined with multi-view synthesis module, the problem of lack of multi-view data in single ancient landscape paintings is solved, and high-quality multi-view synthesis is achieved, which enhances the realism and dynamic expressiveness of the image.
Patent Information
- Application Number
- CN202510122838.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-06
AI Technical Summary
The existing technology has difficulty in generating multi-view data when dealing with traditional Chinese landscape paintings, especially the lack of multi-view data for single ancient landscape paintings, which hinders the multi-view synthesis of ancient paintings based on NeRF method.
A method for generating multi-view data of landscape painting is designed, including image preprocessing, depth estimation, self-supervised perspective generation module and multi-view synthesis module. Through these steps, multi-view images are generated to achieve multi-view synthesis of ancient paintings.
It realizes multi-perspective synthesis of ancient landscape paintings, improves the quality and reality of the composite images, can dynamically change consistent with the user's perspective, and enhances the expressiveness of the work.
Smart Images

Figure BDA0005259273400000021 
Figure BDA0005259273400000022 
Figure BDA0005259273400000051
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of image processing, and in particular relates to a method for generating multi-viewing angle data of a landscape painting. Background Art
[0002] In today's era of booming digital technology, the demand for three-dimensional image expression in fields such as VR / AR technology and driverless cars is growing. Since the introduction of neural radiance field (NeRF), significant progress has been made in image three-dimensional reconstruction and new perspective synthesis, but the early NeRF had problems such as long calculation time, poor generalization, and high requirements for input data. Although there are subsequent improved methods such as Instant-NGP and Mip-NeRF, there is still a problem of lack of multi-perspective data for a single ancient landscape painting, which seriously hinders the multi-perspective synthesis of ancient paintings based on the NeRF method.
[0003] With the rise of AIGC, although it has brought new opportunities for data generation, the existing 3D reconstruction methods for single images still have many limitations when dealing with traditional Chinese paintings, such as high requirements for input data, poor generalization, difficulty in handling complex scenes and maintaining 3D consistency; for example, some methods are limited to the 3D reconstruction of specific objects and are difficult to expand to complex scenes; some methods have poor consistency when generating 3D views and are prone to flickering and blurring in occluded areas; although the method based on multi-plane images (MPI) has certain advantages, it faces problems such as difficulty in depth stratification, over-parameterization and difficulty in ensuring 3D consistency when dealing with the special perspective of landscape paintings and complex scenes.
[0004] Therefore, it is necessary to design an efficient multi-view data generation method for landscape paintings. Summary of the invention
[0005] In order to solve the problems of the prior art, the present invention provides a method for generating multi-view data of landscape painting, comprising the following steps: Step 1: Obtaining image I s Dataset, image I s The images in the dataset are preprocessed as original images to obtain image I p , for image I p Perform depth estimation and generate a high-resolution depth image D s , get the depth image D s Dataset, and the depth image D s Part of the dataset is used as a training set, and the rest is used as a test set;
[0006] Step 2: Image I in step 1 s and its corresponding depth image D s The images in the dataset are used as input; the landscape painting multi-view synthesis network is trained to obtain a training model;
[0007] The landscape painting multi-view synthesis network structure includes a self-supervised view generation module and a multi-view synthesis module. When training the landscape painting multi-view synthesis network, first image I s and its depth image D s As input, the self-supervised view generation module is used to obtain image-depth pairs (I s ,I′ t ,D s ,D′ t ), where image I′ t is the hole image, the depth image D′ t is the depth map of the hole, image I s is the repaired hole image, the depth image D s Depth map of the repaired holes;
[0008] Then, the image-depth pair (I s ,I′ t ,D s ,D′ t ) is input into the multi-view synthesis module to convert the image-depth pair (I s ,I′ t ,D s ,D′ t ) to perform multi-view synthesis and generate a multi-view image I o ;
[0009] Step 3: Input the training set into the training model obtained in step 2 to learn and obtain a landscape painting multi-view synthesis model. The test set is used to test the landscape painting multi-view synthesis model.
[0010] Step 4: Input the landscape painting image into the trained traditional painting image restoration model to obtain a multi-view landscape painting image.
[0011] Furthermore, the self-supervised perspective generation module is used to obtain image-depth pairs (I s ,I′ t ,D s ,D′ t ), including:
[0012] First, in the visual tracing stage, the input image I s and its depth image D s The camera motion model is used to perform perspective transformation operations to obtain an image with holes caused by perspective occlusion. t and the depth image D t ; Using visual backtracking, image I t and the depth image D t Back to image I's and depth image D′ s ;
[0013] In the image restoration stage, the input image I s and its depth image D s Then rotate the camera motion model to generate image I t and the depth image D t , through the image restoration network, image I t and the depth image D t Back to image I' t and depth image D′ t .
[0014] Further, in step 1, image I s The images in the data set are used as original images and preprocessed by using Gaussian difference and image enhancement. The preprocessing of the original images by using Gaussian difference specifically includes: firstly calculating the original image I s The grayscale image average μ is then set to a threshold of ε = μ + 3, where μ represents the average pixel value of the image and 3 is a compensation hyperparameter for noise filtering. γ,ε , set the background to white and retain the depth information, the expression is:
[0015]
[0016] Among them, γ = 255, a ij Represents a pixel in an image, or an element in a matrix; through threshold judgment and mapping calculation, the pixel information regarded as noise is removed, and the background is unified to white, and the image I is obtained after removing irrelevant noise and unifying the background color p ;
[0017] Image enhancement processing for the original image includes: enhancing edge contrast through nonlinear transformation, the expression is:
[0018]
[0019] Among them, φ is the parameter that controls the mapping range, Take 0.015.
[0020] Further, in step 1, the preprocessed image I p and image I i Perform depth estimation, including:
[0021] First, the preprocessed image I p and image I iThe residual connection module in the input depth estimation module is used to extract and enhance features, wherein the depth estimation module is used to generate a depth map of traditional landscape painting, and enhance the high-resolution details of the low-resolution depth image by integrating the local feature information of the high-resolution image into the low-resolution depth map with structural information; the depth estimation module includes a residual connection module, and the residual connection module includes a plurality of residual units connected in sequence, which are used to generate a depth map of the traditional landscape painting ... p and image I i Performing feature extraction and feature enhancement, the residual unit includes a convolution layer, a batch normalization layer and an activation function layer, which are used to gradually extract image features at different levels;
[0022] Then the extracted features are output to the relative depth estimation module, the Midas depth image generation module and the depth measurement interval module to obtain depth-related features from different sources. The depth-related features from different sources are fused through the feature fusion module, and the fused depth feature map is subjected to dimensionality reduction processing through the convolution layer module to generate a depth feature map after dimensionality reduction. Finally, the depth feature map after dimensionality reduction, the high-resolution depth image and the low-resolution depth image are input into the high- and low-resolution fusion module. The image fusion algorithm based on deep learning is used to fuse the local feature information of the high-resolution image after dimensionality reduction and the low-resolution depth estimation result, so as to enhance the high-frequency details of the low-resolution depth image and generate a depth map I that takes into account both structural information and detail information. D .
[0023] Furthermore, the self-supervised perspective generation module is used to generate new perspective images through visual backtracking, including:
[0024] In the estimated depth image D s After that, the new perspective image is generated by perspective backtracking, thereby obtaining the image-depth pair (I s ,I′ t ,D s ,D′ t ), as part of the view traceback input data, is formally expressed as:
[0025] D′ t ,I′ t =Γ(D s ,I s )#(3)
[0026] Where Γ(·) represents the view backtracking algorithm;
[0027] Among them, in the visual tracing stage, the input image I is first s and its depth image D s The camera motion model is used to simulate the change of viewing angle and generate the deformed image I t and the depth image Dt :
[0028] I t ,D t =η (R,t) (I s ,D s )#(4)
[0029] η (R,t) Represents the role of the camera motion matrix (R, t);
[0030] Secondly, using perspective retrieval, image I t and the depth image D t Shrink to image I' s and depth image D′ s .
[0031] Furthermore, the image is inpainted through the image inpainting network, which is used to fill the input image I s ' and depth image D s ′, generating an image I that is consistent with the surrounding visual environment s and the depth image D s ; Wherein, the image restoration network includes an image restoration module and an edge extraction module;
[0032] The image restoration process is as follows: first, the input image I s ' and depth image D s 'The input edge extraction module uses edge detection algorithm to extract image I s ' and depth image D s ′’s edge information to generate an edge image; then the edge image is input into the generation module, which includes a dilated convolution layer and a residual block. The dilated convolution layer is used to expand the receptive field and capture a wider range of contextual information. The residual block is used to learn the deep features of the image and enhance the expression ability of the network. The generation module performs a multi-step extraction of the input image I s ' and depth image D s ' and edge image feature extraction and fusion, generate edge-related feature representation; then send it to the first discrimination module, output the real / non-real result and feature matching loss, and finally image I s ' and depth image D s 'The output result of the first discrimination module is input into the image restoration module to generate the restored image I s and the depth image D s , and input it into the second discrimination module to output a true / false result;
[0033] Among them, the image restoration network uses an enhanced Gaussian difference edge detection algorithm combined with Gaussian blur of different scales to improve the effect of edge detection and adapt to complex line structures. Its expression is as follows:
[0034] ΔG σ,k =G kσ -G σ #(5)
[0035] Among them, G σ and G kσ is a Gaussian blur function with standard deviation σ and kσ, where kσ is a predefined constant used to control the difference between two blurs;
[0036] Image inpainting network fills in the missing areas of the color image I pred , whose resolution is the same as the input image resolution, is expressed as:
[0037] I pred =G 2 (I inc ,C comp )#(6)
[0038] Further, in step 2, multi-view synthesis is performed by a multi-view synthesis module to represent the three-dimensional scene in the landscape painting based on discrete multi-plane images, specifically including: first, the generated image-depth pair (I s ,I′ t ,D s ,D′ t ) is input to the multi-view synthesis module, which extracts depth information from the depth image, segments the original image, and projects it onto multiple planes evenly distributed along the depth, and then generates a multi-view image I through a preset camera posture rotation. o , realizing multi-view synthesis.
[0039] Furthermore, in the process of training the landscape painting multi-view synthesis model, a loss function including reconstruction loss, adversarial loss, perceptual loss and style loss is used for optimization, where:
[0040] Reconstruction loss is used to reduce reconstruction loss by adjusting model parameters. The generated image can retain the artistic style and details of the original as much as possible. The expression is:
[0041]
[0042] Among them, E o represents the restored image features, E gt is the original image feature;
[0043] Adversarial loss is used to motivate the generative network to produce a feature distribution similar to the original image. By gradually optimizing the generated image, it is more consistent with the structure and semantics of the original image. The expression is:
[0044]
[0045] Among them, D represents the discriminator, E gt is the real image feature, E o It is to generate image features;
[0046] Perceptual loss is used to align the high-level semantic features between the repaired image and the original image. By using a pre-trained convolutional neural network to compare the feature differences between the two images in the feature space, the perceptual quality of the repair result is improved. The expression is:
[0047]
[0048] Among them, φ represents the feature extraction layer of the VGG network;
[0049] Style loss is used to ensure that the restored image is consistent with the original image in artistic style. The expression is:
[0050]
[0051] Where G represents the Gram matrix.
[0052] Beneficial effects of the present invention:
[0053] The present invention first pre-processes the landscape painting input into the model by denoising, unifying the background and enhancing the edge contrast, and then the depth prediction generates a depth map by fusing high and low resolution information through the depth estimation module. Next, the landscape painting and its corresponding depth map are processed by backtracking and repairing to generate a pair of images with repaired holes. Among them, the image repair module is combined with the edge detection algorithm for image repair. Subsequently, the multi-view synthesis module is trained using the original image-depth image pair, and discrete multi-plane images are used to represent the three-dimensional scene. Finally, a continuous three-dimensional landscape scene is generated, and a dynamic perspective is introduced to achieve multi-perspective synthesis of a single ancient landscape painting, improve the quality and realism of the synthesized image, and when the user moves the viewpoint, the picture level changes accordingly, giving the work more dynamic expressiveness. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a flow chart of a method for generating multi-viewing angle data of landscape painting according to the present invention;
[0055] Figure 2 It is a flow chart of the depth processing of an image by the CLPDepth module structure adopted by the depth estimation part of the present invention;
[0056] Figure 3 It is a framework diagram of the entire MVSM-CLP including perspective tracing of the present invention;
[0057] Figure 4 Schematic diagram of the comparison between the edge extraction result obtained by using Gaussian difference in the present invention and the traditional Canny method;
[0058] Figure 5 It is a framework diagram of the image restoration network of the present invention;
[0059] Figure 6 is a schematic diagram comparing different depth estimation methods of the present invention;
[0060] Figure 7 It is an ablation result diagram comparing the multi-view synthesis results of the depth estimation module and the restoration module of the present invention;
[0061] Figure 8 , Fig. 9 , Fig.10 There are three groups of qualitative comparative analysis of landscape painting generation results by the method of the present invention and other advanced methods;
[0062] Fig.11 Schematic diagram of the preprocessing process of traditional ancient painting image enhancement according to the present invention. DETAILED DESCRIPTION
[0063] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0064] The present invention provides a method for generating multi-viewing data of a landscape painting, comprising the following steps:
[0065] Step 1: Get image I s Dataset, image I s The images in the dataset are preprocessed as original images to obtain image I p , for image I p Perform depth estimation and generate a high-resolution depth image D s , get the depth image D s Dataset, and the depth image D s Part of the dataset is used as a training set, and the rest is used as a test set;
[0066] Step 2: Image I in step 1 s and its corresponding depth image D sThe images in the dataset are used as input; the landscape painting multi-view synthesis network is trained to obtain a training model;
[0067] The landscape painting multi-view synthesis network structure includes a self-supervised view generation module and a multi-view synthesis module. When training the landscape painting multi-view synthesis network, first image I s and its depth image D s As input, the self-supervised view generation module is used to obtain the image-depth pair I through two stages of visual backtracking and image restoration. s ,I′ t ,D s ,D′ t ), where image I′ t is the hole image, the depth image D′ t is the depth map of the hole, image I s is the repaired hole image, the depth image D s Depth map of the repaired holes;
[0068] Then, the image-depth pair (I s ,I′ t ,D s ,D′ t ) is input into the multi-view synthesis module to convert the image-depth pair (I s ,I′ t ,D s ,D′ t ) to perform multi-view synthesis and generate a multi-view image I o ;
[0069] Step 3: Input the training set into the training model obtained in step 2 to learn and obtain a landscape painting multi-view synthesis model. The test set is used to test the landscape painting multi-view synthesis model.
[0070] Step 4: Input the landscape painting image into the trained traditional painting image restoration model to obtain a multi-view landscape painting image.
[0071] It should be noted that image I s The dataset includes digitized traditional Chinese landscape paintings. The dataset contains 7,324 landscape paintings from different dynasties. The dataset is divided into a training set (14,448 images, including 7,224 ancient paintings and their corresponding 7,224 depth images) and a test set (200 images, including 100 ancient paintings and their corresponding 100 depth images);
[0072] Furthermore, the self-supervised perspective generation module is used to obtain image-depth pairs (I s ,I′ t ,D s ,D′t ), including:
[0073] First, in the visual tracing stage, the input image I s and its depth image D s The camera motion model is used to perform perspective transformation operations to obtain an image with holes caused by perspective occlusion. t and the depth image D t ; Using visual backtracking, image I t and the depth image D t Rotate to image I′ s and depth image D′ s ;
[0074] In the image restoration stage, the input image I s and its depth image D s Then rotate the camera motion model to generate image I t and the depth image D t , through the image restoration network, image I t and the depth image D t Back to image I' t and depth image D′ t .
[0075] Further, in step 1, image I s The images in the data set are used as original images and preprocessed by using Gaussian difference and image enhancement. The preprocessing of the original images by using Gaussian difference specifically includes: firstly calculating the original image I s The grayscale image average μ is then set to a threshold of ε = μ + 3, where μ represents the average pixel value of the image and 3 is a compensation hyperparameter for noise filtering. γ,ε , set the background to white and retain the depth information, the expression is:
[0076]
[0077] Among them, γ = 255, a ij Represents a pixel in an image, or an element in a matrix; through threshold judgment and mapping calculation, the pixel information regarded as noise is removed, and the background is unified to white, and the image I is obtained after removing irrelevant noise and unifying the background color p ; Makes the image visually clearer and more regular, reduces the interference of irrelevant information, and is more conducive to subsequent depth estimation, image restoration and other operations;
[0078] Among them, the original image is preprocessed as follows Fig.11 As shown, I s represents the original image, I pIt represents the result after removing irrelevant noise and unifying the background color;
[0079] Image enhancement processing for the original image includes: enhancing edge contrast through nonlinear transformation, the expression is:
[0080]
[0081] Among them, φ is the parameter that controls the mapping range, Take 0.015. φ is a parameter that controls the mapping range. This parameter is selected to capture the general structure of landscape paintings. This step helps to enhance the visual contrast of lines in landscape paintings, which is beneficial for subsequent depth prediction;
[0082] In the specific implementation, in the dataset construction part, due to the technical limitations of ancient painting protection, the backgrounds of many surviving works contain unwanted noise. After screening out such images, refer to Gaussian difference and try to preprocess the input images by reducing or eliminating irrelevant noise and unifying the background color. Specifically, first calculate the mean value μ of the grayscale image. Then, set the threshold ε=μ+3, where 3 is a hyperparameter for compensation, to remove noise in the image. Secondly, according to the mapping Φ γ,∈ =a ij γ / ∈ sets the background to white and retains the depth information. Where γ = 255, the value of white in RGB. In order to amplify the contrast between lines, a mapping needs to be designed so that points greater than μ should be mapped to positions close to γ, while points less than μ should be mapped to positions close to 0. Based on this, a mapping is designed to amplify the image color contrast, as shown in formula (11).
[0083]
[0084] Where φ is the parameter that controls the mapping range. Here, φ=0.0065 is selected according to the characteristics of landscape painting. The effect of this transformation is as follows Fig.11 shown.
[0085] Further, in step 1, the preprocessed image I p and image I i Depth estimation is performed, specifically including: first, the preprocessed image I p and image I i The residual connection module in the input depth estimation module is used to extract and enhance features, wherein the depth estimation module is used to generate a depth map of traditional landscape painting, and enhance the high-resolution details of the low-resolution depth image by integrating the local feature information of the high-resolution image into the low-resolution depth map with structural information; the depth estimation module includes a residual connection module, and the residual connection module includes a plurality of residual units connected in sequence, which are used to generate a depth map of the traditional landscape painting ... pand image I i Performing feature extraction and feature enhancement, the residual unit includes a convolution layer, a batch normalization layer and an activation function layer, which are used to gradually extract image features at different levels;
[0086] Then the extracted features are output to the relative depth estimation module, the Midas depth image generation module and the depth measurement interval module to obtain depth-related features from different sources. The depth-related features from different sources are fused through the feature fusion module, and the fused depth feature map is subjected to dimensionality reduction processing through the convolution layer module to generate a depth feature map after dimensionality reduction. Finally, the depth feature map after dimensionality reduction, the high-resolution depth image and the low-resolution depth image are input into the high- and low-resolution fusion module. The image fusion algorithm based on deep learning is used to fuse the local feature information of the high-resolution image after dimensionality reduction and the low-resolution depth estimation result, so as to enhance the high-frequency details of the low-resolution depth image and generate a depth map I that takes into account both structural information and detail information. D .
[0087] Among them, the depth estimation module (CLPDepth) is used to pre-process the image I p and image I i Perform depth estimation, set up the depth estimation module, and convert the high-resolution image I p and image I i The local feature information of the deep scene is integrated into the fusion process, and the high-resolution details of the low-resolution depth image are enhanced through multi-resolution fusion. This process also involves adaptively adjusting the depth data so that the depth estimation result can enrich the detail level while retaining the overall scene structure. By combining the detailed high-resolution depth estimation with the structurally consistent low-resolution depth estimation, comprehensive depth information is achieved for the landscape painting. The final depth estimation has both accurate details and coherent scene structure. The overall framework of the depth estimation module is shown in Figure 1. Figure 2 As shown;
[0088] CLPDepth Module uses the existing advanced depth estimation method ZoeDepth to estimate the depth of the input image I, and further enhances the depth estimation result by using the high-low resolution fusion method to obtain a depth map D that matches the ancient painting. Since there is no true value for the depth of the painting, for the blurred ancient painting, the results after image enhancement, depth estimation, and depth enhancement are as follows Figure 6 As shown in the figure, from left to right, the depth estimation results of Depth Anything, ZoeDepth, image enhancement combined with ZoeDepth, ZoeDepth combined with depth map enhancement, and the method of combining the three adopted by the present invention are compared. The qualitative results show that after this process, the effect of depth estimation is more obvious, which is more conducive to subsequent further processing.
[0089] The relative depth information of the input image is estimated by the relative depth estimation module; a relative depth image is generated; the input features are processed based on the Midas depth estimation model by the depth image generation module to generate a Midas depth image with certain depth information;
[0090] The feature is mapped to the depth measurement interval through the neural network layer by the depth measurement interval module to generate a feature representation related to the depth measurement interval. The output ends of the relative depth estimation module, the Midas depth image generation module and the depth measurement interval module are connected to the feature fusion module, which is used to fuse the depth-related features from different sources in a weighted fusion manner; the output end of the feature fusion module is connected to the input end of the convolutional layer module, which is used to perform dimensionality reduction processing on the fused depth feature map to generate a depth feature map after dimensionality reduction. The output end of the convolutional layer module is connected to the high-resolution depth image and the low-resolution depth image and the input end of the high- and low-resolution fusion module, which is used to fuse the local feature information of the high-resolution image after dimensionality reduction with the low-resolution depth estimation result by using an image fusion algorithm based on deep learning, enhance the high-frequency details of the low-resolution depth image, and generate a depth map I that takes into account both structural information and detail information. D .
[0091] It should be noted that the dataset for constructing DepCLP includes digitized traditional Chinese landscape paintings and their corresponding depth estimation maps. This dataset contains 7,324 landscape paintings from different dynasties, and the depth maps corresponding to these images are generated by the method proposed in this paper. The dataset is divided into a training set (14,448 images, including 7,224 ancient paintings and their corresponding 7,224 depth maps) and a test set (200 images, including 100 ancient paintings and their corresponding 100 depth images), which will be used for subsequent training.
[0092] The deep image fusion module adopts an image fusion algorithm based on deep learning (Pix2pix Merge module). It generates a high-quality fused depth image by extracting, matching and fusing the input depth image. The fused depth image not only retains the detail information of the high-resolution depth image, but also integrates the overall structural information of the low-resolution depth image and the depth features predicted by the network, thereby achieving accurate depth estimation of the input image. The Pix2pix Merge module adopts a generative adversarial network architecture, which includes a generator and a discriminator. The generator is used to fuse the input depth image to generate the final depth image, and the discriminator is used to distinguish the difference between the generated depth image and the real depth image. The loss functions used to train this module include adversarial loss, reconstruction loss and perceptual loss. The adversarial loss is used to make the generated depth image as realistic as possible, the reconstruction loss is used to ensure the consistency of the generated depth image with the input depth image at the pixel level, and the perceptual loss is used to make the generated depth image similar to the real depth image in semantics and structure. Through the synergy of these loss functions, the quality and effect of depth image fusion are improved.
[0093] Furthermore, the self-supervised perspective generation module is used to generate new perspective images through visual backtracking, including:
[0094] In the estimated depth image D s After that, the new perspective image is generated by perspective backtracking, thereby obtaining the image-depth pair (I s ,I′ t ,D s ,D′ t ), as part of the view traceback input data, is formally expressed as:
[0095] D′ t ,I′ t =Γ(D s ,I s )#(3)
[0096] Where Γ(·) represents the view backtracking algorithm;
[0097] Among them, in the visual tracing stage, the input image I is first s and its depth image D s The camera motion model is used to simulate the change of viewing angle and generate the deformed image I t and the depth image D t :
[0098] I t ,D t =η (R,t) (I s ,D s )#(4)
[0099] η( R,t ) represents the role of the camera motion matrix (R, t);
[0100] Secondly, using perspective retrieval, image I t and the depth image D t Shrink to image I' s and depth image D′ s .
[0101] It should be noted that if Figure 6 As shown in the figure, the multi-view dataset of traditional paintings is processed in two stages. In the first visual tracing stage, the input image I s and its depth image D s , generate an image I under a reasonable camera motion (R, t) t and the depth image D t ; Secondly, use the perspective tracing to transform image I t and the depth image D t Shrink to image I' s and depth image D′ s Since the camera movement will produce some invisible edge holes, the distribution of these holes is very different from the random masks used by the general repair network. Therefore, we use image I′ s and depth image D′ s , as input, image I s and its depth image D s As the true value, the image restoration module is selected as the basic network model, and a restoration algorithm specifically for the occluded part caused by camera twisting is trained. In the second image restoration stage, a similar retraction idea is also adopted. For the input image I s and its depth image D s , in the same way, generate an image I under a reasonable camera motion t and the depth image D t , use the hole repair network obtained in the first stage to fill the holes in this stage and obtain the image I′ under the camera’s perspective t and depth image D′ t , obtained training data from two perspectives, solving the problem of a single figure lacking pairing from other perspectives. Through supervised training, the continuity and consistency of 3D scene construction are further optimized. A multi-perspective dataset of traditional paintings is constructed.
[0102] Image I s and its depth image D s In the input restoration module, the camera motion model is used again to generate image I s ′ and depth image D s′, and use the image restoration module combined with the XDoG edge detection algorithm to process the occluded area for image restoration to generate the original view image I s and the depth image D s ; This process facilitates the creation of multi-view datasets by overcoming the challenge of insufficient paired viewpoints in a single image. Through self-supervised learning, viewpoint tracing improves the continuity and consistency of 3D scene reconstruction and generates a dedicated multi-view dataset for traditional painting; The image restoration module LInpainting: Figure 5 As shown in Figure 1, Chinese landscape paintings are known for their fine line composition, and using line drawing priors or edge content priors can significantly improve image restoration. In order to deal with the complex texture and random hole problems in multi-view synthesis, the LInpainting module combines the XDoG edge detection algorithm for image restoration.
[0103] Furthermore, the image is inpainted through the image inpainting network, which is used to fill the input image I s ' and depth image D s ′, generating an image I that is consistent with the surrounding visual environment s and the depth image D s ; Wherein, the image restoration network includes an image restoration module and an edge extraction module;
[0104] The image restoration process is as follows: first, the input image I s ' and depth image D s 'The input edge extraction module uses edge detection algorithm to extract image I s ' and depth image D s ′’s edge information to generate an edge image; then the edge image is input into the generation module, which includes a dilated convolution layer and a residual block. The dilated convolution layer is used to expand the receptive field and capture a wider range of contextual information. The residual block is used to learn the deep features of the image and enhance the expression ability of the network. The generation module performs a multi-step extraction of the input image I s ' and depth image D s ' and edge image feature extraction and fusion, generate edge-related feature representation; then send it to the first discrimination module, output the real / non-real result and feature matching loss, and finally image I s ′ and depth image D s 'The output result of the first discrimination module is input into the image restoration module to generate the restored image I s and the depth image D s , and input it into the second discrimination module to output a true / false result;
[0105] Among them, the image restoration network uses an enhanced Gaussian difference edge detection algorithm combined with Gaussian blur of different scales to improve the effect of edge detection and adapt to complex line structures. Its expression is as follows:
[0106] ΔG σ,k =G kσ -G σ #(5)
[0107] Among them, G σ and G kσ is a Gaussian blur function with standard deviation σ and kσ, where kσ is a predefined constant used to control the difference between two blurs;
[0108] Image inpainting network fills in the missing areas of the color image I pred , whose resolution is the same as the input image resolution, is expressed as:
[0109] I pred =G 2 (I inc ,C comp )#(6).
[0110] It should be noted that the image restoration network includes an image restoration module and an edge extraction module. First, the input image I s ' and depth image D s The input edge extraction module uses an edge detection algorithm (XDoG: Extended Difference of Gaussians) to extract image I s ' and depth image D s ' edge information, generate XDoG edge image; the output end of the edge extraction module (XDoG) is connected to the generation module (Edge-G) to generate edge-related feature maps. The generation module (Edge-G) includes a dilated convolution layer and a residual block. The dilated convolution layer is used to expand the receptive field and capture a wider range of contextual information. The residual block is used to learn the deep features of the image and enhance the expression ability of the network. The generation module performs a multi-layered ... s ' and depth image D s ' and edge image feature extraction and fusion, generate edge-related feature representation; the output end of the generation module is connected to the first discrimination module (Edge-D), which is used to discriminate the difference between the generated edge feature map and the real edge feature map, and output the real / fake result (Real / Fake) and the feature matching loss; Image I s ' and depth image D s'And the output end of the first discrimination module is connected to the image restoration module (Inpainting-G), which is used to s ' and depth image D s ′, for image I s ' and depth image D s The missing area or the area that needs to be repaired in ′ is repaired to generate the repaired image I s and the depth image D s The output end of the image restoration module is connected to the second discrimination module (Inpainting-D) to discriminate the difference between the restored image and the real image and output the real / fake result (Real / Fake). It should be noted that by extracting and discriminating the features of the restored image, the image restoration module is prompted to generate a restoration result that is closer to the real image, thereby improving the quality and fidelity of the image restoration.
[0111] When the image restoration network performs image restoration, the input image I s ′ and depth image D s 'Apply two Gaussian blurs with different standard deviations, the smaller σ retains more details, while the larger kσ removes more details; in the restoration stage, the image restoration network uses the incomplete color image I inc As input, and through the synthetic edge map C comp Conditional processing is performed, and the synthetic edge map is constructed by combining the background area of the ground truth edge map with the edges of the damaged area generated in the visual tracing stage; the image inpainting network fills in the color image I of the missing area pred , whose resolution is the same as the input image resolution, and the expression is: pred =G 2 (I inc ,C comp )#(6);Therefore, the image restoration (LInpainting) module combined with the XDoG edge detection algorithm can efficiently and accurately repair the missing parts in the landscape painting and generate a high-quality multi-view image dataset;
[0112] Furthermore, in step 2, multi-view synthesis is performed by a multi-view synthesis module to represent the three-dimensional scene in the landscape painting based on discrete multi-plane images, specifically including:
[0113] First, the generated image-depth pair (I s ,I′ t ,D s ,D′ t) is input to the multi-view synthesis module, which extracts depth information from the depth image, segments the original image, and projects it onto multiple planes evenly distributed along the depth, and then generates a multi-view image I through a preset camera posture rotation. o , realizing multi-view synthesis.
[0114] It should be noted that since a single-view image cannot fully represent the 3D information of a scene, a discrete multi-plane image is used to represent the 3D scene in a landscape painting. First, the generated image-depth pair (I s ,I′ t ,D s ,D′ t ) is input to the multi-view synthesis module. The multi-view synthesis module is used to obtain the depth image D s Extract the depth information from the original image I s The image is segmented and projected onto multiple planes evenly distributed along the depth, thus representing the 3D scene as a series of parallel planes within the field of view. Figure 3 As shown in the figure. The input landscape painting is first preprocessed, and then the depth image is generated by the CLPDepth module. Then, the landscape painting and its corresponding depth image are processed by view backtracking to generate a pair of images with repaired holes. Subsequently, the multi-view synthesis module uses the depth image to perform multi-plane image processing and generate an MPI representation. Finally, multi-view synthesis is achieved through the preset camera posture rotation;
[0115] Furthermore, in the process of training the landscape painting multi-view synthesis model, a loss function including reconstruction loss, adversarial loss, perceptual loss and style loss is used for optimization, where:
[0116] Reconstruction loss is used to reduce reconstruction loss by adjusting model parameters. The generated image can retain the artistic style and details of the original as much as possible. The expression is:
[0117]
[0118] Among them, E o represents the restored image features, E gt is the original image feature;
[0119] Adversarial loss is used to motivate the generative network to produce a feature distribution similar to the original image. By gradually optimizing the generated image, it is more consistent with the structure and semantics of the original image. The expression is:
[0120]
[0121] Among them, D represents the discriminator, E gtis the real image feature, E o is the generated image feature. This loss function evaluates the discriminator's ability to distinguish between real images and generated images, and encourages the generator to generate more realistic images.
[0122] Perceptual loss is used to align the high-level semantic features between the repaired image and the original image. By using a pre-trained convolutional neural network to compare the feature differences between the two images in the feature space, the perceptual quality of the repair result is improved. The expression is:
[0123]
[0124] Where φ represents the feature extraction layer of the VGG network. This loss function measures the difference in feature representation between the generated image and the original image, ensuring that high-level perceptual and semantic content is maintained during the reconstruction process.
[0125] Style loss is used to ensure that the restored image is consistent with the original image in artistic style. The expression is:
[0126]
[0127] Where G represents the Gram matrix;
[0128] The multi-view synthesis method for Chinese landscape painting (MVSM-CLP) can perform depth estimation and multi-view synthesis on traditional ancient paintings with good generalization.
[0129] Effect test: First, qualitative comparison: Compared with MINE, VMPI, 3D Photo, AdaMPI and other models, the MVSM-CLP model has obvious advantages in retaining details and ensuring natural transition of depth relationships;
[0130] Since the generated content is the multi-view result of a single landscape painting, the video stream is pre-made and output as a video format. MINE, VMPI, AdaMPI and MVSM-CLP generate the first frame in the sequence and compare it with the groundtruth. Considering the characteristics of the multi-view views generated by 3D Photo technology, a specific frame in the sequence, namely the view at the 4th / 10th position, is selected as a representative to ensure the comprehensiveness and fairness of the comparison. Figure 8 Fig. 9 and Fig.10The generated results of landscape paintings are shown for comparison, aiming to intuitively show the performance differences of different models when processing such artworks. Through detailed observation and comparison, it is obvious that MINE has certain limitations in generating landscape paintings. Specifically, Fig. 9 and Fig.10 As shown in the figure, the images generated by MINE exhibit significant texture blurring, especially unexpected blurring at the edges. This is due to the failure of MINE to fully capture the complex texture of landscape paintings during feature extraction, and the inadequacy of its depth prediction algorithm in processing details. Although VMPI successfully generates multi-view images, the generated images are often too bright and lack the rustic tones and textures unique to traditional landscape paintings, reducing the artistic appeal of the works. More importantly, the multi-view content generated by VMPI appears to be conservative in terms of camera displacement, with small changes in views when observed from different angles, which reduces the richness and immersion of the multi-view experience.
[0131] About 3D Photo Model, Figure 8 The model shows a significant deficiency in detail preservation. The model tends to over-sharpen edges, which can enhance image clarity in some scenes but often results in loss of detail, especially when dealing with fine textures. This trade-off is detrimental to the overall natural appearance of the image.
[0132] like Figure 8 As shown in Figure 2, when the image contains a large number of irregular holes, the restoration effect of the AdaMPI model fails to reach the expected ideal level. Specifically, the model fails to fully recover the lost texture details when filling these holes, resulting in poor overall image consistency. This shows that although the AdaMPI model can provide satisfactory results when processing simple or less textured landscape images, such as Fig.10 As shown in the figure, the effect of its repair algorithm is significantly weakened when facing scenes with dense textures and complex structures.
[0133] In contrast, the MVSM-CLP model shows significant advantages. By setting up the filling model and depth prediction module, the rich details of the landscape painting are effectively preserved during the generation process, while ensuring the natural transition of the depth relationship, which is closely aligned with human visual perception. These advantages are particularly prominent in the multi-view synthesis task, and the landscape paintings generated by MVSM-CLP appear more coherent and realistic.
[0134] Second, quantitative evaluation: widely recognized image quality metrics, including peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and perceptual image quality evaluation (LPIPS), are used to systematically evaluate the visual quality and authenticity of the generated images. The image quality under new perspectives is evaluated using the currently popular image quality evaluation indicators LPIPS, SSIM, and PSNR. Among them, LPIPS is a measure of the perceptual similarity between images (Learned Perceptual Image Patch Similarity). It is a perceptual similarity index between images learned by neural networks. Lower LPIPS values indicate that the images are closer and more similar. SSIM is the structural similarity index (Structural Similarity Index), which is used to measure the similarity between two images. PSNR is the peak signal-to-noise ratio (Peak Signal-to-Noise Ratio), which is used to measure the reconstruction quality of the image. The higher the PSNR value, the better the image quality, which is usually used to evaluate the performance of the compression algorithm. In addition, consider introducing FID to measure the realism of the generated images. Combined with multiple indicators, the generated multi-view images are evaluated to measure the performance of the model. The PSNR, SSIM and LPIPS are compared with the current advanced multi-view synthesis methods. As shown in Table 1, the model shows excellent advantages in various evaluation indicators, surpassing existing similar technologies.
[0135] Table 1 Quantitative comparison results with advanced multi-view synthesis methods
[0136] Method PSNR(↑) SSIM(↑) LPIPS(↓) MINE 17.028 0.655 0.639 VMPI 13.638 0.550 0.364 3DPhoto 21.389 0.876 0.205 AdaMPI 23.704 0.799 0.183 MVSM-CLP 29.224 0.926 0.145
[0137] In terms of PSNR, an indicator that measures the degree of image distortion, the model achieved 29.22448dB, which is 23% higher than the second-ranked AdaMPI. This excellent result shows that the difference between the synthesized image and the original image is minimal, which strongly confirms the high fidelity effect achieved in the multi-view image synthesis task. In terms of SSIM, a high score of 0.92680 was achieved, which is 5.7% higher than the 3D Photo ranked first. This significant improvement reflects the high degree of restoration of the structural details of the generated image, which is almost seamless with the real image. For LPIPS, the model achieved a low value of 0.14503, which is 21.1% lower than AdaMPI. The low value of LPIPS not only indicates that the image of the present invention performs outstandingly under objective evaluation standards, but also is closer to the authenticity of human eye perception in subjective visual experience, demonstrating the absolute advantage of this method in visual fidelity.
[0138] In summary, the experimental results fully demonstrate the outstanding contribution of the model in the field of multi-view image synthesis, and have achieved comprehensive leadership both in terms of technical accuracy and visual experience.
[0139] To verify the effectiveness of this method, ablation experiments were conducted on the restoration module and depth prediction module, and their multi-view synthesis was used to perform qualitative and quantitative analysis. Figure 7 The qualitative results of the depth estimation module (CLPDepth) for multi-view synthesis are shown. For the blurred ancient painting, the results after image enhancement, depth estimation, and depth enhancement are shown in the figure. The figure compares the depth perception (Depth Anything), the monocular depth estimation method (ZoeDepth), image enhancement combined with ZoeDepth, ZoeDepth combined with depth map enhancement, and the depth estimation results of CLPDepth used in this study.
[0140] As can be seen from the figure, using ZoeDepth alone and ZoeDepth combined with depth map enhancement will cause a large range of image contours to be lost, and the boundary distinction is poor. For darker images, such as Figure 7 In the middle and lower part, DepthAnything can show the general outline of the image while the details disappear completely. For lighter images, such as Figure 7 In the upper middle part, the outline and boundary of the original image are basically lost. However, the proposed CLPDepth method retains the outline of the image better and has more details than the results obtained by the previous four methods.
[0141] Table 2 Ablation study of network design
[0142] Method PSNR(↑) SSIM(↑) LPIPS(↓) w / o CLPDepth 23.604 0.783 0.186 w / o LInpainting 21.763 0.782 0.181 MVSM-CLP 29.224 0.926 0.145
[0143] In order to further quantify the effectiveness of the CLPDepth module, a series of image quality indicators are used for evaluation, as shown in Table 2. The experimental results show that there are significant improvements in all test samples. Compared with the Pre-ZeoDepth results, the MVSM-CLP model improves PSNR by 23.8%, SSIM by 18.3%, and LPIPS by 22%, which proves the key role of the CLPDepth module in improving the restoration quality and visual effects. In summary, the improvement of the CLPDepth model in depth prediction accuracy, especially in edge detail processing, significantly improves the realism and visual quality of multi-view landscape paintings. This not only verifies the advantages of the model in depth estimation, but also provides a solid foundation for subsequent multi-view image synthesis.
[0144] In order to evaluate the role of the image inpainting network (LInpainting Module) in inpainting the hole areas generated when the view is rotated, a detailed comparative experiment was conducted. A control group without the integrated LInpainting Module was selected, and the landscape paintings generated by it were compared with the MVSM-CLP model. Through a careful visual inspection of the multi-view landscape paintings generated by the two models, a significant difference was observed: when the view angle changes, the model without the LInpainting Module appears stiff and unnatural in the inpainting effect of the hole area. This unnaturalness is mainly reflected in the lack of smooth transition between the inpainted texture and the surrounding environment, as well as the low color matching, which makes the inpainted area visually abrupt. This phenomenon is particularly evident in complex texture areas, such as the surface of rocks and the branches of trees.
[0145] On the contrary, when the LInpainting Module is introduced into the MVSM-CLP model, it shows excellent ability in repairing hole areas. The repaired texture is not only highly integrated with the surrounding natural landscape, but also achieves almost seamless connection in color and details, greatly improving the overall quality of the image. It is noted that this improvement is particularly reflected in the processing of scenes with strong contrast and complex structures. The LInpainting Module can effectively restore the missing information while maintaining visual continuity and authenticity.
[0146] As shown in Table 2, in order to further quantify the effectiveness of the LInpainting module, a series of image quality indicators are used for evaluation. Experimental results show that in all test samples, the MVSM-CLP model shows significant improvements in these indicators. Specifically, compared with the model without the LInpainting module, it improves by 34.3% on PSNR, 18.4% on SSIM, and 20.0% on LPIPS. This confirms the key role of the image restoration (LInpainting) module in improving the restoration quality and visual effects. In summary, the integration of the LInpainting module not only solves the challenge of hole filling during view conversion, but also significantly improves the overall visual quality of multi-view landscape painting synthesis, demonstrating its indispensable value in the MVSM-CLP model;
[0147] Quantitative evaluation: Using PSNR, SSIM, LPIPS and other indicators for evaluation, the MVSM-CLP model surpasses existing similar technologies in various indicators. The ablation experiment also verifies the effectiveness of the CLPDepth and LInpainting modules.
[0148] It is worth noting that the method for generating multi-view data of landscape paintings of the present invention first pre-processes the input landscape painting by denoising, unifying the background and enhancing the edge contrast, and then the depth estimation is performed by fusing high and low resolution information through the CLPDepth module to generate a depth map. Next, the landscape painting and its corresponding depth map are processed by the self-supervised perspective generation module in two stages of hole image generation and hole repair to generate a pair of images with repaired holes. Among them, the image repair uses the LInpainting module combined with the XDoG edge extraction algorithm. Subsequently, the multi-view synthesis module uses the image pair of the original image-depth map to perform perspective synthesis processing based on multi-plane images to generate a continuous multi-plane image representation. Finally, through the preset camera posture rotation, multi-view images are generated to achieve multi-view synthesis and improve the quality and realism of the synthesized image; and the corresponding loss functions (reconstruction, adversarial, perception, style loss) are given to optimize the model;
[0149] It should be noted that this invention synthesizes multi-view data for the Chinese landscape painting dataset for the first time, realizing a depth estimation and image restoration method that is more suitable for traditional solidification. Good semantic segmentation results were achieved while maintaining high speed and low hardware configuration. As a result, our model PSNR reached 29.224, confirming the high fidelity effect we achieved in the multi-view image synthesis task. The score of 0.926 on SSIM reflects the high restoration of the structural details of the generated image, which is almost seamless with the real image. LPIPS reached 0.14503, indicating that our images performed outstandingly under objective evaluation criteria, and were closer to the authenticity of human eye perception in subjective visual experience. It has promoted the inheritance, promotion, restoration and protection of the cultural heritage of traditional Chinese landscape paintings.
[0150] This method uses multi-plane image MPI technology to perform multi-view synthesis of a single landscape painting, improving the quality and realism of the synthesized image. First, depth estimation (CLPDepth Module) lays a solid foundation for multi-view synthesis by accurately predicting the depth information of the landscape painting. Secondly, image restoration (LInpainting Module) effectively solves the problem of visual discontinuity that occurs during the multi-view synthesis process, and ensures the integrity and coherence of the final synthesized image through intelligent restoration technology. Finally, the view tracing mechanism not only enhances the natural distortion effect of the synthesized perspective, but also significantly improves the consistency between views, thereby generating a more realistic and immersive multi-view experience.
[0151] Experimental verification: By comparing with MINE, VMPI, 3D Photo, AdaMPI and other methods, the MVSM-CLP model performs better in retaining details, colors, depth transitions and visual fidelity when generating multi-view images of landscape paintings. Ablation experiments show that the CLPDepth and LInpainting modules play a key role in improving image quality and visual effects in depth prediction and hole repair, respectively, and all evaluation indicators confirm the advantages of the model.
[0152] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for generating multi-view data of landscape painting, characterized in that: The following steps are involved: Step 1: Get image I s Dataset, image I s The images in the dataset are preprocessed as original images to obtain image I p , for image I p Perform depth estimation and generate a high-resolution depth image D s , get the depth image D s Dataset, and the depth image D s Part of the dataset is used as a training set, and the rest is used as a test set; Step 2: Image I in step 1 s and its corresponding depth image D s The images in the dataset are used as input; the landscape painting multi-view synthesis network is trained to obtain a training model; The landscape painting multi-view synthesis network structure includes a self-supervised view generation module and a multi-view synthesis module. When training the landscape painting multi-view synthesis network, first image I s and its depth image D s As input, the self-supervised view generation module is used to obtain image-depth pairs (I s ,I′ t ,D s ,D′ t ), where image I′ t is the hole image, the depth image D′ t is the depth map of the hole, image I s is the repaired hole image, the depth image D s Depth map of the repaired holes; Then, the image-depth pair (I s ,I′ t ,D s ,D′ t ) is input into the multi-view synthesis module to convert the image-depth pair (I s ,I′ t ,D s ,D′ t ) to perform multi-view synthesis and generate a multi-view image I o ; Step 3: Input the training set into the training model obtained in step 2 to learn and obtain a landscape painting multi-view synthesis model. The test set is used to test the landscape painting multi-view synthesis model. Step 4: Input the landscape painting image into the trained traditional painting image restoration model to obtain a multi-view landscape painting image.
2. The method for generating multi-view data of landscape painting according to claim 1, characterized in that: The self-supervised view generation module is used to obtain image-depth pairs (I s ,I′ t ,D s ,D′ t ), including: First, in the visual tracing stage, the input image I s and its depth image D s The camera motion model is used to perform perspective transformation operations to obtain an image with holes caused by perspective occlusion. t and the depth image D t ; Using visual backtracking, image I t and the depth image D t Back to image I' s and depth image D′ s ; In the image restoration stage, the input image I s and its depth image D s Then rotate the camera motion model to generate image I t and the depth image D t , through the image restoration network, image I t and the depth image D t Back to image I' t and depth image D′ t .
3. The method for generating multi-view data of landscape painting according to claim 1, characterized in that: In step 1, image I s The images in the data set are used as original images and preprocessed by using Gaussian difference and image enhancement. The preprocessing of the original images by using Gaussian difference specifically includes: firstly calculating the original image I s The grayscale image average μ is then set to a threshold of ε = μ + 3, where μ represents the average pixel value of the image and 3 is a compensation hyperparameter for noise filtering. γ,ε , set the background to white and retain the depth information, the expression is: Among them, γ = 255, a ij Represents a pixel in an image, or an element in a matrix; through threshold judgment and mapping calculation, the pixel information regarded as noise is removed, and the background is unified to white, and the image I is obtained after removing irrelevant noise and unifying the background color p ; Image enhancement processing for the original image includes: enhancing edge contrast through nonlinear transformation, the expression is: Among them, φ is the parameter that controls the mapping range, Take 0.
015.
4. The method for generating multi-view data of landscape painting according to claim 1, characterized in that: In step 1, the preprocessed image I p and image I i Perform depth estimation, including: First, the preprocessed image I p and image I i The residual connection module in the input depth estimation module is used to extract and enhance features, wherein the depth estimation module is used to generate a depth map of traditional landscape painting, and enhance the high-resolution details of the low-resolution depth image by integrating the local feature information of the high-resolution image into the low-resolution depth map with structural information; the depth estimation module includes a residual connection module, and the residual connection module includes a plurality of residual units connected in sequence, which are used to generate a depth map of the traditional landscape painting ... p and image I i Performing feature extraction and feature enhancement, the residual unit includes a convolution layer, a batch normalization layer and an activation function layer, which are used to gradually extract image features at different levels; Then the extracted features are output to the relative depth estimation module, the Midas depth image generation module and the depth measurement interval module to obtain depth-related features from different sources. The depth-related features from different sources are fused through the feature fusion module, and the fused depth feature map is subjected to dimensionality reduction processing through the convolution layer module to generate a depth feature map after dimensionality reduction. Finally, the depth feature map after dimensionality reduction, the high-resolution depth image and the low-resolution depth image are input into the high- and low-resolution fusion module. The image fusion algorithm based on deep learning is used to fuse the local feature information of the high-resolution image after dimensionality reduction and the low-resolution depth estimation result, so as to enhance the high-frequency details of the low-resolution depth image and generate a depth map I that takes into account both structural information and detail information. D .
5. The method for generating multi-view data of landscape painting according to claim 1, characterized in that: The self-supervised perspective generation module is used to generate new perspective images through visual backtracking, including: In the estimated depth image D s After that, the new perspective image is generated by perspective backtracking, thereby obtaining the image-depth pair (I s ,I′ t ,D s ,D′ t ), as part of the view traceback input data, is formally expressed as: D′ t ,I′ t =Γ(D s ,I s )#(3) Where Γ(·) represents the view backtracking algorithm; Among them, in the visual tracing stage, the input image I is first s and its depth image D s The camera motion model is used to simulate the change of viewing angle and generate the deformed image I t and the depth image D t : I t ,D t =η (R,t) (I s ,D s )#(4) η (R,t) Represents the role of the camera motion matrix (R, t); Secondly, using perspective retrieval, image I t and the depth image D t Shrink to image I' s and depth image D′ s .
6. The method for generating multi-view data of landscape painting according to claim 1, characterized in that: Image inpainting is performed through an image inpainting network, which is used to fill in the input image I s ' and depth image D s ′, generating an image I that is consistent with the surrounding visual environment s and the depth image D s ; Wherein, the image restoration network includes an image restoration module and an edge extraction module; The image restoration process is as follows: first, the input image I s ' and depth image D s 'The input edge extraction module uses edge detection algorithm to extract image I s ' and depth image D s ′’s edge information to generate an edge image; then the edge image is input into the generation module, which includes a dilated convolution layer and a residual block. The dilated convolution layer is used to expand the receptive field and capture a wider range of contextual information. The residual block is used to learn the deep features of the image and enhance the expression ability of the network. The generation module performs a multi-step extraction of the input image I s ' and depth image D s ' and edge image feature extraction and fusion, generate edge-related feature representation; then send it to the first discrimination module, output the real / non-real result and feature matching loss, and finally image I s ' and depth image D s 'The output result of the first discrimination module is input into the image restoration module to generate the restored image I s and the depth image D s , and input it into the second discrimination module to output a true / false result; Among them, the image restoration network uses an enhanced Gaussian difference edge detection algorithm combined with Gaussian blur of different scales to improve the effect of edge detection and adapt to complex line structures. Its expression is as follows: ΔG σ,k =G kσ -G σ #(5) Among them, G σ and G kσ is a Gaussian blur function with standard deviation σ and kσ, where kσ is a predefined constant used to control the difference between two blurs; Image inpainting network fills in the missing areas of the color image I pred , whose resolution is the same as the input image resolution, is expressed as: I pred =G2(I inc ,C comp )#(6)。 7. The method for generating multi-view data of landscape painting according to claim 1, characterized in that: In step 2, a multi-view synthesis module is used to perform multi-view synthesis, and a three-dimensional scene in a landscape painting is represented based on a discrete multi-plane image. Specifically, the steps include: first, the generated image-depth pair (I s ,I′ t ,D s ,D′ t ) is input to the multi-view synthesis module, which extracts depth information from the depth image, segments the original image, and projects it onto multiple planes evenly distributed along the depth, and then generates a multi-view image I through a preset camera posture rotation. o , realizing multi-view synthesis.
8. The method for generating multi-view data of landscape painting according to claim 1, characterized in that: In the process of training the landscape painting multi-view synthesis model, a loss function including reconstruction loss, adversarial loss, perceptual loss and style loss is used for optimization, where: Reconstruction loss is used to reduce reconstruction loss by adjusting model parameters. The generated image can retain the artistic style and details of the original as much as possible. The expression is: Among them, E o represents the restored image features, E gt is the original image feature; Adversarial loss is used to motivate the generative network to produce a feature distribution similar to the original image. By gradually optimizing the generated image, it is more consistent with the structure and semantics of the original image. The expression is: Among them, D represents the discriminator, E gt is the real image feature, E o It is to generate image features; Perceptual loss is used to align the high-level semantic features between the repaired image and the original image. By using a pre-trained convolutional neural network to compare the feature differences between the two images in the feature space, the perceptual quality of the repair result is improved. The expression is: Among them, φ represents the feature extraction layer of the VGG network; Style loss is used to ensure that the restored image is consistent with the original image in artistic style. The expression is: Where G represents the Gram matrix.