Ink and wash painting depth estimation method based on diffusion model
Through the combination of diffusion model and depth estimation model, the accuracy of ink painting depth estimation is solved, and the precise capture and detail retention of ink painting depth information is achieved, and the estimation accuracy and generalization ability are improved.
Patent Information
- Application Number
- CN202510445744.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-08
AI Technical Summary
The existing traditional depth estimation methods are difficult to accurately obtain the depth information of ink painting. The traditional methods rely on the clear contours and geometric shapes of objects, while the blurred and abstract characteristics of ink painting lead to inaccurate depth estimation results; the method based on convolutional neural network lacks understanding of ink color changes and brushstroke texture when dealing ink painting.
The diffusion model is used for style transfer, combined with the depth estimation model, and accurately capture the depth information of ink painting through steps such as feature coding, noise addition, residual operation and attention mechanism.
It significantly improves the accuracy and generalization ability of ink painting depth estimation, and can efficiently and accurately estimate depth information with few samples, avoiding the loss of detailed information.
Smart Images

Figure CN120279075A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image depth estimation in computer vision, and specifically relates to a method for estimating the depth of Chinese ink paintings based on a diffusion model. Background Art
[0002] As a precious traditional Chinese art, Chinese ink paintings demonstrate unique artistic charm with their distinct ink rhythms, brushstrokes, and artistic conceptions. In the current era of rapid digital development, estimating the depth of Chinese ink paintings is of great significance. It not only helps with the digital preservation, restoration, and display of Chinese ink paintings but also provides crucial depth information support for creative works and artistic research based on Chinese ink paintings. However, accurately obtaining the depth information of Chinese ink paintings poses many challenges. Chinese ink paintings mainly create a sense of layering and spatiality in the picture through the changes in ink density, the density of brushstrokes, and unique composition methods. Unlike realistic paintings or photographic works, they do not have obvious geometric features and light and shadow cues that can be directly used for depth analysis.
[0003] Traditional depth estimation methods have obvious limitations when dealing with Chinese ink paintings. Depth estimation methods based on geometric feature analysis often rely on clear outlines, definite geometric shapes, and quantifiable spatial relationships of objects to infer depth. However, the freehand style of Chinese ink paintings makes the object forms appear blurred and abstract, and the outlines and lines do not strictly follow realistic geometric rules, resulting in the difficulty of effectively capturing the depth information under the unique artistic expressions of Chinese ink paintings, and the estimated results deviate greatly from the actual depth. Depth estimation methods based on general machine learning, such as common convolutional neural network (CNN) models, although they have achieved certain results in the field of general image depth estimation, also have many deficiencies when applied to Chinese ink paintings. On the one hand, general image datasets are difficult to cover the unique artistic characteristics of Chinese ink paintings, such as the changes in ink color and the texture of brushstrokes, making the trained models not deeply and accurately understand the characteristics of Chinese ink paintings. On the other hand, the subtle ink color gradients, brushstroke overlays, and the resulting complex spatial relationships in Chinese ink paintings exceed the scope that can be accurately processed by conventional machine learning models based on general feature learning, thus leading to poor accuracy in depth estimation. With the continuous development of deep learning technology, diffusion models have demonstrated powerful capabilities in fields such as image style transfer. It can effectively capture and transfer the style features of images, providing new ideas for subsequent more accurate processing. At the same time, depth estimation models have learned some feature patterns related to image depth to a certain extent.
[0004] Through the above analysis of the research status, existing traditional depth estimation methods are difficult to meet the special requirements of ink painting depth estimation. The advantages of diffusion models in style transfer and the existing feature learning foundation of depth estimation models provide an opportunity to design an ink painting depth estimation method based on diffusion models. The method proposed in this patent, that is, first using a diffusion model for style transfer and then using a depth estimation model for depth estimation, can more accurately obtain the depth information of ink paintings, thus providing strong support for applications such as digitization and virtual reality related to ink paintings. Summary of the Invention
[0005] To overcome the deficiencies in existing solutions, the object of the present invention is to provide an ink painting depth estimation method based on a diffusion model, including the following steps:
[0006] Step 1, given an ink painting image I c and a style image I s , respectively obtain the initial latent representation of the ink painting image and the initial latent representation of the style image
[0007]
[0008] where F() is a feature encoder;
[0009] Step 2, in the forward noise addition process, given the initial latent representation of the ink painting image and the initial latent representation of the style image add Gaussian noise through Equation (2),
[0010]
[0011] where is the noisy version of, β t represents a fixed sequence value used to control the noise intensity t represents the diffusion step, and its value range is 1 ≤ t ≤ T, T is the maximum diffusion step, and ε follows a Gaussian distribution with a mean of 0 and a variance of 1;
[0012] Step 3, let α t = 1 - β t , the initial latent representation of the ink painting image obtain the latent noise representation of the ink painting at time T through Equation (3) the style image obtain the latent noise representation of the style at time T through Equation (3)
[0013]
[0014] Step 4: Use the potential noise representation obtained in Step 2 and the potential noise representation to obtain the stylized initial potential noise representation through Equation (4)
[0015]
[0016] where represents the channel-wise standard deviation of the style latent noise at time T, represents the channel-wise mean of the style latent noise at time T, represents the channel-wise standard deviation of the ink-wash painting latent noise at time T, represents the channel-wise mean of the ink-wash painting latent noise at time T;
[0017] Step 5: Use the and obtained in Step 2 and the obtained in Step 4 to obtain the style feature content feature and the stylized feature at time T through Equation (5)
[0018]
[0019] where ResBlk() represents the residual operation;
[0020] Step 6: Operate on the and obtained in Step 5 respectively through Equation (6) to obtain the query of the ink-wash painting feature at time T and the query of the stylized feature at time T
[0021]
[0022] where rearrange() is the rearrangement operation and linear Q () is the linear transformation operation;
[0023] Step 7: Operate on the obtained in Step 6 according to Equation (7) to obtain the key of the style feature at time T and the value of the style feature at time T
[0024]
[0025] where rearrange() is the rearrangement operation and linearK,V () is a linear transformation operation;
[0026] Step 8. Operate on the query obtained in Step 6 and the query according to Equation (8) to obtain the query of the stylized latent noise representation at time T after fusion
[0027]
[0028] where γ is the query retention rate;
[0029] Step 9. Operate on the query obtained in Step 8 and the key and value obtained in Step 7 according to Equation (9) to obtain the output feature of the attention mechanism at time T
[0030]
[0031] where τ is the temperature weight, d is the dimension of the feature vectors of the query and the key, and softmax() is the normalization operation;
[0032] Step 10. Operate on the feature obtained in Step 9 according to Equation (10) to obtain the stylized latent noise representation at time T after fusion
[0033]
[0034] Step 11. Use the ink-wash painting latent noise representation obtained in Step 3 as the initial feature X 0 , perform sequential operations according to Equation (11), then add the time embedding vector and perform sequential operations again, and finally perform a residual connection operation with the input feature X 0 to obtain the feature X 1′ ,
[0035] X 1′ = Circulat(Seq(Seq(X 0 )) + Te(T)) + shortcut(X 0 )) (11)
[0036] where Circulat() is the operation of looping twice, shortcut() is the operation of matching the number of channels, Te() is the time embedding operation, and Seq() is the sequential operation;
[0037] Step 12. According to Equation (12), operate on the feature X obtained in Step 111′ Perform downsampling operation to obtain feature X 1 ,
[0038] X 1 = Downsample(X 1′ ) (12)
[0039] where Downsample() is the downsampling operation;
[0040] Step 13, for the feature X obtained in step 12 1 Recursively repeat the operations of step 11 and step 12 to obtain feature X 2 , then feature X 2 Obtain feature X through equation (13) 3 ,
[0041] X 3 = Circulat(X k-1 ), 2 ≤ k ≤ 3 (13)
[0042] where Circulat() is the operation of circulating twice;
[0043] Step 14, according to equation (14), perform operations on the feature X obtained in step 13 3 to obtain feature X 4 ;
[0044] X 4 = Circulat(Seq(Seq(X 3 ) + Te(T)) + shortcut(X 3 )) (14)
[0045] where Circulat() is the operation of circulating twice, shortcut() is the operation of matching the number of channels, Te() is the time embedding operation, and Seq() is the sequence operation;
[0046] Step 15, according to equation (15), perform operations on the feature X obtained in step 14 4 to obtain feature X middle ,
[0047] X middle = Circulat(Seq(Seq(X 4 ) + Te(T)) + shortcut(X 4 )) (15)
[0048] Then, according to equation (16), perform a concatenation operation on feature X middle and X 4 to obtain the initial decoder feature
[0049]
[0050] where Concat() is the concatenation operation;
[0051] Step 16, for the initial encoder features obtained in Step 15 Perform operations according to Equation (17) to obtain features
[0052]
[0053] Step 17, the features obtained in Step 16 are concatenated along the channel dimension with the features X of the same size corresponding in the encoder path 3 and the operations of Equation (17) are repeated to obtain features Then features are obtained according to Equation (18)
[0054]
[0055] where Upsample() is the upsampling operation;
[0056] Step 18, for the features obtained in Step 17 Recursively repeat the operations of Step 16 to Step 17 twice to obtain the final upsampled features Then, after group normalization and activation by the activation function according to Equation (19), the latent representation noise of the ink painting at time T is obtained through a convolution operation
[0057]
[0058] where Conv2d() is the convolution operation, SiLU() is the activation operation, and GroupNorm() is the group normalization operation;
[0059] Step 19, in the reverse diffusion process, approximately calculate the posterior diffusion conditional probability P(z t-1 |z t , z0) through Equation (20),
[0060]
[0061] where ∝ denotes proportional to, exp is the exponential function, and are all known quantities in Step 2, z t is the latent noise representation at time t, and then z t-1 is obtained according to Equation (21),
[0062]
[0063] The noise obtained through step 18 and the potential noise representation of the Chinese ink painting obtained in step 3 Obtain the potential noise representation of the Chinese ink painting at time T-1 according to Equation (22)
[0064]
[0065] Step 20. For the result obtained in step 3 Repeat the operations from step 11 to step 19 to obtain For the result obtained in step 9 Repeat the operations from step 11 to step 19 to obtain
[0066] Step 21. For the result obtained in step 19 Repeat the operations of step 5, step 6, step 8, step 9 and steps 11 to 19 for T-1 times. For the result obtained in step 20 Repeat the operations of step 5, step 7, step 9 and steps 11 to 19 for T-1 times. For the result obtained in step 20 Repeat the operations of step 5, step 6, step 8, step 9, step 10 and steps 11 to 19 for T-1 times to obtain the stylized potential representation
[0067] Step 22. For the result obtained in step 21 Perform an operation according to Equation (23) to obtain the stylized Chinese ink painting image I cs ,
[0068]
[0069] where D() is the feature decoder;
[0070] Step 23. For I obtained in step 22 cs Perform an operation according to Equation (24) to obtain the stylized image feature Y1,
[0071] Y1 = Downsample(Conv2d(I cs )) (24)
[0072] where Conv2D() is the convolution operation and Downsample() is the downsampling operation;
[0073] Step 24. Iteratively calculate the features Y2, Y3 and Y4 for Y1 obtained in step 23 according to Equation (25)
[0074] Yn = Downsample(Conv2d(I n-1 ), 2 ≤ n ≤ 4 (25)
[0075] where Conv2D() is the convolution operation and Downsample() is the downsampling operation;
[0076] Step 25: Perform operations on Y1, Y2, Y3, and Y4 obtained in Step 24 respectively through Equation (26) to obtain features Y1, Y2, Y3, and Y4,
[0077] Y i = Fusion(Y i ), i ∈ {1, 2, 3, 4} (26)
[0078] where Fusion() is the feature fusion operation;
[0079] Step 26: Perform operations on the features Y1, Y2, Y3, and Y4 obtained in Step 25 respectively through Equation (27) to obtain features Y1', Y2', Y3', and Y4',
[0080] Y i ' = Upsample(Y i ), i ∈ {1, 2, 3, 4} (27)
[0081] where Upsample() is the upsampling operation;
[0082] Step 27: Perform an operation on the features Y1', Y2', Y3', and Y4' obtained in Step 25 through Equation (28) to obtain the multi-scale fusion feature Y RefinedFeatures ,
[0083] Y RefinedFeatures = Aggre(Y1', Y2', Y3', Y4') (28)
[0084] where Aggre() is the feature aggregation operation;
[0085] Step 28: Perform an operation on the aggregated feature Y RefinedFeatures obtained in Step 27 through Equation (29) to obtain the depth map D,
[0086] D = Softplus(Y RefinedFeatures ) (29)
[0087] where Softplus() is the normalization operation.
[0088] Compared with the prior art, the present invention has the following advantages:
[0089] (1) A method for estimating the depth of Chinese ink paintings based on a diffusion model proposed by the present invention can accurately capture the depth information contained in artistic elements such as the shade of ink color and the density of brushstrokes in Chinese ink paintings, avoiding inaccurate depth estimation and loss of detailed information caused by the inability of traditional methods to adapt to the style of Chinese ink paintings, and significantly improving the accuracy of depth estimation of Chinese ink paintings.
[0090] (2) A method for estimating the depth of Chinese ink paintings based on a diffusion model proposed by the present invention has excellent generalization ability, can achieve few-shot learning, and can still efficiently and accurately estimate the depth information under the condition of limited samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] Figure 1 is a flowchart of a method for estimating the depth of Chinese ink paintings based on a diffusion model of the present invention;
[0092] Figure 2 is a network diagram of a method for estimating the depth of Chinese ink paintings based on a diffusion model of the present invention;
[0093] Figure 3 are the Chinese ink painting image and the style image in step 1 of Embodiment 1 of the present invention;
[0094] Figure 4 is the stylized image generated in step 22 of Embodiment 1 of the present invention;
[0095] Figure 5 is the Chinese ink painting depth map generated in step 28 of Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0096] Embodiment 1
[0097] As Figure 1 、 Figure 2 shown, a method for estimating the depth of Chinese ink paintings based on a diffusion model includes the following steps:
[0098] Step 1, given a Chinese ink painting image I c and a style image I s , as Figure 3 shown, the initial latent representation of the Chinese ink painting image and the initial latent representation
[0099]
[0100] of the style image are respectively obtained through Equation (1), where F() is a feature encoder;
[0101] Step 2, in the forward noise addition process, given the initial latent representation of the Chinese ink painting image and the initial latent representation Add Gaussian noise through Equation (2),
[0102]
[0103] where is the noisy version of t represents a fixed sequence value used to control the noise intensity t represents the diffusion step number, whose value range is 1 ≤ t ≤ T, T is the maximum diffusion step number, and ε follows a Gaussian distribution with a mean of 0 and a variance of 1;
[0104] Step 3, let α t = 1 - β t , the initial latent representation of the ink-wash painting image obtain the latent noise representation of the ink-wash painting at time T through Equation (3) the style image obtain the latent noise representation of the style at time T through Equation (3)
[0105]
[0106] Step 4, combine the latent noise representation obtained in Step 2 and the latent noise representation
[0107]
[0108] where represents the channel-wise standard deviation of the latent noise of the style at time T, represents the channel-wise mean of the latent noise of the style at time T, represents the channel-wise standard deviation of the latent noise of the ink-wash painting at time T, represents the channel-wise mean of the latent noise of the ink-wash painting at time T;
[0109] Step 5, combine the and obtained in Step 2 with the obtained in Step 4 to obtain the style feature the content feature and the stylized feature
[0110]
[0111] where ResBlk() represents the residual operation;
[0112] Step 6: Operate on and obtained in Step 5 respectively through Equation (6) to obtain the query of the ink-wash painting features at time T and the query of the stylized features at time T .
[0113]
[0114] where rearrange() is the rearrangement operation and linear Q () is the linear transformation operation;
[0115] Step 7: Operate on obtained in Step 6 according to Equation (7) to obtain the key of the style features at time T and the value of the style features at time T.
[0116]
[0117] where rearrange() is the rearrangement operation and linear K,V () is the linear transformation operation;
[0118] Step 8: Operate on the query and the query obtained in Step 6 according to Equation (8) to obtain the query of the fused stylized latent noise representation at time T.
[0119]
[0120] where γ is the query retention rate;
[0121] Step 9: Operate on the query obtained in Step 8, the key and the value obtained in Step 7 according to Equation (9) to obtain the output feature of the attention mechanism at time T.
[0122]
[0123] where τ is the temperature weight, d is the dimension of the feature vectors of the query and the key, and softmax() is the normalization operation;
[0124] Step 10: Operate on the feature obtained in Step 9 according to Equation (10) to obtain the fused stylized latent noise representation at time T.
[0125]
[0126] Step 11: For the potential noise representation of the ink-wash painting obtained in Step 3 as the initial feature X 0 , perform sequential operations according to Equation (11), then add a time embedding vector and perform sequential operations again, and finally perform a residual connection operation with the input feature X 0 to obtain the feature X 1′ ,
[0127] X 1 ′ = Circulat(Seq(Seq(X 0 )) + Te(T)) + shortcut(X 0 )) (11)
[0128] where Circulat() is an operation of looping twice, shortcut() is an operation of matching the number of channels, Te() is a time embedding operation, and Seq() is a sequential operation;
[0129] Step 12: Perform a downsampling operation on the feature X obtained in Step 11 according to Equation (12) to obtain the feature X 1′ , 1 ,
[0130] X 1 = Downsample(X 1′ ) (12)
[0131] where Downsample() is a downsampling operation;
[0132] Step 13: Recursively repeat the operations of Step 11 and Step 12 on the feature X obtained in Step 12 to obtain the feature X 1 , and then the feature X 2 obtains the feature X through Equation (13) 2 , 3 ,
[0133] X 3 = Circulat(X k-1 ), 2 ≤ k ≤ 3 (13)
[0134] where Circulat() is an operation of looping twice;
[0135] Step 14: Perform an operation on the feature X obtained in Step 13 according to Equation (14) to obtain the feature X 3 , 4 ;
[0136] X 4 = Circulat(Seq(Seq(X3 ) + Te(T)) + shortcut(X 3 )) (14)
[0137] Among them, Circulat() is an operation that loops twice, shortcut() is an operation to match the number of channels, Te() is a time embedding operation, and Seq() is a sequence operation;
[0138] Step 15, perform an operation on the feature X obtained in step 14 according to formula (15) 4 to obtain the feature X middle ,
[0139] X middle = Circulat(Seq(Seq(X 4 ) + Te(T)) + shortcut(X 4 )) (15)
[0140] Then, perform a concatenation operation on the feature X middle and X 4 to obtain the initial decoder feature
[0141]
[0142] Among them, Concat() is a concatenation operation;
[0143] Step 16, perform an operation on the initial encoder feature obtained in step 15 according to formula (17) to obtain the feature
[0144]
[0145] Step 17, concatenate the feature obtained in step 16 with the feature X of the same size corresponding in the encoder path along the channel dimension, and repeat the operation of formula (17) to obtain the feature 3 Then, obtain the feature according to formula (18)
[0146]
[0147] Among them, Upsample() is an upsampling operation;
[0148]
[0148] Step 18, recursively repeat the operations of step 16 to step 17 on the feature obtained in step 17 twice to obtain the final upsampled feature Then, after performing group normalization according to Equation (19) and activation using an activation function, a convolutional operation is carried out to obtain the latent representation noise of the ink-wash painting at time T.
[0149]
[0150] Where Conv2d() is the convolutional operation, SiLU() is the activation operation, and GroupNorm() is the group normalization operation;
[0151] Step 19, in the inverse diffusion process, the posterior diffusion conditional probability P(z t-1 |z t , z0) is approximately calculated through Equation (20),
[0152]
[0153] where ∝ denotes proportional to, and exp is the exponential function, and are all known quantities in Step 2, z t is the latent noise representation at time t, and then z t-1 is obtained according to Equation (21),
[0154]
[0155] the noise obtained through Step 18 and the latent noise representation of the ink-wash painting obtained through Step 3 are used to obtain the latent noise representation of the ink-wash painting at time T-1 according to Equation (22)
[0156]
[0157] Step 20, for the result obtained in Step 3 Repeat the operations from Step 11 to Step 19, and then for the result obtained in Step 9 Repeat the operations from Step 11 to Step 19, and then
[0158] Step 21, for the result obtained in Step 19 Repeat the operations of Step 5, Step 6, Step 8, Step 9, and Step 11 to Step 19 for T-1 times. For the result obtained in Step 20 Repeat the operations of Step 5, Step 7, Step 9, and Step 11 to Step 19 for T-1 times. For the result obtained in Step 20 Repeat the operations of Step 5, Step 6, Step 8, Step 9, Step 10, and Step 11 to Step 19 for T-1 times, and then the stylized latent representation can be obtained
[0159] Step 22, for the result obtained in Step 21 Perform operations through Equation (23) to obtain the ink-wash style image I cs , as Figure 4 shown,
[0160]
[0161] where D() is the feature decoder;
[0162] Step 23, for I obtained in Step 22 cs Perform operations through Equation (24) to obtain the stylized image feature Y1,
[0163] Y1 = Downsample(Conv2d(I cs )) (24)
[0164] where Conv2D() is the convolution operation and Downsample() is the downsampling operation;
[0165] Step 24, iteratively calculate features Y2, Y3, and Y4 for Y1 obtained in Step 23 according to Equation (25),
[0166] Y n = Downsample(Conv2d(I n-1 ), 2 ≤ n ≤ 4 (25)
[0167] where Conv2D() is the convolution operation and Downsample() is the downsampling operation;
[0168] Step 25, perform operations on Y1, Y2, Y3, and Y4 obtained in Step 24 through Equation (26) respectively to obtain features Y1, Y2, Y3, and Y4 respectively,
[0169] Y i = Fusion(Y i ), i ∈ {1, 2, 3, 4} (26)
[0170] where Fusion() is the feature fusion operation;
[0171] Step 26, perform operations on features Y1, Y2, Y3, and Y4 obtained in Step 25 through Equation (27) respectively to obtain features Y1', Y2', Y3', and Y4' respectively,
[0172] Y i ' = Upsample(Y i ), i ∈ {1, 2, 3, 4} (27)
[0173] Among them, Upsample() is an upsampling operation;
[0174] Step 27: Perform an operation on the features Y1', Y2', Y3 ' and Y4' through Equation (28) to obtain a multi-scale fusion feature Y RefinedFeatures ,
[0175] Y RefinedFeatures = Aggre(Y1', Y2', Y3', Y4') (28)
[0176] Among them, Aggre() is a feature aggregation operation;
[0177] Step 28: Perform an operation on the aggregated feature Y RefinedFeatures obtained in Step 27 through Equation (29) to obtain a depth map D, as Figure 5 shown,
[0178] D = Softplus(Y RefinedFeatures ) (29)
[0179] Among them, Softplus() is a normalization operation.
Claims
1. A method for estimating the depth of Chinese ink paintings based on a diffusion model, characterized in that: It includes the following steps: Step 1, given the ink-wash painting image I c and the style image I s , respectively obtain the initial latent representation of the ink-wash painting image and the initial latent representation of the style image where F() is the feature encoder; Step 2, during the forward noise addition process, given the initial latent representation of the ink-wash painting image and the initial latent representation of the style image add Gaussian noise through Equation (2). where is a noisy version of, and β t represents a fixed sequence value used to control the noise intensity t represents the diffusion step number, where the value range is 1 ≤ t ≤ T, T is the maximum diffusion step number, and ε follows a Gaussian distribution with a mean of 0 and a variance of 1; Step 3, let α t = 1 - β t , The initial latent representation of the ink painting image Obtain the latent noise representation of the ink painting at time T through Equation (3) Style image Obtain the latent noise representation of the style at time T through Equation (3) Step 4, take the potential noise representation obtained in Step 2 and the potential noise representation to obtain a stylized initial potential noise representation through Equation (4) Among them represents the channel direction standard deviation of the style latent noise at time T, represents the channel direction mean of the style latent noise at time T, represents the channel direction standard deviation of the ink painting latent noise at time T, represents the channel direction mean of the ink painting latent noise at time T; Step 5, take the and obtained in Step 2 and the obtained in Step 4, and obtain the style feature content feature and the stylized feature where ResBlk() represents the residual operation; Step 6, the and obtained in Step 5 are respectively operated through Equation (6) to obtain the query of the ink painting features at time T and the query of the stylized features at time T Among them, rearrange() is the rearrangement operation, and linear Q () is the linear transformation operation; Step 7, the operation is performed according to Equation (7) to obtain the key of the style feature at time T and the value of the style feature at time T Among them, rearrange() is the rearrangement operation, and linear K,V () is the linear transformation operation; Step 8. Operate on the query obtained in Step 6 and the query according to Equation (8) to obtain the query of the stylized latent noise representation at time T after fusion where γ is the query save rate; Step 9, perform operations on the query obtained in Step 8 and the key obtained in Step 7 and the value according to Equation (9) to obtain the output feature of the attention mechanism at time T where τ is the temperature weight, d is the dimension of the feature vectors of the query and the key, and softmax() is the normalization operation; Step 10, the feature obtained in Step 9 is operated according to Equation (10) to obtain the stylized latent noise representation at time T after fusion Step 11, for the potential noise representation of the ink-wash painting obtained in Step 3 as the initial feature X 0 , perform sequential operations according to Equation (11), then add the time embedding vector and perform sequential operations again, and finally perform a residual connection operation with the input feature X 0 and repeat the above operations to obtain the feature X 1′ , X 1′ = Circulat(Seq(Seq(X 0 ) + Te(T)) + shortcut(X 0 )) (11) where Circulat() is the operation of looping twice, shortcut() is the operation of matching the number of channels, Te() is the time embedding operation, and Seq() is the sequence operation; Step 12, perform downsampling on the feature X obtained in Step 11 according to Equation (12) 1′ to obtain the feature X 1 , X 1 = Downsample(X 1′ ) (12) where Downsample() is the downsampling operation; Step 13, for the feature X obtained in Step 12 1 Recursively repeat the operations of Step 11 and Step 12 to obtain the feature X 2 , and then the feature X 2 Obtain the feature X through Equation (13) 3 , X 3 = Circulat(X k-1 ), 2 ≤ k ≤ 3 (13) where Circulat() is the operation of looping twice; Step 14, operate on the feature X obtained in Step 13 according to Equation (14) 3 to obtain the feature X 4 ; X 4 = Circulat(Seq(Seq(X 3 ) + Te(T)) + shortcut(X 3 )) (14) where Circulat() is the operation of looping twice, shortcut() is the operation of matching the number of channels, Te() is the time embedding operation, and Seq() is the sequence operation; Step 15, according to formula (15), perform operations on the feature X obtained in step 14 4 to obtain the feature X middle , X middle =Circulat(Seq(Seq(X 4 )+Te(T))+shortcut(X 4 )) (15) Then, according to Equation (16), for feature X middle and X 4 perform a splicing operation to obtain the initial feature of the decoder where Concat() is the concatenation operation; Step 16, for the initial encoder features obtained in Step 15 Perform operations according to Equation (17) to obtain features Step 17: Take the feature obtained in Step 16 and the feature X of the same size corresponding in the encoder path 3 concatenate them along the channel dimension, and repeat the operation of Equation (17) to obtain a feature Then, obtain a feature according to Equation (18) where Upsample() is the upsampling operation; Step 18, for the features obtained in Step 17 Recursively repeat the operations in Step 16 to Step 17 twice to obtain the final upsampled features Then, after performing group normalization and activation using the activation function according to Equation (19), the potential representation noise of the ink-wash painting at time T is obtained through a convolution operation where Conv2d() is the convolution operation, SiLU() is the activation operation, and GroupNorm() is the group normalization operation; Step 19, during the inverse diffusion process, the posterior diffusion conditional probability P(z t-1 |z t, z0) is approximately calculated by Equation (20), where ∝ represents proportional to, exp is the exponential function, α t , and are all known quantities in step 2, z t is the potential noise representation at time t, and then z t-1 is obtained according to Equation (21). The noise obtained through step 18 and the potential noise representation of the Chinese ink painting obtained in step 3 Obtain the potential noise representation of the Chinese ink painting at time T - 1 according to Equation (22) Step 20, for what is obtained in Step 3 Repeat the operations in Step 11 to Step 19, and then what can be obtained is For what is obtained in Step 9 Repeat the operations in Step 11 to Step 19, and then what can be obtained is Step 21, for what is obtained in Step 19 Repeat the operations of Step 5, Step 6, Step 8, Step 9, and Step 11 to Step 19 for T - 1 times, for what is obtained in Step 20 Repeat the operations of Step 5, Step 7, Step 9, and Step 11 to Step 19 for T - 1 times, for what is obtained in Step 20 Repeat the operations of Step 5, Step 6, Step 8, Step 9, Step 10, and Step 11 to Step 19 for T - 1 times, and the stylized latent representation can be obtained Step 22, for the result obtained in Step 21 perform operations through formula (23) to obtain the ink-wash style image I cs , where D() is the feature decoder; Step 23, perform operations on I obtained in Step 22 cs through Equation (24) to obtain the stylized image feature Y1 Y1 = Downsample(Conv2d(I cs )) (24) where Conv2D() is the convolution operation and Downsample() is the downsampling operation; Step 24, iteratively calculate the features Y2, Y3, and Y4 for Y1 obtained in Step 23 according to Equation (25), Y n = Downsample(Conv2d(I n-1 ), 2 ≤ n ≤ 4 (25) where Conv2D() is the convolution operation and Downsample() is the downsampling operation; Step 25, perform operations on Y1, Y2, Y3, and Y4 obtained in Step 24 respectively through Equation (26) to obtain the features Y1, Y2, Y3, and Y4 respectively, Y i = Fusion(Y i ), i ∈ {1, 2, 3, 4} (26) where Fusion() is the feature fusion operation; Step 26, perform operations on the features Y1, Y2, Y3, and Y4 obtained in Step 25 respectively through Equation (27) to obtain the features Y1', Y2', Y3', and Y4' respectively, Y i ' = Upsample(Y i ), i ∈ {1, 2, 3, 4} (27) where Upsample() is the upsampling operation; Step 27, perform operations on the feature Y1', Y2', Y3' and Y4' obtained in step 25 through formula (28) to obtain a multi-scale fusion feature Y RefinedFeatures , Y RefinedFeatures = Aggre(Y1', Y2', Y3', Y4') (28) where Aggre() is the feature aggregation operation; Step 28, perform an operation on the aggregated feature Y obtained in Step 27 RefinedFeatures to obtain a depth map D through Equation (29). D = Softplus(Y RefinedFeatures ) (29) where Softplus() is the normalization operation.