Slate picture real-time effect preview system
Through slate image acquisition, depth mapping, progressive style transfer and real-time rendering modules, combined with convolutional neural network and WebGL technology, the visual style difference problem of the real-time effect preview system of slate painting is solved, real-time preview and adaptive optimization of various artistic styles are realized, and creative efficiency and flexibility are improved.
Patent Information
- Application Number
- CN202510403183.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing real-time preview system for slate paintings cannot effectively reflect the visual style differences after etching, and cannot present the expected effects in multiple artistic styles in the same interface. It also has high hardware requirements or high learning costs, and lacks flexibility and instant feedback.
The slate graph acquisition unit, depth mapping module, progressive style transfer module and real-time rendering module are used, combined with convolutional neural network and WebGL technology to generate interactive three-dimensional model previews, support multiple styles and depth adjustments, and adaptive optimization is performed through feedback correction modules.
It realizes instant feedback on the browser side of the visual differences in different etching depths and styles, reduces trial and error costs, improves creative flexibility and interactive experience, and eliminates the high hardware requirements of three-dimensional modeling.
Smart Images

Figure CN120339375A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of real-time effect preview of stone slab paintings, and in particular to a real-time effect preview system for stone slab paintings. Background Art
[0002] As a traditional craft that combines painting art with stone carving, the creation process of stone slab paintings is essentially a visual transformation from a planar concept to a three-dimensional finished product.
[0003] Currently, most real-time effect preview systems still remain at the level of two-dimensional depth maps or texture maps: what designers see on the screen is only the distribution of black, white, and gray shades, or even just a static line drawing overlay. Such a preview cannot truly restore the texture changes on the surface of the stone slab after acid etching, is difficult to reflect the subtle concave and convex effects brought by light and shadow interaction, and cannot display the layering of the natural color and texture of the stone at different etching depths; traditional previews only provide a single visual style. When creators want to try different carving expressions such as raised reliefs, soft bas-reliefs, or minimalist line carvings, they need to repeatedly adjust parameters, export multiple texture maps, and recheck the results, which is a cumbersome and inflexible process;
[0004] Some solutions build three-dimensional models at the display end and render the stone surface, but this method has high hardware requirements and a large learning cost, which does not conform to the usage habits of most studios or craftsmen; while some relatively common planar texture maps, although simple, cannot capture the detailed differences in the natural texture and luster of the stone as the depth changes; therefore, an ideal real-time preview system not only needs to provide instant feedback on the etching depth, but also should be able to present the expected effects in multiple artistic styles in the same interface, allowing designers to intuitively feel the visual differences brought by different carving expressions at the initial stage, rather than finding that it does not match the expectation after the processing is completed. Summary of the Invention
[0005] In view of the above existing problems, the present invention is proposed.
[0006] The present invention provides a real-time effect preview system for stone slab paintings to solve the problem that the existing real-time effect preview solutions for stone slab paintings cannot effectively reflect the visual style differences after etching.
[0007] To solve the above technical problems, the present invention provides the following technical solutions:
[0008] An embodiment of the present invention provides a real-time effect preview system for stone slab paintings, which includes,
[0009] A stone slab image acquisition unit for acquiring a two-dimensional depth map and an original texture map of the stone slab to be processed;
[0010] A depth mapping module that fuses the acquired depth map and texture map to generate an intermediate three-dimensional model that can represent different acid etching depths;
[0011] The progressive style transfer module, based on a pre-trained style conversion network, applies multiple predefined slate painting style mappings to the intermediate 3D model and outputs a transfer result, including a preview of the real stone texture corresponding to the layer-by-layer etching depth;
[0012] The real-time rendering module, based on WebGL technology, converts the transfer result into an interactive 2D preview image, supports the user to adjust the etching depth and style parameters through a slider, and synchronously displays the style differences at different depth stages on the same interface; supports generating a rendering effect including ambient occlusion and specular reflection in real time on the browser side, and allows the user to freely switch the light source direction and intensity in the interface;
[0013] The feedback correction module records the depth and style preferences adjusted by the user, and dynamically updates the transfer network weights in combination with historical actual etching results, gradually realizing adaptive optimization.
[0014] As a preferred solution of the real-time effect preview system for slate paintings of the present invention, wherein: the depth mapping module adopts a U-Net structure based on a convolutional neural network, uses the input depth map and texture map as multi-channel features, and outputs a normal map and a roughness map.
[0015] As a preferred solution of the real-time effect preview system for slate paintings of the present invention, wherein: in the depth mapping module, the method for generating the intermediate 3D model is:
[0016] Concatenate the collected depth map and texture map in the channel dimension to form a multi-channel input X:
[0017] X = Concat(D, T) ∈ R H×W×4 ,
[0018] where D represents the collected 2D depth map, T represents the collected original texture map, H represents the height of the image, W represents the width of the image, and 4 represents the number of channels of the concatenated image, including 1 depth channel and 3 texture channels;
[0019] The network mapping function is defined as: (N, R) = f θ (X), where N represents the predicted normal map, R represents the predicted roughness map, and f θ represents a convolutional neural network mapping function based on the U-Net structure, θ represents the network weight parameter, and X is the multi-channel input;
[0020] Generate the normal map, and the formula is:
[0021]
[0022] where Ni,j Indicates the normal vector at position (i, j), Indicates the gradient of the depth map D in the horizontal direction at the pixel position (i, j), Indicates the gradient of the depth map D in the vertical direction at the pixel position (i, j), 1 is a constant term, and i and j respectively represent the row and column indices of the image;
[0023] The roughness map maps the depth sensitivity coefficient through the texture grayscale value, and the roughness map formula is:
[0024]
[0025] Among them, R i,j Indicates the roughness value at position (i, j), α represents the texture correction constant, Indicates the grayscale mean value of the texture map at position ((i, j), β represents the depth correction constant, D i,j Indicates the depth value at the pixel position (i, j), and the grayscale mean value of the texture is expressed as:
[0026] Among them, Indicates the pixel value of the red channel of the texture map at position (i, j), Indicates the pixel value of the green channel of the texture map at position (i, j), Indicates the pixel value of the blue channel of the texture map at position (i, j), 3 is the normalization factor, and the mean value of the three channels is calculated.
[0027] As a preferred solution of the real-time effect preview system for stone slab paintings described in the present invention, wherein: the training data annotation method of the intermediate 3D model includes:
[0028] In the training stage, a dedicated measurement device is used to obtain the stone slab depth map and surface texture data;
[0029] For the normal map, calculate the depth map gradient to obtain the true normal as the supervision label,
[0030] For the roughness map, according to the value range of the standardized surface roughness measurement value, calibrate different etching depths, and map shallow carving, medium carving, and deep carving to continuous roughness values;
[0031] The loss function is used to adjust the network parameters through backpropagation to make the predicted output closer to the true label. The loss function is a weighted combination and is expressed as:
[0032]
[0033] Among them, Indicates the total loss value, λ N Indicates the weight coefficient of the normal map loss part, Denote the normal map obtained by network prediction as N * Denote the ground-truth normal map obtained by supervision annotation as λ R Denote the weight coefficient of the roughness map loss part Denote the roughness map obtained by network prediction as R * Denote the ground-truth roughness map obtained by supervision annotation
[0034] As a preferred solution of the real-time effect preview system for stone slab paintings according to the present invention, wherein: the progressive style transfer module uses the CycleGAN architecture, pairs different etching depths with real stone samples through contrastive learning, performs progressive mapping from three styles of shallow carving, medium carving and deep carving, and uses the photo of the actual finished stone slab painting as the supervision target during training
[0035] As a preferred solution of the real-time effect preview system for stone slab paintings according to the present invention, wherein: in the progressive style transfer module, perform pre-training style conversion and CycleGAN paired training, and the steps include
[0036] Define the mapping relationship between the intermediate 3D model and the physical photo. Let Denote the intermediate 3D model, where M represents the model, H represents the image height, W represents the image width, and C M Denote the number of channels. Let Denote the physical photo of the real stone slab painting, where Y represents the photo and C Y Denote the number of photo channels
[0037] Introduce the generator G and the inverse generator F. The generator G maps the intermediate model to the style transfer result, which is expressed as Wherein Denote the generated style image, and G represents the style conversion network
[0038] The inverse generator F maps the physical photo back to the intermediate model domain, and the process is expressed as Wherein Denote the model obtained by inverse generation, and F represents the inverse style conversion network
[0039] As a preferred solution of the real-time effect preview system for stone slab paintings according to the present invention, wherein: in the progressive style transfer module, the steps of performing pre-training style conversion and CycleGAN paired training further include
[0040] Define the cycle consistency loss as to ensure the consistency of the content before and after conversion
[0041]
[0042] Wherein denotes the cyclic consistency loss, |·|1 denotes the L1 norm, which is used to measure the pixel-level difference between the generated image and the original image. F(G(M)) represents the result after converting the generated image back to the intermediate model domain in reverse, and G(F(Y)) represents the result of converting the reverse-generated result into a style image;
[0043] The generator is constrained by the adversarial loss, and the adversarial loss is for the generator G and the discriminator D Y It is defined as:
[0044]
[0045] where denotes the adversarial loss between the generator G and the discriminator D Y D Y denotes the discriminator used to distinguish between real photos and generated images, E denotes the expectation operation, log denotes the natural logarithm, D Y (Y) represents the probability that the discriminator determines a real photo, and D Y (G(M)) represents the probability that the discriminator determines a generated image; similarly, the inverse mapping introduces the discriminator D M and defines the adversarial loss
[0046] To enhance the feature correlation between different depth stages and real photos, a contrastive loss is introduced. Let the extracted feature be denoted as z, and the similarity is calculated using the cosine similarity sim(a, b) = a·b / |a||b|. Then the contrastive loss is defined as:
[0047]
[0048] where denotes the contrastive loss, z M denotes the intermediate model feature extracted by the generator G, z Y denotes the feature corresponding to the real photo, τ denotes the temperature parameter, sim(z M , z Y ) represents the cosine similarity between z M and z Y , and k represents other features in the negative sample set;
[0049] Furthermore, a style consistency loss is designed. The Gram matrix is used to capture the image style information. Let the feature map at the l-th layer be φ l (·), and its Gram matrix is defined as:
[0050] G l (·) = φ l (·) φ l (·) T ,
[0051] The style consistency loss is:
[0052]
[0053] where represents the style consistency loss, and G l (·) represents the Gram matrix of the l-th layer, and |·| F represents the Frobenius norm, which is used to measure the difference between matrices; the summation operation accumulates over the selected convolutional layers;
[0054] Combining all losses, the overall training objective is:
[0055]
[0056] where represents the total loss, λ adv represents the weight of the adversarial loss, λ cyc represents the weight of the cycle consistency loss, λ ctr represents the weight of the contrastive loss, λ style represents the weight of the style consistency loss.
[0057] As a preferred solution of the real-time effect preview system for stone slab paintings described in the present invention, wherein: the feedback correction module adopts a weight update mechanism based on incremental learning according to the preview depth and style selected by the user to update the parameters of the progressive style transfer network, so as to improve the consistency between subsequent previews and actual etching results.
[0058] As a preferred solution of the real-time effect preview system for stone slab paintings described in the present invention, wherein: in the feedback correction module, the triggering condition of the weight update mechanism is:
[0059] Let the feedback sample set be where i represents the sample index, and N f represents the cumulative number of feedback samples;
[0060] In the feedback samples, is the preference parameter selected by the user in the i-th sample, including the depth preference d u and the style preference s u , is the actually measured etching result. When the cumulative number of samples N f reaches the preset threshold T f or the mean prediction error in the samples exceeds the threshold ∈, the incremental learning update is triggered;
[0061] In the feedback sample set the user selection parameter for each sample is defined as represents the depth preference parameter selected by the user in the i-th feedback sample, reflecting the acid-etching depth expected by the user. This parameter is used to guide the depth adjustment of the intermediate 3D model in the depth mapping module. It is normalized to the interval [0, 1], where 0 represents the shallowest etching and 1 represents the deepest etching. The specific value is calibrated according to the actual etching effect and user habits. represents the style preference parameter selected by the user in the i-th feedback sample, reflecting the user's preference for the style details of the slate painting, such as texture fineness, color tone, etc. This is used to guide the style transfer module to achieve the progressive mapping of the predefined style.
[0062] If it adopts the form of an embedded vector, then it can be a vector with a fixed dimension, and its specific value is obtained by offline style feature learning and is kept consistent with the style feature dimension in the pre-trained style conversion network.
[0063] In the feedback correction module, to ensure the effectiveness of the feedback samples, each sample is screened, and the screening conditions are:
[0064]
[0065] Among them, is a function that maps the actual measurement result to the parameter space, ∈ s is the screening threshold, and the samples that meet this condition will be used for subsequent weight update.
[0066] The mapping function adopts the form of a linear model and is expressed as:
[0067]
[0068] Among them, represents the estimated value that maps the etching result R a to the user preference parameter space. W is the mapping weight matrix, b is the mapping bias vector, ∈ s is the sample screening threshold.
[0069] As a preferred solution of the real-time effect preview system for slate paintings described in the present invention, among them: in the weight update mechanism:
[0070] For the screened sample set, the number is denoted as N s and the incremental learning loss function is defined as:
[0071]
[0072] Among them, M i represents the intermediate 3D model corresponding to the i-th sample, G θis a pre-trained style transfer network, and θ is its weight. represents the incremental learning loss, γ represents the weight of the etching result fitting part, δ represents the weight of the user preference fitting part, and G θ (M i ) represents the style transfer output of the network for the intermediate model M i ;
[0073] To prevent drastic fluctuations caused by excessive update of network parameters, a regularization term
[0074]
[0075] is added. Among them, θ0 is the weight after the last update, and λ reg represents the regularization weight, and θ represents the current network weight;
[0076] The comprehensive incremental learning total loss is defined as:
[0077]
[0078] The weight update mechanism is updated using the gradient descent method, and the weights before and after the update are set as θ old and θ new :
[0079]
[0080] Among them, η represents the learning rate in incremental learning;
[0081] In the feedback correction module, the user's preference selection S u =(d u , s u ) establishes a mapping relationship with the physical measurement data R a . Through the mapping function , the actual etching result is converted into the estimated user preference parameters (d m , s m ), and then compared with S u to guide the update of the network weights. This mapping relationship is obtained through linear regression with multiple groups of samples in the offline stage.
[0082] Among them, W is the mapping weight matrix, and its dimension is set to [d s ×d r , where d r represents the dimension of the actual etching result measurement vector R a . For example, if the actual measurement result includes multiple indicators such as depth and roughness, then d r is the sum of the number of these indicators, and d sDenotes the dimension of the user preference parameter space. For example, if the user preference parameter only includes depth preference and style preference, then d s = 2. If other dimensions are included, then d s should be adjusted accordingly. Structurally, each row of W corresponds to a component in the user preference parameter. By weighted combination of R a 's various indicators, the corresponding predicted value is formed. b is the bias vector, whose dimension is [d s , which is consistent with the user preference parameter space. Its role is to provide translation compensation for the mapping function to ensure that the mapped result can more accurately match the user's actual choice.
[0083] The beneficial effects of the present invention are as follows: The present invention fuses a two-dimensional depth map and an original texture map into a three-dimensional representation that can reflect the real concave and convex structure, and combines a progressive style transfer module to output color previews of three predefined styles: shallow carving - medium carving - deep carving in real time; it is not necessary to repeatedly export multiple texture maps or rely on 3D modeling to quickly compare the visual differences of different etching depths and presentation methods, and it can effectively evaluate the effect of slate paintings in the initial stage of creation, reducing the trial-and-error cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0085] Figure 1 It is a schematic framework diagram of the real-time effect preview system of the slate painting of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0086] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific embodiments of the present invention with reference to the drawings in the specification.
[0087] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0088] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that mutually excludes other embodiments.
[0089] Example 1, refer to Figure 1 , this example provides a real-time effect preview system for slate paintings, including:
[0090] A slate image acquisition unit for acquiring the two-dimensional depth map and the original texture map of the slate to be processed;
[0091] A depth mapping module that fuses the acquired depth map and texture map to generate an intermediate three-dimensional model that can represent different etching depths;
[0092] The depth mapping module adopts a U-Net structure based on a convolutional neural network, uses the input depth map and texture map as multi-channel features, and outputs a normal map and a roughness map;
[0093] In the depth mapping module, the method for generating the intermediate three-dimensional model is:
[0094] Concatenate the acquired depth map and texture map in the channel dimension to form a multi-channel input X:
[0095] X = Concat(D, T) ∈ R H×W×4 ,
[0096] where D represents the acquired two-dimensional depth map, T represents the acquired original texture map, H represents the height of the image, W represents the width of the image, and 4 represents the number of channels of the concatenated image, including 1 depth channel and 3 texture channels;
[0097] The network mapping function is defined as: (N, R) = f θ (X), where N represents the predicted normal map, R represents the predicted roughness map, f θ represents the convolutional neural network mapping function based on the U-Net structure, θ represents the network weight parameter, and X is the multi-channel input;
[0098] To generate the normal map, the formula is:
[0099]
[0100] where N i,j represents the normal vector at position (i, j), represents the gradient of the depth map D in the horizontal direction at the pixel position (i, j), represents the gradient of the depth map D in the vertical direction at the pixel position (i, j), 1 is a constant term, and i and j respectively represent the indices of the row and column in the image;
[0101] The roughness mapping maps the depth sensitivity coefficient through the texture gray value, and the roughness mapping formula is:
[0102]
[0103] Among them, R i,j represents the roughness value at the position (i, j), α represents the texture correction constant, represents the grayscale mean of the texture map at the position (i, j), β represents the depth correction constant, D i,j represents the depth value at the pixel position (i, j), and the grayscale mean of the texture is expressed as:
[0104] Among them, represents the pixel value of the red channel of the texture map at the position (i, j), represents the pixel value of the green channel of the texture map at the position (i, j), represents the pixel value of the blue channel of the texture map at the position (i, j), and 3 is the normalization factor to calculate the mean value of the three channels;
[0105] The annotation method of the training data of the middle three-dimensional model includes:
[0106] In the training stage, a dedicated measuring device is used to obtain the depth map and surface texture data of the slate;
[0107] For the normal map, the depth map gradient is calculated to obtain the true normal as the supervision label,
[0108] For the roughness map, according to the value range of the standardized surface roughness measurement value, different etching depths are calibrated, and shallow carving, medium carving, and deep carving are mapped to continuous roughness values;
[0109] The loss function is used to adjust the network parameters through backpropagation to make the predicted output closer to the true label. The loss function is a weighted combination and is expressed as:
[0110]
[0111] Among them, represents the total loss value, λ N represents the weight coefficient of the normal map loss part, represents the normal map obtained by network prediction, N * represents the true normal map of the supervision annotation, λ R represents the weight coefficient of the roughness map loss part, represents the roughness map obtained by network prediction, R * represents the true roughness map of the supervision annotation;
[0112] Specifically, here the depth map and texture map are concatenated to form a multi-channel input, and the U-Net network is used to achieve multi-task prediction; the gradient normalization technology is adopted to reflect the fine changes on the slate surface, and the roughness mapping combines the texture grayscale and depth values and performs continuous mapping by adjusting constants;
[0113] In this application, the shallow engraving style refers to the effect formed when the acid etching depth is relatively shallow. Its characteristics are clear lines and relatively subtle concave and convex changes, mainly showing a delicate engraving feeling and the smooth texture of traditional stone carvings. The etching thickness of the latent engraving is generally 0.2 - 0.5 mm;
[0114] The medium engraving style refers to the effect formed when the acid etching depth is moderate. It can not only retain the original texture of the stone material but also form a gradual three-dimensional sense at the etched edge, showing the soft shadows and specular reflection effects under the interaction of light and shadow; the etching thickness of the medium engraving is generally 0.5 - 1.5 mm;
[0115] The deep engraving style refers to the effect when the acid etching depth is relatively deep. Its surface shows obvious concave and convex structures and a rough texture, emphasizing the contrast between heavy shadows and distinct textures, thus producing a strong three-dimensional visual impact. The etching thickness of the deep engraving is generally above 1.5 mm.
[0116] The progressive style transfer module, based on a pre-trained style conversion network, applies multiple predefined stone painting style mappings to the intermediate 3D model and outputs the transfer result, including a preview of the real stone material texture corresponding to the layer-by-layer etching depth;
[0117] The progressive style transfer module uses the CycleGAN architecture. By contrastive learning, it pairs different etching depths with real stone material samples, performs progressive mapping from three styles: shallow engraving, medium engraving, and deep engraving, and uses the photos of the actual completed stone paintings as the supervision target during training;
[0118] In the progressive style transfer module, pre-trained style conversion and CycleGAN paired training are carried out. The steps include:
[0119] Define the mapping relationship between the intermediate 3D model and the physical photo. Let represent the intermediate 3D model, where M represents the model, H represents the image height, W represents the image width, and C M represents the number of channels. Let represent the real physical photo of the stone painting, where Y represents the photo and C Y represents the number of photo channels;
[0120] Introduce the generator G and the inverse generator F. The generator G maps the intermediate model to the style transfer result, which is expressed as: where represents the generated style image and G represents the style conversion network;
[0121] The inverse generator F maps the physical photo back to the intermediate model domain, and the process is expressed as: where represents the model obtained by inverse generation and F represents the inverse style conversion network;
[0122] In the progressive style transfer module, the steps of pre-training style conversion and paired training with CycleGAN also include:
[0123] Define the cycle consistency loss to ensure the consistency of the content before and after conversion as
[0124]
[0125] Among them, represents the cycle consistency loss, |·|1 represents the L1 norm, which is used to measure the difference in pixels between the generated image and the original image. F(G(M)) represents the result after converting the generated image back to the intermediate model domain reversely, and G(F(Y)) represents the result of converting the reverse generation result into a style image;
[0126] Use adversarial loss to constrain the generator. The adversarial loss is for the generator G and the discriminator D Y Define as:
[0127]
[0128] Among them, represents the adversarial loss between the generator G and the discriminator D Y D Y represents the discriminator used to distinguish real photos from generated images, E represents the expectation operation, log represents the natural logarithm, D Y (Y) represents the judgment probability of the discriminator for real photos, and D Y (G(M)) represents the judgment probability of the discriminator for the generated image; similarly, the inverse mapping introduces the discriminator D M And define the adversarial loss
[0129] To enhance the feature correlation between different depth stages and real photos, introduce the contrast loss. Let the extracted feature be represented as z, and calculate the similarity using the cosine similarity sim(a,b) = a·b / |a||b|. Then the contrast loss is defined as:
[0130]
[0131] Among them, represents the contrast loss, z M represents the intermediate model features extracted by the generator G, and z Y represents the features corresponding to real photos, τ represents the temperature parameter, and sim(z M ,z Y ) represents the cosine similarity between z M and z Y , and k represents other features in the negative sample set;
[0132] Further design the style consistency loss, capture the image style information using the Gram matrix, and let the feature map of the l-th layer be φ l (·), and its Gram matrix is defined as:
[0133] G l (·) = φ l (·)φ l (·) T ,
[0134] The style consistency loss is:
[0135]
[0136] Among them, represents the style consistency loss, G l (·) represents the Gram matrix of the l-th layer, |·| F represents the Frobenius norm, which is used to measure the difference between matrices; the summation operation accumulates over the selected convolutional layers;
[0137] Combining all losses, the overall training objective is:
[0138]
[0139] Among them, represents the total loss, λ adv represents the weight of the adversarial loss, λ cyc represents the weight of the cycle consistency loss, λ ctr represents the weight of the contrastive loss, λ style represents the weight of the style consistency loss;
[0140] Specifically, here the pre-trained style transfer network is applied to the intermediate 3D model, and the CycleGAN architecture is used to achieve the inter-domain mapping;
[0141] In the paired training process, construct the generator and the inverse generator to achieve the inverse mapping, use the cycle consistency to ensure the content retention, the adversarial loss guides the generator to generate realistic images, the contrastive learning strengthens the feature alignment between different depth stages and real photos, and the style consistency loss constrains the style features through the Gram matrix. The overall loss function is a weighted combination of multiple losses, and the realism and detail consistency of the style transfer results are improved through various constraints;
[0142] The real-time rendering module, based on WebGL technology, converts the migration result into an interactive 2D preview image, supports users to adjust the etching depth and style parameters through sliders, and synchronously displays the style differences at different depth stages on the same interface; supports generating real-time rendering effects including ambient occlusion and specular reflection on the browser side, and allows users to freely switch the light source direction and intensity in the interface;
[0143] The feedback correction module records the depth and style preferences adjusted by the user, and dynamically updates the migration network weights in combination with historical actual etching results to gradually achieve adaptive optimization;
[0144] According to the preview depth and style selected by the user, the feedback correction module adopts a weight update mechanism based on incremental learning to update the parameters of the progressive style transfer network, so as to improve the consistency between subsequent previews and actual etching results;
[0145] In the feedback correction module, the trigger condition of the weight update mechanism is:
[0146] Let the feedback sample set be where i represents the sample index, and N f represents the cumulative number of feedback samples;
[0147] In the feedback sample, is the preference parameter selected by the user in the i-th sample, including the depth preference d u and the style preference s u , is the actually measured etching result. When the cumulative number of samples N f reaches the preset threshold T f or the mean prediction error in the samples exceeds the threshold ∈, trigger the incremental learning update;
[0148] In the feedback sample set the user-selected parameters for each sample are defined as represents the depth preference parameter selected by the user in the i-th feedback sample, reflecting the etching depth expected by the user. This parameter is used to guide the depth adjustment of the intermediate 3D model in the depth mapping module, normalized to the interval [0,1], where 0 represents the shallowest etching and 1 represents the deepest etching. The specific value is calibrated according to the actual etching effect and user habits, represents the style preference parameter selected by the user in the i-th feedback sample, reflecting the user's preference for the style details of the stone slab painting, such as texture fineness, color tone, etc. This is used to guide the style transfer module to achieve progressive mapping of the predefined style;
[0149] If in the form of an embedding vector, then It can be a vector of fixed dimension, whose specific values are obtained through offline style feature learning and are kept consistent with the style feature dimension in the pre-trained style transfer network;
[0150] In the feedback correction module, to ensure the effectiveness of feedback samples, each sample is screened, and the screening conditions are:
[0151]
[0152] Among them, is a function that maps the actual measurement result to the parameter space, ∈ s is the screening threshold, and the samples that meet this condition will be used for subsequent weight update;
[0153] The mapping function adopts the form of a linear model and is expressed as:
[0154]
[0155] Among them, represents the estimated value that maps the etching result R a to the user preference parameter space, W is the mapping weight matrix, b is the mapping bias vector, ∈ s is the sample screening threshold;
[0156] In the weight update mechanism:
[0157] For the screened sample set, the number is denoted as N s , and the incremental learning loss function is defined as:
[0158]
[0159] Among them, M i represents the intermediate three-dimensional model corresponding to the i-th sample, G θ is the pre-trained style transfer network, θ is its weight, represents the incremental learning loss, γ represents the weight of the etching result fitting part, δ represents the weight of the user preference fitting part, G θ (M i ) represents the style transfer output of the network for the intermediate model M i ;
[0160] To prevent excessive update of network parameters from causing violent fluctuations, a regularization term
[0161]
[0162] Among them, θ0 is the weight after the last update, λ reg represents the regularization weight, and θ represents the current network weight;
[0163] The total loss of comprehensive incremental learning is defined as:
[0164]
[0165] The weight update mechanism is updated using the gradient descent method, and the weights before and after the update are set to θ old and θ new :
[0166]
[0167] where η represents the learning rate in incremental learning;
[0168] In the feedback correction module, the user's preference selection S u =(d u , s u ) establishes a mapping relationship with the physical measurement data R a . Through the mapping function , the actual etching result is converted into the estimated user preference parameter (d m , s m ), and then compared with S u to guide the network weight update. This mapping relationship is obtained through linear regression with multiple groups of samples in the offline stage.
[0169] where W is the mapping weight matrix, and its dimension is set to [d s ×d r . Among them, d r represents the dimension of the actual etching result measurement vector R a . For example, if the actual measurement result includes multiple indicators such as depth and roughness, then d r is the sum of the number of these indicators. d s represents the dimension of the user preference parameter space. For example, if the user preference parameter only includes depth preference and style preference, then d s = 2. If there are other dimensions, then d s should be adjusted accordingly. Structurally, each row of W corresponds to a component in the user preference parameter. After weighting and combining the indicators of R a , the corresponding predicted value is formed. b is the bias vector, and its dimension is [d s . It is consistent with the user preference parameter space, and its role is to provide translation compensation for the mapping function to ensure that the mapped result can more accurately match the user's actual selection;
[0170] Specifically, the feedback correction module records the user's adjustment preferences for depth and style and the actual etching measurement results, and constructs an incremental learning framework;
[0171] Update is initiated when the cumulative number of samples reaches a certain threshold or the prediction error is too large. Invalid data is filtered out through a sample screening mechanism, and the physical etching results are converted into expected parameters using a linear mapping function. The incremental learning loss function is divided into two parts: one part measures the difference between the network output and the actual results, and the other part reflects the inconsistency between the user preference and the mapped value. The regularization term effectively controls the amplitude of parameter update to prevent model oscillation;
[0172] By recording the user's adjustment preferences for depth and style and combining with historical actual etching results, an incremental learning strategy is adopted to dynamically update the weights of the style transfer network. After recording the feedback samples, valid samples are screened according to the triggering conditions, and then the designed loss function is used to update the network parameters, so as to realize the mapping correction between the preview effect and the actual etching results;
[0173] The present invention fundamentally changes the creative process: reduces the number of trial and errors, speeds up the decision-making speed, and anticipates in advance the true texture of the finished product on the screen, thus truly realizing the seamless connection between design and the finished product.
[0174] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A real-time effect preview system for slate paintings, characterized in that: Comprising, A slate image acquisition unit for acquiring a two-dimensional depth map and an original texture map of a slate to be processed; A depth mapping module that fuses the acquired depth map and texture map to generate an intermediate three-dimensional model that can represent different etching depths; A progressive style transfer module that, based on a pre-trained style conversion network, applies a variety of predefined slate painting style mappings to the intermediate three-dimensional model and outputs a transfer result, including a preview of the true stone texture corresponding to the layer-by-layer etching depth; A real-time rendering module that, based on WebGL technology, converts the transfer result into an interactive two-dimensional preview image; A feedback correction module that records the user's adjusted depth and style preferences and dynamically updates the weights of the transfer network in combination with historical actual etching results.
2. The real-time effect preview system for slate paintings according to claim 1, wherein: The depth mapping module adopts a U-Net structure based on a convolutional neural network, using the input depth map and texture map as multi-channel features and outputting a normal map and a roughness map.
3. The real-time effect preview system for slate paintings according to claim 1, characterized in that: In the depth mapping module, the method for generating the intermediate three-dimensional model is as follows: The acquired depth map and texture map are concatenated in the channel dimension to form a multi-channel input X: X = Concat(D, T) ∈ R H×W×4 , Where D represents the acquired two-dimensional depth map, T represents the acquired original texture map, H represents the height of the image, W represents the width of the image, and 4 represents the number of channels of the concatenated image, including 1 depth channel and 3 texture channels; The network mapping function is defined as: (N, R) = f θ (X), where N represents the predicted normal map, R represents the predicted roughness map, and f θ represents the convolutional neural network mapping function based on the U-Net structure, θ represents the network weight parameter, and X is the multi-channel input; Generate a normal map, and the formula is: Among them, N i,j represents the normal vector at the position (i, j), represents the gradient of the depth map D in the horizontal direction at the pixel position (i, j), represents the gradient of the depth map D in the vertical direction at the pixel position (i, j), 1 is a constant term, and i and j respectively represent the indices of the rows and columns in the image; The roughness mapping maps the depth sensitivity coefficient through the texture gray value, and the roughness mapping formula is: Among them, R i,j represents the roughness value at the position (i, j), α represents the texture correction constant, represents the gray-scale mean value of the texture map at the position (i, j), β represents the depth correction constant, D i,j represents the depth value at the pixel position (i, j), and the gray-scale mean value of the texture is expressed as: Among them, represents the pixel value of the red channel of the texture map at position (i, j), represents the pixel value of the green channel of the texture map at position (i, j), represents the pixel value of the blue channel of the texture map at position (i, j), and 3 is the normalization factor for calculating the mean value of the three channels.
4. The real-time effect preview system of a stone slab painting according to claim 3, characterized in that: The training data annotation method of the intermediate three-dimensional model includes: In the training stage, a dedicated measuring device is used to obtain the slate depth map and surface texture data; For the normal map, calculate the depth map gradient to obtain the true normal as the supervision label; For the roughness map, calibrate different etching depths according to the value range of the standardized surface roughness measurement value, and map shallow carving, medium carving, and deep carving to continuous roughness values; Use a loss function to adjust the network parameters through backpropagation. The loss function is a weighted combination and is expressed as: Among them, represents the total loss value, and λ N represents the weight coefficient of the normal map loss part, represents the normal map obtained by network prediction, N * represents the true normal map of the supervised annotation, λ R represents the weight coefficient of the roughness map loss part, represents the roughness map obtained by network prediction, R * represents the true roughness map of the supervised annotation.
5. The real-time effect preview system of a slate painting as described in claim 1, characterized in that: The progressive style transfer module uses the CycleGAN architecture, pairs different etching depths with real stone samples through contrastive learning, performs progressive mapping from three styles of shallow carving, medium carving, and deep carving, and uses the actual completed slate painting photo as the supervision target during training.
6. The real-time effect preview system of a stone slab painting according to claim 5, wherein: In the progressive style transfer module, the steps for pre-training style conversion and CycleGAN paired training include: Define the mapping relationship between the intermediate 3D model and the physical photo. Let represent the intermediate 3D model, where M represents the model, H represents the image height, W represents the image width, and C M represents the number of channels. Let represent the physical photo of the real stone slab painting, where Y represents the photo and G Y represents the number of photo channels; The generator G and the inverse generator F are introduced. The generator G maps the intermediate model to the style transfer result, which is expressed as: where represents the generated style image, and G represents the style conversion network; The inverse generator F maps the physical object photo back to the intermediate model domain, and the process is expressed as: Among them, represents the model obtained by inverse generation, and F represents the inverse style conversion network.
7. The real-time effect preview system of a slate painting according to claim 6, wherein: In the progressive style transfer module, the steps for pre-training style conversion and CycleGAN paired training further include: Define the cyclic consistency loss as Among them, represents the cycle consistency loss, |·|1 represents the L1 norm, which is used to measure the difference in pixels between the generated image and the original image. F(G(M)) represents the result after converting the generated image back to the intermediate model domain in reverse, and G(F(Y)) represents the result of converting the reverse generation result into a style image. The generator is constrained by the adversarial loss, and the adversarial loss is for the generator G and the discriminator D Y It is defined as: Among them, represents the adversarial loss between the generator G and the discriminator D Y , D Y represents the discriminator used to distinguish real photos from generated images, E represents the expectation operation, log represents the natural logarithm, D Y (Y) represents the determination probability of the discriminator for real photos, D Y (G(M)) represents the determination probability of the discriminator for the generated image; the inverse mapping introduces the discriminator D M and defines the adversarial loss Introduce a contrastive loss. Let the extracted feature representation be z, and calculate the similarity using the cosine similarity sim(a,b)=a·b / ||a||b|. Then the contrastive loss is defined as: Among them, represents the contrastive loss, z M represents the intermediate model features extracted by the generator G, z Y represents the features corresponding to the real photos, τ represents the temperature parameter, sim(z M , z Y ) represents z M and z Y cosine similarity, k represents other features in the negative sample set; Capture the image style information using the Gram matrix. Let the feature map at the l-th layer be φ l (·), and its Gram matrix is defined as: G l (·) = φ l (·)φ l (·) T , The style consistency loss is: Among them, represents the style consistency loss, G l (·) represents the Gram matrix of the l-th layer, |·| F represents the Frobenius norm, which is used to measure the difference between matrices; the summation operation accumulates over the selected convolutional layers; Combining various losses, the overall training objective is: Among them, represents the total loss, λ adv represents the weight of the adversarial loss, λ cyc represents the weight of the cycle consistency loss, λ ctr represents the weight of the contrastive loss, λ style represents the weight of the style consistency loss.
8. The real-time effect preview system for slate paintings according to claim 1, wherein: The feedback correction module updates the parameters of the progressive style transfer network according to the preview depth and style selected by the user, using a weight update mechanism based on incremental learning.
9. The real-time effect preview system for slate paintings according to claim 8, wherein: In the feedback correction module, the triggering condition of the weight update mechanism is: Let the feedback sample set be where i represents the sample index and N f represents the cumulative number of feedback samples; In the feedback sample, is the preference parameter selected by the user in the i-th sample, including the depth preference d u and the style preference s u , is the actually measured etching result. When the cumulative sample number N f reaches the preset threshold T f or the mean prediction error in the sample exceeds the threshold ∈, incremental learning update is triggered; In the feedback correction module, each sample is screened, and the screening condition is: Among them, is a function that maps the actual measurement results to the parameter space, ∈ s is the screening threshold, and the samples that meet this condition will be used for subsequent weight updates; The mapping function adopts a linear model form and is expressed as: Among them, represents the estimated value of mapping the etching result R a to the user preference parameter space, W is the mapping weight matrix, b is the mapping bias vector, ∈ s is the sample screening threshold.
10. A real-time effect preview system for slate paintings according to claim 9, characterized in that: In the weight update mechanism: For the filtered sample set, the quantity is denoted as N s , the incremental learning loss function is defined as: Among them, M i represents the intermediate 3D model corresponding to the i-th sample, G θ is the pre-trained style transfer network, and θ is its weight, represents the incremental learning loss, γ represents the weight of the etching result fitting part, δ represents the weight of the user preference fitting part, and G θ (M i ) represents the style transfer output of the network for the intermediate model M i ; Add a regularization term Among them, θ0 is the weight after the last update, and λ reg represents the regularization weight, and θ represents the current network weight; The comprehensive incremental learning total loss is defined as: The weight update mechanism is updated using the gradient descent method, and the weights before and after the update are set to θ old and θ new : where η represents the learning rate in incremental learning; In the feedback correction module, the user's preference selection S u =(d u , s u ) and the physical measurement data R a establish a mapping relationship, and through the mapping function convert the actual etching result into an estimated user preference parameter (d m , s m ), and then compare it with S u to guide the update of the network weights. This mapping relationship is obtained by linear regression with multiple groups of samples in the offline stage. where, W is the mapping weight matrix, and its dimension is set to [d s × d r , where d r represents the dimension of the actual etching result measurement vector R a , and d s represents the dimension of the user preference parameter space. Structurally, each row of W corresponds to a component in the user preference parameters. After weighted combination of the various indicators of R a , the corresponding predicted value is formed, and b is the bias vector.