Single-image semitransparent material editing method based on generative adversarial network

The GAN-based method addresses the challenge of editing transparency in single images by using an encoder, semantic encoder, and parameter editor to adjust transparency parameters, achieving precise control and semantic editing of material properties.

CN120318403APending Publication Date: 2025-07-15BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510091966.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively edit the translucency of an object in a single picture, especially in the absence of multi-angle images and lighting information, and the optical properties of the translucent material cannot be accurately restored.

Method used

Generative adversarial network is adopted to extract and adjust the transparency parameters in the image to achieve editing of semi-transparent material properties through image encoder, appearance semantic encoder, model parameter prediction subnet and model parameter editing subnet.

Benefits of technology

The translucency of objects in the image is accurately controlled under unknown lighting conditions, so as to realize continuous and semantic editing of the transparency of objects in a single image, and enhance the visual performance of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318403A_ABST
    Figure CN120318403A_ABST
Patent Text Reader

Abstract

The invention provides a single-image semitransparent material editing method based on a generative adversarial network, and belongs to the technical field of computer vision, and the method comprises the steps: inputting a target image into an image encoder, and obtaining a plurality of first space codes outputted by the image encoder; inputting the plurality of first spatial codes into an appearance semantic encoder to obtain a plurality of second spatial codes output by the appearance semantic encoder in an encoding space; inputting the plurality of second space codes into a model parameter prediction sub-network to obtain a transparency parameter; after receiving the updated transparency parameter, inputting the updated transparency parameter and the plurality of second space codes into a model parameter editing sub-network to obtain a plurality of code vectors; returning the plurality of coding vectors to the appearance semantic encoder to obtain a plurality of third space codes output by the appearance semantic encoder; and inputting the plurality of third spatial codes into an image generator to obtain an image result. According to the invention, continuous and semantic editing of the transparency of the object in a single image can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method for editing the semi-transparent material of a single image based on a generative adversarial network. Background Art

[0002] Translucency is a fundamental visual property of an object's appearance and is often used to evaluate material characteristics, such as the turbidity of a fluid or the clarity of a gemstone. Editing the translucency of an object in a single image is of great significance, especially in scenarios where it is necessary to display and adjust the texture of the material. For example, when photographing some objects (such as jade, glass products, or semi-transparent plastics), the scattering and transmission of light will significantly affect the visual effect of the object. However, the captured image is often limited by the shooting conditions (such as the lighting direction, intensity, and background environment) and cannot fully display the true texture of the object itself. By editing the translucency of the image, the scattering and absorption behavior of light inside the object can be simulated and adjusted, thereby enhancing or changing the visual performance of the material. Such editing can not only highlight the high-quality sense of the product but also adjust its appearance according to needs to adapt to different uses such as advertising and design. For example, enhancing the transparency of jade or highlighting the luster of glass products makes the final presented image more attractive and persuasive.

[0003] The optical properties of semi-transparent materials are divided into two parts: surface scattering characterized by the bidirectional reflectance distribution function (BRDF) and subsurface scattering characterized by the bidirectional subsurface scattering reflectance distribution function (BSSRDF). Synthesizing images of semi-transparent material objects is already a solved problem. However, in image editing applications, directly adjusting the translucency of an object in an image obtained by shooting or rendering is a common operation in many design-related fields. Editing the translucency of the material based on only a single image is very difficult because the volume scattering of semi-transparent materials may scatter the incident light in unexpected directions, display a blurred background through transmission, and affect the surrounding environment through caustics and blurred shadows.

[0004] For the method of editing semi-transparent materials based on images, currently, it is usually to estimate scene data such as geometric models, lighting conditions, imaging processes, and object materials from the input image, and then adjust the translucency through an image rendering method. This method usually requires collecting multiple images from multiple perspectives or lighting conditions. If only a single image is used and there is a lack of illumination, observation points, and object geometry information, restoring the BRDF and BSSRDF parameters will become an ill-posed problem. Moreover, this method only estimates the lighting conditions and cannot well edit the object material (including translucency).

[0005] Therefore, how to edit the translucency of an object in a single image has become an urgent technical problem to be solved. Summary of the Invention

[0006] The present invention provides a method for editing the semi - transparent material of a single image based on a generative adversarial network, so as to solve the defect in the prior art that the transparency of an object in a single picture cannot be edited.

[0007] The present invention provides a method for editing the semi - transparent material of a single image based on a generative adversarial network, including the following steps: Input the target image into an image encoder to obtain a plurality of first spatial encodings output by the image encoder; Input the plurality of first spatial encodings into an appearance semantic encoder to obtain a plurality of second spatial encodings in the encoding space output by the appearance semantic encoder, where the appearance semantic encoder is used to extract semantic information related to transparency; Input the plurality of second spatial encodings into a model parameter prediction sub - network to obtain transparency parameters output by the model parameter prediction sub - network, where the model parameter prediction sub - network is used to predict the transparency parameters of the object in the target image; After receiving the updated transparency parameters, input the updated transparency parameters and the plurality of second spatial encodings into a model parameter editing sub - network to obtain a plurality of encoded vectors output by the model parameter editing sub - network, where the model parameter editing sub - network is used to map the updated transparency parameters back to the encoding space; Return the plurality of encoded vectors to the appearance semantic encoder to obtain a plurality of third spatial encodings output by the appearance semantic encoder; Input the plurality of third spatial encodings into an image generator to obtain an image result output by the image generator.

[0008] According to the method for editing the semi - transparent material of a single image based on a generative adversarial network provided by the present invention, the appearance semantic encoder includes a plurality of parallel auto - encoders and a plurality of parallel decoders; Each of the auto - encoders is used to connect two first spatial encodings and reduce the dimension of the connected encoded vector to obtain the corresponding second spatial encoding; Each of the decoders is used to decode based on the second spatial encoding to obtain two corresponding third spatial encodings.

[0009] According to the method for editing the semi - transparent material of a single image based on a generative adversarial network provided by the present invention, the auto - encoder includes an input layer and a hidden layer. The input layer is used to reduce the dimension of the connected encoded vector to obtain a first intermediate vector, and the hidden layer of the auto - encoder is used to reduce the dimension of the first intermediate vector to obtain the second spatial encoding; The decoder includes a hidden layer and an output layer. The hidden layer of the decoder is used to increase the dimension based on the second spatial encoding to obtain a second intermediate vector, and the output layer is used to increase the dimension of the second intermediate vector to obtain two third spatial encodings.

[0010] According to a single-image semi-transparent material editing method based on a generative adversarial network provided by the present invention, the transparency parameters include surface roughness, attenuation coefficient of the medium, and single-scattering albedo; the inputting the plurality of second spatial encodings into the model parameter prediction sub-network to obtain the transparency parameters output by the model parameter prediction sub-network includes: For each second spatial encoding in the plurality of second spatial encodings, after splicing the second spatial encoding into a vector of a preset dimension, input it into the input layer of the model parameter prediction sub-network to obtain a third intermediate vector output by the input layer; Input the third intermediate vector into the hidden layer of the model parameter prediction sub-network to obtain a fourth intermediate vector output by the hidden layer; Input the fourth intermediate vector into the output layer of the model parameter prediction sub-network to obtain the transparency parameters output by the output layer.

[0011] According to a single-image semi-transparent material editing method based on a generative adversarial network provided by the present invention, the single-image semi-transparent material editing method based on a generative adversarial network further includes: Using a semantic factorization algorithm to determine the relationship between each dimension in the encoding space and the semi-transparency; Wherein, the relationship between each dimension and the semi-transparency includes a plurality of feature direction vectors and the feature change intensity of each feature direction vector.

[0012] According to a single-image semi-transparent material editing method based on a generative adversarial network provided by the present invention, after receiving the updated transparency parameters, inputting the updated transparency parameters and the plurality of second spatial encodings into the model parameter editing sub-network to obtain a plurality of encoded vectors output by the model parameter editing sub-network includes: For each second spatial encoding in the plurality of second spatial encodings, after receiving the updated transparency parameters, input the transparency parameters and the second spatial encoding into the model parameter editing sub-network to obtain the encoded vector corresponding to the second spatial encoding output by the model parameter editing sub-network, and the encoded vector represents the feature direction vector for adjusting the transparency parameters.

[0013] The present invention also provides a single-image semi-transparent material editing device based on a generative adversarial network, including the following modules: An image encoding module, configured to: input a target image into an image encoder to obtain a plurality of first spatial encodings output by the image encoder; A semantic encoding module, configured to: input the plurality of first spatial encodings into an appearance semantic encoder to obtain a plurality of second spatial encodings in a coding space output by the appearance semantic encoder, where the appearance semantic encoder is used to extract semantic information related to transparency; A parameter prediction module, configured to: input the plurality of second spatial encodings into a model parameter prediction sub-network to obtain transparency parameters output by the model parameter prediction sub-network, where the model parameter prediction sub-network is used to predict the transparency parameters of an object in the target image; A parameter editing module, configured to: after receiving the updated transparency parameters, input the updated transparency parameters and the plurality of second spatial encodings into a model parameter editing sub-network to obtain a plurality of coding vectors output by the model parameter editing sub-network, where the model parameter editing sub-network is used to map the updated transparency parameters back to the coding space; A semantic decoding module, configured to: return the plurality of coding vectors to the appearance semantic encoder to obtain a plurality of third spatial encodings output by the appearance semantic encoder; An image generation module, configured to: input the plurality of third spatial encodings into an image generator to obtain an image result output by the image generator.

[0014] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, it implements the method for editing a single-image semi-transparent material based on a generative adversarial network as described in any one of the above.

[0015] The present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method for editing a single-image semi-transparent material based on a generative adversarial network as described in any one of the above.

[0016] The present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method for editing a single-image semi-transparent material based on a generative adversarial network as described in any one of the above.

[0017] The single - image semi - transparent material editing method based on a generative adversarial network provided by the present invention inputs a target image into an image encoder to obtain a plurality of first spatial encodings output by the image encoder; inputs the plurality of first spatial encodings into an appearance semantic encoder to obtain a plurality of second spatial encodings in the encoding space output by the appearance semantic encoder, where the appearance semantic encoder is used to extract semantic information related to transparency; inputs the plurality of second spatial encodings into a model parameter prediction sub - network to obtain transparency parameters output by the model parameter prediction sub - network, where the model parameter prediction sub - network is used to predict the transparency parameters of the object in the target image; after receiving the updated transparency parameters, inputs the updated transparency parameters and the plurality of second spatial encodings into a model parameter editing sub - network to obtain a plurality of coding vectors output by the model parameter editing sub - network, where the model parameter editing sub - network is used to map the updated transparency parameters back to the encoding space; returns the plurality of coding vectors to the appearance semantic encoder to obtain a plurality of third spatial encodings output by the appearance semantic encoder; and inputs the plurality of third spatial encodings into an image generator to obtain an image result output by the image generator. The present invention extracts semantic information related to semi - transparency from a single input image, and by adjusting the semi - transparency parameters, realizes the editing of the properties of semi - transparent materials, solves the problem of precisely controlling the semi - transparency of an object in an image under unknown lighting conditions and geometric information, and realizes continuous and semantic editing of the transparency of an object in a single image. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 is a schematic flowchart of the single - image semi - transparent material editing method based on a generative adversarial network provided by the present invention; Figure 2 is a schematic structural diagram of the overall network provided by the present invention; Figure 3 is a schematic structural diagram of the appearance semantic encoder provided by the present invention; Figure 4 is a schematic structural diagram of the model parameter prediction sub - network provided by the present invention; Figure 5 is a schematic structural diagram of the model parameter editing sub - network provided by the present invention; Figure 6It is a schematic structural diagram of a single-image semi-transparent material editing device based on a generative adversarial network provided by the present invention; Figure 7 It is a schematic structural diagram of an electronic device provided by the present invention; Figure 8 It is a schematic diagram of the result of single-image semi-transparent material editing based on a generative adversarial network provided by the present invention. Detailed implementation manners

[0020] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] It should be noted that in the description of the embodiments of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitations, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element. The orientation or positional relationship indicated by terms such as "upper", "lower", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be construed as a limitation of the present invention. Unless otherwise clearly defined and limited, the terms "mounted", "connected" and "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0022] The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are usually of the same type, and do not limit the number of objects. For example, the first object can be one or multiple. In addition, "and / or" means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.

[0023] Figure 1 is a schematic flow chart of a single-image semi-transparent material editing method based on a generative adversarial network provided by the present invention. As Figure 1 shown, the method includes the following: S110, input the target image into an image encoder to obtain a plurality of first spatial encodings output by the image encoder; S120, input the plurality of first spatial encodings into an appearance semantic encoder to obtain a plurality of second spatial encodings in the encoding space output by the appearance semantic encoder, and the appearance semantic encoder is used to extract semantic information related to transparency; S130, input the plurality of second spatial encodings into a model parameter prediction sub-network to obtain transparency parameters output by the model parameter prediction sub-network, and the model parameter prediction sub-network is used to predict the transparency parameters of the objects in the target image; S140, after receiving the updated transparency parameters, input the updated transparency parameters and the plurality of second spatial encodings into a model parameter editing sub-network to obtain a plurality of coding vectors output by the model parameter editing sub-network, and the model parameter editing sub-network is used to map the updated transparency parameters back to the encoding space; S150, return the plurality of coding vectors to the appearance semantic encoder to obtain a plurality of third spatial encodings output by the appearance semantic encoder; S160, input the plurality of third spatial encodings into an image generator to obtain an image result output by the image generator.

[0024] It should be noted that the execution subject of the single-image semi-transparent material editing method based on a generative adversarial network provided by the embodiments of this application can be a server, a computer device, such as a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc.

[0025] Figure 2 is a schematic diagram of the overall network provided by the present invention. As Figure 2 shown, the network structure is based on a combination of an encoder and a generator, and a semi-transparency editing sub-network obtained through training is embedded between these two sub-networks for semi-transparency editing. The semi-transparency editing sub-network includes three parts: an appearance semantic encoder, a model parameter prediction sub-network, and a model parameter editing sub-network.

[0026] As Figure 2 shown, the target image (input image) is input into the image encoder for encoding to obtain 2k first spatial encodings W1, W2... W in the W+ space output by the image encoder 2k , and these W+ space encoding vectors capture the high-dimensional feature information of the target image, including features such as material, texture, and transparency; Then, these 2k first spatial encodings are input into the appearance semantic encoder to further reduce the dimension of the high-dimensional features of the image, extract the appearance and material semantic information in the image, and provide a basis for subsequent transparency editing, obtaining k second spatial encodings in the encoding space (T space) output by the appearance semantic encoder; the appearance semantic encoder will convert the W+ space encoding vector of the image into a new semantic vector, which contains the visual features of the image, such as transparency, glossiness, reflectivity, etc.; After that, the k second spatial encodings are input into the model parameter prediction sub-network to obtain the transparency parameters corresponding to the output of the model parameter prediction sub-network, and the user can adjust the transparency parameters according to needs, thereby changing the transparency of the object in the target image; The edited transparency parameters by the user and the second spatial encodings in the T space are input into the model parameter editing sub-network, and the model parameter editing sub-network generates k new encoding vectors corresponding to the second spatial encodings; The k new encoding vectors are returned to the appearance semantic encoder, and through decoding, the corresponding 2k third spatial encodings in the W+' space are obtained; The 2k third spatial encodings are input into the image generator, and the image generator generates a modified image according to the new encoding vectors while maintaining the integrity of other details of the image.

[0027] Optionally, the Pixel2Style2Pixel (pSp) network is used as the image encoder, and the StyleGAN2 is used as the image generator. The encoder can also use other encoders and generators based on the generative adversarial network (GAN) type.

[0028] The single-image semi-transparent material editing method based on a generative adversarial network provided by the present invention inputs a target image into an image encoder to obtain a plurality of first spatial encodings output by the image encoder; inputs the plurality of first spatial encodings into an appearance semantic encoder to obtain a plurality of second spatial encodings in the encoding space output by the appearance semantic encoder, where the appearance semantic encoder is used to extract semantic information related to transparency; inputs the plurality of second spatial encodings into a model parameter prediction sub-network to obtain transparency parameters output by the model parameter prediction sub-network, where the model parameter prediction sub-network is used to predict the transparency parameters of an object in the target image; after receiving the updated transparency parameters, inputs the updated transparency parameters and the plurality of second spatial encodings into a model parameter editing sub-network to obtain a plurality of encoding vectors output by the model parameter editing sub-network, where the model parameter editing sub-network is used to map the updated transparency parameters back to the encoding space; returns the plurality of encoding vectors to the appearance semantic encoder to obtain a plurality of third spatial encodings output by the appearance semantic encoder; inputs the plurality of third spatial encodings into an image generator to obtain an image result output by the image generator. The present invention extracts semantic information related to semi-transparency from a single input image, and through the adjustment of semi-transparency parameters, realizes the editing of the properties of semi-transparent materials, solves the problem of precisely controlling the semi-transparency of an object in an image under unknown lighting conditions and geometric information, and realizes continuous and semantic editing of the transparency of an object in a single image.

[0029] In an alternative embodiment, the appearance semantic encoder includes a plurality of parallel autoencoders and a plurality of parallel decoders; Each of the autoencoders is used to connect two first spatial encodings and reduce the dimension of the connected encoding vector to obtain a corresponding second spatial encoding; Each of the decoders is used to decode based on the second spatial encoding to obtain two corresponding third spatial encodings.

[0030] In an embodiment of the present invention, the appearance semantic encoder is used to extract semantic information related to transparency, and is a multi-layer autoencoder architecture that further maps the first spatial encoding in the W+ space to a new low-dimensional latent space T. The appearance semantic encoder consists of a plurality of parallel autoencoders, and each two latent encodings in the W+ space correspond to an autoencoder. It connects the 512-dimensional W+ space encodings of adjacent layers together and uses a matching autoencoder to reduce their dimension to a 40-dimensional vector in the T space.

[0031] In an alternative embodiment, the autoencoder includes an input layer and a hidden layer. The input layer is used to reduce the dimension of the concatenated encoded vector to obtain a first intermediate vector, and the hidden layer of the autoencoder is used to reduce the dimension of the first intermediate vector to obtain a second spatial encoding; The decoder includes a hidden layer and an output layer. The hidden layer of the decoder is used to increase the dimension based on the second spatial encoding to obtain a second intermediate vector, and the output layer is used to increase the dimension of the second intermediate vector to obtain two third spatial encodings.

[0032] Figure 3 is a schematic structural diagram of the appearance semantic encoder provided by the present invention. As Figure 3 shown, the encoder of each autoencoder is respectively composed of an input layer and a hidden layer, and the decoder is respectively composed of a hidden layer and an output layer. A 1024-dimensional vector is input, reduced to 128 dimensions through the input layer, and then reduced to 40 dimensions through the hidden layer to obtain the T spatial encoding; the T spatial encoding is input into the hidden layer to be increased to 128 dimensions, and finally increased to 1024 dimensions through the output layer. Finally, this 1024-dimensional vector is split into two 512-dimensional W+' spatial encodings, which are respectively connected to the AdaIN (Adaptive Instance Normalization) style modules of the corresponding layers to achieve the conversion of the image style by adjusting the mean and variance of the feature map.

[0033] In an alternative embodiment, the single-image semi-transparent material editing method based on a generative adversarial network further includes: using a semantic factorization algorithm to determine the relationship between each dimension in the encoding space and the semi-transparency; wherein, the relationship between each dimension and the semi-transparency includes a plurality of feature direction vectors and the feature change intensity of each of the feature direction vectors.

[0034] In the embodiment of the present invention, the T encoding space extracts features related to semi-transparency through self-supervised training, and analyzes the association between each dimension in the T space and the semi-transparency through a semantic factorization (SeFa) algorithm.

[0035] Here, SeFa (Semantic Factorization) is a method for analyzing and controlling the semantic attributes of the latent space in Generative Adversarial Networks (GANs). The core idea of SeFa is to find the directions in the latent space of the GAN that represent specific semantic attributes. These directions can be obtained by decomposing the weight matrix of the intermediate layer of the GAN model. Specifically, it is usually the weight of the fully connected layer or the convolutional layer. Through Singular Value Decomposition (SVD) or other matrix decomposition methods, the dominant components can be extracted from these weight matrices, and each component may correspond to a specific semantic attribute. Using SeFa, a discovered direction can be manipulated individually to achieve fine control over the semantic attribute represented by that direction.

[0036] In the embodiments of the present invention, for each layer, 8 feature direction vectors are obtained. These vectors define the directions of movement in the latent space, and each direction theoretically represents a specific semantic attribute or variation pattern; for each feature direction vector, a feature change intensity is obtained, which corresponds to the degree of influence on the image when moving along each feature direction vector.

[0037] The single-image semi-transparent material editing method based on a generative adversarial network provided by the embodiments of the present invention uses the SeFa algorithm to generate multiple feature direction vectors and corresponding feature change intensities, so as to achieve semantic-level editing of an image by individually manipulating one direction.

[0038] In an optional embodiment, the transparency parameters include surface roughness, attenuation coefficient of the medium, and single-scattering albedo; the step of inputting the multiple second spatial encodings into the model parameter prediction sub-network to obtain the transparency parameters output by the model parameter prediction sub-network includes: For each of the multiple second spatial encodings, after splicing the second spatial encoding into a vector of a preset dimension, it is input into the input layer of the model parameter prediction sub-network to obtain a third intermediate vector output by the input layer; Inputting the third intermediate vector into the hidden layer of the model parameter prediction sub-network to obtain a fourth intermediate vector output by the hidden layer; Inputting the fourth intermediate vector into the output layer of the model parameter prediction sub-network to obtain the transparency parameters output by the output layer.

[0039] Figure 4 is a schematic structural diagram of the model parameter prediction sub-network provided by the present invention, as Figure 4As shown in the figure, the model parameter prediction sub-network is used to predict the opacity parameters of the objects in the input image, including surface roughness (α), attenuation coefficient of the medium (σt), and single-scattering albedo (ρa). The model parameter prediction sub-network is a multi-layer perceptron (MLP) network for opacity prediction. For each T-space vector, it is concatenated into a 280-dimensional vector as the input of the MLP. Finally, the MLP outputs a three-dimensional vector, and the three dimensions respectively represent α, σt, and ρa.

[0040] In an alternative embodiment, after receiving the updated transparency parameter, the updated transparency parameter and the plurality of second spatial encodings are input into the model parameter editing sub-network, and a plurality of encoded vectors output by the model parameter editing sub-network are obtained, including: For each of the plurality of second spatial encodings, after receiving the updated transparency parameter, the transparency parameter and the second spatial encoding are input into the model parameter editing sub-network, and the encoded vector corresponding to the second spatial encoding output by the model parameter editing sub-network is obtained. The encoded vector represents the eigen-direction vector for adjusting the transparency parameter.

[0041] In the embodiment of the present invention, the model parameter editing sub-network is used to achieve dynamic adjustment of opacity by predicting the transparency change direction vector of the input image and performing a directional operation in the T space. This operation allows for adjusting the object transparency by adjusting specific dimensions in the latent space while keeping other image attributes unchanged.

[0042] Figure 5 is a schematic structural diagram of the model parameter editing sub-network provided by the present invention. As Figure 5 shown, the model parameter editing sub-network is an MLP network for opacity editing. For each T-space vector, it is concatenated into a 280-dimensional vector as the input. Finally, the MLP outputs a 40-dimensional normalized vector, representing the direction vector for adjusting specific model parameters in the predicted T space.

[0043] The following gives a specific implementation of the single-image semi-transparent material editing method based on a generative adversarial network provided by the present invention.

[0044] In the specific implementation process, the solution includes a training dataset construction stage, an image encoder and image generator training stage, an appearance semantic encoder training stage, and a model parameter prediction sub-network and model parameter editing sub-network training stage.

[0045] Training dataset construction phase: Ambient light can greatly affect the adjustment of transparency and the generation results. Here, 9 different environment maps were selected to simulate natural light sources in the virtual scene, and spherical water droplets (blob) were used as the geometric model. The spherical water droplet model was obtained by modulating a sphere with a two-dimensional Gaussian noise map, where the value of amplitude A ∈ {0.5, 1}; the value of frequency f ∈ {0.0, 0.2, 0.4, 0.7, 0.9, 1.2, 1.5}. A total of 14 different geometric models were modulated. For each rendering, one of the 9 different environment maps was randomly selected as the ambient light; one of the 14 spherical water droplet models was selected as the geometric model, and a total of 18,196 256×256 images were rendered.

[0046] Image encoder and image generator training phase: pixel2Style2pixel (pSp) was used as the image encoder, and the dataset images were divided into a training set and a test set in a ratio of 4:1. For the training set, 14,556 images were randomly selected, and for the test set, the remaining 3,640 images were used. Three losses in pSp, namely lpips loss, l2 loss, and w_norm loss, were used for training. It was set to output the evaluation loss every 1,000 training steps and the weights every 10,000 training steps, and a total of 20,000 steps were trained. For the image generator training, the StyleGAN2 architecture was used for training. Since the network does not require the test set, the entire dataset was used for training. The loss function of StyleGAN mainly consists of two parts: the loss function of GAN generator G and the loss function of GAN discriminator D. In this experiment, StyleGAN combined all loss functions into one loss function, called fid50k_full. It was set to train 10,000 images per round, and a total of 1,000 rounds were trained (that is, a total of 10 million images were repeatedly trained); the evaluation loss was output every 100 rounds.

[0047] Appearance semantic encoder training stage: The 18,196 rendered images are fed into the trained image encoder for dimensionality reduction, and finally 18,196 corresponding W+ space encoding vectors are obtained. The size of each W+ space encoding vector is 14×512. For the i-th autoencoder, the sub-encodings of the (2i - 1)-th and 2i-th layers of each W+ space encoding vector are concatenated as its dataset (so for each autoencoder, there are 18,196 1024-dimensional vectors as the dataset). 16,000 vectors are selected as the training set, and the remaining 2,196 vectors are used as the test set. At the same time, the SeFa algorithm is used to analyze the contribution degree of each dimension in the W+ space to the image features. For each layer, 8 eigen-direction vectors are obtained; for each eigen-direction vector, a feature change intensity is obtained. The k-th dimension of the j-th eigen-direction vector of the i-th layer is denoted as ; the feature intensity of the j-th eigen-direction vector of the i-th layer is denoted as ; the feature contribution of the k-th dimension of the Code of the i-th layer in the W+ space is denoted as ; the j-th dimension of the Code of the i-th layer in the W+ space is denoted as , then the calculation formula of the final loss function is as follows: ; ; Thus, the loss function of each autoencoder is obtained and trained according to this. The training was carried out for a total of 200 steps, and the loss evaluation was output once every 100 vectors were trained in each step.

[0048] Training stage of the model parameter prediction sub-network and the model parameter editing sub-network: Divide 18,196 images of 256×256 into two parts, the training set and the test set, with a ratio of 4:1. For the training set, randomly select 14,556 images, and for the test set, use the remaining 3,640 images. Push the above images into the trained appearance semantic encoder architecture respectively to obtain the corresponding 280-dimensional vectors as the input of the model parameter prediction sub-network; at the same time, connect the rendering parameters α, σt, and ρa of each image into a three-dimensional vector as the label for supervised learning. In addition, use L1 Loss as the loss function for MLP training. A total of 20 steps are trained, 1,000 vectors are trained in each step, and the loss evaluation is output once. For the training of the model parameter editing sub-network, obtain 18,196 images and the corresponding three parameters α, σt, and ρa. For each image I1, appropriately increase the values of the above three parameters and re-render to obtain a new image I2, and push I1 and I2 into the appearance semantic encoder to obtain the corresponding T-space encodings T1 and T2. For the parameters α and σt, obtain the 5th layer codes T1’ and T2’ of T1 and T2 respectively, and take the difference and normalize T1’ and T2’ to finally obtain the training label T; for the parameter ρa, take the difference and normalize the 6th layer codes T1’ and T2’ of T1 and T2 respectively to finally obtain the training label T. Take T1 as the input of the MLP and T as the training label, and train for the three parameter change situations of α, σt, and ρa respectively. Use L1 Loss as the loss function for MLP training, and a total of 30 rounds are trained, 1,000 samples are trained in each round, and the loss evaluation is output once.

[0049] According to the above steps, a deep learning network for transparency editing in images based on a generative adversarial network (GAN) can be obtained, including modules such as an image encoder, an image generator, an appearance semantic encoder, a model parameter prediction sub-network, and a model parameter editing sub-network.

[0050] According to the above steps, a deep learning network for transparency editing in images based on a generative adversarial network (GAN) can be obtained, including modules such as an image encoder, an image generator, an appearance semantic encoder, a model parameter prediction sub-network, and a model parameter editing sub-network.

[0051] Furthermore, the encoded vectors in the image will be passed to the appearance semantic encoder module. The role of this module is to further reduce the dimension of the high-dimensional features of the image, extract the appearance and material semantic information in the image, and provide a basis for subsequent transparency editing. The appearance semantic encoder will convert the W+ space encoding vector of each image into a new semantic vector, which contains the visual features of the image, such as transparency, glossiness, reflectivity, etc.

[0052] After being processed by the appearance semantic encoder module, the obtained semantic vectors will be fed as inputs into the model parameter prediction sub-network. This sub-network will predict the transparency-related parameters of the objects in the image based on the input features, such as the rendering parameters for translucency (α, σt, and ρa). These parameters control the transparency effect, scattering, and absorption characteristics of the objects in the image.

[0053] Users can adjust these parameters according to their needs. For example, the transparency of an object can be changed by increasing or decreasing the value of α, or the light scattering and absorption characteristics of the object can be controlled by adjusting parameters such as σt and ρa. This process achieves precise control of the translucent effect.

[0054] After the transparency parameters are adjusted, the new transparency parameters will be passed to the model parameter editing sub-network. This network will map the adjusted parameters back to the encoding space and generate new encoding vectors. Subsequently, these new encoding vectors will be input into the image generator, and through adversarial generation network structures such as StyleGAN2, they will be transformed into new images. The image generator will generate a modified image based on the new encoding vectors while maintaining the integrity of other details of the image.

[0055] Figure 8 is a schematic diagram of the result of single-image translucent material editing based on the generative adversarial network provided by the present invention. As Figure 8 shown, to verify the quality of the material editing, several new images were re-rendered as input images using the same rendering settings as those used to construct the rendering dataset in the training dataset, and their translucency was linearly changed at equal intervals along the direction vectors predicted by the MLP, thereby generating a set of new images I with a fixed translucent difference. In addition, the material appearance model parameters of the image set I were predicted, and a second set of images I' was drawn based on these predicted parameters. The generated image set I and the newly rendered image set I' are compared in Figure 8 Figure a.

[0056] As Figure 8 shown, all three MLPs are capable of adjusting the translucency of the objects in the input image, but in different forms. This phenomenon is consistent with the different effects of the three model parameters on the material appearance. For example, the RMS slope α tends to change the blurring degree of the object through the background, and the extinction coefficient σt tends to change the brightness of the object through the background, while the single-scattering albedo ρa of the medium a tends to change the turbidity degree of the object.

[0057] Furthermore, a test for adjusting the translucency of real photos was conducted, and the editing results are as Figure 8As shown in Figure b. The results show that the trained neural network can edit the translucency of real objects captured under real lighting conditions. In addition, reasonable effects of caustics and soft shadows can also be observed.

[0058] In summary, the single-image translucent material editing method based on generative adversarial network provided by the present invention uses the latent space of the generative adversarial network (GAN) to decouple material parameters, and solves the problem of precisely controlling the translucency of objects in an image under unknown lighting conditions and geometric information. This method extracts semantic information related to translucency from a single input image, and realizes the editing of the properties of translucent materials by adjusting the parameters directionally in a specific latent space, specifically including the control of the scattering, light transmittance and blur degree of the object.

[0059] Next, the single-image translucent material editing device based on generative adversarial network provided by the embodiments of the present application will be described. The single-image translucent material editing device based on generative adversarial network described below can be correspondingly referred to the single-image translucent material editing method based on generative adversarial network described above.

[0060] Figure 6 is a schematic structural diagram of the single-image translucent material editing device based on generative adversarial network provided by the present invention. As Figure 6 shown, the single-image translucent material editing device based on generative adversarial network may include, but is not limited to; An image encoding module 610, configured to: input a target image into an image encoder to obtain a plurality of first spatial encodings output by the image encoder; A semantic encoding module 620, configured to: input the plurality of first spatial encodings into an appearance semantic encoder to obtain a plurality of second spatial encodings in the encoding space output by the appearance semantic encoder, where the appearance semantic encoder is used to extract semantic information related to transparency; A parameter prediction module 630, configured to: input the plurality of second spatial encodings into a model parameter prediction sub-network to obtain transparency parameters output by the model parameter prediction sub-network, where the model parameter prediction sub-network is used to predict the transparency parameters of the object in the target image; A parameter editing module 640, configured to: after receiving the updated transparency parameters, input the updated transparency parameters and the plurality of second spatial encodings into a model parameter editing sub-network to obtain a plurality of encoding vectors output by the model parameter editing sub-network, where the model parameter editing sub-network is used to map the updated transparency parameters back to the encoding space; A semantic decoding module 650, configured to: return the plurality of encoding vectors to the appearance semantic encoder to obtain a plurality of third spatial encodings output by the appearance semantic encoder; An image generation module 660 is configured to: input the multiple third spatial encodings into an image generator to obtain an image result output by the image generator.

[0061] It should be noted that the single-image semi-transparent material editing device based on a generative adversarial network provided in the embodiments of the present invention can, when specifically running, execute the single-image semi-transparent material editing method based on a generative adversarial network described in any of the above embodiments, and details thereof are not elaborated in this embodiment.

[0062] Figure 7 An example of a schematic physical structure diagram of an electronic device is shown as Figure 7 shown. The electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communication interface 720, and the memory 730 complete mutual communication through the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute the single-image semi-transparent material editing method based on a generative adversarial network. The method includes: inputting a target image into an image encoder to obtain multiple first spatial encodings output by the image encoder; inputting the multiple first spatial encodings into an appearance semantic encoder to obtain multiple second spatial encodings in the encoding space output by the appearance semantic encoder, where the appearance semantic encoder is used to extract semantic information related to transparency; inputting the multiple second spatial encodings into a model parameter prediction sub-network to obtain transparency parameters output by the model parameter prediction sub-network, where the model parameter prediction sub-network is used to predict the transparency parameters of an object in the target image; after receiving the updated transparency parameters, inputting the updated transparency parameters and the multiple second spatial encodings into a model parameter editing sub-network to obtain multiple coding vectors output by the model parameter editing sub-network, where the model parameter editing sub-network is used to map the updated transparency parameters back to the encoding space; returning the multiple coding vectors to the appearance semantic encoder to obtain multiple third spatial encodings output by the appearance semantic encoder; inputting the multiple third spatial encodings into an image generator to obtain an image result output by the image generator.

[0063] In addition, when the logical instructions in the above-mentioned memory 730 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0064] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the single-image semi-transparent material editing method based on a generative adversarial network provided by the above-mentioned various methods. The method includes: inputting a target image into an image encoder to obtain a plurality of first spatial encodings output by the image encoder; Inputting the plurality of first spatial encodings into an appearance semantic encoder to obtain a plurality of second spatial encodings in the encoding space output by the appearance semantic encoder. The appearance semantic encoder is used to extract semantic information related to transparency; Inputting the plurality of second spatial encodings into a model parameter prediction sub-network to obtain transparency parameters output by the model parameter prediction sub-network. The model parameter prediction sub-network is used to predict the transparency parameters of the object in the target image; After receiving the updated transparency parameters, inputting the updated transparency parameters and the plurality of second spatial encodings into a model parameter editing sub-network to obtain a plurality of coding vectors output by the model parameter editing sub-network. The model parameter editing sub-network is used to map the updated transparency parameters back to the encoding space; Returning the plurality of coding vectors to the appearance semantic encoder to obtain a plurality of third spatial encodings output by the appearance semantic encoder; Inputting the plurality of third spatial encodings into an image generator to obtain an image result output by the image generator.

[0065] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for editing the semi-transparent material of a single image based on a generative adversarial network provided by the above various methods. The method includes: inputting a target image into an image encoder to obtain a plurality of first spatial encodings output by the image encoder; Inputting the plurality of first spatial encodings into an appearance semantic encoder to obtain a plurality of second spatial encodings in the encoding space output by the appearance semantic encoder, where the appearance semantic encoder is used to extract semantic information related to transparency; Inputting the plurality of second spatial encodings into a model parameter prediction sub-network to obtain transparency parameters output by the model parameter prediction sub-network, where the model parameter prediction sub-network is used to predict the transparency parameters of an object in the target image; After receiving the updated transparency parameters, inputting the updated transparency parameters and the plurality of second spatial encodings into a model parameter editing sub-network to obtain a plurality of coding vectors output by the model parameter editing sub-network, where the model parameter editing sub-network is used to map the updated transparency parameters back to the encoding space; Returning the plurality of coding vectors to the appearance semantic encoder to obtain a plurality of third spatial encodings output by the appearance semantic encoder; Inputting the plurality of third spatial encodings into an image generator to obtain an image result output by the image generator.

[0066] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0067] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for editing the translucent material of a single image based on a generative adversarial network, characterized in that, Including: Input the target image into an image encoder to obtain a plurality of first spatial encodings output by the image encoder; Input the plurality of first spatial encodings into an appearance semantic encoder to obtain a plurality of second spatial encodings in the encoding space output by the appearance semantic encoder, where the appearance semantic encoder is used to extract semantic information related to transparency; Input the plurality of second spatial encodings into a model parameter prediction sub-network to obtain transparency parameters output by the model parameter prediction sub-network, where the model parameter prediction sub-network is used to predict the transparency parameters of the object in the target image; After receiving the updated transparency parameters, input the updated transparency parameters and the plurality of second spatial encodings into a model parameter editing sub-network to obtain a plurality of encoded vectors output by the model parameter editing sub-network, where the model parameter editing sub-network is used to map the updated transparency parameters back to the encoding space; Return the plurality of encoded vectors to the appearance semantic encoder to obtain a plurality of third spatial encodings output by the appearance semantic encoder; Input the plurality of third spatial encodings into an image generator to obtain an image result output by the image generator.

2. The method for editing a semi-transparent material of a single image based on a generative adversarial network according to claim 1, wherein The appearance semantic encoder includes a plurality of parallel auto-encoders and a plurality of parallel decoders; Each of the auto-encoders is used to connect two first spatial encodings and reduce the dimension of the connected encoded vector to obtain the corresponding second spatial encoding; Each of the decoders is used to decode based on the second spatial encoding to obtain two corresponding third spatial encodings.

3. The method for editing the translucent material of a single image based on a generative adversarial network according to claim 2, wherein The auto-encoder includes an input layer and a hidden layer. The input layer is used to reduce the dimension of the connected encoded vector to obtain a first intermediate vector, and the hidden layer of the auto-encoder is used to reduce the dimension of the first intermediate vector to obtain the second spatial encoding; The decoder includes a hidden layer and an output layer. The hidden layer of the decoder is used to increase the dimension based on the second spatial encoding to obtain a second intermediate vector, and the output layer is used to increase the dimension of the second intermediate vector to obtain two third spatial encodings.

4. The method for editing a semi-transparent material of a single image based on a generative adversarial network according to claim 1, wherein The transparency parameters include surface roughness, attenuation coefficient of the medium, and single-scattering albedo. The inputting the plurality of second spatial encodings into the model parameter prediction sub-network to obtain the transparency parameters output by the model parameter prediction sub-network includes: For each second spatial encoding in the plurality of second spatial encodings, splice the second spatial encoding into a vector of a preset dimension and input it into the input layer of the model parameter prediction sub-network to obtain a third intermediate vector output by the input layer; Input the third intermediate vector into the hidden layer of the model parameter prediction sub-network to obtain a fourth intermediate vector output by the hidden layer; Input the fourth intermediate vector into the output layer of the model parameter prediction sub-network to obtain the transparency parameters output by the output layer.

5. The method for editing the semi-transparent material of a single image based on a generative adversarial network according to claim 1, wherein The single-image semi-transparent material editing method based on a generative adversarial network further includes: Using a semantic factorization algorithm to determine the relationship between each dimension in the encoding space and semi-transparency; Among them, the relationship between each dimension and the transparency includes a plurality of feature direction vectors and the feature change intensity of each of the feature direction vectors.

6. The method for editing a single - image semi - transparent material based on a generative adversarial network according to claim 5, wherein, After receiving the updated transparency parameter, inputting the updated transparency parameter and the plurality of second spatial encodings into the model parameter editing sub-network, and obtaining a plurality of encoded vectors output by the model parameter editing sub-network, including: For each of the plurality of second spatial encodings, after receiving the updated transparency parameter, inputting the transparency parameter and the second spatial encoding into the model parameter editing sub-network, and obtaining the encoded vector corresponding to the second spatial encoding output by the model parameter editing sub-network, where the encoded vector represents the feature direction vector for adjusting the transparency parameter.

7. A single-image semi-transparent material editing device based on a generative adversarial network, characterized in that, Including: An image encoding module, configured to: input a target image into an image encoder, and obtain a plurality of first spatial encodings output by the image encoder; A semantic encoding module, configured to: input the plurality of first spatial encodings into an appearance semantic encoder, and obtain a plurality of second spatial encodings in the encoding space output by the appearance semantic encoder, where the appearance semantic encoder is used to extract semantic information related to transparency; A parameter prediction module, configured to: input the plurality of second spatial encodings into a model parameter prediction sub-network, and obtain the transparency parameter output by the model parameter prediction sub-network, where the model parameter prediction sub-network is used to predict the transparency parameter of an object in the target image; A parameter editing module, configured to: after receiving the updated transparency parameter, input the updated transparency parameter and the plurality of second spatial encodings into a model parameter editing sub-network, and obtain a plurality of encoded vectors output by the model parameter editing sub-network, where the model parameter editing sub-network is used to map the updated transparency parameter back to the encoding space; A semantic decoding module, configured to: return the plurality of encoded vectors to the appearance semantic encoder, and obtain a plurality of third spatial encodings output by the appearance semantic encoder; An image generation module, configured to: input the plurality of third spatial encodings into an image generator, and obtain an image result output by the image generator.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the single-image semi-transparent material editing method based on a generative adversarial network according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the single-image semi-transparent material editing method based on a generative adversarial network according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the single-image semi-transparent material editing method based on a generative adversarial network according to any one of claims 1 to 6.