VAE training method and device, storage medium and electronic equipment
By obtaining sample images in the VAE model, determining the latent variable distribution and generating reconstructed images, calculating the perceptual loss and combining multiple losses for optimization, the difficulty of training high-precision VAE is solved and high-quality image generation is achieved.
Patent Information
- Application Number
- CN202510886855.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-17
AI Technical Summary
How to train a high-precision variational autoencoder (VAE) model to improve image generation quality.
By obtaining sample images and inputting them into the VAE model to be trained, the distribution of latent variables is determined, reconstructed images are generated, image features are extracted, perceptual loss is calculated, and the model is trained based on the loss. Comprehensive optimization is performed by combining KL divergence loss, reconstruction loss, and distillation loss.
Accurately train high-precision VAE models to improve the quality and efficiency of image generation.
Smart Images

Figure CN120808106A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present specification relates to the technical field of computer technology, and particularly relates to a VAE training method and device, a storage medium and an electronic device. BACKGROUND
[0002] At present, image generation models have been widely applied in various fields, and the variational autoencoder (VAE) is one of the important components in various image generation models, such as the Stable Diffusion model and the Flux model.
[0003] The VAE is used for encoding and decoding in the image generation process, the encoding is to convert the original image into a representation in the latent space, and the decoding is to convert the vector in the latent space back to the image. It can be seen that the accuracy of the VAE determines the quality of the images generated by the image generation model to a great extent.
[0004] Therefore, how to train a VAE with high accuracy is a problem to be solved. SUMMARY
[0005] The embodiments of the present specification provide a VAE training method, device, storage medium and electronic device to partially solve the problems existing in the prior art.
[0006] The embodiments of the present specification adopt the following technical solutions:
[0007] The VAE training method provided by the present specification comprises:
[0008] Obtaining a VAE model to be trained and a sample image;
[0009] Inputting the sample image into the VAE model to be trained, determining the latent variable distribution corresponding to the sample image through the encoding module of the VAE model to be trained;
[0010] Inputting the latent variable distribution into the decoding module of the VAE to be trained, and generating a reconstructed image of the sample image according to the latent variable distribution through the decoding module;
[0011] Extracting first image features of the reconstructed image and second image features of the sample image;
[0012] Determining a perceptual loss according to the difference between the first image features and the second image features;
[0013] Training the VAE model to be trained according to at least the perceptual loss.
[0014] The VAE training device provided by the present specification comprises:
[0015] an acquisition module configured to acquire a VAE model to be trained and a sample image;
[0016] a mapping module configured to input the sample image into the VAE model to be trained, and determine a latent variable distribution corresponding to the sample image by an encoding module of the VAE model to be trained;
[0017] a reconstruction module configured to input the latent variable distribution into a decoding module of the VAE to be trained, and generate a reconstructed image of the sample image according to the latent variable distribution by the decoding module;
[0018] an extraction module configured to extract a first image feature of the reconstructed image and a second image feature of the sample image;
[0019] a loss determination module configured to determine a perceptual loss according to a difference between the first image feature and the second image feature;
[0020] a training module configured to train the VAE model to be trained according to at least the perceptual loss.
[0021] The specification provides a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the VAE training method.
[0022] The specification provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the VAE training method when executing the program.
[0023] The above at least one technical solution adopted by the embodiments of the specification can achieve the following beneficial effects:
[0024] The embodiments of the specification disclose a VAE training method, which inputs a sample image into a VAE model to be trained, maps the sample image to a latent space by an encoding module of the VAE model, obtains a latent variable distribution corresponding to the sample image, converts the latent variable distribution into an image by a decoding module of the VAE model to generate a reconstructed image, extracts a first image feature from the reconstructed image and a second image feature from the sample image, determines a perceptual loss according to a difference between the two image features, and trains the VAE model to be trained according to the perceptual loss. The above method can accurately train a high-precision VAE model. BRIEF DESCRIPTION OF DRAWINGS
[0025] The accompanying drawings, which are included to provide a further understanding of the specification, constitute a part of the specification, and the illustrative embodiments of the specification and their description serve to explain the specification, and do not constitute an improper limitation on the specification. In the drawings:
[0026] Figure 1 A flow chart of a training method of a VAE provided for an embodiment of the present specification;
[0027] Figure 2 A schematic diagram of a training device of a VAE provided for an embodiment of the present specification;
[0028] Figure 3 A structural schematic diagram of an electronic device provided for an embodiment of the present specification. DETAILED DESCRIPTION
[0029] In order to make the purpose, technical scheme and advantages of the present specification clearer, the technical scheme of the present specification will be described clearly and completely in combination with the embodiments of the present specification and corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present specification, not all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present specification.
[0030] The technical scheme provided by each embodiment of the present specification will be described in detail below in combination with the drawings.
[0031] Figure 1 A flow chart of a training method of a VAE provided for an embodiment of the present specification, comprising the following steps:
[0032] S100: obtaining a VAE model to be trained and sample images.
[0033] In the embodiments of the present specification, the method shown as Figure 1 The subject that trains the VAE to be trained by performing the method can be any electronic device, such as a personal computer, a server, a server cluster, etc. The following will be described only by taking the server as an example.
[0034] The server can first obtain a VAE model to be trained and sample images.
[0035] For the VAE model to be trained, the server can directly construct the model structure of the VAE model, and randomly initialize the model parameters in the constructed VAE model as the VAE model to be trained. In order to improve the training efficiency, the server can also obtain a pre-trained VAE model as the VAE model to be trained. Specifically, the server can obtain the VAE network in the trained image generation model as the VAE model to be trained, wherein the trained image generation model can include at least one of a stable diffusion model, a Flux model.
[0036] If the obtained to-be-trained VAE model is a pre-trained VAE model, subsequent training of the to-be-trained VAE model is essentially fine-tuning training.
[0037] For sample images, in order to enable the trained VAE model to adapt to various images, in the embodiments of the present specification, each sample image can be obtained through an image enhancement method.
[0038] Specifically, the server can first obtain a standard image, then perform image enhancement on the standard image to obtain at least one enhanced image corresponding to the standard image, and finally take the standard image and each enhanced image as the obtained sample image.
[0039] In the embodiments of the present specification, when the server performs image enhancement on the standard image, the server can enhance the standard image in terms of space, color, and noise.
[0040] For space enhancement, the server can perform scaling processing on the size of the standard image, and / or perform random cropping processing on the standard image. The scaling processing on the size of the standard image can enable the trained VAE model to adapt to different image sizes. The random cropping processing on the standard image, i.e., randomly cropping a part of the standard image, can enable the VAE model to adapt to local features in the image.
[0041] For color enhancement, the server can perform random color jitter processing, and / or random Gamma change processing, and / or random color channel processing, and / or contrast enhancement processing on the standard image. The random color jitter processing is random adjustment of parameters such as brightness, saturation, and contrast of the standard image, simulates different lighting conditions, and enables the trained VAE model to adapt to images under different lighting conditions. The random Gamma change processing is used to adjust the brightness and darkness of the standard image in a non-linear manner, simulates overexposure or underexposure environments, and enables the trained VAE model to adapt to images under different exposure environments. The random color channel processing is used to shuffle the order of color channels (such as RGB channels) in the standard image, preventing the model from relying on a specific color pattern (such as considering that the sky must be blue). The contrast enhancement processing is used to increase or equalize the contrast of the standard image, enabling the trained VAE model to adapt to images under dark light or low-contrast scenes.
[0042] For noise enhancement, the server can add blur noise and / or random image compression noise to the standard image. The added blur noise can specifically include superposition of one or more of Gaussian blur noise and motion blur noise, enabling the trained VAE model to adapt to image quality loss caused by lens defocus or object movement. The added random image compression noise can specifically include compression artifact noise, enabling the trained VAE model to adapt to image quality loss caused by image compression during image transmission or storage.
[0043] S102: input the sample image into the VAE model to be trained, and determine, by an encoding module of the VAE model to be trained, a latent variable distribution corresponding to the sample image.
[0044] The VAE model generally includes an encoding module and a decoding module. The encoding module is configured to map the features of the sample image to a lower-dimensional latent space to obtain a distribution of the features of the sample image in the lower-dimensional latent space, that is, a latent variable distribution. Since the features of the sample image are mapped to the lower-dimensional latent space, after the VAE is deployed into an image generation model, other networks in the image generation model can operate on the lower-dimensional features of the sample image in the latent space. For example, in a stable diffusion model, the lower-dimensional image features of the image in the latent space can be denoised, which can effectively improve the efficiency of the image generation model in generating images.
[0045] S104: input the latent variable distribution into a decoding module of the VAE to be trained, and generate, by the decoding module, a reconstructed image of the sample image according to the latent variable distribution.
[0046] In contrast to the encoding module, the decoding module in the VAE model is configured to regenerate a reconstructed image according to the latent variable distribution of the features of the sample image in the latent space. For example, after the lower-dimensional image features of the image in the latent space are denoised in the stable diffusion model, the denoised image features can be restored by the encoding module of the VAE to reconstruct a high-dimensional image. In this way, the denoising processing of the image features in the stable diffusion model is performed in the lower-dimensional latent space, and the high-dimensional image is reconstructed after the processing, which can improve the efficiency of generating images.
[0047] S106: extract first image features of the reconstructed image and second image features of the labeled image of the sample image.
[0048] In the embodiments of the present disclosure, since the sample image input into the VAE to be trained can include the standard image or the enhanced image obtained by performing image enhancement on the standard image, when a sample image is a standard image, the labeled image of the sample image is the standard image, that is, the sample image and the labeled image are the same at this time. When a sample image is an enhanced image obtained by performing image enhancement on a standard image, the labeled image of the sample image is the standard image corresponding to the enhanced image, that is, the original standard image before image enhancement.
[0049] The server can extract the first image features of the reconstructed image and the second image features of the labeled image by using a pre-trained feature extraction model.
[0050] In order to make the VAE model pay attention to both the global features in the sample image for representing the semantic of the image and the local features in the sample image for representing the details of the image, in the embodiments of the present specification, the first image features of the reconstructed image at different scales can be extracted by the feature extraction model, and the second image features of the labeled image of the sample image at different scales can be extracted. That is, the first image features and the second image features are both image feature pyramids composed of image features at different scales.
[0051] Specifically, the pre-trained feature extraction model can be a resnet model, such as resnet18.
[0052] S108: determining a perception loss according to the difference between the first image features and the second image features.
[0053] After the server extracts the first image features from the reconstructed image and extracts the second image features from the labeled image, the server can determine the difference between the first image features and the second image features, and determine the perception loss according to the difference. Specifically, the similarity, such as the cosine similarity, between the first image features and the second image features can be determined. The greater the similarity, the smaller the difference between the first image features and the second image features, and the smaller the perception loss. Conversely, the smaller the similarity, the greater the difference between the first image features and the second image features, and the greater the perception loss.
[0054] When the first image features and the second image features are both multi-scale image features, the server can determine the difference between the first image features at each scale and the second image features at the scale, and determine the perception loss between the reconstructed image and the labeled image of the sample image according to the difference determined for each scale. Specifically, the difference determined for each scale can be weighted to determine the comprehensive difference between the reconstructed image and the labeled image, and then the perception loss can be determined according to the comprehensive difference. The greater the comprehensive difference, the greater the perception loss. The smaller the comprehensive difference, the smaller the perception loss.
[0055] S110: training the VAE model to be trained at least according to the perception loss.
[0056] In the embodiments of the present specification, the goal of training the VAE model is to make the reconstructed image output by the VAE as close as possible to the labeled image of the sample image. That is, the more the reconstructed image output by the VAE model approaches the labeled image, the higher the accuracy of the VAE model. Therefore, the server can adjust the model parameters of the VAE model to be trained at least with the goal of reducing the perception loss determined in step S108, so as to complete one iteration of training. The training process of the VAE model can be repeated by performing steps S100-S110 to perform multiple iterations of training.
[0057] To further improve the training accuracy of the VAE model, in addition to adjusting the model parameters of the to-be-trained VAE model according to the above perception loss, the server can also adjust the model parameters of the to-be-trained VAE model according to the KL divergence loss and the reconstruction loss, in combination with the above perception loss.
[0058] Specifically, the server can determine the KL divergence loss between the latent variable distribution determined in step S102 and a preset distribution according to the latent variable distribution and the preset distribution. The KL divergence loss represents the difference between the latent variable distribution and the preset distribution. The preset distribution can be a standard normal distribution. The server can also determine the reconstruction loss between the reconstructed image and the labeled image of the sample image according to the reconstructed image and the labeled image of the sample image. The reconstruction loss can be the mean square error between the reconstructed image and the labeled image.
[0059] After determining the above perception loss, KL divergence loss and reconstruction loss, the server can determine a comprehensive loss according to the perception loss, KL divergence loss and reconstruction loss, and adjust the model parameters of the to-be-trained VAE model with the training target of reducing the comprehensive loss. In determining the comprehensive loss, the server can weight the perception loss, KL divergence loss and reconstruction loss to obtain the comprehensive loss.
[0060] It should be noted that the order of determining the above three losses by the server is not sequential, and the above three losses can be determined simultaneously.
[0061] In addition to the above three losses, a distillation loss can also be introduced in this specification to train the to-be-trained VAE model. Specifically, the to-be-trained VAE model obtained by the server in step S100 can be used as a student model, and a VAE model with higher accuracy can also be obtained as a teacher model. For example, the VAE network obtained from the stable diffusion model is used as the to-be-trained VAE model, that is, the student model, and the VAE network obtained from the Flux model is used as the teacher model with higher accuracy. In steps S102-S104, in addition to using the student model to obtain the reconstructed image, the sample image can also be input into the teacher model, and the teacher model also uses the same way as the student model to reconstruct the sample image to obtain the reconstructed image. In step S108, the distillation loss can be determined according to the reconstructed image output by the student model and the reconstructed image output by the teacher model, and the perception loss, KL divergence loss, reconstruction loss and distillation loss can be weighted to obtain the comprehensive loss. Finally, in step S110, the model parameters of the to-be-trained VAE model are adjusted with the training target of reducing the comprehensive loss.
[0062] In addition, when the to-be-trained VAE model obtained by the server in step S100 is a pre-trained VAE model extracted from a trained image generation model, since improving the accuracy of the to-be-trained VAE model at this time mainly depends on the decoding module of the to-be-trained VAE model, when step S110 is performed, the server can keep the model parameters in the encoding module of the to-be-trained VAE model unchanged, and only fine-tune the model parameters of the decoding module of the to-be-trained VAE model.
[0063] In order to speed up the training speed of the to-be-trained VAE model, in the embodiment of the present specification, the server can continuously reduce the learning rate based on the iterative training in the process of iterative training of the to-be-trained VAE model. Specifically, the server can first use the method shown in the following formula (1) to perform iterative training on the to-be-trained VAE model based on a preset learning rate, and in the process of iterative training, every time a preset number of iterations is performed, the learning rate based on the iterative training is reduced. Figure 1 The learning rate determines the amplitude of adjusting the model parameters of the to-be-trained VAE model in each iteration training process. In the initial training stage, since the accuracy of the to-be-trained VAE model is far from meeting the requirements, the model parameters of the VAE model can be greatly adjusted through a larger learning rate. With the continuous progress of iterative training, the accuracy of the to-be-trained VAE model gradually improves, and its model parameters are more suitable for more fine adjustment. Therefore, with the increase of the number of iterations, the learning rate can be continuously reduced to gradually reduce the amplitude of adjusting the model parameters, which can effectively improve the training speed.
[0064] The above-mentioned preset number can be set as needed, for example, it can be set to reduce the learning rate every 1000 iterations. When reducing the learning rate, a polynomial learning rate reduction strategy can be used to determine the amplitude of reducing the learning rate each time.
[0065] In the embodiment of the present specification, in order to ensure the training accuracy while trying to reduce the memory occupation of the server and improve the computing efficiency, the server can also use a half-precision optimization method to train the to-be-trained VAE model when training the to-be-trained VAE model. Specifically, FP16 (i.e. 16-bit floating point number) can be used to represent the model parameters (including weights and activation values) of the to-be-trained VAE model, gradients and the above-mentioned three kinds of losses in the training process.
[0066] In order to minimize the loss of training accuracy caused by gradient explosion or gradient disappearance when using half-precision optimization, the server can use the method of combining half-precision optimization with exponential moving average (EMA) to train the VAE model to be trained. That is, FP16 (i.e., 16-bit floating point number) is used to represent the model parameters (including weights and activation values), gradients and the above three losses of the VAE model to be trained during training. When updating the model parameters each time, the updated model parameters are not directly used, but the updated model parameters are smoothed by calculating the weighted average of the changes of the model parameters in the iteration process, and the historical model parameters (i.e., the model parameters before updating) can be given a higher weight to suppress the instability in the training process caused by the error brought by using FP16 to represent the model parameters, gradients and the above three losses.
[0067] The above is a VAE training method provided by an embodiment of the present specification. Based on the same idea, the present specification also provides a corresponding device, a storage medium and an electronic device.
[0068] Figure 2 A VAE training device provided by an embodiment of the present specification is shown in the schematic diagram, and the device comprises:
[0069] The acquisition module 201 is configured to acquire a VAE model to be trained and a sample image.
[0070] The mapping module 202 is configured to input the sample image into the VAE model to be trained, and determine the latent variable distribution corresponding to the sample image by using the encoding module of the VAE model to be trained.
[0071] The reconstruction module 203 is configured to input the latent variable distribution into the decoding module of the VAE to be trained, and generate a reconstructed image of the sample image according to the latent variable distribution by using the decoding module.
[0072] The extraction module 204 is configured to extract the first image feature of the reconstructed image and the second image feature of the labeled image of the sample image.
[0073] The loss determination module 205 is configured to determine the perceptual loss according to the difference between the first image feature and the second image feature.
[0074] The training module 206 is configured to train the VAE model to be trained according to at least the perceptual loss.
[0075] Optionally, the obtaining module 201 is specifically configured to obtain a standard image; perform image enhancement on the standard image to obtain at least one enhanced image corresponding to the standard image; and take the standard image and / or the enhanced image as the obtained sample image; wherein when the sample image is the standard image, a labeled image of the sample image is the standard image, and when the sample image is an enhanced image, the labeled image of the sample image is a standard image corresponding to the enhanced image.
[0076] Optionally, the obtaining module 201 is specifically configured to perform at least one of spatial enhancement, color enhancement and noise enhancement on the standard image; the spatial enhancement includes at least one of scaling processing on a size of the standard image and random cropping processing on the standard image; the color enhancement includes at least one of random color dithering processing, random Gamma change processing, random color channel processing and enhanced contrast processing on the standard image; and the noise enhancement includes at least one of adding blur noise and random image compression noise to the standard image.
[0077] Optionally, the extracting module 204 is specifically configured to extract, by a pre-trained feature extraction model, first image features of the reconstructed image at different scales and second image features of a labeled image of the sample image at the different scales.
[0078] The loss determining module 205 is specifically configured to determine, for each scale, a difference between the first image features at the scale and the second image features at the scale; and determine a perceptual loss between the reconstructed image and the labeled image of the sample image according to the differences respectively determined for the different scales.
[0079] Optionally, the training module 206 is specifically configured to determine a KL divergence loss between the latent variable distribution and a preset distribution according to the latent variable distribution and the preset distribution; determine a reconstruction loss between the reconstructed image and the labeled image of the sample image according to the reconstructed image and the labeled image of the sample image; determine a comprehensive loss according to the perceptual loss, the KL divergence loss and the reconstruction loss; and adjust model parameters of the to-be-trained VAE model with the comprehensive loss as a training target.
[0080] Optionally, the obtaining module 201 is specifically configured to obtain a pre-trained VAE model as the to-be-trained VAE model.
[0081] The training module 206 is specifically configured to keep model parameters of the encoding module unchanged and perform fine-tuning training on model parameters of the decoding module.
[0082] Optionally, the training module 206 is specifically configured to perform iterative training on the VAE model to be trained based on a preset learning rate, and during the iterative training process, reduce the learning rate on which the iterative training is based after each preset number of iterative trainings.
[0083] Optionally, the training module 206 is specifically used to train the VAE model to be trained using a half-precision optimization method.
[0084] This specification also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can be used to perform the VAE training method provided above.
[0085] based on Figure 1 The VAE training method shown in this specification also provides Figure 3 The structural diagram of the electronic device shown in FIG. Figure 3 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, memory, and non-volatile storage, and may also include other hardware required for its operations. The processor reads the corresponding computer program from the non-volatile storage into the memory and then runs it to implement the aforementioned VAE training method.
[0086] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.
Claims
1. A method for training a variational autoencoder (VAE), comprising: Get the VAE model to be trained and the sample images; Inputting the sample image into the VAE model to be trained, and determining the latent variable distribution corresponding to the sample image through the encoding module of the VAE model to be trained; Inputting the latent variable distribution into the decoding module of the VAE to be trained, and generating a reconstructed image of the sample image according to the latent variable distribution by the decoding module; extracting a first image feature of the reconstructed image and a second image feature of the annotated image of the sample image; determining a perceptual loss based on a difference between the first image feature and the second image feature; The VAE model to be trained is trained at least according to the perceptual loss.
2. The method according to claim 1, wherein obtaining a sample image comprises: Acquire standard images; Performing image enhancement on the standard image to obtain at least one enhanced image corresponding to the standard image; The standard image and / or the enhanced image is used as the acquired sample image; wherein, when the sample image is the standard image, the annotated image of the sample image is the standard image; when the sample image is the enhanced image, the annotated image of the sample image is the standard image corresponding to the enhanced image.
3. The method according to claim 2, further comprising: performing image enhancement on the standard image; performing at least one of spatial enhancement, color enhancement, and noise enhancement on the standard image; The spatial enhancement comprises at least one of scaling the size of the standard image and randomly cropping the standard image; The color enhancement includes performing at least one of random color dithering processing, random gamma change processing, random color channel processing, and contrast enhancement processing on the standard image; The noise enhancement includes adding at least one of blur noise and random image compression noise to the standard image.
4. The method according to claim 1, wherein extracting the first image feature of the reconstructed image and the second image feature of the annotated image of the sample image comprises: Extracting first image features of the reconstructed image at different scales and extracting second image features of the annotated image of the sample image at different scales through a pre-trained feature extraction model; Determining the perceptual loss according to the difference between the first image feature and the second image feature specifically includes: For each scale, determining a difference between a first image feature at the scale and a second image feature at the scale; A perceptual loss between the reconstructed image and the annotated image of the sample image is determined based on the differences determined for each scale respectively.
5. The method of claim 1, wherein training the VAE model to be trained is performed at least based on the perceptual loss, specifically comprising: Determining a KL divergence loss between the latent variable distribution and the preset distribution according to the latent variable distribution and the preset distribution; and, determining a reconstruction loss between the reconstructed image and the annotated image according to the reconstructed image and the annotated image of the sample image; Determining a comprehensive loss based on the perceptual loss, the KL divergence loss, and the reconstruction loss; With the reduction of the comprehensive loss as the training goal, the model parameters of the VAE model to be trained are adjusted.
6. The method according to claim 1, wherein obtaining the VAE model to be trained comprises: Get the pre-trained VAE model as the VAE model to be trained; Training the VAE model to be trained specifically includes: The model parameters of the encoding module are kept unchanged, and the model parameters of the decoding module are fine-tuned and trained.
7. The method according to claim 1, wherein training the VAE model to be trained comprises: Based on a preset learning rate, the VAE model to be trained is iteratively trained, and during the iterative training process, the learning rate based on the iterative training is reduced after each preset number of iterative trainings.
8. The method according to claim 1, wherein training the VAE model to be trained comprises: The VAE model to be trained is trained using a semi-precision optimization method.
9. A training device for a variational autoencoder (VAE), comprising: The acquisition module is used to obtain the VAE model to be trained and the sample image; A mapping module, configured to input the sample image into the VAE model to be trained, and determine the latent variable distribution corresponding to the sample image through the encoding module of the VAE model to be trained; A reconstruction module, configured to input the latent variable distribution into a decoding module of the VAE to be trained, and generate a reconstructed image of the sample image according to the dependent variable distribution through the decoding module; an extraction module, configured to extract a first image feature of the reconstructed image and a second image feature of the annotated image of the sample image; a loss determination module, configured to determine a perceptual loss based on a difference between the first image feature and the second image feature; A training module is used to train the VAE model to be trained based on at least the perceptual loss.
10. A computer-readable storage medium storing a computer program, wherein the computer program implements the method according to any one of claims 1 to 8 when executed by a processor.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 8 when executing the program.