Method and System for Generating Esophageal Optical Coherence Tomography Images

By confronting the variational autoencoding network model V', the encoder E is responsible for the discrimination and encoding tasks at the same time, solving the problem of insufficient image quality and diversity in the esophageal OCT image processing in the prior art, and achieving efficient generation of virtual esophageal OCT images close to real images.

CN114022422BActive Publication Date: 2025-05-27SUZHOU INST OF BIOMEDICAL ENG & TECH CHINESE ACADEMY OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111233747.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-22
Publication Date
2025-05-27
Estimated Expiration
2041-10-22

AI Technical Summary

Technical Problem

The prior art is difficult to generate high-quality and diverse virtual images in the esophageal OCT image processing. Traditional data augmentation methods lead to data overfitting and loss of global topological information. Deep learning models such as VAE and GAN have limitations in training stability and image quality.

Method used

Using the adversarial variational autoencoding network model V, the encoder E undertakes the discrimination and encoding tasks at the same time. By alternately optimizing the encoder E and the generator G, an adversarial variational autoencoding network model V' is constructed to generate a virtual esophageal OCT image close to the real image.

Benefits of technology

It realizes the generation of virtual esophageal OCT images close to real images, simplifies the network structure, improves training efficiency, and produces results better than PGGAN models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114022422B_ABST
    Figure CN114022422B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for generating esophageal optical coherence tomography images. The method includes the following steps: 1) Collect esophageal OCT images and perform preprocessing to serve as training samples; 2) Construct an adversarial variational autoencoder network model V, which includes an encoder E and a generator G; 3) Train the adversarial variational autoencoder network model V. During the training process, the optimization processes of the encoder E and the generator G are alternately carried out. After convergence, the final adversarial variational autoencoder network model V' is obtained; 4) Use the trained adversarial variational autoencoder network model V' to generate virtual esophageal OCT images. The present invention can generate virtual esophageal OCT images close to real images. Among them, the encoder undertakes both the discrimination and encoding tasks at the same time, which can simplify the network structure, thereby making the training of the network more efficient; the present invention can obtain better results in generating esophageal OCT images than the PGGAN model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image processing, and particularly to a method and system for generating esophageal optical coherence tomography images. Background Art

[0002] Optical Coherence Tomography (OCT) is a relatively new non-invasive imaging modality. By combining an endoscopic probe, an OCT device can access the upper digestive tract and provide high-resolution images of biological tissues under a microscope. Benefiting from this, endoscopic OCT devices can visualize morphological changes caused by esophageal lesions and are used for the diagnosis of esophageal diseases. The healthy esophageal wall in OCT images has clearly visible layered tissues, while for patients with lesions, structural abnormalities may occur. Therefore, the automatic analysis of esophageal tissues plays an important role in the diagnosis of esophageal diseases and has received increasing attention in recent years.

[0003] In recent years, deep learning has been widely applied in the field of medical imaging. In the processing of esophageal OCT images, there are also studies showing that deep learning methods can better analyze data, such as obtaining more accurate segmentation results. A sufficient amount of data is a necessary condition for the full training of deep learning models, which is generally achieved through data augmentation techniques. However, the data generated by traditional data augmentation techniques (e.g., cropping, rotation, elastic deformation) are highly correlated with each other, making it difficult to truly ensure data diversity. In addition, these methods are all processed at the whole image level, which results in the enhanced images losing the global topological information of the input, thus leading to the generated samples containing unreasonable features.

[0004] Another way of data augmentation is to adopt a deep neural network with a specific structure. Among them, the most representative methods are the variational autoencoder (VAE) (D.P. Kingma and M. Welling, “Auto-encoding variational bayes,” in International Conference on Learning Representations (Iclr), 2014.) and the generative adversarial networks (GAN) (I.J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” Advances in Neural Information Processing Sys-tems 27 (Nips 2014), vol. 27, pp. 2672–2680, 2014.). These two models are widely used, but still have certain limitations. VAE is easy to train, but will generate blurred images lacking details. GAN can obtain higher image quality, but faces challenges in training stability and mode collapse. In response to the problems existing in GAN, researchers have proposed various solutions to address these challenges. A commonly used method is the coarse-to-fine generation strategy, that is, starting from generating low-resolution images and gradually improving the image quality. The representative model is PGGAN (T. Karras, T. Aila, S. Laine, and J. Lehtinen, “Progressive growing of gans for improved quality, stability, and variation,” in International Conference on Learning Representations (Iclr), 2018.). However, this improvement greatly increases the computational burden. In addition, some research has proposed a hybrid model that combines VAE and GAN to make the network easy to train and have relatively high image generation ability. VAEGAN is a typical case of the hybrid model (A.B.L. Larsen, S.K. and O. Winther, “Autoencoding beyond pixels using a learned similarity metric,” CoRR, vol. abs / 1512.09300, 2015. [Online]. Available: http: / / arxiv.org / abs / 1512.09300.), However, the hybrid model usually contains three networks: an encoder, a decoder, and a discriminator, with a relatively complex structure, and the quality of the generated images is often inferior to that of GANs.

[0005] Therefore, a more reliable solution is needed now. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide a method and system for generating esophageal optical coherence tomography images in view of the above-mentioned deficiencies in the prior art.

[0007] To solve the above technical problem, the technical solution adopted by the present invention is: A method for generating esophageal optical coherence tomography images, comprising the following steps:

[0008] 1) Collect esophageal OCT images and perform preprocessing as training samples;

[0009] 2) Construct an adversarial variational autoencoder network model V, which includes an encoder E and a generator G, where the encoder E undertakes both discrimination and encoding tasks;

[0010] 3) Use the training samples obtained in step 1) to train the adversarial variational autoencoder network model V constructed in step 2). During the training process, the optimization processes of the encoder E and the generator G are alternated. After convergence, the final adversarial variational autoencoder network model V' is obtained;

[0011] 4) Use the trained adversarial variational autoencoder network model V' to generate virtual esophageal OCT images.

[0012] Preferably, in the adversarial variational autoencoder network model V in step 2), L like terms are used to measure the reduction degree of the network, L prior terms are used to measure the closeness of the latent vector to the Gaussian distribution, and L gan represents the adversarial loss of the network model V;

[0013] Use E 1 to represent the encoding output of the encoder, E 2 to represent the discrimination output of the encoder; x represents the input real sample, x rDenotes the generated data reconstructed by the generator G from the encoding vector z, as shown in Equation (1):

[0014]

[0015] The meaning of Equation (1) is that the encoding z output by the real sample x through the encoder E follows the distribution q(z|x), and the reconstructed sample x obtained by the hidden vector z through the generator G r follows the distribution p(x|z);

[0016] Adopt z p Denotes a random vector sampled from a Gaussian distribution, z p has the same dimension as z, x p Denotes the sampled sample reconstructed by the generator G through z p ;

[0017] L like The calculation formula of the term is:

[0018]

[0019] L prior The calculation formula of the term is:

[0020] L prior = D KL (q(z∣x)||p(z)) (3);

[0021] Among them, D KL is the KL divergence, p(z) = N(0,1) represents the standard normal distribution;

[0022] L gan The calculation formula of the term is:

[0023]

[0024] Among them, in, E represents the expectation, and the subscript indicates that the random variable is x, and this variable x follows the distribution of the sample data, that is, x is sampled from the real sample dataset; in, E represents the expectation, and the subscript indicates that z is the random variable, and z follows the p z (z) distribution, that is, sampled from the normal distribution.

[0025] Preferably, in step 3), the loss function L for optimizing the encoder E E The calculation formula is:

[0026]

[0027] The loss function L for optimizing the generator G G The calculation formula is:

[0028]

[0029] Among them, is the currently used generator, is the currently used encoder, λ 1 and λ 2 are weight parameters.

[0030] Preferably, in step 3), the steps of training the anti-variational auto-encoding network model V specifically include:

[0031] Ⅰ. Training the encoder E:

[0032] Ⅰ-1. Keep the parameters of the generator G fixed, extract samples from the training samples to form real samples x, these samples pass through the encoder E to obtain the hidden vector z, and according to the sample data x and the hidden vector z, combine formulas (1) and (3) to obtain the L prior term;

[0033] Ⅰ-2. The hidden vector z generates the reconstructed sample x r through the generator G with fixed parameters, and calculate the L like term in combination with formula (2) to measure the similarity between x and x r ;

[0034] Ⅰ-3. z P represents a vector sampled from the standard normal distribution with the same dimension as the hidden vector. Input z p into the generator G to generate a fake sample x p ; Among them, the encoder E also undertakes the task of the discriminator and can distinguish real samples x, reconstructed samples x r and fake samples x p , and calculate the L gan term through formula (4) to measure the discrimination ability of the encoder E;

[0035] Ⅰ-4. Finally, calculate the loss function L E of the encoder according to formula (5), and finally complete the optimization of the encoder E through the gradient descent method;

[0036] Ⅱ. Training the generator G:

[0037] Ⅱ-1. Keep the parameters of the encoder E fixed, extract samples from the training samples to form real samples x, obtain the hidden vector z through the encoder E, and according to the sample data x and the hidden vector z, combine formulas (1) and (3) to obtain the L prior term;

[0038] Ⅱ-2. The hidden vector z generates the reconstructed sample x r, calculate L by combining with formula (2) like term

[0039] Ⅱ-3, z P denotes a vector sampled from the standard normal distribution and having the same dimension as the hidden vector. Input z p into the generator G to generate a fake sample x p ; calculate the L gan term by formula (4) to measure the discrimination ability of the encoder E

[0040] Ⅱ-4. Finally, calculate the loss function L G of the encoder according to formula (6), and finally complete the optimization of the generator G through the gradient descent method

[0041] Ⅲ. The optimization processes of the encoder E and the generator G are alternately performed. After convergence, the final adversarial variational auto-encoding network model V' is obtained

[0042] Preferably, the method for preprocessing esophageal OCT images is as follows

[0043] First, normalize the collected esophageal OCT images: for any image, subtract the gray value of any pixel of the image from the mean value of the gray values of all pixels of the image and then divide by the standard deviation of the gray values of all pixels of the image

[0044] Then set the image input size. Images larger than the input size are downsampled to the input size, and images smaller than the input size are padded with zeros to the input image size

[0045] The present invention also provides a system for generating esophageal optical coherence tomography images, which generates virtual esophageal OCT images by using the method described above

[0046] The present invention also provides a storage medium, on which a computer program is stored, and when the program is executed, it is used to implement the method described above

[0047] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described above is implemented

[0048] The beneficial effects of the present invention are as follows: The method and system for generating esophageal optical coherence tomography images provided by the present invention can generate virtual esophageal OCT images close to real images. Among them, the encoder undertakes both discrimination and encoding tasks at the same time, which can simplify the network structure, thereby making the training of the network more efficient; the present invention can obtain better esophageal OCT image generation results than the PGGAN model Brief Description of the Drawings

[0049] Figure 1 It is a schematic structural diagram of the adversarial variational auto - encoder network model V of the present invention;

[0050] Figure 2 It is the generation result of esophageal OCT images in Embodiment 3 of the present invention. Detailed implementation manners

[0051] The following further elaborates on the present invention in conjunction with embodiments, so that those skilled in the art can implement it with reference to the text of the specification.

[0052] It should be understood that terms such as "having", "comprising", and "including" used herein do not exclude the presence or addition of one or more other elements or their combinations.

[0053] Embodiment 1

[0054] A method for generating esophageal optical coherence tomography images in this embodiment includes the following steps:

[0055] 1) Collect esophageal OCT images and perform pre - processing as training samples.

[0056] The pre - processing method is as follows: First, normalize the collected esophageal OCT images: For any image, subtract the gray - scale value of any pixel of the image from the mean value of all pixel gray - scale values of the image and then divide by the standard deviation of all pixel gray - scale values of the image;

[0057] Then set the image input size. Images larger than the input size are down - sampled to the input size, and images smaller than the input size are padded with zeros to the input image size.

[0058] 2) Construct an adversarial variational auto - encoder network model V, which includes an encoder E and a generator G, where the encoder E undertakes both discrimination and encoding tasks.

[0059] The network structure of the adversarial variational auto - encoder network model V in this embodiment is as Figure 1 shown, where the L like term is used to measure the reduction degree of the network, the L prior term is used to measure the closeness of the latent vector to the Gaussian distribution, and the L gan represents the adversarial loss of the network model V;

[0060] Use E 1 to represent the encoding output of the encoder, and E 2 to represent the discrimination output of the encoder; x represents the input real sample, and x r represents the generated data reconstructed by the generator G through the encoding vector z, as shown in formula (1):

[0061]

[0062] The meaning of formula (1) is that the encoded z output by the real sample x through the encoder E follows the distribution q(z|x), and the reconstructed sample x obtained by the hidden vector z through the generator G r follows the distribution p(x|z);

[0063] Using z p represents a random vector sampled from a Gaussian distribution, and z p has the same dimension as z, and x p represents the sampled sample reconstructed by the generator G through z p ;

[0064] L like The calculation formula for the term is:

[0065]

[0066] L prior The calculation formula for the term is:

[0067] L prior = D KL (q(z∣x)||p(z)) (3);

[0068] where D KL is the KL divergence, and p(z) = N(0,1) represents the standard normal distribution;

[0069] L gan The calculation formula for the term is:

[0070]

[0071] where , E represents the expectation, and the subscript indicates that the random variable is x, and this variable x follows the distribution of the sample data, that is, x is sampled from the real sample dataset; , E represents the expectation, and the subscript indicates that z is the random variable, and z follows the p z (z) distribution, that is, sampled from the normal distribution.

[0072] In a preferred embodiment, the structures of the encoder E and the generator G are shown in Table 1. The encoder E includes an input image layer, several downsampling units (each downsampling unit includes several convolutional blocks and a downsampling layer), a global pooling layer (Global Average), and finally two 1×1 convolutional blocks (respectively outputting E 1 and E 2 ), and the generator G includes several upsampling units. The first upsampling unit includes Latent Vector E 1 (E 1is a hidden vector, i.e., the output result of the encoder E, which is the input of the decoder G), convolutional blocks, Reshape function, upsampling layers. The last upsampling unit includes a 1×1 convolutional block, and the remaining upsampling units include two 3×3 convolutional blocks and an upsampling layer.

[0073] The network depths of the encoder and decoder can be adjusted according to the resolution of the input image. The adjustment method is to change the number of downsamplings so that the network reduces the minimum value of the image size to 2.

[0074] Table 1 Encoder and Decoder Structure Parameters

[0075]

[0076] 3) Use the training samples obtained in step 1) to train the anti-variational auto-encoding network model V constructed in step 2). During the training process, the optimization processes of the encoder E and the generator G are alternately carried out. After convergence, the final adversarial variational auto-encoding network model V' is obtained.

[0077] In the anti-variational auto-encoding network model V provided by the present invention, for the encoder E, it has three optimization objectives: 1) The encoded output E 1 can follow the prior Gaussian distribution; 2) The encoded output E 1 has a low reconstruction error; 3) The discriminative output E 2 can distinguish real data and generated samples. For the generator G, it has two optimization objectives: 1) Generate samples that are more similar to the real data distribution; 2) Have a small reconstruction error.

[0078] Based on this, the loss function L E for optimizing the encoder E is defined, and its calculation formula is:

[0079]

[0080] The loss function L G for optimizing the generator G is defined, and its calculation formula is:

[0081]

[0082] Among them, is the currently used generator, is the currently used encoder, and λ 1 and λ 2 are weight parameters. During the training process, the E network and the G network are trained in an alternating iterative optimization manner until the model converges or reaches the specified number of iterations.

[0083] Among them, the steps for training the anti-variational auto-encoding network model V specifically include:

[0084] Ⅰ. Training the encoder E:

[0085] Ⅰ-1. Keep the parameters of the generator G fixed, extract samples from the training samples to form real samples x. These samples pass through the encoder E to obtain the hidden vector z. According to the sample data x and the hidden vector z, combine formula (1) and formula (3) to obtain the L prior term;

[0086] Ⅰ-2. The hidden vector z passes through the generator G with fixed parameters to generate the reconstructed sample x r , and calculate L like term according to formula (2) to measure the similarity between x and x r ;

[0087] Ⅰ-3. z P represents a vector sampled from the standard normal distribution with the same dimension as the hidden vector. Input z p into the generator G to generate the fake sample x p ; Among them, the encoder E also undertakes the task of the discriminator and can distinguish real samples x, reconstructed samples x r and fake samples x p . The encoder distinguishes different samples by the differences in the encodings of the samples, and calculates L gan term according to formula (4) to measure the discriminative ability of the encoder E;

[0088] Ⅰ-4. Finally, calculate the loss function L E of the encoder according to formula (5), and finally complete the optimization of the encoder E through the gradient descent method;

[0089] Ⅱ. Training the generator G:

[0090] Ⅱ-1. Keep the parameters of the encoder E fixed, extract samples from the training samples to form real samples x, pass through the encoder E to obtain the hidden vector z. According to the sample data x and the hidden vector z, combine formula (1) and formula (3) to obtain the L prior term;

[0091] Ⅱ-2. The hidden vector z passes through the generator G to generate the reconstructed sample x r , and calculate L like term according to formula (2);

[0092] Ⅱ-3. z P represents a vector sampled from the standard normal distribution with the same dimension as the hidden vector. Input z p into the generator G to generate the fake sample x p ; Calculate L gan term according to formula (4) to measure the discriminative ability of the encoder E;

[0093] Ⅱ-4. Finally, the loss function L of the encoder is calculated according to formula (6). G , and through the gradient descent method, the optimization of the generator G is finally completed;

[0094] Ⅲ. The optimization processes of the encoder E and the generator G are alternately performed. After convergence, the final adversarial variational auto-encoding network model V' is obtained.

[0095] 4) In application, any hidden vector z with a specified dimension can be input to the generator G, and the trained adversarial variational auto-encoding network model V' is used to generate virtual esophageal OCT images that are similar to the training samples and close to real images.

[0096] The present invention can be used for generating esophageal OCT images. The encoder of the present invention undertakes both discrimination and encoding tasks, which can simplify the network structure, thereby making the training of the network more efficient.

[0097] Embodiment 2

[0098] This embodiment provides a system for generating esophageal optical coherence tomography images, which uses the method of Embodiment 1 to generate virtual esophageal OCT images.

[0099] This embodiment also provides a storage medium, on which a computer program is stored, and when the program is executed, it is used to implement the method of Embodiment 1.

[0100] This embodiment also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method of Embodiment 1 is implemented.

[0101] Embodiment 3

[0102] In this embodiment, the method of Embodiment 1 is used to generate esophageal OCT images, and a comparison is made with the images generated using the PGGAN model in the conventional technology, as follows.

[0103] In this embodiment, endoscopic OCT images of the esophagus of 9 healthy mice are collected to form an experimental data set. The training data set of the model comes from 7 of these mice, and a total of 10,500 esophageal images of mice are included (each individual contributes 1,500); the test data set of the model comes from the other 2 mice, and a total of 1,000 esophageal images of mice are included (each individual contributes 500). The test data set is used to evaluate the generation effect of the model. The image size is 256×256, and the pixel intensity is normalized to 0-1.

[0104] As Figure 2 shown, Figure 2 (a) is an endoscopic OCT image of the esophagus selected from the test data set,Figure 2 (b) shows the generation effect of PGGAN. It can be seen that the generated results of PGGAN can reflect the basic structure and texture of the esophagus, but there are still artifacts and noises in the generated images. In contrast, Figure 2 (c) shows the generation effect of the present invention. It can be seen that the generated images are visually closer to real images.

[0105] The quantitative evaluation indicators used in this embodiment are: Wasserstein distance (WD), Inception score (IS), Mode Score (MS), Fréchet Inception Distance (FID), and Maximum Mean Discrepancy (MMD). Among them, the larger the IS and MS indicators, the better the generation quality of the images, and the opposite is true for the other indicators. 1000 samples are generated by the PGGAN network and the network proposed by the present invention respectively, and the above evaluation indicators are calculated. The results are shown in Table 2. Among them, the evaluation indicators of 1000 samples in the test data are calculated as the gold standard of this embodiment. It can be seen that the evaluation indicators of the images generated by the network proposed by the present invention are all better than those of PGGAN and are closer to the gold standard.

[0106] Table 2 Numerical evaluation indicators of the generation results

[0107]

[0108] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to specific details.

Claims

1. A method for generating esophageal optical coherence tomography images, characterized in that, it includes the following steps: 1) Collect esophageal OCT images and perform preprocessing to obtain training samples; 2) Construct an adversarial variational autoencoder network model V, which includes an encoder E and a generator G, where the encoder E undertakes both discrimination and encoding tasks; 3) Use the training samples obtained in step 1) to train the adversarial variational autoencoder network model V constructed in step 2). During the training process, the optimization processes of the encoder E and the generator G are alternately performed. After convergence, the final adversarial variational autoencoder network model V' is obtained; 4) Use the trained adversarial variational autoencoder network model V' to generate virtual esophageal OCT images; In the adversarial variational autoencoder network model V in step 2), L like term is used to measure the restoration degree of the network, and L prior term is used to measure the closeness between the latent vector and the Gaussian distribution. L gan represents the adversarial loss of the network model V; Adopt E 1 represents the encoded output of the encoder, and E 2 represents the discriminative output of the encoder; x represents the input true sample, and x r represents the generated data reconstructed by the generator G through the hidden vector z, as shown in formula (1): The meaning of formula (1) is as follows: the hidden vector z output by the real sample x after passing through the encoder E follows the distribution q(z|x), and the reconstructed sample x obtained by the hidden vector z passing through the generator G r follows the distribution p(x|z); Enc represents the encoder, and Gen represents the generator; Adopt z p denotes a vector sampled from the standard normal distribution and having the same dimension as the hidden vector, z p has the same dimension as z, x p denotes the sampled sample reconstructed by the generator G through z p ; L like The calculation formula for the item is as follows: where E represents expectation; L prior The calculation formula for the item is as follows: L prior = D KL (q(z|x) || p(z)) (3); Among them, D KL is the KL divergence, and p(z) = N(0, 1) represents the standard normal distribution; L gan The calculation formula for the item is as follows: Among them, where E represents expectation, the subscript indicates that the random variable is x, and this variable x follows the distribution of the sample data, and x is sampled from the true sample dataset; where E represents expectation, the subscript indicates that z is a random variable, and z follows the p z (z) distribution, and the p z (z) distribution is sampled from a normal distribution; E 2 represents a part of the encoder output, G represents the generator, and pdata() represents the probability distribution of the sample data; In step 3), the loss function L for optimizing the encoder E E is calculated as follows: The loss function L for optimizing the generator G G has the following calculation formula: Among them, is the currently used generator, is the currently used encoder, λ 1 and λ 2 are weight parameters.

2. The method for generating esophageal optical coherence tomography images according to claim 1, characterized in that, in step 3), the steps of training the adversarial variational autoencoder network model V specifically include: Ⅰ. Train the encoder E: Ⅰ-1. Keep the parameters of the generator G fixed, extract samples from the training samples to form real samples x, these samples are passed through the encoder E to obtain the hidden vector z, and according to the sample data x and the hidden vector z, combine formula (1) and formula (3) to obtain L prior item; Ⅰ-2. The hidden vector z generates a reconstructed sample x through a generator G with fixed parameters r , and calculates L by combining formula (2) like terms to measure the similarity between x and x r ; Ⅰ-3, z p denotes a vector sampled from the standard normal distribution with the same dimension as the hidden vector. Input z p into the generator G to generate the sampled sample x p ; among them, the encoder E also undertakes the task of the discriminator and can distinguish the real sample x, the reconstructed sample x r and the sampled sample x p , and calculate the L gan term through formula (4) to measure the discrimination ability of the encoder E; Ⅰ-4. Finally, calculate the loss function L of the encoder according to formula (5). E Through the gradient descent method, the optimization of the encoder E is finally completed. Ⅱ. Train the generator G: Ⅱ-1. Keep the parameters of encoder E fixed, extract samples from the training samples to form the real sample x, obtain the hidden vector z through the encoder E, and obtain L according to the sample data x and the hidden vector z, combining formula (1) and formula (3). prior item; Ⅱ-2. The latent vector z generates a reconstructed sample x through the generator G r , and calculates L according to formula (2) like term Ⅱ-3, z p denotes a vector sampled from the standard normal distribution with the same dimension as the hidden vector. Input z p into the generator G to generate a fake sample x p ; calculate the L gan term by formula (4) to measure the discriminative ability of the encoder E; Ⅱ-4. Finally, calculate the loss function \(L\) of the encoder according to formula (6), G and finally complete the optimization of the generator \(G\) through the gradient descent method; Ⅲ. The optimization processes of the encoder E and the generator G are alternately performed. After convergence, the final adversarial variational autoencoder network model V' is obtained.

3. The method for generating esophageal optical coherence tomography images according to claim 1, characterized in that, the method for preprocessing esophageal OCT images is: First, normalize the collected esophageal OCT images: for any image, subtract the gray value of any pixel of the image from the mean value of the gray values of all pixels of the image and then divide by the standard deviation of the gray values of all pixels of the image; Then set the image input size. Images larger than the input size are downsampled to the input size, and images smaller than the input size are padded with zeros to the input image size.

4. A system for generating esophageal optical coherence tomography images, characterized in that, it uses the method described in any one of claims 1-3 to generate virtual esophageal OCT images.

5. A storage medium, on which a computer program is stored, characterized in that, when the program is executed, it is used to implement the method described in any one of claims 1-3.

6. A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements the method described in any one of claims 1-3.