A method for repairing film movie patches
By using U-net network to generate plaque masks and VAE methods based on three-domain conversion in film film plaque repair, the problems of high error detection rate and unnatural repair effects in the prior art are solved, and efficient and natural plaque repair effects are achieved.
Patent Information
- Application Number
- CN202510132881.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-06
AI Technical Summary
The existing film film plaque repair methods have problems such as high error detection rate and unnatural repair effects, especially when large-area plaque repair is poor.
The plaque mask generation technology based on U-net network is adopted, combined with the plaque repair method based on three-domain conversion, and the plaque area is segmented through the U-net network to generate an accurate plaque mask, and the damaged picture is encoded and restored by VAE encoder and generator to achieve natural and smooth repair of the image.
It significantly reduces the error detection rate, improves the accurate segmentation and repair effect of the plaque area, makes the repaired image more natural and smooth, and improves the overall quality of the movie picture.
Smart Images

Figure CN119559098B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image restoration, and particularly relates to a method for restoring film movie patches. Background Art
[0002] The restoration of film movie patches mainly relies on two methods: physical and digital. Traditional physical restoration methods mainly repair movie films manually, such as manual smearing and cutting and pasting. This method is not only time-consuming and laborious, but may also cause additional damage to the films. Digital restoration methods use film-to-digital conversion devices to digitally store film movies, and then use digital algorithms to repair the patches in the movie images.
[0003] Digital restoration methods are mainly divided into two types: one-stage and two-stage. One-stage methods directly use filters, such as median filtering, multi-level median filtering, LUM filtering, topological median filtering, etc., which have good restoration effects on small-area patches, but the restoration effects on large-area patches are very limited. Two-stage methods divide the restoration process into two stages: detection and restoration. First, the patch area is segmented to form a patch area mask, and then the patch mask is used to repair the patch area.
[0004] Typical patch mask generation methods include: 1) Manually drawing the patch mask, drawing a rectangular area at the patch, and forming a binary mask of the patch in this area. The mask generated by this method usually contains a large number of normal pixel points, affecting the restoration speed and effect of subsequent restoration work. 2) Mask generation method based on edge detection. This method uses the gray mutation and discontinuity of the patch to generate the patch mask. This method has poor segmentation ability for patches with unclear edges and small areas, and may cause misjudgment for some scenes with physical characteristics similar to the patches. 3) Patch detection algorithm based on the difference between consecutive frames. Since there is only a slight movement between consecutive adjacent frames in the movie projection image, the single-frame suddenness of the patch can be used to segment the patch area. Typical algorithms include SDI (Spike Detection Index) and ROD (Rank Ordered Difference). Due to the influence of compound degradation of film movies, a large number of false detections often occur in the mixed degradation area and complex scene area of the image.
[0005] Typical patch repair methods include: 1) The repair method based on bilinear interpolation calculates three single linear interpolations in the horizontal and vertical directions, and uses four surrounding adjacent pixel points to calculate the interpolated pixel points to be filled. This method has a good repair effect for small areas with small pixel differences, but for large areas with large pixel differences, the repair effect is relatively blurred. 2) The TELEA repair method is based on the (FMM) fast marching algorithm, which starts from known pixels and advances towards the missing area for repair. Therefore, the closer to the center of the missing area, the blurrier the repaired image. This method has an average repair effect when repairing large missing areas. 3) The Navier-Stokes repair algorithm uses the equations of fluid mechanics to restore the image by solving partial differential equations, which can preserve texture detail information but is difficult to complete high-quality repair tasks.
[0006] In view of the above problems, the present invention proposes a method for repairing patches in film movies. Summary of the Invention
[0007] The purpose of the present invention is to provide a method for repairing patches in film movies, aiming to solve the problems proposed in the above background technology.
[0008] The purpose of the present invention is achieved through the following technical solutions:
[0009] A method for repairing patches in film movies includes the following steps:
[0010] Step 1: Obtain the frame of the film movie to be repaired, and use a film-to-video conversion device to digitally transcribe the film movie.
[0011] Step 2: Use the U-net network to segment the patch area in the picture to form a patch mask.
[0012] Step 3: Use the patch repair method based on three-domain conversion to repair and reconstruct the damaged patch picture.
[0013] Step 4: Re-add the reconstructed image frame to the movie clip to complete the repair work.
[0014] Furthermore, the specific operations of step 2 are as follows:
[0015] Step 21: Input image preprocessing;
[0016] Image standardization: Set the scale_tensor parameter to standardize the size of the input image to 256×256 pixels.
[0017] Step 22: Network structure configuration;
[0018] Depth setting: Set the downsampling and upsampling depths of the U-net network to 4;
[0019] Channel setting: Set both the input channel and the output channel to 1;
[0020] Number of feature channels setting: Set the initial number of feature layers of the U-net network to 64;
[0021] Convolution layer configuration: In each convolution block, set the number of convolution layers to 2;
[0022] Upsampling method: During upsampling, use the transposed convolution method;
[0023] Step 23, Network training;
[0024] Activation function: Select ReLU as the activation function;
[0025] Normalization: After each convolution layer, use the Batch Norm normalization method to normalize the data after convolution;
[0026] Select and define the loss function: Select BCEWithLogitsLoss as the loss function to solve the binary classification problem of pixel points. The loss function combines the Sigmoid function and BCELoss, and the expression is as follows:
[0027] ;
[0028] where x is the original output of the model; y is the target value; sigmoid is the sigmoid function;
[0029] Optimizer configuration: Use the Adam optimizer for network training, and set the learning rate α to 0.001, the exponential decay rate of the first moment estimate β 1 to 0.9, and the exponential decay rate of the second moment estimate β 2 to 0.999;
[0030] Step 24, Plaque segmentation: Use the trained U-net network to segment the plaques in the input damaged image to obtain a binary mask for plaque segmentation.
[0031] Furthermore, the specific operations of step 3 are as follows:
[0032] Step 31, Domain definition;
[0033] Define the true damaged scene domain R , the masked covered scene domain X and the intact undamaged scene domain Y, where the intact and undamaged picture area Y contains adjacent and similar frames of the damaged picture, which are used to provide the image features required for restoration;
[0034] Step 32: Use the VAE encoder E to encode the picture into the latent space. The VAE encoder E includes E R , E X and E Y ; Use the VAE generator G to restore the image encoding. The VAE generator G includes G R , G X and G Y ;
[0035] Encode the image in the real damaged picture area R into the latent space r through the encoder E R to obtain the encoded representation Z R ; Encode the image in the masked covered picture area z r through the encoder X into the latent space x through the encoder E X to obtain the encoded representation Z X ; Encode the image in the intact and undamaged picture area z x through the encoder Y into the latent space y through the encoder E Y to obtain the encoded representation Z Y ; z y ;
[0036] Inside the latent space Z R and Z X The intersection of the domains forms a shared domain. The encoded representation of the damaged image in the shared domain is z x∩r ;
[0037] G R , GX and G Y perform decoding operations on the encoded representations of the images respectively to assist the encoder E R 、 E X and E Y in training, specifically: G R Decode the encoded representation z r to restore the image r ; G X Decode the encoded representation z x to restore the image x ; G Y Decode the encoded representation z y to restore the image y ;
[0038] Step 33, Feature mapping;
[0039] Independently train the mapping Z R and Z X between the shared domain and the intact domain Z Y of the damaged domain T Z , and the mapping T Z is used to convert the encoded information in the shared domain to the intact domain Z Y ;
[0040] Step 34, Use the VAE generator G to restore the information mapped by the mapping T Z to Z Y domain into an image: Through the mapping T Z to z x∩r perform a fitting mapping, and then decode by z y to achieve the restoration of the image; G Y The repair process of the damaged image
[0041] is represented by the following formula: r
[0042] ;
[0043] Among them, is the result of transforming the damaged image r from the R domain to the Y domain; is the VAE generator; is the feature mapping transformation; is the VAE encoder; " " indicates the connection of each process.
[0044] Furthermore, in the VAE encoder, assuming that the picture data in the latent space follows a Gaussian distribution, the VAE loss function of the real damaged picture domain R is expressed as:
[0045] ;
[0046] Among them, is the VAE loss function of the real damaged picture domain R ; The first term of the loss function is the encoder loss, KL is the KL divergence, is the mathematical expectation of encoding under the condition of the sample image r, is a normal distribution with a mean of 0 and a variance of ; The second term of the loss function is the generator loss, is the weight coefficient of the generator loss, is the mathematical expectation, is the generation result of the image, is the image sample; The third term of the loss function is , indicating the LSGAN loss.
[0047] Furthermore, a discriminator is introduced to expand the shared domain, and the total objective function for encoding the images r and x is:
[0048] ;
[0049] Among them, represents the encoder loss of the real damaged picture domain R、 mask-covered picture domain X ; represents the generator loss of the real damaged picture domain R、 mask-covered picture domain X ; is the discriminator loss of the real damaged picture domain R、 mask-covered picture domain X ; is the VAE loss function of the real damaged picture domain R ; is the VAE loss function of the mask-covered picture domain X ; The LSGAN loss introduced into the latent space.
[0050] Furthermore, by learning the mapping z x∩r , z y} between the training-coded image pairs T Z , image restoration and reconstruction are performed; the loss function of the mapping T Z is shown as follows:
[0051] ;
[0052] wherein, z x∩r is the coded representation of the damaged image in the shared domain; is the weight coefficient of the L1 loss of the mapping T; is the L1 loss of the mapping T, , which is constrained by the L1 norm; is the mapping transformation of the coded latent space, E is the mathematical expectation; is the introduced least squares loss LSGAN; is the weight coefficient of the feature matching loss; is the introduced feature matching loss, matching the discriminator and multiple-level feature layers in the VGG network, and the expression is as follows:
[0053] ;
[0054] wherein, represents the i-th layer feature layer of the discriminator ; represents the i-th layer feature layer of the VGG network; represents the activation times of the i-th layer feature layer of the discriminator ; represents the activation times of the i-th layer feature layer of the VGG network.
[0055] Compared with the prior art, the beneficial effects of the present invention are:
[0056] The present invention proposes a method for repairing film movie patches. This method first adopts a patch mask generation technique based on the U-net network, which can effectively reduce the false detection rate, accurately identify the patch areas in the picture, and generate a patch mask that accurately segments the patch areas. Further, the present invention also introduces a patch repair method based on three-domain transformation. This strategy first encodes the damaged picture, and then fits and maps the encoded information with the encoding of adjacent intact pictures to achieve image restoration. This strategy makes full use of the effective information in the time domain and the spatial domain, making the restored image more natural and smooth, and significantly improving the overall quality of the movie picture. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 is a flowchart of the present invention.
[0058] Figure 2 is a structural diagram of the U-net network.
[0059] Figure 3 is a schematic diagram of the patch repair method based on three-domain transformation.
[0060] Figure 4 is a result comparison between the patch mask generation method based on U-net proposed by the present invention and traditional detection methods; where (A) is the test image, (B) is the SDI result, (C) is the ROD detection, (D) is the edge detection, and (E) is the U-net detection.
[0061] Figure 5 is a result comparison between the patch repair method based on three-domain transformation proposed by the present invention and traditional repair methods; where (A) is the test image, (B) is the TELEA repair algorithm, (C) is the Navier-Stokes repair algorithm, and (D) is the patch repair method based on three-domain transformation.
[0062] Figure 6 is Figure 5 the enlarged view in the circle; where (A) is Figure 5 the enlarged view in the circle (A) in Figure 5 the enlarged view in the circle (B) in Figure 5 the enlarged view in the circle (C) in Figure 5 the enlarged view in the circle (D) in DETAILED DESCRIPTION OF THE INVENTION
[0063] For a clearer understanding of the technical features, objectives, and beneficial effects of the present invention, the following provides a detailed description of the technical solution of the present invention, but it should not be construed as a limitation on the scope of implementation of the present invention. The experimental methods described in the following embodiments are all conventional methods unless otherwise specified; the reagents and materials, unless otherwise specified, can be obtained from commercial sources.
[0064] The following describes the specific implementation of the present invention in detail in conjunction with specific embodiments.
[0065] As Figure 1 shown, a method for repairing film movie patches provided by an embodiment of the present invention includes the following steps:
[0066] Step 1: Obtain the film movie frame to be repaired, and use a film-to-digital conversion device to perform digital transcription on the film movie.
[0067] Step 2: Use the U-net network to segment the patch area in the picture to form a patch mask.
[0068] The structure of the U-net network is as Figure 2 shown. It will be Figure 2 described in two parts, the left part is the downsampling process of the network. Input an image with a size of 256×256, perform convolution feature extraction operations through 64 convolutional kernels with a size of 3×3, and obtain 64 feature layers of 256×256×1. Then, use convolutional kernels with a size of 3×3 to perform convolution operations on the 64 feature layers, and obtain 64 feature layers of 256×256×1. After the first layer of processing, a feature layer of 256×256×64 is finally obtained. Then, perform downsampling operations on the feature layer of 256×256×64 through a pooling layer with a size of 2×2, so that the size of the feature layer becomes half of the original, and then perform the same feature extraction operations on the obtained feature layer. Each time it is downsampled by one layer, the number of feature layers doubles and the size is halved, and so on. Finally, a feature layer with 512 layers and a size of 16×16 is obtained. The right part is the upsampling process of the network. Starting from the lower right corner, perform deconvolution operations on the feature layer of 16×16×1024 using 512 convolutional kernels with a size of 2×2 to expand the image to 32×32×512. Since deconvolution can only expand the feature layer but cannot fill in data. To reduce data loss, the method of skip connection is used to supplement data. After matrix splicing, the feature layer becomes 32×32×1024 in size, and then convolution is performed using convolutional kernels to fuse the feature layers. Each time it is upsampled by one layer, the size of the feature layer doubles and the number decreases by half. The final upsampling result is 256×256×64. In the last step, use convolutional kernels to change the 64 feature layers into 2, complete the binary classification operation of the image, and divide the entire picture into two categories: background and target.
[0069] The specific operations of step 2 are as follows:
[0070] Step 21: Input image preprocessing;
[0071] Adjust the input size of the U-net network to 256×256 pixels;
[0072] Image normalization: For segmented images of different sizes, set the scale_tensor parameter to normalize the size of the input image to make it suitable for network processing.
[0073] Step 22: Network structure configuration;
[0074] Depth setting: Set the downsampling and upsampling depth depth of the U-net network to 4;
[0075] Channel setting: Both the input channel (in_channels) and the output channel (out_channels) are set to 1, that is, the input is a single-channel grayscale image, and the output is a single-channel binary segmentation mask;
[0076] Feature channel number setting: The initial wf (width factor) is set to 6, that is, the initial number of feature layers = 2 6 = 64;
[0077] Convolution layer configuration: In each convolution block, set the number of convolution layers conv_num to 2 to increase the complexity of feature extraction and enable the network to learn more complex feature representations.
[0078] Upsampling method: In the upsampling process, use the transposed convolution method instead of bilinear interpolation;
[0079] Step 23: Network training;
[0080] Activation function: Select ReLU as the activation function;
[0081] Normalization: After each convolution layer, use the Batch Norm normalization method to normalize the data after convolution;
[0082] Select and define the loss function: Select BCEWithLogitsLoss as the loss function to solve the binary classification problem of plaque pixels. The loss function combines the Sigmoid function and BCELoss, and the expression is as follows:
[0083] ;
[0084] where x is the original output of the model; y is the target value (0 or 1); sigmoid is the sigmoid function;
[0085] Optimizer configuration: The Adam optimizer is used for network training, and the learning rate is set α to 0.001, and the exponential decay rate of the first moment estimate β 1 is 0.9, and the exponential decay rate of the second moment estimate β 2 is 0.999;
[0086] Step 24, plaque segmentation: Use the trained U-net network to segment the plaques in the input damaged image to obtain a binary mask for plaque segmentation.
[0087] Step 3, Use the plaque repair method based on three-domain transformation to repair and reconstruct the plaque damaged image, and the principle is as Figure 3 shown;
[0088] The specific operations of the said Step 3 are as follows:
[0089] Step 31, Define the domain;
[0090] Define the real damaged image (instance) domain R , the masked covered image (label) domain X and the intact undamaged image (feature) domain Y , where the intact undamaged image (feature) domain Y contains adjacent and similar frames of the damaged image, which are used to provide the image features required for repair;
[0091] Step 32, Use the VAE encoder E to encode the picture into the latent space. The VAE encoder E includes E R , E X and E Y ; Use the VAE generator G to restore the image encoding. The VAE generator G includes G R , G X and G Y ;
[0092] Encode the images in the real damaged image domain R through the encoder r R into the latent space E R to obtain the encoded representation Z r z r; Cover the picture area with a mask X the image in x through the encoder E X encode it into the latent space Z X , obtaining the encoded representation z x ; For the intact and undamaged picture area Y the image in y through the encoder E Y encode it into the latent space Z Y , obtaining the encoded representation z y ;
[0093] Inside the latent space Z R and Z X the intersection of the domains forms a shared domain, and the encoded representation of the damaged image in the shared domain is z x∩r ;
[0094] G R , G X and G Y respectively perform decoding operations on the encoded representations of the images to assist in the training of the encoders E R , E X and E Y , specifically: G R Decode the encoded representation z r to restore the image r ; G X Decode the encoded representation z x to restore the image x ; G Y Decode the encoded representation z y to restore the image y ;
[0095] Step 33, Feature mapping;
[0096] Independently train the damaged domain Z R and Z XShared domain and intact domain Z Y Mapping between T Z , the mapping T Z is used to convert the encoded information of the damaged domain to the intact domain Z Y ;
[0097] Step 34, use the VAE generator G to restore the information mapped by T Z mapped to Z Y the domain into an image: through the mapping T Z to z x∩r perform a fitting mapping to z y and then decode by G Y to achieve the restoration of the image;
[0098] Damaged image r The repair process is represented by the following formula:
[0099] ;
[0100] where, is the result of the transformation of the damaged image r from the R domain to the Y domain; is the VAE generator; is the feature mapping transformation; is the VAE encoder; " " represents the connection of each process.
[0101] VAE encoder E, generator G;
[0102] We first encode the damaged picture into the latent space. Assuming that the picture data in the latent space follows a Gaussian distribution, we use the reparameterization trick to sample the picture data in the latent space so that the image can be reconstructed in the latent space.
[0103] Among them, the loss function for encoding the images r, x, and y can be expressed by the following formula, taking the real damaged picture (instance) domain as an example:
[0104] ;
[0105] The loss function consists of three parts: the first part uses the KL divergence to measure the error between two probability distributions and constrains the z r, making it obey the Gaussian distribution. The second part is to constrain the image information restored by the generator so that the encoder can extract the main features in the image. The third part introduces the least squares loss LSGAN to solve the problem of over-smoothing that often occurs in the VAE encoder and improve the quality of the encoded image.
[0106] Among them, is the loss function of the real damaged picture domain R; the first term of the loss function is the encoder loss, and KL is the KL divergence; is the mathematical expectation of encoding under the condition of the sample image r, is a normal distribution with a mean of 0 and a variance of ; the second term of the loss function is the generator loss, is the weight coefficient of the generator loss, is the mathematical expectation, is the generation result of the image, is the image sample, and the third term of the loss function is , representing the LSGAN loss.
[0107] To expand the shared domain, a discriminator is introduced, and the loss function is as follows:
[0108] ;
[0109] Among them, is the LSGAN loss introduced in the latent space; is the mathematical expectation of the loss of the masked covered picture domain X; is the real damaged picture domain R the mathematical expectation of the loss; represents the real damaged picture domain R、 the masked covered picture domain X the discriminator loss; is the real damaged picture domain R、 the masked covered picture domain X the encoding of the image x within; is the real damaged picture domain R、 the masked covered picture domain X the encoding of the image r within;
[0110] The total objective function for encoding the images r and x is obtained as:
[0111] ;
[0112] Among them, represents the encoder loss of the real damaged picture domain R、 the masked covered picture domain X ; Represent the real damaged picture area R、 Mask-covered picture area X Generator loss; Represent the real damaged picture area R、 Mask-covered picture area X Discriminator loss; Is the loss function for the real damaged picture area R ; Is the loss function for the mask-covered picture area X ; Is the LSGAN loss introduced in the latent space.
[0113] (2) Mapping T Z ;
[0114] By learning the mapping z x∩r , z y} between the training-coded image pairs T Z , to perform image restoration and reconstruction; the mapping T Z 's loss function is shown as follows:
[0115] ;
[0116] Among them, z x∩r Is the damaged image coding representation in the shared domain; Is the weight coefficient of the L1 loss of the mapping T; Is the L1 loss of the mapping T, , constrained by the L1 norm; Is the mapping transformation of the coded latent space; E Is the mathematical expectation; Is the introduced least squares loss LSGAN; Is the weight coefficient of the feature matching loss; Is the introduced feature matching loss, Match the discriminator And the multi-level feature layers in the VGG network, and the expression is shown as follows:
[0117] ;
[0118] Among them, Represents the i-th layer feature layer of the discriminator ; Represents the i-th layer feature layer of the VGG network; Represents the activation times of the i-th layer feature layer of the discriminator ; Represents the number of activations of the i-th feature layer of the VGG network.
[0119] Step 4: Add the reconstructed image frame back to the movie clip to complete the restoration work.
[0120] Example 1: Comparison of mask generation effects;
[0121] like Figure 4 As shown in (A)-(E), traditional detection methods (SDI detection, ROD detection, edge detection) are easily affected by complex background environments and noise, and there are a lot of false detections in the generated patch mask images. Figure 4 As shown in (E), the patch mask generation technology based on U-net proposed in the present invention can effectively reduce false detection, accurately identify the patch area in the picture, and generate a patch mask that accurately segments the patch area.
[0122] like Figure 5 As shown in (A)-(D), the details are magnified to show Figure 6 From the results of (A)-(D), we can see that traditional repair methods (TELEA repair algorithm, Navier-Stokes repair algorithm) have good repair effects for small spots, but when repairing large patches, the repair effect is poor due to the lack of sufficient image information. The underlying principle of the algorithm is to start from the edge of the image and repair towards the center of the image. Therefore, the closer to the center of the patch, the worse the repair effect. In the repaired area, unnatural, disconnected, and non-smooth areas can be observed. Figure 6 As shown in (D), the plaque repair method based on three-domain conversion proposed in the present invention utilizes the intact picture information of adjacent frames to fit and map the damaged image to the intact image, which has a better image repair effect and makes the repaired image more natural and smooth.
[0123] We use three indicators, PSNR, SSIM, and MAE, to measure the image restoration effect. The indicators are shown in Table 1.
[0124] Table 1 Image restoration effects
[0125] Method PSNR SSIM MAE TELEA 28.7934 0.6403 0.0627 Navier-Stokes 29.7781 0.6714 0.0586 Patch Repair Method Based on Three-Domain Transformation 30.6683 0.6819 0.0558
[0126] It can be seen from Table 1 that the plaque repair method based on three-domain conversion proposed in the present invention has the highest PSNR peak signal-to-noise ratio, the highest SSIM structural similarity and the lowest MAE mean absolute error, and the overall repair performance is the best.
[0127] The above are only the preferred embodiments of the present invention. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several modifications and improvements can be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicability of the patent.
Claims
1. A method for repairing film plaques, characterized in that: The following steps are involved: Step 1: Obtain the film frame to be restored, and digitally transcribe the film using telecine equipment; Step 2: Use the U-net network to segment the patch area in the image to form a patch mask; Step 3: Use a plaque repair method based on three-domain conversion to repair and reconstruct the plaque damage image; Step 4: Add the reconstructed image frame back to the movie clip to complete the restoration work; The specific operations of step 3 are as follows: Step 31, define the domain; Defining the Real Damaged Image Domain R , mask covers the screen area X and intact image area Y , where the intact and undamaged image area Y It contains adjacent and close frames of the damaged picture, which are used to provide image features required for restoration; Step 32: Using VAE encoder E Encode the image into the latent space, the VAE encoder E include E R , E X and E Y ; Using VAE generator G The image encoding is restored, and the VAE generator G include G R , G X and G Y ; The real damaged image domain R Images in r Through the encoder E R Encoding into latent space Z R , and get the encoded representation z r ; Cover the screen area with mask X Images in x Through the encoder E X Encoding into latent space Z X , and get the encoded representation z x ; The intact and undamaged image area Y Images in y Through the encoder E Y Encoding into latent space Z Y , and get the encoded representation z y ; In hidden space Z R and Z X The intersection of the domains forms a shared domain, and the damaged image encoding in the shared domain is represented as z x∩r ; G R , G X and G Y Decode the encoded representation of the image to assist the encoder E R , E X and E Y The training is as follows: G R The encoded representation z r Decode and restore the image r ; G X The encoded representation z x Decode and restore the image x ; G Y The encoded representation z y Decode and restore the image y ; step 33. Feature mapping; Independent training of the damaged domain Z R and Z X Shared domains and intact domains Z Y The mapping between T Z , mapping T Z Used to convert the encoded information of the shared domain to the intact domain Z Y ; Step 34: Using the VAE generator G Y Will be mapped T Z Map to Z Y The information of the domain is restored to an image: through mapping T Z Will z x∩r Towards z y Perform fitting mapping, and then G Y Decoding is performed to restore the image; Damaged image r The repair process is expressed as follows: ; in, is the result of transforming the damaged image r from the R domain to the Y domain; is the VAE generator; is the feature mapping transformation; is the VAE encoder; " indicates the connection of each process.
2. The method for repairing film spots according to claim 1, characterized in that: The specific operation of step 2 is as follows: Step 21: preprocessing of input image; Image normalization: Set the scale_tensor parameter to normalize the size of the input image to 256×256 pixels; Step 22: Network structure configuration; Depth setting: Set the downsampling and upsampling depth of the U-net network to 4; Channel setting: both input channel and output channel are set to 1; Setting the number of feature channels: Set the initial number of feature layers of the U-net network to 64; Convolution layer configuration: In each convolution block, set the number of convolution layers to 2; Upsampling method: In the upsampling process, the deconvolution method is used; Step 23: Network training; Activation function: Select ReLU as the activation function; Normalization: After each convolution layer, the Batch Norm normalization method is used to normalize the convolutional data; Select and define the loss function: Select BCEWithLogitsLoss as the loss function to solve the binary classification problem of pixels. The loss function combines the Sigmoid function and BCELoss, and the expression is as follows: ; Among them, x is the original output of the model; y is the target value; sigmoid is the sigmoid function; Optimizer configuration: Use Adam optimizer for network training and set learning rate α is 0.001, the exponential decay rate of the first-order moment estimate β 1 is 0.9, the exponential decay rate of the second-order moment estimate β 2 is 0.999; Step 24: Plaque segmentation: Use the trained U-net network to segment the plaques in the input damaged image to obtain a binary mask for plaque segmentation.
3. The method for repairing film spots according to claim 1, characterized in that: In the VAE encoder, assuming that the image data in the latent space obeys Gaussian distribution, the real damaged picture domain R The VAE loss function is expressed as: ; in, For the real damaged image domain R VAE loss function; the first term of the loss function is the encoder loss, KL is the KL divergence, Under the condition of sample image r The mathematical expectation of The mean is 0 and the variance is I The second term of the loss function is the generator loss, is the weight coefficient of the generator loss, is the mathematical expectation, is the generated result of the image, is an image sample; the third term of the loss function is , represents the LSGAN loss.
4. The method for repairing film spots according to claim 1, characterized in that: The discriminator is introduced to expand the shared domain, and the overall objective function for encoding the image r, x is obtained as: ; in, Represents the real damaged image domain R、 Mask covers the screen area X The encoder loss of Represents the real damaged image domain R、 Mask covers the screen area X The generator loss of For the real damaged image domain R、 Mask covers the screen area X The discriminator loss of For the real damaged image domain R VAE loss function; Mask covers the screen area X VAE loss function; is the LSGAN loss introduced in the latent space.
5. The method for repairing film spots according to claim 1, characterized in that: By learning to train the encoded image pairs { z x∩r , z y } T Z , to repair and reconstruct the image; mapping T Z The loss function is as follows: ; in, z x∩r Encode representations for corrupted images in a shared domain; is the weight coefficient of the L1 loss of mapping T; is the L1 loss of mapping T, , using L1 norm for constraint; is the mapping transformation of the encoding latent space, E is the mathematical expectation; The least squares loss LSGAN is introduced; is the weight coefficient of feature matching loss; is the introduced feature matching loss, Matching Discriminator And the multi-level feature layer in the VGG network, the expression is as follows: ; in, Representation Discriminator The i-th feature layer; Represents the i-th feature layer of the VGG network; Representation Discriminator The number of activations of the i-th feature layer; Represents the number of activations of the i-th feature layer of the VGG network.