Transformer Infrared Image Style Transfer Method and System Based on Cyclic Adversarial Network
By applying the GCN-CycleGAN network structure to perform style migration in transformer infrared image detection, the problem of low image detection accuracy under different lighting conditions is solved, and more efficient image style conversion and detection accuracy is achieved.
Patent Information
- Application Number
- CN202410279401.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-12
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-03-12
AI Technical Summary
Under different lighting conditions, the target object in the infrared image of the transformer is difficult to be accurately detected and analyzed, resulting in a decrease in the accuracy of fault recognition.
The GCN-CycleGAN network structure based on a cyclic adversarial network is used to transfer the style of the transformer infrared image. Through the adversarial loss, cyclic consistency loss and Identity loss functions of the generator and discriminator, the image style is transformed, so that the image can generate samples under the same lighting conditions.
It effectively improves the detection accuracy of transformer infrared images under different lighting conditions, enhances the capture and information transmission of image structure information, is suitable for large-scale infrared image data processing, and maintains good generalization ability when data is sparse or noise is high.
Smart Images

Figure CN118154407B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence vision detection, and particularly relates to a method and system for style transfer of transformer infrared images based on a cyclic adversarial network. Background Art
[0002] Infrared image diagnosis is to judge the working state and existing faults of a transformer through the temperature distribution. How to efficiently and accurately identify infrared images is the basis for ensuring the safe and reliable operation of the transformer. Infrared identification technology has the advantages of non-contact, fast response speed, wide measurement range, intuitive and vivid measurement results, etc., and its application in the operation and maintenance of power systems is becoming increasingly important. By real-time monitoring and analyzing the thermal map data, evaluating the load balance of the transformer, and analyzing the high-temperature areas in the thermal map, the fault points can be accurately located, which is not only more efficient but also greatly protects the personal safety of operators. Due to the strong sunlight during actual operation and maintenance, a large amount of infrared light will be generated, and the shielding effect causes the background light source to be stronger than other heat sources. The continuous light causes the environmental temperature to rise, making it difficult to accurately detect and analyze the target objects in the infrared images. To realize the detection of transformer infrared images under different lighting conditions, a style transfer method is used to generate samples of transformer infrared images under the same lighting to improve the accuracy of transformer fault identification. Summary of the Invention
[0003] The present invention aims to solve the deficiencies of the prior art, and the present invention provides the following solutions: A method for style transfer of transformer infrared images based on a cyclic adversarial network, comprising the following steps:
[0004] S1. Collect sample image data, preprocess the sample image data to obtain fourth image data;
[0005] S2. Screen the fourth image data to obtain fifth image data;
[0006] S3. Perform style transfer on the fifth image data to obtain engineering images.
[0007] Further preferably, S1 includes:
[0008] S11. Collect the sample image data, perform format conversion on the sample image data to obtain a first RGB image;
[0009] S12. Use the resize function based on linear interpolation to adjust the size of the first RGB image to obtain a second RGB image;
[0010] S13. Add a re-batch-size dimension to the second RGB image to obtain a third RGB image;
[0011] S14. Normalize the third RGB image to obtain the fourth image data.
[0012] Further preferably, S2 includes:
[0013] S21. Perform a visualization operation on the fourth image data to obtain a first topological image containing the RGB values of the image;
[0014] S22. Extract features from the first topological image to obtain a first feature image;
[0015] S23. Perform feature matching on the first feature image to obtain a feature matching image;
[0016] S24. Screen and measure the temperature of the feature matching image according to the RGB values to obtain the fifth image data.
[0017] Further preferably, perform style transfer on the fifth image data using a style transfer network;
[0018] The style transfer network includes: a first generator, a second generator, a first discriminator, and a second discriminator;
[0019] The method for obtaining the engineering image includes:
[0020] Input the fifth image data into the first generator to obtain a first generated image;
[0021] Input the first generated image into the first discriminator. The first discriminator is used to determine whether the style of the first generated image conforms to the style of the second image set and input it into the second generator;
[0022] The second generator obtains a second generated image based on the first generated image;
[0023] Input the second generated image into the second discriminator to determine whether it conforms to the style of the fifth image data;
[0024] Finally, output the engineering image.
[0025] Further preferably, the loss function of the style transfer network includes:
[0026] L total (G, F, D A , D B ) = L GAN (G, D G , A, B) + L GAN (F, D F , B, A) + λ 1 L cyc(G, F) + λ 2 L Idenfity (G, F),
[0027] Wherein, L total (G, F, D A , D B ) represents the total loss function of the style transfer network; L GAN (G, D G , A, B) represents the adversarial loss of the first discriminator; L GAN (F, D F , B, A) represents the adversarial loss of the second discriminator; λ 1 、λ 2 respectively represent the CycleConsistency loss eigenvalue and the Identity loss eigenvalue; L cyc (G, F) represents the Cycle Consistency loss; L Identity (G, F) represents the Identity loss function.
[0028] Further preferably, the adversarial loss of the first discriminator includes:
[0029]
[0030] Wherein, represents the optimization loss function of the atlas B generator; D B (B) represents the optimization target of the atlas B; represents the optimization loss function of the atlas A generator; D G (G(A)) represents the optimization target of the generator atlas A.
[0031] Further preferably, the adversarial loss of the second discriminator includes:
[0032]
[0033] Wherein, represents the optimization loss function of the set A generator; D A (A) represents the simulation target of the atlas A; D A (F(B)) represents the simulation target of the discriminator atlas B.
[0034] Further preferably, the Identity loss includes:
[0035]
[0036] Wherein, It represents the Identity loss function; F(A) represents the discriminator's simulation target for image set A; A represents image set A; B represents image set B; G(B) represents the generator's optimization target for image set B.
[0037] The present invention also provides a transformer infrared image style transfer system based on a cyclic adversarial network. The system is used to implement the above method and includes: an acquisition system, a screening system, and a transfer system.
[0038] The acquisition system is used to collect sample image data, preprocess the sample image data, and obtain fourth image data.
[0039] The screening system is connected to the acquisition system and is used to screen the fourth image data to obtain fifth image data.
[0040] The transfer system is connected to the screening system and is used to perform style transfer on the fifth image data to obtain engineering images.
[0041] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0042] The present invention uses a GCN-CycleGAN network structure for style transfer. Compared with the traditional CycleGAN network structure, the GCN network structure is used to replace the original convolutional layer. Such an improvement helps to better capture the image structure information, enhance the information transmission and aggregation between each information point, and effectively adapt to the application situation of the infrared image set. Secondly, the GCN network has good scalability and can process large-scale graph data. By utilizing the local properties of graph convolution operations, GCN can effectively reduce the computational and storage complexity while maintaining the model performance. In addition, the graph convolutional network also has good generalization ability and can effectively learn and infer the features of graph data in the case of sparse data or more noise. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0044] Figure 1 It is a schematic flowchart of the transformer infrared image style transfer method based on a cyclic adversarial network according to an embodiment of the present invention.
[0045] Figure 2 It is a schematic flowchart of data preprocessing according to an embodiment of the present invention;
[0046] Figure 3Schematic diagram of the data screening process in the embodiment of the present invention;
[0047] Figure 4 Schematic diagram of the GCN-CycleGAN neural network structure in the embodiment of the present invention;
[0048] Figure 5 Schematic diagram of the mechanism of the generator in the embodiment of the present invention;
[0049] Figure 6 Schematic diagram of the mechanism of the discriminator in the embodiment of the present invention;
[0050] Figure 7 Schematic diagram of the GCN convolutional network structure in the embodiment of the present invention. Detailed implementation manners
[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0052] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0053] Embodiment 1:
[0054] As Figure 1 shown, this embodiment provides a method for style transfer of transformer infrared images based on a cyclic adversarial network, including the following steps:
[0055] S1. Collect sample image data, preprocess the sample image data to obtain fourth image data. The preprocessing process is as Figure 2 shown.
[0056] Specifically, S1 includes:
[0057] S11. Collect sample image data, perform format conversion on the sample image data to obtain a first RGB image.
[0058] Collect the transformer infrared sample image data before style transfer to form a sample image data set. And classify the transformer infrared sample images according to the lighting conditions to obtain the transformer infrared image training sets under different lighting conditions and the transformer infrared image training sets under the same lighting condition. Perform format conversion on the sample image data to obtain an image with a unified format, that is, the first RGB image.
[0059] S12. Use the resize function based on linear interpolation to adjust the size of the first RGB image to obtain the second RGB image.
[0060] S13. Add the re - batch - size dimension to the second RGB image through batch - size to obtain the third RGB image. Batch - size can reduce the training loss, lower the minimum validation loss, and the time required for training in each epoch.
[0061] S14. Perform data normalization on the third RGB image to obtain the fourth image data. Data normalization can effectively improve the model accuracy. After normalization, the features between different dimensions are numerically comparable, which can greatly improve the accuracy of the classifier. It can effectively improve the model convergence. After standardization, the optimization process becomes smoother and is more likely to correctly converge to the optimal solution.
[0062] S2. Perform data screening on the fourth image data to obtain the fifth image data.
[0063] As Figure 3 shown, specifically, S2 includes:
[0064] S21. Perform visualization operations on the fourth image data to obtain the first topological image containing the RGB values of the image pixel points.
[0065] S22. Extract the feature points of the first topological image based on the SIFT algorithm to obtain the first feature image.
[0066] S23. Perform feature matching based on the RGB of the feature points extracted from the first feature image to obtain the feature matching image.
[0067] S24. Screen and measure the feature matching image according to the RGB values, screen out the image set within a certain temperature range, and integrate the screened image set to obtain the fifth image data. In this embodiment, the image set with temperatures between 40°C and 70°C is retained.
[0068] S3. Perform style transfer on the fifth image data to obtain the engineering image. The engineering image is the infrared image of the transformer after style transfer.
[0069] In this embodiment, a style transfer network is used to perform style transfer on the fifth image data; the style transfer network is mainly composed of the GCN - CycleGAN network structure; As Figure 4As shown. The CycleGAN (Cycle Generative Adversarial Network) does not need to establish a one-to-one mapping between the training data in the source domain and the target domain. Instead, it directly utilizes cycle consistency to achieve the style conversion between different images while maintaining content similarity. The CycleGAN realizes the inter-domain image conversion by optimizing the loss functions of the generator and the discriminator, enabling the generator and the discriminator to cooperate with each other and gradually improving the quality and consistency of the generated images. Compared with the traditional CycleGAN network structure, the GCN-CycleGAN network structure uses the GCN network structure to replace the original convolutional layer and transposed convolutional layer. The GCN network structure is as Figure 7 shown. The style transfer network includes: the first generator (generator G), the second generator (generator F), the first discriminator (discriminator DG), and the second discriminator (discriminator DF). In addition, a loss function for cycle consistency is added. Figure 5 , Figure 6 are respectively the schematic diagrams of the generator and discriminator mechanisms. The first generator and the second generator each include: the first decoder, the first encoder, the second decoder, and the second encoder. It also includes a style converter, which adopts the residual neural network structure and includes the first image enhancement module, the downsampling module, the residual module, the upsampling module, and the second image enhancement module. The discriminators all adopt the idea of PatchGAN. After the input feature image passes through a series of convolutional layers, it enters a convolutional layer with 1 channel number, mapping the feature map to an N*N matrix. Each point in this N*N matrix represents the evaluation value of a small area in the original image. The second discriminator includes an input convolutional layer, three CIL modules, and an output convolutional module.
[0070] The method for obtaining the engineering image includes:
[0071] The fifth image data is input into the first generator, and the first enhanced image is obtained through the first image enhancement module. The first enhanced image is input into the first CIL module, and the first convolutional image is obtained through the first GCN convolutional module. After that, the first convolutional image passes through the first regularization layer to obtain the first regularized image. The first regularized image is input into the first LeakyRelu layer to obtain the first activated image. The first activated image passes through the second CIL module and the second GCN convolutional module to obtain the second convolutional image. The second convolutional image passes through the second regularization layer to obtain the second regularized image. The second regularized image is input into the second Leaky Relu layer to obtain the second activated image. The second activated image is input into the third CIL module and the third GCN convolutional module to obtain the third convolutional image. The third convolutional image passes through the third regularization layer to obtain the third regularized image. The third regularized image is input into the third Leaky Relu layer to obtain the third activated image. Among them, the size of the first CIL module is 256×256×64, the size of the second CIL module is 128×128×128, and the size of the third CIL module is 64×64×256. After that, the third activated image passes through the first residual structure to perform restoration and enhancement on it to obtain the first restored and enhanced image. The first residual structure includes 9 repeated residual modules. Then, the first restored and enhanced image is input into the first encoder, and the first deconvolutional image is obtained through the first CTIR module and the first GCN deconvolutional module. The first deconvolutional image is input into the first IN normalization layer to obtain the first normalized image. The first normalized image is input into the first RELU layer to perform size restoration processing on it to obtain the fourth activated image. The fourth activated image is input into the second CTIR module and the second GCN deconvolutional module to obtain the second deconvolutional image. The second deconvolutional image is input into the second IN normalization layer to obtain the second normalized image. The second normalized image is input into the second RELU layer to perform size restoration processing on it to obtain the fifth activated image. Among them, the size of the first CTIR module is 128×128×128, and the size of the second CTIR module is 256×256×64. After that, the fifth activated image passes through the second enhancement module to obtain the second enhanced image. The second enhanced image passes through the fourth GCN convolutional module to obtain the first generated image (atlas A).
[0072] The first generated image is input into the first discriminator for the concept of global comparison, considering the difference in global receptive field information, to determine whether the generated image conforms to the style of the second image set (atlas B), and is input into the second generator.
[0073] For the mapping function G, the adversarial loss between A→B and the first discriminator (DG) is expressed as follows:
[0074]
[0075] In the formula, Denote the optimization loss function of the atlas B generator; D B (B) Denote the optimization objective of the atlas B; Denote the optimization loss function of the atlas A generator; D G (G(A)) Denote the optimization objective of the generator atlas A.
[0076] Input the first generated image into the second generator, and obtain the third enhanced image through the third image enhancement module. Input the third enhanced image into the fourth CIL module, and obtain the fifth convolutional image through the fifth GCN convolutional module. The fifth convolutional image passes through the fourth regularization layer to obtain the fourth regularized image. Input the fourth regularized image into the fourth Leaky Relu layer to obtain the sixth activated image. Input the sixth activated image into the fifth CIL module, and obtain the sixth convolutional image through the sixth GCN convolutional module. The sixth convolutional image passes through the fifth regularization layer to obtain the fifth regularized image. Input the fifth regularized image into the fifth LeakyRelu layer to obtain the seventh activated image. Input the seventh activated image into the sixth CIL module, and obtain the seventh convolutional image through the seventh GCN convolutional module. The seventh convolutional image passes through the sixth regularization layer to obtain the sixth regularized image. Input the sixth regularized image into the sixth Leaky Relu layer to obtain the eighth activated image. Among them, the size of the fourth CIL module is 256×256×64, the size of the fifth CIL module is 128×128×128, and the size of the sixth CIL module is 64×64×256. The eighth activated image passes through the second residual structure for restoration enhancement to obtain the second restored and enhanced image. The second residual structure is 9 repeated residual modules. Input the second restored and enhanced image into the second encoder, and obtain the third deconvolutional image through the third CTIR module and the third GCN deconvolutional module. Input the third deconvolutional image into the third IN normalization layer to obtain the third normalized image. Input the third normalized image into the third RELU layer to perform size restoration processing on the image to obtain the ninth activated image. Input the ninth activated image into the fourth CTIR module and the fourth GCN deconvolutional module to obtain the fourth deconvolutional image. Input the fourth deconvolutional image into the fourth IN normalization layer to obtain the fourth normalized image. Input the fourth normalized image into the fourth RELU layer to perform size restoration processing on it to obtain the tenth activated image. Among them, the size of the third CTIR module is 128×128×128, and the size of the fourth CTIR module is 256×256×64. The tenth activated image passes through the fourth image enhancement module to obtain the fourth enhanced image. The fourth enhanced image passes through the eighth GCN convolutional module to obtain the second generated image.
[0077] The concept of inputting the second generated image into the second discriminator for global comparison, considering the differences in global receptive field information, and determining whether it conforms to the style of the first image set (image set A). Image set A and image set B are image sets obtained by grouping sample images according to image style, and the image styles in each image set are consistent.
[0078] The adversarial loss for the mapping function F: B → A and the second discriminator (DF) is expressed as follows:
[0079]
[0080] In the formula, represents the optimization loss function of the generator of set A; D A (A) represents the simulation target of image set A; D A (F(B)) represents the simulation target of the discriminator image set B.
[0081] In addition to the above two adversarial loss functions, in CycleGAN, there is also an Identity loss. This loss means that when the images of the first image set (A) are fed into the generator G, it should still be itself and no other transformation is done. The role of the Identity loss is mainly to restrict the generator G from autonomously modifying the color of the input image. The Identity loss is expressed as follows:
[0082]
[0083] In the formula, represents the Identity loss function; F(A) represents the simulation target of the discriminator image set A; A represents image set A; B represents image set B; G(B) represents the optimization target of the generator image set B.
[0084] After reaching the number of cycles, the first discriminator and the second discriminator respectively output a sequence of N×N matrices; In summary, the GCN-CycleGAN loss function is obtained by summing the adversarial loss, the cycle consistency loss, and the Identity loss, and the total loss function is expressed as follows:
[0085] L total (G, F, D A , D B ) = L GAN (G, D G , A, B) + L GAN (F, D F , B, A) + λ 1 L cyc (G, F) + λ 2 L Idenfity (G, F),
[0086] In the formula, Ltotal (G, F, D A , D B ) represents the total loss function of the style transfer network; L GAN (G, D G , A, B) represents the adversarial loss of the first discriminator; L GAN (F, D F , B, A) represents the adversarial loss of the second discriminator; λ 1 、λ 2 respectively represent the Cycle Consistency loss eigenvalue and the Identity loss eigenvalue; L cyc (G, F) represents the Cycle Consistency loss; L Identity (G, F) represents the Identity loss.
[0087] Through the above operations, the style transfer process of the sample image set is realized, and the image set after style transfer, that is, the engineering sample set, is obtained; finally, the infrared image of the transformer, that is, the engineering image, is output.
[0088] Embodiment 2:
[0089] This embodiment provides a transformer infrared image style transfer system based on a cyclic adversarial network, including: an acquisition system, a screening system, and a transfer system.
[0090] The acquisition system is used to acquire sample image data, preprocess the sample image data, and obtain the fourth image data.
[0091] Specifically, the acquisition system includes: an image format conversion module, a resize function size adjustment module, a batch-size dimension addition module, and a digital normalization module.
[0092] The image format conversion module is used to acquire sample image data, convert the format of the sample image data, and obtain the first RGB image.
[0093] Acquire the infrared sample image data of the transformer before style transfer to form a sample image data set. And classify the infrared sample images of the transformer according to the illumination conditions to obtain the infrared image training set of the transformer under different illuminations and the infrared image training set of the transformer under the same illumination. Convert the format of the sample image data to obtain an image with a unified format, that is, the first RGB image.
[0094] The resize function size adjustment module is used to adjust the size of the first RGB image by using the resize function based on linear interpolation to obtain the second RGB image.
[0095] The batch-size dimension adding module is used to add a re-batch-size dimension to the second RGB image to obtain a third RGB image. The batch-size can reduce the training loss, lower the minimum validation loss, and the time required for training in each epoch.
[0096] The digital normalization module is used to perform data normalization processing on the third RGB image to obtain a fourth image data. The data normalization processing can effectively improve the model accuracy. After normalization, the features between different dimensions are numerically comparable, which can greatly improve the accuracy of the classifier. It can effectively improve the model convergence. After standardization, the optimization process becomes smoother and it is easier to correctly converge to the optimal solution.
[0097] The screening system is connected to the acquisition system and is used to screen the fourth image data to obtain a fifth image data. Specifically, the screening system includes: a data visualization module, a feature extraction module, a feature matching module, and a temperature screening module.
[0098] The data visualization module is used to perform a visualization operation on the fourth image data to obtain a first topological image containing the RGB values of the image pixels.
[0099] The feature extraction module extracts feature points of the first topological image based on the SIFT algorithm to obtain a first feature image.
[0100] The feature matching module is used to perform feature matching based on the RGB of the feature points extracted from the first feature image to obtain a feature matching image.
[0101] The temperature screening module is used to perform temperature screening and measurement on the feature matching image according to the RGB values, screen out the image set within a certain temperature range, and integrate the screened image set to obtain a fifth image data. In this embodiment, the image set with a temperature between 40°C and 70°C is retained.
[0102] The migration system is connected to the screening system and is used to perform style migration on the fifth image data to obtain an engineering image. The engineering image is the infrared image of the transformer after style migration.
[0103] In this embodiment, the migration system uses a style migration network to perform style migration on the fifth image data; the style migration network is mainly composed of a GCN-CycleGAN network structure. Compared with the traditional CycleGAN network structure, the GCN-CycleGAN network structure uses a GCN network structure to replace the original convolutional layer and transposed convolutional layer. The GCN network structure is as Figure 7As shown in the figure. The style transfer network includes: a first generator (generator G), a second generator (generator F), a first discriminator (discriminator DG), a second discriminator (discriminator DF), and in addition, a loss function for cyclic consistency is added. The first generator and the second generator each include: a first decoder, a first encoder, a second decoder, and a second encoder. It also includes a style converter, and the style converter adopts a residual neural network structure including a first image enhancement module, a downsampling module, a residual module, an upsampling module, and a second image enhancement module. The discriminators all adopt the idea of PatchGAN. After the input feature image passes through a series of convolutional layers, it enters a convolutional layer with 1 channel number, and the feature map is mapped into an N*N matrix. Each point in this N*N matrix represents the evaluation value of a small area in the original image. The second discriminator includes an input convolutional layer, three CIL modules, and an output convolutional module.
[0104] The method for the migration system to obtain the engineering image includes:
[0105] The fifth image data is input into the first generator, and the first enhanced image is obtained through the first image enhancement module. The first enhanced image is input into the first CIL module, and the first convolutional image is obtained through the first GCN convolutional module. After that, the first convolutional image passes through the first regularization layer to obtain the first regularized image. The first regularized image is input into the first LeakyRelu layer to obtain the first activated image. The first activated image passes through the second CIL module and the second GCN convolutional module to obtain the second convolutional image. The second convolutional image passes through the second regularization layer to obtain the second regularized image. The second regularized image is input into the second Leaky Relu layer to obtain the second activated image. The second activated image is input into the third CIL module and the third GCN convolutional module to obtain the third convolutional image. The third convolutional image passes through the third regularization layer to obtain the third regularized image. The third regularized image is input into the third Leaky Relu layer to obtain the third activated image. Among them, the size of the first CIL module is 256×256×64, the size of the second CIL module is 128×128×128, and the size of the third CIL module is 64×64×256. After that, the third activated image passes through the first residual structure to be restored and enhanced to obtain the first restored and enhanced image. The first residual structure includes 9 repeated residual modules. Then the first restored and enhanced image is input into the first encoder, and the first deconvolutional image is obtained through the first CTIR module and the first GCN deconvolutional module. The first deconvolutional image is input into the first IN normalization layer to obtain the first normalized image. The first normalized image is input into the first RELU layer to perform size restoration processing to obtain the fourth activated image. The fourth activated image is input into the second CTIR module and the second GCN deconvolutional module to obtain the second deconvolutional image. The second deconvolutional image is input into the second IN normalization layer to obtain the second normalized image. The second normalized image is input into the second RELU layer to perform size restoration processing to obtain the fifth activated image. Among them, the size of the first CTIR module is 128×128×128, and the size of the second CTIR module is 256×256×64. After that, the fifth activated image passes through the second enhancement module to obtain the second enhanced image. The second enhanced image passes through the fourth GCN convolutional module to obtain the first generated image (Atlas A).
[0106] The first generated image is input into the first discriminator for the concept of global comparison, considering the difference in global receptive field information, determining whether the generated image conforms to the style of the second image set (Atlas B), and input into the second generator.
[0107] For the mapping function G, the adversarial loss between A→B and the first discriminator (DG) is expressed as follows:
[0108]
[0109] In the formula, Denote the optimization loss function of the atlas B generator; D B (B) denotes the optimization objective of the atlas B; Denote the optimization loss function of the atlas A generator; D G (G(A)) denotes the optimization objective of the generator atlas A.
[0110] Input the first generated image into the second generator, and obtain the third enhanced image through the third image enhancement module. Input the third enhanced image into the fourth CIL module, and obtain the fifth convolutional image through the fifth GCN convolutional module. The fifth convolutional image passes through the fourth regularization layer to obtain the fourth regularized image. Input the fourth regularized image into the fourth Leaky Relu layer to obtain the sixth activated image. Input the sixth activated image into the fifth CIL module, and obtain the sixth convolutional image through the sixth GCN convolutional module. The sixth convolutional image passes through the fifth regularization layer to obtain the fifth regularized image. Input the fifth regularized image into the fifth LeakyRelu layer to obtain the seventh activated image. Input the seventh activated image into the sixth CIL module, and obtain the seventh convolutional image through the seventh GCN convolutional module. The seventh convolutional image passes through the sixth regularization layer to obtain the sixth regularized image. Input the sixth regularized image into the sixth Leaky Relu layer to obtain the eighth activated image. Among them, the size of the fourth CIL module is 256×256×64, the size of the fifth CIL module is 128×128×128, and the size of the sixth CIL module is 64×64×256. The eighth activated image passes through the second residual structure for restoration and enhancement to obtain the second restored and enhanced image. The second residual structure is 9 repeated residual modules. Input the second restored and enhanced image into the second encoder, and obtain the third deconvolutional image through the third CTIR module and the third GCN deconvolutional module. Input the third deconvolutional image into the third IN normalization layer to obtain the third normalized image. Input the third normalized image into the third RELU layer to perform size restoration processing on the image to obtain the ninth activated image. Input the ninth activated image into the fourth CTIR module and the fourth GCN deconvolutional module to obtain the fourth deconvolutional image. Input the fourth deconvolutional image into the fourth IN normalization layer to obtain the fourth normalized image. Input the fourth normalized image into the fourth RELU layer to perform size restoration processing on it to obtain the tenth activated image. Among them, the size of the third CTIR module is 128×128×128, and the size of the fourth CTIR module is 256×256×64. The tenth activated image passes through the fourth image enhancement module to obtain the fourth enhanced image. The fourth enhanced image passes through the eighth GCN convolutional module to obtain the second generated image.
[0111] Input the second generated image into the second discriminator, perform the concept of global comparison, consider the difference in global receptive field information, and judge whether it conforms to the style of the first image set (image set A).
[0112] The adversarial loss for the mapping function F: B → A and the second discriminator (DF) is expressed as follows:
[0113]
[0114] where represents the optimization loss function of the generator of set A; D A (A) represents the simulation target of image set A; D A (F(B)) represents the simulation target of the discriminator for image set B.
[0115] In addition to the above two adversarial loss functions, in CycleGAN, there is also an Identity loss. This loss means that when the images of the first image set (A) are fed into the generator G, the output should be the same images without any other transformations. The main role of the Identity loss is to restrict the generator G from autonomously modifying the colors of the input images. The Identity loss is expressed as follows:
[0116]
[0117] where represents the Identity loss function; F(A) represents the simulation target of the discriminator for image set A; A represents image set A; B represents image set B; G(B) represents the optimization target of the generator for image set B.
[0118] After reaching the number of cycles, the first discriminator and the second discriminator respectively output a sequence of N×N matrices; In summary, the GCN - CycleGAN loss function is obtained by summing the adversarial loss, the cycle consistency loss, and the Identity loss. The total loss function is expressed as follows:
[0119] L total (G, F, D A , D B ) = L GAN (G, D G , A, B) + L GAN (F, D F , B, A) + λ 1 L cyc (G, F) + λ 2 L Identity (G, F),
[0120] where L total (G, F, D A , D B ) represents the total loss function of the style transfer network; L GAN (G, D G , A, B) represents the adversarial loss of the first discriminator; L GAN (F, DF , B, A) represents the adversarial loss of the second discriminator; λ 1 , λ 2 respectively represent the CycleConsistency loss eigenvalue and the Identity loss eigenvalue; L cyc (G, F) represents the Cycle Consistency loss; L Identity (G, F) represents the Identity loss.
[0121] Through the above operations, the style transfer processing of the sample image set is realized, and the image set after style transfer, that is, the engineering sample set, is obtained; finally, the infrared image of the transformer, that is, the engineering image, is output.
[0122] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A transformer infrared image style transfer method based on a recurrent adversarial network, characterized in that: The following steps are involved: S1, collecting sample image data, and preprocessing the sample image data to obtain fourth image data; S2, performing data screening on the fourth image data to obtain fifth image data; S2 includes: S21, performing a visualization operation on the fourth image data to obtain a first topological image containing image RGB values; S22, extracting features from the first topological image to obtain a first feature image; S23, performing feature matching on the first feature image to obtain a feature matching image; S24, screening and measuring the temperature of the feature matching image according to the RGB value to obtain the fifth image data; S3. Perform style migration on the fifth image data to obtain an engineering image.
2. According to claim 1, the transformer infrared image style transfer method based on recurrent adversarial network is characterized in that: S1 includes: S11, collecting the sample image data, and performing format conversion on the sample image data to obtain a first RGB image; S12, resizing the first RGB image using a resize function based on linear interpolation to obtain a second RGB image; S13, adding a rebatch-size dimension to the second RGB image to obtain a third RGB image; S14. Normalize the third RGB image to obtain the fourth image data.
3. According to claim 1, the transformer infrared image style transfer method based on recurrent adversarial network is characterized in that: Using a style transfer network to perform style transfer on the fifth image data; The style transfer network includes: a first generator, a second generator, a first discriminator, and a second discriminator; The method for obtaining the engineering image includes: Inputting the fifth image data into the first generator to obtain a first generated image; Inputting the first generated image into the first discriminator, the first discriminator is used to determine whether the style of the first generated image conforms to the style of the second image set, and inputting the first generated image into the second generator; The second generator obtains a second generated image based on the first generated image; Inputting the second generated image into a second discriminator to determine whether it conforms to the style of the fifth image data; Finally, the engineering image is output.
4. According to claim 3, the transformer infrared image style transfer method based on recurrent adversarial network is characterized in that: The loss function of the style transfer network includes: L total (G,F,D A ,D B )=L GAN (G,D G ,A,B)+L GAN (F,D F ,B,A)+λ1L cyc (G,F)+λ2L Identity (G,F), Where, L total (G,F,D A ,D B ) represents the total loss function of the style transfer network; L GAN (G,D G ,A,B) represents the adversarial loss of the first discriminator; L GAN (F,D F ,B,A) represents the adversarial loss of the second discriminator; λ1 and λ2 represent the CycleConsistency loss eigenvalue and Identity loss eigenvalue respectively; L cyc (G, F) indicates cycle consistency loss; L Identity (G, F) indicates Identity loss.
5. According to claim 4, the transformer infrared image style transfer method based on recurrent adversarial network is characterized in that: The adversarial loss of the first discriminator includes: In the formula, Denotes the loss function optimized by the atlas B generator; D B (B) represents the optimization target of atlas B; Denotes the optimization loss function of the atlas A generator; D G (G(A)) represents the optimization objective of the generator graph A.
6. According to claim 5, the transformer infrared image style transfer method based on recurrent adversarial network is characterized in that: The adversarial loss of the second discriminator includes: In the formula, Denotes the optimization loss function of the atlas A generator; D A (A) represents the simulated target of atlas A; D A (F(B)) represents the discriminator atlas B simulation target.
7. The transformer infrared image style transfer method based on recurrent adversarial network according to claim 6 is characterized in that: The Identity loss includes: Where, L Identity (G,F) represents the Identity loss function; F(A) represents the simulation target of the discriminator atlas A; A represents image set A; B represents image set B; G(B) represents the optimization target of the generator atlas B.
8. A transformer infrared image style transfer system based on a recurrent adversarial network, the system being used to implement the method according to any one of claims 1 to 7, characterized in that: include: Collection systems, screening systems, and migration systems; The acquisition system is used to acquire sample image data, and pre-process the sample image data to obtain fourth image data; The screening system is connected to the acquisition system and is used to screen the fourth image data to obtain fifth image data; The migration system is connected to the screening system and is used to perform style migration on the fifth image data to obtain an engineering image.
Citation Information
Patent Citations
Transform-based cross-modal image matching and positioning method and device
CN116168221A
Image ink-wash style migration method and system based on cyclic generative adversarial network
CN116310712A