Water meter image style migration algorithm based on reversible residual network
By using a reversible residual network in image style transfer, image features are extracted and fused, the problems of information loss and gradient disappearance are solved, and high-quality style transfer effect is achieved.
Patent Information
- Application Number
- CN202510205230.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has problems of information loss and gradient disappearance in image style migration, resulting in insufficient content retention and insufficient style transmission.
Using an image style transfer method based on reversible residual network, image features are extracted through cascade of reversible residual blocks and squeeze modules, feature fusion is performed by combining adaptive instance normalization layers, and image reconstruction is performed by optimizing loss function to obtain high-quality style transfer images.
It effectively solves the problems of information loss and gradient disappearance, improves the quality of image style transfer, and ensures the integrity of content and the accurate transmission of style.
Smart Images

Figure CN119991415A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision and deep learning technology, and in particular to an image style transfer method based on a reversible residual network. Background Art
[0002] In the process of water meter recognition, sometimes there will be recognition errors. One of the reasons is the influence of the water meter condition. Dirt, scale and other conditions on the water meter will affect the recognition of the water meter, but not every condition can be captured. This requires the use of image style transfer algorithms to simulate various possible conditions of water meters, so as to achieve the purpose of expanding the water meter data set through data enhancement through the style transfer algorithm, and finally put it into the recognition algorithm for training to improve the recognition accuracy.
[0003] Early traditional image style transfer methods were mainly based on manual features. For example, style transfer is achieved through texture synthesis algorithms. These methods use local texture features of images, such as gray-level co-occurrence matrix to describe texture, and then transplant the texture features of one image to another. However, this manual feature method has limitations, because it is difficult for hand-designed features to fully capture the complex image style and high-level semantic information of the image content. With the rise of deep learning, methods based on convolutional neural networks have become mainstream. This method defines content loss and style loss functions, and uses back propagation to optimize the generated image, achieving good results in content preservation and style transfer. However, traditional architectures have some problems in information transfer, such as information loss during forward propagation, because the pooling operation will cause some information loss, and the calculation of gradients during back propagation may encounter problems such as gradient vanishing.
[0004] The reversible residual network is a novel neural network architecture. Its design concept is to solve the problem of information loss in traditional convolutional neural networks. Unlike traditional convolutional neural networks, the reversible residual network adopts a reversible network structure, which allows information to flow better during forward and backward propagation. During forward propagation, information will not be lost due to operations such as pooling. During backward propagation, gradients can also be calculated more effectively, reducing the possibility of gradient disappearance. This feature of the reversible residual network provides a new idea for image style transfer. In image style transfer, accurately retaining the content information of the content image and effectively transferring the style information of the style image are the key. Traditional style transfer methods based on convolutional neural networks may have defects in content retention or inaccurate style transfer due to information loss. The reversible structure of the reversible residual network can better process this information and improve the quality of image style transfer. Summary of the invention
[0005] In order to address the shortcomings of the prior art, the present invention provides an image style transfer method based on a reversible residual network, which mainly includes two parts: image feature extraction based on a reversible residual network and feature fusion based on adaptive weight allocation, thereby realizing image style transfer and ensuring that the content of the transferred image will not be distorted and the style of the style image can be retained.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions.
[0007] Step 1: Dataset preparation: construct content image dataset and style image dataset;
[0008] Step 2: Image input: input content image and style image, and unify the tensor shapes of all input images;
[0009] Step 3, feature extraction: image feature extraction is completed through the cascade of reversible residual block and squeeze module;
[0010] Step 4: Channel refinement: Connect the cascaded reversible residual blocks through the channel refinement module to remove redundant information and obtain the final content feature Z c and style features Z s ;
[0011] Step 5: Feature fusion: The content feature Z obtained from step 4 c and style features Z s , feature fusion is performed through the adaptive instance normalization layer to obtain the fused feature image Z cs ;
[0012] Step 6: Loss calculation: The loss function L total Defined as L total =L s +αL l +βL cyc Among them, L s Represents the calculation style loss, L l Indicates the calculation of Laplace loss L cyc Represents the calculation of cycle consistency loss;
[0013] Step 7, image reconstruction: Based on the total loss obtained by combining the style loss, Laplace loss, and cycle consistency loss, the optimizer is used to derive and update the loss, thereby updating the decoder parameters, and using the reversible residual network reverse reasoning to reversely map the stylized representation back to the stylized image to obtain the corresponding final style transfer image. In the optimization process, the optimizer Adam is used for optimization processing to obtain the final model parameters.
[0014] The above-mentioned image style transfer method based on a reversible residual network. Compared with the prior art, the present invention has obvious beneficial effects. It can be seen from the above scheme that firstly, the corresponding content and style image data sets are constructed, and the content image and style image are mapped to the latent space, using the zero-padding method; then, the corresponding content features and style features of the content image and style image are extracted by the forward propagation of the reversible residual network, and the corresponding feature fusion is performed using the corresponding extracted content feature map and style feature map, and the adaptive instance normalization method is used to combine the features of the two; the fused feature map is reconstructed by the back propagation of the reversible residual network to obtain the corresponding style transfer image. The reversible residual network is used to extract the feature map, and the style loss, Laplace loss, cycle consistency loss, etc. are calculated according to the obtained feature map, and then the loss is derived and updated by the optimizer to update the network model parameters in the back propagation process, and the final parameters are obtained, and then the style transfer image is obtained. This invention is a method for image style transfer using a feedforward neural network. The mean and variance are used to measure the loss in the style loss calculation. At the same time, the Laplace operator is introduced for structure refinement to better preserve the structural integrity of the style transferred image.
[0015] In summary, the main advantages of the present invention are:
[0016] 1. Using the mean as the style loss measurement can be faster than using Gram for loss measurement, achieving faster style transfer training and making the effect more significant.
[0017] 2. Using the Laplace operator to perform corresponding structural enhancement can make the resulting image more excellent.
[0018] 3. Use reversible networks to achieve lossless results as much as possible, retain the integrity of the content to the greatest extent, and avoid content distortion. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0020] Figure 1 This is the technical roadmap for this approach.
[0021] Figure 2 Design graph for reversible residual block.
[0022] Figure 3 Designing a plot for the residual function
[0023] Figure 4 This is the style transfer effect picture. DETAILED DESCRIPTION
[0024] The present invention proposes a method for image style transfer based on a reversible residual network. First, the method needs to perform feature extraction through the forward propagation of the reversible residual network, and then use an adaptive instance normalization method to perform feature fusion, and finally obtain the stylized image through the back propagation of the reversible residual network.
[0025] Figure 1 This is the technical roadmap of the present invention, and the specific implementation of the present invention will be described below.
[0026] Step 1: Dataset preparation: First, construct content image datasets and style image datasets. For content images, water meter images with clean dials and clear numbers are preferred. Then, image data is cropped to reduce the distortion of scaled images caused by excessive aspect ratio during the input process, and the dataset is expanded at the same time.
[0027] Step 2: Image input: First, the input content image and style image are mapped to the latent space through an injection padding module, which increases the input dimension by zero padding along the channel dimension;
[0028] Step 3, feature extraction: through the reversible residual block such as Figure 2 The cascade of the squeeze module completes the extraction of image features, and the residual function F is as follows: Figure 3 It is implemented by consecutive transformation layers with a kernel size of 3. Each convolution layer is followed by an activation function, except the last one. A large receptive field is obtained by stacking multiple layers and blocks to capture dense pairwise relationships. In order to capture large-scale style information, a squeeze module is used to reduce the spatial information by 2 times and increase the channel dimension by 4 times. A multi-scale architecture is implemented by combining reversible residual blocks and squeeze modules to complete the extraction of content features and style features;
[0029] Step 4: Channel refinement: Use the channel refinement module to connect the cascaded reversible residual blocks. After removing redundant information, the final content feature Z is obtained. c and style features Z s ;
[0030] Step 5: Feature fusion: The content feature Z obtained from step 4 c and style features Z c , feature fusion is performed through the adaptive instance normalization layer to obtain the fused feature image Z cs ;
[0031] Step 6: Loss calculation: The loss function L total Defined as:
[0032] L total =L s +αL l +βL cyc(1)
[0033] L total =L s +αL l +βL cyc Among them, L s Represents the calculation of style loss, which is used to measure the difference in style features between the generated image and the style image.
[0034] L l Represents the calculation of Laplace loss. Introducing Laplace loss in reversible residual network. Directly introducing Laplace loss in network training may lead to blurred images because Laplace loss forces the network to smooth the image instead of maintaining pixel affinity. However, there is no such problem when introducing Laplace loss in reversible network. This is because the dual-objective transformation in our reversible network requires that all information be retained during forward and backward reasoning. The reversible network will not cheat the loss by smoothing the image because it will cause information loss. Since the transformation of the reversible network is deterministic, only a few style images with smooth textures can smooth the content structure.
[0035] L cyc Indicates that the calculation of cycle consistency loss, the reversible network has numerical errors and may cause obvious artifacts. Therefore, the cycle consistency loss is introduced to improve the robustness of the network.
[0036] The style loss is defined as:
[0037]
[0038] Among them, std and mean represent the calculated variance and mean, respectively, and each represents φ i Layers used to compute style loss in VGG-19. In our experiments, we use relu1_1, relu2_1, relu3_1, relu4_1 layers with the same weights.
[0039] The Laplace loss is defined as:
[0040]
[0041] Where, N is the number of image pixels, V c [I cs ] is the stylized image in channel C. Stylized image I cs The vectorization of L is the content image I c The Laplace matrix of .
[0042] The cycle consistency loss is defined as:
[0043] I cyc =||Ic1 -I c || (4)
[0044] Among them, I c1 To reconstruct the content image in the loop, I c is the input content image. We should be able to pass the style information of the content image I c Transfer to stylized image I cs , loop to reconstruct the content image I c1 .
[0045] Step 7, image reconstruction: Based on the total loss obtained by combining the style loss, Laplace loss, and cycle consistency loss, the optimizer is used to derive and update the loss, thereby updating the decoder parameters, and using the reversible residual network reverse reasoning to reversely map the stylized representation back to the stylized image to obtain the corresponding final style transfer image. In the optimization process, the optimizer Adam is used for optimization processing to obtain the final model parameters.
[0046] The above disclosure is only a specific embodiment of the present invention. According to the technical concept provided by the present invention, any changes that can be thought of by those skilled in the art should fall within the protection scope of the present invention.
Claims
1. A method for image style transfer based on a reversible residual network, characterized in that The steps include: Step 1: Dataset preparation: First, construct the content image dataset and style image dataset, perform image correction processing to reduce the distortion of the original image, and then crop the image data to reduce the distortion of the scaled image caused by the large aspect ratio during the input process, and expand the dataset at the same time; Step 2: Image input: First, the input content image and style image are mapped to the latent space through an injection padding module, which increases the input dimension by zero padding along the channel dimension; Step 3, feature extraction: The image features are extracted by cascading the reversible residual block as shown in Figure 2 and the squeeze module. The residual function F is implemented by continuous transformation layers with a kernel size of 3 as shown in Figure 3. Each convolutional layer is followed by an activation function, except the last one. A large receptive field is obtained by stacking multiple layers and blocks to capture dense pairwise relationships. In order to capture large-scale style information, a squeeze module is used to reduce the spatial information by 2 times and increase the channel dimension by 4 times. The reversible residual block and the squeeze module are combined to implement a multi-scale architecture to extract content features and style features; Step 4: Channel refinement: Use the channel refinement module to connect the cascaded reversible residual blocks. After removing redundant information, the final content feature Z is obtained. c and style features Z s ; Step 5: Feature fusion: The content feature Z obtained from step 4 c and style features Z s , feature fusion is performed through the adaptive instance normalization layer to obtain the fused feature image Z cs ; Step 6: Loss calculation: The loss function L total Defined as: L total =L s +αL l +βL cyc (1) Among them, L s Represents the calculation of style loss, which is used to measure the difference in style features between the generated image and the style image. The purpose of style loss is to guide the generated image to be as close to the style image as possible, while combining content loss to generate an image with both original content and target style. L l Represents the calculation of Laplace loss. Introducing Laplace loss in reversible residual network. Directly introducing Laplace loss in network training may lead to blurred images because Laplace loss forces the network to smooth the image instead of maintaining pixel affinity. However, there is no such problem when introducing Laplace loss in reversible network. This is because the dual-objective transformation in our reversible network requires that all information be retained during forward and backward reasoning. The reversible network will not cheat the loss by smoothing the image because it will cause information loss. Since the transformation of the reversible network is deterministic, only a few style images with smooth textures can smooth the content structure. L cyc Indicates that the calculation of cycle consistency loss, the reversible network has numerical errors and may cause obvious artifacts. Therefore, the cycle consistency loss is introduced to improve the robustness of the network. The style loss is defined as: Among them, std and mean represent the calculated variance and mean, respectively, and each represents φ i Layers used to compute style loss in VGG-19. In our experiments, we use relu1_1, relu2_1, relu3_1, relu4_1 layers with the same weights. The Laplace loss is defined as: Where, N is the number of image pixels, V c [I cs ] is the stylized image in channel C. Stylized image I cs The vectorization of L is the content image I c The Laplace matrix of . The cycle consistency loss is defined as: L cyc =||I c1 -I c || (4) Among them, I c1 To reconstruct the content image in the loop, I c is the input content image. We should be able to pass the style information of the content image I c Transfer to stylized image I cs , loop to reconstruct the content image I c1 . Step 7, image reconstruction: Based on the total loss obtained by combining the style loss, Laplace loss, and cycle consistency loss, the optimizer is used to derive and update the loss, thereby updating the decoder parameters, and using the reversible residual network reverse reasoning to reversely map the stylized representation back to the stylized image to obtain the corresponding final style transfer image. In the optimization process, the optimizer Adam is used for optimization processing to obtain the final model parameters.