Railway intrusion target image super-resolution reconstruction method

Through improved generative adversarial network model and composite loss function, the problem of image resolution and detail reconstruction of railway invasion targets in traditional methods is solved, and higher quality super-resolution images are generated, which improves the performance of railway safety monitoring system.

CN120339073APending Publication Date: 2025-07-18CHANGCHUN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510465447.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional image super-resolution reconstruction methods are difficult to improve resolution while ensuring image details, and cannot generate railway invasion target images with higher quality and retain more details.

Method used

The improved generative adversarial network model is adopted, combining advanced degradation model, generator model, discriminator model and composite loss function, and the super-resolution reconstruction process of railway invasion target images is optimized through convolutional layers, dual attention mechanism modules and ERSDB dense residual blocks.

Benefits of technology

Generating more realistic super-resolution images improves image quality and accuracy of detail reconstruction, reduces network computing and improves training efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339073A_ABST
    Figure CN120339073A_ABST
Patent Text Reader

Abstract

The invention discloses a railway intrusion target image super-resolution reconstruction method, which mainly comprises a degradation model, a generator model, a discriminator model and a loss function. Comprising the following steps: designing a simulated railway intrusion target image high-order degradation model, inputting a high-resolution image IHR in a high-resolution railway intrusion target image data set into the railway intrusion target image high-order degradation model, and generating a real railway intrusion target low-resolution image ILR; constructing a generator model, inputting the railway intrusion target low-resolution image ILR into the designed generator model, and outputting to obtain a super-resolution image ISR; constructing a discriminator model, inputting the high-resolution image IHR and the super-resolution image ISR into the discriminator model, and judging whether the image is true or false; optimizing parameters of a remote sensing image super-resolution reconstruction network by calculating a loss function, enabling the network to converge, completing training, and obtaining a remote sensing image super-resolution reconstruction model; and evaluating the generator model by using a quantitative index, and performing railway intrusion target image super-resolution reconstruction by using the evaluated generator model. Through the improved generative adversarial network, texture details of the reconstructed image are richer, and a clearer high-resolution railway intrusion target image is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and specifically to a super-resolution reconstruction method for railway intrusion target images. Background Art

[0002] In a railway safety monitoring system, railway intrusion targets seriously endanger the safe operation of railways. Enhancing the clarity and resolution of railway intrusion target images is crucial for railway safety. It not only improves the performance of railway monitoring but also provides a good operating environment for subsequent intrusion target detection and tracking. Image super-resolution technology can effectively improve image resolution without upgrading hardware performance. Therefore, we propose a super-resolution reconstruction method to solve the above problems.

[0003] Traditional image super-resolution reconstruction methods are difficult to improve resolution while ensuring image details. As a new type of deep learning algorithm, generative adversarial networks have demonstrated their superior performance in image super-resolution reconstruction. Summary of the Invention (1) Technical Problems to be Solved

[0004] In view of the deficiencies of the prior art, the present invention provides a super-resolution reconstruction method for railway intrusion target images, which introduces an improved generative adversarial network model to make the reconstructed image have richer details and can generate super-resolution railway intrusion target images with higher quality and more retained details, thus solving the problems raised in the above background art. (2) Technical Solutions

[0005] The present invention, in order to achieve the above object, provides a super-resolution reconstruction method for railway intrusion target images, including the following steps:

[0006] (1) Input the high-resolution image I in the high-resolution railway intrusion target image dataset into the high-order degradation model of railway intrusion target images to generate a real low-resolution image I of railway intrusion targets; HR LR ;

[0007] (2) Input the low-resolution image I of railway intrusion targets into the generator model to output a super-resolution image I; LR SR ;

[0008] (3) Input the high-resolution image I and the super-resolution image I into the discriminator model, and the discriminant network compares the high-resolution image I and the super-resolution image I to judge true or false; HR SR HR SR ; ​​​​​

[0009] (4) Optimize the parameters of the remote sensing image super-resolution reconstruction network by calculating the loss function, converge the network, complete the training, and obtain the remote sensing image super-resolution reconstruction model;

[0010] (5) Evaluate the generator model using quantitative metrics, and use the evaluated generator model to perform super-resolution reconstruction on the railway intrusion target image.

[0011] Further, the operation process of the high-order degradation model of the railway intrusion target image in step (1) is as follows: The high-resolution image I HR performs bilinear interpolation downsampling; adds Gaussian noise to the downsampled image; and finally performs brightness change to obtain the real low-resolution railway intrusion target image I LR .

[0012] Further, the railway intrusion target described. The generator model in step (2) includes a convolutional layer, a dual attention mechanism module, an ERSDB dense residual block, and an upsampling layer. The convolutional layer includes a pre-convolution and a post-convolution. The pre-convolution is an asymmetric convolution for extracting features from the input low-resolution image, and the post-convolutional layer is used to generate the final high-resolution image; the attention mechanism module, that is, an attention module is added in the middle of the network to enhance the network's recovery of key features and the network operation speed; the ERSDB dense residual block, that is, a dense residual block is added in the middle of the network and is connected in parallel with the attention mechanism module to increase the texture details of the network and improve the ability of detail reconstruction; the upsampling layer is used to increase the resolution of the image.

[0013] Further, the asymmetric convolution is composed of three convolutional kernels of 1X3, 3X1, and 3X3 in parallel.

[0014] Further, the dual attention mechanism module is composed of 8 basic blocks. Each basic block consists of a convolution, a LeakyRelu activation function, a local residual connection, a channel attention block, and a spatial attention module.

[0015] Further, the channel attention block is composed of global average pooling, global maximum pooling, a fully connected layer, and a Sigmoid activation function; the spatial attention module is composed of global average pooling, global maximum pooling, and a Sigmoid activation function.

[0016] Further, the operation process of the channel attention mechanism is as follows:

[0017] Global pooling, perform global average pooling and global maximum pooling on each channel to obtain F avg and F max .

[0018] Feature fusion, the obtained pooling value F avg and F max are connected to obtain the vector F concat

[0019] F concat is input into the fully connected layer, and through the LeakyRelu activation function, the intermediate feature vector F fc1 is obtained.

[0020] Then, through another fully connected layer, the output feature vector F fc2 is obtained.

[0021] Activation function, using the Sigmoid function to activate F fc2 to obtain the channel weight F channel is obtained.

[0022] Channel weighting, the weight F channelF is reallocated to each channel of the input feature map to obtain the weighted feature map F channel_out is obtained.

[0023] Furthermore, the operation process of the spatial attention mechanism is as follows:

[0024] Global pooling, performing global average pooling and global max pooling on each channel.

[0025] Convolution operation, using a 7×7 convolutional kernel to perform convolution operation on the pooling value.

[0026] Activation function, using the Sigmoid function for activation.

[0027] Furthermore, the ERSDB dense residual module consists of a convolutional layer and a LeakyReLU activation function.

[0028] Furthermore, the design of the ERSDB dense residual module is as follows: adopting a branch structure, one branch is densely connected by 4 convolutional activation blocks in a residual manner, and the other branch is a 1×3 and 3×1 convolutional branch. After aggregating the features extracted by the two branches, a convolution output is performed.

[0029] Furthermore, the discriminator model in step (3) includes an initial convolutional layer, a LeakyReLU activation function, a depth convolutional layer, a fully connected layer, and a Sigmoid activation function. The initial convolutional layer is an asymmetric convolution; the depth convolutional layer includes 7 basic modules, and each basic module consists of a convolutional block, a Leaky ReLU activation function, and a dual attention mechanism module.

[0030] Furthermore, the discriminator model structure is designed as follows: First, the first part consists of 1 asymmetric convolution block and 1 Leaky ReLU activation function. The second part is composed of 7 basic modules, and each basic module consists of 1 convolution block, 1 Leaky ReLU activation function, and 1 dual attention mechanism module. The third part consists of 1 fully connected layer with a dimension of 1024 followed by 1 Sigmaid activation function.

[0031] Furthermore, step (4) is specifically as follows: Select random images and train the generator and the discriminator. When the discriminator cannot distinguish the generated super-resolution image I SR and the original high-resolution image I HR , it means that the training is completed. During the training process, the Adam optimizer is used for optimization, and the network parameters are adjusted by means of weighted fusion of adversarial loss, reconstruction loss, contrast loss, and total variation loss. The formula is expressed as:

[0032] L object = L GAN + αL rec + L comp + L TV

[0033] In the formula, L GAN is the adversarial loss, L erc is the reconstruction loss, L comp is the contrast loss, L TV is the total variation loss, and α is a constant coefficient for controlling the reconstruction loss, which is set to 50 or 100 respectively according to the super-resolution magnification factor during the experiment;

[0034] Furthermore, in step (5), the peak signal-to-noise ratio and structural similarity between the super-resolution image I HR and the original high-resolution image I HR are calculated to evaluate the model performance. If the performance requirements are met, deployment is carried out; otherwise, the parameters are readjusted and retrained. (III) Beneficial Effects

[0035] Compared with the prior art, the present invention provides a method for super-resolution reconstruction of railway intrusion target images, having the following beneficial effects:

[0036] The high-order degradation model designed by the present invention can simulate the adverse factors of railway intrusion target images, generate low-resolution images that are closer to the actual situation, can provide more real data samples for the generator, enable it to be more adapted to the railway intrusion target scenario, and thus generate more real super-resolution images.

[0037] The present invention designs asymmetric convolution, replaces some convolutions in the model, reduces the computational amount of network parameters, and improves the efficiency of network training while maintaining the size of the receptive field unchanged.

[0038] The present invention introduces a dual attention mechanism module and an ERSDB dense residual module into the model, enabling the network to perform deeper training and more accurately locate the fine details in the railway intrusion target image, thereby improving the image quality and the authenticity of the reconstructed target.

[0039] The present invention provides a composite loss function that can comprehensively evaluate the quality of the generated image, thereby improving the quality of the generated image. Brief Description of the Drawings

[0040] Figure 1 is the overall flowchart of the present invention;

[0041] Figure 2 is the schematic diagram of the image degradation model of the present invention;

[0042] Figure 3 is the structural diagram of the ERSDB dense residual module of the present invention;

[0043] Figure 4 is the structural diagram of the dual attention mechanism module of the present invention.

[0044] Figure 5 is the structural diagram of the generator and discriminator networks of the present invention. Detailed Embodiments

[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Embodiment

[0046] As Figure 1 shown, a super-resolution reconstruction method for railway intrusion target images proposed in an embodiment of the present invention includes the following steps:

[0047] (1) Input the high-resolution image I in the dataset HR into the high-order degradation model of the railway intrusion target image to generate a low-resolution railway intrusion target image I LR ;

[0048] (2) Input the low-resolution image I of the railway intrusion target image LR into the generator model, and output the super-resolution image I SR;

[0049] (3) Input the high-resolution image I HR and the super-resolution image I SR into the discriminator model, and the discriminative network compares the high-resolution image I HR and the super-resolution image I SR to determine true or false;

[0050] (4) Optimize the parameters of the remote sensing image super-resolution reconstruction network by calculating the loss function, converge the network, complete the training, and obtain the remote sensing image super-resolution reconstruction model;

[0051] (5) Evaluate the generator model using quantization metrics, and use the evaluated generator model for super-resolution reconstruction of railway intrusion target images.

[0052] As Figure 2 shown, the operation process of the railway intrusion target image degradation model in step (1) is as follows: The high-resolution image I HR is downsampled by bilinear interpolation; Gaussian noise is added to the downsampled image; and finally, brightness change is performed to obtain the real low-resolution railway intrusion target image I LR .

[0053] As Figure 5 shown, the generator model in step (2) includes a convolutional layer, a dual attention mechanism module, an ERSDB dense residual block, and an upsampling layer. The convolutional layer includes a pre-convolution and a post-convolution. The pre-convolution is an asymmetric convolution for extracting features from the input low-resolution image, and the post-convolutional layer is used to generate the final high-resolution image; the attention mechanism module, that is, an attention module is added in the middle of the network to enhance the network's recovery of key features and the network operation speed; the ERSDB dense residual block, that is, a dense residual block is added in the middle of the network and connected in parallel with the attention mechanism module to increase the texture details of the network and improve the ability of detail reconstruction; the upsampling layer is used to increase the resolution of the image.

[0054] The asymmetric convolution is composed of three convolutional kernels of 1X3, 3X1, and 3X3 in parallel.

[0055] As Figure 3 shown, the ERSDB dense residual module is composed of a convolutional layer and a LeakyReLU activation function. The design is as follows: A branch structure is adopted. One branch is densely connected by 4 convolutional activation blocks in parallel, and the other branch is a 1×3, 3×1 convolutional branch. The features extracted by the two branches are aggregated and then a convolution output is performed.

[0056] As Figure 4As shown, the dual attention mechanism module is composed of 8 basic blocks. Each basic block consists of a convolution, a LeakyRelu activation function, a local residual connection, a channel attention block, and a spatial attention module.

[0057] The operation process of the channel attention mechanism is as follows:

[0058] Global pooling, performing global average pooling and global max pooling on each channel to obtain F avg and F max .

[0059] Feature fusion, connecting the obtained pooling values F avg and F max to obtain the vector F concat

[0060] Input F concat into the fully connected layer, and after passing through the LeakyRelu activation function, obtain the intermediate feature vector F fc1 .

[0061] Then pass through another fully connected layer to obtain the output feature vector F fc2 .

[0062] Activation function, using the Sigmoid function to activate F fc2 to obtain the channel weight F channel .

[0063] Channel weighting, reassigning the weight F channelF to each channel of the input feature map to obtain the weighted feature map F channel_out .

[0064] The operation process of the spatial attention mechanism is as follows:

[0065] Global pooling, performing global average pooling and global max pooling on each channel.

[0066] Convolution operation, using a 7×7 convolutional kernel to perform convolution operation on the pooling value.

[0067] Activation function, using the Sigmoid function to activate the output result.

[0068] As Figure 5 shown, the discriminator model in step (3) includes an initial convolutional layer, a LeakyReLU activation function, a depth convolutional layer, a fully connected layer, and a Sigmoid activation function. The initial convolutional layer is an asymmetric convolution; the depth convolutional layer includes 7 basic modules, and each basic module consists of a convolutional block, a Leaky ReLU activation function, and a dual attention mechanism module.

[0069] The discriminator model structure is designed as follows: First, the first part consists of 1 asymmetric convolution block and 1 Leaky ReLU activation function. The second part is composed of 7 basic modules, and each basic module consists of 1 convolution block, 1 Leaky ReLU activation function, and 1 dual attention mechanism module. The third part consists of a fully connected layer with a dimension of 1024 followed by 1 Sigmoid activation function.

[0070] Step (4) is specifically as follows: Select random images to train the generator and the discriminator. When the discriminator cannot distinguish the generated super-resolution image I SR and the original high-resolution image I HR , it means that the training is completed. During the training process, the Adam optimizer is used for optimization, and the adversarial loss, reconstruction loss, contrast loss, and total variation loss are weighted and fused to guide the adjustment of network parameters. The formula is expressed as:

[0071] L object = L GAN + αL rec + L comp + L TV

[0072] In the formula, L GAN is the adversarial loss, L erc is the reconstruction loss, L comp is the contrast loss, L TV is the total variation loss, and α is a constant coefficient that controls the reconstruction loss. It is set to 50 or 100 respectively according to the super-resolution magnification factor during the experiment.

[0073] Furthermore, in step (5), the peak signal-to-noise ratio and structural similarity between the super-resolution image I HR and the original high-resolution image I HR are calculated to evaluate the model performance. If the performance requirements are met, deployment is carried out. Otherwise, the parameters are readjusted and retrained.

[0079] In this embodiment, a large number of experiments are carried out to verify the effectiveness of the method of the present invention, and the performance of the method of the present invention is compared with that of other existing algorithms. The comparative experiments use the same training and test data sets, and all the networks involved are retrained and tested. The detailed results of the experiments are shown in Table 1:

[0080] Table 1 Comparison of quantitative indicators of reconstruction effects of different methods

[0081] SR algorithm Bicubic SRCNN SRGAN ESRGAN Ours PSNR / dB 25.36 28.43 28.31 28.73 29.21 SSIM 0.5711 0.5821 0.5896 0.6131 0.6329

[0082] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A super-resolution reconstruction method for railway intrusion target images, characterized in that Including: (1) Design a high-order degradation model for simulating railway intrusion target images, and input the high-resolution image I in the high-resolution railway intrusion target image dataset into the high-order degradation model of railway intrusion target images to generate a real low-resolution image I of railway intrusion targets HR ; LR ; (2) Design the generator model of the railway intrusion target image reconstruction network, and input the low-resolution image I of the railway intrusion target image into the generator model to output the super-resolution image I LR ; SR ; (3) Design the discriminator model of the railway intrusion target image reconstruction network, and input the high-resolution image I HR and the super-resolution image I SR into the discriminator model. The discriminator network compares the high-resolution image I HR and the super-resolution image I SR to judge true or false; (4) Optimize the parameters of the remote sensing image super-resolution reconstruction network by calculating the loss function to make the network converge, complete the training, and obtain the remote sensing image super-resolution reconstruction model; (5) Evaluate the generator model using quantitative metrics, and use the evaluated generator model to perform super-resolution reconstruction on the railway intrusion target image.

2. The super-resolution reconstruction method of railway intrusion target images according to claim 1, wherein The operation process of the high-order degradation model of the railway intrusion target image in step (1) is as follows: perform bilinear interpolation downsampling on the high-resolution image I HR ; add Gaussian noise to the downsampled image; and finally perform brightness change to obtain the real low-resolution image I LR .

3. The super-resolution reconstruction method for railway intrusion target images according to claim 1, characterized in that: The generator model in step (2) includes a convolutional layer, a dual attention mechanism module, an ERSDB dense residual block, and an upsampling layer. The convolutional layer includes a pre-convolution and a post-convolution. The pre-convolution is an asymmetric convolution used to extract features from the input low-resolution image, and the post-convolutional layer is used to generate the final high-resolution image; the attention mechanism module, that is, an attention module is added in the middle of the network to enhance the network's recovery of key features and the network operation speed; the ERSDB dense residual block, that is, a dense residual block is added in the middle of the network and connected in parallel with the attention mechanism module to increase the texture details of the network and improve the ability of detail reconstruction; the upsampling layer is used to increase the resolution of the image.

4. A super-resolution reconstruction method for railway intrusion target images according to claim 3, characterized in that: The asymmetric convolution consists of 3 parallel convolutional kernels.

5. The super-resolution reconstruction method for railway intrusion target images according to claim 3, characterized in that: The dual attention mechanism module consists of 8 basic blocks. Each basic block consists of a convolution, a LeakyRelu activation function, a local residual connection, a channel attention block, and a spatial attention module.

6. The super-resolution reconstruction method of railway intrusion target images according to claim 5, characterized in that: The channel attention block consists of global average pooling, global max pooling, a fully connected layer, and a Sigmoid activation function; the spatial attention module consists of global average pooling, global max pooling, and a Sigmoid activation function.

7. A super-resolution reconstruction method for railway intrusion target images according to claim 3, characterized in that: The ERSDB dense residual module consists of a convolutional layer and a LeakyReLU activation function.

8. A super-resolution reconstruction method for railway intrusion target images according to claim 1, characterized in that The discriminator model in step (3) includes an initial convolutional layer, a LeakyReLU activation function, a depth convolutional layer, a fully connected layer, and a Sigmoid activation function. The initial convolutional layer is an asymmetric convolution; the depth convolutional layer includes 7 groups of basic modules, and each basic module consists of a convolutional block, a Leaky ReLU activation function, and a dual attention mechanism module.