EDSR image super-resolution reconstruction method based on particle swarm optimization
By modeling the EDSR network structure as a six-dimensional optimization problem and using the particle swarm optimization algorithm to automatically optimize the network structure, the problem of insufficient adaptability of the EDSR network is solved, and the quality and efficiency of image super-resolution reconstruction are improved.
Patent Information
- Application Number
- CN202510795972.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-26
AI Technical Summary
The existing EDSR network structure is fixed, with poor adaptability and insufficient ability to restore image details. In addition, traditional hyperparameter tuning methods are inefficient and difficult to optimize the key components of the network structure.
The EDSR network structure is modeled as a six-dimensional optimization problem, and the particle swarm optimization algorithm is used to automatically find the optimal network configuration in the structure search space. Through structural search, self-similar feature modeling and group sparse feature enhancement, the residual block configuration, convolutional layer parameters, upsampling method and attention mechanism are optimized to improve reconstruction performance and computational efficiency.
It achieves automatic optimization of the network structure while maintaining model performance, improves the quality and practicality of image super-resolution reconstruction, and enhances the model's adaptability and computational efficiency.
Smart Images

Figure CN120707391A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing and deep learning technologies, and in particular to an image super-resolution reconstruction method combining a particle swarm optimization algorithm with an enhanced residual network (EDSR). Background Art
[0002] Image super-resolution (SR) aims to restore high-resolution images from low-resolution ones. It is an important research direction in computer vision and is widely used in remote sensing, medical imaging, video surveillance, and other fields. In recent years, deep convolutional neural networks (CNNs) have made significant progress in image super-resolution tasks. In particular, the Enhanced Deep Super-Resolution Network (EDSR) has achieved leading performance on multiple benchmark datasets by removing batch normalization layers and increasing model capacity.
[0003] However, the traditional EDSR network structure is fixed, and it is difficult to adaptively adjust its structural parameters for specific tasks or data scenarios, resulting in limited model generalization ability and the lack of texture details in the reconstructed image. In addition, deep models such as EDSR still face challenges in practical applications, such as complex model structure design, difficult parameter selection, and high computing resource consumption. Traditional hyperparameter tuning methods, such as grid search and random search, are inefficient and difficult to cope with high-dimensional parameter spaces. To this end, researchers began to explore the use of intelligent optimization algorithms to automatically design and optimize network structures. Particle Swarm Optimization (PSO), as a global optimization algorithm, has been successfully applied to weight optimization and structure search of neural networks, showing good performance. (Reference: Yu Xuebin, Jia Yuchen, Gao Li'ai, et al. Considering the improved particle swarm algorithm to optimize the BP neural network soft sensor prediction model for biogas production [J]. Acta Energiae Solaris Sinica, 2024, 45(8): 643-650.)
[0004] In the field of image super-resolution, some studies have attempted to combine PSO with lightweight network structures to optimize network parameters and structural configurations, thereby improving reconstruction quality and model efficiency. For example, researchers have proposed a PSO-based convolutional neural network training method that maintains the diversity of the particle swarm and avoids falling into local optimality by introducing a cosine similarity mutation strategy. (Reference: Zhou C, Xiong A. Fast image super-resolution using particle swarm optimization-based convolutional neural networks[J]. Sensors, 2023, 23(4): 1923.) In addition, there are studies that use evolutionary algorithms to search for efficient residual dense block structures to build fast, lightweight and accurate image super-resolution networks. (Reference: An Z, Zhang J, Sheng Z, et al. RBDN: Residual bottleneck dense network for image super-resolution[J]. IEEE Access, 2021, 9: 103440-103451)
[0005] However, existing methods still lack the granularity and flexibility of network structure search, making it difficult to simultaneously optimize multiple aspects, such as the number of residual blocks, partitioning methods, convolutional layer parameters, upsampling strategies, and the introduction of attention mechanisms. Therefore, a more systematic and efficient approach combining deep learning models with intelligent optimization algorithms is urgently needed. This approach can automatically optimize key components of the network structure while maintaining model performance to improve image super-resolution reconstruction. Summary of the Invention
[0006] In order to solve the problems of fixed EDSR network structure, poor adaptability and insufficient image detail restoration ability in the existing technology, the present invention proposes an EDSR super-resolution reconstruction method based on particle swarm optimization, which effectively improves the image reconstruction quality through structure search, self-similar feature modeling and group sparse feature enhancement.
[0007] The specific steps are as follows:
[0008] The EDSR network structure is modeled as a six-dimensional structural optimization problem. The six-dimensional structure is disassembled and optimized separately, including the residual block configuration, convolution layer parameters, upsampling method, activation function type, and the introduction of the attention mechanism. This method uses PSO to automatically find the optimal network configuration in the structural search space, while improving reconstruction performance while taking into account model complexity and computational efficiency, thereby improving the quality and practicality of image super-resolution reconstruction.
[0009] S1: Input the original low-resolution image data, prepare the training set and validation set, and provide training samples and performance measurement basis for the evaluation of the network structure;
[0010] S2: Define the EDSR network architecture optimization space. Each particle represents a candidate 6-dimensional EDSR network structure, encoded as a vector X = [x1, x2, x3, x4, x5, x6]. The meaning of each optimization dimension is as follows:
[0011] x1 represents the input convolution layer configuration, and the optimized sub-parameters include the size of the input convolution kernel.
[0012] Number of channels and activation function type;
[0013] x2 represents the residual block group structure configuration. The derived parameters include the total number of residual blocks R and the number of residual block stacks S, the convolution kernel size, number of channels, and activation function type (such as ReLU, PReLU, and SiLU) of the two layers of convolution inserted in the residual block and the inter-stage convolution inserted between the residual blocks;
[0014] x3 represents the convolution configuration inserted after the residual block group, and the derived parameters are the same as x1;
[0015] x4 represents the attention module configuration, and the derived parameters include SE, CBAM, and ECA types;
[0016] x5 represents the upsampling module configuration, and the derived parameters include the upsampling method (optional: PixelShuffle, transposed convolution, interpolated convolution) and the amplification factor;
[0017] x6 represents the output convolution layer configuration, and the optimized sub-parameters include the size of the input convolution kernel.
[0018] Number of channels and activation function type;
[0019] All derived parameters (such as convolution kernel size) are not optimized independently, but are derived from the main dimension values through mapping tables or calculation rules to achieve "dependent structural configuration modeling";
[0020] S3: Initialize the parameters of the particle swarm optimization algorithm to provide a search basis for particles, enabling them to have the ability to conduct parallel and global searches in complex structure spaces. The parameters include the number of particles, the maximum number of iterations T, the inertia factor ω, the individual learning factor c1, the group learning factor c2, the speed range, the position range, the early stopping iteration threshold K, and the fitness improvement threshold ΔF_th;
[0021] S4: Each particle dynamically constructs the EDSR network structure according to its position code. The construction steps are carried out in the following order:
[0022] S41: The first stage (residual skeleton construction) divides the residual block groups according to the total number of residual blocks R derived from x2 and the number of residual block stacks S, and inserts convolutional layers in and between residual blocks according to the rules derived from x2;
[0023] S42: The second stage (convolutional feature processing) configures the input convolution layer, the residual block group post-convolution layer, and the output convolution layer according to the parameters derived from x1, x3, and x6;
[0024] S43: The third stage (attention enhancement) sets the type of attention module inserted according to the x4 derivative parameter setting to enhance feature expression capabilities, but the attention module may introduce a slight increase in reasoning complexity;
[0025] S44: The fourth stage (up-sampling reconstruction) configures the up-sampling module according to the up-sampling method and amplification factor derived from x5;
[0026] S5: Fitness function evaluation. After training (or short-term fine-tuning) each network structure, calculate its performance index on the validation set and construct the fitness function F(X). The fitness function is:
[0027] F(X) = α PSNR(X) - β FLOPs(X)
[0028] Where: X is the network structure generated by particle coding; PSNR(X) is the peak signal-to-noise ratio after lightweight training on the validation set; FLOPs(X) represents the inference computational complexity of the network structure; α and β are weight coefficients (balancing performance and efficiency), which adjust the importance of indicators and are used to balance performance and complexity;
[0029] S6: Update individual optimality and global optimality;
[0030] S61: Individual Optimal Update (pBest)
[0031] For each particle Xi, if the current fitness is better than the historical optimal, the individual optimal is updated:
[0032]
[0033] otherwise:
[0034]
[0035] S62: Global Optimal Update (gBest)
[0036] Find the particle j with the best fitness among all particles. If its current fitness is better than the global optimal:
[0037]
[0038] otherwise:
[0039] gBest t+1 =gBest t
[0040] S7: Use the particle swarm update strategy to iterate the particle position and velocity, that is, the current value and change trend of the key structural parameters in the EDSR network structure:
[0041] S71: Continuous / ordered discrete parameter update (such as the number of channels, convolution kernel size, number of residual blocks R, number of stages S):
[0042]
[0043] The variables are described as follows:
[0044] t represents the current iteration round of the particle (integer);
[0045] represents the particle i-th dimension velocity, indicating the trend and magnitude of changes in each structural parameter in the next iteration;
[0046] Represents the current encoding of the particle, that is, the complete structural configuration of the EDSR network corresponding to the particle;
[0047] represents the best performing structure in the particle's own search history;
[0048] g best Represents the structure with the best fitness among all particles;
[0049] ω represents the inertia factor, which controls the tendency of the particle to maintain the current search direction;
[0050] c1 and c2 represent learning factors, which control the degree to which particles are affected by their own experience and group experience;
[0051] r1, r2 represent two random numbers between [0, 1], which are used to increase the randomness of the search;
[0052] Indicates the updated structural configuration;
[0053] S72: Discrete / classification parameter update (such as activation function type, upsampling method, attention type): adopt the probability jump strategy to generate a random number rand∈[0,1]
[0054]
[0055] If rand() <P flip , then the current value is randomly switched between optional values;
[0056] S8: Determine whether the termination conditions are met (termination occurs if any one of them is met):
[0057] Reach the maximum number of iterations T;
[0058] The global optimal fitness F(gBest) is improved by less than the preset threshold in K consecutive iterations (early stopping); the global optimal fitness F(gBest) reaches the preset target threshold; if the termination condition is not met, return to step S4 to continue iteration;
[0059] S9: Output the EDSR network structure corresponding to the global optimal particle position gBest as the optimal architecture obtained by the particle swarm search;
[0060] S10: Use the full training set to fully train the optimal EDSR network structure and apply the trained optimal EDSR model to super-resolution reconstruct the input low-resolution image to generate high-quality image results:
[0061] I_SR=EDSR * (I_LR)
[0062] Where I_LR is the input low-resolution image, I_SR is the output super-resolution image, and EDSR * () Train the final EDSR network for the optimal network architecture. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 This is a flow chart of a particle swarm optimized EDSR image super-resolution reconstruction method of the present invention;
[0064] Figure 2 This is a schematic diagram of the EDSR network structure and the internal structure of the module of the present invention. The present invention mainly applies the particle swarm algorithm to the residual layer skeleton construction part; DETAILED DESCRIPTION
[0065] See below Figure 1 、 Figure 2 , a particle swarm optimized EDSR image super-resolution reconstruction method described in the present invention is described in detail.
[0066] like Figure 1 As shown, in order to obtain a better image super-resolution reconstruction effect, the present invention provides an image super-resolution method combining particle swarm optimization strategy and EDSR network structure adjustment, which specifically includes the following steps:
[0067] The EDSR network structure is modeled as a six-dimensional structural optimization problem. The six-dimensional structure is disassembled and optimized separately, including the residual block configuration, convolution layer parameters, upsampling method, activation function type, and the introduction of the attention mechanism. This method uses PSO to automatically find the optimal network configuration in the structural search space, while improving reconstruction performance while taking into account model complexity and computational efficiency, thereby improving the quality and practicality of image super-resolution reconstruction.
[0068] S1: Input the original low-resolution image data and prepare the training set and validation set to provide training samples and performance measurement basis for network structure evaluation. The low-resolution image size usually used for training is 32x32 to 128x128, and the training set / validation set is divided in a ratio of 8:2;
[0069] S2: Define the EDSR network architecture optimization space. Each particle represents a candidate 6-dimensional EDSR network structure, encoded as a vector X = [x1, x2, x3, x4, x5, x6]. The meaning of each optimization dimension is as follows:
[0070] x1 represents the input convolution layer configuration. The optimized sub-parameters include the size of the input convolution kernel (set to 3x3), the number of channels C (the selection range is C∈[64,256], that is, the number of channels can be adjusted from 64 to 256), and the activation function type (such as ReLU, PReLU, SiLU);
[0071] x2 represents the residual block group structure configuration. The derived parameters include the total number of residual blocks R (selected in the range of R∈[16,48]), the number of residual block stacks S (selected in the range of S∈[1,4], the number of blocks per stage is R / S), the convolution kernel size, number of channels, and activation function type of the two layers of convolution inserted in the residual block and the inter-stage convolution inserted between residual blocks are set the same as x1;
[0072] x3 represents the convolution configuration inserted after the residual block group, and the derived parameters are: the size of the convolution kernel (set to 3x3), the number of channels C (selected in the range of C∈[128,512]);
[0073] x4 represents the attention module configuration, and the derived parameters can be selected from SE / CBAM / ECA types;
[0074] x5 represents the upsampling module configuration, and the derived parameters include the upsampling method (you can choose PixelShuffle, transposed convolution, interpolated convolution) and the amplification factor (select x4);
[0075] x6 represents the output convolution layer configuration, optimizing sub-parameters: the size of the convolution kernel (set to 3x3), and the number of channels C set to 3;
[0076] All derived parameters (such as convolution kernel size) are not optimized independently, but are derived from the main dimension values through mapping tables or calculation rules to achieve "dependent structural configuration modeling";
[0077] S3: Initialize the parameters of the particle swarm optimization (PSO) algorithm to provide a search basis for particles, enabling them to have the ability to conduct parallel and global searches in complex structure spaces. The parameters include the number of particles N∈[30,50], the maximum number of iterations T_max∈[80,150], the inertia factor ω initially 0.8, linearly decaying to 0.4 (ω(t)=0.8-0.4*(t / T_max)), the individual learning factor c1∈[1.0,2.0], the group learning factor c2∈[1.5,2.5], the speed range [v_min,v_max]∈[-3,3], the position range as defined in S2, the early stopping iteration threshold K∈[10,20], and the fitness improvement threshold ΔF_th∈[0.05,0.2]dB;
[0078] S4: Each particle dynamically constructs the EDSR network structure according to its position code. The construction steps are carried out in the following order:
[0079] S41: The first stage (residual skeleton construction) divides the residual blocks into groups according to the total number of residual blocks R derived from x2 and the number of residual block stacks S (e.g., R = 32, S = 4, i.e., 4 groups of residual blocks, each with 8 blocks), and inserts convolutional layers within and between residual blocks according to the rules derived from x2;
[0080] S42: The second stage (convolutional feature processing) configures the input convolution layer, the residual block group post-convolution layer, and the output convolution layer according to the parameters derived from x1, x3, and x6;
[0081] S43: The third stage (attention enhancement) sets the type of attention module inserted according to the x4 derivative parameter setting to enhance feature expression capabilities, but the attention module may introduce a slight increase in reasoning complexity;
[0082] S44: The fourth stage (upsampling reconstruction) configures the upsampling module according to the upsampling method and amplification factor derived from x5 (upsampling method priority: PixelShuffle: 70% transposed convolution: 20% interpolated convolution: 10%);
[0083] S5: Fitness function evaluation. After training (or short-term fine-tuning) each network structure, calculate its performance index on the validation set and construct the fitness function F(X). The fitness function is:
[0084] F(X) = α PSNR(X) - β FLOPs(X)
[0085] Where: X is the network structure generated by particle encoding; PSNR(X) is the peak signal-to-noise ratio after lightweight training on the validation set; FLOPs(X) represents the inference computational complexity of the network structure; α and β are weight coefficients (balancing performance and efficiency) that adjust the importance of indicators and are used to balance performance and complexity, such as: high precision priority: α = 0.01; balanced mode: β = 0.03;
[0086] S6: Update individual optimality and global optimality;
[0087] S61: Individual Optimal Update (pBest)
[0088] For each particle Xi, if the current fitness is better than the historical optimal, the individual optimal is updated:
[0089]
[0090] otherwise:
[0091]
[0092] S62: Global Optimal Update (gBest)
[0093] Find the particle j with the best fitness among all particles. If its current fitness is better than the global optimal:
[0094]
[0095] otherwise:
[0096] gBest t+1 =gBest t
[0097] S7: Use the particle swarm update strategy to iterate the particle position and velocity, that is, the current value and change trend of the key structural parameters in the EDSR network structure:
[0098] S71: Continuous / ordered discrete parameter update (such as the number of channels, convolution kernel size, number of residual blocks R, number of stages S):
[0099]
[0100] The variables are described as follows:
[0101] t represents the current iteration round of the particle (integer);
[0102] represents the particle i-th dimension velocity, indicating the trend and magnitude of changes in each structural parameter in the next iteration;
[0103] Represents the current encoding of the particle, that is, the complete structural configuration of the EDSR network corresponding to the particle;
[0104] represents the best performing structure in the particle's own search history;
[0105] g best Represents the structure with the best fitness among all particles;
[0106] ω represents the inertia factor, which controls the tendency of the particle to maintain the current search direction (e.g., ω = 0.7);
[0107] c1 and c2 represent learning factors, respectively, which control the degree to which a particle is influenced by its own experience and the group's experience (e.g., c1 = 1.5, c2 = 1.7);
[0108] r1 and r2 represent two random numbers between [0, 1], which are used to increase the randomness of the search (e.g., r1 = 0.5, r2 = 0.8);
[0109] Indicates the updated structural configuration;
[0110] S72: Discrete / classification parameter update (such as activation function type, upsampling method, attention type): adopt the probability jump strategy to generate a random number rand∈[0,1]
[0111]
[0112] If rand() <P flip , then the current value is randomly switched between optional values;
[0113] S8: Determine whether the termination conditions are met (termination occurs if any one of them is met):
[0114] Reach the maximum number of iterations T (e.g. T = 100);
[0115] The global optimal fitness F(gBest) is improved by less than the preset threshold in K consecutive iterations (early stopping); the global optimal fitness F(gBest) reaches the preset target threshold; if the termination condition is not met, return to step S4 to continue iteration;
[0116] S9: Output the EDSR network structure corresponding to the global optimal particle position gBest as the optimal architecture obtained by the particle swarm search;
[0117] S10: Use the full training set to fully train the optimal EDSR network structure and apply the trained optimal EDSR model to super-resolution reconstruct the input low-resolution image to generate high-quality image results:
[0118] I_SR=EDSR * (I_LR)
[0119] Where I_LR is the input low-resolution image, I_SR is the output super-resolution image, and EDSR * () Train the final EDSR network for the optimal network architecture.
Claims
1. A particle swarm optimized EDSR image super-resolution reconstruction method, characterized in that: The steps include: S1: Input the original low-resolution image data, prepare the training set and validation set, and provide training samples and performance measurement basis for the evaluation of the network structure; S2: Define the EDSR network architecture optimization space. Each particle represents a candidate 6-dimensional EDSR network structure, encoded as a vector X = [x1, x2, x3, x4, x5, x6]. The meaning of each optimization dimension is as follows: x1 represents the input convolution layer configuration, and the optimized sub-parameters include the size of the input convolution kernel, the number of channels, and the activation function type; x2 represents the residual block group structure configuration. The derived parameters include the total number of residual blocks R and the number of residual block stacks S, the convolution kernel size, number of channels, and activation function type of the two layers of convolution inserted in the residual block and the inter-stage convolution inserted between the residual blocks; x3 represents the convolution configuration inserted after the residual block group, and the derived parameters are the same as x1; x4 represents the attention module configuration, and the derived parameters include types; x5 represents the upsampling module configuration, and the derived parameters include the upsampling method and amplification factor; x6 represents the output convolution layer configuration, and the optimized sub-parameters include the size of the input convolution kernel, the number of channels, and the activation function type; All derived parameters (such as convolution kernel size) are not optimized independently, but are derived from the main dimension values through mapping tables or calculation rules to achieve "dependent structural configuration modeling"; S3: Initialize the parameters of the particle swarm optimization algorithm to provide a search basis for particles, enabling them to have the ability to conduct parallel and global searches in complex structure spaces. The parameters include the number of particles, the maximum number of iterations T, the inertia factor ω, the individual learning factor c1, the group learning factor c2, the speed range, the position range, the early stopping iteration threshold K, and the fitness improvement threshold ΔF_th; S4: Each particle dynamically constructs the EDSR network structure according to its position code. The construction steps are carried out in the following order: S41: The first stage (residual skeleton construction) divides the residual block groups according to the total number of residual blocks R derived from x2 and the number of residual block stacks S, and inserts convolutional layers in and between residual blocks according to the rules derived from x2; S42: The second stage (convolutional feature processing) configures the input convolution layer, the residual block group post-convolution layer, and the output convolution layer according to the parameters derived from x1, x3, and x6; S43: The third stage (attention enhancement) sets the type of attention module inserted according to the x4 derivative parameter setting to enhance feature expression capabilities, but the attention module may introduce a slight increase in reasoning complexity; S44: The fourth stage (up-sampling reconstruction) configures the up-sampling module according to the up-sampling method and amplification factor derived from x5; S5: Fitness function evaluation. After training (or short-term fine-tuning) each network structure, calculate its performance index on the validation set and construct the fitness function F(X). The fitness function is: F(X) = α PSNR(X) - β FLOPs(X) Where: X is the network structure generated by particle coding; PSNR(X) is the peak signal-to-noise ratio after lightweight training on the validation set; FLOPs(X) represents the inference computational complexity of the network structure; α and β are weight coefficients (balancing performance and efficiency), which adjust the importance of indicators and are used to balance performance and complexity; S6: Update individual optimality and global optimality; S61: Individual Optimal Update (pBest) For each particle Xi, if the current fitness is better than the historical optimal, the individual optimal is updated: otherwise: S62: Global Optimal Update (gBest) Find the particle j with the best fitness among all particles. If its current fitness is better than the global optimal: otherwise: gBest t+1 =gBest t S7: Use the particle swarm update strategy to iterate the particle position and velocity, that is, the current value and change trend of the key structural parameters in the EDSR network structure: S71: Continuous / ordered discrete parameter update (such as the number of channels, convolution kernel size, number of residual blocks R, number of stages S): The variables are described as follows: t represents the current iteration round of the particle (integer); represents the particle i-th dimension velocity, indicating the trend and magnitude of changes in each structural parameter in the next iteration; Represents the current encoding of the particle, that is, the complete structural configuration of the EDSR network corresponding to the particle; represents the best performing structure in the particle's own search history; g best Represents the structure with the best fitness among all particles; ω represents the inertia factor, which controls the tendency of the particle to maintain the current search direction; c1 and c2 represent learning factors, which control the degree to which particles are affected by their own experience and group experience; r1, r2 represent two random numbers between [0, 1], which are used to increase the randomness of the search; Indicates the updated structural configuration; S72: Discrete / classification parameter update (such as activation function type, upsampling method, attention type): adopt the probability jump strategy to generate a random number rand∈[0,1] If rand() <P flip , then the current value is randomly switched between optional values; S8: Determine whether the termination conditions are met (termination occurs if any one of them is met): Reach the maximum number of iterations T; The global optimal fitness F(gBest) is improved by less than the preset threshold in K consecutive iterations (early stopping); the global optimal fitness F(gBest) reaches the preset target threshold; if the termination condition is not met, return to step S4 to continue iteration; S9: Output the EDSR network structure corresponding to the global optimal particle position gBest as the optimal architecture obtained by the particle swarm search; S10: Use the full training set to fully train the optimal EDSR network structure and apply the trained optimal EDSR model to super-resolution reconstruct the input low-resolution image to generate high-quality image results: I_SR=EDSR * (I_LR) Where I_LR is the input low-resolution image, I_SR is the output super-resolution image, and EDSR * () Train the final EDSR network for the optimal network architecture.
Citation Information
Cited By
Neural network partial discharge identification method fused with polarization particle swarm optimization
CN122132786A
A neural network partial discharge identification method fusing polarized particle swarm optimization
CN122132786B