A Method for Identifying Long-Distance Dial Information Based on PSO and ESRGAN

By introducing the PSO-based ESRGAN method in the super-resolution image reconstruction technology, the problem of low information recognition rate in long-distance shooting is solved, and the lossless amplification of the image and the improvement of recognition accuracy is achieved.

CN118429187BActive Publication Date: 2025-07-01GUANGZHOU INST OF MEASURING & TESTING TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410560073.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-08
Publication Date
2025-07-01
Estimated Expiration
2044-05-08

AI Technical Summary

Technical Problem

In long-distance shooting scenes, the narrowing of the field of view of the camera leads to a low information recognition rate, and the existing super-resolution reconstruction methods are difficult to effectively extract small details, resulting in low recognition accuracy.

Method used

The Enhanced Super-Resolution Generative Adversarial Networks (ESRGAN) method based on particle swarm optimization (PSO) is adopted to build an ESRGAN model, and image feature similarity calculation is performed using the generative adversarial network and perceptual loss function, and the perceptual function weight in the PSO optimization generator is achieved to achieve lossless amplification of the image.

Benefits of technology

Lossless amplification of long-distance images is achieved, the clarity and recognition accuracy of images are improved, the presence of image distortion is reduced, and the weight parameters can be adjusted quickly and automatically, and it is suitable for shooting scenes of different instruments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118429187B_ABST
    Figure CN118429187B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of super-resolution image reconstruction, and discloses a method for identifying long-distance dial information based on PSO and ESRGAN, including Step 1: setting up an operation site; Step 2: establishing an ESRGAN model and calculating the feature similarity between the real high-resolution image and the image based on the generative adversarial network GAN and the perceptual loss function; Step 3: magnifying a new image through the best model obtained by training. The present invention solves the problem that when identifying the dial information on the surface of a measured instrument at a long distance, the image is unclear and the features are not obvious, resulting in a poor recognition rate. Through this method, it is possible to reconstruct a photo with a low resolution and overall blurriness taken at a long distance into a high-resolution state, ensuring that as many details in the image as possible can be obtained during subsequent image recognition and improving the recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of super-resolution image reconstruction, and particularly relates to a method for identifying long-distance dial information based on PSO-ESRGAN. Background Art

[0002] Generally, when taking long-distance photos, a lens with a longer focal length is usually selected. However, as the focal length increases, the field of view (FOV) of the lens becomes narrower, which results in very limited information that can be captured at a fixed camera position. If the camera is made position-adjustable, it will cause changes in the position and direction of the camera, affecting the internal and external parameters of the camera and increasing the difficulty of the calibration algorithm. The reconstruction of the present invention is mainly based on the identification of long-distance information on the dial of the instrument to be measured, which is an important technology used for subsequent automatic testing of the robotic arm. Currently, a fixed camera can achieve a higher recognition rate, but the disadvantage is that since the camera cannot be installed on the robotic arm, a large space is required between the camera position and the instrument to be measured for the movement of the robotic arm. This leads to a problem that the font size of the image captured by the camera is small and the recognition rate is low.

[0003] In previous super-resolution reconstruction methods, traditional interpolation methods lack the reconstruction of image details and are prone to distortion. Super-resolution methods based on ordinary neural networks, such as SRCNN and VDSR, do not have very ideal magnification effects on small-size images and are prone to phenomena such as jagged edges, which have a greater impact on the recognition accuracy. In addition, the core of most current reconstruction technologies is to reconstruct blurred images into clear states, rather than magnifying images, so there are deficiencies in the extraction of minute details. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for identifying long-distance dial information based on PSO-ESRGAN to solve the above technical problems.

[0005] To solve the above technical problems, the specific technical solution of a method for identifying long-distance dial information based on PSO-ESRGAN of the present invention is as follows:

[0006] A method for identifying long-distance dial information based on PSO-ESRGAN includes the following steps:

[0007] Step 1: Set up the working site;

[0008] Step 2: Establish an ESRGAN model, and calculate the feature similarity between the real high-resolution and the image based on the generative adversarial network GAN and the perceptual loss function;

[0009] Step 3: Magnify the new image with the best model obtained through training.

[0010] Further, in step 1, the camera is installed at a high place, and an industrial camera is used for shooting to ensure that the entire instrument can be completely captured by the camera with a fixed position, and the size of the instrument to be measured does not exceed the field of view angle range of the camera.

[0011] Further, step 2 specifically includes:

[0012] Step 2.1: Construct the generator loss function;

[0013] Step 2.2: Construct the discriminator loss function;

[0014] Step 2.3: Select the method of PSO to optimize the selection of the weights of the perception function in the generator, and use the peak signal-to-noise ratio PSNR as the objective function to evaluate whether the iteration stops.

[0015] Further, step 2.1 includes constructing the loss function. The training process of ESRGAN involves adversarial training between the generator and the discriminator. The goal of the generator is to convert the low-resolution image LR into the high-resolution image HR. The loss function of the generator consists of adversarial loss and perceptual loss. The formula for the generator adversarial loss is:

[0016]

[0017] where D(G(x)) represents the discrimination result of the discriminator on the image G(x) generated by the generator, indicating the probability that the discriminator believes that the generated image G(x) is a real high-resolution image. represents the expected value of the probability distribution P of the random variable x for the low-resolution image LR. LR of;

[0018] The formula for the perceptual loss function is:

[0019]

[0020] In the formula is the selected feature extraction network. represents the gap between the features of the generated image and the features of the real high-resolution image. The perceptual loss function is used to calculate the feature similarity between the generated image and the real high-resolution image.

[0021] Then the total loss of the generator is:

[0022] L G = L adv (G) + λ per L per (G) (3)

[0023] where λ peris a hyperparameter and the optimal value needs to be found in specific repeated training. Here, the Particle Swarm Optimization (PSO) algorithm is combined to obtain the optimal value more quickly.

[0024] Furthermore, the loss adversarial function of the discriminator in step 2.2 is:

[0025]

[0026] Distinguish the real high-resolution image HR from the image G(R) generated by the generator.

[0027] Furthermore, in step 2.3, the Peak Signal-to-Noise Ratio (PSNR) is used as the objective function to evaluate whether the iteration stops:

[0028]

[0029] where 8 is the number of gray bits and MSE is the Mean Squared Error, and its formula is:

[0030]

[0031] Calculate the pixel point difference between the generated image and the real high-resolution image.

[0032] Substitute into the PSO algorithm formula, and its formula is

[0033]

[0034] where w is the inertia weight, which controls the contribution of the previous moment's speed to the current speed, c1 and c2 are learning factors, which respectively control the influence of individual experience and group collaboration on the speed. These three values are determined by the user according to the actual situation. r1 and r2 are random numbers, usually taking values between [0,1]. Let p g , p i be the global optimal position and the individual optimal position respectively.

[0035] Then, by combining PSNR as the objective function, we obtain:

[0036]

[0037] where argmin means selecting the minimum value among the objective functions.

[0038] Substitute the generated image calculated in the first iteration into the formula, that is, let After that, the iteration can start. Continuously update the speed and position of the particle swarm according to formulas (6) and (7), and re-obtain the individual and global optimal positions.

[0039] Then use the standard deviation formula

[0040]

[0041] To evaluate the stability of the global optimal solution, set the standard deviation = 0.01, and set t to 10. That is, if the standard deviation is less than 0.01 in 10 iterations, it is considered that the iteration enters the convergence state, and the iteration can be stopped. At this time, the global solution obtained is regarded as the optimal solution.

[0042] Furthermore, step 2.3 also includes that when the output of the loss function stabilizes at 50%, it can be considered convergent. After that, perform PSNR evaluation on the finally output generated image and input it into PSO to readjust λ in the generator. per , until λ per has the best weight for PSRN, and output the image with the best signal-to-noise ratio: The specific process is as follows:

[0043] Step a: First complete steps 2.1 and 2.2 once. At this time, the initially generated image obtained for the first time is obtained. Take the weight for generating this image as the initial weight and substitute it into the objective function to form formula (7), that is, the initial p i = λ0, and λ0 can be set to any random value, which is specifically selected by the user; Step b: Since all variables except the global optimal position in formula (6) are known at this time, calculate the global optimal position at this time, and reverse the value of the weight λ according to the new global position at this time. per ; Step c: Substitute the new weight and re-perform the processes in steps 2.1 and 2.2. Then repeat the above steps in a loop until the value calculated by formula (8) converges. At this time, regard the global position obtained as the optimal, that is, the weight λ per at this time is the optimal;

[0044] Then, the image generated using the optimal weight at this time obtains the best PSRN.

[0045] Furthermore, the method for obtaining the low-resolution and high-resolution images in step 3 is as follows: The instrument information intercepted from the picture taken at 60 cm is the low-resolution image, and the instrument information taken at a close distance is the high-resolution image. The difference between them is that the instrument information taken at 60 cm only occupies a small number of pixels in the picture, while the instrument information taken at a close distance occupies the vast majority of the pixels in the picture. In this way, the difference between the low-resolution image and the high-resolution image is distinguished.

[0046] A method for identifying long-distance dial information based on PSO of the present invention has the following advantages:

[0047] 1. The present invention realizes the lossless amplification of the picture taken in a special-purpose scenario where the camera position must be fixed, far from the object to be measured, and the whole body of the object to be measured must be photographed, without the need for additional operations such as repeatedly adjusting the angle or position to obtain a clearer picture.

[0048] 2. The ESRGAN method combined with PSO realizes the rapid automatic optimization of adjusting the weight parameters in ESRGAN, enabling it to be quickly deployed for shooting with different instruments. The ESRGAN lossless magnification method improved by PSO further improves the image quality and reduces image distortion. Based on the feature of removing the BN layer in ESRGAN, the existence of artifacts can be minimized as much as possible, making the magnified image clearer and closer to the real situation. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a schematic flow chart of establishing the ESRGAN model structure of the present invention;

[0050] Figure 2 It is a flow chart of obtaining the global optimal solution of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] In order to better understand the purpose, structure and function of the present invention, the following further describes in detail a method for identifying long-distance dial information based on ESRGAN of the present invention with reference to the accompanying drawings.

[0052] A method for identifying long-distance dial information based on ESRGAN of the present invention includes the following steps:

[0053] Step 1: Set up the working site;

[0054] Considering that the minimum movement size of a normal general robotic arm on the market is within a 50-cm semi-spherical range, the camera is installed at a height of 60 cm. The instrument to be photographed is about 10 cm long and 5 cm wide. At this time, a common industrial camera with a focal length of 2.8 and 12 cm is used for shooting to ensure that the entire instrument can be completely captured by the camera with a fixed position. The above are the application environmental conditions of the method of the present invention. The size of the instrument to be measured can be larger than the example above, provided that it does not exceed the field of view angle range of the camera.

[0055] Step 2: Establish an ESRGAN model;

[0056] ESRGAN (Enhanced Super-Resolution Generative Adversarial Networks) is a method for image super-resolution reconstruction. It is based on the generative adversarial network (GAN) and the perceptual loss function to calculate the feature similarity between the real high-resolution and the image to restore more image details and structures. Specifically, it includes:

[0057] Step 2.1: Construct the generator loss function, which is composed of adversarial loss and perceptual loss respectively;

[0058] Construct the loss function. The training process of ESRGAN involves adversarial training between the generator and the discriminator. The goal of the generator is to convert the low-resolution image LR into the high-resolution image HR. The loss function of the generator consists of the adversarial loss and the perceptual loss.

[0059] Among them, the formula for the generator adversarial loss is:

[0060]

[0061] This function mainly attempts to minimize the loss of the discriminator for the generated image, in an attempt to make the generated image as close as possible to the real high-resolution image. Among them, D(G(x)) in the formula represents the discrimination result of the discriminator for the image G(x) generated by the generator, indicating the probability that the discriminator believes the generated image G(x) is a real high-resolution image. represents the expected value of the probability distribution P of the random variable x for the low-resolution image LR LR of.

[0062] The formula for the perceptual loss function is:

[0063]

[0064] In the formula is the feature representation extracted by the selected feature extraction network, such as VGG or ResNet. represents the gap between the features of the generated image and the features of the real high-resolution image. The perceptual loss function is used to calculate the feature similarity between the generated image and the real high-resolution image.

[0065] Then the total loss of the generator is:

[0066] L G = L abv (G) + λ per L per (G) (3)

[0067] Among them, λ per is a hyperparameter that needs to find the optimal value in specific repeated training. Here, the particle swarm optimization algorithm (PSO) is combined to obtain the optimal value faster.

[0068] Step 2.2: Construct the discriminator loss function;

[0069] The loss adversarial function of the discriminator is:

[0070]

[0071] Distinguish the real high-resolution image HR from the image G(R) generated by the generator.

[0072] The overall process is asFigure 1 As shown in Figure 1 , for ESRGAN, the input of its generator is the low-resolution image LR, and the output is the reconstructed high-resolution image SR; the input of the discriminator is the original high-resolution image HR and SR, and the output is a probability value. The generator and the discriminator have their own optimization objectives, and these optimization objectives constitute the loss function of the entire model. Then, the generator and the discriminator are updated through the loss function, where the optimized parameters are the weight matrices and bias vectors inside the generator and discriminator networks.

[0073] After the loss functions of the generator and the discriminator both converge, it can be considered that the image generated by the generator at this time is judged as a real high-resolution image.

[0074] Step 2.3: To further optimize the quality of the generated image, it is also necessary to optimize the selection of the weights of the perceptual function in the generator. Here, the method of PSO is used for optimization in the present invention. The peak signal-to-noise ratio (PSNR) is used as the objective function to evaluate whether the iteration stops.

[0075]

[0076] where 8 is the number of gray bits, and MSE is the mean square error, and its formula is:

[0077]

[0078] It mainly calculates the pixel point differences between the generated image and the real high-resolution image.

[0079] Substitute it into the PSO algorithm formula, and its formula is

[0080]

[0081] where w is the inertia weight, which controls the contribution of the previous moment's speed to the current speed, c1 and c2 are the learning factors, which respectively control the influence of individual experience and group cooperation on the speed. These three values are determined by the user according to the actual situation. r1 and r2 are random numbers, usually taking values between [0,1]. Let p g , p i be the global optimal position and the individual optimal position respectively.

[0082] Then, combining PSNR as the objective function, we can obtain

[0083]

[0084] where argmin means to select the minimum value among the objective functions.

[0085] Substitute the generated image calculated in the first iteration into the formula, that is, let After that, the iteration can be started, and the velocity and position of the particle swarm are continuously updated according to formulas (6) and (7), and the individual and global optimal positions are obtained again.

[0086] Then, the standard deviation formula

[0087]

[0088] is used to evaluate the stability of the global optimal solution. Generally, in this method, the standard deviation is set to 0.01, and t is set to 10. That is, if the standard deviation is less than 0.01 in 10 iterations, it is considered that the iteration enters the convergence state, and the iteration can be stopped. The global solution obtained at this time is regarded as the optimal solution.

[0089] If the output of the loss function stabilizes at 50%, it can be considered convergent. After that, the PSNR evaluation is performed according to the finally output generated image and input into the PSO to readjust λ in the generator per , until the weight of λ per makes the PSRN the best, and the image with the best signal-to-noise ratio is output. The specific process is as Figure 2 shown:

[0090] Step a: First, complete the steps in Step 2.1 and Step 2.2 once. At this time, the initial generated image obtained for the first time is obtained. The weight for generating this image is used as the initial weight and substituted into the objective function to form formula (7). That is, the initial p i = λ0, and λ0 can be set to any random value, which is specifically selected by the user.

[0091] Step b: Since all variables in formula (6) except the global optimal position are known at this time, the global optimal position at this time can be calculated. According to the new global position at this time, the value of the weight λ per is deduced backward.

[0092] Step c: Substitute the new weight and re-perform the Figure 1 process. Then, the above steps are repeated until the value calculated by formula (8) converges. At this time, the global position obtained is regarded as the optimal, that is, the weight λ per at this time is the optimal.

[0093] Then, the image generated using the optimal weight can theoretically obtain the best PSRN.

[0094] It should be noted separately that the convergence of the loss function adjusts the parameters of the generator and discriminator itself, and the convergence of the objective function adjusts the weights in the composition of the loss function, which are two different levels of things. The ultimate goal is to make the finally output image have a better effect.

[0095] Step 3: Use the best model obtained through training to magnify the new image. Generally, if the true resolution image used during training is several times the low-resolution image at the initial input, the model result obtained through training can achieve the same multiple of magnification. Therefore, in the present invention, the low-resolution and high-resolution images are obtained in the following way: the instrument information intercepted from the picture taken at 60 cm is the low-resolution image, and the instrument information taken at a close distance is the high-resolution image. The difference between them is that the instrument information taken at 60 cm only occupies a small number of pixels in the picture, while the instrument information taken at a close distance occupies the vast majority of the pixels in the picture, so as to distinguish between the low-resolution image and the high-resolution image.

[0096] It can be understood that the present invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. Additionally, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.

Claims

1. A PSO-based ESRGAN long-distance dial information recognition method, characterized in that: The steps include: Step 1: Set up the work site; Step 2: Establish the ESRGAN model to calculate the feature similarity between the real high-resolution image and the image based on the generative adversarial network GAN and the perceptual loss function; Step 2.1: Construct the generator loss function; The step 2.1 includes forming a loss function. The training process of ESRGAN involves adversarial training between the generator and the discriminator. The goal of the generator is to convert the low-resolution image LR into a high-resolution image HR. The loss function of the generator consists of adversarial loss and perceptual loss, where the generator adversarial loss formula is: Where D(G(x)) represents the discriminator's judgment result on the image G(x) produced by the generator, which means the probability that the discriminator believes that the generated image G(x) is a real high-resolution image. represents the probability distribution P of the random variable x for the low-resolution image LR LR Expected value; The perceptual loss function formula is: In the formula is the selected feature extraction network, Represents the gap between the generated image features and the real high-resolution image features. The perceptual loss function is used to calculate the feature similarity between the generated image and the real high-resolution image. Then the total loss of the generator is: L G =L adv (G)+λ per L per (G) (3) where λ per It is a hyperparameter, and the optimal value needs to be found in specific repeated training. Here, the particle swarm algorithm PSO is combined to obtain the optimal value faster; Step 2.2: Construct the discriminator loss function; The loss adversarial function of the discriminator in step 2.2 is: Distinguish the real high-resolution image HR from the image G(R) generated by the generator; Step 2.3: Use the PSO method to optimize the selection of the perceptual function weights in the generator, and use the peak signal-to-noise ratio (PSNR) as the objective function to evaluate whether the iteration should be stopped; The step 2.3 uses the peak signal-to-noise ratio (PSNR) as the objective function for evaluating whether the iteration is stopped: Among them, 8 is the number of grayscale bits, and MSE is the mean square error, and its formula is: Calculate the pixel difference between the generated image and the real high-resolution image, Substitute into the PSO algorithm formula, the formula is Where w is the inertia weight, which controls the contribution of the previous speed to the current speed. c1 and c2 are learning factors, which respectively control the influence of individual experience and group collaboration on speed. These three values ​​are determined by the user according to the actual situation. r1 and r2 are random numbers, usually between [0,1]. Let p g , p i are the global optimal position and the individual optimal position respectively, Then combine PSNR as the objective function to obtain: Among them, argmin means selecting the minimum value between the objective functions. Substitute the generated image calculated in the first iteration into the formula, that is, After that, the iteration can be started, and the speed and position of the particle swarm are continuously updated according to formulas (6) and (7), and the individual and global optimal positions are obtained again. Then use the standard deviation formula To evaluate the stability of the global optimal solution, set the standard deviation = 0.01, and t is set to 10. That is, if the standard deviation is less than 0.01 in 10 iterations, the iteration is considered to have entered a convergence state, and the iteration can be stopped. The global solution obtained at this time is considered to be the optimal solution. Step 2.3 also includes the loss function output being considered converged when it stabilizes at 50%. After that, the PSNR is evaluated based on the final output generated image and input into the PSO to readjust the λ in the generator. per , until λ per The weights make PSRN optimal and output the image with the best signal-to-noise ratio: The specific process is as follows: Step a: First complete steps 2.1 and 2.2 once, and then get the initial generated image. Substitute the weight of the generated image into the objective function as the initial weight, and form formula (7), that is, the initial p i =λ0, λ0 can be set to any random value, which is selected by the user; Step b: Since all variables except the global optimal position in formula (6) are known at this time, the global optimal position at this time is calculated, and the weight value λ is inferred from the new global position at this time per ; Step c: Substitute the new weights and repeat the process in steps 2.1 and 2.2; Then the above steps are repeated until the value calculated by formula (8) converges. The global position obtained at this time is considered to be optimal, that is, the weight λ at this time is per is the best; At this time, the image generated by using the optimal weights obtains the best PSRN; Step 3: Use the best model obtained through training to enlarge the new image.

2. The PSO-based ESRGAN long-distance dial information recognition method according to claim 1 is characterized in that: In the step 1, the camera is installed at a high place, and an industrial camera is used for shooting to ensure that the entire instrument can be completely captured by the camera at a fixed position, and the size of the instrument of the measured object does not exceed the field of view of the camera.

3. The PSO-based ESRGAN long-distance dial information recognition method according to claim 1 is characterized in that: The method for obtaining low-resolution and high-resolution images in step 3 is that the instrument information captured in the picture taken at 60 cm is a low-resolution image, and the corresponding instrument information taken at a close distance is a high-resolution image. The difference between them is that the instrument information taken at 60 cm only occupies a small number of pixels in the picture, while the instrument information taken at a close distance occupies the vast majority of the pixels in the picture, so as to distinguish the difference between low-resolution images and high-resolution images.

Citation Information

Patent Citations

  • Image super-resolution reconstruction method based on improved ESRGAN

    CN114463176A

  • ESRGAN-based dual-perception loss image super-resolution reconstruction method

    CN114663289A