Binocular stereo matching distance measurement precision improving method and device
Through deep learning blind super-resolution network, the resolution of binocular images is improved, and combined with the stereo matching algorithm, the shortcomings of the stereo matching algorithm in the existing technology in subpixel accuracy are solved, higher precision ranging is achieved, and the safety of autonomous driving and assisted driving is improved.
Patent Information
- Application Number
- CN202510185611.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-13
AI Technical Summary
There are limitations in the breakthrough of existing stereo matching algorithms in subpixel accuracy, and it is impossible to effectively display the information of tiny objects between pixel blocks, resulting in insufficient ranging accuracy.
Deep learning blind super-resolution network is used to map low-resolution binocular images into high-resolution images, and combined with stereo matching algorithms, a parallax map is generated and super-resolution processing is performed, and finally converted into a depth map to obtain the real distance between the object and the camera.
It significantly improves the accuracy of binocular stereo matching distance measurement, and can more accurately calculate the distance between the object and the camera, thereby improving the safety of autonomous driving and assisted driving.
Smart Images

Figure CN120147393A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of assisted driving, and particularly to a method and device for improving the ranging accuracy of binocular stereo matching. Background Art
[0002] In the scenarios of autonomous driving and assisted driving, the binocular stereo matching algorithm plays a crucial role. The binocular camera can calculate the real distance between the objects in the scene and the camera through the stereo matching algorithm. When the vehicle is driving on the road surface, the stereo matching algorithm can calculate the specific positions of the vehicles, pedestrians, obstacles, etc. in front, ensuring the safety of the driver. Therefore, it is very necessary and meaningful to improve the ranging accuracy of binocular stereo matching.
[0003] In the existing stereo matching algorithms, generally, the left image and the right image captured by the binocular camera are directly input into the stereo matching algorithm to calculate the distance. However, due to the limitations of the camera's photosensitive element, the stereo images we captured can only be presented in the form of pixel blocks, and it is impossible to show the more subtle object information between the pixel blocks in sub-pixel units, which may limit the breakthrough of the stereo matching algorithm in sub-pixel accuracy. Currently, a method of using a high-resolution camera has been proposed to solve this problem, but this will increase the manufacturing cost of the camera.
[0004] In view of this, the present invention is proposed. Summary of the Invention
[0005] The main purpose of the present invention is to disclose a method and device for improving the ranging accuracy of binocular stereo matching, which is used to solve the problem that the breakthrough of the stereo matching algorithm in sub-pixel accuracy is limited in the prior art.
[0006] To achieve the above object, according to one aspect of the present invention, a method for improving the ranging accuracy of binocular stereo matching is provided, and the following technical solutions are adopted:
[0007] A method for improving the ranging accuracy of binocular stereo matching includes: acquiring the left image and the right image captured by the binocular camera; inputting the left image and the right image into a pre-trained deep learning blind super-resolution network to obtain a high-resolution left image and a high-resolution right image; inputting the high-resolution left image and the high-resolution right image into the stereo matching algorithm for calculation to generate a disparity map; performing super-resolution multiplication processing on the disparity map, and based on the parameters of the camera, converting the processed disparity map into a depth map to obtain the real distance between the camera and the object.
[0008] Further, the deep learning blind super-resolution network includes: a generator and a discriminator. The generator is used to map the low-resolution image to a high-resolution image, and the discriminator is used to help the generator generate a more realistic high-resolution image.
[0009] Further, the process of the generator is as follows: The low-resolution image first enters the convolutional layer to extract shallow feature information, and then enters a group of dense residual blocks to extract deep feature information. After the extraction is completed, it enters another convolutional layer for processing, and at the same time, a long residual connection is added to ensure the convergence and stability of training. Next, the extracted deep feature information is sent to the upsampling layer, where the low-resolution features are converted into high-resolution features. Then, the high-resolution features are successively sent to three convolutional layers for feature compression, and finally, a high-resolution image is obtained.
[0010] Further, the process of the discriminator is as follows: The high-resolution image generated by the generator is used as the input and sent to the discriminator for determination. First, shallow feature extraction is performed through the convolutional layer, and then it enters the deep feature extraction layer to extract more information. Finally, the true and false probability of the image is obtained through the sigmoid function.
[0011] Further, inputting the high-resolution left image and the high-resolution right image into the stereo matching algorithm for calculation to generate a delicate and accurate disparity map includes: sending the high-resolution left image and the high-resolution right image to a 2D convolutional layer for feature extraction; on the one hand, the extracted features are input to a 3D convolutional layer to calculate the disparity cost, thereby obtaining an initial disparity map and a geometric encoding volume. On the other hand, the extracted features calculate the correlation of the image, and the calculated correlation is combined with the geometric encoding volume to obtain a combined encoding volume; the combined geometric encoding volume and the initial disparity Figure 1 are sent to the disparity optimization layer for the final disparity optimization to obtain the disparity map.
[0012] Further, performing super-resolution multiplication processing on the disparity map includes: uniformly dividing the value of the disparity map calculated by the stereo matching algorithm by the magnification factor of the image, and the calculated disparity is twice the original disparity.
[0013] According to another aspect of the present invention, a binocular stereo matching ranging accuracy improvement device is provided, and the following technical solutions are adopted:
[0014] A binocular stereo matching ranging accuracy improvement device includes: an acquisition module for acquiring the left image and the right image captured by the binocular camera; an input module for inputting the left image and the right image into a pre-trained deep learning blind super-resolution network to obtain a high-resolution left image and a high-resolution right image; a calculation module for inputting the high-resolution left image and the high-resolution right image into the stereo matching algorithm for calculation to generate a disparity map; a processing module for performing super-resolution multiplication processing on the disparity map and converting the processed disparity map into a depth map based on the parameters of the camera to obtain the true distance between the camera and the object.
[0015] Further, the input module includes a generator and a discriminator. The generator is used to map a low-resolution image into a high-resolution image, and the discriminator is used to help the generator generate a more realistic high-resolution image.
[0016] Further, the operations of the generator are as follows: The low-resolution image first enters a convolutional layer to extract shallow feature information, and then enters a group of dense residual blocks to extract deep feature information. After the extraction is completed, it enters another convolutional layer for processing, and a long residual connection is added to ensure the convergence and stability of training. Next, the extracted deep feature information is sent to an upsampling layer, where the low-resolution features are converted into high-resolution features. Then, the high-resolution features are sequentially sent to three convolutional layers for feature compression, and finally a high-resolution image is obtained.
[0017] Further, the operations of the discriminator are as follows: The high-resolution image generated by the generator is used as the input and sent to the discriminator for determination. First, shallow feature extraction is performed through a convolutional layer, and then more information is extracted in a deep feature extraction layer. Finally, the true / false probability of the image is obtained through a sigmoid function.
[0018] The present invention innovatively uses a super-resolution network to simulate the sub-pixel effect, and proposes a new process that combines a super-resolution network and a stereo matching algorithm to improve the ranging accuracy of stereo matching, thereby enhancing the safety of autonomous driving and assisted driving and ensuring the safety of drivers. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0020] Figure 1 It is the overall flowchart of the method for improving the ranging accuracy of stereo matching provided by the present invention;
[0021] Figure 2 It is the overall structure diagram of the method for improving the ranging accuracy of stereo matching provided by the present invention;
[0022] Figure 3 It is the structure diagram of the generator of the blind super-resolution network BSRGAN adopted by the present invention;
[0023] Figure 4 It is the structure diagram of the discriminator of the blind super-resolution network BSRGAN adopted by the present invention;
[0024] Figure 5The structural diagram of the stereo matching algorithm IGEV-Stereo adopted by the present invention;
[0025] Figure 6 The comparison chart of disparities before and after the blind super-resolution network processing;
[0026] Figure 7 The comparison chart of disparity errors before and after the blind super-resolution network processing; and
[0027] Figure 8 The structural schematic diagram of a binocular stereo matching ranging accuracy improvement device. Specific embodiments
[0028] The following will explain the embodiments of the present invention in detail with reference to the accompanying drawings. However, the present invention can be implemented in many different ways defined and covered by the claims.
[0029] First, some terms involved in the present invention are explained:
[0030] Stereo matching: Inferring the geometry of a three-dimensional scene from captured images is a fundamental task in computer vision and graphics, and its application scope includes three-dimensional reconstruction, robotics, and autonomous driving. Stereo matching is to reconstruct a dense three-dimensional representation from two images using calibrated cameras, and it is a key technology for reconstructing the geometry of a three-dimensional scene. The goal of stereo matching is to calculate the disparity d of each pixel in the reference image. Disparity refers to the horizontal displacement between a pair of corresponding pixels on the left and right images. For the pixel (x, y) in the left image, if its corresponding point is found at (x - d, y) in the right image, then the disparity of this pixel is d. If we obtain the disparity, we can calculate the depth value corresponding to this pixel through a formula. Therefore, the key to improving the accuracy of stereo matching lies in calculating the accurate disparity.
[0031] Super-resolution network: The image super-resolution network has very important application value in the field of computer vision. The essence of image super-resolution is to express the original object in a more detailed and dense manner. After the image is processed by a 2-fold super-resolution network, one pixel of the original image will be represented by four new pixels, and a certain object in the original image will be represented by 4 times as many pixel points as before. Essentially, these new pixels still represent the object in the original image, but only make the original object more refined. Currently, the super-resolution algorithm has developed to the blind super-resolution stage and is moving towards real-world super-resolution.
[0032] Blind Super-Resolution Network: Since there are many and complex degradations in the real world, it is not enough to simply use bicubic interpolation to simulate the super-resolution network trained from high-resolution to low-resolution images. Therefore, in order to better simulate the degradations in the real world and make the algorithm perform better on real-world images, the idea of blind super-resolution is proposed. This idea concatenates the input image with various degradation information (blur kernel, noise) and inputs it into the super-resolution model. We can understand the degradation model as a process where the image is first blurred by the blur kernel, then downsampled, and finally noise is added. The blind super-resolution network BSRGAN achieves very good results on different types of real degradation data by randomly permuting various blur kernels, downsamplings, and noise degradations.
[0033] The following combines with the legends to give a detailed explanation of the entire process of the present invention, the blind super-resolution network, and the binocular stereo matching algorithm:
[0034] Figure 1 It is the overall flowchart of the method for improving the ranging accuracy of stereo matching provided by the present invention.
[0035] See Figure 1 As shown, the present invention provides a method for improving the ranging accuracy of binocular stereo matching. The specific steps include:
[0036] S110: Obtain the left image and the right image captured by the binocular camera;
[0037] S120: Input the left image and the right image into a pre-trained deep learning super-resolution network to obtain high-resolution left and right images;
[0038] S130: Input the high-resolution left and right images into the stereo matching algorithm to generate a delicate and accurate disparity map; perform super-resolution magnification processing on the disparity map, that is, divide the disparity values uniformly by the magnification factor of super-resolution;
[0039] S140: According to the parameters of the camera, convert the processed disparity map into a depth map to obtain the real distance between the camera and the object.
[0040] Refer to Figure 2 , the implementation steps of the present invention are as follows:
[0041] Step 1, obtain binocular images.
[0042] The left image and the right image in the same scene can be obtained through the binocular camera. Through epipolar rectification, the same object corresponding to the left and right images is in the same pixel row, which is convenient for subsequent stereo matching.
[0043] Step 2, obtain high-resolution binocular images.
[0044] We input the binocular images into the blind super-resolution network, as shown in Figure 3 to obtain the high-resolution left and right images. The process is as follows:
[0045]
[0046] Among them, are the low-resolution left and right images, is the blind super-resolution algorithm, are the high-resolution left and right images.
[0047] Step 3, calculate the disparity map
[0048] We input the obtained high-resolution left and right images into the stereo matching algorithm IGEV-Stereo to obtain a dense and detailed disparity map. The process is as follows:
[0049]
[0050] Among them, is the stereo matching algorithm, is the high-resolution disparity map.
[0051] Step 4, process the disparity map
[0052] Divide the obtained disparity values by the magnification factor of the super-resolution network. Stereo matching is performed based on pixels. Therefore, after the image is processed by the blind super-resolution network, the matching step size will increase by N times, where N is the magnification factor of the super-resolution. As a result, the corresponding calculated disparity will also be magnified by N times. However, this does not correctly represent the original scene because our purpose is to calculate the disparity more accurately rather than change the disparity distribution. Therefore, we finally need to divide the disparity values by N. The process is as follows:
[0053]
[0054] Among them, is the disparity map that correctly represents the original scene, is the magnification factor.
[0055] Step 5, calculate the depth map
[0056] Calculate the depth map through the camera parameters. The camera parameters, focal length f and baseline B, are very important information. Through the camera parameters, we can establish the relationship between the disparity d and the depth D, and calculate the depth value corresponding to this pixel to achieve the ranging effect.
[0057]
[0058] Among them, is the distance that correctly represents the original scene.
[0059] Refer to Figure 3 And Figure 4 , the blind super-resolution network used in the present invention will be described in detail:
[0060] This network consists of a generator and a discriminator. Figure 3 represents the generator of the network. The generator is used to map low-resolution images to high-resolution images. Figure 4 represents the discriminator of the network. The discriminator is used to help the generator generate more realistic high-resolution images. That is to say, the generator has to deceive the discriminator, while the discriminator has to find a way to determine that the picture is generated by the generator. The two complement each other and promote each other.
[0061] According to Figure 3 , the main process of the generator is as follows: The low-resolution image first enters the convolutional layer to extract shallow feature information, and then enters a group of dense residual blocks to extract deep feature information. After the extraction is completed, it will enter another convolutional layer for processing, and at the same time, a long residual connection is added to ensure the convergence and stability of training. Next, the extracted deep feature information will be sent to the upsampling layer, where the low-resolution features will be converted into high-resolution features. Then, the high-resolution features will be sent to three convolutional layers in sequence for feature compression, and finally a high-resolution image will be obtained.
[0062] According to Figure 4 , the main process of the discriminator is as follows: The high-resolution image generated by the generator is used as the input and sent to the discriminator for determination. First, shallow feature extraction is performed through the convolutional layer, and then more information is extracted in the deep feature extraction layer. Finally, the true and false probability of the image is obtained through the sigmoid function.
[0063] If represents the original image, represents the blur kernel, represents downsampling, represents noise, then the definition formula of the general image degradation model is:
[0064]
[0065] The blind super-resolution network BSRGAN complicates (practicalizes) the training data set in terms of blurring, downsampling, and noise. Blurring: Two types of blurring are used, namely isotropic Gaussian blurring and anisotropic Gaussian blurring; Downsampling: Nearest neighbor interpolation, bilinear interpolation, bicubic interpolation, and size scaling; Noise: 3D Gaussian noise, JPEG noise, camera noise. Finally, the order of blurring, downsampling, and noise is randomly shuffled. Therefore, the generalization ability of this algorithm is stronger than that of general super-resolution networks.
[0066] In specific training, BSRGAN uses relatively large low-resolution patches with a size of 72 × 72, enabling the deep model to capture more information. The loss functions are: L1 loss, VGG perceptual loss, and spectral-norm-based least-squares PatchGAN loss, with weights of 1, 1, and 0.1 respectively. Finally, the optimizer Adam is used to train BSRGAN with a fixed learning rate of , a batch size of 48 per round, and finally the final model is obtained after training on a deep learning V100 graphics card for 10 days.
[0067] Refer to Figure 5 for a detailed description of the stereo matching algorithm IGEV-Stereo used in the present invention:
[0068] According to Figure 5 , the main process of this algorithm is as follows: After obtaining the left and right images through a binocular camera, the binocular images are sent to a 2D convolutional layer for feature extraction. Then, on the one hand, the extracted features are input into a 3D convolutional layer to calculate the disparity cost, thereby obtaining an initial disparity map and a geometric encoding volume. On the other hand, the extracted features are used to calculate the image correlation, and the calculated correlation is combined with the geometric encoding volume to obtain a combined encoding volume. After that, the combined geometric encoding volume and the initial disparity Figure 1 are sent to the disparity optimization layer for the final step of disparity optimization, and finally the final disparity map is obtained.
[0069] 2D convolutional layer. The main purpose of the 2D convolutional layer is to extract the feature information of binocular images to facilitate subsequent disparity cost calculation.
[0070] 3D convolutional layer. The 3D convolutional layer mainly constructs the extracted features into a disparity cost volume to facilitate subsequent fitting of the disparity map.
[0071] Combined geometric encoding volume. The combined geometric encoder is a combination of a geometric encoder and image correlation. Since only using the cost information of features lacks global geometric information, using the combined geometric encoder is more conducive to cost aggregation.
[0072] Disparity optimization layer. The disparity optimization layer starts from the initialized input disparity d, and the paper uses a combination of a convolutional neural network and a recurrent neural network to iteratively update the disparity.
[0073] The results of the present invention can be illustrated by the following experiments:
[0074] Experimental conditions:
[0075] The present invention runs on a windows11 system with 8GB of memory and a graphics calculator GTX 4060, using the software editor pycharm and the environment pytorch.
[0076] Experimental content and results:
[0077] In the experiment, the low-resolution left and right images are input into the blind super-resolution network BSRGAN to obtain the high-resolution left and right images magnified by two times. At this time, the obtained high-resolution left and right images are input into the stereo matching algorithm IGEV-Stereo to obtain the disparity map of this scene. Finally, the disparity map is divided by 2 to obtain a dense disparity map that correctly represents the original scene.
[0078] The dataset used in the experiment comes from the official Middlebury. This dataset contains left and right images and depth maps of different resolutions for the same scene. In this experiment, the low-resolution left and right images and the corresponding ground truth come from a scene with a resolution of 750×500, and the disparity ground truth corresponding to the super-resolved images comes from a scene with a resolution of 1500×1000.
[0079] According to Figure 6 , this figure describes the comparison diagram of the disparity effects of binocular stereo matching before and after using the blind super-resolution network BSRGAN:
[0080] From a visual perspective, the super-resolution algorithm improves the defective parts of the stereo matching algorithm and the originally calculated disparity map, and at the same time improves the ability of the stereo matching algorithm to capture details, improving the visual effect of the disparity.
[0081] According to Figure 7 , this figure describes the comparison diagram of the disparity error effects of binocular stereo matching before and after using the blind super-resolution network BSRGAN:
[0082] When the blind super-resolution network BSRGAN is combined with the binocular stereo matching IGEV-Stereo, the calculated disparity error is smaller, and the areas with poor original disparity calculation are improved through the blind super-resolution network.
[0083] According to the results of the Middlebury2014 test, compared with only using the stereo matching algorithm IGEV-Stereo, we found that using the combination of BSRGAN and the stereo matching algorithm IGEV-Stereo can reduce the relative error rate and the matching relative error rate, and improve the ranging accuracy of stereo matching.
[0084] Figure 8 It is a schematic diagram of the structure of a device for improving the ranging accuracy of binocular stereo matching.
[0085] See Figure 8As shown in the figure, a device for improving the ranging accuracy of binocular stereo matching includes: an acquisition module 80, configured to acquire a left image and a right image captured by a binocular camera; an input module 82, configured to input the left image and the right image into a pre-trained deep learning blind super-resolution network to obtain a high-resolution left image and a high-resolution right image; a calculation module 84, configured to input the high-resolution left image and the high-resolution right image into a stereo matching algorithm for calculation to generate a disparity map; and a processing module 86, configured to perform super-resolution multiplication processing on the disparity map, and based on the parameters of the camera, convert the processed disparity map into a depth map to obtain the true distance between the camera and the object.
[0086] Preferably, the input module 82 includes: a generator and a discriminator. The generator is used to map a low-resolution image to a high-resolution image, and the discriminator is used to help the generator generate a more realistic high-resolution image.
[0087] Preferably, the generator is configured to: the low-resolution image first enters a convolutional layer to extract shallow feature information, and then enters a group of dense residual blocks to extract deep feature information. After the extraction is completed, it will enter another convolutional layer for processing, and at the same time, a long residual connection is added to ensure the convergence and stability of the training. Then, the extracted deep feature information will be sent to an upsampling layer, where the low-resolution features will be converted into high-resolution features. Next, the high-resolution features will be sequentially sent to three convolutional layers for feature compression, and finally a high-resolution image is obtained.
[0088] Preferably, the discriminator is configured to: take the high-resolution image generated by the generator as input and send it to the discriminator for determination. First, shallow feature extraction is performed through a convolutional layer, and then more information is extracted in a deep feature extraction layer. Finally, the true and false probability of the image is obtained through a sigmoid function.
[0089] Only some exemplary embodiments of this embodiment have been described by way of illustration above. Undoubtedly, for those of ordinary skill in the art, without departing from the spirit and scope of the present invention, the described embodiments can be modified in various different ways. Therefore, the above drawings and descriptions are illustrative in nature and should not be construed as limiting the protection scope of the claims of the present invention.
Claims
1. A binocular stereo matching ranging accuracy improvement method, characterized in that: include: Get the left and right images taken by the binocular camera; Inputting the left image and the right image into a pre-trained deep learning blind super-resolution network to obtain a high-resolution left image and a high-resolution right image; Input the high-resolution left image and the high-resolution right image into the stereo matching algorithm for calculation to generate a disparity map; The disparity map is processed by super-resolution multiplication, and based on the parameters of the camera, the processed disparity map is converted into a depth map to obtain the real distance between the camera and the object.
2. The accuracy improvement method according to claim 1, characterized in that: The deep learning blind super-resolution network includes: Generator and discriminator: the generator is used to map low-resolution images to high-resolution images, and the discriminator is used to help the generator generate more realistic high-resolution images.
3. The accuracy improvement method according to claim 2, characterized in that: The generator process is: The low-resolution image first enters the convolution layer to extract shallow feature information, and then enters groups of dense residual blocks to extract deep feature information. After the extraction is completed, it will enter a convolution layer for processing, and a long residual connection is added to ensure the convergence and stability of the training. The extracted deep feature information will be sent to the upsampling layer. At this time, the low-resolution features will be converted into high-resolution features. Next, the high-resolution features will be sent to three convolution layers in turn for feature compression, and finally a high-resolution image will be obtained.
4. The accuracy improvement method according to claim 3, characterized in that: The process of the discriminator is as follows: The high-resolution image generated by the generator is sent as input to the discriminator for judgment. First, the shallow feature extraction is performed through the convolution layer, and then it enters the deep feature extraction layer to extract more information. Finally, the true or false probability of the image is obtained through the sigmiod function.
5. The accuracy improvement method according to claim 4, characterized in that: The step of inputting the high-resolution left image and the high-resolution right image into the stereo matching algorithm for calculation to generate a delicate and accurate disparity map includes: Send the high-resolution left image and high-resolution right image to the 2D convolution layer for feature extraction; The extracted features will be input into the 3D convolution layer to calculate the disparity cost, thereby obtaining the initial disparity map and geometric code body. On the other hand, the extracted features will calculate the correlation of the image, and the calculated correlation will be combined with the geometric code body to obtain the combined code body. The combined geometric coding experience is sent to the disparity optimization layer together with the initial disparity map for the final step of disparity optimization to obtain the disparity map.
6. The accuracy improvement method according to claim 5, wherein the super-resolution multiplication of the disparity map comprises: The values of the disparity map calculated by the stereo matching algorithm are uniformly divided by the magnification of the image, and the calculated disparity is twice the original disparity.
7. A binocular stereo matching ranging accuracy improvement device, characterized in that: include: An acquisition module is used to acquire the left image and the right image taken by the binocular camera; An input module, used to input the left image and the right image into a pre-trained deep learning blind super-resolution network to obtain a high-resolution left image and a high-resolution right image; A calculation module, used for inputting the high-resolution left image and the high-resolution right image into a stereo matching algorithm for calculation to generate a disparity map; The processing module is used to perform super-resolution multiplication processing on the disparity map, and based on the parameters of the camera, convert the processed disparity map into a depth map to obtain the real distance between the camera and the object.
8. The precision improvement device according to claim 7, characterized in that: The input module comprises: Generator and discriminator: the generator is used to map low-resolution images to high-resolution images, and the discriminator is used to help the generator generate more realistic high-resolution images.
9. The precision improvement device according to claim 8, characterized in that: The generator is used to: The low-resolution image first enters the convolution layer to extract shallow feature information, and then enters groups of dense residual blocks to extract deep feature information. After the extraction is completed, it will enter a convolution layer for processing, and a long residual connection is added to ensure the convergence and stability of the training. The extracted deep feature information will be sent to the upsampling layer. At this time, the low-resolution features will be converted into high-resolution features. Next, the high-resolution features will be sent to three convolution layers in turn for feature compression, and finally a high-resolution image will be obtained.
10. The precision improvement device according to claim 9, characterized in that: The discriminator is used to: The high-resolution image generated by the generator is sent as input to the discriminator for judgment. First, the shallow feature extraction is performed through the convolution layer, and then it enters the deep feature extraction layer to extract more information. Finally, the true or false probability of the image is obtained through the sigmiod function.