Image reconstruction method and device, computer equipment and storage medium
By performing channel transformation and feature upsampling on image features, combined with residual learning algorithms and multiple convolution functions, the problem of poor image reconstruction quality of neural network models is solved, and high-quality image reconstruction is achieved to meet the needs of machine vision tasks.
Patent Information
- Application Number
- CN202410648767.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-22
- Filing Date
- 2024-05-23
- Publication Date
- 2025-09-23
AI Technical Summary
Existing neural network models rely on the feature similarity of the training sample space and the test sample space during image reconstruction, resulting in poor image reconstruction quality and making it difficult to meet the requirements of machine vision tasks.
By performing channel transformation and feature upsampling on the target image features, combined with residual learning algorithms and multiple convolution functions, the image reconstruction quality is improved to meet the needs of human eye review.
It improves the image reconstruction quality, can restore images that meet machine vision requirements, provides an accurate image data basis for intelligent tasks, and improves the output quality of reconstructed images.
Smart Images

Figure CN120689224A_ABST
Abstract
Description
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on March 22, 2024, with application number 2024103368324 and application name “Image reconstruction method, device, computer equipment and storage medium”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of image reconstruction technology, and in particular to an image reconstruction method, apparatus, computer equipment, storage medium, and computer program product. Background Art
[0003] With the growing application of machine learning, numerous intelligent platforms have been adopted in areas such as the Internet of Vehicles, video surveillance, and smart cities. These platforms generate massive amounts of data communication between numerous sensors. This surge in data volume has led to the inefficiency of previous encoding methods designed for human vision, which also struggled to meet real-world requirements in terms of latency and scale. Therefore, feature encoding for intelligent machines has become increasingly important.
[0004] The neural network model in related technologies mainly learns network weights through a large amount of training data and then performs online model inference. It essentially relies on the feature similarity of the training sample space and the test sample space and the generalization ability of the model, resulting in poor quality of image reconstruction. Summary of the Invention
[0005] Based on this, it is necessary to provide an image reconstruction method, device, computer equipment, computer-readable storage medium and computer program product that can restore the original image as much as possible on the basis of meeting the machine vision task, meet the needs of human eye review, and improve the quality of image reconstruction.
[0006] In a first aspect, the present application provides an image reconstruction method. The method comprises:
[0007] Performing channel transformation processing on the target image feature to obtain a first intermediate image feature;
[0008] Feature upsampling is performed based on the first intermediate image features, the second convolutional layer, and the second activation function to obtain a target reconstructed image.
[0009] In one embodiment, the target image feature is an original image feature.
[0010] In one embodiment, the method further comprises:
[0011] Based on a preset residual learning processing algorithm, a first convolutional layer, and a first activation function, residual learning processing is performed on the original image features to obtain target image features.
[0012] In one embodiment, performing feature upsampling based on the first intermediate image features, the second convolutional layer, and the second activation function to obtain a target reconstructed image includes:
[0013] Based on a preset residual learning processing algorithm, a first convolutional layer, and a first activation function, performing residual learning processing on the first intermediate image feature to obtain a second intermediate image feature;
[0014] Based on the second convolutional layer and the second activation function, feature upsampling processing is performed on the second intermediate image features to obtain a target reconstructed image.
[0015] In one embodiment, performing channel transformation processing on the target image feature to obtain the first intermediate image feature includes:
[0016] Channel transformation processing is performed on the target image feature through a two-dimensional convolution function and a target output channel value to obtain a first intermediate image feature whose channel is the target output channel value.
[0017] In one embodiment, performing feature upsampling based on the first intermediate image features, the second convolutional layer, and the second activation function to obtain a target reconstructed image includes:
[0018] performing feature upsampling processing on the first intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial reconstructed image;
[0019] The pixel points included in the initial reconstructed image are screened by using the first pixel value and the second pixel value to obtain a target reconstructed image.
[0020] In one embodiment, performing feature upsampling processing on the second intermediate image features based on the second convolutional layer and the second activation function to obtain a target reconstructed image includes:
[0021] performing feature upsampling processing on the second intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial reconstructed image;
[0022] The pixel points included in the initial reconstructed image are screened by using the first pixel value and the second pixel value to obtain a target reconstructed image.
[0023] In one embodiment, performing feature upsampling processing on the first intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial reconstructed image includes:
[0024] Based on a preset residual learning processing algorithm, a two-dimensional convolution function, an upsampling function, and a second activation function, feature upsampling processing is performed on the first intermediate image features to obtain an initial reconstructed image.
[0025] In one embodiment, performing feature upsampling processing on the second intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial reconstructed image includes:
[0026] performing convolution processing on the second intermediate image feature based on a first two-dimensional convolution function to obtain a first convolution result, and performing upsampling processing on the first convolution result using an upsampling function to obtain a first upsampling result;
[0027] Performing convolution processing on the first upsampling result based on a depthwise separable convolution function to obtain a second convolution result, and processing the second convolution result through a leaky activation function to obtain a first activation result;
[0028] Performing splicing processing on the first upsampling result and the first activation result to obtain a splicing feature;
[0029] Based on the grouped convolution function, the second two-dimensional convolution function, the leaky activation function and the initial two-dimensional convolution function, the splicing features are processed to obtain initial output features, and the splicing features are reconstructed by the upsampling function to obtain an initial reconstructed image.
[0030] In one embodiment, the size of the first two-dimensional convolution function is greater than the size of the second two-dimensional convolution function.
[0031] In one embodiment, performing feature upsampling processing on the second intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial reconstructed image includes:
[0032] performing upsampling processing on the second intermediate image feature based on an upsampling function to obtain a first upsampling result; performing convolution processing on the first upsampling result based on a third two-dimensional convolution function to obtain a first convolution result;
[0033] The first convolution result is upsampled based on an upsampling function to obtain a second upsampling result; and the second upsampling result is convolved based on a fourth two-dimensional convolution function to obtain an initial reconstructed image.
[0034] In one embodiment, the input channel value of the third two-dimensional convolution function is different from the input channel value of the fourth two-dimensional convolution function, and the output channel value of the third two-dimensional convolution function is different from the output channel value of the fourth two-dimensional convolution function.
[0035] In one embodiment, performing feature upsampling processing on the second intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial reconstructed image includes:
[0036] performing iterative processing on the second intermediate image feature a second target number of times based on the initial two-dimensional convolution function, the upsampling function, and the second activation function to obtain an iterative output feature;
[0037] The iterative output features are processed based on a two-dimensional convolution function to obtain an initial reconstructed image.
[0038] In one embodiment, the iterative processing of the second intermediate image feature a second target number of times based on the initial two-dimensional convolution function, the upsampling function, and the second activation function to obtain the iterative output feature includes:
[0039] For the i-th time, the i-th input feature is convolved by a two-dimensional convolution function to obtain a first convolution result, the first convolution result is upsampled by an upsampling function to obtain a first upsampling result, the first upsampling result is activated by an activation function to obtain the i-th output feature, and the i-th output feature is used as the input feature of the i+1-th iterative processing, the 1st input feature is the second intermediate image feature, and the output feature of the second target number is the iterative output feature.
[0040] In one embodiment, performing feature upsampling processing on the second intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial reconstructed image includes:
[0041] performing feature upsampling processing on the second intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain a target upsampling output result;
[0042] If the channel value of the up-sampled output result does not match the target channel value, a channel transformation process is performed on the up-sampled output result to obtain an initial reconstructed image.
[0043] In one embodiment, performing feature upsampling processing on the second intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain a target upsampling output result includes:
[0044] performing feature upsampling processing on the second intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial upsampling output result;
[0045] Based on a preset residual learning processing algorithm, residual learning processing is performed on the initial up-sampling output result to obtain a target up-sampling output result.
[0046] In one embodiment, filtering each pixel point included in the initial reconstructed image by using the first pixel value and the second pixel value to obtain the target reconstructed image includes:
[0047] Filtering each pixel point included in the initial reconstructed image by using the first pixel value and the second pixel value to obtain a first reconstructed image;
[0048] Performing format processing on the first reconstructed image to obtain a target reconstructed image.
[0049] In one embodiment, performing format processing on the first reconstructed image to obtain a target reconstructed image includes:
[0050] Performing edge clipping processing on the first reconstructed image based on a preset edge clipping algorithm to obtain a clipped reconstructed image;
[0051] Based on a preset format conversion algorithm, the format of the cropped reconstructed image is converted to obtain a target reconstructed image.
[0052] In one embodiment, the method further comprises:
[0053] Based on the preset residual learning processing algorithm, the first convolutional layer and the first activation function, residual learning processing is performed on the target object to obtain output features.
[0054] In one embodiment, the target object is a first intermediate image feature, the output feature is a second intermediate image feature, and the residual learning processing is performed on the target object based on a preset residual learning processing algorithm, a first convolutional layer, and a first activation function to obtain the output feature, including:
[0055] Based on the first convolutional layer and the first activation function, residual learning processing is performed on the first intermediate image feature to obtain a second intermediate image feature.
[0056] In one embodiment, the first convolution layer is a two-dimensional convolution function, and the performing residual learning processing on the first intermediate image feature based on the first convolution layer and the first activation function to obtain the second intermediate image feature includes:
[0057] Based on the two-dimensional convolution function and the activation function, the first intermediate image feature is iteratively processed a first target number of times to obtain an output feature.
[0058] In one embodiment, the method further comprises:
[0059] For the i-th time, the i-th input features are processed by residual learning through the two-dimensional convolution function and the first activation function to obtain the i-th residual learning result;
[0060] The residual learning result of the i-th iteration and the output feature of the i-th iteration are spliced to obtain a splicing result, and the splicing result is used as the output feature of the i-th iteration and as the input feature of the i+1-th iteration. The input feature of the 1st iteration is the first intermediate image feature, and the output feature of the first target number of times is the second intermediate image feature.
[0061] In one embodiment, the step size of the two-dimensional convolution function is configured to be 1, and the size of the two-dimensional convolution function is configured to be 3.
[0062] In one embodiment, performing residual learning processing on the first intermediate image feature based on the first convolutional layer and the first activation function to obtain the second intermediate image feature includes:
[0063] Performing residual learning processing on the first intermediate image feature through the target convolution function to obtain a first convolution result;
[0064] Processing the first convolution result by a two-dimensional convolution function to obtain a second convolution result;
[0065] The second convolution result is activated by a target activation function to obtain a second intermediate image feature.
[0066] In one embodiment, the target convolution function is a grouped convolution function, the grouping value of the grouped convolution function is determined based on the number of channels of the first intermediate image feature, and the target activation function is a LeakyReLU function.
[0067] In one embodiment, the target convolution function is a depthwise separable convolution function, the grouping value of the depthwise separable convolution function is determined based on the number of channels of the first intermediate image feature, and the target activation function is a LeakyReLU function.
[0068] In a second aspect, the present application further provides an image reconstruction device. The device comprises:
[0069] An image reconstruction device, comprising:
[0070] a channel transformation module, configured to perform channel transformation processing on the target image feature to obtain a first intermediate image feature;
[0071] A feature upsampling module is used to perform feature upsampling processing based on the first intermediate image features, the second convolutional layer and the second activation function to obtain a target reconstructed image.
[0072] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are performed:
[0073] Performing channel transformation processing on the target image feature to obtain a first intermediate image feature;
[0074] Feature upsampling is performed based on the first intermediate image features, the second convolutional layer, and the second activation function to obtain a target reconstructed image.
[0075] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the following steps:
[0076] Performing channel transformation processing on the target image feature to obtain a first intermediate image feature;
[0077] Feature upsampling is performed based on the first intermediate image features, the second convolutional layer, and the second activation function to obtain a target reconstructed image.
[0078] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the following steps:
[0079] Performing channel transformation processing on the target image feature to obtain a first intermediate image feature;
[0080] Feature upsampling is performed based on the first intermediate image features, the second convolutional layer, and the second activation function to obtain a target reconstructed image.
[0081] The above-mentioned image reconstruction method, apparatus, computer device, storage medium, and computer program product include: performing channel transformation processing on target image features to obtain first intermediate image features; and performing feature upsampling processing based on the first intermediate image features, a second convolutional layer, and a second activation function to obtain a target reconstructed image. By employing this method, image features can be reconstructed to improve image reconstruction quality, meeting the requirements of human eye review, and images that meet machine vision requirements can be restored, providing an accurate image data foundation for intelligent image input tasks. The reconstructed image can be processed based on machine vision coding preprocessing in conjunction with a super-resolution network, further improving the output quality of the reconstructed image. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 is a schematic flow chart of an image reconstruction method in one embodiment;
[0083] Figure 2 is a schematic flow chart of an image reconstruction method in one embodiment;
[0084] Figure 3 is a schematic flow chart of an image reconstruction method in one embodiment;
[0085] Figure 4 is a schematic flow chart of an image reconstruction method in one embodiment;
[0086] Figure 5 is a schematic flow chart of an image reconstruction method in one embodiment;
[0087] Figure 6 is a schematic flow chart of an image reconstruction method in one embodiment;
[0088] Figure 7 A schematic flow chart of an image reconstruction method according to an embodiment;
[0089] Figure 8 is a schematic flow chart of an image reconstruction method in one embodiment;
[0090] Figure 9 is a schematic flow chart of an image reconstruction method in one embodiment;
[0091] Figure 10 A schematic flow chart of an image reconstruction method according to an embodiment;
[0092] Figure 11 A schematic flow chart of an image reconstruction method according to an embodiment;
[0093] Figure 12 is a schematic flow chart of an image reconstruction method in one embodiment;
[0094] Figure 13 is a schematic flow chart of an image reconstruction method in one embodiment;
[0095] Figure 14 is a schematic flow chart of an image reconstruction method in one embodiment;
[0096] Figure 15 is a structural block diagram of an image reconstruction device in one embodiment;
[0097] Figure 16 is a diagram of the internal structure of a computer device in one embodiment;
[0098] Figure 17 FIG. 4 is a flow chart of an image reconstruction method in one embodiment. DETAILED DESCRIPTION
[0099] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0100] In one embodiment, Figure 1 As shown, an image reconstruction method is provided. This embodiment uses the method applied to a terminal as an example for illustration. It is understandable that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The above-mentioned terminals can be, but are not limited to, various personal computers, laptops, smart phones, tablet computers, Internet of Things devices and portable wearable devices. Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server can be implemented as an independent server or a server cluster composed of multiple servers. In this embodiment, the image reconstruction method includes the following steps:
[0101] Step 102: Perform channel transformation processing on the target image feature to obtain a first intermediate image feature.
[0102] Specifically, the terminal may transform the channel values of the target image feature to obtain a first intermediate image feature after the channel value transformation. In other words, the terminal may convolve the target image using a two-dimensional convolution function and the channel values of the pre-configured output feature, and determine the convolution result as the first intermediate image feature. In one example, the size of the target image feature may be 128x112x56.
[0103] Step 104: Perform feature upsampling based on the first intermediate image features, the second convolutional layer, and the second activation function to obtain a target reconstructed image.
[0104] Among them, the second convolution layer can be a convolution module including multiple convolution functions, and the multiple convolution functions can be convolution functions of the same type, or convolution functions of different types, or convolution functions of the same type but with different configuration parameters, and so on.
[0105] Specifically, the terminal can perform feature upsampling processing based on the first intermediate image features, the second convolutional layer and the second activation function, and output the target reconstructed image.
[0106] In the above-mentioned image reconstruction method, channel transformation processing is performed on the target image features to obtain first intermediate image features; feature upsampling processing is performed based on the first intermediate image features, the second convolutional layer, and the second activation function to obtain the target reconstructed image. By adopting this method, image features can be reconstructed to improve image reconstruction quality, meet the needs of human eye review, and restore images that meet machine vision requirements, providing an accurate image data foundation for intelligent image input tasks. The reconstructed image can be processed based on machine vision coding preprocessing in combination with a super-resolution network, further improving the output quality of the reconstructed image.
[0107] In one embodiment, the target image feature is the original image feature.
[0108] Specifically, the original image features may be image features obtained after feature extraction processing is performed on the input image. That is, the terminal may obtain the input image, perform feature extraction on the input image, and obtain the original image features. The terminal may use the original image features as target image features, perform channel transformation processing on the target image features, and obtain the first intermediate image features.
[0109] In this embodiment, the terminal can directly perform channel transformation processing on the image features to ensure the efficiency of image processing.
[0110] In one embodiment, the method further comprises:
[0111] Based on the preset residual learning processing algorithm, the first convolutional layer and the first activation function, residual learning processing is performed on the original image features to obtain the target image features.
[0112] Among them, the preset residual learning processing algorithm can be a process of feature learning based on the first convolutional layer and the first activation function. The first convolutional layer can be a convolutional layer containing multiple identical convolution functions, or a convolutional layer containing multiple different convolution functions.
[0113] Specifically, after the terminal performs feature extraction processing on the target image, it obtains the original image features. The terminal can enhance the original image features based on a preset residual learning processing algorithm. In fact, it performs residual learning processing on the original image features to obtain the output result, and determines that the output result is the target image feature.
[0114] In this embodiment, the terminal performs residual learning on the image features and then performs channel transformation processing, and can achieve channel transformation after feature enhancement on the image features, thereby ensuring the effectiveness of the channel transformation.
[0115] In one embodiment, Figure 2As shown in FIG, the specific execution process of the step “performing feature upsampling processing based on the first intermediate image feature, the second convolutional layer, and the second activation function to obtain the target reconstructed image” includes:
[0116] In step 202 , based on a preset residual learning processing algorithm, a first convolutional layer, and a first activation function, residual learning processing is performed on the first intermediate image feature to obtain a second intermediate image feature.
[0117] The first convolutional layer may include multiple convolution functions, and the first activation function may include a common activation function or a leaky activation function, etc.
[0118] Specifically, the terminal can perform residual learning processing on the first intermediate image feature based on the preset residual learning processing algorithm, the first convolutional layer and the first activation function, and use the output result of the preset residual learning processing algorithm as the second intermediate image feature.
[0119] Step 204: perform feature upsampling processing on the second intermediate image features based on the second convolutional layer and the second activation function to obtain a target reconstructed image.
[0120] The second convolutional layer may include multiple convolution functions of the same type, or multiple convolution functions of different types, and the second activation function may be a common activation function or a leaky activation function.
[0121] Specifically, the terminal may perform feature upsampling processing on the second intermediate image features through a second convolutional layer and a second activation function to obtain a target reconstructed image.
[0122] In this embodiment, the terminal can perform residual learning after channel transformation and before feature upsampling, thereby ensuring the flexibility of residual learning, further improving the effectiveness of feature encoding, and improving the quality of the output reconstructed image.
[0123] In one embodiment, the step of “performing channel transformation processing on the target image feature to obtain the first intermediate image feature” may include:
[0124] Channel transformation processing is performed on the target image feature through a two-dimensional convolution function and a target output channel value to obtain a first intermediate image feature whose channel is the target output channel value.
[0125] The target output channel value may be a preconfigured channel value of an output feature, or a channel value after channel transformation that meets business scenario requirements based on an actual application scenario, or an output channel value determined based on the channel value of a target image feature. In one example, the target output channel value (NumChanne) may be 64. Specific parameters of the two-dimensional convolution function may be configured based on an actual application scenario.
[0126] For example, the two-dimensional convolution function can be Conv(c_in=128,c_out=64,s=1,k_ver=1,k_hor=1), where the number of channels (channel value) of the input feature of the two-dimensional convolution function can be 128, the channel value of the output feature can be 64, the convolution step can be 1, the vertical degree of the convolution kernel can be 1, the horizontal degree of the convolution kernel can be 1, and so on.
[0127] Specifically, the terminal can configure the specific parameters of the two-dimensional convolution function through the target output channel value to obtain the configured two-dimensional convolution function, and perform channel transformation processing on the target image feature based on the configured two-dimensional convolution function, that is, transform the number of channels of the target image feature to obtain a first intermediate image feature whose channel number is the pre-configured target output channel value.
[0128] In this embodiment, the number of channels can be configured by configuring the two-dimensional convolution function, thereby ensuring the flexibility of the channels and the convenience of channel transformation of image features. It can also further improve the effectiveness of the output image features and provide a reliable data basis for subsequent image reconstruction.
[0129] In one embodiment, Figure 3 As shown in FIG, the specific implementation process of the step of “performing feature upsampling processing based on the first intermediate image feature, the second convolutional layer, and the second activation function to obtain the target reconstructed image” includes:
[0130] Step 302 : Perform feature upsampling processing on the first intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial reconstructed image.
[0131] The second activation function may be a leaky activation function or a common activation function. The upsampling function may be a function for upsampling pixel features, such as nn.PixelShuffle(2);
[0132] Specifically, the terminal can perform feature upsampling processing on the first intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial reconstructed image.
[0133] Step 304 : Filter each pixel point included in the initial reconstructed image using the first pixel value and the second pixel value to obtain a target reconstructed image.
[0134] The first pixel value may be smaller than the second pixel value. In one example, the first pixel value may be 0 and the second pixel value may be 255.
[0135] Specifically, the terminal can filter the pixel points contained in the initial reconstructed image based on the first pixel value, the second pixel value and the pixel value of each pixel point contained in the initial reconstructed image. The terminal can compare the pixel value of each pixel point based on the first pixel value, and update the pixel value of the pixel point whose pixel value is less than the first pixel value to the first pixel value. The terminal can also compare the pixel value of each pixel point based on the second pixel value, and update the pixel value of the pixel point whose pixel value is greater than the second pixel value to the second pixel value. Based on this, the terminal can obtain the target reconstructed image.
[0136] In one example, the terminal can filter the initial reconstructed image based on a pixel value filtering function, a first pixel value, and a second pixel value. The pixel value filtering function can be clamp (first pixel value, second pixel value, initial reconstructed image), that is, clamp (0, 255, rec_Image0). The terminal can update the pixel values of pixel points in the initial reconstructed image whose pixel values are less than 0 to 0, and update the pixel values of pixel points in the initial reconstructed image whose pixel values are greater than 255 to 255. Based on this, the target reconstructed image is obtained.
[0137] In this embodiment, after performing feature upsampling on the image features, the terminal may update the initial reconstructed image based on the first pixel value and the second pixel value, remove invalid pixels, and improve the quality of the output image.
[0138] In one embodiment, Figure 4 As shown, the specific implementation process of the step of "performing feature upsampling processing on the second intermediate image features based on the second convolutional layer and the second activation function to obtain the target reconstructed image" may include:
[0139] Step 402 : Perform feature upsampling processing on the second intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial reconstructed image.
[0140] The second activation function may be a leaky activation function or a common activation function. The upsampling function may be a function for upsampling pixel features, such as nn.PixelShuffle(2);
[0141] Specifically, the terminal can perform feature upsampling processing on the second intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial reconstructed image.
[0142] Step 404 : Filter each pixel point included in the initial reconstructed image by using the first pixel value and the second pixel value to obtain a target reconstructed image.
[0143] The first pixel value may be smaller than the second pixel value. In one example, the first pixel value may be 0 and the second pixel value may be 255.
[0144] Specifically, the terminal can filter the pixel points contained in the initial reconstructed image based on the first pixel value, the second pixel value and the pixel value of each pixel point contained in the initial reconstructed image. The terminal can compare the pixel value of each pixel point based on the first pixel value, and update the pixel value of the pixel point whose pixel value is less than the first pixel value to the first pixel value. The terminal can also compare the pixel value of each pixel point based on the second pixel value, and update the pixel value of the pixel point whose pixel value is greater than the second pixel value to the second pixel value. Based on this, the terminal can obtain the target reconstructed image.
[0145] In one example, the terminal can filter the initial reconstructed image based on a pixel value filtering function, a first pixel value, and a second pixel value. The pixel value filtering function can be clamp (first pixel value, second pixel value, initial reconstructed image), that is, clamp (0, 255, rec_Image0). The terminal can update the pixel values of pixel points in the initial reconstructed image whose pixel values are less than 0 to 0, and update the pixel values of pixel points in the initial reconstructed image whose pixel values are greater than 255 to 255. Based on this, the target reconstructed image is obtained.
[0146] In this embodiment, after performing feature upsampling on the image features, the terminal may update the initial reconstructed image based on the first pixel value and the second pixel value, remove invalid pixels, and improve the quality of the output image.
[0147] In one embodiment, the specific implementation process of the step of "performing feature upsampling processing on the first intermediate image features based on the two-dimensional convolution function, the upsampling function, and the second activation function to obtain the initial reconstructed image" may include:
[0148] Based on a preset residual learning processing algorithm, a two-dimensional convolution function, an upsampling function, and a second activation function, feature upsampling processing is performed on the first intermediate image features to obtain an initial reconstructed image.
[0149] Specifically, the terminal can perform feature upsampling processing on the first intermediate image features based on a preset residual learning processing algorithm, a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial reconstructed image. In an example, the terminal can perform feature upsampling processing including multiple steps, and can perform residual learning processing based on a preset residual learning processing algorithm before and after one step or before and after multiple steps. This application does not specifically limit the specific location, specific steps, and specific number of times for performing residual learning. There may be multiple implementation methods, and this application cannot give examples one by one.
[0150] In this embodiment, before and after each step in the feature upsampling process, residual learning processing can be performed on the input data or output data based on a preset residual learning processing algorithm, thereby ensuring the flexibility of residual learning processing and further enhancing the data quality of the output feature upsampling results.
[0151] In one embodiment, Figure 5 As shown, the second activation function may include a leaky activation function, and the specific implementation process of the step of "performing feature upsampling processing on the second intermediate image features based on the two-dimensional convolution function, the upsampling function, and the second activation function to obtain an initial reconstructed image" may include:
[0152] In step 502 , a convolution process is performed on the second intermediate image feature based on the first two-dimensional convolution function to obtain a first convolution result, and an upsampling process is performed on the first convolution result through an upsampling function to obtain a first upsampling result.
[0153] The first two-dimensional convolution function and the second two-dimensional convolution function have different input channel numbers, and the input channel numbers are different. In one example, the first two-dimensional convolution function can be Conv(c_in=64,c_out=256,s=1,k_ver=3,k_hor=3), where the number of channels (channel values) of the input features of the two-dimensional convolution function can be 64, the channel value of the output features can be 256, the convolution stride can be 1, the vertical extent of the convolution kernel can be 3, the horizontal extent of the convolution kernel can be 3, and so on. The upsampling function can be nn.PixelShuffle(2).
[0154] Specifically, the terminal can perform two-dimensional convolution processing on the second intermediate image feature through a first two-dimensional convolution function to obtain the output result of the first two-dimensional convolution function, that is, the first convolution result; the terminal can also perform upsampling processing on the first convolution result through an upsampling function to obtain the first upsampling result output by the upsampling function.
[0155] In step 504 , the first upsampling result is convolved based on a depthwise separable convolution function to obtain a second convolution result, and the second convolution result is processed using a leaky activation function to obtain a first activation result.
[0156] Among them, the depth-wise separable convolution function can be Conv(c_in=64, c_out=64, s=1, k_ver=1, k_hor=1), the number of channels (channel value) of the input feature of the depth-wise separable convolution function can be 64, the channel value of the output feature can be 64, the convolution step can be 1, the vertical degree of the convolution kernel can be 1, and the horizontal degree of the convolution kernel can be 1.
[0157] Specifically, the terminal can convolve the first upsampling result through a depth-separable convolution function to obtain a second convolution result, and activate the second convolution result through a leaky activation function to obtain a first activation result. The activation result can be expressed as an image feature vector, the upsampling result can also be expressed as an image feature vector, and the convolution result can also be expressed as an image feature vector.
[0158] Step 506: perform splicing processing on the first upsampling result and the first activation result to obtain a splicing feature.
[0159] Specifically, the terminal can perform splicing processing on the first upsampling result and the first activation result to obtain a splicing feature. In one example, the terminal can perform sum processing on the first upsampling result and the first activation result to obtain a sum vector, and use the sum vector as the splicing feature.
[0160] In step 508 , the spliced features are processed based on the grouped convolution function, the second two-dimensional convolution function, the leaky activation function, and the initial two-dimensional convolution function to obtain initial output features, and the spliced features are reconstructed using an upsampling function to obtain an initial reconstructed image.
[0161] For example, the grouped convolution function is Conv(c_in=64,c_out=64,s=1,k_ver=3,k_hor=3,groups=64). The number of channels (channel value) of the input feature of the grouped convolution function can be 64, the channel value of the output feature can be 64, the convolution step can be 1, the vertical extent of the convolution kernel can be 3, the horizontal extent of the convolution kernel can be 3, and the group value can be 64.
[0162] The second two-dimensional convolution function can be Conv(c_in=64,c_out=64,s=1,k_ver=1,k_hor=1), where the number of channels (channel value) of the input feature of the two-dimensional convolution function can be 64, the channel value of the output feature can be 64, the convolution step can be 1, the vertical degree of the convolution kernel can be 1, the horizontal degree of the convolution kernel can be 1, and so on.
[0163] The initial two-dimensional convolution function can be Conv(c_in=64,c_out=12,s=1,k_ver=3,k_hor=3), where the number of channels (channel value) of the input feature of the two-dimensional convolution function can be 64, the channel value of the output feature can be 12, the convolution step can be 1, the vertical degree of the convolution kernel can be 3, and the horizontal degree of the convolution kernel can be 3.
[0164] Specifically, the terminal can convolve the splicing features based on the grouped convolution function to obtain a second convolution result, and convolve the second convolution result through a second two-dimensional convolution function to obtain a third convolution result. The terminal can activate the third convolution result through a leaky activation function to obtain an activation processing result. The terminal can convolve the activation processing result through an initial two-dimensional convolution function to obtain an initial output feature; based on this, the terminal can upsample the splicing features through an upsampling function to obtain an initial reconstruction result.
[0165] In one example, the terminal can upsample the splicing features through an upsampling function, and use the upsampling result output by the upsampling function as the initial reconstructed image. In another example, the terminal can also perform channel transformation on the upsampling result output by the upsampling function, and use the result after channel transformation as the initial reconstructed image. Alternatively, the terminal can perform residual learning on the upsampling result output by the upsampling function based on a preset residual learning processing algorithm, and then perform channel transformation, and use the result after channel transformation as the initial reconstructed image, and so on.
[0166] In this embodiment, the second intermediate image features are processed by a depthwise separable convolution function, a plurality of different two-dimensional convolution functions, and an activation function, thereby enriching the image feature processing method and improving the quality of image reconstruction.
[0167] In one embodiment, the size of the first two-dimensional convolution function is larger than the size of the second two-dimensional convolution function, and the vertical size of the convolution kernel of the first two-dimensional convolution function is larger than the vertical size of the convolution kernel of the second two-dimensional convolution function, and the horizontal size of the convolution kernel of the first two-dimensional convolution function is larger than the horizontal size of the convolution kernel of the second two-dimensional convolution function.
[0168] In one embodiment, Figure 6As shown, the specific implementation process of the step of "performing feature upsampling processing on the second intermediate image features based on the two-dimensional convolution function, the upsampling function, and the second activation function to obtain the initial reconstructed image" may include:
[0169] Step 602: Upsample the second intermediate image feature based on the upsampling function to obtain a first upsampled result. Convolve the first upsampled result based on a third two-dimensional convolution function to obtain a first convolution result.
[0170] The upsampling function may be nn.PixelShuffle(2). The third two-dimensional convolution function may be Conv(c_in=16,c_out=12,s=1,k_ver=1,k_hor=1), wherein the number of channels (channel value) of the input feature of the third two-dimensional convolution function may be 16, the channel value of the output feature may be 12, the convolution step may be 1, the vertical extent of the convolution kernel may be 1, and the horizontal extent of the convolution kernel may be 1;
[0171] Specifically, the terminal can perform upsampling processing on the second intermediate image feature through the upsampling function to obtain a first upsampling result, and convolve the first upsampling result through a third two-dimensional convolution function to obtain a first convolution result output by the third two-dimensional convolution function.
[0172] Step 604: Perform upsampling processing on the first convolution result based on the upsampling function to obtain a second upsampling result. Perform convolution processing on the second upsampling result based on a fourth two-dimensional convolution function to obtain an initial reconstructed image.
[0173] Among them, the fourth two-dimensional convolution function can be Conv(c_in=4,c_out=3,s=1,k_ver=1,k_hor=1), wherein the number of channels (channel value) of the input feature of the fourth two-dimensional convolution function can be 4, the channel value of the output feature can be 3, the convolution step can be 1, the vertical degree of the convolution kernel can be 1, and the horizontal degree of the convolution kernel can be 1.
[0174] Specifically, the terminal may perform upsampling processing on the first convolution result outputted in step 602 through an upsampling function to obtain a second upsampling result. Based on this, the terminal may perform convolution on the second upsampling result through a fourth two-dimensional convolution function to obtain an initial reconstructed image.
[0175] In this embodiment, the terminal can process the second intermediate image features through an upsampling function and multiple different two-dimensional convolution functions, further enriching the processing method of the second intermediate image features, improving processing efficiency, and ensuring the quality of the output reconstructed image.
[0176] In one embodiment, the input channel value of the third two-dimensional convolution function is different from the input channel value of the fourth two-dimensional convolution function, and the output channel value of the third two-dimensional convolution function is different from the output channel value of the fourth two-dimensional convolution function.
[0177] In one embodiment, Figure 7 As shown, the step of “performing feature upsampling processing on the second intermediate image features based on the two-dimensional convolution function, the upsampling function, and the second activation function to obtain an initial reconstructed image” includes:
[0178] Step 702 : Perform iterative processing on the second intermediate image feature for a second target number of times based on the initial two-dimensional convolution function, the upsampling function, and the second activation function to obtain an iterative output feature.
[0179] The second target number can be a preconfigured number of iterations, and the initial two-dimensional convolution function can be Conv(c_in=rc_in,c_out=rc_out,s=rs,k_ver=rk_ver,k_hor=rk_hor). The number of input feature channels, the channel values of the output features, the convolution step size, the vertical extent of the convolution kernel, and the horizontal extent of the convolution kernel of the initial two-dimensional convolution function can be configured based on the actual application scenario. The upsampling function can be nn.PixelShuffle(2), and the second activation function can be a common activation function nn.Relu(). In one example, the second target number can be log2(Scale) times, for example, Scale can be 4, etc.
[0180] Specifically, the terminal can iteratively process the second intermediate image features a second target number of times through a two-dimensional convolution function, an upsampling function, and a second activation function from front to back to obtain an iterative output feature.
[0181] Step 704: Process the iterative output features based on a two-dimensional convolution function to obtain an initial reconstructed image.
[0182] The two-dimensional convolution function may be a two-dimensional convolution function with a configured number of output channels, for example, Conv(c_in=rc_out,c_out=3,s=rs,k_ver=rk_ver,k_hor=rk_hor).
[0183] Specifically, the terminal may convolve the iterative output features through a two-dimensional convolution function with a configured number of output channels to obtain an initial reconstructed image.
[0184] In this embodiment, by iteratively processing the second intermediate image features, the flexibility of configuring the number of iterations and the effectiveness of feature data processing can be guaranteed, thereby improving the degree of adaptability to the human eye visual scene in the machine vision feature encoding scenario.
[0185] In one embodiment, the step of “iteratively processing the second intermediate image feature a second target number of times based on the initial two-dimensional convolution function, the upsampling function, and the second activation function to obtain an iterative output feature” includes:
[0186] For the i-th time, the i-th input feature is convolved by the initial two-dimensional convolution function to obtain the first convolution result, the first convolution result is upsampled by the upsampling function to obtain the first upsampling result, the first upsampling result is activated by the activation function to obtain the i-th output feature, and the i-th output feature is used as the input feature of the i+1-th iterative processing, the first input feature is the second intermediate image feature, and the output feature of the second target number is the iterative output feature.
[0187] The second target number may be log2(Scale) times.
[0188] Specifically, for the first iterative processing, the terminal can convolve the second intermediate image feature through the initial two-dimensional convolution function to obtain a first convolution result, and upsample the first convolution result through the upsampling function to obtain a first upsampling result. Based on this, the terminal can activate the first upsampling result through the ordinary activation function nn.Relu() to obtain the first output feature, and the terminal can use the first output feature as the input feature of the second iteration; similarly, the terminal can convolve the second input feature through the initial two-dimensional convolution function to obtain the first convolution result, and upsample the first convolution result through the upsampling function to obtain the first upsampling result, and process the first upsampling result in turn through the second activation function to obtain the output feature of the second iteration, and use the output feature of the second iteration as the input feature of the third iteration, until the log2(Scale)th iteration is performed, and the terminal can use the log2(Scale)th output feature as the iterative output feature.
[0189] In this embodiment, by iteratively processing image features, the efficiency of feature processing can be improved and the quality of the output reconstructed image can be improved.
[0190] In one embodiment, Figure 8 As shown, the step of “performing feature upsampling processing on the second intermediate image features based on the two-dimensional convolution function, the upsampling function, and the second activation function to obtain an initial reconstructed image” includes:
[0191] Step 802 : Perform feature upsampling processing on the second intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain a target upsampling output result.
[0192] Specifically, the terminal can perform feature upsampling processing on the second intermediate image features through a two-dimensional convolution function, an upsampling function, and a second activation function to obtain a target upsampling output result.
[0193] Step 804: If the channel value of the target up-sampling output result does not match the target channel value, a channel transformation process is performed on the target up-sampling output result to obtain an initial reconstructed image.
[0194] Specifically, the terminal can perform a matching judgment on the channel value of the target upsampling output result and the target channel value. If it is determined that the channel value of the current target upsampling output result is inconsistent with the target value, the terminal can perform channel transformation processing on the target upsampling output result to obtain an initial reconstructed image whose channel value is consistent with the target channel value. In one example, the target channel value can be 3.
[0195] In this embodiment, the channel values may be transformed before outputting the reconstructed image to ensure the consistency of the channel values of the output reconstructed image.
[0196] In one embodiment, Figure 9 As shown, the step of “performing feature upsampling processing on the second intermediate image features based on the two-dimensional convolution function, the upsampling function, and the second activation function to obtain a target upsampling output result” includes:
[0197] Step 902 : Perform feature upsampling processing on the second intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial upsampling output result.
[0198] Specifically, the terminal can perform feature upsampling processing on the second intermediate image features through a two-dimensional convolution function, an upsampling function, and a second activation function to obtain a target upsampling output result.
[0199] Step 904 : Based on a preset residual learning processing algorithm, perform residual learning processing on the initial up-sampling output result to obtain a target up-sampling output result.
[0200] Specifically, the terminal can perform residual learning processing on the initial upsampling output result through a preset residual learning processing algorithm, a first convolutional layer, and a first activation function to obtain a target upsampling output result.
[0201] In this embodiment, the terminal can re-perform residual learning processing on the result of feature upsampling output after feature upsampling, further enhancing the image features and also providing a reliable data basis for subsequent output reconstructed images.
[0202] In one embodiment, the step of “screening each pixel point included in the initial reconstructed image using the first pixel value and the second pixel value to obtain a target reconstructed image” includes:
[0203] The pixel points included in the initial reconstructed image are screened by using the first pixel value and the second pixel value to obtain a first reconstructed image, and the first reconstructed image is formatted to obtain a target reconstructed image.
[0204] The first pixel value may be smaller than the second pixel value. In one example, the first pixel value may be 0 and the second pixel value may be 255. The format processing may include edge clipping processing, format conversion processing, and the like.
[0205] Specifically, the terminal can screen the pixels included in the initial reconstructed image based on the first pixel value, the second pixel value, and the pixel value of each pixel included in the initial reconstructed image. The terminal can compare the pixel value of each pixel with the first pixel value, and update the pixel value of the pixel whose pixel value is less than the first pixel value to the first pixel value. The terminal can also compare the pixel value of each pixel with the second pixel value, and update the pixel value of the pixel whose pixel value is greater than the second pixel value to the second pixel value. Based on this, the terminal can obtain a first reconstructed image. The terminal can perform format processing on the first reconstructed image to obtain a target reconstructed image after format processing.
[0206] In one example, the terminal can filter the initial reconstructed image based on a pixel value filtering function, a first pixel value, and a second pixel value. The pixel value filtering function can be clamp (first pixel value, second pixel value, initial reconstructed image), that is, clamp (0, 255, rec_Image0). The terminal can update the pixel values of pixel points in the initial reconstructed image whose pixel values are less than 0 to 0, and update the pixel values of pixel points in the initial reconstructed image whose pixel values are greater than 255 to 255. Based on this, the target reconstructed image is obtained.
[0207] In this embodiment, the reconstructed image may be post-processed based on the machine vision coding pre-processing to further remove invalid areas in the reconstructed image, thereby further improving the quality of the output reconstructed image.
[0208] In one embodiment, Figure 10 As shown, the step of “performing format processing on the first reconstructed image to obtain a target reconstructed image” includes:
[0209] Step 1002 : Based on a preset edge clipping algorithm, edge clipping processing is performed on the first reconstructed image to obtain a clipped reconstructed image.
[0210] The preset edge clipping algorithm may be an algorithm for clipping image edges, an image segmentation algorithm, and the like.
[0211] Specifically, the terminal may perform a shearing process on the edge area of the first reconstructed image based on a preset edge shearing algorithm to obtain a sheared reconstructed image.
[0212] Step 1004 : performing format conversion on the cropped reconstructed image based on a preset format conversion algorithm to obtain a target reconstructed image.
[0213] The preset format conversion algorithm may be an algorithm for converting the image format of an image. For example, an image may include multiple image formats, such as png, jpg, etc. The preset format conversion algorithm may be an algorithm for converting between multiple image formats.
[0214] Specifically, the terminal may perform format conversion on the cropped reconstructed image based on a preset format conversion algorithm, and convert the cropped reconstructed image in a first format into a reconstructed image in a second format, that is, obtain a target reconstructed image.
[0215] In this embodiment, edge processing and format processing may be performed on the output image to further improve the quality of the output reconstructed image.
[0216] In one embodiment, Figure 11 As shown, the image reconstruction method also includes:
[0217] Step 1102: Based on a preset residual learning processing algorithm, a first convolutional layer, and a first activation function, residual learning processing is performed on the target object to obtain output features.
[0218] Among them, the target object can be an object for residual learning, and the target object can be one or more of the original image features, the first intermediate image features, the second intermediate image features, and the initial upsampling output results in the above embodiments. The process of the terminal performing residual learning processing on different target objects through the preset residual learning processing algorithm is similar, and the specific process of residual learning processing for each target object will not be repeated here. The following embodiments take the target object as the first intermediate image feature as an example for illustration. The process of residual learning processing for other types of target objects based on the preset residual learning processing algorithm is similar to that when the target object is the first intermediate image feature, and will not be repeated here.
[0219] Specifically, the terminal can perform residual learning processing on the target object through a preset residual learning processing algorithm, a first convolutional layer, and a first activation function to obtain output features output by the preset residual learning processing algorithm.
[0220] In this embodiment, the target object can be various types of image features, ensuring the flexibility of residual learning processing.
[0221] In one embodiment, the target object is the first intermediate image feature, and the output feature is the second intermediate image feature, such as Figure 12 As shown, step 1102 "performing residual learning processing on the target object based on the preset residual learning processing algorithm, the first convolutional layer and the first activation function to obtain output features" includes:
[0222] Step 1202: Perform residual learning processing on the first intermediate image feature based on the first convolutional layer and the first activation function to obtain a second intermediate image feature.
[0223] Specifically, the terminal can perform residual learning processing on the first intermediate image feature through the convolution function included in the first convolution layer and the first activation function to obtain the second intermediate image feature.
[0224] In one embodiment, the first convolution layer is a two-dimensional convolution function. Specifically, the two-dimensional convolution function can be Conv(c_in=rc_in,c_out=rc_out,s=rs,k_ver=rk_ver,k_hor=rk_hor). The specific configuration parameters of the two-dimensional convolution function can be determined based on the actual application scenario.
[0225] Step 1202, “performing residual learning processing on the first intermediate image feature based on the first convolutional layer and the first activation function to obtain a second intermediate image feature,” includes:
[0226] Based on a two-dimensional convolution function and an activation function, the first intermediate image feature is iteratively processed a first target number of times to obtain a second intermediate image feature.
[0227] The activation function may be a common activation function Relu, and the first target number may be configured based on an actual application scenario. In an example, the first target number may be 16.
[0228] Specifically, the terminal can perform a first target number of iterative processes on the first intermediate image feature through a two-dimensional convolution function and an activation function Relu from front to back to obtain a second intermediate image feature.
[0229] In this embodiment, the first intermediate image features may be iteratively processed based on a configurable two-dimensional convolution function to ensure the effectiveness of residual learning.
[0230] In one embodiment, Figure 13 As shown, the image reconstruction method also includes:
[0231] Step 1302: For the i-th time, perform residual learning processing on the i-th input feature through a two-dimensional convolution function and a first activation function to obtain the i-th residual learning result.
[0232] In step 1304, the residual learning result of the i-th iteration and the output feature of the i-th iteration are spliced to obtain a spliced result, and the spliced result is used as the output feature of the i-th iteration and as the input feature of the i+1-th iteration. The input feature of the 1st iteration is the first intermediate image feature, and the output feature of the first target number of times is the second intermediate image feature.
[0233] Specifically, for the first iterative processing, the first input feature input can be the first intermediate image feature; the terminal can first convolve the first intermediate image feature through a two-dimensional convolution function to obtain a first convolution result, and activate the first convolution result through the activation function Relu to obtain a first activation result. Based on this, the terminal can convolve the first activation result through a two-dimensional convolution function to obtain a first residual learning result. The first residual learning result is a feature resF with an output channel number of rc_out.
[0234] Based on this, the terminal can concatenate the residual learning results of the first iteration and the input features of the first iteration to obtain the concatenated result outF (outF=input+resF), and use the concatenated result as the output feature of the first iteration and the input feature of the second iteration.
[0235] For the second time, the second input feature input can be the output feature of the first iteration; the terminal can first convolve the output feature of the first iteration through a two-dimensional convolution function to obtain a first convolution result, and activate the first convolution result through the activation function Relu to obtain a first activation result. Based on this, the terminal can convolve the first activation result through a two-dimensional convolution function to obtain a second residual learning result. The second residual learning result is the feature resF with an output channel number of rc_out.
[0236] Based on this, the terminal can splice the residual learning results of the second iteration and the input features of the second iteration to obtain the splicing result outF (outF=input+resF), and use the splicing result as the output feature of the second iteration and the input feature of the third iteration until the first target number of iterations is performed. The terminal can use the output features of the first target number of iterations as the second intermediate image features.
[0237] In this embodiment, by iteratively processing image features, the efficiency of feature processing can be improved and the quality of the output reconstructed image can be improved.
[0238] In one embodiment, the step size of the two-dimensional convolution function is configured to 1, the size of the two-dimensional convolution function is configured to 3, the horizontal extent of the convolution kernel of the two-dimensional convolution function is configured to 3, and the vertical extent of the convolution kernel of the two-dimensional convolution is configured to 3.
[0239] In one embodiment, Figure 14 As shown, the step of “performing residual learning processing on the first intermediate image feature based on the first convolutional layer and the first activation function to obtain the second intermediate image feature” includes:
[0240] Step 1402: Perform residual learning processing on the first intermediate image feature through the target convolution function to obtain a first convolution result.
[0241] Specifically, the terminal may perform residual learning processing on the first intermediate image feature through the convolution function included in the target convolution function to obtain a first convolution result.
[0242] Step 1404: Process the first convolution result through a two-dimensional convolution function to obtain a second convolution result.
[0243] Among them, the two-dimensional convolution function can be Conv(c_in=64,c_out=64,s=1,k_ver=1,k_hor=1), wherein the number of channels (channel value) of the input feature of the two-dimensional convolution function can be 64, the channel value of the output feature can be 64, the convolution step size can be 1, the vertical degree of the convolution kernel can be 1, and the horizontal degree of the convolution kernel can be 1.
[0244] Alternatively, the two-dimensional convolution function can also be Conv(c_in=128,c_out=64,s=1,k_ver=1,k_hor=1), where the number of channels (channel value) of the input feature of the two-dimensional convolution function can be 128, the channel value of the output feature can be 64, the convolution step can be 1, the vertical degree of the convolution kernel can be 1, and the horizontal degree of the convolution kernel can be 1.
[0245] Specifically, the terminal may process the first convolution result through a two-dimensional convolution function to obtain a second convolution result.
[0246] Step 1406: Activate the second convolution result using a target activation function to obtain a second intermediate image feature.
[0247] Among them, the target activation function is the LeakyReLU function.
[0248] Specifically, the terminal can activate the second convolution result through the LeakyReLU function to obtain the second intermediate image feature.
[0249] In this embodiment, the first intermediate image features can be processed in sequence by the target convolution function, the two-dimensional convolution function and the activation function, and the residual learning of the image features can be realized through multiple processing methods to ensure the effectiveness of the residual learning and provide a reliable data basis for subsequent image reconstruction.
[0250] In one embodiment, the target convolution function is a grouped convolution function, the grouping value of the grouped convolution function is determined based on the number of channels of the first intermediate image feature, and the target activation function is a LeakyReLU function.
[0251] Specifically, the grouped convolution function can be Conv(c_in=64,c_out=64,s=1,k_ver=3,k_hor=3,groups=64), where the number of channels (channel values) of the input features of the two-dimensional convolution function can be 64, the channel value of the output feature can be 64, the convolution step can be 1, the vertical extent of the convolution kernel can be 3, the horizontal extent of the convolution kernel can be 3, and the number of groups can be 64.
[0252] When the target convolution function is a grouped convolution function, the two-dimensional convolution function can be Conv(c_in=64,c_out=64,s=1,k_ver=1,k_hor=1), where the number of channels (channel value) of the input feature of the two-dimensional convolution function can be 64, the channel value of the output feature can be 64, the convolution step can be 1, the vertical degree of the convolution kernel can be 1, and the horizontal degree of the convolution kernel can be 1.
[0253] In this embodiment, residual learning processing is performed on the first intermediate image features by using a grouped convolution function and a two-dimensional convolution function, which can improve the efficiency of residual learning.
[0254] In one embodiment, the target convolution function is a depthwise separable convolution function, the grouping value of the depthwise separable convolution function is determined based on the number of channels of the first intermediate image feature, and the target activation function is a LeakyReLU function.
[0255] Specifically, the depth-wise separable convolution function can be Conv(c_in=64,c_out=64,s=1,k_ver=3,k_hor=3,groups=64), the number of channels (channel value) of the input feature of the convolution function can be 64, the channel value of the output feature can be 64, the convolution step can be 1, the vertical degree of the convolution kernel can be 3, the horizontal degree of the convolution kernel can be 3, and the grouping value can be 64.
[0256] When the target convolution function is a grouped convolution function, the two-dimensional convolution function can be Conv(c_in=64,c_out=64,s=1,k_ver=1,k_hor=1), the number of channels (channel value) of the input feature of the two-dimensional convolution function can be 64, the channel value of the output feature can be 64, the convolution step can be 1, the vertical degree of the convolution kernel can be 1, and the horizontal degree of the convolution kernel can be 1.
[0257] In this embodiment, residual learning processing is performed on the first intermediate image features through a depthwise separable convolution function and a two-dimensional convolution function, which can improve the efficiency of residual learning.
[0258] The specific implementation process of the above image reconstruction method is described in detail below with reference to a specific embodiment:
[0259] This embodiment provides a method for reconstructing an image based on image features, performing image reconstruction on decoded and reconstructed machine vision features. The image reconstruction method provided in this embodiment may include a feature reconstruction image network, which includes at least a channel transformation module, a feature upsampling module, and may also include a feature learning module (residual learning module) and a post-processing module. The channel transformation module performs channel transformation by using a two-dimensional convolution function to align the input features; the feature learning module uses a residual learning module; the feature upsampling module may use a log2(N) upsampling module, a two-dimensional convolution feature-to-image transformation, threshold filtering, and other methods; and the post-processing module performs edge acquisition and format conversion on the features.
[0260] In one example, a feature reconstruction image network may include a channel transformation module, a feature learning module, a feature upsampling module, and a post-processing module.
[0261] Specifically, the input of the channel transformation process in the channel transformation module can be the image feature ImageFeature[N][:][:], and the output of the channel transformation module can be the intermediate feature warpF[NumChannel][:][:]. It includes:
[0262] Use two-dimensional convolution Conv(c_in=N,c_out=NumChannel,s=1,k_ver=3,k_hor=3).
[0263] Specifically, the feature learning module uses a residual learning module. The input of this process is the intermediate feature warpF[NumChannel][:][:] output by the channel transformation process, and the output is the intermediate feature outResF[NumChannel][:][:]. Unless otherwise specified, all two-dimensional convolution parameters are set according to the following default settings: rc_in is set to NumChannel; rc_out is set to NumChannel; rs is set to 1; rk_ver is set to 3; rk_hor is set to 3.
[0264] The residual learning process of the feature learning module may include steps 1 to 4, which are repeated M times:
[0265] Step 1: Take input as input and use the two-dimensional convolution function:
[0266] Conv(c_in=rc_in,c_out=rc_out,s=rs,k_ver=rk_ver,k_hor=rk_hor);
[0267] Step 2: Use activation function Relu
[0268] Step 3, use the two-dimensional convolution function:
[0269] Conv(c_in=rc_in,c_out=rc_out,s=rs,k_ver=rk_ver,k_hor=rk_hor), the output channel number is the feature resF of rc_out;
[0270] Step 4: Output the feature outF=input+resF with the number of output channels rc_out, which is used as the input feature for the next cycle.
[0271] The input of the feature upsampling module can be the intermediate feature outresF output by the feature learning module, and the output can be the intermediate feature outImage. Unless otherwise specified, all 2D convolution parameters are set as follows: rc_in is set to NumChannel; rc_out is set to 4*NumChannel; rs is set to 1; rk_ver is set to 3; and rk_hor is set to 3.
[0272] The execution process of the feature upsampling module includes the following steps:
[0273] Step 1. Loop log2(Scale) times and follow steps 1.1 to 1.3:
[0274] Step 1.1. Use the two-dimensional convolution function:
[0275] Conv(c_in=rc_in,c_out=rc_out,s=rs,k_ver=rk_ver,k_hor=rk_hor);
[0276] Step 1.2. Use nn.PixelShuffle(2);
[0277] Step 1.3. Use nn.Relu();
[0278] Step 2. Use the two-dimensional convolution function:
[0279] Conv(c_in=rc_out,c_out=3,s=rs,k_ver=rk_ver,k_hor=rk_hor), output reconstructed image rec_Image0;
[0280] Step 3. Use clamp(0,255,rec_Image0) to generate the reconstructed image rec_Image1;
[0281] Based on this, the terminal can sequentially perform edge clipping and format conversion through the post-processing module. The input of this process is rec_Image1, and the output is the final reconstructed image rec_Image.
[0282] In another example, the feature reconstruction image network includes a channel transformation module, a feature learning module, a feature upsampling module, and a post-processing module. Take the reconstruction image network input feature 128x112x56 as an example.
[0283] Specifically, the channel transformation module can be a channel transformation process in which the input is the image feature ImageFeature
[128] [:][:] and the output is the intermediate feature warpF
[64] [:][:], specifically the channel transformation is achieved through two-dimensional convolution Conv(c_in=128,c_out=64,s=1,k_ver=1,k_hor=1).
[0284] Specifically, the feature learning module can use a residual learning module. The input of this process is the intermediate feature warpF
[64] [:][:] output by the channel transformation process, and the output is the intermediate feature outResF
[64] [:][:]. Here, NumChannel is 64.
[0285] The residual learning process of the feature learning module may include the following steps:
[0286] Step 1. Take input as input and use the grouped convolution function:
[0287] Conv(c_in=64,c_out=64,s=1,k_ver=3,k_hor=3,groups=64);
[0288] Step 2. Use 2D convolution Conv(c_in=64,c_out=64,s=1,k_ver=1,k_hor=1)
[0289] Step 3. Use the activation function LeakyReLU
[0290] Specifically, the input of the feature upsampling module may be the intermediate feature outresF output by the feature learning module, and the output of the feature upsampling module may be the intermediate feature outImage.
[0291] The processing of this feature upsampling module includes the following steps:
[0292] Step 1. Use 2D convolution Conv(c_in=64,c_out=256,s=1,k_ver=3,k_hor=3)
[0293] Step 2. Use nn.PixelShuffle(2);
[0294] Step 3. Use depth-wise separable convolution Conv(c_in=64,c_out=64,s=1,k_ver=1,k_hor=1);
[0295] Step 4. Use the activation function LeakyReLU;
[0296] Step 5. Sum the output of step 2 and the output of step 4;
[0297] Step 6. Use the grouped convolution function:
[0298] Conv(c_in=64,c_out=64,s=1,k_ver=3,k_hor=3,groups=64);
[0299] Step 7. Use 2D convolution Conv(c_in=64,c_out=64,s=1,k_ver=1,k_hor=1);
[0300] Step 8. Use the activation function LeakyReLU;
[0301] Step 9. Use 2D convolution Conv(c_in=64,c_out=12,s=1,k_ver=3,k_hor=3)
[0302] Step 10. Use nn.PixelShuffle(2) to output the reconstructed image rec_Image0;
[0303] Step 11. Use clamp(0,255,rec_Image0) to generate the reconstructed image rec_Image1;
[0304] Specifically, the terminal can sequentially perform edge clipping and format conversion through the post-processing module. The input of this process is rec_Image1, and the output is the final reconstructed image (target reconstructed image) rec_Image.
[0305] In another example, the feature reconstruction image network includes a channel transformation module, a feature learning module, a feature upsampling module and a post-processing module, taking the reconstruction image network input feature 128x112x56 as an example.
[0306] Specifically, the channel transformation module can be a channel transformation process in which the input is the image feature ImageFeature
[128] [:][:] and the output is the intermediate feature warpF
[64] [:][:], specifically the channel transformation is achieved by the two-dimensional convolution function Conv(c_in=128,c_out=64,s=1,k_ver=1,k_hor=1).
[0307] Specifically, the feature learning module can use a residual learning module. The input of this process is the intermediate feature warpF
[64] [:][:] output by the channel transformation process, and the output is the intermediate feature outResF
[64] [:][:]. Here, NumChannel is 64.
[0308] The residual learning process of the feature learning module may include the following steps:
[0309] Step 1. Take input as input and use the depth-wise separable convolution function:
[0310] Conv(c_in=64,c_out=64,s=1,k_ver=3,k_hor=3,groups=64);
[0311] Step 2. Use 2D convolution Conv(c_in=64,c_out=64,s=1,k_ver=1,k_hor=1)
[0312] Step 3. Use the activation function LeakyReLU
[0313] Specifically, the input of the feature upsampling module can be the intermediate feature outresF output by the feature learning module, and the output of the feature upsampling module is the intermediate feature outImage.
[0314] The processing of this feature upsampling module includes the following steps:
[0315] Step 1. Use nn.PixelShuffle(2);
[0316] Step 2. Use 2D convolution Conv(c_in=16,c_out=12,s=1,k_ver=1,k_hor=1)
[0317] Step 3. Use nn.PixelShuffle(2),
[0318] Step 4. Use two-dimensional convolution Conv(c_in=4,c_out=3,s=1,k_ver=1,k_hor=1) to output the reconstructed image rec_Image0;
[0319] Step 5. Use clamp(0,255,rec_Image0) to generate the reconstructed image rec_Image1;
[0320] Specifically, the terminal can sequentially perform edge clipping and format conversion through the post-processing module. The input of this process is rec_Image1, and the output is the final reconstructed image (target reconstructed image) rec_Image.
[0321] In another example, the feature reconstruction image network includes a channel transformation module, a feature learning module, a feature upsampling module, and a post-processing module. Take the reconstruction image network input feature 128x112x56 as an example.
[0322] Specifically, the channel transformation module can be a channel transformation process in which the input is the image feature ImageFeature
[128] [:][:], the output is the intermediate feature warpF
[64] [:][:], and the channel transformation can be implemented by the terminal using two-dimensional convolution Conv(c_in=128,c_out=64,s=1,k_ver=3,k_hor=3).
[0323] Specifically, the feature learning module can adopt a residual learning module. The input of this process is the intermediate feature warpF
[64] [:][:] output by the channel transformation process, and the output is the intermediate feature outResF
[64] [:][:]. Here, NumChannel is 64. If not specified, the parameters of all two-dimensional convolutions are set according to the following default method: rc_in is set to NumChannel; rc_out is set to NumChannel; rs is set to 1; rk_ver is set to 3; and rk_hor is set to 3.
[0324] The residual learning process of the feature learning module may include steps 1 to 4, which are repeated 16 times:
[0325] Step 1. Take input as input and use the two-dimensional convolution function:
[0326] Conv(c_in=rc_in,c_out=rc_out,s=rs,k_ver=rk_ver,k_hor=rk_hor);
[0327] Step 2. Use activation function Relu
[0328] Step 3. Use the two-dimensional convolution function:
[0329] Conv(c_in=rc_in,c_out=rc_out,s=rs,k_ver=rk_ver,k_hor=rk_hor), the output channel number is the feature resF of rc_out;
[0330] Step 4. The feature outF = input + resF with the number of output channels rc_out is used as the input feature of the next cycle.
[0331] The input of the feature upsampling module is the intermediate features outresF output by the feature learning module, and the output is the intermediate features outImage. Here, NumChannel is set to 64 and Scale is set to 4. Unless otherwise specified, all 2D convolution parameters are set to the following defaults: rc_in is set to NumChannel; rc_out is set to 4*NumChannel; rs is set to 1; rk_ver is set to 3; and rk_hor is set to 3.
[0332] The execution process of the feature upsampling module includes the following steps:
[0333] Step 1. Loop log2(Scale)=log2(4)=2 times Steps 1.1 to 1.3:
[0334] Step 1.1. Use the two-dimensional convolution function:
[0335] Conv(c_in=rc_in,c_out=rc_out,s=rs,k_ver=rk_ver,k_hor=rk_hor);
[0336] Step 1.2. Use nn.PixelShuffle(2);
[0337] Step 1.3. Use nn.Relu();
[0338] Step 2. Use the two-dimensional convolution function:
[0339] Conv(c_in=rc_out,c_out=3,s=rs,k_ver=rk_ver,k_hor=rk_hor), output reconstructed image rec_Image0;
[0340] Step 3. Use clamp(0,255,rec_Image0) to generate the reconstructed image rec_Image1;
[0341] The post-processing module sequentially performs edge clipping and format conversion. The input of this process is rec_Image1, and the output is the final reconstructed image (target reconstructed image) rec_Image.
[0342] The specific implementation process of the above image reconstruction method is described in detail below with reference to a specific embodiment:
[0343] like Figure 17 As shown, before the step of "performing channel transformation processing on the target image feature to obtain a first intermediate image feature", the method further includes: step a, performing two-dimensional residual convolution processing on the input image feature to obtain a first residual result; and step b, performing two-dimensional weighted convolution on the first residual result to obtain a first weighted result. The input image feature is r[:C][:rH][:rW], where r is the reconstruction tensor, C is the reconstruction tensor channel, rH is the reconstruction tensor height, and rW is the reconstruction tensor width.
[0344] Specifically include:
[0345] Step a: perform two-dimensional residual convolution, which can be specifically determined by the following formula:
[0346] RT1[:C][:rH][:rW]=ResConv(C,C,1)(r[:C][:rH][:rW]), the first residual result is RT1[:C][:rH][:rW]
[0347] Step b: Perform two-dimensional weighted convolution processing to obtain a first weighted result. Specifically, it can be determined by the following formula:
[0348] RT2[:C][:rH][:rW]=MaskConv(C)(RT1[:C][:rH][:rW])
[0349] The step of "performing channel transformation processing on the target image feature to obtain the first intermediate image feature" may include: step c, performing two-dimensional convolution on the first weighted result to obtain the first intermediate image feature, which may be specifically determined by the following formula:
[0350] RT3[ :C / 2 ][ :rH ][ :rW ] = Conv( C, C / 2, 1, 3, 3 )( RT2[ :C ][ :rH ][ :rW ] )
[0351] The specific implementation process of the step of “performing feature upsampling processing based on the first intermediate image features, the second convolutional layer, and the second activation function to obtain a target reconstructed image” may include:
[0352] Step d: Perform two-dimensional convolution on the first intermediate image feature to obtain a first convolution result (which can be recorded as RT4[:C*2][:rH][:rW]);
[0353] RT4[:C*2][:rH][:rW]=Conv(C / 2,C*2,1,3,3)(RT3[:C / 2][:rH][:rW])
[0354] Step e: Super-reconstruct the first convolution result to obtain the first reconstructed result (which can be recorded as RT5[:C / 2][:rH*2][:rW*2]).
[0355] RT5[:C / 2][:rH*2][:rW*2]=Shuffle(2)(RT4[:C*2][:rH][:rW])
[0356] Step f, perform two-dimensional residual convolution on the first reorganized result to obtain the second residual result (RT6[:C / 2][:rH*2][:rW*2]):
[0357] RT6[:C / 2][:rH*2][:rW*2]=ResConv(C / 2,C / 2,1)(RT5[:C / 2][:rH*2][:rW*2])
[0358] Step g, perform two-dimensional residual convolution on the second residual result to obtain the third residual result (which can be recorded as RT7[:C / 2][:rH*2][:rW*2]):
[0359] RT7[:C / 2][:rH*2][:rW*2]=ResConv(C / 2,C / 2,1)(RT6[:C / 2][:rH*2][:rW*2])
[0360] In step h, perform two-dimensional weighted convolution on the third residual result to obtain the second weighted result (RT8[:C / 2][:rH*2][:rW*2]):
[0361] RT8[:C / 2][:rH*2][:rW*2]=MaskConv(C / 2)(RT7[:C / 2][:rH*2][:rW*2])
[0362] Step i, perform two-dimensional residual convolution on the second weighted result to obtain the fourth residual result (RT9[:C / 2][:rH*2][:rW*2]):
[0363] RT9[:C / 2][:rH*2][:rW*2]=ResConv(C / 2,C / 2,1)(RT8[:C / 2][:rH*2][:rW*2])
[0364] In step j, perform two-dimensional residual convolution on the fourth residual result to obtain the fifth residual result (which can be recorded as RT10[:C / 2][:rH*2][:rW*2]):
[0365] RT10[:C / 2][:rH*2][:rW*2]=ResConv(C / 2,C / 2,1)(RT9[:C / 2][:rH*2][:rW*2])
[0366] Step k: perform tensor summation. For example, the first reorganization result and the fifth residual result may be summed to obtain a first sum vector (which can be expressed as RT11[:C / 2][:rH*2][:rW*2]):
[0367] RT11[:C / 2][:rH*2][:rW*2]=
[0368] Add(RT5[:C / 2][:rH*2][:rW*2],RT10[:C / 2][:rH*2][:rW*2])
[0369] Step 1: Perform two-dimensional convolution on the first sum vector to obtain the first convolution result (which can be recorded as RT12[:C*2][:rH*2][:rW*2]):
[0370] RT12[:C*2][:rH*2][:rW*2]=Conv(C / 2,C*2,1,3,3)(RT11[:C / 2][:rH*2][:rW*2])
[0371] Step m, perform super-recombination on the first convolution result to obtain the second recombined result (which can be recorded as RT13[:C / 2][:rH*4][:rW*4]):
[0372] RT13[:C / 2][:rH*4][:rW*4]=Shuffle(2)(RT12[:C*2][:rH*2][:rW*2])
[0373] Step n, perform two-dimensional residual convolution on the second reorganized result to obtain the sixth residual result (which can be recorded as RT14[:C / 2][:rH*4][:rW*4]):
[0374] RT14[:C / 2][:rH*4][:rW*4]=ResConv(C / 2,C / 2,1)(RT13[:C / 2][:rH*4][:rW*4])
[0375] In step o, perform two-dimensional convolution on the sixth residual result to obtain the second convolution result (which can be recorded as RT15[:3][:rH*4][:rW*4]):
[0376] RT15[:3][:rH*4][:rW*4]=Conv(C / 2,3,1,3,3)(RT14[:C / 2][:rH*4][:rW*4])
[0377] Step p: crop the second convolution result to obtain a cropped result.
[0378] Step q: performing format conversion on the cropping result. For example, when the value of rec_image_format_id is 0, 1, or 2, performing format conversion to obtain a format conversion result.
[0379] Step r, round and truncate the format conversion result to obtain the reconstructed image (target reconstructed image):
[0380] Specifically, based on the value of rec_image_format_id, a component determination strategy corresponding to the value can be determined. The component determination strategy is used to determine the luminance component im0[:riH][:riW], the chrominance Cr component im1[:(riH+1) / 2][:(riW+1) / 2], and the chrominance Cb component im2[:(riH+1) / 2][:(riW+1) / 2] of the target reconstructed image. The value may include 0, 1, 2, and 3. Those skilled in the art may configure the component determination strategy corresponding to each value based on actual application scenarios, and this embodiment does not limit this.
[0381] The specific process of performing the two-dimensional residual convolution processing may include, for example, performing the two-dimensional residual convolution processing through ResConv(c_in, c_out, type)(input, weight1, weight2, weight3, bias1, bias2, bias3), where:
[0382] c_in represents the number of channels of the input tensor; c_out represents the number of channels of the output tensor; type represents the type of ResConv, and its value is 0 or 1.
[0383] The input of ResConv includes: input tensor input[:c_in][:h][:w; convolution kernel weights weight1[:c_in][:3][:3], weight2[:c_out][:c_in][:1][:1], weight3[:c_out][:c_in][:1][:1]; bias bias1[c_in], bias2[c_out], bias3[c_out], and the default value of bias is 0.
[0384] The output of ResConv is: tensor output[:c_out][:h][:w].
[0385] Specifically, the parsing and decoding process loads the entire set of model parameters of the neural network. ResConv's weight1, weight2, and weight3, as well as bias1, bias2, and bias3 are part of the entire set of model parameters and no longer change during the parsing and decoding process. At this time, ResConv is simply expressed as ResConv(c_in, c_out, s, k_ver, k_hor)(input). Optionally, if the ResConv defined in this embodiment is called in the current operation and the input tensor is not explicitly set, the input tensor is set to the output tensor of the previous operation of the current operation in the series of operations by default. At this time, ResConv is simply expressed as ResConv(c_in, c_out, s, k_ver, k_hor).
[0386] ResConv performs the following steps in sequence:
[0387] In step 0, if the value of type is 1, the leaky activation function is used to process the input tensor to obtain the first activation result (which can be recorded as tmp0[:c_in][:h][:w], that is, the output of step 0). For example, it can be processed by the following formula:
[0388] tmp0[:c_in][:h][:w]=LeakyReLU(input[:c_in][:h][:w])
[0389] If the value of type is 0, let tmp0 = input;
[0390] Step 1: Perform a two-dimensional depthwise convolution on the output of step 0 to obtain the first depthwise convolution result (which can be recorded as tmp1[:c_in][:h][:w], which is the output of step 1):
[0391] tmp1[:c_in][:h][:w]=
[0392] DepthConv(c_in,1,3,3)(tmp0[:c_in][:h][:w],weight1[:c_in][:3][:3],bias1[c_in])
[0393] Step 2: Perform a two-dimensional convolution on the output of step 1 to obtain the output of step 2 (which can be recorded as tmp2[:c_out][:h][:w]):
[0394] tmp2[:c_out][:h][:w]=
[0395] Conv(c_in,c_out,1,1,1)(tmp1[:c_in][:h][:w],weight2[:c_out][:c_in][:1][:1],bias2[c_out])
[0396] Step 3, if the value of type is 1, set tmp3 = tmp2;
[0397] If the value of type is 0, the output of step 2 is activated using the leaky activation function to obtain the output of step 3 (which can be recorded as tmp3[:c_out][:h][:w]):
[0398] tmp3[:c_out][:h][:w]=LeakyReLU(tmp2[:c_out][:h][:w])
[0399] Step 4: For the output result of step 3, if c_in == c_out, then the output result of step 4 is determined to be tmp4[:c_out][:h][:w] = input[:c_in][:h][:w]; if c_in is not equal to c_out, the output result of step 4 can be calculated using the following formula:
[0400] tmp4[:c_out][:h][:w]=
[0401] Conv(c_in,c_out,1,1,1)(input[:c_in][:h][:w],weight3[:c_out][:c_in][:1][:1],bias3[c_out])
[0402] Step 5: Perform tensor summation. For example, the output result of step 3 and the output result of step 4 may be summed to obtain the output result of ResConv, i.e., the output tensor (output[:c_out][:h][:w]).
[0403] output[:c_out][:h][:w]=Add(tmp4[:c_out][:h][:w],tmp3[:c_out][:h][:w])
[0404] The specific process of performing the two-dimensional weighted convolution processing may include, for example, performing the two-dimensional weighted convolution processing through MaskConv(c)(input, weight1, weight2, bias1, bias2), wherein:
[0405] The input of MaskConv includes: input tensor input[:c][:h][:w]; convolution kernel weights weight1[:c][:c][:3][:3], weight2[:c][:c][:1][:1]; bias bias1[c], bias2[c], the default value of bias is 0.
[0406] The result of MaskConv is: output tensor output[:c][:h][:w].
[0407] Specifically, the parsing and decoding process loads the entire set of model parameters of the neural network. The weight and bias of MaskConv are part of the entire set of model parameters and no longer change during the parsing and decoding process. At this time, MaskConv is simply expressed as MaskConv(c)(input). Optionally, if the MaskConv defined in this embodiment is called in the current operation and the input tensor is not explicitly set, the input tensor is set to the output tensor of the previous operation of the current operation in the series of operations by default. At this time, MaskConv is simply expressed as MaskConv(c).
[0408] MaskConv performs the following steps in sequence:
[0409] In step 1, the input tensor is processed using the leaky activation function to obtain the first activation result (which can be recorded as tmp1[:c][:h][:w], which is the output of step 1). For example, it can be processed by the following formula:
[0410] tmp1[:c][:h][:w]=LeakyReLU(input[:c][:h][:w])
[0411] Step 2: Perform a two-dimensional depthwise convolution on the first activation result to obtain the first depthwise convolution result (i.e., the output of step 2, i.e., tmp2[:c][:h][:w]):
[0412] tmp2[:c][:h][:w]=DepthConv(c,1,3,3)(tmp1[:c][:h][:w],weight1[:c][:c][:3][:3],bias1[c])
[0413] Step 3: Perform a two-dimensional convolution on the first depthwise convolution result to obtain the output of step 3 (i.e., tmp3[:c][:h][:w]):
[0414] tmp3[:c][:h][:w]=Conv(c,c,1,1,1)(tmp2[:c][:h][:w],weight2[:c][:c][:1][:1],bias2[c])
[0415] Step 4: Perform jumper connection processing on the output result of step 3 and the input tensor to obtain the output result of step 4, that is, the output tensor of the MaskConv function (which can be recorded as output[i][j][k]):
[0416] for(i=0;i <c;i++){
[0417] for(j=0;j <h;j++){
[0418] for(k=0;k <w;k++){
[0419] output[i][j][k]=input[i][j][k]*(1+tmp3[i][j][k])
[0420] Specifically, the output result of step 4, i.e., the output tensor of the MaskConv function, can be determined based on the sum vector by calculating the product vector of the output result of step 3 and the input tensor, and calculating the sum vector of the product vector and the input tensor. The first target value can be 1.
[0421] The structure of the image reconstruction network involved in the image reconstruction method provided in this embodiment is relatively simple and has low complexity. It can also be decoupled from the feature encoding and decoding network to perform image reconstruction based on decoding features. In other words, the image reconstruction method provided in this embodiment can be decoupled from the machine vision feature encoding and decoding model, combine features with the super-resolution network, and post-process the reconstructed image based on machine vision encoding preprocessing, thereby further improving the degree of adaptability to the human eye visual scene under the machine vision feature encoding and decoding scenario.
[0422] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0423] Based on the same inventive concept, embodiments of the present application also provide an image reconstruction device for implementing the aforementioned image reconstruction method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations in one or more of the following image reconstruction device embodiments can be found in the above-described limitations on the image reconstruction method and will not be further elaborated here.
[0424] In one embodiment, Figure 15 As shown, an image reconstruction device 1500 is provided, comprising:
[0425] The channel transformation module 1502 is configured to perform channel transformation processing on the target image feature to obtain a first intermediate image feature.
[0426] The feature upsampling module 1504 is used to perform feature upsampling processing based on the first intermediate image features, the second convolutional layer and the second activation function to obtain a target reconstructed image.
[0427] In one embodiment, the target image feature is an original image feature.
[0428] In one embodiment, the apparatus further comprises:
[0429] The residual learning module is used to perform residual learning processing on the original image features based on a preset residual learning processing algorithm, a first convolutional layer and a first activation function to obtain target image features.
[0430] In one embodiment, the feature upsampling module specifically:
[0431] Based on a preset residual learning processing algorithm, a first convolutional layer, and a first activation function, performing residual learning processing on the first intermediate image feature to obtain a second intermediate image feature;
[0432] Based on the second convolutional layer and the second activation function, feature upsampling processing is performed on the second intermediate image features to obtain a target reconstructed image.
[0433] In one embodiment, the channel conversion module is specifically configured to:
[0434] Channel transformation processing is performed on the target image feature through a two-dimensional convolution function and a target output channel value to obtain a first intermediate image feature whose channel is the target output channel value.
[0435] In one embodiment, the feature upsampling module is specifically configured to:
[0436] performing feature upsampling processing on the first intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial reconstructed image;
[0437] The pixel points included in the initial reconstructed image are screened by using the first pixel value and the second pixel value to obtain a target reconstructed image.
[0438] In one embodiment, the feature upsampling module is specifically configured to:
[0439] performing feature upsampling processing on the second intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial reconstructed image;
[0440] The pixel points included in the initial reconstructed image are screened by using the first pixel value and the second pixel value to obtain a target reconstructed image.
[0441] In one embodiment, the feature upsampling module is specifically configured to:
[0442] Based on a preset residual learning processing algorithm, a two-dimensional convolution function, an upsampling function, and a second activation function, feature upsampling processing is performed on the first intermediate image features to obtain an initial reconstructed image.
[0443] In one embodiment, the feature upsampling module is specifically configured to:
[0444] performing convolution processing on the second intermediate image feature based on a first two-dimensional convolution function to obtain a first convolution result, and performing upsampling processing on the first convolution result using an upsampling function to obtain a first upsampling result;
[0445] Performing convolution processing on the first upsampling result based on a depthwise separable convolution function to obtain a second convolution result, and processing the second convolution result through a leaky activation function to obtain a first activation result;
[0446] Performing splicing processing on the first upsampling result and the first activation result to obtain a splicing feature;
[0447] Based on the grouped convolution function, the second two-dimensional convolution function, the leaky activation function and the initial two-dimensional convolution function, the splicing features are processed to obtain initial output features, and the splicing features are reconstructed by the upsampling function to obtain an initial reconstructed image.
[0448] In one embodiment, the size of the first two-dimensional convolution function is greater than the size of the second two-dimensional convolution function.
[0449] In one embodiment, the feature upsampling module is specifically configured to:
[0450] performing upsampling processing on the second intermediate image feature based on an upsampling function to obtain a first upsampling result; performing convolution processing on the first upsampling result based on a third two-dimensional convolution function to obtain a first convolution result;
[0451] The first convolution result is upsampled based on an upsampling function to obtain a second upsampling result; and the second upsampling result is convolved based on a fourth two-dimensional convolution function to obtain an initial reconstructed image.
[0452] In one embodiment, the input channel value of the third two-dimensional convolution function is different from the input channel value of the fourth two-dimensional convolution function, and the output channel value of the third two-dimensional convolution function is different from the output channel value of the fourth two-dimensional convolution function.
[0453] In one embodiment, the feature upsampling module is specifically configured to:
[0454] performing iterative processing on the second intermediate image feature a second target number of times based on the initial two-dimensional convolution function, the upsampling function, and the second activation function to obtain an iterative output feature;
[0455] The iterative output features are processed based on a two-dimensional convolution function to obtain an initial reconstructed image.
[0456] In one embodiment, the feature upsampling module is specifically configured to:
[0457] For the i-th time, the i-th input feature is convolved by the initial two-dimensional convolution function to obtain a first convolution result, the first convolution result is upsampled by the upsampling function to obtain a first upsampling result, the first upsampling result is activated by the activation function to obtain the i-th output feature, and the i-th output feature is used as the input feature of the i+1-th iterative processing, the 1st input feature is the second intermediate image feature, and the output feature of the second target number is the iterative output feature.
[0458] In one embodiment, the feature upsampling module is specifically configured to:
[0459] performing feature upsampling processing on the second intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain a target upsampling output result;
[0460] If the channel value of the target up-sampling output result does not match the target channel value, a channel transformation process is performed on the up-sampling output result to obtain an initial reconstructed image.
[0461] In one embodiment, the feature upsampling module specifically includes:
[0462] an upsampling unit, configured to perform feature upsampling processing on the second intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial upsampling output result;
[0463] The residual learning unit is used to perform residual learning processing on the initial up-sampling output result based on a preset residual learning processing algorithm to obtain a target up-sampling output result.
[0464] In one embodiment, the feature upsampling module specifically includes:
[0465] a pixel screening unit, configured to screen each pixel point included in the initial reconstructed image by using the first pixel value and the second pixel value to obtain a first reconstructed image;
[0466] A format processing unit is used to perform format processing on the first reconstructed image to obtain a target reconstructed image.
[0467] In one embodiment, the format processing unit is specifically configured to:
[0468] Performing edge clipping processing on the first reconstructed image based on a preset edge clipping algorithm to obtain a clipped reconstructed image;
[0469] Based on a preset format conversion algorithm, the format of the cropped reconstructed image is converted to obtain a target reconstructed image.
[0470] In one embodiment, the apparatus further comprises:
[0471] The residual learning module is used to perform residual learning processing on the target object based on a preset residual learning processing algorithm, a first convolutional layer, and a first activation function to obtain output features.
[0472] In one embodiment, the target object is a first intermediate image feature, the output feature is a second intermediate image feature, and the residual learning processing is performed on the target object based on a preset residual learning processing algorithm, a first convolutional layer, and a first activation function to obtain the output feature, including:
[0473] Based on the first convolutional layer and the first activation function, residual learning processing is performed on the first intermediate image feature to obtain a second intermediate image feature.
[0474] In one embodiment, the first convolution layer is a two-dimensional convolution function, and the residual learning module is specifically configured to:
[0475] Based on the two-dimensional convolution function and the activation function, the first intermediate image feature is iterated a first target number of times to obtain a second intermediate image feature.
[0476] In one embodiment, the residual learning module included in the apparatus is further specifically configured to:
[0477] For the i-th time, the i-th input features are processed by residual learning through the two-dimensional convolution function and the first activation function to obtain the i-th residual learning result;
[0478] The residual learning result of the i-th iteration and the output feature of the i-th iteration are spliced to obtain a splicing result, and the splicing result is used as the output feature of the i-th iteration and as the input feature of the i+1-th iteration. The input feature of the 1st iteration is the first intermediate image feature, and the output feature of the first target number of times is the second intermediate image feature.
[0479] In one embodiment, the step size of the two-dimensional convolution function is configured to be 1, and the size of the two-dimensional convolution function is configured to be 3.
[0480] In one embodiment, the residual learning module is specifically used to:
[0481] Performing residual learning processing on the first intermediate image feature through the target convolution function to obtain a first convolution result;
[0482] Processing the first convolution result by a two-dimensional convolution function to obtain a second convolution result;
[0483] The second convolution result is activated by a target activation function to obtain a second intermediate image feature.
[0484] In one embodiment, the target convolution function is a grouped convolution function, the grouping value of the grouped convolution function is determined based on the number of channels of the first intermediate image feature, and the target activation function is a LeakyReLU function.
[0485] In one embodiment, the target convolution function is a depthwise separable convolution function, the grouping value of the depthwise separable convolution function is determined based on the number of channels of the first intermediate image feature, and the target activation function is a LeakyReLU function.
[0486] Each module in the above-mentioned image reconstruction device may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0487] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 16 As shown. The computer device includes a processor, a memory, and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store image feature data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, an image reconstruction method is implemented.
[0488] Those skilled in the art will understand that Figure 16 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0489] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0490] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0491] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0492] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0493] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.
[0494] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0495] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. An image reconstruction method, characterized in that: The method comprises: Performing channel transformation processing on the target image feature to obtain a first intermediate image feature; Feature upsampling is performed based on the first intermediate image features, the second convolutional layer, and the second activation function to obtain a target reconstructed image.
2. The method according to claim 1, characterized in that The target image features are original image features.
3. The method according to claim 2, characterized in that The method further comprises: Based on a preset residual learning processing algorithm, a first convolutional layer, and a first activation function, residual learning processing is performed on the original image features to obtain target image features.
4. The method according to claim 1, wherein The performing feature upsampling processing based on the first intermediate image feature, the second convolutional layer, and the second activation function to obtain a target reconstructed image includes: Based on a preset residual learning processing algorithm, a first convolutional layer, and a first activation function, performing residual learning processing on the first intermediate image feature to obtain a second intermediate image feature; Based on the second convolutional layer and the second activation function, feature upsampling processing is performed on the second intermediate image features to obtain a target reconstructed image.
5. The method according to claim 1, wherein The performing channel transformation processing on the target image feature to obtain the first intermediate image feature includes: Channel transformation processing is performed on the target image feature through a two-dimensional convolution function and a target output channel value to obtain a first intermediate image feature whose channel is the target output channel value.
6. The method according to claim 1, characterized in that The performing feature upsampling processing based on the first intermediate image feature, the second convolutional layer, and the second activation function to obtain a target reconstructed image includes: performing feature upsampling processing on the first intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial reconstructed image; The pixel points included in the initial reconstructed image are screened by using the first pixel value and the second pixel value to obtain a target reconstructed image.
7. The method according to claim 4, characterized in that The performing feature upsampling processing on the second intermediate image features based on the second convolutional layer and the second activation function to obtain a target reconstructed image includes: performing feature upsampling processing on the second intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial reconstructed image; The pixel points included in the initial reconstructed image are screened by using the first pixel value and the second pixel value to obtain a target reconstructed image.
8. The method according to claim 6, characterized in that The performing feature upsampling processing on the first intermediate image features based on the two-dimensional convolution function, the upsampling function, and the second activation function to obtain an initial reconstructed image includes: Based on a preset residual learning processing algorithm, a two-dimensional convolution function, an upsampling function, and a second activation function, feature upsampling processing is performed on the first intermediate image features to obtain an initial reconstructed image.
9. The method according to claim 7, characterized in that The performing feature upsampling processing on the second intermediate image features based on the two-dimensional convolution function, the upsampling function, and the second activation function to obtain an initial reconstructed image includes: performing convolution processing on the second intermediate image feature based on a first two-dimensional convolution function to obtain a first convolution result, and performing upsampling processing on the first convolution result using an upsampling function to obtain a first upsampling result; Performing convolution processing on the first upsampling result based on a depthwise separable convolution function to obtain a second convolution result, and processing the second convolution result through a leaky activation function to obtain a first activation result; Performing splicing processing on the first upsampling result and the first activation result to obtain a splicing feature; Based on the grouped convolution function, the second two-dimensional convolution function, the leaky activation function and the initial two-dimensional convolution function, the splicing features are processed to obtain initial output features, and the splicing features are reconstructed by the upsampling function to obtain an initial reconstructed image.
10. The method according to claim 9, characterized in that The size of the first two-dimensional convolution function is greater than the size of the second two-dimensional convolution function.
11. The method according to claim 7, characterized in that The performing feature upsampling processing on the second intermediate image features based on the two-dimensional convolution function, the upsampling function, and the second activation function to obtain an initial reconstructed image includes: performing upsampling processing on the second intermediate image feature based on an upsampling function to obtain a first upsampling result; performing convolution processing on the first upsampling result based on a third two-dimensional convolution function to obtain a first convolution result; The first convolution result is upsampled based on an upsampling function to obtain a second upsampling result; and the second upsampling result is convolved based on a fourth two-dimensional convolution function to obtain an initial reconstructed image.
12. The method according to claim 11, characterized in that The input channel value of the third two-dimensional convolution function is different from the input channel value of the fourth two-dimensional convolution function, and the output channel value of the third two-dimensional convolution function is different from the output channel value of the fourth two-dimensional convolution function.
13. The method according to claim 7, characterized in that The performing feature upsampling processing on the second intermediate image features based on the two-dimensional convolution function, the upsampling function, and the second activation function to obtain an initial reconstructed image includes: performing iterative processing on the second intermediate image feature a second target number of times based on the initial two-dimensional convolution function, the upsampling function, and the second activation function to obtain an iterative output feature; The iterative output features are processed based on a two-dimensional convolution function to obtain an initial reconstructed image.
14. The method according to claim 13, characterized in that The iterative processing of the second intermediate image feature a second target number of times based on the initial two-dimensional convolution function, the upsampling function, and the second activation function to obtain an iterative output feature includes: For the i-th time, the i-th input feature is convolved by the initial two-dimensional convolution function to obtain a first convolution result, the first convolution result is upsampled by the upsampling function to obtain a first upsampling result, the first upsampling result is activated by the activation function to obtain the i-th output feature, and the i-th output feature is used as the input feature of the i+1-th iterative processing, the 1st input feature is the second intermediate image feature, and the output feature of the second target number is the iterative output feature.
15. The method according to claim 7, characterized in that The performing feature upsampling processing on the second intermediate image features based on the two-dimensional convolution function, the upsampling function, and the second activation function to obtain an initial reconstructed image includes: performing feature upsampling processing on the second intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain a target upsampling output result; If the channel value of the target up-sampling output result does not match the target channel value, a channel transformation process is performed on the up-sampling output result to obtain an initial reconstructed image.
16. The method according to claim 15, characterized in that The performing feature upsampling processing on the second intermediate image feature based on the two-dimensional convolution function, the upsampling function, and the second activation function to obtain a target upsampling output result includes: performing feature upsampling processing on the second intermediate image features based on a two-dimensional convolution function, an upsampling function, and a second activation function to obtain an initial upsampling output result; Based on a preset residual learning processing algorithm, residual learning processing is performed on the initial up-sampling output result to obtain a target up-sampling output result.
17. The method according to claim 6 or 7, characterized in that The step of filtering each pixel point included in the initial reconstructed image by using the first pixel value and the second pixel value to obtain a target reconstructed image includes: Filtering each pixel point included in the initial reconstructed image by using the first pixel value and the second pixel value to obtain a first reconstructed image; Performing format processing on the first reconstructed image to obtain a target reconstructed image.
18. The method according to claim 17, characterized in that The performing format processing on the first reconstructed image to obtain a target reconstructed image includes: Performing edge clipping processing on the first reconstructed image based on a preset edge clipping algorithm to obtain a clipped reconstructed image; Based on a preset format conversion algorithm, the format of the cropped reconstructed image is converted to obtain a target reconstructed image.
19. The method according to claim 4, characterized in that The method further comprises: Based on the preset residual learning processing algorithm, the first convolutional layer and the first activation function, residual learning processing is performed on the target object to obtain output features.
20. The method according to claim 19, characterized in that The target object is a first intermediate image feature, the output feature is a second intermediate image feature, and the residual learning processing is performed on the target object based on a preset residual learning processing algorithm, a first convolutional layer, and a first activation function to obtain the output feature, including: Based on the first convolutional layer and the first activation function, residual learning processing is performed on the first intermediate image feature to obtain a second intermediate image feature.
21. The method according to claim 20, characterized in that The first convolution layer is a two-dimensional convolution function, and the performing residual learning processing on the first intermediate image feature based on the first convolution layer and the first activation function to obtain the second intermediate image feature includes: Based on the two-dimensional convolution function and the activation function, the first intermediate image feature is iterated a first target number of times to obtain a second intermediate image feature.
22. The method according to claim 21, characterized in that The method further comprises: For the i-th time, the i-th input features are processed by residual learning through the two-dimensional convolution function and the first activation function to obtain the i-th residual learning result; The residual learning result of the i-th iteration and the output feature of the i-th iteration are spliced to obtain a splicing result, and the splicing result is used as the output feature of the i-th iteration and as the input feature of the i+1-th iteration. The input feature of the 1st iteration is the first intermediate image feature, and the output feature of the first target number of times is the second intermediate image feature.
23. The method according to claim 21, characterized in that The step size of the two-dimensional convolution function is configured as 1, and the size of the two-dimensional convolution function is configured as 3.
24. The method according to claim 20, characterized in that The performing residual learning processing on the first intermediate image feature based on the first convolutional layer and the first activation function to obtain the second intermediate image feature includes: Performing residual learning processing on the first intermediate image feature through the target convolution function to obtain a first convolution result; Processing the first convolution result by a two-dimensional convolution function to obtain a second convolution result; The second convolution result is activated by a target activation function to obtain a second intermediate image feature.
25. The method according to claim 24, characterized in that The target convolution function is a grouped convolution function, the grouping value of the grouped convolution function is determined based on the number of channels of the first intermediate image feature, and the target activation function is a LeakyReLU function.
26. The method according to claim 24, characterized in that The target convolution function is a depthwise separable convolution function, the grouping value of the depthwise separable convolution function is determined based on the number of channels of the first intermediate image feature, and the target activation function is a LeakyReLU function.
27. An image reconstruction device, characterized in that: The device comprises: a channel transformation module, configured to perform channel transformation processing on the target image feature to obtain a first intermediate image feature; A feature upsampling module is used to perform feature upsampling processing based on the first intermediate image features, the second convolutional layer and the second activation function to obtain a target reconstructed image.
28. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 26 are implemented.
29. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 26 are implemented.
30. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 26 are implemented.