Image super-division method and device and lightweight image super-division model

The low-resolution images are super-resolved by a lightweight image super-resolution model, which solves the problems of large computational complexity and long time consumption in the existing technology and achieves efficient image super-resolution effect, which is suitable for corn seedling detection tasks.

CN120634863APending Publication Date: 2025-09-12BEIJING QDING INTERCONNECTION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510927786.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing image super-resolution algorithms are difficult to strike a balance between super-resolution effect and super-resolution efficiency. They have large computational complexity and are time-consuming, which affects detection accuracy, especially in corn seedling detection tasks.

Method used

A lightweight image super-resolution model is adopted, including a first convolution module, a combined basic network module, a large separation convolution attention module, a second convolution module, a first pixel reconstruction module, a third convolution module and a feature map fusion module connected in sequence. The low-resolution image is super-resolved through the trained model to enhance the model's ability to capture long-distance dependencies and reduce the amount of computation.

Benefits of technology

Efficient image super-resolution is achieved, and the resolution of the super-reconstructed image is at least twice that of the image to be processed. It has low computational complexity and short time consumption, and is suitable for crop quantity detection scenarios such as corn seedling detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634863A_ABST
    Figure CN120634863A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image processing, and provides an image super-division method, an image super-division device and a lightweight image super-division model. The method comprises the steps that an image to be processed is input into a trained lightweight image super-division model, a super-division reconstructed image is output, and the resolution of the image to be processed is lower than that of the super-division reconstructed image; the lightweight image super-division model comprises a first convolution module, a combined basic network module, a large-scale separation convolution attention module, a second convolution module, a first pixel recombination module, a third convolution module and a feature map fusion module which are connected in sequence; the second pixel recombination module is connected with the first convolution module; the fourth convolution module is connected with the second pixel recombination module; and the fourth convolution module is connected with the feature map fusion module. The method is good in image super-division effect, small in calculation amount, short in time consumption, high in super-division efficiency and suitable for crop number detection scenes (such as corn seedling detection scenes).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular to an image super-resolution method, device and lightweight image super-resolution model. Background Art

[0002] In the fields of computer vision, image and video processing, image super-resolution technology (also known as image super-resolution technology) is usually used to reconstruct corresponding high-resolution images from observed low-resolution images, thereby improving the clarity and resolution of the image.

[0003] For example, in target detection tasks (such as corn seedling detection tasks), the existing super-resolution algorithms with good super-resolution effects still have problems of large computational complexity and long time consumption, while the existing super-resolution algorithms with faster super-resolution inference speed have poor super-resolution effects. Summary of the Invention

[0004] In view of this, the embodiments of the present application provide an image super-resolution method, device and lightweight image super-resolution model to at least solve the problem that the super-resolution algorithm in the related art cannot take into account both image super-resolution efficiency and super-resolution effect.

[0005] A first aspect of the embodiments of the present application provides an image super-resolution method, comprising:

[0006] Input the image to be processed into a trained lightweight image super-resolution model and output a super-resolution reconstructed image, wherein the resolution of the image to be processed is lower than the resolution of the super-resolution reconstructed image;

[0007] The lightweight image super-resolution model includes: a first convolution module, a combined basic network module, a large separation convolution attention module, a second convolution module, a first pixel reconstruction module, a third convolution module and a feature map fusion module connected in sequence;

[0008] a second pixel reassembly module connected to the first convolution module, and a fourth convolution module connected to the second pixel reassembly module;

[0009] The fourth convolution module is connected to the feature map fusion module.

[0010] According to a second aspect of the embodiments of the present application, an image super-resolution apparatus is provided, wherein the image super-resolution apparatus is configured to:

[0011] Input the image to be processed into a trained lightweight image super-resolution model and output a super-resolution reconstructed image, wherein the resolution of the image to be processed is lower than the resolution of the super-resolution reconstructed image;

[0012] The lightweight image super-resolution model includes:

[0013] The first convolution module, the combined basic network module, the large separation convolution attention module, the second convolution module, the first pixel reconstruction module, the third convolution module and the feature map fusion module are connected in sequence;

[0014] A second pixel reconstruction module connected to the first convolution module, a fourth convolution module connected to the second pixel reconstruction module, and the fourth convolution module connected to the feature map fusion module.

[0015] A third aspect of the embodiments of the present application provides a lightweight image super-resolution model, including:

[0016] The first convolution module, the combined basic network module, the large separation convolution attention module, the second convolution module, the first pixel reconstruction module, the third convolution module and the feature map fusion module are connected in sequence;

[0017] a second pixel reassembly module connected to the first convolution module, and a fourth convolution module connected to the second pixel reassembly module;

[0018] The fourth convolution module is connected to the feature map fusion module.

[0019] According to a fourth aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.

[0020] According to a fifth aspect of the embodiments of the present application, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.

[0021] Compared with the prior art, the embodiments of the present application have at least the following beneficial effects: by using the trained lightweight image super-resolution model provided by the embodiments of the present application to perform image super-resolution processing on low-resolution images to be processed, not only is the super-resolution effect good, but the computational complexity is small, the time consumption is short, and the super-resolution efficiency is high, making it suitable for crop quantity detection scenarios (such as corn seedling detection scenarios). The resolution of the super-reconstructed image obtained after processing using the lightweight image super-resolution model of the embodiments of the present application is at least twice the resolution of the image to be processed. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0023] Figure 1This is a flow chart of an image super-resolution method provided in an embodiment of the present application;

[0024] Figure 2 Schematic diagram of the structure of a lightweight image super-resolution model provided in an embodiment of the present application;

[0025] Figure 3 This is a schematic diagram of the structure of a basic network module provided in an embodiment of the present application;

[0026] Figure 4 is a structural diagram of a first basic residual unit provided in an embodiment of the present application;

[0027] Figure 5 This is a structural diagram of a high-frequency extraction unit provided in an embodiment of the present application;

[0028] Figure 6 This is a schematic structural diagram of a downsampling unit provided in an embodiment of the present application;

[0029] Figure 7 Schematic diagram of the structure of a channel attention unit provided in an embodiment of the present application;

[0030] Figure 8 This is a schematic diagram of the structure of a large-scale separation convolution attention module provided in an embodiment of the present application;

[0031] Figure 9 This is a schematic diagram of the structure of a large-scale separation convolutional attention layer provided in an embodiment of the present application;

[0032] Figure 10 This is a structural diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0033] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0034] An image super-resolution method, device, and lightweight image super-resolution model according to an embodiment of the present application will be described in detail below with reference to the accompanying drawings.

[0035] Counting corn seedlings in a cornfield is crucial for agriculture. In this scenario, target detection is often used. However, this method requires the drone to be kept at a low altitude, as the corn seedlings appear very small in the image, image clarity deteriorates, and the seedlings become difficult to detect. If the drone flies too low, collecting image data for the entire field takes a long time, resulting in high labor and time costs. Furthermore, the amount of collected image data is large, and uploading it to the cloud takes a long time. Therefore, super-resolution algorithms are often used to improve image clarity and resolution to prevent the target detection algorithm from losing accuracy.

[0036] However, existing super-resolution algorithms that achieve good results still suffer from high computational complexity and long processing times, while existing super-resolution algorithms with faster inference speeds often produce poor results. In other words, existing super-resolution algorithms still struggle to strike a balance between super-resolution effectiveness and efficiency.

[0037] In view of this, an embodiment of the present application proposes an image super-resolution method, which uses a trained lightweight image super-resolution model designed in this application to perform super-resolution processing on low-resolution, poor-quality images to be processed, and can obtain high-resolution, good-quality super-resolution reconstructed images. The model has a fast super-resolution inference speed and high super-resolution efficiency.

[0038] Figure 1 This is a flowchart of an image super-resolution method provided in an embodiment of the present application. Figure 1 The image super-resolution method can be used in a terminal device or a server. The terminal device can be a smartphone, tablet computer, laptop computer, desktop computer, etc. The server can be a server that provides various services, specifically a single server or a server cluster consisting of multiple servers, such as a backend server or a cloud server.

[0039] See also Figure 1 , the image super-resolution method comprises the following steps:

[0040] Step S101: input the image to be processed into a trained lightweight image super-resolution model, and output a super-resolution reconstructed image.

[0041] The resolution of the image to be processed is lower than the resolution of the super-resolution reconstructed image.

[0042] The image to be processed is generally a low-resolution image. A low-resolution image refers to an image with a resolution lower than a resolution threshold, and is perceived by humans as having low clarity and poor quality.

[0043] The super-resolution reconstructed image is an image with higher resolution reconstructed from the image to be processed with lower resolution.

[0044] Super-resolution reconstructed images are generally high-resolution images. High-resolution images refer to images with a resolution higher than the resolution threshold, and are perceived by humans as images with high clarity and good quality.

[0045] Figure 2 This is a structural diagram of a lightweight image super-resolution model provided in an embodiment of the present application.

[0046] See also Figure 2 The lightweight image super-resolution model provided in the embodiment of the present application includes: a first convolution module 201, a combined basic network module 202, a large separation convolution attention module (LSKAModule) 203, a second convolution module 204, a first pixel reorganization module (Pixel Shuffle_1) 205, a third convolution module 206 and a feature map fusion module (add) 207 connected in sequence; a second pixel reorganization module (Pixel Shuffle_2) 208 connected to the first convolution module 201, and a fourth convolution module 210 connected to the second pixel reorganization module 209; the fourth convolution module 210 is connected to the feature map fusion module 207.

[0047] Among them, the first convolution module 201, the second convolution module 204, the third convolution module 206 and the fourth convolution module 210 can be convolution networks with the same convolution kernel size, for example, they can be full convolution networks with a convolution kernel size of 3×3 (which can be marked as Conv-3).

[0048] The main function of the first pixel reassembly module 205 and the second pixel reassembly module 208 is to obtain a high-resolution feature map by reassembling the low-resolution feature map through convolution and multi-channel. For example, if you want to magnify the original image by 3 times, you need to generate 3 pixels based on the original image. 2 = 9 feature maps of the same size, and then the 9 feature maps of the same size are pieced together into a ×3 large image, so that the high-resolution feature map can be recombined from the low-resolution feature map.

[0049] The main function of the combined basic network module 202 is to reduce the resolution of the feature map, thereby reducing the amount of calculation, increasing the speed of super-resolution inference, and further improving the super-resolution efficiency.

[0050] The main function of the large separation convolution attention module 203 is to enhance the model's ability to capture long-distance dependencies, thereby improving the image super-resolution effect.

[0051] The technical solution provided by the embodiments of this application uses the trained lightweight image super-resolution model provided by the embodiments of this application to perform image super-resolution processing on low-resolution images to be processed. This not only achieves good super-resolution results, but also reduces computational complexity, shortens processing time, and achieves high super-resolution efficiency. It is suitable for crop quantity detection scenarios (such as corn seedling detection scenarios). The resolution of the super-reconstructed image obtained after processing using the lightweight image super-resolution model of the embodiments of this application is at least twice the resolution of the image to be processed.

[0052] In some embodiments, inputting the image to be processed into a trained lightweight image super-resolution model and outputting a super-resolution reconstructed image includes:

[0053] Performing convolution processing on the image to be processed through a first convolution module to obtain a first convolution feature map;

[0054] The first convolution feature map is processed by combining basic network modules to obtain a reduced-resolution feature map;

[0055] The reduced-resolution feature map is processed by a large separable convolutional attention module to obtain an attention feature map;

[0056] Performing convolution processing on the attention feature map through the second convolution module to obtain a second convolution feature map;

[0057] Performing pixel reassembly processing on the second convolution feature map through a first pixel reassembly module to obtain a first pixel reassembly feature map;

[0058] Performing convolution processing on the first pixel reorganization feature map through a third convolution module to obtain a third convolution feature map;

[0059] Performing pixel reassembly processing on the first convolution feature map through a second pixel reassembly module to obtain a second pixel reassembly feature map;

[0060] performing convolution processing on the second pixel reorganization feature map through a fourth convolution module to obtain a fourth convolution feature map;

[0061] The third convolution feature map and the fourth convolution feature map are fused through the feature map fusion module to obtain a super-resolution reconstructed image.

[0062] As an example, see Figure 2First, the image to be processed (such as a low-resolution, poor-quality cornfield image) is input into the first convolution module 201 (such as a convolution network with a convolution kernel size of 3×3) for convolution processing to obtain a first convolution feature map; then, the first convolution feature map is input into the combined basic network module 202 for processing to obtain a reduced-resolution feature map (the spatial resolution of the reduced-resolution feature map is lower than the spatial resolution of the first convolution feature map); then, the reduced-resolution feature map is input into the large-scale separation convolution attention module (LSKAModule) 203 for processing to obtain Attention feature map; then, the attention feature map is input into the second convolution module 204 (such as a convolution network with a convolution kernel size of 3×3) for convolution processing to obtain a second convolution feature map; thereafter, the second convolution feature map is input into the first pixel reorganization module 205 for pixel reorganization processing (specifically, it can be through convolution operation and multi-channel reorganization operation) to obtain a first pixel reorganization feature map; then, the first pixel reorganization feature map is input into the third convolution module 206 (such as a convolution network with a convolution kernel size of 3×3) for convolution processing to obtain a third convolution feature map. The first convolution feature map is input into the second pixel reorganization module 208 for pixel reorganization processing to obtain a second pixel reorganization feature map, and then the second pixel reorganization feature map is input into the fourth convolution module 210 (such as a convolution network with a convolution kernel size of 3×3) for convolution processing to obtain a fourth convolution feature map. The third convolution feature map and the fourth convolution feature map are input into the feature map fusion module 207 for fusion (specifically, the third convolution feature map and the fourth convolution feature map are added together) to obtain a super-resolution reconstructed image (a high-resolution, good-quality cornfield image) and output it.

[0063] The resolution of the super-reconstructed image obtained after the above-mentioned image super-resolution processing is at least twice the resolution of the image to be processed. At the same time, the lightweight image super-resolution model designed in the embodiment of the application has a small amount of calculation, a short time consumption, a fast super-resolution inference speed, and a high super-resolution efficiency.

[0064] In some embodiments, see Figure 2 The combined basic network module 202 may include n basic network modules (BNMs), wherein the n basic network modules are connected in series, and n is an integer ≥ 3. For example, the combined basic network module 202 may include a basic network module 2021 (BNM 1), a basic network module 2021 (BNM 2), ..., a basic network module 202n (BNMn) connected in sequence, where n is an integer ≥ 3. Generally, the value of n is in the range of 3 to 8.

[0065] Figure 3 This is a schematic diagram of the structure of a basic network module provided by an embodiment of the present application. Figure 3 , each basic network module (BNM) provided in the embodiment of the present application includes: a first basic residual unit (Basicresidual unit 1, BRU_1), a high-frequency extraction unit (High-frequency extraction module, HEM), a second basic residual unit (Basic residual unit 2, BRU_2), a feature map splicing unit (Concat), a first channel adjustment unit, a channel attention unit (CALayer), a third basic residual unit (Basic residual unit 3, BRU_3) and a feature map fusion unit connected in sequence; each basic network module also includes: a downsampling unit connected to the first basic residual unit, a combined basic residual unit connected to the downsampling unit, a second channel adjustment unit connected to the combined basic residual unit, and an upsampling unit (Upsampling) connected to the second channel adjustment unit; the upsampling unit is connected to the feature map splicing unit; wherein the combined basic residual unit includes m fourth basic residual units, the m fourth basic residual units are in a serial structure, and m is an integer ≥3. Generally, the value of m is 3 to 8.

[0066] The first channel adjustment unit and the second channel adjustment unit are mainly used to adjust the number of channels of the feature map. The first channel adjustment unit and the second channel adjustment unit can be a combination of a convolutional network with the same convolution kernel size and the same activation function. For example, the first channel adjustment unit and the second channel adjustment unit can both be a combination of a convolutional network with a convolution kernel size of 1×1 and a ReLU activation function (which can be marked as Conv-ac-1).

[0067] The first basic residual unit, the second basic residual unit, the third basic residual unit, and the fourth basic residual unit may be residual units with the same structure or residual units with different structures.

[0068] As an example, combining Figure 2 and Figure 3 , assuming that the combined basic network module 202 includes three basic network modules (i.e., n=3), namely basic network module 2021, basic network module 2022 and basic network module 2023; then the first convolution feature map output from the first convolution module 201 can be input into the basic network module 2021, the basic network module 2022 and the basic network module 2023 in sequence for processing to obtain a reduced resolution feature map.

[0069] The following is a detailed explanation using the example of inputting the first convolutional feature map into the basic network module 2021 for resolution reduction processing. First, the first convolution feature map is sequentially input into the first basic residual unit of the basic network module 2021 for feature extraction to obtain a first feature map; then, the first feature map is sequentially input into the high-frequency extraction unit and the second basic residual unit for feature extraction to obtain a second feature map; at the same time, the first feature map is sequentially input into the downsampling unit, the combined basic residual unit, the second channel adjustment unit and the upsampling unit to obtain a third feature map; thereafter, the second feature map and the third feature map are input into the feature map splicing unit (Concat) for feature map splicing processing to obtain a spliced ​​feature map; the spliced ​​feature map is sequentially input into the first channel adjustment unit for channel number adjustment (for example, reducing the number of channels of the spliced ​​feature map by half) to obtain a channel-adjusted feature map; the channel-adjusted feature map is input into the channel attention unit (CALayer) to obtain a channel attention feature map; the channel attention feature map is input into the third basic residual unit for processing to obtain a fourth feature map; the fourth feature map and the first convolution feature map are input into the feature map fusion unit for feature map addition operation to obtain a first resolution feature map. The feature map addition operation refers to performing an addition operation on elements at corresponding positions of two or more feature maps with the same shape.

[0070] Next, the first resolution feature map is input into the basic network module 2022 for resolution reduction processing to obtain a second resolution feature map; the second resolution feature map is input into the basic network module 2023 for resolution reduction processing to obtain the reduced resolution feature map. The operation process is basically the same as the operation process of inputting the first convolution feature map into the basic network module 2021 for processing to obtain the first resolution feature map, and will not be repeated here. Among them, the resolution of the first resolution feature map is smaller than that of the second resolution feature map, and the resolution of the second resolution feature map is smaller than that of the reduced resolution feature map.

[0071] By combining the basic network modules, the resolution of the first convolutional feature map can be reduced, thereby reducing the amount of computation. Furthermore, by connecting a high-frequency extraction unit in series after the first basic residual unit, the high-frequency information of the feature map can be extracted and preserved, reducing the degradation of the reconstructed image quality caused by the reduction of the first feature map.

[0072] The above-mentioned channel attention unit (CALayer) is a channel attention network layer that can increase the weight of channels with high activation values.

[0073] Figure 4 This is a schematic diagram of the structure of a first basic residual unit provided in an embodiment of the present application. Figure 4The first basic residual unit provided in the embodiment of the present application includes: a first convolution activation layer, a second convolution activation layer and a first feature map fusion layer connected in sequence; wherein the convolution kernel sizes of the first convolution activation layer and the second convolution activation layer are different; the first feature map fusion layer is connected to the high-frequency extraction unit; and the first convolution activation layer is connected to the feature fusion module.

[0074] As an example, the first convolution activation layer can be a combination of a convolution network with a convolution kernel size of 1x1 and a relu activation function. The second convolution activation layer can be a combination of a convolution network with a convolution kernel size of 3x3 and a relu activation function.

[0075] As an example, combining Figure 3 and Figure 4 , the first convolution feature map is subjected to feature extraction through the first basic residual unit, and the process of obtaining the first feature map includes the following steps: the first convolution feature map is sequentially input into the first convolution activation layer (such as the combination of a convolution network with a convolution kernel size of 1x1 + a relu activation function) and the second convolution activation layer (such as the combination of a convolution network with a convolution kernel size of 3x3 + a relu activation function) for processing to obtain a convolution activation feature map; the convolution activation feature map and the first convolution feature map are input into the first feature map fusion layer for feature map addition operation to obtain the first feature map.

[0076] Figure 5 This is a schematic diagram of the structure of a high-frequency extraction unit provided in an embodiment of the present application. Figure 5 The high frequency extraction unit (HEM) may include: a first average pooling layer (avgpool_1), an upsampling layer, and a second feature map fusion layer connected in sequence; wherein the first average pooling layer is connected to the first feature map fusion layer; and the second feature map fusion layer is connected to the second basic residual unit.

[0077] As an example, combining Figure 4 and Figure 5 The step of extracting high-frequency information from the first feature map by the high-frequency extraction unit includes: inputting the first feature map (TL) output by the first feature map fusion layer of the first basic residual unit into the first average pooling layer of the high-frequency extraction unit for average pooling processing, and then inputting it into the upsampling layer for upsampling processing (specifically, bilinear interpolation upsampling processing can be used) to obtain average information (TU) of the first feature map (TL); then, inputting the average information (TU) and the first feature map (TL) into the second feature map fusion layer for feature map subtraction operation, and outputting high-frequency information. The feature map subtraction operation refers to subtracting the elements at corresponding positions of the two feature maps.

[0078] High-frequency information primarily encompasses areas of image variation that are most dramatic, reflecting features such as image details, edges, texture, and noise. From a frequency perspective, it corresponds to areas of the image where grayscale values ​​rapidly change. For example, in an image of a cornfield, the edges of the corn plants, the veins on the leaves, and the details of the corn cobs all constitute high-frequency information. This high-frequency information makes the image appear clearer and more vivid, revealing the specific shapes and subtle features of the objects.

[0079] By extracting the high-frequency information of the first feature map through the high-frequency extraction unit, the details and edge information of the feature map can be better preserved.

[0080] Figure 6 This is a schematic diagram of the structure of a downsampling unit provided in an embodiment of the present application. Figure 6 The downsampling unit provided in the embodiment of the present application can be a Haar wavelet transform downsampling unit, which includes: a two-dimensional discrete decomposition layer, a feature map splicing layer connected to the two-dimensional discrete decomposition layer, and a full convolution layer connected to the feature map splicing layer; the two-dimensional discrete decomposition layer is connected to the first basic residual unit; the full convolution layer is connected to the combined basic residual unit.

[0081] The two-dimensional discrete decomposition layer can be a two-dimensional discrete wavelet first-order decomposition network layer. The input of the two-dimensional discrete decomposition layer includes the first feature map (TL) output by the first feature map fusion layer of the first basic residual unit. The full convolution layer can be a convolutional network with a convolution kernel size of 1×1.

[0082] As an example, combining Figure 6 The downsampling process of the first feature map (TL) includes the following steps: the first feature map (of size h×w×c, where h represents the feature map height, w represents the feature map width, and c represents the number of feature map channels) is input into a two-dimensional discrete decomposition layer. Through two-dimensional discrete decomposition, it is decomposed into one low-frequency sub-map (LL sub-map) and three high-frequency sub-maps (LH sub-map, HL sub-map, and HH sub-map) of size h / 2×w / 2×c. This can obtain a downsampled lossless encoding of the first feature map. Afterwards, the LL sub-map, LH sub-map, HL sub-map, and HH sub-map are input into the feature map splicing layer for feature map splicing processing, and then input into a full convolutional layer (e.g., a convolutional network with a convolution kernel size of 1×1) for convolution learning of frequency domain features.

[0083] Feature map concatenation is the process of concatenating and merging multiple feature maps along a specific dimension. In convolutional neural networks, feature maps are typically tensors with multiple dimensions, most commonly height (h), width (w), and number of channels (c). Generally speaking, feature map concatenation is performed along the channel dimension, concatenating channels from different feature maps.

[0084] The LL sub-image is an approximate representation of the original image (i.e., the feature map (TL)), which is the wavelet coefficient obtained by convolution using low-pass wavelet filters in the horizontal and vertical directions.

[0085] The HL subimage is a horizontal detail subimage of the original image, which is used to highlight the singular characteristics of the image in the horizontal direction. It is a wavelet coefficient obtained by convolving a high-pass wavelet filter in the horizontal direction and then convolving a low-pass wavelet filter in the vertical direction.

[0086] The LH subimage is a vertical detail subimage of the original image, which is used to highlight the unique characteristics of the image in the vertical direction. It is a wavelet coefficient obtained by convolving a low-pass wavelet filter in the horizontal direction and then convolving a high-pass wavelet filter in the vertical direction.

[0087] The HH subimage is a diagonal detail subimage of the original image, which is used to highlight the diagonal edge characteristics of the image. It is a wavelet coefficient obtained by convolution using a high-pass wavelet filter in the horizontal and vertical directions.

[0088] Compared with the commonly used downsampling methods of pooling and strided convolution, the embodiment of the present application uses Haar wavelet transform downsampling to downsample the first feature map (TL), which can reduce the spatial resolution of the first feature map (TL), which is beneficial to reducing the amount of calculation, while retaining more feature information, which is beneficial to improving the image super-resolution effect.

[0089] Figure 7 This is a schematic diagram of the structure of a channel attention unit provided in an embodiment of the present application. Figure 7 The channel attention unit (CALayer) provided in the embodiment of the present application includes: a second average pooling layer, k first convolution activation layers connected in series (k is an integer ≥1), a second convolution activation layer and a third feature map fusion layer; wherein the first convolution activation layer and the second convolution activation layer use different activation functions; the third feature map fusion layer is connected to the third basic residual unit; the second average pooling layer is connected to the first channel adjustment unit.

[0090] Among them, the first convolution activation layer can be a convolution + relu activation function with a convolution kernel size of 1x1 (which can be marked as Conv-ac-1).

[0091] The second convolution activation layer can be a convolution with a convolution kernel size of 1x1 + sigmoid activation function (which can be marked as Conv-sigmoid-1).

[0092] Combine Figure 3 and Figure 7As an example, the process of extracting features from the channel-adjusted feature map output by the first channel adjustment unit through the channel attention unit (CALayer) is as follows: the channel-adjusted feature map is sequentially input into the second average pooling layer, k concatenated first convolution activation layers (k is an integer ≥ 1, for example, k = 2, 3, 4...), and the second convolution activation layer for feature extraction to obtain a fifth feature map; the fifth feature map and the channel-adjusted feature map are input into the third feature map fusion layer for feature map multiplication to obtain a channel attention feature map. Feature map multiplication refers to the multiplication of elements in corresponding positions of two or more feature maps.

[0093] In some embodiments, the large-scale separated convolutional attention module includes: an attention unit, a first feature map superposition unit, a feedforward unit, and a second feature map superposition unit connected in sequence; the attention unit is connected to the combined basic network module;

[0094] The feedforward unit includes: a first fully convolutional layer, a second fully convolutional layer, a first activation function layer, and a third fully convolutional layer connected in sequence; the convolution kernel sizes of the first fully convolutional layer and the second fully convolutional layer are different, and the convolution kernel sizes of the first fully convolutional layer and the third fully convolutional layer are the same;

[0095] The attention unit includes: the fourth fully convolutional layer, the second activation function layer, the large separated convolutional attention layer and the fifth fully convolutional layer connected in sequence; the convolution kernels of the fourth and fifth fully convolutional layers have the same size.

[0096] Figure 8 This is a schematic diagram of the structure of a large-scale separation convolution attention module provided by the embodiment of this application. Figure 8 The large separable kernel attention module (LSKAModule) provided in the embodiment of the present application includes: an attention unit (Attention), a first feature map superposition unit, a feed-forward network (FFN), and a second feature map superposition unit connected in sequence.

[0097] The first feature map superposition unit and the second feature map superposition unit may be feature map superposition units with the same structure or with different structures, and they are mainly used to perform addition operations on two or more feature maps.

[0098] The feed-forward unit (FFN) includes: a first fully convolutional layer (for example, a fully convolutional network with a convolution kernel size of 1×1), a second fully convolutional layer (for example, a fully convolutional network with a convolution kernel size of 3×3), a first activation function layer (GELU_1, for example, a GELU activation function), and a third fully convolutional layer (for example, a fully convolutional network with a convolution kernel size of 1×1), which are connected in sequence; the convolution kernel sizes of the first fully convolutional layer and the second fully convolutional layer are different, and the convolution kernel sizes of the first fully convolutional layer and the third fully convolutional layer are the same;

[0099] The attention unit includes: a fourth fully convolutional layer (for example, a fully convolutional network with a convolution kernel size of 1×1), a second activation function layer (GELU_2, for example, a GELU activation function), a large separable convolutional attention layer (LSKA) and a fifth fully convolutional layer (for example, a fully convolutional network with a convolution kernel size of 1×1) connected in sequence; the convolution kernel sizes of the fourth and fifth fully convolutional layers are the same.

[0100] As an example, combining Figure 2 and Figure 8 The process of extracting features from the reduced-resolution feature map output by the combined basic network module 202 through the large-scale separated convolutional attention module 203 includes the following steps: first, inputting the reduced-resolution feature map into the attention unit (Attention) for feature extraction to obtain a sixth feature map; multiplying the sixth feature map by the modulation parameter β to obtain a seventh feature map; adjusting the reduced-resolution feature map using the modulation parameter α to obtain an eighth feature map; inputting the seventh and eighth feature maps into the first feature map superposition unit for feature map addition to obtain a ninth feature map. Inputting the ninth feature map into the feedforward unit (FFN) for feature extraction to obtain a tenth feature map; multiplying the tenth feature map by the modulation parameter δ to obtain an eleventh feature map; adjusting the ninth feature map using the modulation parameter γ to obtain a twelfth feature map; inputting the eleventh and twelfth feature maps into the second feature map superposition unit for feature map addition to obtain an attention feature map. Among them, the above-mentioned γ, β, α, and δ are all modulation parameters that can be obtained through model training. These modulation parameters can control the style, texture, color and other attributes of the generated image.

[0101] As an example, the process of extracting features from a reduced-resolution feature map through an attention unit (Attention) includes the following steps: inputting the reduced-resolution feature map into the fourth full convolutional layer (for example, it can be a full convolutional network with a convolution kernel size of 1×1) and the second activation function layer (GELU_2, for example, it can be a GELU activation function) for processing in sequence to obtain a thirteenth feature map; inputting the thirteenth feature map into the large separated convolutional attention layer (LSKA) for feature extraction to obtain a fourteenth feature map; inputting the fourteenth feature map into the fifth full convolutional layer (for example, it can be a full convolutional network with a convolution kernel size of 1×1) for convolution processing to obtain a sixth feature map.

[0102] As an example, the process of extracting features from the ninth feature map through a feed-forward unit (FFN) includes the following steps: inputting the ninth feature map into the first full convolution layer (for example, it can be a full convolution network with a convolution kernel size of 1×1), the second full convolution layer (for example, it can be a full convolution network with a convolution kernel size of 3×3), the first activation function layer (GELU_1, for example, it can be a GELU activation function) and the third full convolution layer (for example, it can be a full convolution network with a convolution kernel size of 1×1) for feature extraction processing to obtain the tenth feature map.

[0103] The second fully convolutional layer of the feedforward unit (FFN) in the embodiment of the present application adopts a fully convolutional network with a convolution kernel size of 3×3, which can reduce information loss and thus help improve the super-resolution effect.

[0104] In some embodiments, the large separable convolutional attention layer includes: a first depthwise separable convolution sublayer, a second depthwise separable convolution sublayer, a first depthwise separable dilated convolution sublayer, a second depthwise separable dilated convolution sublayer, a full convolution sublayer, and a feature fusion sublayer connected in sequence;

[0105] Among them, the first depth-wise separable convolution sublayer is connected to the second activation function layer; the feature fusion sublayer is connected to the fifth full convolution layer.

[0106] Figure 9 This is a schematic diagram of the structure of a large separation convolutional attention layer provided in an embodiment of the present application. Figure 9The large separable kernel attention (LSKA) layer provided in an embodiment of the present application includes: a first depth-wise separable convolution sublayer (1×(2d-1)DW-Conv), a second depth-wise separable convolution sublayer ((2d-1)×1DW-Conv), a first depth-wise separable dilated convolution sublayer (1×[k / d]DW-D-Conv), a second depth-wise separable dilated convolution sublayer ([k / d]×1DW-D-Conv), a full convolution sublayer (1×1Conv) and a feature fusion sublayer connected in sequence; wherein the feature fusion sublayer is connected to the second activation function layer. Among them, 1×(2d-1)DW-Conv represents depthwise separable convolution with a kernel size of 1×(2d-1); (2d-1)×1DW-Conv represents depthwise separable convolution with a kernel size of (2d-1)×1; 1×[k / d]DW-D-Conv represents depthwise separable dilated convolution with a kernel size of 1×[k / d]; [k / d]×1DW-D-Conv represents depthwise separable dilated convolution with a kernel size of [k / d]×1; 1×1Conv represents full convolution with a kernel size of 1×1; k and d are parameters related to the convolution kernel size, and k>d.

[0107] The large-scale separation convolutional attention layer provided in the embodiments of this application reduces the computational complexity and memory usage of 2D large-core convolution by decomposing the traditional 2D convolution kernel into two cascaded 1D convolution kernels (first horizontally and then vertically), thus avoiding the high computational complexity caused by the large convolution kernels in the deep convolution layer of the large-core attention LSKA. This structure can capture long-range dependencies while requiring less computation than traditional large-core attention.

[0108] As an example, combining Figure 8 and Figure 9 The process of extracting features from the thirteenth feature map through the large separable convolutional attention layer includes the following steps: inputting the thirteenth feature map into the first depth-wise separable convolution sublayer (1×(2d-1)DW-Conv), the second depth-wise separable convolution sublayer ((2d-1)×1DW-Conv), the first depth-wise separable dilated convolution sublayer (1×[k / d]DW-D-Conv), the second depth-wise separable dilated convolution sublayer ([k / d]×1DW-D-Conv), and the full convolution sublayer (1×1Conv) in sequence for feature extraction to obtain the fifteenth feature map; inputting the fifteenth feature map and the thirteenth feature map into the feature fusion sublayer for feature map multiplication to obtain the fourteenth feature map.

[0109] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.

[0110] In some embodiments, the training method of the lightweight image super-resolution model provided in the embodiments of the present application includes the following steps:

[0111] A training dataset is obtained, where the training dataset includes a general super-resolution dataset and a high-definition image dataset. The general super-resolution dataset can be a dataset including multiple natural images of different image categories (such as airplanes, cars, birds, trees, etc.) and different natural scenes (such as landscapes, people, buildings, etc.); the high-definition image dataset can be a dataset including multiple high-definition cornfield images collected by a low-altitude drone.

[0112] The initial image super-resolution model is trained using the training dataset until a preset convergence condition is reached, thereby obtaining a trained lightweight image super-resolution model. The model structure of the initial image super-resolution model is the same as that of the trained lightweight image super-resolution model. The preset convergence condition can be achieving a preset model accuracy (e.g., 80%, 90%, etc.) or reaching a preset number of iterations (e.g., 50, 100, etc.).

[0113] By using general super-resolution datasets and high-definition image datasets to train the initial image super-resolution model, the model's image super-resolution effect in the corn seedling detection scenario can be enhanced.

[0114] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0115] The embodiment of the present application further provides an image super-resolution device, which is configured to: input an image to be processed into a trained lightweight image super-resolution model, and output a super-resolved reconstructed image, wherein the resolution of the image to be processed is lower than the resolution of the super-resolved reconstructed image;

[0116] The lightweight image super-resolution model includes: a first convolution module, a combined basic network module, a large separation convolution attention module, a second convolution module, a first pixel reconstruction module, a third convolution module and a feature map fusion module connected in sequence; a second pixel reconstruction module connected to the first convolution module, a fourth convolution module connected to the second pixel reconstruction module, and the fourth convolution module connected to the feature map fusion module.

[0117] In some embodiments, the above-mentioned image super-resolution device may be specifically configured as follows:

[0118] Performing convolution processing on the image to be processed through a first convolution module to obtain a first convolution feature map;

[0119] The first convolution feature map is processed by combining basic network modules to obtain a reduced-resolution feature map;

[0120] The reduced-resolution feature map is processed by a large separable convolutional attention module to obtain an attention feature map;

[0121] Performing convolution processing on the attention feature map through the second convolution module to obtain a second convolution feature map;

[0122] Performing pixel reassembly processing on the second convolution feature map through a first pixel reassembly module to obtain a first pixel reassembly feature map;

[0123] Performing convolution processing on the first pixel reorganization feature map through a third convolution module to obtain a third convolution feature map;

[0124] Performing pixel reassembly processing on the first convolution feature map through a second pixel reassembly module to obtain a second pixel reassembly feature map;

[0125] performing convolution processing on the second pixel reorganization feature map through a fourth convolution module to obtain a fourth convolution feature map;

[0126] The third convolution feature map and the fourth convolution feature map are fused through the feature map fusion module to obtain a super-resolution reconstructed image.

[0127] An embodiment of the present application also provides a lightweight image super-resolution model, comprising: a first convolution module, a combined basic network module, a large separation convolution attention module, a second convolution module, a first pixel reconstruction module, a third convolution module and a feature map fusion module connected in sequence; a second pixel reconstruction module connected to the first convolution module, a fourth convolution module connected to the second pixel reconstruction module; and the fourth convolution module is connected to the feature map fusion module.

[0128] In some embodiments, the combined basic network module includes n basic network modules, the n basic network modules are in a serial structure, and n is an integer ≥3; each basic network module includes: a first basic residual unit, a high-frequency extraction unit, a second basic residual unit, a feature map splicing unit, a first channel adjustment unit, a channel attention unit, a third basic residual unit and a feature map fusion unit connected in sequence; each basic network module also includes: a downsampling unit connected to the first basic residual unit, a combined basic residual unit connected to the downsampling unit, a second channel adjustment unit connected to the combined basic residual unit, and an upsampling unit connected to the second channel adjustment unit; the upsampling unit is connected to the feature map splicing unit; wherein, the combined basic residual unit includes m fourth basic residual units, the m fourth basic residual units are in a serial structure, and m is an integer ≥3.

[0129] In some embodiments, the first basic residual unit includes: a first convolution activation layer, a second convolution activation layer and a first feature map fusion layer connected in sequence; wherein the convolution kernel sizes of the first convolution activation layer and the second convolution activation layer are different; the first feature map fusion layer is connected to the high-frequency extraction unit; the high-frequency extraction unit includes: a first average pooling layer, an upsampling layer and a second feature map fusion layer connected in sequence; wherein the first average pooling layer is connected to the first feature map fusion layer; the second feature map fusion layer is connected to the second basic residual unit.

[0130] In some embodiments, the downsampling unit includes a two-dimensional discrete decomposition layer, a feature map splicing layer connected to the two-dimensional discrete decomposition layer, and a full convolution layer connected to the feature map splicing layer; the two-dimensional discrete decomposition layer is connected to the first basic residual unit; and the full convolution layer is connected to the combined basic residual unit.

[0131] In some embodiments, the channel attention unit includes a second average pooling layer, at least one first convolution activation layer, a second convolution activation layer and a third feature map fusion layer connected in sequence; wherein the first convolution activation layer and the second convolution activation layer use different activation functions; the third feature map fusion layer is connected to the third basic residual unit; and the second average pooling layer is connected to the first channel adjustment unit.

[0132] In some embodiments, the large separated convolutional attention module includes: an attention unit, a first feature map superposition unit, a feedforward unit and a second feature map superposition unit connected in sequence; the attention unit is connected to the combined basic network module; wherein, the feedforward unit includes: a first full convolutional layer, a second full convolutional layer, a first activation function layer and a third full convolutional layer connected in sequence; the convolution kernel sizes of the first full convolutional layer and the second full convolutional layer are different, and the convolution kernel sizes of the first full convolutional layer and the third full convolutional layer are the same; the attention unit includes: a fourth full convolutional layer, a second activation function layer, a large separated convolutional attention layer and a fifth full convolutional layer connected in sequence; the convolution kernel sizes of the fourth full convolutional layer and the fifth full convolutional layer are the same.

[0133] In some embodiments, the large separable convolutional attention layer includes: a first depth-wise separable convolution sublayer, a second depth-wise separable convolution sublayer, a first depth-wise separable dilated convolution sublayer, a second depth-wise separable dilated convolution sublayer, a full convolution sublayer and a feature fusion sublayer connected in sequence; wherein the first depth-wise separable convolution sublayer is connected to the second activation function layer; and the feature fusion sublayer is connected to the fifth full convolution layer.

[0134] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0135] Figure 10Schematic diagram of an electronic device provided by an embodiment of the present application. Figure 10 As shown, the electronic device 1000 of this embodiment includes: a processor 1001, a memory 1002, and a computer program 1003 stored in the memory 1002 and executable by the processor 1001. When the processor 1001 executes the computer program 1003, the steps of the above-described method embodiments are implemented. Alternatively, when the processor 1001 executes the computer program 1003, the functions of the modules / units in the above-described apparatus embodiments are implemented.

[0136] The electronic device 1000 may be a desktop computer, a notebook, a PDA, a cloud server, or other electronic device. The electronic device 1000 may include but is not limited to a processor 1001 and a memory 1002. Those skilled in the art will understand that Figure 10 The electronic device 1000 is merely an example and does not limit the electronic device 1000 . The electronic device 1000 may include more or fewer components than shown in the figure, or different components.

[0137] The processor 1001 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0138] The memory 1002 may be an internal storage unit of the electronic device 1000, such as a hard disk or memory of the electronic device 1000. The memory 1002 may also be an external storage device of the electronic device 1000, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the electronic device 1000. The memory 1002 may also include both an internal storage unit of the electronic device 1000 and an external storage device. The memory 1002 is used to store computer programs and other programs and data required by the electronic device.

[0139] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0140] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. The computer program may include computer program code, which may be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0141] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. An image super-resolution method, characterized in that: include: Inputting the image to be processed into a trained lightweight image super-resolution model and outputting a super-resolution reconstructed image, wherein the resolution of the image to be processed is lower than the resolution of the super-resolution reconstructed image; The lightweight image super-resolution model includes: a first convolution module, a combined basic network module, a large separation convolution attention module, a second convolution module, a first pixel reconstruction module, a third convolution module and a feature map fusion module connected in sequence; a second pixel reassembly module connected to the first convolution module, and a fourth convolution module connected to the second pixel reassembly module; The fourth convolution module is connected to the feature map fusion module.

2. The method according to claim 1, characterized in that The combined basic network module includes n basic network modules, the n basic network modules are in a serial structure, and n is an integer ≥3; Each of the basic network modules includes: a first basic residual unit, a high frequency extraction unit, a second basic residual unit, a feature map splicing unit, a first channel adjustment unit, a channel attention unit, a third basic residual unit and a feature map fusion unit connected in sequence; Each of the basic network modules further includes: a downsampling unit connected to the first basic residual unit, a combined basic residual unit connected to the downsampling unit, a second channel adjustment unit connected to the combined basic residual unit, and an upsampling unit connected to the second channel adjustment unit; the upsampling unit is connected to the feature map splicing unit; The combined basic residual unit includes m fourth basic residual units, the m fourth basic residual units are in a serial connection structure, and m is an integer ≥3.

3. The method according to claim 2, characterized in that The first basic residual unit includes: a first convolution activation layer, a second convolution activation layer, and a first feature map fusion layer connected in sequence; wherein the convolution kernel sizes of the first convolution activation layer and the second convolution activation layer are different; the first feature map fusion layer is connected to the high frequency extraction unit; The high-frequency extraction unit includes: a first average pooling layer, an upsampling layer, and a second feature map fusion layer connected in sequence; wherein the first average pooling layer is connected to the first feature map fusion layer; and the second feature map fusion layer is connected to the second basic residual unit.

4. The method according to claim 2, characterized in that The downsampling unit includes a two-dimensional discrete decomposition layer, a feature map splicing layer connected to the two-dimensional discrete decomposition layer, and a full convolution layer connected to the feature map splicing layer; The two-dimensional discrete decomposition layer is connected to the first basic residual unit; the full convolution layer is connected to the combined basic residual unit.

5. The method according to claim 2, characterized in that The channel attention unit includes a second average pooling layer, at least one first convolution activation layer, a second convolution activation layer and a third feature map fusion layer connected in sequence; The first convolution activation layer and the second convolution activation layer use different activation functions; The third feature map fusion layer is connected to the third basic residual unit; The second average pooling layer is connected to the first channel adjustment unit.

6. The method according to claim 1, characterized in that The large-scale separation convolution attention module includes: an attention unit, a first feature map superposition unit, a feedforward unit, and a second feature map superposition unit connected in sequence; the attention unit is connected to the combined basic network module; The feedforward unit includes: a first full convolutional layer, a second full convolutional layer, a first activation function layer, and a third full convolutional layer connected in sequence; the convolution kernel sizes of the first full convolutional layer and the second full convolutional layer are different, and the convolution kernel sizes of the first full convolutional layer and the third full convolutional layer are the same; The attention unit includes: a fourth full convolutional layer, a second activation function layer, a large separation convolutional attention layer and a fifth full convolutional layer connected in sequence; the convolution kernels of the fourth full convolutional layer and the fifth full convolutional layer have the same size.

7. The method according to claim 6, characterized in that The large separable convolutional attention layer includes: a first depth-wise separable convolution sublayer, a second depth-wise separable convolution sublayer, a first depth-wise separable dilated convolution sublayer, a second depth-wise separable dilated convolution sublayer, a full convolution sublayer, and a feature fusion sublayer connected in sequence; Among them, the first depth-wise separable convolution sublayer is connected to the second activation function layer; and the feature fusion sublayer is connected to the fifth full convolution layer.

8. The method according to claim 1, characterized in that Input the image to be processed into the trained lightweight image super-resolution model and output the super-resolution reconstructed image, including: Performing convolution processing on the image to be processed by the first convolution module to obtain a first convolution feature map; Processing the first convolution feature map by combining basic network modules to obtain a reduced-resolution feature map; Processing the reduced-resolution feature map through the large separable convolutional attention module to obtain an attention feature map; Performing convolution processing on the attention feature map through the second convolution module to obtain a second convolution feature map; Performing pixel reassembly processing on the second convolution feature map by the first pixel reassembly module to obtain a first pixel reassembly feature map; Performing convolution processing on the first pixel reconstruction feature map by the third convolution module to obtain a third convolution feature map; Performing pixel reassembly processing on the first convolution feature map by the second pixel reassembly module to obtain a second pixel reassembly feature map; Performing convolution processing on the second pixel reconstruction feature map by the fourth convolution module to obtain a fourth convolution feature map; The third convolution feature map and the fourth convolution feature map are fused by the feature map fusion module to obtain a super-resolution reconstructed image.

9. An image super-resolution device, characterized in that: The image super-resolution device is configured as follows: Inputting the image to be processed into a trained lightweight image super-resolution model and outputting a super-resolution reconstructed image, wherein the resolution of the image to be processed is lower than the resolution of the super-resolution reconstructed image; The lightweight image super-resolution model includes: The first convolution module, the combined basic network module, the large separation convolution attention module, the second convolution module, the first pixel reconstruction module, the third convolution module and the feature map fusion module are connected in sequence; A second pixel reconstruction module connected to the first convolution module, a fourth convolution module connected to the second pixel reconstruction module, and the fourth convolution module is connected to the feature map fusion module.

10. A lightweight image super-resolution model, characterized in that: include: The first convolution module, the combined basic network module, the large separation convolution attention module, the second convolution module, the first pixel reconstruction module, the third convolution module and the feature map fusion module are connected in sequence; a second pixel reassembly module connected to the first convolution module, and a fourth convolution module connected to the second pixel reassembly module; The fourth convolution module is connected to the feature map fusion module.