Automobile part image super-resolution method based on space channel joint attention

By building a spatial channel joint attention network, the high-frequency and low-frequency characteristics of low-resolution automotive parts images are separated and enhanced, the problem of improving low-resolution image quality is solved, and significant improvements in image details and structural information are achieved.

CN120355575AActive Publication Date: 2025-07-22SHANDONG BLUEBIRD IND INTERNET CO LTD

Patent Information

Application Number
CN202510846796.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

In the image processing of automotive parts, low-resolution images have poor image processing and recognition effects due to factors such as shooting angle and lighting changes. How to improve image quality on the basis of ensuring details and structural information has become a research hotspot.

Method used

By building a spatial channel joint attention network, low-resolution images are divided into high-frequency and low-frequency features, processed and enhanced, and then feature fusion and upsampling reconstruction are carried out to improve the image super-resolution effect.

Benefits of technology

Effectively improve image details and structural information, and significantly improve super-resolved image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355575A_ABST
    Figure CN120355575A_ABST
Patent Text Reader

Abstract

The invention provides an automobile part image super-resolution method based on space channel joint attention, which relates to the field of computer vision, and comprises the following steps of: firstly, dividing an image into a high-frequency part and a low-frequency part, and respectively processing the high-frequency part and the low-frequency part; for high-frequency information, the image definition is improved by enhancing detail features of the high-frequency information; for low-frequency information, redundancy is removed and low-frequency features are enhanced, so that the low-frequency information is more prominent; thirdly, fusing the high-frequency features and the low-frequency features, and further enhancing the features through combination of a space channel and an attention block; and finally, recovering the fused features to a high-resolution space through up-sampling operation to obtain a super-resolution image. According to the method, image detail and structure information is effectively improved, and the quality of the super-resolution image is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and particularly relates to a method for super-resolution of automotive component images based on spatial-channel joint attention. Background Art

[0002] With the rapid development of industrial automation and intelligent manufacturing, the inspection and quality control of automotive components have become increasingly important; in order to improve production efficiency and ensure the quality of components, the application of image recognition technology in component inspection has received extensive attention; however, in practical applications, the images of many automotive components often present low resolution or noise interference due to factors such as shooting angle and lighting changes, resulting in poor effects of traditional image processing and recognition methods; therefore, how to improve the quality of low-resolution images while ensuring detail and structural information has become a research hotspot in the field of image processing.

[0003] In recent years, spatial and channel attention mechanisms have been widely applied to image processing tasks, effectively enhancing the representation ability of images by weighting features in different regions and channels; however, how to reasonably integrate spatial and channel attention mechanisms in super-resolution tasks and give full play to their enhancement effects on high-frequency and low-frequency features remains a topic worthy of exploration; the present invention proposes a method for super-resolution of automotive component images based on spatial-channel joint attention, which divides low-resolution images into high-frequency and low-frequency features, processes them separately, then performs feature fusion, and further enhances and up-samples and reconstructs on this basis, which can better enhance the detail and structural information of images and achieve better super-resolution reconstruction effects. Summary of the Invention

[0004] The present invention proposes a method for super-resolution of automotive component images based on spatial-channel joint attention, aiming to construct a spatial-channel joint network, divide low-resolution into high-frequency and low-frequency features, and perform spatial-channel joint attention enhancement after processing, so as to improve the super-resolution effect of images.

[0005] The present invention integrates spatial attention and channel attention, and provides a method for super-resolution of automotive component images based on spatial-channel joint attention, including the following steps:

[0006] S1. Use an industrial camera to capture automotive components to obtain low-resolution automotive component images;

[0007] S2. Construct a high-frequency feature enhancement module, which includes high-frequency feature extraction, 3x3 convolution, and ReLU activation function;

[0008] S3. Construct a low-frequency feature enhancement module, which includes low-frequency feature extraction, thresholding redundancy removal, and local adaption;

[0009] S4. Construct a spatial-channel joint attention block, which consists of upper and lower branches. The upper and lower branches are the improved spatial attention and channel attention respectively. Finally, multiply the outputs of the upper and lower branches by the original feature map through element-wise multiplication;

[0010] S5. Construct a feature processing block. The feature processing block passes through a high-frequency feature enhancement module and a low-frequency feature enhancement module. Connect the outputs of the two modules, then pass through a 1x1 convolution and the spatial-channel joint attention block, and then through a 1x1 convolution and pixel normalization. Finally, connect the output feature with the input through a skip connection;

[0011] S6. Construct a spatial-channel joint attention network, and input the image features into a multi-level spatial-channel joint attention network. The multi-level attention network includes n feature processing blocks and a feature fusion layer;

[0012] S7. Construct an image reconstruction module, which performs convolution and upsampling on the fused features;

[0013] S8. Construct an image super-resolution model, which includes an input, feature extraction, spatial-channel joint network, image reconstruction, and image output;

[0014] S9. Input the automotive part image into the image super-resolution model, and obtain a super-resolution image after processing.

[0015] Preferably, in step S2, construct a high-frequency feature enhancement module, which is characterized in that: Input the initial features of the automotive part image , and pass the initial features of the automotive part image through a Gaussian kernel of size 7x7, and use the initial features to subtract the features processed by the Gaussian kernel. The Gaussian kernel blur approximates the low-frequency part, and subtracting the low-frequency part from the initial features obtains the high-frequency part. , represents a 7x7 Gaussian kernel. Subsequently, pass the high-frequency features through a 3x3 convolution and a ReLU activation function. The 3x3 convolution obtains the features of the high-frequency part, and the ReLU activation function obtains the attention map. Multiply the high-frequency features and the attention map pixel by pixel through a skip connection to further enhance the high-frequency region. To avoid over-enhancement, use a residual connection to combine the enhanced high-frequency features with the original high-frequency features. Finally, pass the combined features through a 3x3 convolution and a ReLU activation function. This process can be expressed as: , represents a 3×3 convolution, represents a ReLU activation function, represents element-wise multiplication, represents element-wise addition, is the output of the high-frequency feature enhancement module.

[0016] Preferably, in step S2, low-frequency features are extracted by Gaussian kernel blur, and high-frequency features are obtained by differential operation, effectively combining the high-frequency features with the attention map to enhance the performance of detailed features; residual connections are used to avoid over-enhancement, retain the original information, and improve the stability and training effect of the network; through element-wise multiplication and convolution processing, the expression ability of the high-frequency part is enhanced, while maintaining the diversity and fineness of the features.

[0017] Preferably, in step S3, a low-frequency feature enhancement module is constructed, which is characterized in that: Input the initial features of the automotive component image The initial features of the automotive component image are passed through a 3x3 convolutional layer to extract low-frequency features, and then by calculating the range of changes in the low-frequency region, redundant information is removed according to a preset threshold. The features after removing redundancy , represents thresholding to remove redundancy, represents the preset threshold, The mathematical model of is , where is the input feature is the number of elements in represents the th element value, , is the input feature is the mean value of is the input feature is the standard deviation of; for the low-frequency features after removing redundancy, the local mean value of the low-frequency region is subtracted using differential operation to avoid loss of details or over-enhancement, and then multiplied by an enhancement factor to enhance the low-frequency features. Subsequently, residual connections are used to combine the enhanced low-frequency features with the original low-frequency features. Finally, through a 3×3 convolution and ReLU activation function, this process can be expressed as: , is the output of the low-frequency feature enhancement module, represents the ReLU activation function, is the local mean value of the low-frequency region, is the enhancement factor, The mathematical model of is , , where is the mean value of the low-frequency features after removing redundancy, is the input feature The number of elements in , Indicates The value of the element, is the standard deviation of the low-frequency features after redundancy removal, is the standard deviation of the input feature.

[0018] Preferably, in step S3, low-frequency features are extracted through 3×3 convolution, and redundancy is removed through thresholding to effectively reduce noise and unnecessary information; the combination of differential operation and enhancement factor ensures that the details of low-frequency features are maintained and enhanced, avoiding the risk of over-enhancement; residual connection helps to retain the original low-frequency information and improve the stability and convergence speed of the model; through the subtraction operation of the local mean, the expression ability of the low-frequency features is further improved, and the image quality is improved.

[0019] Preferably, in step S4, a spatial channel joint attention block is constructed, which includes: S41, the spatial channel joint attention block is divided into an upper branch and a lower branch, and the processed image features are input , the processed image features Divided into The two parts are processed by two branches respectively. For the upper branch, the input feature Max pooling and average pooling are performed separately to obtain two feature maps. The max pooling layer can retain the main features in the feature map and suppress noise and unimportant details. , Represents the maximum pooling operation. The average pooling layer reduces the size of the feature map by calculating the average value of the local area. , Represents the average pooling operation, and then concatenates the pooling results , Represents the connection operation. The connected features pass through a convolution layer with a kernel size of 3×3, which helps to capture the spatial dependencies in the local area. Then, the sigmoid activation function is used to obtain the spatial attention map. In order to enhance the response of the important area, the output map is multiplied element by element with the original feature map. The 3×3 convolution layer and ReLU activation function are used to enhance the expression ability of the local features. This process can be expressed as , represents the ReLU activation function, represents the sigmoid activation function, represents 3×3 convolution, stands for element-wise multiplication; S42. For the lower branch, input feature It is further divided into two branches, the upper branch performs global average pooling, and then uses 1×1 convolution and ReLU activation function to get the attention of each channel. , represents the global average pooling operation. After the global average pooling is used in the lower branch, it is folded back, and the front and back channels are swapped. This process can be described as: , and then a 1×1 convolution and a ReLU activation function are used. Then, the result of the folding back is subjected to skip connection, and finally the obtained result is folded back. This process can be described as: , represents the folding back operation. Finally, the outputs of the two branches are weighted and connected. This process can be described as: , represents the weight, and its value is equal to 0.7, represents element-wise addition; S43. After enhancing the high-frequency and low-frequency features, then the outputs of the two branches and the initial input features are multiplied element-wise to obtain the final output of the spatial-channel attention block , represents element-wise multiplication.

[0020] Preferably, in step S4, the advantage of the spatial-channel joint attention block is that it can effectively fuse high-frequency and low-frequency features. By separately processing spatial and channel attention through the upper and lower branches, the expression ability of the feature map is improved; the upper branch enhances local spatial dependence through max pooling and average pooling, and the lower branch enhances the relationship between channels through global pooling and folding back operations, so as to better capture important feature information; the final element-wise multiplication operation ensures the effective fusion between different features and improves the performance and generalization ability of the model.

[0021] Preferably, in step S5, a feature processing block is constructed, which is characterized in that: For the input image features of automotive parts of the feature processing block , first, they pass through a high-frequency feature enhancement module and a low-frequency feature enhancement module. The enhanced high-frequency features , represents the high-frequency feature enhancement module, and the enhanced low-frequency features , represents the low-frequency feature enhancement module. Then, the enhanced high-frequency and low-frequency features are connected, and the features are fused through a 1×1 convolution. The fused features , represents the connection operation, represents the 1×1 fusion convolution; subsequently, the features are further finely adjusted through the spatial-channel attention joint module, a 1×1 convolution is used for feature transformation, and at the same time, pixel normalization is added to improve the model performance. Finally, a long skip connection is used to enhance the residual learning ability of the model to obtain the final output of the feature processing block , represents the pixel normalization operation, Represents a spatial channel joint attention block.

[0022] Preferably, in step S5, the feature processing block effectively separates and enhances the image information through the high-frequency and low-frequency feature enhancement modules, improving the model's processing ability for different frequency features. The fused features are further optimized through the 1×1 convolution and the spatial channel attention joint module, which can more finely adjust the feature expression and enhance the model's learning ability for local and global information; Pixel normalization and long skip connections help improve the model's stability and residual learning ability, ultimately achieving stronger feature expression and improving the overall performance.

[0023] Preferably, in step S6, a spatial channel joint attention network is constructed, which is characterized in that: For the input feature , we concatenate the outputs of each feature processing block, perform a connection operation on the outputs of each feature processing block, fuse the connected results through 1x1 convolution, then refine the features through a 3x3 convolution, and finally add the original feature to the network output through a skip connection, so that the final output , represents the output of the nth spatial channel attention block.

[0024] Preferably, in step S6, by cascading the outputs of multiple spatial channel attention blocks, the feature information at different levels is fully fused, improving the model's expression ability for multi-scale features; 1×1 convolution is used for feature fusion, effectively reducing the dimension and retaining important information, and 3×3 convolution further refines the features, improving the spatial detail expression; The original feature is directly added through a skip connection, retaining the low-level information, enhancing the residual learning ability, and helping to improve the stability and convergence speed of the model.

[0025] Compared with the prior art, the present invention has the following technical effects: The spatial channel joint attention network provided by the technical solution of the present invention consists of n spatial channel joint attention blocks and a feature fusion module. The spatial channel joint attention block divides the image into high-frequency and low-frequency parts and processes them separately. For high-frequency features, by enhancing their detailed features, the image clarity is improved; for low-frequency features, by removing redundancy and enhancing low-frequency features, they are made more prominent; then, the high-frequency and low-frequency features are fused, and then the features are further enhanced through the spatial channel joint attention block. Finally, the fused features are restored to the high-resolution space through an upsampling operation to obtain a super-resolution image. This method effectively improves the image detail and structural information and significantly improves the quality of the super-resolution image. Brief Description of the Drawings

[0026] Figure 1 is the flowchart of the super-resolution of automotive part images proposed by the present invention.

[0027] Figure 2 is an example diagram of the high-frequency feature enhancement module proposed by the present invention.

[0028] Figure 3 is an example diagram of the low-frequency feature enhancement module proposed by the present invention.

[0029] Figure 4 is an example diagram of the spatial-channel joint attention block proposed by the present invention.

[0030] Figure 5 is the structural diagram of the feature processing block proposed by the present invention.

[0031] Figure 6 is the structural diagram of the spatial-channel joint attention network proposed by the present invention. Detailed implementation manners

[0032] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0033] The present invention proposes a super-resolution method for automotive part images based on spatial-channel joint attention. First, a spatial-channel joint attention network is constructed. The spatial-channel joint attention network consists of n feature processing blocks and a feature fusion module. The spatial-channel joint attention block divides the image into high-frequency and low-frequency parts and processes them separately. For high-frequency features, by enhancing their detailed features, the image clarity is improved. For low-frequency features, by removing redundancy and enhancing low-frequency features, they are made more prominent. Then, the high-frequency and low-frequency features are fused, and then the features are further enhanced through the spatial-channel joint attention block. Finally, the fused features are restored to the high-resolution space through an upsampling operation to obtain a super-resolution image. This method effectively improves the image details and structural information and significantly enhances the quality of the super-resolution image.

[0034] Please refer to Figure 1 as shown, the super-resolution method for automotive part images based on spatial-channel joint attention in the embodiments of the present application:

[0035] S1. Use an industrial camera to capture automotive parts to obtain a low-resolution automotive part image.

[0036] S2. Construct a high-frequency feature enhancement module, which includes high-frequency feature extraction, 3x3 convolution, and ReLU activation function.

[0037] Furthermore, as Figure 2 shown, construct a high-frequency feature enhancement module, and the specific steps are as follows.

[0038] Input the initial features of the automotive component image , and pass the initial features of the automotive component image through a Gaussian kernel of size 7x7, and use the initial features to subtract the features processed by the Gaussian kernel. The Gaussian kernel blurs to approximate the low-frequency part, and subtracting the low-frequency part from the initial features gives the high-frequency part. , denotes the 7x7 Gaussian kernel. Subsequently, pass the high-frequency features through 3x3 convolution and ReLU activation function. The 3x3 convolution obtains the features of the high-frequency part, and the ReLU activation function obtains the attention map. Then, multiply the high-frequency features and the attention map pixel by pixel through skip connection to further enhance the high-frequency region. To avoid over-enhancement, use residual connection to combine the enhanced high-frequency features with the original high-frequency features. Finally, pass the combined features through 3x3 convolution and ReLU activation function. This process can be expressed as: , denotes 3×3 convolution, denotes ReLU activation function, represents element-wise multiplication, represents element-wise addition, is the output of the high-frequency feature enhancement module.

[0039] S3. Construct a low-frequency feature enhancement module, which includes low-frequency feature extraction, thresholding redundancy removal, and local adaption.

[0040] Furthermore, as Figure 3 shown, construct a low-frequency feature enhancement module, and the specific steps are as follows.

[0041] Input the initial features of the automotive component image , and pass the initial features of the automotive component image through a 3x3 convolutional layer to extract low-frequency features, and then calculate the change range of the low-frequency region. Remove redundancy according to a preset threshold. The features after removing redundancy , denotes thresholding redundancy removal, denotes the preset threshold, The mathematical model of is , where is the input feature the number of elements in represents the value of the th element, , which is the mean of the input features , and is the standard deviation of the input features; for removing redundant low-frequency features, a difference operation is used to subtract the local mean of the low-frequency region to avoid detail loss or over-enhancement, and then the low-frequency features are enhanced by multiplying with an enhancement factor. Subsequently, a residual connection is used to combine the enhanced low-frequency features with the original low-frequency features. Finally, a 3×3 convolution and a ReLU activation function are applied. This process can be expressed as: , where represents the ReLU activation function, is the local mean of the low-frequency region, is the enhancement factor, The mathematical model of is , , where is the mean of the low-frequency features after redundancy removal, is the input feature and represents the th element value, is the standard deviation of the low-frequency features after redundancy removal, and

[0042] S4. Construct a spatial-channel joint attention block. The spatial-channel joint attention block consists of upper and lower branches. The upper and lower branches are improved spatial attention and channel attention respectively. Finally, the outputs of the upper and lower branches are multiplied with the original feature map through element-wise multiplication.

[0043] Furthermore, as Figure 4 shown, construct a spatial-channel joint attention block. The specific steps are as follows.

[0044] S41. The spatial-channel joint attention block is divided into an upper branch and a lower branch. The processed image features are input. The processed image features are divided into two parts and processed by the two branches respectively. For the upper branch, the input features are respectively subjected to max pooling and average pooling to obtain two feature maps. The max pooling layer can retain the main features in the feature map and suppress noise and unimportant details , represents the max pooling operation. The average pooling layer reduces the size of the feature map by calculating the average value of the local region , represents an average pooling operation, and then the pooling results are concatenated , represents a concatenation operation. The concatenated features pass through a convolutional layer with a kernel size of 3×3, which helps capture spatial dependencies within the local region. Then, a sigmoid activation function is used to obtain the spatial attention map. To enhance the response of important regions, the output map is multiplied element-wise with the original feature map, and a 3×3 convolutional layer and ReLU activation function are used to enhance the expression ability of local features. This process can be expressed as , represents the ReLU activation function, represents the sigmoid activation function, represents a 3×3 convolution, represents element-wise multiplication;

[0045] S42. For the lower branch, the input features are further divided into upper and lower branches. The upper branch performs global average pooling, and then uses a 1×1 convolution and ReLU activation function to obtain the attention for each channel, , represents the global average pooling operation. The lower branch uses global average pooling followed by transposition, swapping the front and back channels. This process can be expressed as: , and then uses a 1×1 convolution and ReLU activation function. Then, the transposed result is subjected to a skip connection, and finally the obtained result is transposed. This process can be expressed as: , represents the transposition operation. Finally, the outputs of the two branches are weighted and concatenated. This process can be expressed as: , represents the weight, and its value is equal to 0.7, represents element-wise addition;

[0046] S43. After enhancing the high-frequency and low-frequency features, then the outputs of the two branches are multiplied element-wise with the initial input features to obtain the final output of the spatial channel attention block , represents element-wise multiplication.

[0047] S5. Construct a feature processing block. The feature processing block passes through a high-frequency feature enhancement module and a low-frequency feature enhancement module. The outputs of the two modules are concatenated and then passed through a 1x1 convolution and a spatial channel joint attention block, and then through a 1x1 convolution and pixel normalization. Finally, the output features are connected to the input through a skip connection.

[0048] Furthermore, as Figure 5 shown, construct a feature processing block, and the specific steps are as follows.

[0049] The feature processing block processes the input image features of automotive parts , first through a high-frequency feature enhancement module and a low-frequency feature enhancement module. After enhancement, the high-frequency features , denote the high-frequency feature enhancement module, and the low-frequency features after enhancement , denote the low-frequency feature enhancement module. Then, the enhanced high-frequency and low-frequency features are concatenated, and the features are fused through 1×1 convolution. The fused features , denote the concatenation operation, denote the 1×1 fusion convolution; subsequently, the features are further refined through a spatial-channel attention joint module. 1×1 convolution is used for feature transformation, and pixel normalization is added to improve the model performance. Finally, long skip connections are used to enhance the residual learning ability of the model to obtain the final output of the feature processing block , represents the pixel normalization operation, denotes the spatial-channel joint attention block.

[0050] S6. Construct a spatial-channel joint attention network and input the image features into the spatial-channel joint attention network. The spatial-channel joint attention network includes n feature processing blocks and a feature fusion layer.

[0051] S7. Construct an image reconstruction module. The image reconstruction module performs convolution and upsampling on the fused features.

[0052] S8. Construct an image super-resolution model, which includes an input, feature extraction, spatial-channel joint attention network, image reconstruction, and image output.

[0053] S9. Input the automotive parts image into the image super-resolution model, and after processing, obtain a super-resolution image.

[0054] Furthermore, as Figure 6 shown, construct an image super-resolution model. After using an industrial camera to capture a low-resolution automotive parts image, input the image into the feature extraction module to obtain image features , and input the image features into the spatial-channel joint attention network. In the network, cascade the outputs of each feature processing block, and perform a concatenation operation on the outputs of each feature processing block. Fuse the concatenated results through 1×1 convolution, then further refine the features through a 3×3 convolution, and finally add the original features to the network output through skip connections, so as to obtain the final output of the spatial-channel attention network , denotes the output of the nth spatial channel attention block. Finally, the output of the spatial channel attention network is passed through an image reconstruction module to obtain a super-resolution image. This process can be expressed as: , denotes upsampling.

[0055] The above is only the preferred embodiment of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the inventive concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention.

Claims

1. An automotive parts image super-resolution method based on spatial channel joint attention, characterized in that Including the following steps: S1. Use an industrial camera to capture images of automotive parts to obtain low-resolution images of automotive parts; S2. Construct a high-frequency feature enhancement module, which includes high-frequency feature extraction, 3x3 convolution, and ReLU activation function; S3. Construct a low-frequency feature enhancement module, which includes low-frequency feature extraction, thresholding redundancy removal, and local adaption; S4. Construct a spatial-channel joint attention block, which consists of upper and lower branches. The upper and lower branches are improved spatial attention and channel attention respectively. Finally, multiply the outputs of the upper and lower branches with the original feature map through element-wise multiplication; S5. Construct a feature processing block. The feature processing block passes through the high-frequency feature enhancement module and the low-frequency feature enhancement module. Connect the outputs of the two modules, then pass through 1x1 convolution and the spatial-channel joint attention block, then through 1x1 convolution and pixel normalization, and finally make a skip connection between the output feature and the input; S6. Construct a spatial-channel joint attention network and input the image features into the spatial-channel joint attention network, where the spatial-channel joint attention network includes n feature processing blocks and a feature fusion layer; S7. Construct an image reconstruction module, which performs convolution and upsampling on the fused features; S8. Construct an image super-resolution model, which includes input, feature extraction, spatial-channel joint attention network, image reconstruction, and image output; S9. Input the images of automotive parts into the image super-resolution model, and obtain super-resolution images after processing.

2. The method for super-resolving images of automotive parts based on spatial-channel joint attention according to claim 1, in step S2, construct a high-frequency feature enhancement module, characterized in that: Input the initial features of automotive component images , the initial features of automotive component images Pass through a Gaussian kernel of size 7x7 and use the initial features Subtract the features processed by the Gaussian kernel. The Gaussian kernel blurs to approximate the low-frequency part, and subtract the low-frequency part from the initial features to obtain the high-frequency part. , Denote the 7x7 Gaussian kernel. Subsequently, pass the high-frequency features through a 3x3 convolution and the ReLU activation function. The 3x3 convolution obtains the features of the high-frequency part, and the ReLU activation function obtains the attention map. Multiply the high-frequency features and the attention map pixel by pixel through a skip connection to further enhance the high-frequency region. To avoid over-enhancement, use a residual connection to combine the enhanced high-frequency features with the original high-frequency features. Finally, pass the combined features through a 3x3 convolution and the ReLU activation function. This process can be expressed as: , Denote the 3×3 convolution Denote the ReLU activation function Represents element-wise multiplication Represents element-wise addition Is the output of the high-frequency feature enhancement module 3. The method for super-resolving images of automotive parts based on spatial-channel joint attention according to claim 2, in step S3, construct a low-frequency feature enhancement module, characterized in that: Input the initial features of automotive parts images The initial features of the automotive parts images pass through a 3x3 convolutional layer to extract low-frequency features. Then, by calculating the change range of the low-frequency region and removing redundancy according to a preset threshold, the features after redundancy removal , denotes thresholding to remove redundancy, denotes the preset threshold, The mathematical model of which is , is the input feature 's mean value, is the input feature 's standard deviation; for the low-frequency features after redundancy removal, use a difference operation to subtract the local mean of the low-frequency region to avoid loss of details or over-enhancement, then multiply by an enhancement factor to enhance the low-frequency features. Subsequently, use a residual connection to combine the enhanced low-frequency features with the original low-frequency features. Finally, pass through a 3×3 convolution and a ReLU activation function. This process can be expressed as: , is the output of the low-frequency feature enhancement module, denotes the ReLU activation function, is the local mean of the low-frequency region, is the enhancement factor, The mathematical model of which is , is the standard deviation of the low-frequency features after redundancy removal, is the standard deviation of the input features.

4. The method for super-resolving images of automotive parts based on spatial-channel joint attention according to claim 3, in step S4, construct a spatial-channel joint attention block, characterized in that: S41. The spatial channel joint attention block is divided into an upper branch and a lower branch. The processed image features are input , and the processed image features are divided into two parts and processed by the two branches respectively. For the upper branch, the input features are respectively subjected to max pooling and average pooling to obtain two feature maps. The max pooling layer can retain the main features in the feature map, suppress noise and unimportant details , denotes the max pooling operation. The average pooling layer reduces the size of the feature map by calculating the average value of the local region , denotes the average pooling operation. Then the pooling results are concatenated , denotes the concatenation operation. The concatenated features pass through a convolutional layer with a kernel size of 3×3, which helps to capture the spatial dependencies within the local region. Then, a sigmoid activation function is used to obtain the spatial attention map. To enhance the response of the important regions, the output map is multiplied element-wise with the original feature map, and a 3×3 convolutional layer and a ReLU activation function are used to enhance the expression ability of the local features. This process can be expressed as , denotes the ReLU activation function, denotes the sigmoid activation function, denotes the 3×3 convolution, represents the element-wise multiplication; S42. For the lower branch, the input features are further divided into upper and lower branches. The upper branch performs global average pooling, and then uses a 1×1 convolution and the ReLU activation function to obtain the attention for each channel. , denotes the global average pooling operation. The lower branch performs an inverse fold after global average pooling and swaps the front and back channels. This process can be expressed as: , and then uses a 1×1 convolution and the ReLU activation function. Then, the result of the inverse fold is subjected to a skip connection, and finally the obtained result is inverse folded. This process can be expressed as: , denotes the inverse fold operation. Finally, the outputs of the two branches are weighted and connected. This process can be expressed as: , represents the weight, and its value is equal to 0.

7. represents element-wise addition; S43. After enhancing the high-frequency and low-frequency features, subsequently multiply the outputs of the two branches with the initial input features element-wise to obtain the final output of the spatial channel attention block , represents element-wise multiplication.

5. The method for super-resolving images of automotive parts based on spatial-channel joint attention according to claim 4, in step S5, construct a feature processing block, characterized in that: The feature processing block processes the input image features of automotive parts , first through a high-frequency feature enhancement module and a low-frequency feature enhancement module. After enhancement, the high-frequency features , denote the high-frequency feature enhancement module, and the low-frequency features after enhancement , denote the low-frequency feature enhancement module. Then, the enhanced high-frequency and low-frequency features are concatenated, and the features are fused through a 1×1 convolution. The fused features , denote the concatenation operation, denote the 1×1 fusion convolution; Subsequently, the features are further refined through a spatial-channel attention joint module. A 1×1 convolution is used for feature transformation, and pixel normalization is added to improve the model performance. Finally, a long skip connection is used to enhance the residual learning ability of the model to obtain the final output of the feature processing block , represents the pixel normalization operation, denotes the spatial-channel joint attention block.

6. The method for super-resolving images of automotive parts based on spatial-channel joint attention according to claim 5, in step S6, construct a spatial-channel joint attention network, characterized in that: For the input features , we concatenate the outputs of each feature processing block, perform a concatenation operation on the outputs of each feature processing block, fuse the concatenated results through a 1x1 convolution, then refine the features through a 3x3 convolution, and finally add the original features to the network output through skip connections, thereby obtaining the final output of the spatial channel attention network denotes the output of the nth spatial channel attention block.

Citation Information

Patent Citations

  • Lightweight image super-resolution method based on serial high-frequency attention

    CN114897690A

  • Binocular image super-resolution reconstruction method based on multistage intensified attention mechanism

    CN116797461A

  • Mine image super-resolution reconstruction method based on residual mixed attention

    CN117078516A

  • Automobile part image super-resolution method based on dual-channel residual attention

    CN118570065A

  • Single image dehazing method based on detail recovery

    US20240289928A1

Cited By

  • Image super-resolution system and method based on high and low frequency separation sensing Mama

    CN121639473A

  • Image super-resolution system and method based on high-low frequency separation perception mamba

    CN121639473B