Image processing method and system based on adaptive dictionary
Through the combination of adaptive dictionary and convolutional neural network, the spatial and channel self-attention mechanisms are used to fusion and strengthen the feature, which solves the problem of insufficient feature capture in traditional super-resolution technology, and realizes high-quality image reconstruction and feature extraction.
Patent Information
- Application Number
- CN202510377747.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-25
AI Technical Summary
Traditional super-resolution technology lacks the ability to adaptively capture complex textures and special structures of images, resulting in missing features or insufficient expression, affecting the quality of image reconstruction, and the combination of shallow features and deep features lacks sufficient interaction, affecting the accuracy and comprehensiveness of feature extraction.
By initializing the sparse dictionary, an adaptive dictionary is generated, and shallow features are extracted in combination with the convolutional neural network, and multi-dimensional attention features are generated using the spatial window self-attention mechanism and the channel window self-attention mechanism to perform feature fusion and strengthening, and finally image reconstruction is carried out through the spatial gate feedforward network.
It significantly improves the capture ability of complex textures and special structures, enhances the quality of image reconstruction, and improves the accuracy and comprehensiveness of feature extraction.
Smart Images

Figure CN120375111A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to an image processing method and system based on an adaptive dictionary. Background Art
[0002] In the fields of modern image processing and computer vision, super-resolution image processing (SR) is a crucial technology. With the popularization of high-resolution displays and cameras, people's requirements for the quality of images and videos are increasing day by day. However, in practical applications, due to device limitations, environmental factors, or transmission compression, etc., image and video data often suffer from problems such as reduced resolution and blurred details. Therefore, the super-resolution image processing technology has emerged, aiming to convert low-resolution (LR) images or videos into high-resolution (HR) images or videos through software or hardware means, so as to meet people's needs for high-quality visual information.
[0003] Currently, super-resolution technology is widely applied in multiple fields. Super-resolution technology can significantly improve image resolution and enhance the ability to identify and track targets. In the field of medical imaging, high-resolution medical images help doctors diagnose diseases more accurately. Super-resolution technology can improve the resolution of medical images such as B-ultrasound and magnetic resonance imaging, providing more reliable details for doctors' diagnosis. In addition, in the fields of film and television production, daily photography, network video transmission, etc., super-resolution technology also plays an important role.
[0004] In recent years, with the rise of deep learning, learning-based super-resolution methods have become mainstream. These methods use deep learning models such as convolutional neural networks (CNNs) and generative adversarial networks (GANs) to automatically learn the mapping relationship from low resolution to high resolution, and can capture the high-frequency details of images, generating more natural and realistic high-resolution images.
[0005] However, traditional super-resolution technologies lack the ability to adaptively capture complex textures and special structures of images, which easily leads to feature omission or insufficient expression, thus affecting the quality of image reconstruction. At the same time, shallow features and deep features are usually combined in a simple way, lacking sufficient feature interaction, which affects the accuracy and comprehensiveness of image feature extraction. Summary of the Invention
[0006] To solve the technical problems that traditional super-resolution techniques lack the ability to adaptively capture complex textures and special structures in images, which easily leads to feature omission or insufficient expression, thus affecting the quality of image reconstruction. At the same time, shallow features and deep features are usually combined only in a simple way, lacking sufficient feature interaction and affecting the accuracy and comprehensiveness of image feature extraction, the present invention provides an image processing method and system based on an adaptive dictionary.
[0007] The technical solutions provided by the embodiments of the present invention are as follows:
[0008] First aspect:
[0009] An image processing method based on an adaptive dictionary provided by an embodiment of the present invention includes:
[0010] S1: Obtain an image to be processed;
[0011] S2: Initialize a sparse dictionary, and use the sparse dictionary to extract sparse features of the image to be processed;
[0012] S3: Based on the sparse features, perform update and optimization processing on the sparse dictionary to generate an adaptive dictionary;
[0013] S4: Use the adaptive dictionary to extract first shallow features of the image to be processed;
[0014] S5: Input the first shallow features into a convolutional layer neural network to extract second shallow features of the image to be processed;
[0015] S6: Based on the second shallow features, use a spatial window self-attention mechanism and a channel window self-attention mechanism to respectively generate spatial window self-attention features and channel window self-attention features;
[0016] S7: Based on the spatial window self-attention features and the channel window self-attention features, perform feature fusion to obtain fused features;
[0017] S8: Use a spatial feed-forward network to strengthen the fused features to obtain strengthened features;
[0018] S9: According to the first shallow features and the strengthened features, perform image reconstruction on the image to be processed to obtain a reconstructed image.
[0019] Second aspect:
[0020] An image processing system based on an adaptive dictionary provided by an embodiment of the present invention includes:
[0021] A processor;
[0022] A memory stores computer-readable instructions, which, when executed by the processor, implement the image processing method based on an adaptive dictionary as described in the first aspect.
[0023] The third aspect:
[0024] A computer-readable storage medium provided by an embodiment of the present invention stores a computer program, which, when executed by a processor, implements the image processing method based on an adaptive dictionary as described in the first aspect.
[0025] The beneficial effects brought by the technical solution provided by the embodiment of the present invention at least include:
[0026] (1) In the present invention, through the initialization and optimization of the sparse dictionary, an adaptive dictionary is generated, effectively improving the ability to capture complex textures and special structures; combining the adaptive dictionary with the convolutional neural network to deeply extract shallow features and strengthening the representation of key details; introducing the spatial and channel self-attention mechanisms to generate multi-dimensional attention features and further enhancing the feature expression through feature fusion, thereby significantly improving the quality of image reconstruction.
[0027] (2) In the present invention, the second shallow features are extracted through the convolutional network, and combined with the spatial window self-attention mechanism and the channel window self-attention mechanism to generate spatial window self-attention features and channel window self-attention features respectively. Further, the two types of attention features are fused to form a fused feature, and finally the spatial feed-forward network is used to selectively strengthen the fused feature, effectively enhancing the interaction between shallow and deep features and significantly improving the accuracy and comprehensiveness of image feature extraction. Description of the Drawings
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0029] Figure 1 It is a schematic flowchart of an image processing method based on an adaptive dictionary provided by an embodiment of the present invention;
[0030] Figure 2 It is a schematic structural diagram of an image processing system based on an adaptive dictionary provided by an embodiment of the present invention. Detailed Embodiments
[0031] The following describes the technical solutions in the present invention with reference to the drawings.
[0032] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or advantageous than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0033] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.
[0034] In the embodiments of the present invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meanings they express are the same.
[0035] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0036] Refer to the attached Figure 1 , which shows a schematic flowchart of an image processing method based on an adaptive dictionary provided by an embodiment of the present invention.
[0037] The embodiments of the present invention provide an image processing method based on an adaptive dictionary. This method can be implemented by an image processing device based on an adaptive dictionary. The image processing device based on an adaptive dictionary can be a terminal or a server. The processing flow of an image processing method based on an adaptive dictionary can include the following steps:
[0038] S1: Obtain the image to be processed.
[0039] Among them, the image to be processed refers to a low-quality or degraded image received from an external input or a sensor. These images may have poor visual effects due to various reasons (such as blur, noise, low resolution, etc.).
[0040] S2: Initialize the sparse dictionary, and use the sparse dictionary to extract the sparse features of the image to be processed.
[0041] Among them, the sparse dictionary is a tool for signal processing and feature representation, aiming to efficiently represent the input data with a small number of bases (atoms) through the method of sparse representation.
[0042] In a possible implementation manner, S2 is specifically:
[0043] Using a sparse dictionary, the sparse features of the image to be processed are extracted through the following formula:
[0044]
[0045] where α * represents the sparse features of the image to be processed, and arg min α represents the sparse features corresponding to when the objective function obtains the minimum value. represents the sparse dictionary. M represents the number of atoms in the sparse dictionary, d represents the dimension of the sparse dictionary, x represents the image to be processed. represents the square of the two-norm, λ∥α∥1 represents the parameter controlling the sparse coefficient features, λ represents the regularization parameter, ∥α∥1 represents the sum of the absolute values of all elements in α, and α represents the sparse features.
[0046] In the present invention, the sparse dictionary can compactly represent data with a small number of atoms, accurately extract the core features of the image, effectively remove redundant information at the same time, and improve the representation efficiency. At the same time, the sparse features are insensitive to noise, can extract key information in complex environments, and improve the performance and reliability of image processing tasks.
[0047] S3: Based on the sparse features, the sparse dictionary is updated and optimized to generate an adaptive dictionary.
[0048] Among them, the adaptive dictionary is a dictionary matrix generated through iterative optimization, which can be dynamically adjusted to better adapt to the features and distribution of the input data. Compared with the fixed dictionary, the adaptive dictionary can provide a more efficient and accurate representation according to specific task requirements (such as image processing, signal analysis).
[0049] It should be noted that by encoding the sparse dictionary, the input data can be represented as a sparse linear combination of dictionary atoms, so as to capture the key features of the data. Updating and optimizing the dictionary can further improve the accuracy of this representation, making the sparse coefficients better reflect the essential structure of the data.
[0050] In a possible implementation manner, S3 is specifically:
[0051] Based on the sparse features, the sparse dictionary is updated and optimized through the following formula to generate an adaptive dictionary:
[0052]
[0053] where D represents the adaptive dictionary, x represents the image to be processed. represents the sparse dictionary of the t-th iteration. represents the square of the reconstruction error, λ represents the regularization parameter, ∥α∥1 represents the sum of the absolute values of all elements in α, α represents the sparse feature, and arg min D represents the sparse dictionary corresponding to when the objective function achieves the minimum value.
[0054] Specifically, the data preparation stage is responsible for collecting and preprocessing the sample data, providing the basis for model training. Subsequently, the model initialization stage lays the foundation for subsequent iterative optimization by setting the initial dictionary. Entering the iterative optimization stage, this process uses an iterative algorithm to continuously update the dictionary. In each iteration, the algorithm calculates the sparse representation of the samples based on the current dictionary, and then adjusts the dictionary atoms according to the error between the sparse representation and the original samples, aiming to minimize the reconstruction error and enhance sparsity. This process is repeated until the iterative termination condition is met, such as reaching the preset number of iterations or the error converges below a certain threshold. Finally, in the feature output stage, the optimized dictionary is used to perform sparse coding on the input data, and the sparse representation coefficients are extracted as the feature vectors.
[0055] In the present invention, the adaptive dictionary can dynamically adapt to the feature distribution of the input data, provide a more efficient sparse representation, significantly improve the data representation ability and task adaptability. At the same time, by reducing the reconstruction error and strengthening sparsity, the adaptive dictionary can not only highlight the core features of the data, but also effectively ignore redundant information, enhancing the robustness of the model to noise.
[0056] S4: Using the adaptive dictionary, extract the first shallow-layer features of the image to be processed.
[0057] S5: Input the first shallow-layer features into the convolutional layer neural network to extract the second shallow-layer features of the image to be processed.
[0058] Among them, the convolutional neural network is a deep learning model, which is particularly good at processing image data. It extracts the features in the image through operations such as convolution and pooling, and maps these features to a high-dimensional feature space. In the convolutional neural network, the convolutional layer is one of the core components. It performs a sliding operation on the input data using a series of convolutional kernels to extract local features. These local features can be edges, textures, corners, etc. in the image, and they are crucial for subsequent image classification, recognition and other tasks. As the convolutional layer deepens, the network can capture more complex and abstract features, which are usually called deep features.
[0059] In a possible implementation manner, S5 is specifically:
[0060] Input the first shallow-layer features into the convolutional layer neural network, and extract the second shallow-layer features of the image to be processed through the following formula:
[0061]
[0062] Among them, X represents the second shallow feature of the image to be processed, represents the first shallow feature of the image to be processed, Conv 3×3 represents a convolution operation with a convolution kernel size of 3×3.
[0063] Specifically, the image enters the convolutional layer and undergoes three 3×3 convolution operations in sequence. These convolution operations slide over the image through different convolution kernels (filters) to capture local features of the image, such as edges, textures, etc., and generate corresponding feature maps. The feature maps retain the key information of the image while reducing the data dimension. Subsequently, the feature maps are processed through an activation function, such as the ReLU function, to enhance the network's non-linear expression ability, retain important features, and suppress unimportant features. Finally, the processed shallow feature maps are output, and these feature maps contain the basic structure and texture information of the image.
[0064] In the present invention, the shallow features are used as inputs, providing a good starting point for the convolutional neural network, enabling the network to capture key information in the image more quickly. At the same time, since the adaptive dictionary has sparsely represented and updated the features, the features input into the convolutional neural network are more accurate and rich, which helps to improve the effect of subsequent deep feature extraction.
[0065] Furthermore, it is possible to combine the advantages of the adaptive dictionary and the convolutional neural network to achieve more efficient and accurate feature extraction. Through the preprocessing of the adaptive dictionary, the convolutional neural network can learn key features in the image more quickly, thereby improving the efficiency and accuracy of the entire image processing process.
[0066] S6: Based on the second shallow feature, use the spatial window self-attention mechanism and the channel window self-attention mechanism to generate spatial window self-attention features and channel window self-attention features respectively.
[0067] Among them, the spatial window self-attention mechanism is used in computer vision to capture complex dependencies within local regions of an image. By dividing the query matrix, key matrix, and value matrix into non-overlapping windows and performing self-attention calculations on the features within each window respectively, this mechanism can efficiently process large-scale image data.
[0068] Among them, the channel self-attention mechanism is used in the field of computer vision to enhance the model's ability to understand the dependencies between image channels. This mechanism models the channel dimension of the image features, learns the weight distribution between different channels, thereby highlighting features beneficial to the task and suppressing irrelevant features.
[0069] In a possible implementation manner, generating the self-attention features of the spatial window in S6 specifically includes:
[0070] S601: Generate the query matrix, key matrix, and value matrix of the second shallow feature according to the spatial window self-attention mechanism:
[0071] Q s = XW Q
[0072] K s = XW K
[0073] V s = XW V
[0074] Among them, Q s represents the query matrix of the second shallow feature, K s represents the key matrix of the second shallow feature, Vx represents the value matrix of the second shallow feature, X represents the second shallow feature of the image to be processed, and W Q represents the query weight matrix, W K represents the key weight matrix, and W V represents the value weight matrix.
[0075] It should be noted that Q s , K s , V s (Query, Key, Value) are the core components of the self-attention mechanism and the Transformer architecture. Q s , K s , V s are generated from the original input image or shallow feature through linear transformation, representing query, key, and value respectively. The query matrix Q s is used to find relevant information in the image features, the key matrix K s serves as an index to help locate this information, while the value matrix V s contains the actual information content. By calculating the similarity between Q s and K (i.e., the attention weight), and performing weighted summation on V s , the model can dynamically focus on different parts of the input image, capture complex dependencies, and improve performance in various CV tasks.
[0076] S602: Divide the query matrix, key matrix, and value matrix into non-overlapping windows respectively, and flatten each window into multiple attention heads:
[0077]
[0078]
[0079] V s = [Vs 1 , …, V s h
[0080] Among them, represents the query matrix of the i-th attention head, represents the key matrix of the i-th attention head, V s i represents the value matrix of the i-th attention head, i = 1, 2, …, h, and h represents the total number of attention heads.
[0081] Among them, the attention head is the core component of the self-attention mechanism and multi-head attention. Its main role is to calculate the correlation weights between the query, key, and value, and dynamically extract the features of the input data based on these weights.
[0082] In the present invention, the query matrix, key matrix, and value matrix are respectively divided into non-overlapping windows, which can process the information of specific subspaces in parallel. While retaining the global information, it particularly strengthens the capture of local features, enabling the model to take into account both details and overall features. At the same time, by dividing the query, key, and value matrices into non-overlapping windows and performing a flattening operation on each window, the overhead of global matrix calculation can be effectively reduced. While ensuring the feature capture ability, the calculation efficiency is greatly improved.
[0083] S603: Calculate the output features of each attention head:
[0084]
[0085] Among them, represents the output feature of the i-th attention head, represents the query matrix of the i-th attention head, represents the key matrix of the i-th attention head, Vsi represents the value matrix of the i-th attention head, e represents the size of each attention head, C represents the dimension of the channel, h represents the total number of attention heads, softmax() represents the activation function, T represents the transpose operation, and E represents the relative position encoding.
[0086] S604: Concatenate the output features of each attention head to generate the spatial window self-attention feature:
[0087]
[0088] SW - SA(X) = Y s W p
[0089] Among them, Y s Denote the spatial window self-attention feature, and concat denote the concatenation operation. Denote the output feature of the i-th attention head, where i = 1, 2, ..., h, h represents the total number of attention heads, SW-SA represents the operation of spatial window self-attention, X represents the second shallow feature of the image to be processed, and W p Denote the projection matrix for mapping the fused feature.
[0090] In the present invention, by concatenating the output features of multiple attention heads and performing linear projection to generate the spatial window self-attention feature, it is possible to fuse the multi-scale and multi-perspective features extracted by different attention heads, making the generated features more rich and comprehensive.
[0091] In a possible implementation manner, generating the self-attention feature of the channel window in S6 specifically includes:
[0092] S605: According to the channel window self-attention mechanism, generate the channel query matrix, channel key matrix, and channel value matrix of the to-be second shallow feature:
[0093] Q C = XW Q ,
[0094] K C = XW K ,
[0095] V C = XW V ,
[0096] Among them, Q C Denote the channel query matrix of the second shallow feature, K C Denote the channel key matrix of the second shallow feature, V C Denote the channel value matrix of the second shallow feature, X represents the second shallow feature of the image to be processed, and W Q Denote the query weight matrix, and W K Denote the key weight matrix, and W V Denote the value weight matrix.
[0097] S606: Divide the channel query matrix, channel key matrix, and channel value matrix into non-overlapping windows respectively, and flatten each window and decompose it into multiple channel attention heads:
[0098]
[0099] Among them, Denote the channel query matrix in the j-th channel attention head, Denote the channel key matrix in the j-th channel attention head, denotes the channel value matrix in the j-th channel attention head, where j = 1, 2, …, h, and h represents the total number of channel attention heads.
[0100] S607: Calculate the output features of each channel attention head:
[0101]
[0102] where, denotes the output feature of the i-th channel attention head, denotes the channel query matrix of the i-th channel attention head, denotes the channel key matrix of the i-th channel attention head, denotes the channel value matrix of the i-th channel attention head, softmax() represents the activation function, T represents the transpose, and β represents the scaling factor.
[0103] S608: Concatenate the output features of each channel attention head to generate the channel window self-attention feature:
[0104]
[0105] SW - SA(X) = Y C W p
[0106] where, Y C denotes the channel window self-attention feature, concat represents the concatenation operation, h represents the total number of attention heads, SW - SA represents the operation of channel window self-attention, and W p denotes the projection matrix for mapping the fused features.
[0107] Specifically, the model processes the input feature map through the spatial self-attention mechanism. It calculates the correlation between each pixel in the feature map and other pixels to generate a spatial attention weight matrix. Then, this weight matrix is multiplied by the original feature map to obtain the weighted feature map. This step helps the model focus on the key regions in the image and extract the spatial features closely related to the task. Next, the model further processes the weighted feature map using the channel self-attention mechanism. It calculates the correlation between different channels to generate a channel attention weight vector. Then, this weight vector is multiplied by each channel of the feature map to obtain a new weighted feature map.
[0108] In the present invention, through the channel window self-attention mechanism, a channel query matrix, a key matrix, and a value matrix are dynamically generated, and after being divided into multiple attention heads, the attention weights are calculated respectively, which can capture the global dependencies between channels and the correlations between specific channels. At the same time, by weighted summing and concatenating the output features of each attention head, more expressive channel window self-attention features are generated, effectively integrating the important information of different channels.
[0109] Furthermore, by using the channel window self-attention, the modeling ability of the channel dimension information is improved, the expression quality of the features is enhanced, which helps to more accurately extract the key features between channels in complex tasks with multiple scales and multiple perspectives, and at the same time improves the performance of the model and the adaptability to diverse data.
[0110] S7: Based on the spatial window self-attention features and the channel window self-attention features, perform feature fusion to obtain fused features.
[0111] It should be noted that the spatial self-attention mechanism mainly focuses on the dependencies between different positions in an image or feature map. By calculating the similarity between different positions, the model can capture the spatial context information in the image, thereby enhancing the understanding of local features. The channel self-attention mechanism, on the other hand, focuses on the dependencies between different channels. By calculating the correlations between channels, the model can learn which channels are more important for the task, and thus dynamically adjust the channel weights to improve the expression ability of the features. By combining the spatial self-attention and the channel self-attention, the adaptive module can more comprehensively understand the input features and capture the complex dependencies in the image.
[0112] In a possible implementation manner, S7 specifically includes:
[0113] S701: Perform convolution operations on the spatial window self-attention features and the channel window self-attention features respectively, and through an activation function, generate a spatial attention map and a channel attention map:
[0114] S-Map(B) = f(W2σ(W1B))
[0115] C-Map(B) = f(W4σ(W3H GP (B)))
[0116] Wherein, S-Map(B) represents the spatial attention map, f() represents the sigmoid function, σ represents the non-linear activation function, W1B represents the result of multiplying the input feature map B by the first weight matrix W1 of the spatial window, C-Map(B) represents the channel attention map, W2 represents the second weight matrix of the spatial window, W3 represents the third weight matrix of the channel window, W4 represents the fourth weight matrix of the channel window, H GPRepresents global average pooling.
[0117] S702: Re-weight the spatial window self-attention feature and the channel window self-attention feature according to the spatial attention map and the channel attention map:
[0118] S-I(A,B) = A ⊙ S-Map(B)
[0119] C-I(A,B) = A ⊙ C-Map(B)
[0120] Wherein, S-I represents spatial interaction, C-I represents channel interaction, ⊙ represents element-wise multiplication, and A represents the input feature to be weighted.
[0121] In the present invention, spatial attention maps and channel attention maps are generated through convolution operations, important information in the spatial and channel dimensions is extracted respectively, and the saliency is further refined through an activation function (such as sigmoid). This re-weighted feature can effectively highlight the key information in the input image while suppressing unimportant or redundant information, thereby improving the expression quality of the feature.
[0122] S703: Through the spatial and channel interaction attention mechanism, perform feature fusion on the re-weighted spatial window self-attention feature and the channel window self-attention feature to obtain a fused feature:
[0123] AS-SA(X) = (C-I(Y s ,Y w ) + S-I(Y w ,Y s ))W p
[0124] AC-SA(X) = (S-I(Y c ,Y w ) + C-I(Y w ,Y c ))W p
[0125] Wherein, AS-SA represents the adaptive spatial self-attention mechanism, AC-SA represents the adaptive channel self-attention mechanism, Y s represents the spatial window self-attention feature, Y w represents the global feature map, Y c represents the channel window self-attention feature, W p represents the projection matrix for mapping the fused feature.
[0126] In the present invention, through the synergistic effect of spatial interaction and channel interaction, combined with the projection matrix to generate a unified fused feature, this way of feature interaction significantly improves the model's ability to capture multi-scale and multi-dimensional feature relationships.
[0127] S8: Use the spatial gating feed-forward network to enhance the fused features to obtain enhanced features.
[0128] Among them, the spatial gating feed-forward network (SGFN) is a module that combines the spatial attention mechanism and the feed-forward network, and is used to enhance the feature expression ability. It optimizes the feature representation of the model by dynamically selecting and enhancing key features in the spatial dimension while suppressing redundant information.
[0129] In a possible implementation, S8 specifically includes:
[0130] S801: Decompose the fused features along the channel dimension to obtain the first fused feature and the second fused feature:
[0131]
[0132] Among them, represents the final features of the first fused feature and the second fused feature, represents the projection matrix for linearly transforming the fused features along the channel dimension, represents the original fused features, represents the first fused feature, represents the second fused feature, and σ represents the non-linear activation function.
[0133] S802: Use depth convolution to extract the final features of the first fused feature and the second fused feature:
[0134]
[0135] Among them, represents the feature extraction operation on the feature , W p represents the projection matrix for mapping the fused features, W d represents the projection matrix for depth convolution, and ⊙ represents element-wise multiplication.
[0136] S803: According to the final features, use the adaptive self-attention mechanism to determine the enhanced features:
[0137] X′ l = A - SA(LN(X l-1 )) + X l-1
[0138] X l = SGFN(LN(X l ′)) + X l ′
[0139] Among them, X′ lrepresents the feature vector after the linear transformation of the l-th layer, A-SA represents the adaptive self-attention mechanism, LN represents the Layer-Norm layer, and X l-1 represents the final feature of the (l-1)-th layer, and X l represents the enhanced feature of the l-th layer, and SGFN represents the feature extraction operation.
[0140] In the present invention, by decomposing the fused feature into two parts and using the methods of depth convolution and element-wise weighting, the important parts in the feature are effectively learned and enhanced, and redundant information is suppressed. At the same time, through the adaptive self-attention mechanism and the normalization operation, the spatial feed-forward network can dynamically model the context dependencies between the input features.
[0141] Further, the spatial feed-forward network combines the final feature with the enhanced intermediate feature through the residual connection, ensuring that the original information is retained while further optimizing the fused feature. This way significantly improves the overall quality of the feature and provides a more accurate input for subsequent tasks.
[0142] S9: According to the first shallow feature and the enhanced feature, perform image reconstruction on the image to be processed to obtain a reconstructed image.
[0143] In a possible implementation manner, S9 specifically includes:
[0144] S901: Concatenate the first shallow feature and the enhanced feature to generate a concatenated fused feature:
[0145] F fused = Concat(X l )
[0146] where F fused represents the concatenated fused feature, Concat represents the concatenation operation, X represents the first shallow feature, and X l represents the enhanced feature of the l-th layer.
[0147] S902: According to the concatenated fused feature, perform image reconstruction on the image to be processed to obtain a reconstructed image:
[0148] I HR = H RC (F fused )
[0149] where I HR represents the reconstructed image, and H RC represents a reconstruction module composed of a 3×3 convolutional layer and a pixel shuffle convolutional layer.
[0150] Specifically, we generated a feature map containing rich information through splicing operations. Next, we utilized this feature map for image reconstruction. To achieve this goal, we designed a reconstruction module composed of a 3×3 convolutional layer and a pixel shuffle convolutional layer. This reconstruction module can effectively utilize the information in the feature map and gradually restore the details and structure of the image through convolutional operations and pixel shuffle operations, ultimately outputting a high-quality image.
[0151] In the present invention, the first shallow feature is spliced with the enhanced feature to generate a fusion feature map with rich information. The combination of such multi-level features can effectively integrate local and global information, significantly enhancing the ability to restore details and textures. At the same time, the fusion feature map contains important information related to the task. By designing a reconstruction module composed of a 3×3 convolutional layer and a pixel shuffle layer, the features can be flexibly extracted and mapped.
[0152] Furthermore, the convolutional operation strengthens the feature expression, while the pixel shuffle operation gradually improves the resolution, making the details of the reconstructed image more abundant and the structure more complete.
[0153] The beneficial effects brought by the technical solution provided by the embodiments of the present invention at least include:
[0154] (1) In the present invention, through the initialization and optimization of the sparse dictionary, an adaptive dictionary is generated, effectively improving the ability to capture complex textures and special structures; combining the adaptive dictionary with the convolutional neural network to deeply extract shallow features and strengthening the representation of key details; introducing the spatial and channel self-attention mechanisms to generate multi-dimensional attention features, and further enhancing the feature expression through feature fusion, thus significantly improving the quality of image reconstruction.
[0155] (2) In the present invention, the second shallow feature is extracted through the convolutional network, and combined with the spatial window self-attention mechanism and the channel window self-attention mechanism to generate spatial window self-attention features and channel window self-attention features respectively. Further, based on the fusion of the two types of attention features, a fusion feature is formed. Finally, the spatial feed-forward network is used to selectively strengthen the fusion feature, effectively enhancing the interaction between shallow and deep features and significantly improving the accuracy and comprehensiveness of image feature extraction.
[0156] Refer to the attached Figure 2 description, which shows a schematic structural diagram of an image processing system based on an adaptive dictionary provided by the present invention.
[0157] The present invention also provides an image processing system 20 based on an adaptive dictionary, which is applied to the above-mentioned image processing method based on an adaptive dictionary, and includes:
[0158] A processor 201.
[0159] Memory 202 stores computer-readable instructions thereon. When the computer-readable instructions are executed by the processor 201, an image processing method based on an adaptive dictionary as in the method embodiment is implemented.
[0160] The image processing system 20 based on an adaptive dictionary provided by the present invention can execute the above-mentioned image processing method based on an adaptive dictionary and achieve the same or similar technical effects. To avoid repetition, the present invention will not elaborate further.
[0161] The beneficial effects brought by the technical solution provided by the embodiments of the present invention at least include:
[0162] (1) In the present invention, through the initialization and optimization of the sparse dictionary, an adaptive dictionary is generated, effectively improving the ability to capture complex textures and special structures; combining the adaptive dictionary with a convolutional neural network to deeply extract shallow features and strengthening the representation of key details; introducing spatial and channel self-attention mechanisms to generate multi-dimensional attention features, and further improving the feature expressiveness through feature fusion, thereby significantly improving the quality of image reconstruction.
[0163] (2) In the present invention, the second shallow features are extracted through a convolutional network, and combined with the spatial window self-attention mechanism and the channel window self-attention mechanism to generate spatial window self-attention features and channel window self-attention features respectively. Further, the two types of attention features are fused to form a fused feature. Finally, the spatial feed-forward network is used to selectively strengthen the fused feature, effectively enhancing the interaction between shallow and deep features and significantly improving the accuracy and comprehensiveness of image feature extraction.
[0164] It should be understood that the processor in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0165] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0166] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be solid-state drives.
[0167] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be understood specifically by referring to the context before and after.
[0168] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0169] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0170] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0171] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices, apparatuses, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0172] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0173] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0174] In addition, the functional units in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0175] When the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0176] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the image processing method based on an adaptive dictionary as described in the method embodiment.
[0177] The computer-readable storage medium provided by the present invention can implement the steps and effects of the image processing method based on an adaptive dictionary in the above method embodiment. To avoid repetition, the present invention will not elaborate further.
[0178] The beneficial effects brought by the technical solution provided by the embodiments of the present invention at least include:
[0179] (1) In the present invention, through the initialization and optimization of the sparse dictionary, an adaptive dictionary is generated, effectively improving the ability to capture complex textures and special structures; combining the adaptive dictionary with the convolutional neural network to deeply extract shallow features and strengthening the representation of key details; introducing the spatial and channel self-attention mechanisms to generate multi-dimensional attention features, and further improving the feature expression ability through feature fusion, thereby significantly improving the quality of image reconstruction.
[0180] (2) In the present invention, the second shallow features are extracted through the convolutional network, and combined with the spatial window self-attention mechanism and the channel window self-attention mechanism to generate spatial window self-attention features and channel window self-attention features respectively. Further, the two types of attention features are fused to form a fusion feature. Finally, the spatial feed-forward network is used to selectively strengthen the fusion feature, effectively enhancing the interaction between shallow and deep features and significantly improving the accuracy and comprehensiveness of image feature extraction.
[0181] As described above, this is only the specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily conceive of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
[0182] The following points need to be explained:
[0183] (1) The drawings of the embodiments of the present invention only relate to the structures involved in the embodiments of the present invention, and other structures can refer to the general design.
[0184] (2) For clarity, in the drawings used to describe the embodiments of the present invention, the thickness of layers or regions is enlarged or reduced, that is, these drawings are not drawn to actual scale. It can be understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, the element can be "directly" on or under the other element or there can be intervening elements.
[0185] (3) Without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.
[0186] As above, this is only the specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. An image processing method based on an adaptive dictionary, characterized in that, Including: S1: Obtain the image to be processed; S2: Initialize the sparse dictionary, and utilize the sparse dictionary to extract the sparse features of the image to be processed; S3: Based on the sparse features, perform update and optimization processing on the sparse dictionary to generate an adaptive dictionary; S4: Utilize the adaptive dictionary to extract the first shallow features of the image to be processed; S5: Input the first shallow features into a convolutional layer neural network to extract the second shallow features of the image to be processed; S6: Based on the second shallow features, respectively generate spatial window self-attention features and channel window self-attention features by using the spatial window self-attention mechanism and the channel window self-attention mechanism; S7: Based on the spatial window self-attention features and the channel window self-attention features, perform feature fusion to obtain fused features; S8: Utilize a spatial feed-forward network to strengthen the fused features to obtain strengthened features; S9: According to the first shallow features and the strengthened features, perform image reconstruction on the image to be processed to obtain a reconstructed image.
2. The image processing method based on an adaptive dictionary according to claim 1, wherein The specific content of S2 is as follows: Utilize the sparse dictionary to extract the sparse features of the image to be processed through the following formula: Among them, α * represents the sparse feature of the image to be processed, and arg min α represents the sparse feature corresponding to the minimum value of the objective function, represents the sparse dictionary, M represents the number of atoms in the sparse dictionary, d represents the dimension of the sparse dictionary, x represents the image to be processed, represents the square of the two-norm, λ∥α∥1 represents the parameter for controlling the sparse feature, λ represents the regularization parameter, ∥α∥1 represents the sum of the absolute values of all elements in α, and α represents the sparse feature.
3. The image processing method based on an adaptive dictionary according to claim 1, wherein The specific content of S3 is as follows: Based on the sparse features, perform update and optimization processing on the sparse dictionary through the following formula to generate an adaptive dictionary: Among them, D represents the adaptive dictionary, and x represents the image to be processed. represents the sparse dictionary at the t-th iteration. represents the square of the reconstruction error, λ represents the regularization parameter, ∥α∥1 represents the sum of the absolute values of all elements in α, α represents the sparse feature, arg min D represents the sparse dictionary corresponding to when the objective function obtains the minimum value.
4. The image processing method based on an adaptive dictionary according to claim 1, wherein The specific content of S5 is as follows: Input the first shallow features into a convolutional layer neural network to extract the second shallow features of the image to be processed through the following formula: Among them, X represents the second shallow feature of the image to be processed, represents the first shallow feature of the image to be processed, Conv 3×3 represents a convolution operation with a convolution kernel size of 3×3.
5. The image processing method based on an adaptive dictionary according to claim 1, wherein The generation of the spatial window self-attention features in S6 specifically includes: S601: Generate the query matrix, key matrix, and value matrix of the second shallow features according to the spatial window self-attention mechanism: Q s = XW Q K s = XW K V s = XW V Among them, Q s represents the query matrix of the second shallow feature, K s represents the key matrix of the second shallow feature, V s represents the value matrix of the second shallow feature, X represents the second shallow feature of the image to be processed, W Q represents the query weight matrix, W K represents the key weight matrix, W V represents the value weight matrix; S602: Divide the query matrix, the key matrix, and the value matrix into non-overlapping windows respectively, and flatten each window and decompose it into multiple attention heads: Among them, represents the query matrix of the i-th attention head, represents the key matrix of the i-th attention head, represents the value matrix of the i-th attention head, where i = 1, 2, …, h, and h represents the total number of attention heads; S603: Calculate the output features of each of the attention heads: Among them, represents the output feature of the i-th attention head, represents the query matrix of the i-th attention head, represents the key matrix of the i-th attention head, represents the value matrix of the i-th attention head, e represents the size of each attention head, C represents the dimension of the channel, h represents the total number of attention heads, softmax() represents the activation function, T represents the transpose operation, and E represents the relative position encoding; S604: Concatenate the output features of each of the attention heads to generate the spatial window self-attention features: SW-SA(X) = Y s W p Among them, Y s represents the spatial window self-attention feature, and concat represents the concatenation operation. represents the output feature of the i-th attention head, where i = 1, 2,..., h, h represents the total number of attention heads, SW-SA represents the operation of spatial window self-attention, X represents the second shallow feature of the image to be processed, and W p represents the projection matrix for mapping the fused feature.
6. The image processing method based on an adaptive dictionary according to claim 1, wherein The generation of the channel window self-attention features in S6 specifically includes: S605: Generate the channel query matrix, channel key matrix, and channel value matrix of the second shallow features to be processed according to the channel window self-attention mechanism: Q C = XW Q , K C = XW K , V C = XW V , Among them, Q C represents the channel query matrix of the second shallow feature, K C represents the channel key matrix of the second shallow feature, V C represents the channel value matrix of the second shallow feature, X represents the second shallow feature of the image to be processed, W Q represents the query weight matrix, W K represents the key weight matrix, W V represents the value weight matrix; S606: Divide the channel query matrix, the channel key matrix, and the channel value matrix into non-overlapping windows respectively, and flatten each window and decompose it into multiple channel attention heads: Among them, represents the channel query matrix in the j-th channel attention head, represents the channel key matrix in the j-th channel attention head, represents the channel value matrix in the j-th channel attention head, where j = 1, 2, …, h, and h represents the total number of channel attention heads; S607: Calculate the output features of each of the channel attention heads: Among them, represents the output feature of the i-th channel attention head, represents the channel query matrix of the i-th channel attention head, represents the channel key matrix of the i-th channel attention head, represents the channel value matrix of the i-th channel attention head, softmax() represents the activation function, T represents the transpose, and β represents the scaling factor; S608: Concatenate the output features of each of the channel attention heads to generate the channel window self-attention features: SW-SA(X) = Y C W p Among them, Y C represents the channel window self-attention feature, concat represents the concatenation operation, h represents the total number of attention heads, SW-SA represents the operation of channel window self-attention, and W p represents the projection matrix for mapping the fused features.
7. The image processing method based on an adaptive dictionary according to claim 1, wherein The specific content of S7 includes: S701: Respectively perform convolutional operations on the spatial window self-attention features and the channel window self-attention features, and generate a spatial attention map and a channel attention map through an activation function: S-Map(B) = f(W2σ(W1B)) C-Map(B) = f(W4σ(W3H GP (B))) Among them, S-Map(B) represents the spatial attention map, f() represents the sigmoid function, σ represents the non-linear activation function, W1B represents the result of multiplying the input feature map B by the first weight matrix W1 of the spatial window, C-Map(B) represents the channel attention map, W2 represents the second weight matrix of the spatial window, W3 represents the third weight matrix of the channel window, W4 represents the fourth weight matrix of the channel window, H GP represents global average pooling; S702: Re-weight the spatial window self-attention feature and the channel window self-attention feature according to the spatial attention map and the channel attention map: S-I(A,B) = A ⊙ S-Map(B) C-I(A,B) = A ⊙ C-Map(B) where S-I represents spatial interaction, C-I represents channel interaction, ⊙ represents element-wise multiplication, and A represents the input feature to be weighted; S703: Through the spatial and channel interaction attention mechanism, perform feature fusion on the re-weighted spatial window self-attention feature and channel window self-attention feature to obtain a fused feature: AS-SA(X)=(C-I(Y s ,Y w )+S-I(Y w ,Y s ))W p AC-SA(X)=(S-I(Y c ,Y w )+C-I(Y w ,Y c ))W p Among them, AS-SA represents the adaptive spatial self-attention mechanism, AC-SA represents the adaptive channel self-attention mechanism, Y s represents the spatial window self-attention feature, Y w represents the global feature map, Y c represents the channel window self-attention feature, W p represents the projection matrix.
8. The image processing method based on an adaptive dictionary according to claim 1, characterized in that The specific steps of S8 include: S801: Decompose the fused feature along the channel dimension to obtain a first fused feature and a second fused feature: Among them, represents the fused feature obtained through linear transformation, represents the projection matrix used for linear transformation of the fused feature in the channel dimension, represents the original fused feature, represents the first fused feature, represents the second fused feature, and σ represents the non-linear activation function; S802: Use depth convolution to extract the final features of the first fused feature and the second fused feature: Among them, represents the final feature of the first fusion feature and the second fusion feature, W p represents the projection matrix for mapping the fused features, W d represents the projection matrix of depth convolution, and ⊙ represents element-wise multiplication; S803: According to the final features, use the adaptive self-attention mechanism to determine the enhanced feature: X′ = A - SA(LN(X l-1 )) + X l-1 X l = SGFN(LN(X l ′)) + X l ′ Among them, X' l represents the feature vector after the linear transformation of the l-th layer, A-SA represents the adaptive self-attention mechanism, LN represents the Layer-Norm layer, and X l-1 represents the final feature of the (l-1)-th layer, and X l represents the enhanced feature of the l-th layer, and SGFN represents the feature extraction operation.
9. The image processing method based on an adaptive dictionary according to claim 1, wherein, The specific steps of S9 include: S901: Concatenate the first shallow feature and the enhanced feature to generate a concatenated fused feature: F fused = Concat([X,X l ) Among them, F fused represents the splicing and fusion feature, Concat([·,·]) represents the splicing operation, X represents the first shallow feature, and X l represents the enhanced feature of the l-th layer; S902: According to the concatenated fused feature, perform image reconstruction on the image to be processed to obtain a reconstructed image: I HR = H RC (F fused ) Among them, I HR represents the reconstructed image, and H RC represents a reconstruction module composed of a 3×3 convolutional layer and a pixel shuffle convolutional layer.
10. An image processing system based on an adaptive dictionary, characterized in that, Comprising: A processor; A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the image processing method based on an adaptive dictionary as described in any one of claims 1 to 9 is implemented.