Image processing method and device, electronic equipment, storage medium and program product

By performing self-attention mechanism processing and cross-attention enhancement at multiple image scales, combined with a channel gate feedforward network, multi-scale image features are extracted, which solves the shortcomings of single-scale feature extraction and achieves more efficient and accurate image processing.

CN120707880APending Publication Date: 2025-09-26BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510859060.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The existing image processing method based on the self-attention mechanism adopts single-scale feature extraction, which results in the extracted features being relatively single and lacking in richness and comprehensiveness, thereby affecting the accuracy and robustness of image processing.

Method used

The original image features are processed by the self-attention mechanism at multiple image scales to obtain multi-scale image features, and image processing is performed through multi-scale features, including linear mapping, aggregation and cross-attention enhancement, combined with the channel gate feedforward network for feature extraction.

Benefits of technology

The accuracy and robustness of image processing are improved, and the image processing effects are enhanced, especially in detail restoration, edge enhancement and noise removal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707880A_ABST
    Figure CN120707880A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing method and device, electronic equipment, a storage medium and a program product. The image processing method comprises the following steps: acquiring original image features of a to-be-processed image; under multiple image scales, self-attention mechanism processing is carried out on the original image features to obtain multi-scale image features, and the multiple image scales are different from the scale of the image to be processed; and processing the to-be-processed image at least according to the multi-scale image features to obtain a target image. According to the technical scheme, the accuracy and robustness of image processing can be improved, and the image processing effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and in particular to an image processing method, device, electronic device, storage medium, and program product. Background Art

[0002] As image processing technology has advanced, so too has the self-attention mechanism. The self-attention mechanism is an attention mechanism that associates different positions in a single sequence to compute a representation of the same sequence. It allows the model to dynamically adjust the focus on each element when processing sequential data, thereby capturing complex dependencies within the sequence. Consequently, the self-attention mechanism can be applied to the Transformer model architecture to extract image features. Summary of the Invention

[0003] To overcome the problems existing in the related art, the present disclosure provides an image processing method, apparatus, electronic device, storage medium and program product.

[0004] According to a first aspect of an embodiment of the present disclosure, there is provided an image processing method, comprising: obtaining original image features of an image to be processed; performing self-attention mechanism processing on the original image features at multiple image scales to obtain multi-scale image features, wherein the multiple image scales are different from the scales of the image to be processed; and processing the image to be processed at least according to the multi-scale image features to obtain a target image.

[0005] By acquiring the original image features of the image to be processed, the self-attention mechanism is applied to these features at multiple scales different from the scale of the image to obtain multi-scale image features. Based on these multi-scale image features, the image to be processed is processed to obtain the target image. The self-attention mechanism can extract richer, multi-layered image features at different image scales. Compared to feature extraction at a single scale, this method can produce richer, more comprehensive, and deeper multi-scale image features. Furthermore, using these multi-scale image features for image processing can improve the accuracy and robustness of image processing, enhancing the image processing effect.

[0006] In some possible implementations, the self-attention mechanism is performed on the original image features at multiple image scales to obtain multi-scale image features, including: linearly mapping the original image features at multiple image scales to obtain linear mapping vectors corresponding to the multiple image scales; aggregating the linear mapping vectors corresponding to the multiple image scales to obtain initial multi-scale image features; and cross-attention enhancement processing is performed on the initial multi-scale image features to obtain target multi-scale image features.

[0007] Through linear mapping of multiple image scales, linear mapping vector aggregation and cross-attention enhancement, rich, comprehensive and deep multi-scale image features can be obtained.

[0008] In some possible implementations, the linear mapping vectors corresponding to the multiple image scales include: a query vector, a first key vector, and a first value vector corresponding to the first image scale, and a second key vector and a second value vector corresponding to the second image scale. The linear mapping vectors corresponding to the multiple image scales are aggregating to obtain initial multi-scale image features, including: determining a first self-attention map based on the query vector and the first key vector; determining a second self-attention map based on the query vector and the second key vector; determining a multi-scale self-attention map based on the first self-attention map and the second self-attention map; determining a multi-scale eigenvalue based on the first value vector and the second value vector; and determining the initial multi-scale image features based on the multi-scale self-attention map and the multi-scale eigenvalue.

[0009] The first self-attention map can be understood as the self-attention map between image patches of the same scale, and the second self-attention map can be understood as the self-attention map between image patches of different scales. Therefore, multi-scale self-attention maps can reflect the relationship between features from both the same scale and cross-scale dimensions; and multi-scale eigenvalues ​​can reflect the relationship between features of different scales. This allows for the identification of richer, more comprehensive, and deeper multi-scale image features.

[0010] In some possible implementations, the cross-attention enhancement processing is performed on the initial multi-scale image features to obtain target multi-scale image features, including: determining a third self-attention map based on the second self-attention map; determining new multi-scale image features based on the third self-attention map and the second value vector; and determining target multi-scale image features based on the initial multi-scale image features and the new multi-scale image features.

[0011] Through cross-attention enhancement processing, the initial multi-scale image features can be further refined to achieve deep feature extraction.

[0012] In some possible implementations, processing the image to be processed at least based on the multi-scale image features to obtain a target image includes: determining a first image feature at least based on the multi-scale image features; processing the first image feature through a channel gate feedforward network to obtain a second image feature; determining a target image feature based on the first image feature and the second image feature; and processing the image to be processed based on the target image feature to obtain a target image.

[0013] Through the channel gate feedforward network, global features can be extracted at different channel dimensions, and then global features and multi-scale image features can be combined to obtain richer and more comprehensive image features.

[0014] In some possible implementations, determining the first image feature at least based on the multi-scale image feature includes: performing convolution processing on the original image feature to obtain a third image feature; performing channel attention mechanism processing on the original image feature to obtain a fourth image feature; and determining the first image feature based on the multi-scale image feature, the third image feature, and the fourth image feature.

[0015] Through channel attention mechanism processing, the ability to focus on long-distance global information can be improved, and through deep convolution processing, local information can be emphasized and the perception of local features can be enhanced.

[0016] In some possible implementations, the first image feature is processed through a channel gate feedforward network to obtain a second image feature, and the second image feature is obtained, including: dividing the first image feature according to multiple channel dimensions to obtain image features corresponding to multiple channel dimensions respectively; processing the image features corresponding to the multiple channel dimensions respectively through a channel gate feedforward network to obtain the second image feature.

[0017] By first dividing the first image features in multiple channel dimensions, the channel gate feedforward network can process them separately based on the divided features, thereby improving computational efficiency and improving feature extraction effects in different channel dimensions.

[0018] In some possible embodiments, the channel gate feedforward network includes a first network branch and a second network branch, and the image features corresponding to the multiple channel dimensions are processed through the channel gate feedforward network to obtain the second image feature. Obtaining the second image feature includes: performing local feature extraction on the image features corresponding to the first channel dimension through the first network branch to obtain local image features; performing global feature extraction on the image features corresponding to the second channel dimension through the second network branch to obtain global image features; and determining the second image feature based on the local image features and the global image features.

[0019] Different network branches are used to extract local features and global features respectively. Then, the input feature map is divided into two branches along the channel dimension, each branch occupies half of the feature channel, reducing channel redundancy and making it easier to capture nonlinear feature information.

[0020] In some possible implementations, the image to be processed is an image obtained by decoding the video to be played, and processing the image to be processed at least based on the multi-scale image features to obtain the target image includes: performing image enhancement processing on the image to be processed at least based on the multi-scale image features to obtain the target image.

[0021] Image enhancement processing through richer, more comprehensive and deeper image features can improve the resolution and quality of the image and solve the limitations in detail restoration, edge enhancement and noise removal.

[0022] According to a second aspect of an embodiment of the present disclosure, an image processing device is provided, for executing the image processing method described in the first aspect.

[0023] According to a third aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the program instructions are executed by a processor, the image processing method described in the first aspect of the present disclosure is implemented.

[0024] According to a fourth aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to: execute the executable instructions to implement the image processing method as described in the first aspect of the present disclosure.

[0025] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the image processing method as described in the first aspect of the present disclosure.

[0026] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0028] Figure 1 The figure is a flowchart of an image processing method according to an exemplary embodiment.

[0029] Figure 2 A framework for self-attention mechanism processing is shown according to an example embodiment.

[0030] Figure 3 It is a block diagram of a multi-type feature integration module according to an exemplary embodiment.

[0031] Figure 4It is a block diagram of a deep feature extraction module according to an exemplary embodiment.

[0032] Figure 5 The figure is a block diagram of a channel gate feedforward network according to an exemplary embodiment.

[0033] Figure 6 FIG. 4 is a schematic diagram of a Transformer network architecture according to an exemplary embodiment.

[0034] Figure 7 The figure is a block diagram of an image processing apparatus according to an exemplary embodiment.

[0035] Figure 8 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0036] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0037] It should be noted that all actions of acquiring signals, information or data in the present disclosure are carried out in compliance with the corresponding data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.

[0038] As image processing technology has advanced, so too has image processing technology based on the self-attention mechanism. The self-attention mechanism is an attention mechanism that associates different positions in a single sequence to compute a representation of the same sequence. It allows the model to dynamically adjust the focus on each element when processing sequential data, thereby capturing complex dependencies within the sequence. Consequently, the self-attention mechanism can be applied to the Transformer model architecture to extract image features.

[0039] In related technologies, the self-attention module based on the self-attention mechanism is used as the core attention module of the Transformer model architecture, and features are extracted based on the original image scale to obtain corresponding image features.

[0040] In this image processing approach, the self-attention module uses a single-scale feature extraction method. The extracted features are relatively simple and lack richness and comprehensiveness. Furthermore, the image processing effect of using this single-scale image feature is also poor.

[0041] Based on this, the embodiment of the present disclosure provides a technical solution by obtaining the original image features of the image to be processed, performing self-attention mechanism processing on the original image features at multiple image scales different from the scale of the image to be processed, and obtaining multi-scale image features. Based on the multi-scale image features, the image to be processed is processed to obtain a target image.

[0042] At different image scales, the self-attention mechanism can extract richer multi-level image features. Compared with the feature extraction of a single scale, richer, comprehensive and deep multi-scale image features can be obtained. Furthermore, using this multi-scale image feature for image processing can improve the accuracy and robustness of image processing and enhance the image processing effect.

[0043] The technical solutions of the disclosed embodiments can be applied to multiple scenarios, including medical imaging, satellite and remote sensing image processing, digital entertainment and content creation, security and monitoring, consumer electronics, autonomous driving, document and OCR (Optical Character Recognition) processing, scientific research and academic analysis, e-commerce and retail.

[0044] In these scenarios, the image processing technology of the embodiments of the present disclosure can solve the limitations in detail restoration, edge enhancement, and noise removal, achieve more efficient and accurate image reconstruction and enhancement, and provide reliable technical support for various industries in detail preservation and visual optimization.

[0045] Taking low-resolution images and image environments interfered with by noise as an example, the technical solutions of the embodiments of the present disclosure can significantly improve image quality and meet the needs of image detail restoration and visual effect enhancement in actual scenarios.

[0046] Figure 1 is a flowchart of an image processing method according to an exemplary embodiment. Figure 1 As shown, the image processing method is used in a terminal and includes the following steps: Step S11: obtaining original image features of the image to be processed.

[0047] Step S12: performing a self-attention mechanism on the original image features at multiple image scales to obtain multi-scale image features, wherein the multiple image scales are different from the scales of the image to be processed.

[0048] Step S13: Process the image to be processed at least according to the multi-scale image features to obtain a target image.

[0049] Since the technical solutions of the embodiments of the present disclosure can be applied to different image processing scenarios, the types of images to be processed and the processing requirements may be different in different application scenarios.

[0050] For example, in a video playback scenario, the image to be processed can be a video frame obtained by decoding the video to be played. In a security and surveillance scenario, the image to be processed can be a surveillance image that needs to be analyzed. In an autonomous driving scenario, the image to be processed can be an environmental image used for vehicle perception.

[0051] In the above application scenarios, the image to be processed may be a low-quality image with low resolution or low image quality.

[0052] In some embodiments, the original image features of the image to be processed may be image features obtained by performing shallow feature extraction on the image to be processed. Shallow feature extraction may be implemented through a convolutional layer or a deep convolutional layer.

[0053] As an example, suppose the image to be processed is processed by I LR Represents, then, through shallow feature extraction, the shallow feature (i.e., original image feature) F0 can be obtained Where H and W represent the height and width of the image respectively, 3 represents the number of channel input images, and C represents the number of channel feature maps (features).

[0054] In step S12, the original image features are processed by the self-attention mechanism at multiple image scales. It can also be understood that the original image features are processed by the self-attention mechanism according to multiple image scales. The obtained image features are multi-scale image features, which are deeper, richer and more comprehensive features compared to the original image features.

[0055] In some embodiments, a multi-scale feature attention module (or network) can be used to perform self-attention mechanism processing on the original image features at multiple image scales, wherein the multi-scale feature attention module can adopt the self-attention mechanism processing logic of the embodiment of the present disclosure.

[0056] In some embodiments, the multi-scale feature attention module can be a separate processing module or a processing module integrated into a multi-type feature integration module. The multi-type feature integration module may involve the extraction of multiple types of image features. Therefore, in addition to the extraction of multi-scale features, other types of feature extraction may also be involved.

[0057] It can be understood that whether the multi-scale feature attention module is a separate processing module or a processing module integrated in the multi-type feature integration module, its processing logic is the same. It’s just that in different scenarios, the input of the multi-scale feature attention module may be different, and the application of its output may be different. Therefore, the processing logic of the self-attention mechanism will be introduced next.

[0058] As an optional implementation, step S12 includes: performing linear mapping on the original image features at multiple image scales to obtain linear mapping vectors corresponding to the multiple image scales; aggregating the linear mapping vectors corresponding to the multiple image scales to obtain initial multi-scale image features; and performing cross-attention enhancement processing on the initial multi-scale image features to obtain target multi-scale image features.

[0059] In this implementation, rich, comprehensive, and deep multi-scale image features can be obtained through linear mapping of multiple image scales, linear mapping vector aggregation, and cross-attention enhancement.

[0060] In some embodiments, linear mapping is performed on the original image features at multiple image scales, which may include: downsampling the original image features to multiple image scales to obtain image features at multiple image scales, and then linearly mapping the image features at multiple image scales to obtain linear mapping vectors corresponding to the multiple image scales.

[0061] In some embodiments, the multiple image scales are different from the scale of the image to be processed. For example, the image scale is determined by r i , r1 can represent image scale 1, and r2 can represent image scale 2. If the original image scale is 1, r1 can be 2, and r2 can be 4. Here, 2 and 4 can be understood as image scales compared to the original image scale. In different application scenarios, other image scales can be set according to requirements and are not limited here.

[0062] In some embodiments, the original image features can be downsampled to different scales through average pooling operations with different step lengths (asynchronous long pooling). This feature downsampling method can reduce the computational complexity of self-attention and extract rich pooling features for different receptive fields.

[0063] In the self-attention mechanism, when performing linear mapping, features can be mapped to three vectors: Q, K, and V. Among them, Q represents the query vector, K represents the key vector, and V represents the value vector.

[0064] The original image features For example, by downsampling the original image features to different scales, we can get X i , where r i Represents a specific image scale. For example, , .

[0065] Furthermore, linear mapping vectors of different scales can be expressed as:

[0066]

[0067]

[0068]

[0069]

[0070] Among them, Q1, K1 and V1 are three linear mapping vectors corresponding to image scale 1, K2 and V2 are two linear mapping vectors corresponding to image scale 2, and image scale 1 and image scale 2 can share a Q vector, so the Q vector corresponding to image scale 2 is also Q1.

[0071] as well as, , , They are the parameters of the linear projection, which are preset parameters; and represent image scale 1 and image scale 2 respectively.

[0072] In some embodiments, linear mapping vectors corresponding to multiple image scales are aggregated to obtain initial multi-scale image features.

[0073] As an optional implementation, the linear mapping vectors corresponding to multiple image scales include: a query vector (Q1), a first key vector (K1) and a first value vector (V1) corresponding to a first image scale, and a second key vector (K2) and a second value vector (V2) corresponding to a second image scale.

[0074] Correspondingly, the linear mapping vectors corresponding to multiple image scales are aggregated to obtain initial multi-scale image features, including: determining a first self-attention map based on the query vector and the first key vector; determining a second self-attention map based on the query vector and the second key vector; determining a multi-scale self-attention map based on the first self-attention map and the second self-attention map; determining a multi-scale eigenvalue based on the first value vector and the second value vector; and determining the initial multi-scale image features based on the multi-scale self-attention map and the multi-scale eigenvalue.

[0075] In some embodiments, the first self-attention map can be understood as a self-attention map between image blocks of the same scale, and the second self-attention map can be understood as a self-attention map between image blocks of different scales.

[0076] Therefore, the multi-scale self-attention map can reflect the relationship between features from both the same scale and cross-scale dimensions; and the multi-scale eigenvalues ​​can reflect the relationship between features at different scales; thus, richer, more comprehensive and deeper multi-scale image features can be determined.

[0077] In some embodiments, the first self-attention map can be represented as the product of the query vector and the transposed vector of the first key vector, and the second self-attention map can be represented as the product of the query vector and the steered vector of the second key vector.

[0078] For example, , , where S1AP represents the first self-attention map and S2AP represents the second self-attention map. represents the transposed vector of the first bond vector, A vector representing the transposed vector of the second key vector.

[0079] In some embodiments, a multi-scale self-attention map can be obtained by concatenating the first self-attention map and the second self-attention map.

[0080] In some embodiments, the multi-scale feature value may be obtained by connecting the first value vector and the second value vector.

[0081] In some embodiments, the first self-attention map and the second self-attention map may be concatenated along a first dimension, and the first value vector and the second value vector may be concatenated along a second dimension to obtain multi-scale eigenvalues. The first dimension and the second dimension are different dimensions and correspond to the dimensions of width W. For example, the dimensions corresponding to width W include two dimensions, and the self-attention maps may be concatenated along the first dimension, and the value vectors may be concatenated along the second dimension to obtain multi-scale eigenvalues.

[0082] In some embodiments, matrix multiplication is calculated between the multi-scale self-attention map and the multi-scale eigenvalues ​​to obtain a multi-scale feature aggregation result, that is, an initial multi-scale image feature.

[0083] As an example, the multi-scale self-attention map MSFAP is expressed as: Concatenation

[0084] The multi-scale eigenvalue MSFV is expressed as: Concatenation

[0085] Then, the initial multi-scale image feature Y is expressed as:

[0086] Among them, Softmax is a function that converts any real vector into a probability distribution, and D represents the dimension of each self-attention head.

[0087] In some embodiments, the initial multi-scale image features can be further refined through cross-attention enhancement processing.

[0088] In some embodiments, cross-attention enhancement processing is performed on the initial multi-scale image features to obtain target multi-scale image features, including: determining a third self-attention map based on the second self-attention map; determining new multi-scale image features based on the third self-attention map and the second value vector; and determining the target multi-scale image features based on the initial multi-scale image features and the new multi-scale image features.

[0089] In this embodiment, a new self-attention map can be formed based on the second self-attention map.

[0090] For example, the most relevant k image blocks can be selected from the second self-attention map, and these k image blocks can be used to form a new self-attention map to obtain a third self-attention map.

[0091] In some embodiments, the third self-attention map may be matrix multiplied with the second value vector to obtain a new multi-scale image feature.

[0092] In some embodiments, the initial multi-scale image feature and the new multi-scale image feature may be added element by element to obtain the target multi-scale image feature.

[0093] For example, the new multi-scale image feature Z can be expressed as:

[0094] in, Represents the third self-attention map formed by selecting the most relevant k image patches from the second self-attention map.

[0095] Then, the target multi-scale image features It can be expressed as: .

[0096] Figure 2 is a framework diagram of a self-attention mechanism processing according to an exemplary embodiment. Figure 2 As shown, the input feature X , first through feature downsampling, get , Features at two image scales.

[0097] For X1, through linear mapping, we get the query vector (Q1), the first key vector (K1) and the first value vector (V1); for X2, the second key vector (K2) and the second value vector (V2).

[0098] Next, calculations are performed based on Q1 and K1 to obtain the self-attention map S1AP between image blocks of the same scale; and calculations are performed based on Q1 and K2 to obtain the self-attention map S2AP between image blocks of different scales.

[0099] Next, S1AP and S2AP are concatenated to obtain the multi-scale self-attention map MSFAP; and V1 and V2 are concatenated to obtain the multi-scale feature value MSFV.

[0100] Next, the multi-scale self-attention map MSFAP and the multi-scale eigenvalue MSFV are matrix multiplied to obtain the initial multi-scale image feature Y.

[0101] Then, a new self-attention map is formed through S2AP, and the new self-attention map is matrix multiplied with V2 to obtain a new multi-scale image feature Z.

[0102] Finally, by accumulating the initial multi-scale image feature Y and the new multi-scale image feature Z, the target multi-scale image feature can be obtained. .

[0103] The method of determining multi-scale image features based on the self-attention mechanism adopted in the embodiment of the present disclosure can not only reduce the complexity of self-attention calculation, but also extract rich, comprehensive and deep features at multiple scales.

[0104] In step S13, the image to be processed is processed at least according to the multi-scale image features to obtain a target image.

[0105] As an optional implementation, step S13 includes: determining a first image feature based on at least a multi-scale image feature; processing the first image feature through a channel gate feedforward network to obtain a second image feature; determining a target image feature based on the first image feature and the second image feature; and processing the image to be processed based on the target image feature to obtain a target image.

[0106] In this embodiment, the channel-gated feedforward network can extract features at different channel dimensions. Thus, based on the first image features, the channel-gated feedforward network can extract second image features, which are then integrated with the first image features to obtain target image features. Furthermore, the image to be processed can be processed based on the target image features.

[0107] Through the channel gate feedforward network, global features can be extracted at different channel dimensions, and then global features and multi-scale image features can be combined to obtain richer and more comprehensive image features.

[0108] In some embodiments, multi-scale image features may be used as a type of feature. In addition to multi-scale image features, other types of image features may also be involved.

[0109] As an optional implementation, determining the first image feature at least based on the multi-scale image feature includes: performing convolution processing on the original image feature to obtain a third image feature; performing channel attention mechanism processing on the original image feature to obtain a fourth image feature; and determining the first image feature based on the multi-scale image feature, the third image feature and the fourth image feature.

[0110] In this embodiment, the third image feature can be obtained through convolution processing, and the fourth image feature can be obtained through channel attention mechanism processing.

[0111] In some embodiments, when performing self-attention mechanism processing, convolution processing, and channel attention mechanism processing based on the original image features, the original image features can be layer-normalized first, and then based on the image features after layer normalization processing, self-attention mechanism processing, convolution processing, and channel attention mechanism processing can be performed respectively.

[0112] In some embodiments, the convolution processing may be a depthwise convolution processing, which may be implemented by a depthwise convolution module.

[0113] In some embodiments, channel attention mechanism processing can be implemented by a channel attention network.

[0114] Through channel attention mechanism processing, the ability to focus on long-distance global information can be improved, and through deep convolution processing, local information can be emphasized and the perception of local features can be enhanced.

[0115] In some embodiments, the normalization module, the self-attention mechanism module, the deep convolution module, the channel attention network and the channel gate feedforward network can be integrated into a multi-type feature integration module, through which the integrated extraction of multi-type features can be achieved.

[0116] Figure 3 is a block diagram of a multi-type feature integration module according to an exemplary embodiment. Figure 3 As shown in the figure, the multi-type feature integration module includes: normalization module, self-attention mechanism module, deep convolution module, channel attention network and channel gate feedforward network.

[0117] The input of the normalization module is the input of the entire multi-type feature integration module. The output of the normalization module is input into the self-attention mechanism module, the deep convolution module, and the channel attention network. Through separate processing, three features can be obtained. By integrating these three features, the integrated features are obtained. Next, the integrated features are processed using the channel gate feedforward network. The features output by the channel gate feedforward network are then integrated with the integrated features to obtain the target image features output by the multi-type feature integration module.

[0118] Input features For example, some features involved in the multi-type feature integration module can be expressed as: X LN =LN(X) X INTER =MSA(X LN )+CA(X LN )+DWconv(X LN ) Y=CGFN(LN(X INTER ))+X INTER in, and They represent intermediate features, and Y represents the output of the multi-type feature integration module. LN represents layer normalization, CGFN represents channel-gated feedforward network processing, MSA represents multi-scale self-attention mechanism processing, Dwconv represents deep convolution processing, and CA represents channel attention mechanism processing.

[0119] In some embodiments, the multi-type feature integration module may be part of a residual multi-type feature integration group, and the residual multi-type feature integration group may be part of a deep feature extraction module.

[0120] Figure 4 is a block diagram of a deep feature extraction module according to an exemplary embodiment. Figure 4 ,The deep feature extraction module includes multiple residual multi-type feature ,integration groups and a 3×3 convolutional layer to further optimize the ,intermediate features and achieve deep feature extraction.

[0121] Each residual multi-type feature integration group includes multiple multi-type feature integration modules and a 3×3 convolutional layer.

[0122] For any residual multi-type feature integration group, some of the features involved can be expressed as: F i,0 =F i-1 F i,j =MFIB i,j (F i,j-1 ), j=1,2,…N2 F i =Conv i (F i ,N2)+F i-1 in, represents the input features of the i-th residual multi-type feature integration group, and represents the output features of the jth multi-type feature integration module in the i-th residual multi-type feature integration group module, and N2 represents the number of multi-type feature integration modules. A 3×3 convolutional layer is added at the end of the residual multi-type feature integration group. Finally, residual connections can be used to better fuse intermediate features, making the entire network more stable during application, for example, making the network training process more stable.

[0123] In some embodiments, the first image feature is processed through a channel gate feedforward network to obtain a second image feature. Obtaining the second image feature may include: dividing the first image feature according to multiple channel dimensions to obtain image features corresponding to the multiple channel dimensions; processing the image features corresponding to the multiple channel dimensions through a channel gate feedforward network to obtain the second image feature.

[0124] In this embodiment, the first image features can be divided into multiple channel dimensions first, so that the channel gate feedforward network can process them separately based on the divided features, thereby improving computational efficiency and improving feature extraction effects in different channel dimensions.

[0125] In some embodiments, the channel-gate feed-forward network may include a first network branch and a second network branch.

[0126] Correspondingly, the image features corresponding to multiple channel dimensions are processed through the channel gate feedforward network to obtain the second image feature. The second image feature includes: performing local feature extraction on the image features corresponding to the first channel dimension through the first network branch to obtain the local image feature; performing global feature extraction on the image features corresponding to the second channel dimension through the second network branch to obtain the global image feature; and determining the second image feature based on the local image feature and the global image feature.

[0127] In this implementation, different network branches are used to extract local features and global features respectively. Thus, the input feature map is divided into two branches along the channel dimension, with each branch occupying half of the feature channel, reducing channel redundancy while making it easier to capture nonlinear feature information.

[0128] In some embodiments, the channel gate feedforward network may further include a layer normalization module and a feature partitioning module, wherein the input of the layer normalization module is the input of the entire network, the output of the layer normalization module is the input of the feature partitioning module, and the output of the feature partitioning module is respectively input into the two network branches.

[0129] In some embodiments, the channel gate feed-forward network may further include a 1×1 convolutional layer for further processing the features output by the two network branches.

[0130] In some embodiments, the channel gate feedforward network may further include a residual learning module for performing residual connections on the input of the entire network and the features processed by the 1×1 convolutional layer to obtain the features of the final output of the network.

[0131] Figure 5 is a block diagram of a channel gate feedforward network according to an exemplary embodiment, such as Figure 5 As shown in the figure, the channel gate feedforward network includes: a layer normalization module, a feature partitioning module, a first network branch, a second network branch, a 1×1 convolution, and a residual connection module.

[0132] like Figure 5 As shown, the first network branch may include, in sequence: a global average pooling layer, a 1×1 convolution, a 3×3 convolution, a GELU (Gaussian Error Linear Units, an activation function based on Gaussian error) activation function, a 1×1 convolution, a 3×3 convolution, and a Sigmoid (S-type growth curve) activation function.

[0133] like Figure 5 As shown, the second network branch may sequentially include: 1×1 convolution and 3×3 convolution.

[0134] like Figure 5 As shown, the first network branch and the second network branch are respectively connected to the feature division module, and the feature division module is connected to the layer normalization module.

[0135] like Figure 5 As shown in Figure 1, the residual learning module performs residual learning on the output of the 1×1 convolution module and the input of the entire network to obtain the final features.

[0136] Combine Figure 5 The network structure shown in the figure, the processing process of the channel gate feedforward network can include:

[0137] Among them, LN represents layer normalization processing, Represents dividing the input features into two parts along the channel dimension, Represents global average pooling. represents convolution, represents the depthwise convolution, where Represents the size of the kernel and stride, for example: represents a convolution with a kernel size of 1 and a stride size of 1, represents a depthwise convolution with a kernel size of 3 and a stride size of 3. represents the GELU activation function, Represents the Sigmoid activation function, and ⊙ represents the element-by-element multiplication operation.

[0138] as well as, and Represent the features of the two network branches input, and Represent the features of the two network branches output, The integrated features representing the features output by the two network branches.

[0139] Furthermore, by introducing residual learning, the final output feature of the channel gate feedforward network is It can be expressed as:

[0140] Through this channel-gated feedforward network, not only can the utilization of global information in different channel dimensions be improved, but also the redundancy between channels can be reduced.

[0141] Furthermore, according to the target image features, various processing such as image reconstruction and image enhancement can be performed on the image to be processed.

[0142] In the disclosed embodiment, the channel gate feedforward network may also be replaced by other forms of networks, such as an MLP (Multilayer Perceptron) network.

[0143] It is understandable that in different application scenarios, different image processing methods can be used according to different processing requirements.

[0144] As an example, in a video playback scenario, the image processing method of the embodiment of the present disclosure can be used for post-processing at the decoding end, where the decoding end can be a terminal that can play videos, such as a mobile phone, tablet, or TV.

[0145] As an example, after shooting a video, the terminal encodes and stores the video. When the user plays the encoded and stored video through the album, the terminal decodes the video to obtain video frame images. Then, these video frame images are processed using the image processing method of the embodiment of the present disclosure to obtain enhanced video frame images. Then, the video is played based on the enhanced video frame images. The image processing method of the embodiment of the present disclosure can improve the resolution and quality of the image, and solve the limitations in detail restoration, edge enhancement, and noise removal.

[0146] As an example, consider a terminal equipped with a video playback application. When playing videos in this application, there are low-bitrate videos that are limited by bandwidth transmission. Directly decoding and playing such videos will reduce image quality, resulting in mosaics, smearing, and other issues. By performing image enhancement processing on decoded video frames using the image processing method of the disclosed embodiments, the image quality impact caused by encoding can be effectively reduced, thereby improving the user's video viewing experience.

[0147] Therefore, in the case where the image to be processed is an image obtained by decoding the video to be played, step S13 may include: performing image enhancement processing on the image to be processed at least according to the multi-scale image features to obtain the target image.

[0148] Alternatively, according to the target image features, image enhancement processing is performed on the image to be processed to obtain the target image.

[0149] In addition to the aforementioned application scenarios, the technical solutions of the disclosed embodiments can also be applied to restore 265-degree compressed images. Specifically, the pixel shuffling operation during image reconstruction can be removed, ensuring that the restored image resolution remains consistent with the original image, allowing the network to restore 265-degree compressed images.

[0150] In the embodiments disclosed herein, the various processing logics described above can be integrated into Figure 5 In the Transformer network architecture shown.

[0151] Figure 6 FIG. 1 is a schematic diagram of a Transformer network architecture according to an exemplary embodiment. Figure 6 As shown in the figure, the network architecture includes a shallow feature extraction module, a deep feature extraction module and an image reconstruction module.

[0152] The shallow feature extraction module can be implemented through convolutional layers. The image reconstruction module can be implemented using PixelShuffle, a deep learning-based upsampling method that can be referenced by mature technologies in this field.

[0153] And, regarding the deep feature extraction module, the image processing logic provided by the embodiment of the present disclosure can be adopted. Specifically, the self-attention mechanism processing logic and / or the channel gate feedforward network introduced in the above embodiment can be adopted. Figures 2 to 5 The relevant network architecture is shown.

[0154] In order to verify the technical effects that can be achieved by the image processing method of the embodiment of the present disclosure, the application effect of the deep learning network using the image processing method of the embodiment of the present disclosure is tested. For example, the two indicators of PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity) can be tested.

[0155] PSNR is an objective metric used to measure the difference between two images. It can be used to evaluate the effectiveness of image compression, transmission, or reconstruction algorithms. A higher PSNR value indicates greater similarity between the two images, resulting in less quality loss. SSIM is a metric used to measure the similarity between two images in terms of brightness, contrast, and structure. Unlike PSNR, SSIM is closer to human visual perception and can more accurately reflect image quality.

[0156] In a test scenario, the overall deep learning network was tested at scales of ×2, ×3, and ×4. Compared to other image processing networks, the PSNR and SSIM gains at scales of ×2 and ×4 were 0.33dB and 0.3dB, respectively. These tests demonstrate the effectiveness of the technical solutions of the disclosed embodiments, which can effectively extract texture information from images at different scales.

[0157] In a test scenario, the application effect of the channel gate feedforward network was tested. The technical solution of the embodiment of the present disclosure can achieve a gain of 0.22 dB in PSNR and SSIM compared with other technical solutions, proving that the processing effect based on the channel gate feedforward network is better than the processing effect of other networks that do not use the channel gate feedforward network.

[0158] In a test scenario, the application effect of the self-attention mechanism was tested. Compared with the solution without the self-attention mechanism of the embodiment of the present disclosure, the PSNR and SSIM can obtain a gain of 0.24 dB, proving that the processing effect based on the self-attention mechanism is better than other processing effects without the self-attention mechanism.

[0159] In a test scenario, the complexity of the test model is compared with the complexity of other models, which proves that the model of the embodiment of the present disclosure not only has fewer parameters and more addition counts, but also has better performance.

[0160] In a test scenario, the deep learning network was tested for its effectiveness in restoring 265 compressed images. The network demonstrated stability across all metrics, achieving excellent image restoration capabilities in low- to medium-quality compression scenarios. Furthermore, it demonstrated excellent results in restoring image detail and structure. Overall, the network demonstrated excellent performance in compressed image restoration, demonstrating both stability and restoration capabilities. In low- to medium-quality compression scenarios, it was able to achieve a good balance between peak signal-to-noise ratio, structural similarity, and multi-scale structure preservation.

[0161] By leveraging the Transformer's powerful global dependency modeling capabilities, the technical solutions of the disclosed embodiments enable more efficient and accurate image reconstruction and enhancement, providing reliable technical support for detail preservation and visual optimization. For low-resolution images and those affected by noise, image quality can be significantly improved, meeting the demands for detailed image restoration and enhanced visual effects in real-world scenarios.

[0162] Figure 7 FIG. 7 is a block diagram of an image processing apparatus 700 according to an exemplary embodiment. Figure 7 , the device comprises: The acquisition module 701 is configured to acquire original image features of the image to be processed.

[0163] The processing module 702 is configured to perform a self-attention mechanism on the original image features at multiple image scales to obtain multi-scale image features, wherein the multiple image scales are different from the scales of the image to be processed.

[0164] The processing module 702 is further configured to process the image to be processed at least according to the multi-scale image features to obtain a target image.

[0165] Optionally, the processing module 702 is further configured to: perform linear mapping on the original image features at multiple image scales to obtain linear mapping vectors corresponding to the multiple image scales; aggregate the linear mapping vectors corresponding to the multiple image scales to obtain initial multi-scale image features; and perform cross-attention enhancement processing on the initial multi-scale image features to obtain target multi-scale image features.

[0166] Optionally, the processing module 702 is further configured to: determine a first self-attention map based on the query vector and the first key vector; determine a second self-attention map based on the query vector and the second key vector; determine a multi-scale self-attention map based on the first self-attention map and the second self-attention map; determine a multi-scale eigenvalue based on the first value vector and the second value vector; and determine an initial multi-scale image feature based on the multi-scale self-attention map and the multi-scale eigenvalue.

[0167] Optionally, the processing module 702 is further configured to: determine a third self-attention map based on the second self-attention map; determine a new multi-scale image feature based on the third self-attention map and the second value vector; and determine a target multi-scale image feature based on the initial multi-scale image feature and the new multi-scale image feature.

[0168] Optionally, the processing module 702 is further configured to: determine a first image feature based at least on the multi-scale image feature; process the first image feature through a channel gate feedforward network to obtain a second image feature; determine a target image feature based on the first image feature and the second image feature; and process the image to be processed based on the target image feature to obtain a target image.

[0169] Optionally, the processing module 702 is further configured to: perform convolution processing on the original image feature to obtain a third image feature; perform channel attention mechanism processing on the original image feature to obtain a fourth image feature; and determine the first image feature based on the multi-scale image feature, the third image feature and the fourth image feature.

[0170] Optionally, the processing module 702 is further configured to: divide the first image features according to multiple channel dimensions to obtain image features corresponding to the multiple channel dimensions; and process the image features corresponding to the multiple channel dimensions through a channel gate feedforward network to obtain the second image features.

[0171] Optionally, the processing module 702 is further configured to: perform local feature extraction on the image features corresponding to the first channel dimension through the first network branch to obtain local image features; perform global feature extraction on the image features corresponding to the second channel dimension through the second network branch to obtain global image features; and determine the second image features based on the local image features and the global image features.

[0172] Optionally, the processing module 702 is further configured to: perform image enhancement processing on the image to be processed at least according to the multi-scale image features to obtain the target image.

[0173] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0174] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, which implement the steps of the image processing method provided by the present disclosure when the program instructions are executed by a processor.

[0175] Figure 8 8 is a block diagram of an electronic device 800 according to an exemplary embodiment. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0176] Reference Figure 8 The electronic device 800 may include one or more of the following components: a processing component 802 , a first memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output interface 812 , a sensor component 814 , and a communication component 816 .

[0177] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 802 may include one or more first processors 820 to execute instructions to perform all or part of the steps of the above-described image processing method. In addition, the processing component 802 may include one or more modules to facilitate interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate interaction between the multimedia component 808 and the processing component 802.

[0178] The first memory 804 is configured to store various types of data to support operations on the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The first memory 804 can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0179] The power supply component 806 provides power to the various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 800.

[0180] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, it may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensors can not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide action. In some embodiments, the multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the electronic device 800 is in an operating mode, such as a capture mode or a video mode, the front-facing camera and / or the rear-facing camera can receive external multimedia data. Each front-facing camera and the rear-facing camera can have a fixed optical lens system or have focal length and optical zoom capabilities.

[0181] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, or a voice recognition mode. The received audio signals may be further stored in the first memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0182] The input / output interface 812 provides an interface between the processing component 802 and peripheral interface modules, such as a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.

[0183] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect changes in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and temperature changes of the electronic device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0184] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0185] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-mentioned image processing method.

[0186] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a first memory 804 including instructions. The instructions can be executed by the first processor 820 of the electronic device 800 to perform the above-described image processing method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0187] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program executable by a programmable device, and has a code portion for performing the above-mentioned image processing method when executed by the programmable device.

[0188] Those skilled in the art will also understand that the various illustrative logical blocks and steps listed in the embodiments of this application can be implemented through electronic hardware, computer software, or a combination of both. Whether such functions are implemented through hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art may use various methods to implement the described functions for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of this application.

[0189] It should be understood that, unless otherwise specifically noted, the features of the various embodiments of the present disclosure described herein may be combined with each other. As used herein, the term "and / or" includes any one of the relevant listed items and any combination of any two or more thereof; similarly, "at least one of" includes any one of the relevant listed items and any combination of any two or more thereof.

[0190] It should be understood that, unless otherwise expressly specified or limited, the terms "join," "attach," "install," "connect," "connect," "fix," etc. used in the embodiments of the present disclosure should be understood in a broad sense. For example, they can be fixedly connected, detachably connected, or integrated; they can be mechanically connected, electrically connected, or communicable with each other; they can be directly connected, or indirectly connected through an intermediate medium, and they can be internally connected between two elements or an interactive relationship between two elements, unless otherwise expressly limited. For those skilled in the art, the specific meanings of the above terms in this article can be understood according to specific circumstances.

[0191] Although terms such as "first", "second" and "third" may be used herein to describe various components, parts, regions, layers or sections, these components, parts, regions, layers or sections are not limited to these terms. On the contrary, these terms are only used to distinguish one component, part, region, layer or section from another component, part, region, layer or section. Therefore, without departing from the teachings of each example, the first component, part, region, layer or section mentioned in the examples described herein may also be referred to as the second component, part, region, layer or section. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" can explicitly or implicitly include at least one such feature. In the description herein, the meaning of "multiple" is at least two, for example, two, three, etc., unless otherwise clearly and specifically defined.

[0192] Furthermore, the word "exemplary" is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as advantageous over other aspects or designs. Rather, the use of the word exemplary is intended to present concepts in a concrete manner. As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from the context, "X applies to A or B" is intended to mean any of the natural inclusive permutations. That is, if X applies to A; X applies to B; or X applies to both A and B, then "X applies to A or B" satisfies any of the aforementioned instances. Furthermore, the articles "a" and "an," as used in this application and the appended claims, are generally understood to mean "one or more," unless otherwise specified or clear from the context to refer to the singular form.

[0193] Likewise, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding this specification and the accompanying drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the claims. With particular regard to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, terms used to describe such components are intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if not structurally equivalent to the disclosed structure. In addition, although particular features of the present disclosure may have been disclosed with respect to only one of several implementations, such features may be combined with one or more other features of other implementations as may be desired and advantageous for any given or particular application. Furthermore, to the extent that the terms "include," "have," "have," "have," or variations thereof are used in the detailed description or claims, such terms are intended to be inclusive in a manner similar to the term "comprising."

[0194] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.

[0195] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. An image processing method, characterized in that: include: Obtaining original image features of the image to be processed; Performing a self-attention mechanism on the original image features at multiple image scales to obtain multi-scale image features, wherein the multiple image scales are different from the scale of the image to be processed; The image to be processed is processed at least according to the multi-scale image features to obtain a target image.

2. The image processing method according to claim 1, wherein: The self-attention mechanism is performed on the original image features at multiple image scales to obtain multi-scale image features, including: Performing linear mapping on the original image features at multiple image scales to obtain linear mapping vectors corresponding to the multiple image scales; Aggregating the linear mapping vectors corresponding to the multiple image scales to obtain an initial multi-scale image feature; Cross-attention enhancement processing is performed on the initial multi-scale image features to obtain target multi-scale image features.

3. The image processing method according to claim 2, wherein: The linear mapping vectors corresponding to the multiple image scales respectively include: a query vector, a first key vector, and a first value vector corresponding to a first image scale, and a second key vector and a second value vector corresponding to a second image scale. The linear mapping vectors corresponding to the multiple image scales are aggregated to obtain initial multi-scale image features, including: Determining a first self-attention map based on the query vector and the first key vector; Determining a second self-attention map based on the query vector and the second key vector; Determining a multi-scale self-attention map according to the first self-attention map and the second self-attention map; determining a multi-scale eigenvalue according to the first value vector and the second value vector; An initial multi-scale image feature is determined according to the multi-scale self-attention map and the multi-scale feature value.

4. The image processing method according to claim 3, wherein: The cross-attention enhancement process is performed on the initial multi-scale image features to obtain the target multi-scale image features, including: Determining a third self-attention map according to the second self-attention map; Determining a new multi-scale image feature according to the third self-attention map and the second value vector; A target multi-scale image feature is determined according to the initial multi-scale image feature and the new multi-scale image feature.

5. The image processing method according to claim 1, wherein: The step of processing the image to be processed at least according to the multi-scale image features to obtain a target image includes: determining a first image feature based at least on the multi-scale image feature; Processing the first image feature through a channel gate feedforward network to obtain a second image feature; determining a target image feature according to the first image feature and the second image feature; The image to be processed is processed according to the target image features to obtain a target image.

6. The image processing method according to claim 5, characterized in that The determining of the first image feature at least based on the multi-scale image feature includes: Performing convolution processing on the original image feature to obtain a third image feature; Performing a channel attention mechanism on the original image feature to obtain a fourth image feature; The first image feature is determined according to the multi-scale image feature, the third image feature, and the fourth image feature.

7. The image processing method according to claim 5 or 6, characterized in that: The first image feature is processed by a channel gate feedforward network to obtain a second image feature, and the second image feature is obtained, including: Dividing the first image feature according to multiple channel dimensions to obtain image features corresponding to the multiple channel dimensions respectively; The image features corresponding to the multiple channel dimensions are processed through a channel gate feedforward network to obtain the second image features.

8. The image processing method according to claim 7, wherein: The channel gate feedforward network includes a first network branch and a second network branch. The channel gate feedforward network processes the image features corresponding to the multiple channel dimensions to obtain the second image features. Obtaining the second image features includes: Performing local feature extraction on the image features corresponding to the first channel dimension through the first network branch to obtain local image features; Performing global feature extraction on the image features corresponding to the second channel dimension through the second network branch to obtain global image features; The second image feature is determined according to the local image feature and the global image feature.

9. The image processing method according to claim 1, wherein: The image to be processed is an image obtained by decoding the video to be played, and the processing of the image to be processed at least according to the multi-scale image features to obtain a target image includes: Image enhancement processing is performed on the image to be processed at least according to the multi-scale image features to obtain the target image.

10. An image processing device, characterized in that: Used to execute the image processing method according to any one of claims 1 to 9.

11. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the executable instructions to implement the image processing method according to any one of claims 1 to 9.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image processing method according to any one of claims 1 to 9 is implemented.

13. A computer program product, characterized in that The invention comprises a computer program, which implements the image processing method according to any one of claims 1 to 9 when executed by a processor.