Data processing method and computing device
By dividing the feature map into feature blocks with window size M and using the mask matrix to obtain sub-feature blocks with window size P, combined with the self-attention mechanism, the problem of low computing resource occupation and efficiency of AI models when processing large-scale data is solved, and higher performance and accuracy are achieved.
Patent Information
- Application Number
- PCT/CN2025/071241
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-17
- Filing Date
- 2025-01-08
- Publication Date
- 2025-08-21
AI Technical Summary
When AI models process large-scale data, the data volume of a single process is large, occupy a large amount of computing resources and have low computing efficiency, and cannot obtain long-distance data information, resulting in inaccurate output results.
The feature map is divided into feature blocks with window size M, and a mask matrix with window size M is used to obtain sub-feature blocks with window size P from it, and the calculation is performed through the self-attention mechanism, combining local and global information, and improving the receptive field to mine the relationship between longer-distance data.
Without increasing the data processing volume, the performance and computing efficiency of the AI model are improved to ensure the accuracy of data processing.
Smart Images

Figure CN2025071241_21082025_PF_FP_ABST
Abstract
Description
Data processing method and computing device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on February 17, 2024, with application number 202410180021.X and application name “A Data Processing Method and Computing Device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a data processing method and a computing device. Background Art
[0003] In the AI field, AI models are used to process large amounts of data. Large amounts of data processed at a time consume a significant amount of computing resources and result in low computational efficiency. Therefore, to avoid consuming excessive computing resources and improve data processing efficiency, AI models typically divide the data to be processed into multiple windows and then process the data within each window separately. However, AI models process the data within each window separately, focusing only on the local data within the window and failing to obtain information about long-range data. This can reduce the performance of the AI model and lead to inaccurate output results. Summary of the Invention
[0004] This application provides a data processing method and computing device that can improve the performance of AI models without occupying too many computing resources.
[0005] In a first aspect, the present application provides a data processing method. The method can be applied to a computing device and can include: obtaining a first feature block with a window size of M from a feature map of a first channel of a target image, wherein the first feature block includes N sub-feature blocks each with a window size of P, where M is greater than P; obtaining a second sub-feature block with a window size of P from the first feature block based on a first mask matrix with a window size of M; and determining a super-resolved image corresponding to the target image based on the second sub-feature block, wherein the resolution of the super-resolved image is greater than the resolution of the target image.
[0006] In the above scheme, after dividing the feature map into multiple sub-feature blocks based on a window size P, the sub-feature blocks are expanded into feature blocks with a window size of M. Image super-resolution is performed based on the feature blocks with a window size of M. This increases the receptive field of data processing, allowing relationships between more distant data to be mined from relatively large window data, thereby improving the performance of image super-resolution. Furthermore, the feature blocks are sampled using a mask matrix with a window size of M to obtain a second sub-feature block with a window size of P. The super-resolved image is determined based on this second sub-feature block, which can avoid the increase in data processing required by expanding the window.
[0007] In a possible implementation of the first aspect, based on a first mask matrix with a window size of M, obtaining a second sub-feature block with a window size of P from the first feature block may specifically include multiplying the first mask matrix and data at the same position in the first feature block to obtain new data, where the second sub-feature block is a set of the new data. The first mask matrix includes N sub-matrices corresponding to the N sub-feature blocks, each sub-matrix including one or more invalid values, and the second sub-feature block includes eigenvalues corresponding to valid values in the first mask matrix. Valid values may be set to 1, and invalid values may be set to 0.
[0008] In the above scheme, the effective values in the first mask matrix are dispersed, so that when the first mask matrix samples the feature block once (i.e., the data multiplication described above), the data of any sub-feature block is not sampled in its entirety. In other words, the second sub-feature block includes at least one data point from at least two of the N sub-feature blocks.
[0009] In a possible implementation of the first aspect, the N sub-feature blocks include a first sub-feature block and N-1 reference sub-feature blocks corresponding to the first sub-feature block, the first sub-feature block is the same as the first target sub-feature block, the first target sub-feature block is obtained based on the union of the second sub-feature block and N-1 second target sub-feature blocks with a window size of P, the N-1 second target sub-feature blocks are obtained based on the first mask matrix and N-1 target feature blocks with a window size of M in the feature map of the first channel, and the N-1 target feature blocks all include the first sub-feature block.
[0010] In the above scheme, taking the feature map window expansion S times as an example, M = S × P, N = S × S, each sub-feature block exists in N feature blocks. The first mask matrix samples the N feature blocks that contain the first sub-feature block. The data obtained by taking the union of the N sampled sub-feature blocks is the same as the data in the first sub-feature block, which ensures that no data is lost during the feature map sampling process of the first mask matrix.
[0011] In a possible implementation of the first aspect, determining the super-resolved image corresponding to the target image based on the second sub-feature block includes: processing the second sub-feature block based on a self-attention mechanism to obtain a third sub-feature block with a window size of P, and processing a fourth sub-feature block with a window size of P obtained from a feature map of a second channel of the target image based on a self-attention mechanism to obtain a fifth sub-feature block with a window size of P; and generating the super-resolved image based on the third sub-feature block and the fifth sub-feature block.
[0012] In a possible implementation of the first aspect, the method further includes: obtaining the fourth sub-feature block from a second feature block with a window size of M in a feature map of the second channel based on a second mask matrix with a window size of M, wherein the second feature block includes N sub-feature blocks each having a window size of P, and the N sub-feature blocks in the second feature block include a fourth sub-feature block and N-1 reference sub-feature blocks corresponding to the fourth sub-feature block. The position of the valid values in the second mask matrix in the mask matrix is the same as the position of the fourth sub-feature block in the second feature block.
[0013] In the above scheme, two mask matrices are used to sample the feature maps of different channels, which can effectively combine the local information and global information of the sub-feature blocks, thereby improving the performance of data processing.
[0014] In a possible implementation of the first aspect, the method is applied to a super-resolution model, and the method further includes: determining a delay of the super-resolution model; and updating the first mask matrix when the delay is less than or equal to a delay threshold.
[0015] In the above solution, the first mask matrix is determined by the delay of the model, which can ensure that the model delay is not less than a preset threshold, thereby improving the performance of the model.
[0016] In a possible implementation of the first aspect, the method further includes: performing feature extraction on an initial feature map of the first channel using a first operator to obtain a first feature of the first channel, wherein a frequency of the first feature is greater than a frequency threshold; determining a feature map of the first channel based on the first feature of the first channel; and / or performing feature extraction on an initial feature map of the second channel using the second operator to obtain a first feature of the second channel; and determining a feature map of the second channel based on the first feature of the second channel.
[0017] In a possible implementation of the first aspect, the first operator includes a Sobel operator or a Laplace operator.
[0018] In the above solution, high-frequency features are extracted from the feature map based on the high-frequency feature extraction operator, so that the first feature map can contain more details and improve the data processing effect.
[0019] In a second aspect, the present application provides a data processing device, which includes a processing module, a mask module, and a super-resolution module.
[0020] The processing module is used to obtain a first feature block with a window size of M from a feature map of a first channel of a target image, wherein the first feature block includes N sub-feature blocks with a window size of P, where M is greater than P.
[0021] The super-resolution module is used to obtain a second sub-feature block with a window size of P from the first feature block based on a first mask matrix with a window size of M.
[0022] The super-resolution module is used to determine a super-resolution image corresponding to the target image based on the second sub-feature block, and the resolution of the super-resolution image is greater than the resolution of the target image.
[0023] In a possible implementation of the second aspect, the N sub-feature blocks include a first sub-feature block, the first sub-feature block is the same as the first target sub-feature block, the first target sub-feature block is obtained based on the union of the second sub-feature block and N-1 second target sub-feature blocks with a window size of P, the N-1 second target sub-feature blocks are obtained based on the first mask matrix and N-1 target feature blocks with a window size of M in the feature map of the first channel, and the N-1 target feature blocks all include the first sub-feature block.
[0024] In a possible implementation of the second aspect, the super-resolution module is specifically used to: process the second sub-feature block based on a self-attention mechanism to obtain a third sub-feature block with a window size of P, and process the fourth sub-feature block with a window size of P obtained from the feature map of the second channel of the target image based on the self-attention mechanism to obtain a fifth sub-feature block with a window size of P; and generate the super-resolution image based on the third sub-feature block and the fifth sub-feature block.
[0025] In a possible implementation of the second aspect, the mask module is further used to: obtain the fourth sub-feature block from a second feature block with a window size of M in the feature map of the second channel based on a second mask matrix with a window size of M, the second feature block including N sub-feature blocks with a window size of P, and the N sub-feature blocks in the second feature block including the fourth sub-feature block.
[0026] In a possible implementation of the second aspect, the processing module is further used to: perform feature extraction on the initial feature map of the first channel using a first operator to obtain a first feature of the first channel, wherein the frequency of the first feature is greater than a frequency threshold; determine the feature map of the first channel based on the first feature of the first channel; and / or perform feature extraction on the initial feature map of the second channel using the second operator to obtain a first feature of the second channel; and determine the feature map of the second channel based on the first feature of the second channel.
[0027] In a possible implementation of the second aspect, the first operator includes a Sobel operator or a Laplace operator.
[0028] In a possible implementation of the second aspect, the apparatus is applied to a super-resolution model, and the mask module is further configured to: determine a delay of the super-resolution model; and update the first mask matrix when the delay is less than or equal to a delay threshold.
[0029] In a third aspect, the present application further provides a computing device comprising: a processor and a memory, wherein the processor is configured to execute a computer program stored in the memory to implement the data processing method described in the first aspect or any optional embodiment of the first aspect.
[0030] In a fourth aspect, the present application also provides a computer-readable storage medium comprising instructions, which, when executed on a computer, enable the computer to execute the data processing method as described in implementing the first aspect or any optional embodiment of the first aspect.
[0031] In a fifth aspect, the present application provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the data processing method provided by the first aspect or any possible implementation of the first aspect.
[0032] Any of the above-mentioned devices, computing equipment, computer storage media, or computer program products is used to execute the method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding schemes in the corresponding methods provided above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] FIG1 is a schematic diagram of the structure of a super-resolution model provided in an embodiment of the present application;
[0034] FIG2 is a schematic diagram of the structure of a processing module in the super-resolution model shown in FIG1 according to an embodiment of the present application;
[0035] FIG3 is a schematic diagram of the structure of the mask self-attention module in the super-resolution model shown in FIG1 provided in an embodiment of the present application;
[0036] FIG4 is a flow chart of a data processing method applied to the processing module shown in FIG2 according to an embodiment of the present application;
[0037] FIG5 is a flowchart of a data processing method applied to the masked self-attention module shown in FIG3 , provided by an embodiment of the present application;
[0038] FIG6 is a schematic diagram of the masked self-attention module shown in FIG3 according to an embodiment of the present application dividing the feature map into windows to obtain sub-feature blocks;
[0039] FIG7 is a schematic diagram of the masked self-attention module shown in FIG3 according to an embodiment of the present application, which fills the sub-feature blocks before expanding the feature map window;
[0040] 8 and 9 are schematic diagrams of the masked self-attention module shown in FIG3 according to an embodiment of the present application, which expands the window of the feature map to obtain a feature block;
[0041] FIG10 is a schematic diagram of the masked self-attention module shown in FIG3 using a sparse mask matrix to sample feature blocks according to an embodiment of the present application;
[0042] FIG11 is a schematic diagram of a sparse mask matrix used by the masked self-attention module shown in FIG3 provided in an embodiment of the present application;
[0043] FIG12 is a schematic diagram of the non-volatility principle satisfied by the sparse mask matrix used by the masked self-attention module shown in FIG3 provided in an embodiment of the present application;
[0044] FIG13 is a schematic diagram of the masked self-attention module shown in FIG3 using a dense mask matrix to sample feature blocks according to an embodiment of the present application;
[0045] FIG14 is a flowchart of another data processing method of the masked self-attention module shown in FIG3 provided in an embodiment of the present application;
[0046] FIG15a is a flowchart of a method for determining a sparse mask matrix provided in an embodiment of the present application;
[0047] FIG15 b is a schematic diagram of three sparse mask matrices and their sampling loss rates provided in an embodiment of the present application;
[0048] FIG16a is a flowchart of a method for training the super-resolution model shown in FIG1 provided in an embodiment of the present application;
[0049] FIG16b is a comparison diagram of the effects of the model obtained by the method shown in FIG16a on image super-resolution according to an embodiment of the present application;
[0050] FIG17 is a schematic structural diagram of a data processing device for implementing the method shown in FIG4 according to an embodiment of the present application;
[0051] FIG18 is a schematic structural diagram of a data processing device for implementing the method shown in FIG5 , provided in an embodiment of the present application;
[0052] FIG19 is a schematic structural diagram of a data processing device for implementing the method shown in FIG14 provided in an embodiment of the present application;
[0053] Figure 20 is a structural diagram of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0054] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.
[0055] In the description of the embodiments of the present application, words such as "exemplary," "for example," or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary," "for example," or "for example" in the embodiments of the present application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary," "for example," or "for example" is intended to present the relevant concepts in a concrete manner.
[0056] In the description of the embodiments of the present application, the term "and / or" is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, B exists alone, and A and B exist at the same time. In addition, unless otherwise specified, the term "plurality" means two or more. For example, "multiple systems" refers to two or more systems, and "multiple screen terminals" refers to two or more screen terminals.
[0057] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly identifying the technical features being referred to. Thus, features specified as "first" or "second" may explicitly or implicitly include one or more of such features. The terms "include," "comprising," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.
[0058] Before introducing the embodiments of the present application, the nouns appearing in the embodiments of the present application are first introduced below.
[0059] A feature map is the result of performing a feature representation on the data to be processed. The data to be processed may include, for example, image data, audio data, or text data. Image data and audio data may include data from multiple channels. Therefore, feature representation can be performed on each of the multiple channels to generate feature maps for each channel.
[0060] A feature block is a dataset within a feature map that includes multiple feature values. In the AI field, feature maps can be divided into multiple feature blocks for ease of processing. Similarly, a feature block can be divided into sub-feature blocks, which are subsets of the data within the feature block.
[0061] Window is a tool for preprocessing data. When there is a lot of data to be processed, the window can be used to divide the data to be processed into multiple feature blocks, and then each feature block can be processed separately. This can avoid taking up more computing resources and improve computing efficiency. The length and width of the window can be the same or different. This application is introduced as an example of the same length and width. For example, for the feature map C×H×W (C represents channel, H represents height, and W represents width) of image data, the feature map of each channel can be divided into windows according to window P to obtain multiple feature blocks. Each feature block can be expressed as C×P×P, where the number of feature blocks corresponding to the feature map of each channel is HW / P 2 .
[0062] A mask matrix is used to filter out a portion of the data to be processed, or to extract other data from it. The mask matrix includes valid values and invalid values. The number of valid and invalid values and their positions in the mask matrix are usually set according to the actual application scenario. In the field of AI, invalid values can be set to 0, and valid values can be set to non-zero values (for example, 1).
[0063] High-frequency features refer to feature data that changes dramatically in the feature graph to be processed. In the field of AI, high-frequency features can be extracted from the feature graph through a target operator. The target operator can include the Sobel operator or the Laplace operator. In addition, whether the feature data is a high-frequency feature can be determined by the value of the target parameter of the feature data, where the target parameter of the feature data can include gradient, entropy, frequency, first-order derivative, second-order derivative, etc.
[0064] The transformer structure is a deep learning network based on the self-attention mechanism. The transformer can use the self-attention mechanism to process the input data, thereby mining the relationship between the data in the input data.
[0065] Super-resolution (SA) computing is a data processing method in the field of AI. Specifically, it reconstructs relatively low-resolution (LR) data into relatively high-resolution (high-resolution) data. SA computing has a wide range of applications. For example, SA computing can be performed on LR image data to obtain high-resolution image data, or SA computing can be performed on LR audio data to obtain high-resolution audio data.
[0066] In the field of AI, SA is often performed using a super-resolution model based on deep learning. The super-resolution model can be obtained by training an AI model based on LR data and the HR data corresponding to the LR data. The AI model can use, but is not limited to, a convolutional neural network (CNN) structure or a transformer structure.
[0067] In related art 1, the super-resolution model can use a convolutional neural network (CNN) structure, processing input data through the CNN convolution layer and extracting local features from the input data in parallel using multiple convolution kernels in the convolution layer. This solution can improve data processing efficiency by extracting local features from the input data in parallel using the convolution kernels in the convolution layer. However, this solution is limited by the size of the convolution kernel and cannot explore the relationship between data in the input data, resulting in low performance of the super-resolution model.
[0068] In related technology 2, the super-resolution model can include a convolutional layer and a transformer structure. The convolution kernel of the convolutional layer extracts local features from the input feature map. The self-attention module in the transformer then divides the feature map into sub-feature blocks with a window size of P. Self-attention calculations are performed on each sub-feature block to mine the relationships between the data in each sub-feature block. However, this solution can only mine data relationships from local data and cannot mine long-range data information, which has limited performance improvements for the model.
[0069] To this end, an embodiment of the present application provides a data processing method that can solve the above problems.
[0070] In the data processing method provided in the embodiment of the present application, under the premise of dividing the feature map into multiple sub-feature blocks based on the window size P, each window sub-feature block of the feature map is multiplied by S, and the expanded window size is M (M=S×P). Then, a mask matrix with a window size of M is used to sample sub-feature blocks with a window size of P from each feature block with a window size of M in the feature map, and self-attention calculation is performed on the sampled sub-feature blocks with a window size of P. Among them, the sub-feature block A with a window size of P sampled from the feature block with a window size of M is different from the sub-feature block B before the window expansion. The sub-feature block A contains the global information in the feature block with a window size of M. Therefore, compared with the above-mentioned related technology 2, the embodiment of the present application can enable the self-attention calculation to mine the relationship between data at a relatively longer distance through the two steps of window expansion and mask sampling.
[0071] The data processing method provided in the embodiments of the present application can be applied to a super-resolution model to improve the performance of the super-resolution model.
[0072] Figure 1 is a schematic diagram of the structure of a super-resolution model provided in an embodiment of the present application. As shown in Figure 1, the super-resolution model may include one or more convolution units 210, one or more feature extraction units 220, and one or more reconstruction units 230. Taking image data as an example, the super-resolution model shown in Figure 1 can super-resolve LR image data to obtain HR images.
[0073] The convolution unit 210 can be used to perform convolution calculations on the initial feature maps of multiple channels corresponding to the LR image data to extract local features in the initial feature maps, and then fuse the extracted local features into feature maps 3 of multiple channels. Specifically, the convolution unit 210 may include one or more convolution layers. The convolution layer can perform convolution calculations on the input data of the convolution layer using a convolution kernel of size 3x3 and a stride of 1 to extract features from the input data. Of course, the convolution layer may include multiple convolution kernels, and each convolution kernel may use the same or different sizes and / or strides.
[0074] The feature extraction unit 220 may include one or more processing modules 221 and one or more mask self-attention modules 222 shown in FIG. 1 .
[0075] The processing module 221 is used to perform high-frequency reparameterization calculations on the input data of the processing module 221 to extract high-frequency features from the input data. By extracting high-frequency features, the model's ability to reconstruct details in image data can be improved. Specifically, the processing module 221 may include multiple processing branches and a fusion layer 2219. Each processing branch extracts features from the input data from different dimensions. As shown in Figure 2, the processing module 221 may include five branches (labeled as branches 1 to 5 from left to right). The fusion layer 2219 adds the output data of branches 1 to 5. Among them, the specific processing process of the processing module 221 shown in Figure 2 will be specifically introduced in conjunction with Figure 4 later.
[0076] The masked self-attention module 222 is used to perform self-attention calculation on the feature data in each window of the input data of the masked self-attention module 222, thereby obtaining feature maps 2 of multiple channels. Specifically, as shown in FIG3, the masked self-attention module 222 may include a partitioning layer 2221, an expansion layer 2222, a sparse mask layer 2223, a dense mask layer 2224, multiple self-attention layers 2225, a connection layer 2226, and a convolution layer 2227.
[0077] The partitioning layer 2221 is used to divide the input feature map into windows, and divide the feature map into multiple sub-feature blocks with a window size of P.
[0078] The expansion layer 2222 is used to perform window expansion S times on the multiple sub-feature blocks to obtain multiple feature blocks with a window size of M. The number of feature blocks is the same as the number of sub-feature blocks, where M=S×P.
[0079] The sparse mask layer 2223 is used to sample a portion of the feature blocks with a window size of M using a sparse mask matrix Ms with a window size of M, and obtain a second sub-feature block with a window size of P from each feature block. The sparse mask matrix Ms includes N sub-matrices corresponding to the N sub-feature blocks in the feature block, and each sub-matrix includes one or more invalid values and one or more valid values.
[0080] The dense mask layer 2224 is used to use a dense mask matrix Md with a window size of M to sample another part of the feature blocks with a window size of M, and obtain a fourth sub-feature block with a window size of P from each feature block.
[0081] The self-attention layer 2225 is used to perform self-attention calculation on each second sub-feature block to obtain a third sub-feature block with a window size of P; or to perform self-attention calculation on each fourth sub-feature block to obtain a fifth sub-feature block with a window size of P.
[0082] The connection layer 2226 is used to connect the output results of multiple self-attention layers 2225.
[0083] The convolution layer 2227 is used to perform convolution calculation on the output results of the connection layer 2226 again, extract local features and fuse them, so as to obtain feature maps 2 of multiple channels.
[0084] Among them, the specific processing process of the masked self-attention module 222 shown in Figure 3 will be specifically introduced in conjunction with Figure 5 later.
[0085] The reconstruction unit 230 is used to calculate the input data of the reconstruction unit 230 to obtain super-resolution data. The reconstruction unit 230 may include one or more convolutional layers and one or more permutation layers. The permutation layer may include a shuffle operation for randomly rearranging the elements in the input feature map to improve the generalization ability of the model. Low-resolution image data can be converted into high-resolution image data through convolution operations and shuffle operations. In the case of super-resolution of image data, the number of convolutional layers and permutation layers can be determined according to the super-resolution multiple. For example, when the super-resolution multiple is 2, the number of convolutional layers and permutation layers can be 1; when the super-resolution multiple is 3, the number of convolutional layers and permutation layers can be 2; when the super-resolution multiple is 4, the number of convolutional layers and permutation layers can be 3.
[0086] The specific processing process of the processing module 221 shown in FIG. 2 will be described below in conjunction with FIG. 4 .
[0087] Figure 4 is a flow chart of a data processing method provided by an embodiment of the present application. The method can be applied to a computing device that is equipped with the processing module 221 of the super-resolution model shown in Figure 1. As shown in Figure 4, the method may include steps S401-S402.
[0088] At step S401, the processing module 221 performs multiple feature extraction processes on the feature maps 3 of multiple channels of the LR image to obtain multiple output data. The feature maps 3 of multiple channels of the LR image may be the output data of the convolution unit 201. Specifically, as shown in FIG2 , the processing module 221 may include five branches (labeled as branches 1 to 5 from left to right). In FIG2 , the input data of the processing module 221 may include the feature maps 3 of multiple channels output by the convolution unit 210.
[0089] Branch 1 may include a convolution layer 2211, which performs convolution processing on the input data of convolution layer 2211. Convolution layer 2211 may include one convolution kernel or multiple identical or different convolution kernels. For example, convolution layer 2211 may be configured with multiple convolution kernels of size 1×1 and stride 1, which convolve different data in the input data of convolution layer 2211 in parallel, thereby improving the efficiency of convolution processing. In Figure 2, the input data of convolution layer 2211 is the input data of processing module 221.
[0090] Branch 2 may include a convolution layer 2212 and a convolution layer 2216. Convolution layer 2212 and convolution layer 2216 may include one convolution kernel or multiple identical or different convolution kernels. For example, convolution layer 2212 may be configured with multiple convolution kernels of size 1×1 and stride 1, and perform convolution processing on different data in the input data of convolution layer 2212 in parallel. Similarly, convolution layer 2216 may include multiple convolution kernels of size 3×3 and stride 1, and perform convolution processing on different data in the input data of convolution layer 2216 in parallel. In Figure 2, the input data of convolution layer 2212 is the input data of processing module 221, and the input data of convolution layer 2216 is the output data of convolution layer 2212.
[0091] Branch 3 may include a convolution layer 2213 and an extraction layer 2217. Convolution layer 2213 may include one convolution kernel or multiple identical or different convolution kernels. For example, convolution layer 2213 may include multiple convolution kernels of size 1×1 and stride 1, which perform convolution processing on different data in the input data of convolution layer 2213 in parallel. Extraction layer 2217 may include a high-frequency feature extraction operator for extracting high-frequency features from the input data of extraction layer 2217. In Figure 2, the input data of convolution layer 2213 is the input data of processing module 221, and the input data of extraction layer 2217 is the output data of convolution layer 2213.
[0092] Branch 4 may include a convolution layer 2214 and an average pooling layer 2218. The convolution layer 2214 may include one convolution kernel or multiple identical or different convolution kernels. For example, the convolution layer 2214 may include multiple convolution kernels of size 1×1 and step size 1, and perform convolution processing on different data in the input data of the convolution layer 2214 in parallel. The average pooling layer 2218 may divide the input data of the average pooling layer 2218 into several regions and calculate the average value of all elements in each region. The average pooling layer 2218 can effectively reduce the dimension of the data while retaining local features. In Figure 2, the input data of the convolution layer 2214 is the input data of the processing module 221, and the input data of the average pooling layer 2218 is the output data of the convolution layer 2214.
[0093] Branch 5 may include a convolutional layer 2215. Convolutional layer 2215 may include one convolution kernel or multiple identical or different convolution kernels. Convolutional layer 2215 may include multiple convolution kernels of size 3×3 and stride 1, and concurrently perform convolution processing on different data in the input data of convolutional layer 2215. In Figure 2, the input data of convolutional layer 2215 is the input data of processing module 221.
[0094] S402: The processing module 221 fuses the output data of each branch to obtain the feature maps 1 of multiple channels of the LR image. Specifically, as shown in FIG2 , the fusion layer 2219 adds the output data of branches 1 to 5 to obtain the feature maps 1 of multiple channels.
[0095] In the above method, multiple convolution operations are used in processing module 221, which requires fewer memory access calculations, shortening data processing time and reducing data processing latency. In addition, the use of high-frequency feature extraction operators to extract high-frequency features can enable feature map 1 to retain more details, improving data processing performance.
[0096] The specific processing process of the masked self-attention module 222 shown in FIG3 is introduced below in conjunction with FIG5 .
[0097] FIG5 is a flow chart of a data processing method provided in an embodiment of the present application. This method can be applied to a computing device that is equipped with the masked self-attention module 222 of the super-resolution model shown in FIG1. The masked self-attention module 222 can adopt the structure of the embodiment shown in FIG3. As shown in FIG5, the method may include S501-S506.
[0098] S501, the division layer 2221 divides the feature map 1 of multiple channels of the input LR image into multiple sub-feature blocks with a window size of P, and the expansion layer 2222 expands the multiple sub-feature blocks by S times to obtain multiple feature blocks with a window size of M.
[0099] Each feature block consists of N sub-feature blocks, each with a window size of P, where M = S × P and N = S × S. Taking feature block 101 as an example, the N sub-feature blocks include sub-feature block 11 and its N-1 reference sub-feature blocks, which will be described in detail later. The N-1 reference sub-feature blocks are the sub-feature blocks adjacent to sub-feature block 11 in feature map 1.
[0100] Among them, the feature maps 1 of multiple channels of the LR image can be obtained by processing the feature maps 3 of multiple channels of the LR image by the processing module 221. For details, please refer to the above introduction to Figure 4, which will not be repeated here.
[0101] Specifically, the partitioning layer 2221 performs window partitioning on the feature maps 1 of multiple channels of the LR image based on the window size P, and divides the feature maps 1 of each channel of the LR image into multiple sub-feature blocks with a window size of P. Taking the feature map 1 of a channel shown in Figure 6 as an example, the height H and width W of the feature map 1 are both 32. As shown in Figure 6, when P is 4, H / P = 8, W / 4 = 8, and the feature map 1 can be divided into 64 sub-feature blocks with a window size of 4.
[0102] Specifically, the expansion layer 2222 expands multiple sub-feature blocks with a window size of P by S times to obtain multiple feature blocks with a window size of M.
[0103] Since the number of reference sub-feature blocks corresponding to the sub-feature blocks on both sides of the end point of the expansion direction in the feature map does not meet the requirement of S times expansion, that is, the number of reference sub-feature blocks corresponding to the sub-feature blocks in the Z rows and Z columns at the end point of the expansion direction is less than N-1, the expansion layer 2222 needs to use the sub-feature blocks in the Z rows and Z columns selected from the starting point of the expansion direction in the feature map 1 to fill the gaps at the end point of the expansion direction, and also fill the sub-feature blocks at the starting point of the expansion direction at the end point of the expansion direction. The starting point of the expansion direction can include the sub-feature blocks in the upper left corner, upper right corner, lower left corner, or lower right corner of the feature map 1. Wherein, Z = S-1.
[0104] Taking the expansion of the 64 sub-feature blocks of the feature map 1 shown in Figure 6 by 2 times as an example, as shown in Figure 7, the expansion layer 2222 can select sub-feature block 11 as the starting point of the expansion direction for expansion, and the expansion direction is along the lower right corner, using the sub-feature blocks in the 1st row and 1st column and sub-feature block 11 to fill in the last row and the last column.
[0105] After completing the sub-feature blocks, expansion layer 2222 uses a sliding window of size M to expand feature image 1. After the sub-feature blocks in feature image 1 are expanded S times, the number of sub-feature blocks is the same as the number of feature blocks, with each sub-feature block corresponding to one feature block. Taking the 64 sub-feature blocks shown in Figure 6 as an example, when S = 2 and M = 8, expansion layer 2222 can use the sub-feature block as the starting point and expand it by 2 times toward the lower right corner of feature image 1, resulting in 64 feature blocks.
[0106] Each feature block includes N sub-feature blocks, each of which includes the expanded sub-feature block and its N-1 reference sub-feature blocks. Taking sub-feature block 11 as an example, after the window is expanded by a factor of 2, the feature block corresponding to sub-feature block 11 is shown in Figure 8 . The feature block corresponding to sub-feature block 11 includes sub-feature block 11 and its three reference sub-feature blocks. These three reference sub-feature blocks include sub-feature block 12, sub-feature block 21, and sub-feature block 22. Taking sub-feature block 12 as an example, after the window is expanded by a factor of 2, the feature block corresponding to sub-feature block 12 is shown in Figure 8 . The feature block corresponding to sub-feature block 12 includes sub-feature block 12 and its three reference sub-feature blocks. These three reference sub-feature blocks include sub-feature block 13, sub-feature block 22, and sub-feature block 23. Taking sub-feature block 18 as an example, after the window is expanded by a factor of 2, the feature blocks corresponding to sub-feature block 18 are shown in Figure 9 . The feature blocks corresponding to sub-feature block 18 include sub-feature block 18 and its three reference sub-feature blocks. These three reference sub-feature blocks include sub-feature block 28, sub-feature block 11, and sub-feature block 21. Taking sub-feature block 81 as an example, after the window is expanded by a factor of 2, the feature blocks corresponding to sub-feature block 81 are shown in Figure 9 . The feature blocks corresponding to sub-feature block 81 include sub-feature block 81 and its three reference sub-feature blocks. These three reference sub-feature blocks include sub-feature block 82, sub-feature block 11, and sub-feature block 12.
[0107] After obtaining multiple feature blocks corresponding to each channel, S502 is executed on the feature map 1 of a part of the channels, and sparse sampling is performed to obtain sub-feature blocks. For example, as shown in Figure 3, the sparse mask layer 2223 obtains the feature block 101 from the feature map of channel C1, and sparse sampling is performed on the feature block 101. S504 is executed on the feature map 1 of another part of the channels, and dense sampling is performed to obtain sub-feature blocks. For example, as shown in Figure 3, the dense mask layer 2224 obtains the feature block 201 from the feature map of channel C2, and sparse sampling is performed on the feature block 201. Among them, the multiple channels of the LR image can be divided into two parts, that is, S502 is executed on the feature map 1 of the C / 2 channel (including the channel C1 below), and S504 is executed on the feature map 1 of the other C / 2 channel (including the channel C2 below).
[0108] S502, the sparse mask layer 2223 samples each feature block in the feature map 1 of a part of the channels based on the sparse mask matrix Ms with a window size of M, and obtains multiple sub-feature blocks 2 with a window size of P.
[0109] Specifically, the sparse mask layer 2223 can use the sparse mask matrix Ms to sample each feature block in the feature map 1 of channel C1 with a step size P, and obtain a sub-feature block with a window size of P from each feature block. As shown in Figure 10, the sparse mask layer 2223 can use the sparse mask matrix Ms to sample feature block 101 to obtain a sub-feature block 2 with a window size of P. Then, it moves P steps to the right to sample feature block 102, and moves P steps to sample feature block 103, and so on, until all feature blocks in the feature map 1 of channel C1 are sampled.
[0110] The reason why the window size of the sub-feature blocks sampled from the sparse mask matrix Ms is P is that the number of valid values in the sparse mask matrix Ms is the same as the number in the sub-feature blocks, both P×P. Taking feature block 101 as an example, as shown in Figure 8, feature block 101 includes sub-feature block 11 and its reference sub-feature block. The number of valid values in the sparse mask matrix Ms is the same as the number in sub-feature block 11 and its reference sub-feature block. Valid values in the sparse mask matrix Ms can be 1, and invalid values can be 0. When P = 4 and S = 2, the number of valid values in the sparse mask matrix Ms is 16, and the number of invalid values is 48.
[0111] The sub-feature blocks obtained by sampling the feature block using the sparse mask matrix Ms are different from the sub-feature blocks in the feature block. The reason is that the effective values in the sparse mask matrix Ms are dispersed, so the sparse mask matrix Ms will not sample all the data of the sub-feature blocks in the feature block. Taking the sparse mask matrix Ms shown in Figure 11 as an example, the black squares represent the positions of the effective values 1. The effective values 1 in the sparse mask matrix Ms are dispersed, so when the sparse mask matrix Ms samples the feature block 101 shown in Figure 11, it will not sample all the data in the sub-feature blocks 11, sub-feature blocks 12, sub-feature blocks 21, or sub-feature blocks 22.
[0112] The sparse mask matrix Ms satisfies non-volatility. The determination process of the sparse mask matrix Ms will be introduced in conjunction with Figure 15a later and will not be repeated here. Among them, non-volatility means that the data in the feature map will be sampled by the sparse mask matrix Ms and will not be lost. The reason is that after the sub-feature blocks in the feature block are used to fill, when the sparse mask matrix Ms is used to sample all the feature blocks in the feature map, each sub-feature block in the feature map will be sampled N times. For a sub-feature block in the feature map, the data after the union of the N sub-feature blocks 2 obtained by sampling N times is the same as the sub-feature block. As shown in Figure 12, when S=2, when the sparse mask matrix Ms with a window size of 8 is sampled, the four feature blocks containing the sub-feature block 11 will be sampled, that is, the sparse mask matrix Ms will sample the sub-feature block 11 four times. The sub-feature block 2 obtained by four samplings is the same as the sub-feature block 11 after the union is calculated, that is, the data in the sub-feature block after the union is the same as the data in the sub-feature block 11.
[0113] S503, the self-attention layer 2225 performs self-attention processing on the sub-feature block 2 output by the sparse mask layer 2223 to obtain a sub-feature block 3 with a window size of P. The self-attention layer 2225 can be based on the preset query weight matrix W Q , key weight matrix W K , and value weight matrix W V Calculate with sub-feature block 2 respectively to obtain query matrix Q, key matrix K and value matrix V, then multiply the dot product of query matrix Q and key matrix K with value matrix V to obtain sub-feature block 3 and output it.
[0114] S504, the dense mask layer 2224 samples each feature block in the feature map of another part of the channels based on the dense mask matrix Md with a window size of M, and obtains multiple sub-feature blocks 4 with a window size of P.
[0115] Specifically, the dense mask layer 2224 can use the dense mask matrix Md to sample each feature block in the feature map 1 of the channel C2 with a step size P, and obtain a sub-feature block with a window size of P from each feature block. The sampling step can be the sampling step of the sparse mask layer 2223 shown in Figure 10 until the sampling is completed.
[0116] The reason why the window size of the sub-feature block obtained by sampling the dense mask matrix Md is P is that the number of valid values in the dense mask matrix Md is the same as the number in the sub-feature block, both of which are P×P.
[0117] The sub-feature block obtained by sampling the feature block with the dense mask matrix Md is the same as a sub-feature block in the feature block. The reason is that the position of the valid value in the dense mask matrix Md is the same as the position of a sub-feature block in the feature block. Taking the dense mask matrix Md shown in Figure 13 as an example, the valid values in the dense mask matrix Md are concentrated in the upper left corner. When the dense mask matrix Md samples the feature block 201 shown in Figure 13, all the data in the sub-feature block 11 in the feature block 201 can be sampled, that is, the sub-feature block 11 is the sub-feature block 4. Among them, the feature block 201 and the feature block 101 belong to different feature maps, so the sub-feature block 11 in the feature block 201 and the sub-feature block 11 in the feature block 101 belong to different feature maps. Figure 13 uses the same mark only for the convenience of representation.
[0118] S505, the self-attention layer 2225 performs self-attention processing on the sub-feature block 4 output by the dense mask layer 2224 to obtain a sub-feature block 5 with a window size of P. The self-attention layer 2225 can be based on the preset query weight matrix W Q , key weight matrix W K , and value weight matrix W V Calculate with sub-feature block 4 respectively to obtain query matrix Q, key matrix K and value matrix V, then multiply the dot product of query matrix Q and key matrix K with value matrix V to obtain sub-feature block 5 and output it.
[0119] S506, the connection layer 2226 fuses the sub-feature blocks 3 of a part of the channels output from the attention layer 2225 and the sub-feature blocks 5 of another part of the channels, and then the convolution layer 2227 performs convolution calculation on the output of the connection layer 2226 to obtain the feature maps 2 of multiple channels and outputs the feature maps 2 of the multiple channels.
[0120] Figure 14 is a flow chart of a data processing method provided by an embodiment of the present application. This method can be applied to a computing device that is equipped with the masked self-attention module 222 of the super-resolution model shown in Figure 1. The masked self-attention module 222 can adopt the structure of Figure 3 without the dense mask layer 2224, that is, the sparse mask layer 2223 samples the feature maps of all channels. As shown in Figure 14, the method may include S1401-S1404.
[0121] S1401, the feature map 1 of multiple channels of the input LR image is divided into multiple sub-feature blocks with a window size of P, and the expansion layer 2222 expands the multiple sub-feature blocks by S times to obtain multiple feature blocks with a window size of M.
[0122] S1402: Based on the sparse mask matrix Ms with a window size of M, obtain sub-feature blocks 2 with a window size of P from the feature blocks in the feature maps of each channel. The sparse mask layer 2223 of the masked self-attention module 222 can sample each feature block in the feature maps of each channel to obtain multiple sub-feature blocks 2 with a window size of P.
[0123] S1403: Generate an SR image corresponding to the LR image based on sub-feature block 2. Specifically, S1403 may include the following steps S14031-S14033.
[0124] In S14031, the self-attention layer 2225 performs self-attention processing on the sub-feature block 2 output by the sparse mask layer 2223 to obtain a sub-feature block 3 with a window size of P. The self-attention layer 2225 can be based on the preset query weight matrix W Q , key weight matrix W K , and value weight matrix W V Calculate with sub-feature block 2 respectively to obtain query matrix Q, key matrix K and value matrix V, then multiply the dot product of query matrix Q and key matrix K with value matrix V to obtain sub-feature block 3 and output it.
[0125] Among them, S14032, the connection layer 2226 fuses the sub-feature blocks 3 of each channel output from the attention layer 2225, and the convolution layer 2227 performs convolution calculation on the output of the connection layer 2226 to obtain the feature maps 2 of multiple channels and outputs the feature maps 2 of the multiple channels.
[0126] In S14033, the reconstruction unit 230 generates an SR image corresponding to the LR image based on the feature map 2 of multiple channels.
[0127] The specific process of each step in the above embodiment can refer to the introduction of the embodiments shown in Figures 1 and 5 above, and will not be repeated here.
[0128] Based on the method embodiments shown in FIG. 5 and FIG. 14 , an embodiment of the present application further provides a method for determining a sparse mask matrix.
[0129] Figure 15a is a flow chart of a method for determining a sparse mask matrix provided in an embodiment of the present application. This method can be applied to the computing device described above and used to determine the sparse mask matrix Ms used in Figures 5 and 14 based on the principle of non-volatility. As shown in Figure 15a, the method may include steps S1501-S1503.
[0130] S1501: The computing device randomly generates a sparse mask matrix Ms0 with a window size of M, wherein the number of valid values in the sparse mask matrix Ms0 is 16 and the number of invalid values is 48.
[0131] S1502: The computing device determines N target feature blocks in the target feature map, each having a window size of P. The N target feature blocks all include the first target sub-feature block in the target feature map.
[0132] The N target feature blocks may include the N feature blocks in the feature map of any of the above channels. Taking the feature map shown in FIG12 as an example, if sub-feature block 11 is used as the first target sub-feature block, the N target feature blocks may be the feature block in the upper left corner (including sub-feature block 11, sub-feature block 12, sub-feature block 21, and sub-feature block 22), the feature block in the upper right corner (including sub-feature block 18, sub-feature block 11, sub-feature block 28, and sub-feature block 21), the feature block in the lower left corner (including sub-feature block 81, sub-feature block 82, sub-feature block 11, and sub-feature block 12), and the feature block in the lower right corner (including sub-feature block 88, sub-feature block 81, sub-feature block 18, and sub-feature block 11).
[0133] S1503: The computing device samples the N target feature blocks based on the sparse mask matrix Ms0 to obtain N second target sub-feature blocks with a window size of P.
[0134] S1504: The computing device calculates the union of the N second target sub-feature blocks to obtain a third target sub-feature block with a window size of P.
[0135] S1505. The computing device determines whether the sparse mask matrix Ms0 satisfies the above-mentioned non-volatility principle based on the first target sub-feature block and the third target sub-feature block. If so, the sparse mask matrix Ms0 is used as the sparse mask matrix Ms. Otherwise, S1501 is re-executed.
[0136] The computing device can calculate the sampling loss rate β between the first target sub-feature and the third target sub-feature according to the following formula (1). If the sampling loss rate β is 0, the data at the same position of the first target sub-feature F1 and the third target sub-feature F3 are the same, and the sparse mask matrix Ms0 is non-volatile. For example, the feature block 101 shown in FIG11 is sampled using three sparse mask matrices MS0 (including MS0-1, MS0-2, and MS0-3) as shown in FIG15b. If the sampling loss rates β of the three samples are 0, 25%, and 75% as shown in FIG15b, respectively, the sparse mask matrix MS0-1 with a sampling loss rate β of 0 satisfies the non-volatile principle.
[0137] In formula (1), ‖F1∩F3‖ represents the bimodal value obtained by calculating the intersection of the first target sub-feature F1 and the third target sub-feature F3, and ‖F1‖ represents the bimodal value obtained by calculating the intersection of the first target sub-feature F1 and the third target sub-feature F3.
[0138] Among them, the computing device can repeatedly execute S1501-S1505 to obtain multiple sparse mask matrices Ms0 that meet the non-volatile principle as candidate sparse mask matrices Ms. Then, each candidate sparse mask matrix Ms is applied to the method shown in Figure 1 or Figure 14, and the LR image data is super-resolved using the super-resolution model based on the method shown in Figure 1 or Figure 14, and the delay of the super-resolution model when performing super-resolution calculations based on each candidate sparse mask matrix Ms is determined. Finally, based on the delay corresponding to each candidate sparse mask matrix Ms, the candidate sparse mask matrix Ms that is less than or equal to the delay threshold is selected as the final result.
[0139] Based on the super-resolution model shown in FIG1 , an embodiment of the present application also provides a method for training a super-resolution model.
[0140] FIG16a is a flow chart of a method for training a super-resolution model provided in an embodiment of the present application. As shown in FIG16a , the method may include:
[0141] S1601: A computing device determines a training data set, where the training data set may include one or more LR images and HR images corresponding to the LR images.
[0142] At step S1602, the computing device uses each LR image as input data for the super-resolution model shown in FIG1 , and obtains an SR image corresponding to each LR image output by the super-resolution model. The super-resolution model's processing of the input LR images can be found in the description of FIG1-FIG5 above and is not further described here.
[0143] S1603, the computing device calculates the super-resolution loss value corresponding to each LR image based on the SR image corresponding to each LR image and the HR image corresponding to each LR image using a predetermined loss function, and updates the parameters in the super-resolution model based on the super-resolution loss value corresponding to each LR image until the super-resolution loss value corresponding to each LR image meets the loss threshold.
[0144] Three super-resolution models (labeled M1, M2, and M3, with increasing parameter counts) obtained by training based on the method shown in Figure 16a were trained and tested on the same datasets (Dataset 1 and Dataset 2) with the super-resolution models trained based on Related Techniques 1 and 2. The performance comparisons are shown in Table 1. The performance shown in Table 1 includes latency (ms), parameters (M), floating-point operations per second (FLOPs), and signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM), where PSNR and SSIM are expressed in decibels (dB).
[0145] Table 1 Performance comparison of super-resolution models
[0146] As can be seen from Table 1, the values of d (the two data in the boxes below dataset 1 and dataset 2) of the three models obtained by the embodiment method of the present application at two super-resolution scales (super-resolution 2 times and super-resolution 4 times) are all due to related technology 1 and related technology 2.
[0147] Image 1 in dataset 2 was super-resolved by a factor of 2 using M1 obtained using related technique 1 and an embodiment of the present invention. Image 2 in dataset 2 was super-resolved by a factor of 2 using M3 obtained using related technique 2 and an embodiment of the present invention. The resulting super-resolved (SR) images are shown in Figure 16b. Compared to the HR images corresponding to the original images, the super-resolved results of the embodiments of the present invention can contain more texture.
[0148] In addition, the processing modules 221 in the three models provided in the embodiments of the present application all include a high-frequency extraction branch (branch 3), which can extract high-frequency information from the feature map. When the model is used for image super-resolution, it is beneficial to the reconstruction of images with dense textures and edge details. The high-frequency extraction branch of model M3 is ablated, and the PSNR comparison results shown in Table 2 are obtained based on data set 1 and data 2. As can be seen from Table 2, in the implementation including branch 3, the PSNR when processing data set 1 and data set 2 are improved by 0.05dB and 0.04dB, respectively.
[0149] Table 2 PSNR comparison with and without adding high frequency extraction branch
[0150] Based on the above-mentioned data processing method embodiment, the embodiment of the present application also provides a data processing device.
[0151] FIG17 is a schematic diagram of the structure of a data processing device 1700 provided in an embodiment of the present application. The data processing device 1700 can be used to implement the data processing method shown in FIG4 . As shown in FIG17 , the data processing device 1700 can include a processing module 1701 and a fusion module 1702 .
[0152] The processing module 1701 is used to perform multiple feature extraction processes on the feature maps 3 of multiple channels of the input LR image to obtain multiple output data.
[0153] The fusion module 1702 is used to fuse multiple output data to obtain feature maps 1 of multiple channels of the LR image.
[0154] FIG18 is a schematic diagram of the structure of a data processing device 1800 provided in an embodiment of the present application. The data processing device 1800 can be used to implement the data processing method shown in FIG5 . As shown in FIG18 , the data processing device 1800 may include a processing module 1801 , a masking module 1802 , and a fusion module 1803 .
[0155] Among them, the processing module 1801 is used to divide the feature map 1 of multiple channels of the input LR image into multiple sub-feature blocks with a window size of P, and expand the multiple sub-feature blocks by S times to obtain multiple feature blocks with a window size of M.
[0156] Among them, the mask module 1802 is used to sample each feature block in the feature map of a part of the channels (for example, C / 2) based on the sparse mask matrix Ms with a window size of M, to obtain multiple sub-feature blocks 2 with a window size of P, and to sample each feature block in the feature map of another part of the channels (for example, C / 2) based on the dense mask matrix Md with a window size of M, to obtain multiple sub-feature blocks 4 with a window size of P.
[0157] Among them, the fusion module 1803 is used to fuse the sub-feature blocks 3 of one part of the channels and the sub-feature blocks 5 of another part of the channels, and then perform convolution calculation on the fused data to obtain feature maps 2 of multiple channels.
[0158] FIG19 is a schematic diagram of the structure of a data processing device 1900 provided in an embodiment of the present application. The data processing device 1900 can be used to implement the data processing method shown in FIG14 . As shown in FIG19 , the data processing device 1900 may include a processing module 1901, a masking module 1902, and a super-resolution module 1903.
[0159] Among them, the processing module 1901 is used to divide the feature map 1 of multiple channels of the input LR image into multiple sub-feature blocks with a window size of P, and expand the multiple sub-feature blocks by S times to obtain multiple feature blocks with a window size of M.
[0160] Among them, the mask module 1902 is used to sample each feature block in the feature map of each channel based on the sparse mask matrix Ms with a window size of M, and obtain multiple sub-feature blocks 2 with a window size of P.
[0161] Among them, the super-resolution module 1903 is used to generate the HR image corresponding to the LR image based on the sub-feature block 2.
[0162] It should be noted that the data processing device 1700 provided by the embodiment shown in FIG17, the data processing device 1800 provided by the embodiment shown in FIG18, and the data processing device 1800 provided by the embodiment shown in FIG19, when executing the data processing method, only uses the division of the above-mentioned functional modules as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure can be divided into different functional modules to complete all or part of the functions described above. In addition, the computing device provided by the above-mentioned embodiment and the data processing method embodiment shown in FIG4 or FIG5 or FIG14 belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0163] FIG20 is a schematic diagram of the hardware structure of a computing device 2000 provided in an embodiment of the present application.
[0164] The computing device 2000 can be configured as the computing device configured with the super-resolution model described in the above embodiments. Referring to FIG20 , the computing device 2000 includes a processor 2001, a memory 2002, a communication interface 2003, and a bus 2004. The processor 2001, the memory 2002, and the communication interface 2003 are interconnected via the bus 2004. The processor 2001, the memory 2002, and the communication interface 2003 can also be connected using other connection methods besides the bus 2004.
[0165] The processor 2001 may be a general-purpose processor, which may be a processor that performs specific steps and / or operations by reading and executing the contents stored in a memory (e.g., memory 2002). For example, the general-purpose processor may be a central processing unit (CPU). The processor 2001 may include at least one circuit to perform all or part of the steps of the method provided in the embodiment shown in Figure 4, Figure 5, or Figure 14. The processor 2001 may include one or more cores. When multiple cores are included, the multiple cores may be used to execute all or part of the steps of the method provided in the embodiment shown in Figure 4, Figure 5, or Figure 14, respectively.
[0166] The memory 2002 may be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, optical storage, a hard disk, etc. When the computing device 2000 is deployed with a super-resolution model, the memory 2002 may store the program corresponding to the data processing device 1700 and / or the program corresponding to the data processing device 1800.
[0167] Communication interface 2003 includes input / output (I / O) interfaces, physical interfaces, and logical interfaces, which are used to interconnect components within computing device 2000, as well as interfaces for interconnecting computing device 2000 with other devices (e.g., other computing devices or user devices). Physical interfaces can be Ethernet interfaces, fiber optic interfaces, ATM interfaces, etc.
[0168] The bus 2004 may be any type of communication bus for interconnecting the processor 2001 , the memory 2002 , and the communication interface 2003 , such as a system bus.
[0169] The above-mentioned devices can be provided on separate chips, or at least partially or entirely on the same chip. Whether to provide each device independently on different chips or to integrate them on one or more chips often depends on the product design requirements. The embodiments of this application do not limit the specific implementation of the above-mentioned devices.
[0170] The computing device 2000 shown in FIG20 is merely exemplary. During implementation, the computing device 2000 may further include other components, which are not listed one by one herein.
[0171] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state drive (SSD)).
[0172] It is understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not intended to limit the scope of the embodiments of the present application. It should be understood that in the embodiments of the present application, the order of the sequence numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0173] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of this application. It should be understood that the above description is only the specific implementation methods of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of this application should be included in the scope of protection of this application.
Claims
1. A data processing method, characterized in that: The method comprises: Obtain a first feature block with a window size of M from a feature map of a first channel of a target image, wherein the first feature block includes N sub-feature blocks with window sizes of P, where M is greater than P; Based on a first mask matrix with a window size of M, obtaining a second sub-feature block with a window size of P from the first feature block; A super-resolved image corresponding to the target image is determined based on the second sub-feature block, where a resolution of the super-resolved image is greater than a resolution of the target image.
2. The method according to claim 1, characterized in that The N sub-feature blocks include a first sub-feature block, the first sub-feature block is the same as the first target sub-feature block, the first target sub-feature block is obtained based on the union of the second sub-feature block and N-1 second target sub-feature blocks with a window size of P, the N-1 second target sub-feature blocks are obtained based on the first mask matrix and the N-1 target feature blocks with a window size of M in the feature map of the first channel, and the N-1 target feature blocks all include the first sub-feature block.
3. The method according to claim 1 or 2, characterized in that Determining the super-resolved image corresponding to the target image based on the second sub-feature block includes: Processing the second sub-feature block based on the self-attention mechanism to obtain a third sub-feature block with a window size of P, and processing the fourth sub-feature block with a window size of P obtained from the feature map of the second channel of the target image based on the self-attention mechanism to obtain a fifth sub-feature block with a window size of P; The super-resolved image is generated based on the third sub-feature block and the fifth sub-feature block.
4. The method according to claim 3, characterized in that The method further comprises: Based on a second mask matrix with a window size of M, the fourth sub-feature block is obtained from a second feature block with a window size of M in the feature map of the second channel, the second feature block includes N sub-feature blocks with a window size of P, and the N sub-feature blocks in the second feature block include the fourth sub-feature block.
5. The method according to claim 4, characterized in that The method further comprises: Performing feature extraction on the initial feature map of the first channel using a first operator to obtain a first feature of the first channel, where a frequency of the first feature is greater than a frequency threshold; Determining a feature map of the first channel based on the first feature of the first channel; and / or Performing feature extraction on the initial feature map of the second channel using the second operator to obtain a first feature of the second channel; A feature map of the second channel is determined based on the first feature of the second channel.
6. The method according to any one of claims 1 to 5, characterized in that The first operator includes a Sobel operator or a Laplace operator.
7. The method according to any one of claims 1 to 6, characterized in that The method is applied to a super-resolution model, and the method further comprises: Determining a time delay of the super-resolution model; In a case where the delay is less than or equal to a delay threshold, the first mask matrix is updated.
8. A data processing device, characterized in that: The device comprises: a processing module, configured to obtain a first feature block with a window size of M from a feature map of a first channel of a target image, wherein the first feature block includes N sub-feature blocks with window sizes of P, where M is greater than P; a mask module, configured to obtain a second sub-feature block with a window size of P from the first feature block based on a first mask matrix with a window size of M; A super-resolution module is used to determine a super-resolution image corresponding to the target image based on the second sub-feature block, where the resolution of the super-resolution image is greater than the resolution of the target image.
9. A computing device, characterized in that The computing device comprises: a processor and a memory, wherein the processor is configured to execute a computer program stored in the memory to perform the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The method comprises instructions, and when the instructions are executed on a computer, the computer executes the method according to any one of claims 1 to 7.
11. A computer program product comprising instructions, characterized in that When the instructions are executed on a computer, the computer executes the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data processing method and computing device
CN120510038A
Image information fusion and super-resolution reconstruction method based on feature processing
CN112598575A
Super-resolution amplification method, system and device for image with any scale
CN114092337A
Image reconstruction method, electronic device and computer-readable storage medium
US20220351333A1