Computing systems, methods, devices, and storage media for processing images
Patent Information
- Application Number
- CN202210851223.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-19
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-07-19
AI Technical Summary
在一些图像处理任务中,可能要进行像素级别的数据处理,从而需要较大的计算量
[0007]应当理解,本发明内容部分中所描述的内容并非旨在限定本公开的实施例的关键特征或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的描述而变得容易理解。
Smart Images

Figure CN117475165B_ABST
Abstract
Description
Technical Field
[0001] The exemplary embodiments disclosed herein relate generally to the field of computers, and more particularly to computing systems, methods, apparatuses, and computer-readable storage media for processing images. Background Technology
[0002] In the field of computer vision (CV), various image processing techniques based on artificial intelligence have made significant progress and have wide applications. Computer vision can be applied to a variety of different image processing tasks, such as image classification, image segmentation, and image generation. Some image processing tasks may require pixel-level data processing, thus demanding a large amount of computation. Summary of the Invention
[0003] In a first aspect of this disclosure, a computational system for processing images is provided. The system includes one or more processors; and one or more non-transitory computer-readable media that collectively store an image processing model. The image processing model is configured to process an input image to generate a processing result. The image processing model includes at least one hybrid operation unit, comprising: an attention module configured to apply an attention mechanism to generate a weighted feature map based on a first feature map, wherein the first feature map is obtained based on the input image; a first convolution module configured to generate a first convolutional feature map based on the weighted feature map; and a feature fusion module configured to generate a second feature map based on the first convolutional feature map and the first feature map, wherein the second feature map is used to generate the processing result.
[0004] In a second aspect of this disclosure, a method for processing an image is provided. The method includes: generating a weighted feature map based on a first feature map and weights shared by query information, key information, and value information, according to an attention subunit in a hybrid operation unit. The first feature map is obtained based on an input image. The attention subunit includes a linear mapping layer shared by query information, key information, and value information. The method further includes: generating a first convolutional feature map based on a first convolutional subunit in a hybrid operation unit by performing a convolution operation on the weighted feature map. The method further includes: generating a processing result for the input image based on the first convolutional feature map and the first feature map.
[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the electronic device to perform the method of the second aspect.
[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the second aspect.
[0007] It should be understood that the content described in this summary section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0008] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0009] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;
[0010] Figure 2 A schematic diagram of an example of a hybrid operation unit according to some embodiments of the present disclosure is shown;
[0011] Figure 3 A schematic diagram of another example of a hybrid operation unit according to some embodiments of the present disclosure is shown;
[0012] Figure 4 A schematic diagram of an image generation model according to some embodiments of the present disclosure is shown;
[0013] Figure 5 A flowchart illustrating a process for generating an image according to some embodiments of the present disclosure is shown;
[0014] Figure 6 A block diagram of an apparatus for processing images according to some embodiments of the present disclosure is shown;
[0015] Figure 7 A block diagram of a computing system for processing images according to some embodiments of the present disclosure is shown; and
[0016] Figure 8 A block diagram of an apparatus capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation
[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0018] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.
[0019] As used in this paper, the term "model" refers to a system that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that uses multiple layers of processing units to process inputs and provide corresponding outputs. In this paper, "model" may also be referred to as a "machine learning model," a "machine learning network," or simply a "network," and these terms are used interchangeably. A model can also include different types of processing units or networks.
[0020] As used herein, a “unit,” “operation unit,” or “subunit” can consist of any suitable machine learning model or network. As used herein, a set of elements, group of elements, or similar expression can include one or more such elements. For example, a “group of convolutional units” can include one or more convolutional units.
[0021] As briefly mentioned earlier, computer vision (CV) technology has been applied to various image processing tasks. With the development of CV technology, there is a wide range of demands for image processing tasks across various fields. For example, users can publish images through content-sharing applications. Such images may be generated by terminal devices (e.g., mobile devices) with content-sharing applications installed. Another example is the identification and classification of objects in images captured by terminal devices. It is generally expected that terminal devices can provide excellent image processing results with high computational efficiency (e.g., providing high-quality generated images) to improve the user experience. However, such image processing tasks typically require a large amount of computation.
[0022] Conventionally, image processing models are based on convolutional neural networks (CNNs). For example, generative adversarial networks (GANs) have been proposed for pixel-level image generation tasks. GANs are basically composed of CNNs, resulting in large networks with many parameters. Considering the limited computing power of terminal devices (e.g., mobile phones), such image processing models struggle to quickly provide users with excellent image processing results. This poses a challenge to the deployment of image processing models on terminals. On the other hand, transformers that apply attention mechanisms have achieved better results than CNNs in image processing tasks such as image classification, segmentation, and detection.
[0023] Embodiments of this disclosure propose a scheme for image processing. According to various embodiments of this disclosure, a hybrid operation unit comprising lightweight attention subunits and convolutional subunits is used to process the feature maps of an image. The hybrid operation unit includes attention subunits arranged in series and at least one convolutional subunit. Within the attention subunits, a shared linear mapping layer is used to compute query information, key information, and value information.
[0024] Attention mechanisms capture global correlations of features, while convolution captures local information. Combining attention with convolution can improve image processing performance. Furthermore, utilizing shared linear mapping layers can reduce the computational cost of attention subunits. Therefore, embodiments of this disclosure can implement a lightweight image processing model. This lightweight image processing model achieves excellent image processing performance while maintaining high image processing efficiency. This lightweight image processing model is particularly suitable for deployment on terminal devices.
[0025] In some embodiments, the convolutional subunit performs depthwise separable convolution on the feature map, i.e., performs depthwise convolution and pointwise convolution. Utilizing depthwise separable convolution can significantly reduce the computational cost of the convolutional subunit, which can further improve image processing efficiency.
[0026] Example Environment
[0027] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. In environment 100, an image processing model 120, also simply referred to as model 120, is deployed in computing device 110. Image processing model 120 is configured to generate a processing result 102 based on an input image 101 in an image processing task. Image processing model 120 has been trained using training data before being applied to the image processing task.
[0028] Image processing tasks can include, but are not limited to, image classification, image segmentation, image generation, and object detection. For example, when image processing model 120 is used for image classification, the processing result 102 can be the classification of objects in the input image 101. Similarly, when image processing model 120 is used for image generation, the processing result 102 can include an output image. Image generation tasks can include various transformation tasks from the input image 101 to the output image.
[0029] After the input image 101 is fed into the image processing model 120, the image processing model 120 generates and processes various feature maps to produce the processing result 102. Figure 1 In the example, image processing model 120 includes a hybrid operation unit 130 that receives a first feature map 131 obtained based on the input image 101 and generates a second feature map 132. The second feature map 132 is further used to generate a processing result 102. The hybrid operation unit 130 includes a series of attention sub-units and at least one convolution sub-unit. The term "sequential arrangement" refers to a series of units arranged sequentially, with the output of the previous unit serving as the input of the next unit.
[0030] In environment 100, computing device 110 can be any type of computing-capable device, including terminal devices or server devices. Terminal devices can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof.
[0031] It should be understood that the structure and function of environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0032] Example of a hybrid operation unit
[0033] The hybrid operation unit 130 includes attention sub-units arranged in series and at least one convolution sub-unit. The hybrid operation unit 130 generates a second feature map 132 based on the first feature map 131. Figure 2 An example of a hybrid operation unit 130 according to some embodiments of the present disclosure is shown. Figure 2 In the example, the hybrid operation unit 130 generally includes attention subunits 230 and a first convolution subunit 210 arranged in series.
[0034] Attention subunit 230 generates a weighted feature map 201 based on the first feature map 131. Attention subunit 230 is configured to apply an attention mechanism to generate the weighted feature map 201 based on weights shared by query information, key information, and value information from the first feature map 131. Specifically, attention subunit 230 can extract the global relevance of features from the first feature map 131 and generate a weighted feature map 201 that utilizes global relevance.
[0035] In various embodiments of this disclosure, the attention subunit 230 includes a linear mapping layer shared by query information, key information, and value information. In other words, the weight matrices used when calculating query information, key information, and value information based on the first feature map 131 are the same. Let X represent the first feature map 131, and let Q, K, and V represent query information, key information, and value information, respectively. The query information Q, key information K, and value information V are calculated using equations (1) to (3), respectively:
[0036] Q = XW q (1)
[0037] K = XW k (2)
[0038] V = XW v (3) Among them W q W k and W v These are the weight matrices used to calculate query information, key information, and value information, respectively. For attention subunit 230, W... q =W k =W v .
[0039] By utilizing a shared linear mapping layer, the computation of Q, K, and V can be simplified, enabling lightweight self-attention. This approach allows for a better balance between performance and computational overhead. Such a shared linear mapping layer facilitates the deployment of the hybrid operation unit 130 in terminal devices.
[0040] The weighted feature map 201 is input to the first convolutional subunit 210. The first convolutional subunit 210 generates a first convolutional feature map 202 based on the weighted feature map 201. In some embodiments, the first convolutional subunit 210 may be configured to perform depthwise separable convolution on the received feature map, i.e., perform depthwise convolution and pointwise convolution sequentially. That is, the weighted feature map 201 is generated by performing depthwise convolution and pointwise convolution on the first feature map 131.
[0041] Depthwise separable convolution is a lightweight convolution operation that includes depthwise convolution and pointwise convolution. Depthwise convolution performs convolution on each channel without changing the depth of the feature map. The output feature map obtained by depthwise convolution has the same number of channels as the input feature map. Pointwise convolution is a 1×1 convolution. Pointwise convolution can be used to change the number of channels in the feature map. As an example, the first convolutional subunit 210 can be implemented as or include MobileNet V1, MobileNet V2, MobileNet V3, etc.
[0042] Having already captured global information using the attention subunit 230, the first convolution subunit 210 can supplement local information. Furthermore, the first convolution subunit 210, which performs depthwise separable convolutions, is a lightweight network. This approach saves computational overhead, improves computational efficiency, and further facilitates the deployment of the hybrid operation unit 130 in terminal devices.
[0043] The hybrid operation unit 130 further generates a second feature map 132 based on the first convolutional feature map 202 and the first feature map 131. The first convolutional feature map 202 has already fused the global and local correlations of the features. Combining the first convolutional feature map 202 and the initially input first feature map 131 to generate the second feature map 202 can enhance information aggregation. This information aggregation is beneficial to improving the performance of the hybrid operation unit.
[0044] In some embodiments, the first convolutional feature map 202 and the first feature map 131 can be added together to form the second feature map 132. By combining the first convolutional feature map 202 and the first feature map 131 through addition, information aggregation can be enhanced without significantly increasing computational overhead.
[0045] In some embodiments, the hybrid operation unit 130 may further include additional sub-units to further process the first convolutional feature map 202 and the first feature map 131. Figure 3 Another example of a hybrid operation unit 130 according to some embodiments of the present disclosure is shown. Figure 3 Zhongyu Figure 2 Elements with the same reference numerals are identical, therefore their description will not be repeated.
[0046] exist Figure 3 In the example, the first convolutional feature map 202 and the first feature map 131 are combined to form a hybrid feature map 303. In some embodiments, the first convolutional feature map 202 and the first feature map 131 can be added together to form the hybrid feature map 303. By adding together, information aggregation can be achieved without significantly reducing computational efficiency.
[0047] The hybrid feature map 303 is input to the second convolutional subunit 320. The second convolutional subunit 320 generates a second convolutional feature map 304 based on the hybrid feature map 303. Similar to the first convolutional subunit 210, in some embodiments, the second convolutional subunit 320 can be configured to perform depthwise separable convolution on the received feature map, i.e., perform depthwise convolution and pointwise convolution. For example, the second convolutional subunit 320 can be implemented as or include MobileNet V1, MobileNet V2, MobileNet V3, etc.
[0048] After obtaining the second convolutional feature map 304, it can be combined with the mixed feature map 303 to form the second feature map 132, which serves as the output of the mixing operation unit 130. The mixed feature map 303 integrates the first convolutional feature map 202 and the initially input first feature map 131. Further fusing the second convolutional feature map 304 with the mixed feature map 303 can further enhance information aggregation, thereby improving the performance of the mixing operation unit.
[0049] In some embodiments, the first convolutional feature map 202 and the first feature map 131 can be added together to form the second feature map 132. By combining the first convolutional feature map 202 and the first feature map 131 through addition, information aggregation can be enhanced without significantly increasing computational overhead.
[0050] exist Figure 2 and Figure 3 In the hybrid operation unit 130 shown, the attention subunit 230 implements lightweight self-attention. Such a hybrid operation unit 130 can be viewed as a lightweight Transformer, where lightweight self-attention replaces conventional self-attention or multi-head attention. Furthermore, in embodiments where the convolution subunit is configured to perform depthwise separable convolution, lightweight depthwise separable convolution replaces the multilayer perceptron (MLP) in a conventional Transformer.
[0051] Besides its high computational efficiency, depthwise separable convolutions have a larger receptive field than MLPs. In conventional Transformer architectures (e.g., Vision Transformer architecture), MLPs focus on extracting single-channel information through linear layers, while spatial dimensional information interaction mainly occurs in the self-attention part. In the hybrid operation unit 130 of this disclosure, the outputs of pointwise convolution and depthwise convolution are concatenated to obtain multi-dimensional information in both spatial and channel dimensions. Therefore, the hybrid operation unit 130 according to the embodiments of this disclosure can balance computational efficiency and model performance.
[0052] In addition to the attention subunit 230, the first convolution subunit 210, and the second convolution subunit 320, the hybrid operation unit 130 may also include other layers or subunits, such as a normalization layer and an activation function layer. The normalization layer and activation function layer may be connected after any or each of the attention subunit 230, the first convolution subunit 210, and the second convolution subunit 320. Alternatively, the normalization layer and activation function layer may be implemented within the attention subunit 230, the first convolution subunit 210, and the second convolution subunit 320.
[0053] In some embodiments, the normalization layer may be a batch normalization (BN) layer. Batch normalization layers are more efficient and better suited for deployment on end devices compared to layer normalization used in conventional Transformers. Alternatively or additionally, in some embodiments, the activation function layer may be a modified linear unit (ReLU) activation function layer. In conventional Transformers, the Gaussian error linear unit (GeLU) activation function is used. Compared to the GeLU activation function, the ReLU activation function is more lightweight, which further contributes to improved computational efficiency.
[0054] Example image processing model
[0055] The image processing model 120, including the hybrid operation unit 130, can be used for various image processing tasks. In image generation tasks, the input image 101 needs to be transformed into an output image. Some pixel-level image generation tasks typically require a large amount of computation. GANs have been proposed for these pixel-level image generation tasks. Conventionally, GANs are basically composed of CNNs, with large networks and many parameters. This is not friendly to model deployment, especially for deployment on terminal devices.
[0056] In some embodiments, the image processing model 120 can be implemented as follows: Figure 4 The image generation model 400 shown here for the image generation task is also simply referred to as model 400. The image generation task can include various transformation tasks from input image 101 to output image 450. Image generation tasks can include, but are not limited to, image inpainting, image style transfer, conditional image generation, virtual clothing transformation, etc. As an example, Figure 4 The image inpainting task is illustrated. Input image 101 contains smudges 405, which makes the plant in input image 101 incomplete. Image generation model 400 inpaints input image 101. The plant is complete and clear in output image 450. It should be understood that... Figure 4 The input image 101 and output image 450 shown are merely exemplary and are not intended to limit the scope of this disclosure. The image generation model 400 can be applied to any type of image processing task.
[0057] Overall, the image generation model 400 includes a first convolutional unit group 410, a second convolutional unit group 420, and a cascaded hybrid operation unit group 430. The first convolutional unit group 410 is located near the input of the image generation model 400 and may include one or more convolutional units. In some embodiments, the first convolutional unit group 410 may include multiple convolutional units arranged in series. For example, in Figure 4 In the example, the first convolutional unit group 410 includes convolutional units 401-1, 401-2, and 401-3 arranged in series, which are also collectively referred to or individually as convolutional unit 401. Convolutional unit 401 (e.g., convolutional block) can have any suitable structure to perform convolution operations.
[0058] The first convolutional unit group 410 receives an input image 101 with an initial resolution and generates a feature map 451 with a first resolution based on the input image 101. The first resolution is smaller than the initial resolution. That is, the first convolutional unit group 410 downsamples the input image 101 to generate the feature map 451.
[0059] exist Figure 4 In the example, the first convolutional unit 401-1 in the first convolutional unit group 410 extracts features from the input image 101 and generates a feature map. The remaining convolutional units after the first convolutional unit 401-1 perform convolution operations on the feature map generated by the previous convolutional unit to generate new feature maps, until the last convolutional unit 401-3 in the first convolutional unit group 410 generates feature map 451.
[0060] Despite Figure 4 The image shows the first convolutional unit 401-1 in the first convolutional unit group 410 downsampling the input image 101, but this is merely exemplary and not intended to limit the scope of this disclosure. In some embodiments, one or more convolutional units in the first convolutional unit group near the input may not downsample the input image. Alternatively, in some embodiments, the model 400 may include an operation unit preceding the first convolutional unit group 410, which receives the feature map output by that operation unit.
[0061] Furthermore, despite Figure 4 The illustration shows a first convolutional unit group 410 comprising three convolutional units arranged in series, but this is merely exemplary. In embodiments of this disclosure, the first convolutional unit group may include any suitable number of convolutional units. Additionally, the first convolutional unit group may include convolutional units that do not alter the feature map resolution.
[0062] The hybrid operation unit group 430 is located in the middle portion of the image generation model 400 and is configured to apply attention mechanisms and convolution operations. The hybrid operation unit group 430 may include one or more hybrid operation units. In some embodiments, the hybrid operation unit group 430 may include multiple hybrid operation units arranged in series. Figure 4 In the example, the hybrid operation unit group 430 includes hybrid operation units 403-1, 403-2, 403-3, 403-4, 403-5, and 403-6 arranged in series, which are also collectively or individually referred to as hybrid operation unit 403. The hybrid operation unit group 430 receives the feature map 451 generated by the first convolutional unit group 410 and generates a feature map 452 with a fourth resolution based on the feature map 451.
[0063] exist Figure 4 In the example, the first hybrid operation unit 403-1 receives the feature map 451 output from the first convolutional unit group 410 and applies attention and convolution operations to the feature map 451 to generate a new feature map. Subsequent hybrid operation units after the first hybrid operation unit 403-1 apply attention and convolution operations to the feature maps generated by the previous hybrid operation unit to generate new feature maps, until the last hybrid operation unit 403-6 generates the feature map 452.
[0064] although Figure 4 The diagram illustrates that each blending operation unit upsamples or downsamples the feature map, but this is merely exemplary. A group of blending operation units can include blending operation units that do not change the feature map resolution. Furthermore, although... Figure 4 The illustration shows a hybrid operation unit group 430 comprising six convolutional units arranged in series, but this is merely exemplary. In embodiments of this disclosure, the hybrid operation unit group may include any suitable number of hybrid operation units.
[0065] Hybrid operation unit 403 includes an attention subunit that implements the attention mechanism and a convolution subunit that performs convolution operations. One or more hybrid units in hybrid operation unit group 430 can be implemented as follows: Figure 2 or Figure 3 The hybrid operation unit 130 shown is illustrated. In some embodiments, each hybrid operation unit 430 is implemented as a hybrid operation unit 130. In other words, each hybrid operation unit 430 includes a lightweight attention subunit and a lightweight convolution subunit.
[0066] As an example, when the hybrid operation unit 403-1 is implemented as the hybrid operation unit 130, feature map 451 is the first feature map 131, and the feature map generated by the hybrid operation unit 403-1 is the second feature map 132. Similarly, when the hybrid operation unit 403-6 is implemented as the hybrid operation unit 130, the feature map generated by the hybrid operation unit 403-5 is the first feature map 131, and feature map 452 is the second feature map 132.
[0067] The second convolutional unit group 420 is located near the output of the image generation model 400 and may include one or more convolutional units. In some embodiments, the second convolutional unit group 420 may include multiple convolutional units arranged in series. Figure 4 In the example, the second convolutional unit group 420 includes convolutional units 402-1, 402-2, and 402-3 arranged in series, which are also collectively referred to or individually as convolutional unit 402. Convolutional unit 402 can have any suitable structure to perform convolution operations.
[0068] The second convolutional unit group 420 receives a feature map 452 with a fourth resolution and generates an output image 450 with a target resolution based on the feature map 452. The target resolution is greater than the fourth resolution. That is, the second convolutional unit group 420 upsamples the feature map to generate an output image 450 with the target resolution. Depending on the specific image generation task, the initial resolution and the target resolution may be the same or different.
[0069] exist Figure 4 In the example, the first convolutional unit 402-1 in the second convolutional unit group 420 performs a convolution operation on the feature map 452 and generates a feature map. The remaining convolutional units after the first convolutional unit 402-1 perform convolution operations on the feature maps generated by the previous convolutional unit to generate new feature maps, until the last convolutional unit 402-3 in the second convolutional unit group 420 generates the output image 450.
[0070] although Figure 4 The illustration shows a second convolutional unit group 420 comprising three convolutional units arranged in series, but this is merely exemplary. In embodiments of this disclosure, the second convolutional unit group may include any suitable number of convolutional units. Additionally, the second convolutional unit group may include convolutional units that do not alter the feature map resolution.
[0071] Overall, image generation model 400 can be as follows: Figure 4The network shown first downsamples and then upsamples. From the input to the output, the resolution of the processed feature maps gradually decreases and then gradually increases. Convolutional units are used in the high-resolution regions near the input and output, while hybrid operation units combining attention and convolution are used in the lower-resolution regions in the middle. Therefore, at the stage level, the image generation model 400 consists of convolutional units and hybrid operation units. Simultaneously, at the block level, the hybrid operation units consist of attention sub-units and convolutional sub-units.
[0072] Attention units, such as those in Transformers, are more powerful than convolutional units, such as those in CNNs, but are less computationally efficient because attention units aim to build global relationships between features, while convolutions only capture local information. Statistical analysis of computational density (computation cost / latency) at different resolutions reveals that attention units are extremely inefficient at high resolutions, but their computational efficiency is comparable to that of convolutional units at low resolutions. Furthermore, the efficiency gap between attention units and convolutional units decreases as resolution decreases.
[0073] Therefore, introducing attention units into layers with low resolution can balance the quality and efficiency of image generation. This maintains high computational efficiency while achieving better image generation results than pure convolutional networks.
[0074] In some embodiments, convolutional units 401 and 402 can be configured to perform depthwise separable convolution on the input feature map, i.e., perform depthwise convolution and pointwise convolution. In this embodiment, convolutional units 401 and 402 are depthwise separable convolutional units. For example, convolutional units 401 and 402 can be implemented as or include MobileNet V1, MobileNet V2, MobileNet V3, etc. Utilizing depthwise separable convolution allows model 400 to be more lightweight and reduces computational overhead. In this way, image generation model 400 can have excellent performance on end devices (e.g., mobile devices) and is deployment-friendly.
[0075] In some embodiments, the first resolution of feature map 451 can be equal to the fourth resolution of feature map 452. In this embodiment, for feature maps smaller than the first resolution, the attention unit can have a computational efficiency that is substantially the same as or higher than that of the convolutional unit. In this way, a balance can be achieved between computational efficiency and generation quality.
[0076] The overall structure of the image generation model 400 has been described above. It should be understood that this overall structure is exemplary. In some embodiments, the first convolutional unit group 410 may be replaced by or included in a first feature generation module. This first feature generation module is configured to downsample the input image 101 to generate a feature map 451. Alternatively or additionally, in some embodiments, the second convolutional unit group 420 may be replaced by or included in a second feature generation module. This second feature generation module is configured to upsample the feature map 452 to generate an output image 450 with a target resolution. The first and second feature generation modules can be implemented using any suitable network. Furthermore, the structure of different parts of the image generation model 400 can be refined in various suitable ways.
[0077] In some embodiments, the first convolutional unit group 410 may have a structure with gradually decreasing resolution. Specifically, the resolution of the feature map output by the convolutional units 401 in the first convolutional unit group 410 may be sequentially reduced to a first resolution according to the order in which the convolutional units are arranged in series. Figure 4 In the above, the resolution of the feature map generated by convolutional unit 401-2 is smaller than the resolution of the feature map generated by convolutional unit 401-1. The resolution of the feature map generated by convolutional unit 402-3 (i.e., the first resolution) is smaller than the resolution of the feature map generated by convolutional unit 401-2.
[0078] Accordingly, the second convolutional unit group 420 can have a structure with gradually increasing resolution. Specifically, the resolution of the feature map output by the convolutional unit 402 in the second convolutional unit group 420 can be increased sequentially to the target resolution according to the order in which the convolutional units are arranged in series. Figure 4 In the above, the resolution of the feature map generated by convolutional unit 402-2 is greater than the resolution of the feature map generated by convolutional unit 402-1. The resolution of the output image (i.e., the target resolution) generated by convolutional unit 402-3 is greater than the resolution of the feature map generated by convolutional unit 402-2.
[0079] In some embodiments, the hybrid operation unit group 430 may have a structure of downsampling followed by upsampling. Specifically, the hybrid operation unit group 430 may include a first hybrid operation unit group for downsampling and a second hybrid operation unit group for upsampling. For example... Figure 4As shown, the first group of hybrid operation units includes hybrid operation units 403-1, 403-2, and 403-3, and can generate a feature map 453 with a third resolution based on feature map 451. The third resolution is smaller than the first and fourth resolutions. The second group of hybrid operation units includes hybrid operation units 403-4, 403-5, and 403-6, and can generate a feature map 452 based on feature map 453.
[0080] In some embodiments, the hybrid operation units 403-1, 403-2, and 403-3 in the first hybrid operation unit group may have a structure with gradually decreasing resolution. Specifically, the resolution of the feature maps output by the hybrid operation units 403-1, 403-2, and 403-3 in the first hybrid operation unit group may be sequentially reduced to a third resolution according to the order in which the hybrid operation units are arranged in series. Figure 4 In this process, the resolution of the feature map generated by the hybrid operation unit 403-2 is lower than the resolution of the feature map generated by the hybrid operation unit 403-1. The resolution of the feature map generated by the hybrid operation unit 403-3 (i.e., the third resolution) is lower than the resolution of the feature map generated by the hybrid operation unit 403-2.
[0081] Accordingly, the hybrid operation units 403-4, 403-5, and 403-6 in the second hybrid operation unit group can have a structure with gradually increasing resolution. Specifically, the resolution of the feature maps output by the hybrid operation units 403-4, 403-5, and 403-6 in the second hybrid operation unit group can be increased sequentially to the fourth resolution according to the order in which the hybrid operation units are arranged in series. Figure 2 In this process, the resolution of the feature map generated by the hybrid operation unit 403-5 is greater than the resolution of the feature map generated by the hybrid operation unit 403-4. The resolution of the feature map 452 generated by the hybrid operation unit 403-6 (i.e., the fourth resolution) is greater than the resolution of the feature map generated by the hybrid operation unit 403-5.
[0082] In some embodiments, the image generation model 400 may be based on a GAN. In this embodiment, the image generation model 400 is a generator within a GAN. During model training, the image generation model 400 is trained using a corresponding discriminative model. The parameters of the image generation model 400 can be optimized using a standard binary cross-entropy function. The specific loss function definition may depend on the image generation task to which the model is applied.
[0083] Example process
[0084] Figure 5A flowchart of an image processing procedure 500 according to some embodiments of the present disclosure is shown. Procedure 500 may be implemented at a computing device 110. For ease of discussion, reference will be made to... Figures 1 to 4 To describe process 500.
[0085] In block 510, computing device 110 generates a weighted feature map 201 based on a first feature map 131 and weights shared by query information, key information, and value information, according to attention subunit 230 in hybrid operation unit 130. The first feature map 131 is obtained based on input image 101. Attention subunit 230 includes a linear mapping layer shared by query information, key information, and value information. That is, in attention subunit 230, W... q =W k =W v .
[0086] In block 520, computing device 110 generates a first convolutional feature map 202 based on a first convolutional subunit 210 in hybrid operation unit 130 by performing a convolution operation on the weighted feature map 201. In some embodiments, the first convolutional subunit 210 may be configured to perform depthwise convolution and pointwise convolution on the feature map input to the first convolutional subunit. That is, the convolution operation performed by the first convolutional subunit 210 is a depthwise separable convolution.
[0087] In box 530, computing device 110 generates a processing result 102 for input image 101 based on first convolutional feature map 202 and first feature map 131. In the image generation task, processing result 102 includes an output image.
[0088] In some embodiments, the computing device 110 may generate a second feature map 132 based on the first convolutional feature map 202 and the first feature map 201, according to the mixing operation unit 130. The computing device 110 may generate a processing result 102 for the input image 101 based on the second feature map 132.
[0089] In some embodiments, the mixing operation unit 130 is included in the generator of the GAN, and the processing result is an output image of the input image.
[0090] In some embodiments, the hybrid operation unit 130 further includes a second convolutional subunit 320. To generate the second feature map 132, the first convolutional feature map 202 and the first feature map 131 can be combined to form a hybrid feature map 303. Based on the hybrid feature map 303, a second convolutional feature map 304 can be generated according to the second convolutional subunit 320 in the hybrid operation unit 130. The second convolutional feature map 304 and the hybrid feature map 303 can be combined to form the second feature map 132.
[0091] In some embodiments, the first convolutional feature map 202 and the first feature map 131 can be added together to obtain a mixed feature map 303. The second convolutional feature map 304 and the mixed feature map 303 can be added together to obtain a second feature map 132.
[0092] In some embodiments, the second convolutional subunit 320 can be configured to perform depthwise convolution and pointwise convolution on the feature map input to the second convolutional subunit. That is, the convolution operation performed by the second convolutional subunit 320 is a depthwise separable convolution.
[0093] In some embodiments, the hybrid operation unit 130 further includes at least one of the following: a batch normalization layer and a ReLU activation function layer.
[0094] In some embodiments, the hybrid operation unit 130 can be implemented in the image generation model. For example, the hybrid operation unit 403-1 can be implemented using the hybrid operation unit 130. In this embodiment, the computing device 110 can generate a first feature map 131 with a first resolution based on the input image 101 with an initial resolution, according to the first convolutional unit group 430. The first resolution is smaller than the initial resolution. That is, the first feature map 131 can be... Figure 4 Feature diagram 451.
[0095] In some embodiments, computing device 110 may generate a fourth feature map with a fourth resolution based on the second feature map 132 and according to a group of hybrid operation units. The group of hybrid operation units includes at least one hybrid operation unit. The group of hybrid operation units may be, for example, a series of hybrid operation units 403-2 to 403-6, and the fourth feature map may be, for example, feature map 452. Computing device 110 may generate an output image 450 with a target resolution based on the fourth feature map and according to a second group of convolutional units 420. The target resolution is greater than the fourth resolution.
[0096] In some embodiments, the convolutional units in the first convolutional unit group 410 and the second convolutional unit group 420 are configured to perform depthwise convolution and pointwise convolution on the feature maps input to the convolutional units. That is, the convolutional operation performed by the first convolutional unit group 410 and the second convolutional unit group 420 can be depthwise separable convolution.
[0097] In some embodiments, the first convolutional unit group 410 includes a plurality of convolutional units arranged in series, and the resolution of the feature map output by each convolutional unit in the first convolutional unit group 410 decreases sequentially to a first resolution according to the cascaded arrangement of the convolutional units. The second convolutional unit group 420 includes a plurality of convolutional units arranged in series, and the resolution of the feature map output by each convolutional unit in the second convolutional unit group 420 increases sequentially to a target resolution according to the cascaded arrangement of the convolutional units.
[0098] Example devices, computing systems, and equipment
[0099] Figure 6 A schematic structural block diagram of an apparatus 600 for processing images according to certain embodiments of the present disclosure is shown. The apparatus 600 may be implemented as or included in a computing device 110. The various modules / components in the apparatus 600 may be implemented by hardware, software, firmware, or any combination thereof.
[0100] As shown in the figure, the device 600 includes an attention module 610 configured to generate a weighted feature map based on a first feature map and weights shared by query information, key information, and value information, according to an attention subunit in a hybrid operation unit. The first feature map is obtained based on the input image. The device 600 also includes a first convolution module 620 configured to generate a first convolutional feature map based on a first convolution subunit in a hybrid operation unit by performing a convolution operation on the weighted feature map. The device 600 further includes a result generation module 630 configured to generate a processing result for the input image based on the first convolutional feature map and the first feature map.
[0101] In some embodiments, the result generation module 630 includes: a feature fusion module configured to generate a second feature map based on a first convolutional feature map and a first feature map, according to a mixing operation unit; and a feature map processing module configured to generate a processing result for the input image based on the second feature map.
[0102] In some embodiments, the feature fusion module includes: a first combining module configured to combine a first convolutional feature map and a first feature map into a hybrid feature map; a second convolution module configured to generate a second convolutional feature map based on the hybrid feature map and according to a second convolutional subunit in a hybrid operation unit; and a second combining module configured to combine the second convolutional feature map and the hybrid feature map into a second feature map.
[0103] In some embodiments, the first combining module is further configured to add the first convolutional feature map and the first feature map to obtain a hybrid feature map, and wherein the second combining module is further configured to add the second convolutional feature map and the hybrid feature map to obtain a second feature map.
[0104] In some embodiments, the apparatus 600 further includes an input module configured to generate a first feature map with a first resolution less than the initial resolution based on an input image having an initial resolution and according to a first group of convolutional units.
[0105] In some embodiments, the result generation module 640 includes: an intermediate module configured to generate a fourth feature map with a fourth resolution based on the second feature map and according to a group of hybrid operation units, the group of hybrid operation units including at least one hybrid operation unit; and an output module configured to generate an output image with a target resolution as a processing result based on the fourth feature map and according to a group of second convolutional units, the target resolution being greater than the fourth resolution.
[0106] In some embodiments, the convolutional units in the first and second convolutional unit groups are configured to perform depthwise convolution and pointwise convolution on the feature maps input to the convolutional units.
[0107] In some embodiments, the first convolutional unit group includes a plurality of convolutional units arranged in series, and the resolution of the feature map output by each convolutional unit in the first convolutional unit group is sequentially reduced to a first resolution according to the cascaded arrangement of the convolutional units. The second convolutional unit group includes a plurality of convolutional units arranged in series, and the resolution of the feature map output by each convolutional unit in the second convolutional unit group is sequentially increased to a target resolution according to the cascaded arrangement of the convolutional units.
[0108] In some embodiments, the hybrid operation unit further includes at least one of the following: a batch normalization layer and a modified linear unit activation function layer.
[0109] In some embodiments, the hybrid operation unit is included in the generator of the generative adversarial network, and the processing result includes an output image for the input image.
[0110] In some embodiments, at least one of the first and second convolutional subunits is configured to perform depthwise convolution and pointwise convolution on the feature map input to the convolutional subunit.
[0111] Figure 7 A schematic structural block diagram of a computing system 700 for processing images according to some embodiments of the present disclosure is shown. The computing system 700 may be implemented as or included in a computing device 110.
[0112] like Figure 7 As shown, the computing system 700 includes one or more processors 710 and one or more non-transitory computer-readable media 720. The one or more non-transitory computer-readable media 720 jointly store an image processing model 730. The image processing model 730 is configured to process an input image to generate a processing result. The image processing model 730 includes at least one hybrid operation unit 750. The image processing model 730 is, for example, an image processing model 120.
[0113] The hybrid operation unit 750 includes an attention module 751 configured to apply an attention mechanism to generate a weighted feature map based on a first feature map, obtained from the input image, using weights shared by query information, key information, and value information. The hybrid operation unit 750 also includes a first convolution module 752 configured to generate a first convolutional feature map based on the weighted feature map. The hybrid operation unit 750 further includes a feature fusion module 753 configured to generate a second feature map based on the first convolutional feature map and the first feature map, wherein the second feature map is used to generate the processing result. The hybrid operation unit 750 is, for example, a hybrid operation unit 130.
[0114] In some embodiments, the feature fusion module includes: a first combination module configured to combine a first convolutional feature map and a first feature map into a hybrid feature map; a second convolution module configured to generate a second convolutional feature map based on the hybrid feature map; and a second combination module configured to combine the second convolutional feature map and the hybrid feature map into a second feature map.
[0115] In some embodiments, the first combining module is further configured to add the first convolutional feature map and the first feature map to obtain a hybrid feature map, and wherein the second combining module is further configured to add the second convolutional feature map and the hybrid feature map to obtain a second feature map.
[0116] In some embodiments, the computing system 700 further includes an input module comprising a first convolutional unit group (e.g., a first convolutional unit group 410), the first convolutional unit group including at least one convolutional unit. The input module is configured to generate a first feature map having a first resolution smaller than the initial resolution based on an input image having an initial resolution.
[0117] In some embodiments, the hybrid operation unit 750 is included in the hybrid operation module. The hybrid operation module is further configured to generate a third feature map with a third resolution based on the first feature map, and to generate a fourth feature map with a fourth resolution based on the third feature map, wherein the third resolution is smaller than the fourth resolution. In this embodiment, the hybrid operation unit 750 is, for example, hybrid operation unit 430-1, and the first feature map is, for example, feature map 451, the third feature map is, for example, feature map 453, and the fourth feature map is, for example, feature map 452.
[0118] In some embodiments, the computing system 700 further includes an output module comprising a second convolutional unit group (e.g., second convolutional unit group 420), the second convolutional unit group including at least one convolutional unit. The output module is configured to generate an output image with a target resolution as a processing result based on a second feature map having a second resolution, wherein the target resolution is greater than the second resolution. In this embodiment, the mixing operation unit 750 is, for example, mixing operation units 430-6, and the second feature map is, for example, feature map 452, and the second resolution is equal to the fourth resolution described above.
[0119] In some embodiments, the convolutional units in the first and second convolutional unit groups are configured to perform depthwise convolution and pointwise convolution on the feature maps input to the convolutional units.
[0120] In some embodiments, the first convolutional unit group includes a plurality of convolutional units arranged in series, and the resolution of the feature map output by each convolutional unit of the first convolutional unit group decreases sequentially to a first resolution according to the cascaded order of the convolutional units. In some embodiments, the second convolutional unit group includes a plurality of convolutional units arranged in series, and the resolution of the feature map output by each convolutional unit of the second convolutional unit group increases sequentially to a target resolution according to the cascaded order of the convolutional units.
[0121] In some embodiments, at least one of the first convolution module 752 and the second convolution module is configured to perform depthwise convolution and pointwise convolution on the input feature map.
[0122] Figure 8 A block diagram is shown illustrating a computing device 800 in which one or more embodiments of the present disclosure may be implemented. It should be understood that... Figure 8 The computing device 800 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 8 The computing device 800 shown can be used to implement Figure 1 The computing device 110.
[0123] like Figure 8 As shown, computing device 800 is in the form of a general-purpose computing device. Components of computing device 800 may include, but are not limited to, one or more processors or processing units 810, memory 820, storage devices 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. Processing unit 810 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 820. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 800.
[0124] Computing device 800 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to computing device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 820 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 830 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within computing device 800.
[0125] The computing device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 8 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 820 may include computer program product 825 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.
[0126] The communication unit 840 enables communication with other computing devices via a communication medium. Additionally, the components of the computing device 800 can function as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 800 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or another network node.
[0127] Input device 850 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 860 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 800 can also communicate as needed with one or more external devices (not shown) via communication unit 840. These external devices, such as storage devices, display devices, etc., can communicate with one or more devices that enable user interaction with computing device 800, or with any device (e.g., network card, modem, etc.) that enables computing device 800 to communicate with one or more other computing devices. Such communication can be performed via input / output (I / O) interfaces (not shown).
[0128] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0129] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0130] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0131] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0132] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0133] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A computing system for processing images, comprising: One or more processors; as well as One or more non-transitory computer-readable media, which share the following storage: An image processing model configured to process an input image to generate a processing result, the image processing model including at least one hybrid operation unit, the hybrid operation unit comprising: An attention module is configured to apply an attention mechanism to generate a weighted feature map based on a first feature map, which is obtained based on the input image, and the attention module includes a linear mapping layer shared by the query information, key information, and value information, and the attention module is configured to apply the shared weights through the linear mapping layer. The first convolutional module is configured to generate a first convolutional feature map based on the weighted feature map; and The feature fusion module is configured to generate a second feature map based on the first convolutional feature map and the first feature map, wherein the second feature map is used to generate the processing result.
2. The computing system according to claim 1, wherein the feature fusion module comprises: The first combination module is configured to combine the first convolutional feature map and the first feature map into a hybrid feature map; The second convolutional module is configured to generate a second convolutional feature map based on the hybrid feature map; as well as The second combination module is configured to combine the second convolutional feature map and the hybrid feature map into the second feature map.
3. The computing system according to claim 2, wherein the first combining module is further configured to: add the first convolutional feature map and the first feature map to obtain the mixed feature map, and The second combination module is further configured to add the second convolutional feature map and the mixed feature map to obtain the second feature map.
4. The computing system according to claim 1, further comprising: The input module includes a first convolutional unit group, the first convolutional unit group including at least one convolutional unit, and the input module is configured to generate a first feature map with a first resolution based on the input image having an initial resolution, the first resolution being smaller than the initial resolution.
5. The computing system of claim 4, wherein the hybrid operation unit is included in a hybrid operation module, and the hybrid operation module is further configured to generate a third feature map with a third resolution based on the first feature map, and to generate a fourth feature map with a fourth resolution based on the third feature map, wherein the third resolution is smaller than the fourth resolution.
6. The computing system according to claim 1, further comprising: The output module includes a second convolutional unit group, the second convolutional unit group including at least one convolutional unit, and the output module is configured to generate an output image with a target resolution as the processing result based on the second feature map having a second resolution, wherein the target resolution is greater than the second resolution.
7. The computing system of claim 6, wherein the convolutional units in the first and second convolutional unit groups are configured to perform depthwise convolution and pointwise convolution on feature maps input to the convolutional units.
8. The computing system according to claim 4, wherein the first convolutional unit group comprises a plurality of convolutional units arranged in series, and the resolution of the feature map output by each convolutional unit of the first convolutional unit group is sequentially reduced to the first resolution according to the cascaded arrangement of the convolutional units.
9. The computing system according to claim 6, wherein the second convolutional unit group comprises a plurality of convolutional units arranged in series, and the resolution of the feature map output by each convolutional unit in the second convolutional unit group is sequentially increased to the target resolution according to the cascaded arrangement of the convolutional units.
10. The computing system of claim 2, wherein at least one of the first convolution module and the second convolution module is configured to perform depthwise convolution and pointwise convolution on the input feature map.
11. A method for processing an image, comprising: Based on a first feature map and weights shared by query information, key information, and value information, a weighted feature map is generated according to an attention subunit in a hybrid operation unit. The first feature map is obtained based on an input image. The attention subunit includes a linear mapping layer shared by query information, key information, and value information, and the attention subunit is configured to apply the shared weights through the linear mapping layer. By performing a convolution operation on the weighted feature map, a first convolutional feature map is generated based on the first convolutional subunit in the hybrid operation unit; Based on the first convolutional feature map and the first feature map, a second feature map is generated according to the hybrid operation unit; as well as The processing result for the input image is generated based on the second feature map.
12. The method of claim 11, wherein generating the second feature map comprises: The first convolutional feature map and the first feature map are combined into a hybrid feature map; Based on the hybrid feature map, a second convolutional feature map is generated according to the second convolutional sub-unit in the hybrid operation unit; as well as The second convolutional feature map and the mixed feature map are combined to form the second feature map.
13. The method of claim 12, wherein combining the first convolutional feature map and the first feature map into the hybrid feature map comprises: The first convolutional feature map and the first feature map are added together to form the mixed feature map, and Combining the second convolutional feature map and the mixed feature map into the second feature map includes: The second convolutional feature map and the mixed feature map are added together to form the second feature map.
14. The method of claim 11, further comprising: Based on the input image having an initial resolution, a first feature map having a first resolution is generated according to a first group of convolutional units, wherein the first resolution is smaller than the initial resolution.
15. The method of claim 14, wherein generating the processing result comprises: Based on the second feature map, a fourth feature map with a fourth resolution is generated according to the hybrid operation unit group, wherein the hybrid operation unit group includes at least one of the hybrid operation units; as well as Based on the fourth feature map, and according to the second convolutional unit group, an output image with a target resolution is generated as the processing result, wherein the target resolution is greater than the fourth resolution.
16. The method of claim 15, wherein the first convolutional unit group comprises a plurality of said convolutional units arranged in series, and the resolution of the feature maps output by each convolutional unit in the first convolutional unit group decreases sequentially to the first resolution according to the cascaded arrangement of the convolutional units. The second convolutional unit group includes a plurality of convolutional units arranged in series, and the resolution of the feature map output by each convolutional unit in the second convolutional unit group is sequentially increased to the target resolution according to the cascaded arrangement of the convolutional units.
17. The method of claim 12, wherein at least one of the first convolutional subunit and the second convolutional subunit is configured to perform depthwise convolution and pointwise convolution on the feature map input to the convolutional subunit.
18. An electronic device comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 11 to 17.
19. A computer-readable storage medium having a computer program stored thereon, the computer program being executable by a processor to implement the method according to any one of claims 11 to 17.
Citation Information
Patent Citations
Multi-scale self-attention target detection method based on weight sharing
CN111967480A